How Orcabase works
The architecture of Orcabase on one page: where your data lives, how it flows from your databases to answers, and what each layer does.
Orcabase has four layers. Your data comes in at the top, gets a clear meaning in the middle, and comes out as answers wherever you work. Answers only ever read your data, and with a read-only login nothing in Orcabase can change it.
01 Your data
Wherever it already lives
Your app’s database
PostgreSQL
A warehouse you have
Google BigQuery
Spreadsheets and files
Google Sheets, CSV, Parquet, JSON
Other databases and apps
MySQL, and hundreds of apps via Fivetran
02 Data Warehouse
Data sources in your organization
Connected databases
Queried live, where they are. Tables aren’t copied.
Hosted by Orcabase
Postgres servers and DuckDB warehouses that hold synced and uploaded data.
03 Semantic layer
What your numbers mean
Data models
Each table’s dimensions, measures and joins.
Metrics
Revenue, churn and the rest: defined once, certified by a person.
04 Where you use it
Same data, same definitions, everywhere
Data Agent
Questions in plain words
SQL and notebooks
For people who write SQL
Dashboards
And public share links
Claude and ChatGPT
Through MCP
1. Your data
Data reaches Orcabase in one of three ways, and you can mix them:
- Connected live. Your app’s Postgres database or a BigQuery project stays where it is. Orcabase queries it when someone asks a question, so answers always reflect the data as it is now.
- Synced on a schedule. Google Sheets and MySQL are copied into a warehouse Orcabase hosts, every 15 minutes to once a day. Hundreds of other apps sync the same way through Fivetran. See Sync data.
- Uploaded. A CSV, Parquet or JSON file becomes a table in a few clicks. See Upload files.
2. The Data Warehouse
Everything you connect becomes a data source in your organization. A data source is one database: either one you already run, or one Orcabase hosts for you (a Postgres server, or a DuckDB warehouse built for analysis). Synced and uploaded data always lands in a hosted one.
Every other part of Orcabase works through data sources: Explore browses them, the SQL Editor queries them, and the agent reads them. A DuckDB warehouse can also attach a Postgres data source, so one query can join synced spreadsheets with your app’s live data.
3. The semantic layer
Raw tables don’t say what your numbers mean. Is revenue every order, or only paid ones, minus refunds? The semantic layer is where you write that down, once:
- A data model describes a table: what you can group by (dimensions, like order date or plan), what you can add up (measures, like order amount) and how it joins to other tables.
- A metric is a named business number built on a model, like Revenue or Conversion rate. When a person checks it and clicks Certify, it becomes the definition everyone uses.
When anyone asks for a metric, Orcabase writes the SQL for that data source’s database from the current definition. Fix a definition, and every chart that uses it is fixed on its next run. See How metrics work.
4. Where you use it
| Surface | Who it’s for | What it does |
|---|---|---|
| Data Agent | Everyone | Answers questions in plain words, with a chart and its steps. Builds charts, dashboards and draft metrics when asked. |
| SQL Editor, queries, notebooks | People who write SQL | Run SQL against any data source, save it, chart it. |
| Dashboards | The whole team, and people outside it | The numbers that matter on one page. Shareable with a public link. |
| Claude and ChatGPT | People who already work in an AI assistant | The same tools as the agent, through MCP, from the assistant you already use. |
How a question becomes an answer
Here’s what happens when you ask the agent “How much revenue did we make last month?”
- The agent checks your metrics. There’s a certified Revenue metric, so it uses that.
- Orcabase turns “Revenue, last month” into SQL for the database behind the metric, from the current definition.
- The query runs against your data source, capped at 5,000 rows and 30 seconds.
- The AI model sees a short summary of the result (at most 30 rows). You see the full result, as a chart.
- The answer leads with the number, and “How I got this” lists each step, including the SQL that ran.
With no metric that fits, the agent reads the table’s columns and a few sample rows, then writes the SQL itself, and tells you so. Certified metrics are what make answers repeatable.
What Orcabase stores, and where
| What | Where it’s kept |
|---|---|
| Tables in a database you connect | In your database. Orcabase reads them when a query runs and doesn’t copy them. |
| Synced and uploaded data | In a Postgres server or DuckDB warehouse that Orcabase hosts for your organization. |
| The latest result of each saved query | In Orcabase, so dashboards open instantly. Only the latest run is kept, up to 5,000 rows. |
| Queries, dashboards, notebooks, data models and metrics | In Orcabase, in your organization. |
| Agent chats | In Orcabase, visible only to the person who started them. |
| Database passwords, service account keys and AI keys | Encrypted in Orcabase. After you save one, it’s never shown again. |
Note
Everything belongs to exactly one organization, and members only ever see their own organization’s data. See Security for the details.
Under the hood
For the technically curious:
- The app at
app.orcabase.cotalks to the Orcabase API atapi.tini.so. So do AI assistants, through the MCP endpoint atapi.tini.so/api/mcp. - Three database engines are supported: PostgreSQL, Google BigQuery and DuckDB. Metric queries compile to each one’s own SQL.
- Hosted Postgres servers sit behind a gateway with a connection pooler, and every connection uses TLS. Orcabase reads them as a dedicated read-only user.
- The Data Agent calls AI models through OpenRouter or Vercel AI Gateway, using the same tools that AI assistants get over MCP.
Something unclear or missing? Troubleshooting covers the common errors, or message us.