Dremio
Dremio
Scan Dremio tables and views with lineage across spaces and sources.
- Category
- Warehouse & Lakehouse
- Source type
- DREMIO
- Produces
- table
Dremio sits in front of other systems. Its sources are other people’s databases and object stores, read in place; its spaces are the layer of views an organisation builds on top of them. So a Dremio scan finds sensitive data twice over: in the raw tables it can reach, and in every view that quietly carries a customer email three joins away from where it started.
What you need to connect
First, where Dremio runs:
| Deployment | What you supply |
|---|---|
| Self-hosted Dremio | The URL you open Dremio at |
| Dremio Cloud | The project ID and the region the project lives in |
The URL is the one in your browser’s address bar, port included — usually
:9047. Classifyre talks to Dremio only through that address; it needs no
JDBC driver and no additional open port.
Then, how to sign in:
| Sign-in | Works with |
|---|---|
| Personal access token | Dremio Cloud, and self-hosted Enterprise Edition — the better choice wherever it exists |
| Username and password | Self-hosted Dremio only. Community Edition has no access tokens, so this is its only option |
Give the user SELECT on the spaces and sources you want scanned, and nothing
more. A dedicated service user is the right identity: what it cannot see is not
scanned.
What Classifyre reads
Tables (Dremio’s physical datasets — a database table, an Iceberg table, a promoted file or folder) and views, across every source and space the user can see. Narrow that with a list of paths to include or exclude, written the way Dremio writes them:
Finance— a whole space or sourceMarketing.Campaigns— one folder and everything beneath itSamples."samples.dremio.com"— quote a name that itself contains a dot
Personal home spaces (@username) are skipped unless you ask for them. Files
that nobody has promoted to a table are skipped too: Dremio cannot query them,
so there are no rows to sample.
Shared behaviour · SQL databases
One asset per table or view, never one per row. The asset carries the table's structure — database, schema, table name, object type, its columns and their types, and a row-count estimate — and its content is a sample of real rows, formatted so a detector reads actual values rather than a schema dump.
How many rows, and which ones, is entirely up to the sampling strategy. Large tables are paged through by key rather than by OFFSET, so a scan that stops halfway can resume from where it left off instead of re-reading from the top.
Read-only throughout. The connector issues catalog queries and bounded SELECTs. Nothing is written back, and a read-only account is the right account to give it.
Relationships come out of the engine's own catalog: a view and the tables it reads from are recorded as FLOW — real lineage, with column-level detail parsed out of the view's SQL where the SQL makes that possible. See Lineage.
Metadata on every asset
Asset kind · table
| Field | Type | Always present | What it is |
|---|---|---|---|
| database | string | Yes | Database or catalog name |
| table_name | string | Yes | Table name |
| table_type | string | Yes | Object type (TABLE/VIEW) |
| schema | string | No | Schema name |
| columns | object[] | No | Columns as {name, type} objects |
| row_count | integer | No | Estimated number of rows |
| object_type | string | No | Source object type |
| container_type | string | No | Kind of Dremio container the dataset lives in (SOURCE/SPACE/HOME) |
| source_system | string | No | Type of the external system behind the Dremio source (for example POSTGRES, S3) |
| format | string | No | Storage format of the underlying data (for example Parquet, Iceberg) |
Lineage
Lineage
Two kinds, and together they are the reason to scan Dremio at all.
- View lineage — each view and the tables and views it reads from. On Dremio Cloud and Enterprise Edition this comes from Dremio’s own lineage graph. On Community Edition, which has none, it is read from the view’s SQL instead and marked as SQL parsed, so you can tell the two apart. Either way the SQL is kept with the edge, and column-level detail is recovered from it where the SQL names its columns.
- Source lineage — a table in a Dremio source is really a table somewhere else. For PostgreSQL, MySQL, SQL Server, Oracle and Snowflake sources, Classifyre links the Dremio table to the underlying table by name. If that database is also connected as its own source — now or later — the two meet, and lineage runs from the original table, through Dremio, to the last view.
Tables and views outside your scan scope are not dropped: they are recorded by name and the edge completes itself if they are scanned later.
Worth knowing
- Every sampled table is a Dremio query, so scan cost is engine cost. On Dremio Cloud that means the engine has to be running; the sampling strategy is what keeps the bill small.
- Source lineage needs to see the source’s settings. The host and database behind a source are only visible to a user who may view that source’s configuration. Without that, assets and view lineage are unaffected — only the link through to the underlying database is missing.
- Object-storage sources can be deep. Discovery walks folders, and a data lake has a lot of them. Scope those sources with include paths.
- Dremio tables have no keys, so incremental scans page by row position rather than by key. On a table that is being rewritten between runs, a row can be seen twice or missed until the next full pass.
- Self-signed certificates on a self-hosted Dremio can be accepted with a connection setting. Leave verification on everywhere else.
Configuration
Beyond the fields below, every source also has the settings shared by all of them: the sampling strategy, the augmentation notebook that enriches assets after extraction, the detectors to run, the scan schedule, and the compute limits for its scan jobs.
Required
Without these, the source will not save.
This section depends on which authentication method you pick — one of the following applies.
Self-hosted Dremio
| Field | Type | Required | What it does | Default |
|---|---|---|---|---|
| url | string | Yes | Address you open Dremio at, including the port (for example, https://dremio.example.com:9047) | — |
Dremio Cloud
| Field | Type | Required | What it does | Default |
|---|---|---|---|---|
| project_id | string | Yes | Dremio Cloud project ID (Project Settings > General Information) | — |
| region | enum | Yes | Dremio Cloud control plane the project lives in Allowed: US, EU | "US" |
Secrets
Stored encrypted and never shown again after you save them. See Configuration & Fields.
This section depends on which authentication method you pick — one of the following applies.
Personal Access Token
| Field | Type | Required | What it does | Default |
|---|---|---|---|---|
| token | string | Yes | Dremio personal access token (Account Settings > Personal Access Tokens) | — |
Username & Password
| Field | Type | Required | What it does | Default |
|---|---|---|---|---|
| username | string | Yes | Dremio login username | — |
| password | string | Yes | Dremio login password | — |
Optional
Everything you can tune. Sensible defaults apply when you leave them alone.
| Field | Type | Required | What it does | Default |
|---|---|---|---|---|
| optional | object | No | —no extra properties | — |
| connection | object | No | Dremio API and query tuning options.no extra properties | — |
| connection.query_timeout_seconds | integer | No | Maximum time to wait for one sampling query to finish before it is cancelledmin 10, max 3600 | 300 |
| connection.timeout_seconds | integer | No | HTTP timeout for a single Dremio API callmin 5, max 300 | 30 |
| connection.verify_ssl | boolean | No | Verify the Dremio server's TLS certificate. Turn off only for a self-hosted Dremio with a self-signed certificate. | true |
| extraction | object | No | Lineage extraction controls for Dremio.no extra properties | — |
| extraction.include_source_lineage | boolean | No | Link tables that Dremio reads from an external database (PostgreSQL, MySQL, SQL Server, Oracle, Snowflake) to that database's own table, so lineage continues across systems. Needs permission to view the Dremio source's settings. | true |
| extraction.include_view_lineage | boolean | No | Link each view to the tables and views it reads from. Uses Dremio's own lineage graph where the edition provides it, and the view's SQL otherwise. | true |
| scope | object | No | Which Dremio sources, spaces, folders, and datasets to scan.no extra properties | — |
| scope.exclude_paths | array | No | Sources, spaces, folders, or datasets to skip, written the same way as include_paths. Exclusions win over inclusions. | [] |
| scope.exclude_paths[] | string | No | — | — |
| scope.include_home_spaces | boolean | No | Include personal home spaces (@username) in extraction | false |
| scope.include_paths | array | No | Optional allowlist of sources, spaces, folders, or datasets, written as dot-separated paths (for example, Marketing or Samples."samples.dremio.com"). Everything beneath a listed path is scanned. Wrap a name that itself contains a dot in double quotes. | — |
| scope.include_paths[] | string | No | — | — |
| scope.include_tables | boolean | No | Include tables (physical datasets) in extraction | true |
| scope.include_views | boolean | No | Include views (virtual datasets) in extraction | true |
| scope.table_limit | integer | No | Optional cap on number of table/view assets extractedmin 1 | — |