Skip to Content
SourcesDremio

Dremio

Dremio

Scan Dremio tables and views with lineage across spaces and sources.

Category
Warehouse & Lakehouse
Source type
DREMIO
Produces
table

Dremio sits in front of other systems. Its sources are other people’s databases and object stores, read in place; its spaces are the layer of views an organisation builds on top of them. So a Dremio scan finds sensitive data twice over: in the raw tables it can reach, and in every view that quietly carries a customer email three joins away from where it started.

What you need to connect

First, where Dremio runs:

DeploymentWhat you supply
Self-hosted DremioThe URL you open Dremio at
Dremio CloudThe project ID and the region the project lives in

The URL is the one in your browser’s address bar, port included — usually :9047. Classifyre talks to Dremio only through that address; it needs no JDBC driver and no additional open port.

Then, how to sign in:

Sign-inWorks with
Personal access tokenDremio Cloud, and self-hosted Enterprise Edition — the better choice wherever it exists
Username and passwordSelf-hosted Dremio only. Community Edition has no access tokens, so this is its only option

Give the user SELECT on the spaces and sources you want scanned, and nothing more. A dedicated service user is the right identity: what it cannot see is not scanned.

What Classifyre reads

Tables (Dremio’s physical datasets — a database table, an Iceberg table, a promoted file or folder) and views, across every source and space the user can see. Narrow that with a list of paths to include or exclude, written the way Dremio writes them:

  • Finance — a whole space or source
  • Marketing.Campaigns — one folder and everything beneath it
  • Samples."samples.dremio.com" — quote a name that itself contains a dot

Personal home spaces (@username) are skipped unless you ask for them. Files that nobody has promoted to a table are skipped too: Dremio cannot query them, so there are no rows to sample.

Shared behaviour · SQL databases

One asset per table or view, never one per row. The asset carries the table's structure — database, schema, table name, object type, its columns and their types, and a row-count estimate — and its content is a sample of real rows, formatted so a detector reads actual values rather than a schema dump.

How many rows, and which ones, is entirely up to the sampling strategy. Large tables are paged through by key rather than by OFFSET, so a scan that stops halfway can resume from where it left off instead of re-reading from the top.

Read-only throughout. The connector issues catalog queries and bounded SELECTs. Nothing is written back, and a read-only account is the right account to give it.

Relationships come out of the engine's own catalog: a view and the tables it reads from are recorded as FLOW — real lineage, with column-level detail parsed out of the view's SQL where the SQL makes that possible. See Lineage.

Metadata on every asset

Asset kind · table

FieldTypeAlways presentWhat it is
databasestringYesDatabase or catalog name
table_namestringYesTable name
table_typestringYesObject type (TABLE/VIEW)
schemastringNoSchema name
columnsobject[]NoColumns as {name, type} objects
row_countintegerNoEstimated number of rows
object_typestringNoSource object type
container_typestringNoKind of Dremio container the dataset lives in (SOURCE/SPACE/HOME)
source_systemstringNoType of the external system behind the Dremio source (for example POSTGRES, S3)
formatstringNoStorage format of the underlying data (for example Parquet, Iceberg)

Lineage

Lineage

Two kinds, and together they are the reason to scan Dremio at all.

  • View lineage — each view and the tables and views it reads from. On Dremio Cloud and Enterprise Edition this comes from Dremio’s own lineage graph. On Community Edition, which has none, it is read from the view’s SQL instead and marked as SQL parsed, so you can tell the two apart. Either way the SQL is kept with the edge, and column-level detail is recovered from it where the SQL names its columns.
  • Source lineage — a table in a Dremio source is really a table somewhere else. For PostgreSQL, MySQL, SQL Server, Oracle and Snowflake sources, Classifyre links the Dremio table to the underlying table by name. If that database is also connected as its own source — now or later — the two meet, and lineage runs from the original table, through Dremio, to the last view.

Tables and views outside your scan scope are not dropped: they are recorded by name and the edge completes itself if they are scanned later.

See Lineage & Relationships.

Worth knowing

  • Every sampled table is a Dremio query, so scan cost is engine cost. On Dremio Cloud that means the engine has to be running; the sampling strategy is what keeps the bill small.
  • Source lineage needs to see the source’s settings. The host and database behind a source are only visible to a user who may view that source’s configuration. Without that, assets and view lineage are unaffected — only the link through to the underlying database is missing.
  • Object-storage sources can be deep. Discovery walks folders, and a data lake has a lot of them. Scope those sources with include paths.
  • Dremio tables have no keys, so incremental scans page by row position rather than by key. On a table that is being rewritten between runs, a row can be seen twice or missed until the next full pass.
  • Self-signed certificates on a self-hosted Dremio can be accepted with a connection setting. Leave verification on everywhere else.

Configuration

Beyond the fields below, every source also has the settings shared by all of them: the sampling strategy, the augmentation notebook that enriches assets after extraction, the detectors to run, the scan schedule, and the compute limits for its scan jobs.

Required

Without these, the source will not save.

This section depends on which authentication method you pick — one of the following applies.

Self-hosted Dremio

FieldTypeRequiredWhat it doesDefault
urlstringYesAddress you open Dremio at, including the port (for example, https://dremio.example.com:9047)—

Dremio Cloud

FieldTypeRequiredWhat it doesDefault
project_idstringYesDremio Cloud project ID (Project Settings > General Information)—
regionenumYesDremio Cloud control plane the project lives in Allowed: US, EU"US"

Secrets

Stored encrypted and never shown again after you save them. See Configuration & Fields.

This section depends on which authentication method you pick — one of the following applies.

Personal Access Token

FieldTypeRequiredWhat it doesDefault
tokenstringYesDremio personal access token (Account Settings > Personal Access Tokens)—

Username & Password

FieldTypeRequiredWhat it doesDefault
usernamestringYesDremio login username—
passwordstringYesDremio login password—

Optional

Everything you can tune. Sensible defaults apply when you leave them alone.

FieldTypeRequiredWhat it doesDefault
optionalobjectNo—no extra properties—
connectionobjectNoDremio API and query tuning options.no extra properties—
connection.query_timeout_secondsintegerNoMaximum time to wait for one sampling query to finish before it is cancelledmin 10, max 3600300
connection.timeout_secondsintegerNoHTTP timeout for a single Dremio API callmin 5, max 30030
connection.verify_sslbooleanNoVerify the Dremio server's TLS certificate. Turn off only for a self-hosted Dremio with a self-signed certificate.true
extractionobjectNoLineage extraction controls for Dremio.no extra properties—
extraction.include_source_lineagebooleanNoLink tables that Dremio reads from an external database (PostgreSQL, MySQL, SQL Server, Oracle, Snowflake) to that database's own table, so lineage continues across systems. Needs permission to view the Dremio source's settings.true
extraction.include_view_lineagebooleanNoLink each view to the tables and views it reads from. Uses Dremio's own lineage graph where the edition provides it, and the view's SQL otherwise.true
scopeobjectNoWhich Dremio sources, spaces, folders, and datasets to scan.no extra properties—
scope.exclude_pathsarrayNoSources, spaces, folders, or datasets to skip, written the same way as include_paths. Exclusions win over inclusions.[]
scope.exclude_paths[]stringNo——
scope.include_home_spacesbooleanNoInclude personal home spaces (@username) in extractionfalse
scope.include_pathsarrayNoOptional allowlist of sources, spaces, folders, or datasets, written as dot-separated paths (for example, Marketing or Samples."samples.dremio.com"). Everything beneath a listed path is scanned. Wrap a name that itself contains a dot in double quotes.—
scope.include_paths[]stringNo——
scope.include_tablesbooleanNoInclude tables (physical datasets) in extractiontrue
scope.include_viewsbooleanNoInclude views (virtual datasets) in extractiontrue
scope.table_limitintegerNoOptional cap on number of table/view assets extractedmin 1—
Last updated on