Azure Blob Storage
Azure Blob Storage
Scan blobs from Azure Storage containers with key, SAS, or managed identity auth.
- Category
- Warehouse & Lakehouse
- Source type
- AZURE_BLOB_STORAGE
- Produces
- fileimageaudiovideoarchive
Azure Blob Storage holds the flat files of the Microsoft-estate data platform: landing zones, exports, archives, and whatever a pipeline dropped along the way.
What you need to connect
The account URL (https://<account>.blob.core.windows.net) and a
container. Then pick whichever credential you already have — the source
accepts all the usual Azure options:
| Credential | Notes |
|---|---|
| Connection string | Takes precedence over everything else |
| Account key | Full access to the account — prefer something narrower |
| SAS token | Scope it to read + list on one container |
| Service principal | Entra client ID, secret and tenant ID |
| Managed identity | Leave all credentials empty — the default Azure credential chain is used |
Storage Blob Data Reader is the right role for the Entra-based options.
What Classifyre reads
Every blob in the container, narrowed by prefix and by extension allow- and denylists.
Shared behaviour · File and object sources
Every object in the container becomes one asset. What it is — a PDF, a spreadsheet, a screenshot, a video — is worked out from the bytes themselves, not from the file name, because storage systems routinely label everything as generic binary data. The full list of readable types is on File Formats.
Files hidden inside other files are pulled out and scanned in their own right: images embedded in a document, and every member of a ZIP, TAR, 7z or RAR archive. Images, audio and video need OCR and transcription switched on before their content can be read.
Large objects are streamed rather than loaded whole, so a multi-gigabyte file costs disk rather than memory, and only the part a sampling window asks for is read.
Blobs from this container are read with the shared file pipeline: see Supported File Formats for everything it can open, and OCR & Transcription for reading text out of images, audio and video.
Metadata on every asset
Asset kind · file
| Field | Type | Always present | What it is |
|---|---|---|---|
| size_bytes | integer | Yes | Raw byte size of the content |
| mime_type | string | Yes | Resolved MIME type |
| parse_error | string | No | Set when content extraction failed |
| image_width | integer | No | Width in pixels |
| image_height | integer | No | Height in pixels |
| page_count | integer | No | Number of pages (pdf) |
| paragraph_count | integer | No | Number of paragraphs (docx) |
| table_count | integer | No | Number of tables (docx) |
| row_count | integer | No | Number of data rows |
| columns | object[] | No | Columns as {name, type} objects (type may be empty for csv/xlsx) |
| encoding | string | No | Detected character encoding |
| json_root_type | string | No | Root JSON type: object, array, or scalar |
| top_level_keys | integer | No | Number of top-level keys when the root is an object |
| array_length | integer | No | Length when the root is an array |
| provider | string | Yes | Storage provider label |
| object_key | string | Yes | Object key/path |
| etag | string | No | Object entity tag |
| source_hash | string | No | Hash of the parent asset (embedded files and archive members only) |
| location | string | No | Location within the parent (embedded file location or archive member path) |
Asset kind · image
| Field | Type | Always present | What it is |
|---|---|---|---|
| size_bytes | integer | Yes | Raw byte size of the content |
| mime_type | string | Yes | Resolved MIME type |
| parse_error | string | No | Set when content extraction failed |
| image_width | integer | No | Width in pixels |
| image_height | integer | No | Height in pixels |
| provider | string | No | Storage provider label |
| object_key | string | No | Object key/path |
| etag | string | No | Object entity tag |
| source_hash | string | No | Hash of the parent asset (embedded files and archive members only) |
| location | string | No | Location within the parent (embedded file location or archive member path) |
Asset kind · audio
| Field | Type | Always present | What it is |
|---|---|---|---|
| size_bytes | integer | Yes | Raw byte size of the content |
| mime_type | string | Yes | Resolved MIME type |
| parse_error | string | No | Set when content extraction failed |
| provider | string | Yes | Storage provider label |
| object_key | string | Yes | Object key/path |
| source_hash | string | No | Hash of the parent asset (embedded files and archive members only) |
| location | string | No | Location within the parent (embedded file location or archive member path) |
Asset kind · video
| Field | Type | Always present | What it is |
|---|---|---|---|
| size_bytes | integer | Yes | Raw byte size of the content |
| mime_type | string | Yes | Resolved MIME type |
| parse_error | string | No | Set when content extraction failed |
| provider | string | Yes | Storage provider label |
| object_key | string | Yes | Object key/path |
| source_hash | string | No | Hash of the parent asset (embedded files and archive members only) |
| location | string | No | Location within the parent (embedded file location or archive member path) |
Asset kind · archive
| Field | Type | Always present | What it is |
|---|---|---|---|
| size_bytes | integer | Yes | Raw byte size of the content |
| mime_type | string | Yes | Resolved MIME type |
| parse_error | string | No | Set when content extraction failed |
| provider | string | Yes | Storage provider label |
| object_key | string | Yes | Object key/path |
| etag | string | No | Object entity tag |
| source_hash | string | No | Hash of the parent asset (embedded files and archive members only) |
| location | string | No | Location within the parent (embedded file location or archive member path) |
Lineage
Lineage
This source records no lineage. Nothing in the system it reads describes data moving from one place to another, so no FLOW edges are produced. Related items are still linked — see Lineage & Relationships for what those links mean and how they differ from lineage.
Worth knowing
- Managed identity is the cleanest deployment in AKS — no secret to store or rotate at all.
- Archives and embedded files are expanded into child assets, under configurable caps.
- Turning off content preview gives you a blob inventory with no downloads.
Configuration
Beyond the fields below, every source also has the settings shared by all of them: the sampling strategy, the detectors to run, the scan schedule, and the compute limits for its scan jobs.
Required
Without these, the source will not save.
| Field | Type | Required | What it does | Default |
|---|---|---|---|---|
| required | object | Yes | —no extra properties | — |
| account_url | string | Yes | Azure Blob account URL (for example, https://<account>.blob.core.windows.net)format uri | — |
| container | string | Yes | Azure Blob container name | — |
Secrets
Stored encrypted and never shown again after you save them. See Configuration & Fields.
| Field | Type | Required | What it does | Default |
|---|---|---|---|---|
| masked | object | No | Optional Azure credentials. Leave empty to use managed identity/default credential chain.no extra properties | — |
| azure_account_key | string | No | Azure storage account key | — |
| azure_client_id | string | No | Azure Entra client ID (service principal auth) | — |
| azure_client_secret | string | No | Azure Entra client secret (service principal auth) | — |
| azure_connection_string | string | No | Azure storage connection string (takes precedence over other auth fields) | — |
| azure_sas_token | string | No | Azure SAS token | — |
| azure_tenant_id | string | No | Azure Entra tenant ID (service principal auth) | — |
Optional
Everything you can tune. Sensible defaults apply when you leave them alone.
| Field | Type | Required | What it does | Default |
|---|---|---|---|---|
| optional | object | No | —no extra properties | — |
| connection | object | No | —no extra properties | — |
| connection.max_archive_member_bytes | integer | No | Maximum uncompressed bytes read from a single archive membermin 1024 | 10485760 |
| connection.max_archive_members | integer | No | Maximum member files expanded from one archive into child assets. Bounds fan-out, and is the guard against a zip bomb.min 1, max 10000 | 200 |
| connection.max_archive_total_bytes | integer | No | Maximum uncompressed bytes read across all members of one archive. The decompression-ratio ceiling: a zip bomb hits this before it hits memory.min 1024 | 104857600 |
| connection.max_embedded_files | integer | No | Maximum embedded files (parquet image/audio columns, office media) expanded from one container into child assets per runmin 1, max 10000 | 200 |
| connection.max_file_bytes | integer | No | Refuse any object larger than this many bytes. 0 or unset means no limit: an object above max_object_bytes is spooled to disk rather than held in memory, so file size is bounded by free disk, not by RAM.min 0 | — |
| connection.max_keys_per_page | integer | No | Maximum blobs requested per list pagemin 1, max 1000 | 200 |
| connection.max_object_bytes | integer | No | Maximum bytes of one object held in memory. Larger objects are streamed to a temporary file (or read by byte range where the provider supports it), so this bounds memory rather than the size of file that can be scanned. See max_file_bytes to refuse large objects outright.min 1024 | 5242880 |
| connection.request_timeout_seconds | number | No | Network timeout in seconds for list/download operationsmin 1, max 300 | 30 |
| scope | object | No | Object scope and filtering controls.no extra properties | — |
| scope.exclude_extensions | array | No | Optional extension denylist | — |
| scope.exclude_extensions[] | string | No | — | — |
| scope.include_content_preview | boolean | No | Download object bytes to infer MIME and extract detector-ready text previews | true |
| scope.include_empty_objects | boolean | No | Include zero-byte objects in extraction results | false |
| scope.include_extensions | array | No | Optional extension allowlist (for example, .pdf, .csv, .parquet) | — |
| scope.include_extensions[] | string | No | — | — |
| scope.include_object_metadata | boolean | No | Attach provider metadata (etag, size, content-type hints, timestamps) to asset checksums | true |
| scope.prefix | string | No | Object key prefix filter (for example, exports/2026/) | — |