Test Databricks
CAT tests Delta tables in Databricks, whatever compute serves them — an all-purpose cluster, a SQL warehouse, serverless. The native Databricks@1 provider is the recommended way: nothing to install, a token, a connection string, and the tests that pay off: source against bronze, silver against gold, the lakehouse against Power BI.
Your organisation standardises on the Databricks ODBC driver, or has a DSN already? The ODBC route has its own page: Test Databricks through ODBC.
What you connect to
Any compute that can run SQL over your tables, through the same driver and the same connection string — only two values differ:
| Compute | Where to copy Server hostname and HTTP path |
|---|---|
| SQL warehouse — classic or serverless; the right choice for tests | SQL → SQL Warehouses → the warehouse → Connection details tab |
| All-purpose cluster | Compute → the cluster → Configuration → Advanced options → JDBC/ODBC tab |
A warehouse that is stopped starts when CAT connects (and your run waits for it); a cluster must be running. Unity Catalog tables are addressed with three parts, catalog.schema.table; the adbc.connection.catalog and adbc.connection.db_schema keys of the connection string set the default.
Prerequisites
- A way to authenticate: a personal access token (User settings → Developer → Access tokens → Generate new token) for your own runs; for unattended runs a service principal with an OAuth secret. Either way, the secret goes into an environment variable — never into the project file.
Add the data source
CAT Studio: Data sources → New data source, technology Databricks, the guided form. In YAML, Databricks@1 — everything in the string, nothing to maintain on the machine, and the values that change come from environment variables:
Data sources:
- Name: lakehouse
Provider: Databricks@1
Technology: Databricks
Connection string: >
adbc.spark.host=%DATABRICKS_HOST%;
adbc.spark.path=%DATABRICKS_HTTP_PATH%;
adbc.spark.auth_type=token;
adbc.spark.token=%DATABRICKS_TOKEN%;
Data sources:
- Name: lakehouse
Provider: Databricks@1
Technology: Databricks
Connection string: >
adbc.spark.host=%DATABRICKS_HOST%;
adbc.spark.path=%DATABRICKS_HTTP_PATH%;
adbc.spark.auth_type=oauth;
adbc.databricks.oauth.grant_type=client_credentials;
adbc.databricks.oauth.client_id=%DATABRICKS_CLIENT_ID%;
adbc.databricks.oauth.client_secret=%DATABRICKS_CLIENT_SECRET%;
The service principal’s application id and OAuth secret come from the workspace’s Identity and access → Service principals.
See Databricks@1 for the full key reference.
Write the tests
- Name
- Gold customers table is loaded
- Suite
- Smoke tests
- Data source
- lakehouse
- Query
- SELECT * FROM gold.dim.customer LIMIT 10
- Expectation
- set is not empty
- Name
- Silver orders equal the source system
- Suite
- Source vs lakehouse
- First data source
- erp
- First query
- SELECT order_id, amount FROM dbo.Orders WHERE status = ‘closed’ ORDER BY order_id
- Second data source
- lakehouse
- Second query
- SELECT order_id, amount FROM silver.sales.orders WHERE status = ‘closed’ ORDER BY order_id
- Expectation
- sets match
- Key
- order_id
Tests:
- Name: Gold customers table is loaded
Suite: Smoke tests
Data source: lakehouse
Query: SELECT * FROM gold.dim.customer LIMIT 10
Expectation: set is not empty
- Name: Silver orders equal the source system
Suite: Source vs lakehouse
First data source: erp
First query: SELECT order_id, amount FROM dbo.Orders WHERE status = 'closed' ORDER BY order_id
Second data source: lakehouse
Second query: SELECT order_id, amount FROM silver.sales.orders WHERE status = 'closed' ORDER BY order_id
Expectation: sets match
Key: order_id
The tests that pay off on a lakehouse are the ones across its layers and across systems — exactly what a notebook inside the platform cannot see: the source system against bronze, silver against gold, gold against the Power BI model built on it (Test Power BI — both sides in one test). Both sides of a sets match ordered, or Sort data: true.
Generate tests from Unity Catalog
Unity Catalog’s information_schema is a ready-made metadata source: one template — “every table in silver has rows”, “every table has a _loaded_at column” — plus a metadata query over system.information_schema.tables or …columns gives you one test per table, refreshed on every open. See Generate tests from metadata.
Related
- CAT in Databricks notebooks — when you want to run CAT inside Databricks, and why usually not.
- Databricks@1 — the provider, with the full key reference.
- Test Databricks through ODBC — the ODBC route, for organisations standardised on the driver.
- Work with secrets — the token and the secret.