Get Help

Test Databricks

CAT tests Delta tables in Databricks, whatever compute serves them — an all-purpose cluster, a SQL warehouse, serverless. The native Databricks@1 provider is the recommended way: nothing to install, a token, a connection string, and the tests that pay off: source against bronze, silver against gold, the lakehouse against Power BI.

Your organisation standardises on the Databricks ODBC driver, or has a DSN already? The ODBC route has its own page: Test Databricks through ODBC.

What you connect to

Any compute that can run SQL over your tables, through the same driver and the same connection string — only two values differ:

Compute Where to copy Server hostname and HTTP path
SQL warehouse — classic or serverless; the right choice for tests SQL → SQL Warehouses → the warehouse → Connection details tab
All-purpose cluster Compute → the cluster → Configuration → Advanced options → JDBC/ODBC tab

A warehouse that is stopped starts when CAT connects (and your run waits for it); a cluster must be running. Unity Catalog tables are addressed with three parts, catalog.schema.table; the adbc.connection.catalog and adbc.connection.db_schema keys of the connection string set the default.

Prerequisites

  • A way to authenticate: a personal access token (User settings → Developer → Access tokens → Generate new token) for your own runs; for unattended runs a service principal with an OAuth secret. Either way, the secret goes into an environment variable — never into the project file.

Add the data source

CAT Studio: Data sources → New data source, technology Databricks, the guided form. In YAML, Databricks@1 — everything in the string, nothing to maintain on the machine, and the values that change come from environment variables:

Data sources:
- Name: lakehouse
  Provider: Databricks@1
  Technology: Databricks
  Connection string: >
    adbc.spark.host=%DATABRICKS_HOST%;
    adbc.spark.path=%DATABRICKS_HTTP_PATH%;
    adbc.spark.auth_type=token;
    adbc.spark.token=%DATABRICKS_TOKEN%;
Data sources:
- Name: lakehouse
  Provider: Databricks@1
  Technology: Databricks
  Connection string: >
    adbc.spark.host=%DATABRICKS_HOST%;
    adbc.spark.path=%DATABRICKS_HTTP_PATH%;
    adbc.spark.auth_type=oauth;
    adbc.databricks.oauth.grant_type=client_credentials;
    adbc.databricks.oauth.client_id=%DATABRICKS_CLIENT_ID%;
    adbc.databricks.oauth.client_secret=%DATABRICKS_CLIENT_SECRET%;

The service principal’s application id and OAuth secret come from the workspace’s Identity and access → Service principals.

See Databricks@1 for the full key reference.

Write the tests

Name
Gold customers table is loaded
Suite
Smoke tests
Data source
lakehouse
Query
SELECT * FROM gold.dim.customer LIMIT 10
Expectation
set is not empty
Name
Silver orders equal the source system
Suite
Source vs lakehouse
First data source
erp
First query
SELECT order_id, amount FROM dbo.Orders WHERE status = ‘closed’ ORDER BY order_id
Second data source
lakehouse
Second query
SELECT order_id, amount FROM silver.sales.orders WHERE status = ‘closed’ ORDER BY order_id
Expectation
sets match
Key
order_id
Tests:
- Name: Gold customers table is loaded
  Suite: Smoke tests
  Data source: lakehouse
  Query: SELECT * FROM gold.dim.customer LIMIT 10
  Expectation: set is not empty

- Name: Silver orders equal the source system
  Suite: Source vs lakehouse
  First data source: erp
  First query: SELECT order_id, amount FROM dbo.Orders WHERE status = 'closed' ORDER BY order_id
  Second data source: lakehouse
  Second query: SELECT order_id, amount FROM silver.sales.orders WHERE status = 'closed' ORDER BY order_id
  Expectation: sets match
  Key: order_id

The tests that pay off on a lakehouse are the ones across its layers and across systems — exactly what a notebook inside the platform cannot see: the source system against bronze, silver against gold, gold against the Power BI model built on it (Test Power BI — both sides in one test). Both sides of a sets match ordered, or Sort data: true.

Generate tests from Unity Catalog

Unity Catalog’s information_schema is a ready-made metadata source: one template — “every table in silver has rows”, “every table has a _loaded_at column” — plus a metadata query over system.information_schema.tables or …columns gives you one test per table, refreshed on every open. See Generate tests from metadata.