DATARIG LITE BY EXASCALE — FREE FOREVER

The AI data engineer that runs on your machine.

DataRig Lite is free for your desktop: describe a pipeline in plain English, review exactly what it builds, then approve — governed Iceberg tables, orchestrated dbt models, and live dashboards, without your data ever leaving your laptop.

Free forever · Windows & macOS · no account needed · Cloud & Enterprise coming soon

  • Zero data retention — nothing leaves your machine
  • Review-gated agents — nothing runs without approval
  • Open Iceberg tables — no lock-in
  • All tools built into the rig

Part of the ExaScale DataRig platform: start free on your desktop, scale to managed Cloud, or deploy in your own VPC — same pipelines, same governance, zero rewrites.

Everything you need, packaged in one install. Nothing hand-wired.

  • Apache Iceberg
  • Apache Spark
  • Trino
  • Apache Polaris
  • Kafka + Debezium
  • Apache Airflow
  • dbt Core
  • Apache Superset
How it works

From plain English to production, in four steps

You stay in control: DataRig proposes the exact plan, you approve the exact hash, your machine runs it.

  1. STEP 01

    Connect

    Drop in CSV, JSON, or Parquet files — or connect PostgreSQL / SQL Server for continuous CDC. All local.

  2. STEP 02

    Describe

    Say what you want: “load this into Iceberg,” “add dbt tests,” “build a revenue dashboard.”

  3. STEP 03

    Review & approve

    Inspect source analysis, plan steps, and generated files. Approve binds to an exact content hash.

  4. STEP 04

    Run & monitor

    Watch live Spark logs, staged progress, and status chips — with error-specific recovery if anything fails.

AI Data Explorer

Explore governed data, without leaving the rig

Browse your Iceberg catalog tree, inspect schemas, and run read-only SQL — right beside the chat. AI-suggested queries load here for review; they never auto-run.

ON THE ROADMAP Curated metadata → accurate NL-to-SQL

Annotate your tables and columns once; AI-drafted SQL gets dramatically more accurate. Every draft still passes a validation layer — statement allowlists, reference resolution, forced row limits — before it can run.

ON THE ROADMAP Template chips & semantic layer

Deterministic, zero-model starting points for common queries, built on your governed table metadata — the reliable path to answers before any AI is involved.

CATALOG
▾ local ▾ user · orders▸ customers· order_items· revenue_daily
SQL · READ-ONLY · 200 ROWS MAX
SELECT
  order_date,
  SUM(amount) AS revenue
FROM local.user.orders
GROUP BY order_date
ORDER BY order_date DESC;
SCHEMA
order_id BIGINT customer_id BIGINT order_date DATE amount DECIMAL(12,2)
RESULT · 7 ROWS
order_daterevenue
2026-08-2518,204.50
2026-08-2416,930.00
2026-08-2321,417.75
2026-08-2214,882.10
Preview capped at 200 rows · full power via Trino or DBeaver
Built into the rig

All the tools, pre-connected

Specialist consoles are packaged, wired, and one click away — with the DataRig console as your single home base.

Trino SQL engine
Airflow Orchestration
dbt Docs Models & lineage
Superset Dashboards
Kafka Connect CDC pipeline
Grafana Runtime health
Capabilities

The whole pipeline, one conversation

Sources → Iceberg → Airflow → dbt → Superset. Each stage is a governed skill with its own validated contract.

⇥

Ingestion

Land CSV, JSON, JSONL, and Parquet files — or HTTPS APIs — into governed Iceberg tables. Full load today; XML profiling is review-ready.

⇄

CDC & streaming

Capture changes from PostgreSQL and SQL Server with Kafka + Debezium. Snapshot, updates, deletes, restart recovery — verified end to end.

⚙

Transformation

dbt models generated with tests and docs, executed by Airflow against Trino — every artifact reviewable before it runs.

◷

Orchestration

Airflow DAGs created, scheduled, paused, retried, and refreshed in place. Run-scoped docs and live status inside the console.

✓

Data quality

Freshness checks, row-count parity, and null/duplicate tests proposed as part of the plan — not bolted on afterwards.

◎

Analytics & BI

Superset datasets, charts, and dashboards provisioned from your governed business layer, queried through Trino.

Automations

Works while you sleep

Approve a pipeline once and DataRig keeps it running. Airflow schedules every load, dbt tests guard every refresh, and failed tasks retry on their own — all on your machine, with the full history one click away.

  • Scheduled, not babysat. DAGs are created, scheduled, paused, and refreshed in place.
  • Tested every run. Freshness, row-count parity, and null/duplicate checks run with each refresh.
  • Self-healing by default. Transient failures retry automatically; the run log shows exactly what happened.
One platform · three runtimes

Start on your laptop. Never start over.

The pipelines you build in Lite are the same artifacts every edition runs: Iceberg tables, Airflow DAGs, tested dbt models, Superset dashboards. When you scale, your work scales with you — no rewrites, no migration project.

LITE · FREE FOREVER

Your machine

Individual engineers, pilots, local-first teams

  • Complete local lakehouse + review-gated AI
  • All tools built in — zero setup beyond install
  • Zero data retention, telemetry off
$0 forever
Download Lite
CLOUD · COMING SOON

Our cloud

Teams that outgrew the laptop but not their budget

  • Everything in Lite — fully managed and always on
  • Automated upgrades, backups, and monitoring
  • Transparent usage metering — no opaque credits, no idle cluster sprawl
  • Email + Slack support
Usage-based pay for what you use
Join the waitlist
ENTERPRISE · COMING SOON

Your VPC

Regulated orgs that need sovereign control

  • Everything in Cloud, deployed inside your network perimeter
  • SSO, organization RBAC, audit export
  • Customer-managed encryption + private networking
  • Contractual SLA with dedicated support
Annual contract + SLA
Contact us

Pipeline definitions carry forward across every edition. Your data stays in open Apache Iceberg formats everywhere.

On the roadmap · multi-tool connectivity

Every platform locks its AI inside its own wall

Cortex lives in SnowflakeGenie lives in DatabricksCopilot lives in FabricGemini lives in BigQuery

ExaScale is the bridge when your data
doesn't live in one wall.

⇄

Query across warehouses

Planned: read-only federation to Snowflake, Redshift, Databricks, and BigQuery catalogs — join lakehouse and warehouse data without ETL sprawl, behind the same review-gate stack.

◈

Bring your semantics

Planned: import existing semantic models and metadata from the platforms you already pay for — your definitions travel with you instead of being rebuilt per vendor.

✓

Gated write-back, last

Federation arrives read-only first. Any future write-back adapter lands behind explicit, hash-bound approvals — never silent synchronization.

Designed, not yet scheduled — join the waitlist to shape the priority.

Ship your first governed pipeline today.

Free forever on Windows & macOS. Your data never leaves your machine.