DEn Data Engineering Agent is live.

Messy data in. Clean pipelines out.

DEn audits your dataset, checks it against a schema contract, and writes the SQL and PySpark to fix it. Designed and built by Aaron Tekle.

den / agent run Example
  • profile_datasettypes, nulls, cardinality
  • run_quality_checksissues ranked by severity
  • validate_schemacontract mismatch found
  • run_sqlread-only, DuckDB
  • generate_pipelineSQL + PySpark
-- keep the latest row per order
CREATE OR REPLACE TABLE orders_clean AS
SELECT * FROM dataset
QUALIFY ROW_NUMBER() OVER (
  PARTITION BY order_id
  ORDER BY updated_at DESC) = 1;

Built on an open, inspectable stack

Qwen3-Coder Hugging Face Inference DuckDB SQLGlot PySpark Gradio

The product

From raw file to reviewed pipeline in four steps.

DEn is a data engineering agent. It works through the same steps a careful engineer would, and shows its work at each one.

01

Load

Bring in CSV, Parquet, JSON or Excel. DEn profiles every column: types, nulls, cardinality and sample values.

In the appLoad and Profileor let the example dataset load on its own
02

Audit

Each column is checked for nulls, duplicates and other problems, and every issue is ranked high, medium or low.

In the appRun Quality Auditfindings land in the Detected issues table
03

Validate

The data is tested against a schema contract you can edit, with a clear pass or fail plus warnings.

In the appValidate Schemaedit the contract JSON, then check again
04

Generate

The agent tests read-only SQL, then writes the cleanup pipeline in SQL and as a PySpark equivalent.

In the appRun Agentcopy the SQL and PySpark from their tabs

SQL in the dialect you already run

Pipelines come out in the dialect of your warehouse, so they drop into the stack you have instead of the one a tool prefers.

duckdbsparkpostgressnowflake bigquerymysqltsql
Pick it in the app withSQL dialecton the Pipeline tab

Read-only by design

SQL is validated in a DuckDB sandbox that can read your data but never change it.

Queries run against one table:dataset

Issues ranked by risk

Fix what matters first. Every finding carries a severity.

highmediumlow
Listed inDetected issueson the Audit tab

Every step on the record

A full tool trace shows each call the agent made and what it saw, so nothing is a black box.

OpenTool traceafter any run

PySpark, ready to deploy

Generated as code for your cluster. No Spark runtime is needed to review it.

SQL + PySparkSQL onlyPySpark only

Selected work

Built in the open.

DEn is designed and built end to end by Aaron Tekle, from the agent and its tools to the interface.

Product, design and engineering

DEn, an agentic data engineering workspace

A live workspace where anyone can load a dataset, run a quality audit, validate a schema contract and have an agent write the SQL and PySpark to fix what it finds.

Model
Qwen3-Coder via Hugging Face Inference
SQL engine
DuckDB, read-only sandbox
SQL parsing
SQLGlot, seven dialects
Interface
Gradio, light and dark themes

Questions

Good to know.

What exactly is DEn?

A data engineering agent. It profiles a dataset, audits its quality, validates it against a schema contract, then writes SQL and PySpark pipelines to fix the problems it finds. You can see every tool call it makes.

Which SQL dialects are supported?

DuckDB, Spark SQL, PostgreSQL, MySQL, Snowflake, BigQuery and T-SQL.

Try DEn right now.

Load the example dataset, run a quality audit and let the agent write the fix, all in your browser.