Load
Bring in CSV, Parquet, JSON or Excel. DEn profiles every column: types, nulls, cardinality and sample values.
DEn Data Engineering Agent is live.
DEn audits your dataset, checks it against a schema contract, and writes the SQL and PySpark to fix it. Designed and built by Aaron Tekle.
profile_datasettypes, nulls, cardinalityrun_quality_checksissues ranked by severityvalidate_schemacontract mismatch foundrun_sqlread-only, DuckDBgenerate_pipelineSQL + PySpark-- keep the latest row per order CREATE OR REPLACE TABLE orders_clean AS SELECT * FROM dataset QUALIFY ROW_NUMBER() OVER ( PARTITION BY order_id ORDER BY updated_at DESC) = 1;
Built on an open, inspectable stack
Qwen3-Coder Hugging Face Inference DuckDB SQLGlot PySpark GradioThe product
DEn is a data engineering agent. It works through the same steps a careful engineer would, and shows its work at each one.
Bring in CSV, Parquet, JSON or Excel. DEn profiles every column: types, nulls, cardinality and sample values.
Each column is checked for nulls, duplicates and other problems, and every issue is ranked high, medium or low.
The data is tested against a schema contract you can edit, with a clear pass or fail plus warnings.
The agent tests read-only SQL, then writes the cleanup pipeline in SQL and as a PySpark equivalent.
Pipelines come out in the dialect of your warehouse, so they drop into the stack you have instead of the one a tool prefers.
SQL is validated in a DuckDB sandbox that can read your data but never change it.
Fix what matters first. Every finding carries a severity.
A full tool trace shows each call the agent made and what it saw, so nothing is a black box.
Generated as code for your cluster. No Spark runtime is needed to review it.
Selected work
DEn is designed and built end to end by Aaron Tekle, from the agent and its tools to the interface.
A live workspace where anyone can load a dataset, run a quality audit, validate a schema contract and have an agent write the SQL and PySpark to fix what it finds.
Questions
A data engineering agent. It profiles a dataset, audits its quality, validates it against a schema contract, then writes SQL and PySpark pipelines to fix the problems it finds. You can see every tool call it makes.
DuckDB, Spark SQL, PostgreSQL, MySQL, Snowflake, BigQuery and T-SQL.
Load the example dataset, run a quality audit and let the agent write the fix, all in your browser.