Skip to content
Vibedata

Re-engineer something you own

Convert a Spark ingestion notebook into a dlt pipeline and prove landed parity

  • When a Spark notebook is the only thing that lands a source into bronze and nobody can run it unattended.
  • When a notebook-based ingestion job has no incremental logic and reloads everything every time.

Convert a Spark ingestion notebook into a dlt pipeline that lands the same bronze schema, and prove the landed data matches what the notebook produced. Covers the pipeline conversion and its landed-parity proof, not adding new resources the notebook did not already load.

Area
Ingestion
Runs on
  • Microsoft Fabric Lakehouse
  • MotherDuck
  • DuckDB
Built with
  • dlt
Readiness
SupportedEverything this Recipe composes runs today, without a case that proves this exact shape.

Sample This Recipe has not been materialized in the Cookbook repository yet. Its trigger, description, prompt, agent guidance and acceptance conditions, and the explanation below, are prototype drafts. Its name, job, area and readiness come from the reconciled Cookbook seed snapshot. Readiness is a separate question from this one: it says whether the capability exists, not whether the writing has been reviewed.

Use this Recipe

Use the VibeData Recipe `spark-ingestion-notebook-to-dlt` at https://getvibedata.ai/cookbook/spark-ingestion-notebook-to-dlt Read the Recipe and execute it in the context of the current Intent.

Recipe id spark-ingestion-notebook-to-dlt · Not yet materialized in the Cookbook repository, so the pointer addresses this page.

Verified by

What has to be observably true before this Recipe is finished.

  • landed row counts match the notebook's last accepted run for the same source window
  • the dlt pipeline runs unattended without a manual notebook execution
  • a second run loads only new or changed records
  • the notebook is not retired in the same change that introduces the pipeline
Recipe promptThe task specification the agent reads. Reference only — it is not what you copy.
Deliver: Convert a Spark ingestion notebook into a dlt pipeline and prove landed parity.

Convert a Spark ingestion notebook into a dlt pipeline that lands the same bronze schema, and prove the landed data matches what the notebook produced. Covers the pipeline conversion and its landed-parity proof, not adding new resources the notebook did not already load.

Execute inside the current Intent. Its Domain, repository, platform, environment and attached sources are the context for this work — read them rather than asking for them.

The work is done when:
- landed row counts match the notebook's last accepted run for the same source window
- the dlt pipeline runs unattended without a manual notebook execution
- a second run loads only new or changed records
- the notebook is not retired in the same change that introduces the pipeline

Report the evidence for each condition above with the result. A condition you cannot meet is something to say, not something to work around.
Agent guidanceHow the agent approaches the work, and what it will not do.

Write down the behaviour that has to stay constant before you touch the current form, and get that agreed. Rebuild in an isolated copy, then reconcile old against new at the grain the consumers read. The reconciliation is the deliverable; the diff is not.

Composes

  • source profiling and incremental load
  • dlt in the project
  • isolated-copy execution and gate verification

Guardrails

  • Build in an isolated copy. Production is read, never written.
  • Do not edit the object being held constant. Parity that required a change to the original is not parity.

What you need

  • Microsoft Fabric Lakehouse, MotherDuck, or DuckDB.
  • Read access to the source. Not write access to production.
  • dlt in the project, or the intent to add it.

How it goes

  1. Name the thing you own, and what about its behaviour has to stay the same.
  2. Let it read the current form and write down the behaviour it is holding constant.
  3. Review the design before any SQL is written. Grain first.
  4. Let it rebuild in an isolated copy of your estate. Production is untouched throughout.
  5. Reconcile the new form against the old one, row by row.
  6. Read the acceptance conditions. They are the contract; the diff is not.

What you end up with

The deliverable, in your own repository, as ingestion work a reviewer who knows the project reads as native to it. Alongside it, the evidence for every one of the acceptance conditions above — which is the part that is still there in three weeks when somebody asks.