Writing

Notes from the pipeline

Things I've learned building data platforms — what broke, what worked, and what the hype cycle gets wrong. New posts most months.

  1. The question to ask before you build what someone asked for

    Stakeholders request a solution. The decision underneath it is usually smaller than the thing they asked you to build.

  2. Writing a spec an engineer doesn't have to decode

    A spec exists to remove the follow-up questions. The only honest way to find out whether it worked is to count them.

  3. The dashboard was wrong for six weeks and nobody noticed

    A join quietly started dropping rows and the revenue dashboard kept looking plausible. What I changed so it never happens silently again.

  4. Semantic models are the most underrated work in analytics

    Nobody puts "built a measure library" on a conference slide, but it took our reporting cycle from two days to four hours.

  5. Iceberg won. Now what?

    Everyone agrees on the table format now. That settles much less than people think.

  6. What consolidating 8 source systems taught me about ELT design

    Lessons from pulling product, marketing, and customer data into one Snowflake analytics layer without breaking anything downstream.

  7. dbt tests are unit tests for your beliefs

    You think order_id is unique. The test is how you find out when you stop being right.

  8. What "done" means when the deliverable is a number

    Software features have a demo anyone can judge. Data work ships a number, and "it looks about right" is how wrong numbers survive.

  9. DuckDB is my new scratchpad

    Not everything needs the warehouse. A local engine that eats parquet has quietly changed how I debug and prototype.

  10. Cutting our Snowflake bill without touching a single dashboard

    Full-refresh models were rebuilding the world nightly. Making them incremental cut warehouse compute hard, and nobody upstairs noticed a thing. That was the point.

  11. The dependency map nobody drew

    Consolidating a lot of systems, the scary part was never a single pipeline. It was the graph of what breaks what, and which human owns each edge.

  12. The anomaly detector that cried wolf

    My first risk model was technically excellent and practically ignored. Fixing it had nothing to do with the model.

  13. Data contracts: promising, mostly aspirational

    The idea is right. The org chart is the hard part. Notes from trying a lightweight version in real life.

  14. BigQuery optimization is mostly about not being clever

    How warehouse workloads got ~45% faster with partition pruning, honest data types, and the removal of my own smart ideas.

  15. What credit data taught me about data modeling

    Financial data is audited, incentive-shaped, and unforgiving about "as of when." Two habits I picked up there and refuse to drop.

  16. Nobody reads your data docs. Write them anyway.

    Documentation and lineage felt like chores until self-service went up 60% and my interrupt load collapsed. Turns out docs are read at exactly one moment.

  17. LLMs in the pipeline: where they actually help

    Ignore the text-to-SQL demos. The wins are smaller, weirder, and mostly about unstructured inputs and terrible source data.

  18. From business analytics to data engineering: what actually transferred

    I came in through the analytics door, not the CS door. Less of a handicap than the discourse suggests, with a couple of real gaps to close.

  19. Translation runs both ways. I was only good at one.

    Everyone practices turning business needs into technical requirements. The return trip, explaining a constraint in terms someone actually cares about, gets almost no practice.

  20. Kicking the tires on Microsoft Fabric

    A Snowflake-and-dbt person spends real evenings with Fabric and comes back with mixed, specific feelings.

  21. Boring pipelines are a feature

    The most senior thing a data pipeline can do is nothing interesting. On idempotency, retries, and the luxury of not being needed at 3 a.m.