The Analytics Engineering Podcast

Tristan Handy has been curating the Analytics Engineering Roundup newsletter since 2015, pulling together the internet’s best data science & analytics articles. Tristan and co-host Julia Schottenstein now bring the Roundup to real life, hosting biweekly conversations with data practitioners inventing the future of analytics engineering. You can view full episode summaries and read back issues of the Roundup newsletter at https://roundup.getdbt.com. The podcast is sponsored by dbt labs, makers of the data transformation framework dbt. To reach our team, drop a note to [email protected].

Follow on

Episodes

Julia, Pedram Navid + Taylor Murphy Recap Data Council

Julia just got back from Data Council in Austin, a conference organized by Pete Sonderling, where lots of startups share what they're building, data practitioners go to learn in hands-on workshops, and of course investors go to spot the next big trend. In this episode, Taylor Murphy (Head of Product & Data at Meltano) + Pedram Navid (Founder, West Marin Data) join Julia to recap the conference and have a bit of fun. They talked streaming, how the MDS is growing up, new SQL variants, and, of cour...

Apr 07, 2023•42 min•Season 1Ep. 44

Cloud Warehouse Cost Optimization (w/ Niall Woodward + Brad Culberson)

Brad Culberson is a Principal Architect in the Field CTO’s office at Snowflake. Niall Woodward is a co-founder of SELECT, a startup providing optimization and spend management software for Snowflake customers. In this conversation with Tristan and Julia, Brad and Niall discuss all things cost optimization: cloud vs on-prem, measuring ROI, and tactical ways to get more out of your budget. For full show notes and to read 6+ years of back issues of the podcast's companion newsletter, head to https:...

Mar 24, 2023•46 min•Season 1Ep. 43

dbt Labs + Transform Join Forces on Metrics (w/ Nick Handel + Drew Banin)

Nick Handel, as co-founder at Transform, helped develop the popular open source metrics framework MetricFlow. Drew Banin, a co-founder at dbt Labs, helped build the initial version of the dbt Semantic Layer, which launched last year. Transform was acquired in February by dbt Labs, and in this conversation with Tristan, they talk through their collective plans for the future of the dbt Semantic Layer. For full show notes and to read 7+ years of back issues of the podcast's companion newsletter, h...

Mar 10, 2023•43 min•Season 1Ep. 42

What Can Generative AI Do for Data People? (W/ Sarah Nagy + Chris Aberger)

Sarah and Chris are both at the forefront of bringing the promise of gen AI to our actual work as data people—which is a unique challenge! Precise truth is critical for business questions in a way that it’s not for a consumer search query. Sarah Nagy is the CEO of Seek AI, a startup that aims to use natural language processing to change how professionals work with data. Chris Aberger currently leads Numbers Station AI, a startup focused on data-intensive workflow automation. In this conversation...

Feb 24, 2023•48 min•Season 1Ep. 41

3rd Party Data, Demystified

Auren Hoffman currently serves as the CEO and Chief Historian at SafeGraph, a data-as-a-service company he founded, which provides primarily location data. In this conversation with Tristan and Julia, Auren shares how truly few companies are making use of 3rd-party datasets today, how opening up more datasets to public research could help us solve big problems, and a fun fact about Abraham Lincoln's (!) work in the industry. For full show notes and to read 6+ years of back issues of the podcast'...

Feb 10, 2023•45 min•Season 1Ep. 40

A Romp Through Database History (w/ Postgres co-creator Mike Stonebraker + Andy Palmer)

Mike Stonebraker is a veritable database pioneer and a Turing Award recipient. In addition to teaching at MIT, he is a serial entrepreneur and co-creator of Postgres. Andy Palmer is a veteran business leader who serves as the CEO of Tamr, a company he co-founded with Mike. Through his seed fund Koa Labs, Andy has helped found and/or fund numerous innovative companies in diverse sectors, including health care, technology, and the life sciences. In this conversation with Tristan and Julia, Mike an...

Jan 27, 2023•48 min•Season 1Ep. 39

What Does Apache Arrow Unlock for Analytics? (w/ Wes McKinney)

Wes McKinney is the creator of pandas, co-creator of Apache Arrow, and now Co-founder/CTO at Voltron Data. In this conversation with Tristan and Julia, Wes takes us on a tour of the underlying guts, from hardware to data formats, of the data ecosystem. What innovations, down to the hardware level, will stack to lead to significantly better performance for analytics workloads in the coming years? To dig deeper on the Apache Arrow ecosystem, check out replays from their recent conference at https:...

Jan 06, 2023•47 min•Season 1Ep. 38

Minimum Viable Experimentation

Product experimentation is full of potholes for companies of any size, given the number of pieces (tooling, culture, process, persistence) that need to come together to be successful. Vijaye Raji (currently Statsig, formerly Facebook + Microsoft) and Sean Taylor (currently Motif Analytics, formerly Facebook + Lyft) have navigated these failure modes, and are here to help you (hopefully) do the same. This convo with Tristan + Julia is light on tooling + heavy on process: how to watch out for spil...

Dec 16, 2022•46 min•Season 1Ep. 37

The Data Generalist's Vision Quest (LIVE w/ Stephen Bailey)

The first LIVE IRL episode! Stephen Bailey, data engineer at Whatnot and writer of an incredibly entertaining data substack, joins Tristan for a follow-up conversation to Stephen’s Coalesce talk, “Excel at nothing: how to be an effective generalist.” You can read Stephen’s writing at https://stkbailey.substack.com/ . For full show notes and to read 6+ years of back issues of the podcast's companion newsletter, head to https://roundup.getdbt.com. The Analytics Engineering Podcast is sponsored by ...

Dec 02, 2022•27 min•Season 1Ep. 35

Why You'll Need Data Contracts (w/ Chad Sanderson + Prukalpa)

WARNING: This episode contains detailed discussion of data contracts. The modern data stack introduces challenges in terms of collaboration between data producers and consumers. How might we solve them to ultimately build trust in data quality? Chad Sanderson leads the data platform team at Convoy, a late-stage series-E freight technology startup. He manages everything from instrumentation and data ingestion to ETL, in addition to the metrics layer, experimentation software and ML. Prukalpa Sank...

Nov 18, 2022•49 min•Season 1Ep. 34

How Does Data Drive Growth in Practice? (w/ Abhi Sivasailam)

Abhi is a growth and data leader, and an excellent Twitter follow. Most recently, he was Head of Growth and Analytics at Flexport, where he helped the company to grow 10x over the past 3 years. Previously, Abhi led growth and data teams at Keap, Hustle, and Honeybook. In this conversation with Tristan and Julia, Abhi explains his methodology for setting up a new growth data organization, and how you might be falling victim to the dreaded "arbitrary uniqueness" bug. For full show notes and to rea...

Nov 04, 2022•50 min•Season 1Ep. 33

Katie Bauer: Data Scientists Are Not Pizza

Katie was a founding member of Reddit's data science team and, currently, as Twitter’s Data Science Manager, she leads the company’s infrastructure data science and analytics organization. In this conversation with Tristan and Julia, Katie explores how, as a manager, to help data people (especially those new to the field!) do their best work. For full show notes and to read 6+ years of back issues of the podcast's companion newsletter, head to https://roundup.getdbt.com . The Analytics Engineeri...

Jul 29, 2022•43 min•Season 1Ep. 33

Data Activation Everywhere (w/ Julie Beynon of Clearbit)

As Head of Analytics at Clearbit, Julie serves as a data team of one in a 200+ person company (wow!). In this conversation with Tristan and Julia, Julie dives into how she's helped Clearbit implement data activation throughout the business, and realize the glorious dream of self-serve analytics. For full show notes and to read 6+ years of back issues of the podcast's companion newsletter, head to https://roundup.getdbt.com . The Analytics Engineering Podcast is sponsored by dbt Labs....

Jul 15, 2022•44 min•Season 1Ep. 32

The Personal Data Warehouse (w/ Jordan Tigani of MotherDuck)

Jordan Tigani is an expert in large-scale data processing, having spent a decade+ in the development and growth of BigQuery, and later SingleStore. Today, Jordan and his team at MotherDuck are in the early days of working on commercial applications for the open source DuckDB OLAP database. In this conversation with Tristan and Julia, Jordan dives into the origin story of BigQuery, why he thinks we should do away with the concept of working in files, and how truly performant “data apps” will requ...

Jul 01, 2022•52 min

Making Sense of the Last 2 Years in Data

Matt Bornstein and Jennifer Li (and their co-author Martin Casado) of a16z have compiled arguably the most nuanced diagram of the data ecosystem ever made. They recently refreshed their classic 2020 post, "Emerging Architectures for Modern Data Infrastructure" and in this conversation, Tristan attempts to pin down: what does all of this innovation in tooling mean for data people + the work we're capable of doing? When will the glorious future come to our laptops? For full show notes and to read ...

Jun 17, 2022•47 min

Building an Open Source Company (w/ Aaron Katz of ClickHouse)

ClickHouse, the lightning-fast open source OLAP database, was initially released in 2016 as an open source project out of Yandex, the Russian search giant. In 2021, Aaron Katz helped form a group to spin it out of Yandex as an independent company, dedicated to the development + commercialization of the open source project. In this conversation with Tristan and Julia, Aaron gets into why he believes open source, independent software companies are the future. And of course, this conversation would...

Jun 03, 2022•39 min•Season 1Ep. 29

"To Move, or Not to Move" (Data). That is the Question.

Justin Borgman is the co-founder, Chairman and CEO of Starburst, and has almost a decade spent in senior executive roles building new businesses in the data warehousing and analytics space. In this conversation with Tristan and Julia, Justin dives into the nuts and bolts of Trino, the open source distributed query engine, and explores how teams are adopting a data mesh architecture without making a mess. For full show notes and to read 6+ years of back issues of the podcast's companion newslette...

May 20, 2022•40 min•Season 1Ep. 28

What’s The Role Of AI in BI?

Amit Prakash is Co-founder and CTO at ThoughtSpot. He has a deep background in search, having previously led the AdSense engineering team at Google and served on the early Bing team at Microsoft. In this conversation with Tristan and Julia, Amit gets real about the promise of AI in data: which applications are being widely used today, and which are still a few years out? For full show notes and to read 6+ years of back issues of the podcast's companion newsletter, head to https://roundup.getdbt....

May 06, 2022•45 min

Automating Away Your Work w/ Configuration-as-Code (w/ Sarah Krasnik)

Most recently leading a data engineering team at Perpay, Sarah has built and managed data platforms end to end by working closely with internal engineering, product, and operational teams. She recently left her role to pursue a wide variety of endeavors, including writing on her Substack ( https://sarahsnewsletter.substack.com/ ). In this conversation with Tristan and Julia, Sarah dives into how configuration-as-code can automate away data work, why you might want to consider adding a data lake ...

Apr 22, 2022•44 min

The Hard Problems™️ of Data Observability w/ Kevin Hu of Metaplane

As a PhD candidate at MIT, Kevin (and friends) published Sherlock, a data type detection engine (a surprisingly bedeviling problem) for data cleaning + data discovery. Now as co-founder and CEO of Metaplane, a data observability startup, Kevin applies these same automated data discovery methods to help data teams keep their data healthy. In this conversation with Tristan & Julia, Kevin wins the coveted award for “most crystal-clear explanations of complex technical concepts through physics analo...

Apr 08, 2022•43 min

The Bundling vs Unbundling Debate w/ Tristan, Benn Stancil and David Jayatillake

A debate has erupted on data Twitter and data Substack - should the modern data stack remain unbundled, or should it consolidate? In this conversation, Benn Stancil (Mode), David Jayatillake (Avora) and our host Tristan Handy try to make some sense of this debate, and play with various future scenarios for the modern data stack. For full show notes and to read 6+ years of back issues of the podcast's companion newsletter, head to https://roundup.getdbt.com . The Analytics Engineering Podcast is ...

Mar 25, 2022•43 min

One Database to Rule All Workloads? With Jon "Natty" Natkins of dbt Labs

Will the dream of a mythical database to handle all workloads (transactional + analytical) ever become a reality, or does it violate the laws of physics? This question sparked a hearty debate internally at dbt Labs, and Jon "Natty" Natkins joins Julia here to continue the conversation. Natty knows databases, and this episode will take you on a historical romp through the rise and fall of Hadoop, the transition to cloud data warehouses, and what's waiting for us next in database-land. For full sh...

Mar 11, 2022•36 min

Ashley Sherwood (AE @ Hubspot): Permissionless Innovation for Data Teams

Ashley is a Principal Analytics Engineer at Hubspot, and has helped lead their implementation of dbt. Ashley makes unique connections in her writing and work. On her Substack, "syntax error at or near ❤️," Ashley might be found comparing growing companies to butterflies, or going deep on how to accommodate sensitive people in the workplace. In this conversation with Tristan & Julia, Ashley dives into the nuts and bolts of her trajectory pushing data innovation forward at Hubspot. For full show n...

Feb 25, 2022•46 min

Tristan in the Hot Seat

In this very special episode, we’ll be turning the spotlight on co-host Tristan Handy, the CEO & Co-founder of dbt Labs. In this AMA with Julia, you’ll get to know more about Tristan as a human, as a writer, and as the CEO of dbt Labs helping to push the analytics engineering practice forward. For full show notes and to read 6+ years of back issues of the podcast's companion newsletter, head to https://roundup.getdbt.com.

Dec 17, 2021•39 min•Season 1Ep. 20

[COALESCE] Down With "Data Science" w/ Emilie Schario of Amplify Partners

Your company has one definition for revenue across the organization, one definition of the customer, and one definition of sign-up. For people whose jobs are so defined by ensuring we’re aligned, we can’t seem to standardize on one definition for the Data Scientist. In this talk, Emilie Schario (Data Strategist-in-Residence at Amplify Partners and longtime dbt community member) proposes we lobby against the title Data Scientist, instead choosing some variation of the Core Four Data Roles: Data A...

Dec 10, 2021•46 min•Season 1Ep. 19

[COALESCE] Peeking Into the Future of Data Analytics w/ Julia

How is the data landscape evolving, what trends should you pay attention to and which should you ignore? In this panel, Julia Schottenstein (our fearless co-host and dbt Labs product manager) catches up with Sarah Catanzaro, Jennifer Li and Astasia Myers to dive into the trends playing out in our work. Register to catch the rest of Coalesce, the Analytics Engineering Conference, at https://coalesce.getdbt.com. The Analytics Engineering Podcast is brought to you by dbt Labs....

Dec 09, 2021•45 min•Season 1Ep. 18

[COALESCE] The Modern Data Experience w/ Benn Stancil of Mode

In this talk, former podcast guest Benn Stancil walks through what he believe the next evolution of the modern data stack should look like - and more importantly, how those who use it should experience it. Register to catch the rest of Coalesce, the Analytics Engineering Conference, at https://coalesce.getdbt.com. The Analytics Engineering Podcast is brought to you by dbt Labs.

Dec 09, 2021•30 min•Season 1Ep. 17

[COALESCE] Data Analytics In A Snowflake World ft. Christian Kleinerman

Where does Snowflake go from here? What meta trends and technologies play into that vision? How does that impact the world of data analytics? Christian and Tristan have no shortage of opinions or ideas. This is your chance to hear some of them, live and unfiltered. Register to catch the rest of Coalesce, the Analytics Engineering Conference, at https://coalesce.getdbt.com. The Analytics Engineering Podcast is brought to you by dbt Labs.

Dec 09, 2021•24 min•Season 1Ep. 16

[COALESCE] You Don’t Need Another Database W/ Reynold Xin of Databricks and Drew Banin of dbt Labs

Reynold Xin is a technical co-founder and Chief Architect at Databricks. He’s also a co-creator and the top contributor to the Apache Spark project. In this casual conversation with Drew Banin, co-founder and Chief Product Officer at dbt Labs, the two will be discussing the data infrastructure trends they find most interesting. Register to catch the rest of Coalesce, the Analytics Engineering Conference, at https://coalesce.getdbt.com. The Analytics Engineering Podcast is brought to you by dbt L...

Dec 07, 2021•30 min•Season 1Ep. 15

[COALESCE] How big is this wave? Ft. Martin Casado of a16z

The modern data stack is the third generation of data analysis products to come to prominence since the 90's. The prior waves—data warehouse appliances and then Hadoop—were both big steps forwards but ultimately failed to live up to their initial promise. Is the modern data stack just another iteration in a long string of “trendy technologies” in data––waves that crash upon the shore but ultimately recede? Or is it somehow more permanent? Register to catch the rest of Coalesce, the Analytics Eng...

Dec 07, 2021•45 min•Season 1Ep. 14

← Prev Next →