Data Brew by Databricks - podcast cover

Data Brew by Databricks

Databricksdatabricks.com
Welcome to Data Brew by Databricks with Denny and Brooke! In this series, we explore various topics in the data and AI community and interview subject matter experts in data engineering/data science. So join us with your morning brew in hand and get ready to dive deep into data + AI! For this first season, we will be focusing on lakehouses – combining the key features of data warehouses, such as ACID transactions, with the scalability of data lakes, directly against low-cost object stores.

Episodes

Retrieval, rerankers, and RAG tips and tricks | Data Brew | Episode 39

In this episode, Andrew Drozdov, Research Scientist at Databricks, explores how Retrieval Augmented Generation (RAG) enhances AI models by integrating retrieval capabilities for improved response accuracy and relevance. Highlights include: - Addressing LLM limitations by injecting relevant external information. - Optimizing document chunking, embedding, and query generation for RAG. - Improving retrieval systems with embeddings and fine-tuning techniques. - Enhancing search results using re-rank...

Feb 20, 202545 min

The Power of Synthetic Data | Data Brew | Episode 38

In this episode, Yev Meyer, Chief Scientist at Gretel AI, explores how synthetic data transforms AI and ML by improving data access, quality, privacy, and model training. Highlights include: - Leveraging synthetic data to overcome AI data limitations. - Enhancing model training while mitigating ethical and privacy risks. - Exploring the intersection of computational neuroscience and AI workflows. - Addressing licensing and legal considerations in synthetic data usage. - Unlocking private dataset...

Feb 04, 202542 min

Secret to Production AI: Tools & Infrastructure | Data Brew | Episode 37

In this episode, Julia Neagu, CEO & co-founder of Quotient AI, explores the challenges of deploying Generative AI and LLMs, focusing on model evaluation, human-in-the-loop systems, and iterative development. Highlights include: - Merging reinforcement learning and unsupervised learning for real-time AI optimization. - Reducing bias in machine learning with fairness and ethical considerations. - Lessons from large-scale AI deployments on scalability and feedback loops. - Automating workflows ...

Jan 22, 202537 min

Mixture of Memory Experts (MoME) | Data Brew | Episode 36

In this episode, Sharon Zhou, Co-Founder and CEO of Lamini AI, shares her expertise in the world of AI, focusing on fine-tuning models for improved performance and reliability. Highlights include: - The integration of determinism and probabilism for handling unstructured data and user queries effectively. - Proprietary techniques like memory tuning and robust evaluation frameworks to mitigate model inaccuracies and hallucinations. - Lessons learned from deploying AI applications, including insig...

Jan 10, 202541 min

Mixed Attention & LLM Context | Data Brew | Episode 35

In this episode, Shashank Rajput, Research Scientist at Mosaic and Databricks, explores innovative approaches in large language models (LLMs), with a focus on Retrieval Augmented Generation (RAG) and its impact on improving efficiency and reducing operational costs. Highlights include: - How RAG enhances LLM accuracy by incorporating relevant external documents. - The evolution of attention mechanisms, including mixed attention strategies. - Practical applications of Mamba architectures and thei...

Nov 21, 202439 min

Kumo AI & Relational Deep Learning | Data Brew | Episode 34

In this episode, Jure Leskovec, Co-founder of Kumo AI and Professor of Computer Science at Stanford University, discusses Relational Deep Learning (RDL) and its role in automating feature engineering. Highlights include: - How RDL enhances predictive modeling. - Applications in fraud detection and recommendation systems. - The use of graph neural networks to simplify complex data structures.

Oct 14, 202443 minSeason 6Ep. 28

LLMs: Internals, Hallucinations, and Applications | Data Brew | Episode 33

Our fifth season dives into large language models (LLMs), from understanding the internals to the risks of using them and everything in between. While we're at it, we'll be enjoying our morning brew. In this session, we interviewed Chengyin Eng (Senior Data Scientist, Databricks), Sam Raymond (Senior Data Scientist, Databricks), and Joseph Bradley (Lead Production Specialist - ML, Databricks) on the best practices around LLM use cases, prompt engineering, and how to adapt MLOps for LLM...

Jul 21, 202339 minSeason 5Ep. 4

Demonstrate–Search–Predict Framework | Data Brew | Episode 32

We will dive into LLMs for our fifth season, from understanding the internals to the risks of using them and everything in between. While we’re at it, we’ll be enjoying our morning brew. In this session, we interviewed Omar Khattab - Computer Science Ph.D. Student at Stanford, creator of DSP (Demonstrate–Search–Predict Framework), to discuss DSP, common applications, and the future of NLP.

Jun 29, 202333 minSeason 5Ep. 3

Generative AI Risks | Data Brew | Episode 31

We will dive into LLMs for our fifth season, from understanding the internals to the risks of using them and everything in between. While we’re at it, we’ll be enjoying our morning brew. In this session, we interviewed Yaron Singer, CEO of Robust Intelligence, Professor of Computer Science at Harvard University, and guest of Data Brew Season 3 (our first repeat guest!). In this session, we discuss generative AI, the trends toward embracing LLMs, and how the surface area for vulnerabilities in ge...

Jun 08, 202335 minSeason 5Ep. 2

John Snow Labs & SparkNLP | Data Brew | Episode 30

We are back and we will dive into LLMs from understanding the internals to the risks of using them and everything in between. While we’re at it, we’ll be enjoying our morning brew. In this session, we interviewed David Talby who is the CTO at John Snow Labs; they help healthcare & life science companies put AI to good use. David's interests include natural language processing, applied artificial intelligence in healthcare, and responsible AI.

Jun 01, 202343 minSeason 5Ep. 1

Data Brew Season 4 Episode 6: Professional Athletes

For our fourth season, we focus on connected health and how data & AI augment and improve our daily health. While we’re at it, we’ll be enjoying our morning brew. Shayna Powless and Eli Ankou, professional cyclist for L39ion of Los Angeles and defensive tackle for the Buffalo Bills, respectively, provide valuable insight on how professional athletes leverage data to improve their performance and how they combine their passion for sports with the Dreamcatcher Foundation. See more at databrick...

Jun 09, 202236 minSeason 4Ep. 6

Data Brew Season 4 Episode 5: Public Health: Education, Access, and Policy

For our fourth season, we focus on connected health and how data & AI augment and improve our daily health. While we’re at it, we’ll be enjoying our morning brew. Matt Willis, Marin County Public Health Officer, shares the three pillars of public health: education, access, and policy, and the critical role data plays in addressing the COVID-19 pandemic & opioid epidemic. See more at databricks.com/data-brew...

May 05, 202235 minSeason 4Ep. 5

Data Brew Season 4 Episode 4: 1283 Days of Running (and Counting)

For our fourth season, we focus on connected health and how data & AI augment and improve our daily health. While we’re at it, we’ll be enjoying our morning brew. Running the length of the US every year, Alexandra Matthiesen shares her motivational secrets for running 1,283 consecutive days (and counting!) and redefining physical and mental limits. See more at databricks.com/data-brew

Apr 14, 202236 minSeason 4Ep. 4

Data Brew Season 4 Episode 3: Last Man Standing

For our fourth season, we focus on connected health and how data & AI augment and improve our daily health. While we’re at it, we’ll be enjoying our morning brew. Winner of the infamous Last Man Standing race (running 246 miles in 59 hours), Guillaume merges the world of competitive long-distance running with data science to push the boundaries of body and mind. See more at databricks.com/data-brew

Mar 31, 202241 minSeason 4Ep. 3

Data Brew Season 4 Episode 2: NBA Analytics

For our fourth season, we focus on connected health and how data & AI augment and improve our daily health. While we’re at it, we’ll be enjoying our morning brew. Alexander Powell chronicles the evolution of sports analytics and how professional sports teams use data as a competitive advantage. See more at databricks.com/data-brew

Mar 10, 202230 minSeason 4Ep. 2

Data Brew Season 4 Episode 1: Reducing Injury & Increasing Retention of Industrial Athletes

For our fourth season, we focus on connected health and how data & AI augment and improve our daily health. While we’re at it, we’ll be enjoying our morning brew. Globally, 38,000 people get hurt on the job every hour. In the United States alone, over $250 billion dollars is spent on workplace injury annually. Sean Petterson, founder and CEO of StrongArm Tech, discusses the role of wearable devices to reduce workplace injury and increase retention of industrial athletes. See more at databric...

Feb 24, 202234 minSeason 4Ep. 1

Data Brew Season 3 Episode 6: Open Source

For our third season, we focus on how leaders use data for change. Whether it’s building data teams or using data as a constructive catalyst, we interview subject matter experts from industry to dive deeper into these topics. For our season 3 finale, Nithya Ruff discusses the open-source ecosystem, ways to contribute to open-source projects (hint: it’s not just about the code), and how businesses can balance community and company interests. With 95% of open-source contributions coming from men, ...

Oct 28, 202134 minSeason 3Ep. 6

Data Brew Season 3 Episode 5: Sustainability & Sake

For our third season, we focus on how leaders use data for change. Whether it’s building data teams or using data as a constructive catalyst, we interview subject matter experts from industry to dive deeper into these topics. We interview Junta Nakai in our most unique location yet - Brooklyn Kura - the first non-Japanese sake distillery in New York. In this episode, Junta shares the philosophical, economic, and tactical approaches to sustainability and ESG, as well as the secrets to brewing sak...

Oct 14, 202132 minSeason 3Ep. 5

Data Brew Season 3 Episode 4: Executive Education

For our third season, we focus on how leaders use data for change. Whether it’s building data teams or using data as a constructive catalyst, we interview subject matter experts from industry to dive deeper into these topics. Did you know that the average tenure of a board member is longer than the average tenure of a marriage in the United States? In this episode, Coco Brown discusses the benefits and drawbacks of the long tenures of corporate boards, their current structure, the impact of rece...

Oct 07, 202139 minSeason 3Ep. 4

Data Brew Season 3 Episode 3: 3 T’s to Securing AI Systems: Tests, tests, and more tests

For our third season, we focus on how leaders use data for change. Whether it’s building data teams or using data as a constructive catalyst, we interview subject matter experts from industry to dive deeper into these topics. What does it mean to make your machine learning system “production-ready”? Yaron Singer walks us through the infrastructure, testing procedures, and more that help make ML systems ready for the real world in this episode of Data Brew. See more at databricks.com/data-brew...

Sep 30, 202135 minSeason 3Ep. 3

Data Brew Season 3 Episode 2: Data Culture Outside ‘The Valley’

For our third season, we focus on how leaders use data for change. Whether it’s building data teams or using data as a constructive catalyst, we interview subject matter experts from industry to dive deeper into these topics. Have you ever had a spam call automatically blocked for you? You can thank First Orion for that - in one day they blocked or scam tagged over 108 million calls - just on T-Mobile alone! In this episode, we have the pleasure to chat with Charles Morgan and Kent Welch, CEO an...

Sep 23, 202136 minSeason 3Ep. 2

Data Brew Season 3 Episode 1: Disrupt: Challenge your Business Assumptions

For our third season, we focus on how leaders use data for change. Whether it’s building data teams or using data as a constructive catalyst, we interview subject matter experts from industry to dive deeper into these topics. In this season opener, Elena Donio shares her experience using data and domain knowledge to disrupt the traditional service and sales compensation model. She also discusses how to build companies that scale, manage corporate cultural evolution, and the influence of corporat...

Sep 16, 202130 minSeason 3Ep. 1

Data Brew Season 2 Episode 9: Data Driven Software

For our second season of Data Brew, we will be focusing on machine learning, from research to production. We will interview folks in academia and industry to discuss topics such as data ethics, production-grade infrastructure for ML, hyperparameter tuning, AutoML, and many more. We branch, version, and test our code, but what if we treated data like code? Tim Hunter joins us to discuss the open-source Data-Driven Software (DDS) package and how it leads to immense gains in collaboration and decre...

Jul 21, 202131 minSeason 2Ep. 9

Data Brew Season 2 Episode 8: Feature Engineering

For our second season of Data Brew, we will be focusing on machine learning, from research to production. We will interview folks in academia and industry to discuss topics such as data ethics, production-grade infrastructure for ML, hyperparameter tuning, AutoML, and many more. Is there ever a “one-size fits all” approach for feature engineering? Find out this and more with Amanda Casari and Alice Zheng, co-authors of the Feature Engineering for Machine Learning book. See more at databricks.com...

Jul 09, 202131 minSeason 2Ep. 8

Data Brew Season 2 Episode 7: Interpretable Machine Learning

For our second season of Data Brew, we will be focusing on machine learning, from research to production. We will interview folks in academia and industry to discuss topics such as data ethics, production-grade infrastructure for ML, hyperparameter tuning, AutoML, and many more. What does it mean for a model to be “interpretable”? Ameet Talwalkar shares his thoughts on IML (Interpretable Machine Learning), how it relates to data privacy and fairness, and his research in this field. See more at d...

Jul 01, 202137 minSeason 2Ep. 7

Data Brew Season 2 Episode 6: AutoML

For our second season of Data Brew, we will be focusing on machine learning, from research to production. We will interview folks in academia and industry to discuss topics such as data ethics, production-grade infrastructure for ML, hyperparameter tuning, AutoML, and many more. Erin LeDell shares valuable insight on AutoML, what problems are best solved by it, its current limitations, and her thoughts on the future of AutoML. We also discuss founding and growing the Women in Machine Learning an...

Jun 17, 202136 minSeason 2Ep. 6

Data Brew Season 2 Episode 5: ML Applications

For our second season of Data Brew, we will be focusing on machine learning, from research to production. We will interview folks in academia and industry to discuss topics such as data ethics, production-grade infrastructure for ML, hyperparameter tuning, AutoML, and many more. Good machine learning starts with high quality data. Irina Malkova shares her experience managing and ensuring high-fidelity data, developing custom metrics to satisfy business needs, and discusses how to improve interna...

Jun 10, 202133 minSeason 2Ep. 5

Data Brew Season 2 Episode 4: Hyperparameter and Neural Architecture Search

For our second season of Data Brew, we will be focusing on machine learning, from research to production. We will interview folks in academia and industry to discuss topics such as data ethics, production-grade infrastructure for ML, hyperparameter tuning, AutoML, and many more. Liam Li is a leading researcher in the fields of hyperparameter optimization and neural architecture search, and is the author of the seminal Hyperband paper. In this session, Liam discusses the evolution of hyperparamet...

May 13, 202133 minSeason 2Ep. 4

Data Brew Season 2 Episode 3: Infrastructure for ML

For our second season of Data Brew, we will be focusing on machine learning, from research to production. We will interview folks in academia and industry to discuss topics such as data ethics, production-grade infrastructure for ML, hyperparameter tuning, AutoML, and many more. Adam Oliner discusses how to design your infrastructure to support ML, from integration tests to glue code, the importance of iteration, and centralized vs decentralized data science teams. He provides valuable advice fo...

May 05, 202131 minSeason 2Ep. 3

Data Brew Season 2 Episode 2: Data Ethics

For our second season of Data Brew, we will be focusing on machine learning, from research to production. We will interview folks in academia and industry to discuss topics such as data ethics, production-grade infrastructure for ML, hyperparameter tuning, AutoML, and many more. Have you ever wondered how your purchasing behavior may reveal protected attributes? Or how data scientists and business play a role in combating bias? We discuss with Diana Pfeil recommendations to reduce bias and impro...

Apr 28, 202126 minSeason 2Ep. 2