Going In Deep On Data | YC Paper Club
Score and analysis
Summary
Francois Chaubard introduces a YC Paper Club session focused on data, featuring talks by Vincent Sunn Chen, Volo Kuleshov, and Shayne Longpre.
Vincent Sunn Chen discusses scaling expert supervision and benchmarking AI agents, while Volo Kuleshov explains diffusion language models and their production applications.
Shayne Longpre presents research on multilingual pre-training and practical scaling laws for diverse language models.
Breakdown
Vincent Sunn Chen discusses scaling expert supervision by detailing Snorkel's 'data programming' approach and the Senior SWE-bench benchmark, highlighting real-world applications in coding agents.
Volo Kuleshov presents Inception's diffusion language models, explaining their speed advantage and use in real-time voice applications.
The speakers present detailed methodologies and research findings, such as Vincent Sunn Chen's explanation of weak supervision techniques and Shayne Longpre's empirical analysis of cross-lingual transfer.
Volo Kuleshov backs claims about diffusion models with performance metrics on benchmarks like T-bench.
Francois Chaubard effectively introduces the topic of data's importance, setting a clear agenda for the speakers.
Vincent Sunn Chen, Volo Kuleshov, and Shayne Longpre deliver their presentations with clear articulation and a logical flow, using specific examples and data points to maintain audience interest throughout the long session.
The session covers a broad spectrum of data-related topics, from the fundamental importance of data in AI to advanced concepts like diffusion models and multilingual pre-training.
Actionable insights are provided, such as Shayne Longpre's guidance on mixing data sources and Volo Kuleshov's mention of available models and startup programs.
Vincent Sunn Chen, a founding team member at Snorkel, explains the evolution of data supervision from Data 1. 0 to Data 2. 0, using analogies like 'data programming' and 'labeling functions' to clarify complex concepts.
Shayne Longpre, an MIT PhD student, details multilingual pre-training by comparing language data availability and transfer learning.
The video features a speaker presenting slides in front of an audience.
The lighting is adequate for visibility, and the camera work is steady, capturing the speaker and screen without significant distraction.
The overall production quality is functional for an educational talk.
Click any metric to see why it scored this way
Use VidScore on YouTube
Explore scores and profiles while you browse.
Get the ExtensionRelated videos
More videos we scored on the same topics.

Why Netflix is betting on systems thinkers—not specialists—in the AI era | Elizabeth Stone (CPTO)
Lenny's Podcast
Here's what an AI-first phone might look like | The Vergecast
The Verge
You can't ignore Google Zero anymore | The Vergecast
The Verge
The AI Workflow That Puts You in the Top 1% | Practical Steps to Level Up
Silicon Valley Girl
The Fight Over Open Source AI, Anthropic's $1.5B Payout, NYC Socialists: Evictions = Violence?
All-In Podcast
Google’s AI Brain Drain, SpaceX's Huge Quarter, Airtable’s 90% Collapse, US Data Fuels China AI
All-In Podcast
AI and the Enshittification Era w/ Cory Doctorow | The Weekly Show with Jon Stewart
The Weekly Show with Jon Stewart
The AI content machine that turns ideas into posts that don't sound like slop | Alex Lieberman
How I AI
This AI System Runs My Company for $20/month
Alex Lieberman
Tech MELTDOWN After AI ESCAPE and HACK
Breaking Points
OpenAI Shocks The World With GENIE... Almost Unlimited AI Power
AI Revolution