Welcome to BuffConf!
A two-day technical conference for video and infrastructure engineers.
Connect to WiFi
SSID: MOPOP_Events
Password: MoreFun!

Remember you can use your badge for free access to the entire museum!
Day 1 Agenda
8:00 am - 9:00 am
Breakfast

Add a Title
Add a Title

Add a Title
Add a Title

Add a Title
Add a Title

Add a Title
Add a Title
9:00 am - 9:20 am
104 Matches. Five Billion Viewers. Zero Margin for Error.

Lionel Bringuier
GM Media, Entertainment & Gaming, Momento

Add a Title
Add a Title

Add a Title
Add a Title

Add a Title
Add a Title
The 2026 World Cup shows unprecedented numbers. At 3 am London time, 9.1 million British fans stayed up to watch England beat Mexico, the largest UK audience ever recorded in that timeslot, while BBC iPlayer logged 48 million requests for its biggest day in history. From the other side of the pond, FOX drew 30 million viewers for USA–Belgium, at the time the most-watched soccer telecast in American history. Netflix's Tyson fight peaked at 65 million concurrent streams. Cricket on JioHotstar just hit 72.5 million. But the average global viewership for a FIFA match is 175 million. The 2022 final averaged 571 million global viewers. So, the question is: was this FIFA World Cup the largest event ever streamed? The answer depends entirely on what is measured, whether that's concurrents, average minute audience, requests, or reach. But it mostly reveals about where live video infrastructure is headed, and why the hardest engineering problems in our industry now happen at 3 am.
9:20 am - 9:50 am
From Pixels to Decisions: Real-Time Video Intelligence at Broadcast Scale

Manish Rao
General Manager, AWS Elemental, Amazon Web Services

Add a Title
Add a Title

Add a Title
Add a Title

Add a Title
Add a Title
Audiences are engaging with video across more platforms, formats, and moments than any production team can manually serve — vertical content for social, highlights for engagement, subtitles for accessibility. Until recently, applying AI to live video at broadcast latency was impractical: the computational demands of modern ML models conflicted with the strict real-time constraints of live encoding pipelines. Learn how AWS Elemental has removed that constraint, what becomes possible when content intelligence operates in parallel with live encoding rather than as a post-production step, and how broadcasters are using structured video metadata to drive engagement, unlock new revenue streams, and reach audiences across every screen in real time.
9:50 am - 10:05 am
Simulating and Mitigating Token Abuse in High-Volume Video Pools

Zac Shenker
Senior Director, Media Strategy, Fastly

Add a Title
Add a Title

Add a Title
Add a Title

Add a Title
Add a Title
Validating anti-piracy mechanisms under realistic live conditions is challenging due to the complexity of simulating malicious subscriber pools at scale. Traditional piracy detection relies on evaluating raw CDN logs, introducing data egress costs and delaying mitigation until after the event.
This session demonstrates an edge-native approach to identifying token abuse in real time. We will show how to design a simulation framework to create pools of valid and pirate users. Then, utilizing edge-level token parsing and time-limited window tracking, we will flag bad actors with high confidence. You'll learn how to design the simulation framework and corresponding orchestration layer with real-time alerts, protecting stream monetization and avoiding overloading the operations center.
10:05 am - 10:25 am
Beyond the Feed, Building CBS Sports AI-Powered "For You Page" for Modern Fandom

Edwin Rivera
Principal Architect, CBS Interactive

Add a Title
Add a Title

Add a Title
Add a Title

Add a Title
Add a Title
As sports streaming environments become increasingly saturated, static editorial curation is no longer enough to capture and sustain fan attention. To drive meaningful user retention and deeper daily engagement, the CBS Sports app successfully integrated a highly responsive, AI-powered personalized recommendation engine alongside its traditional editorial coverage.
This session breaks down the technical and architectural realities of designing and launching the new CBS Sports "For You Page" feature. We will explore how to blend algorithmic machine learning models with human-curated editorial content in real time, using embedding models to represent content as vectors. Finally, we will cover the infrastructure required to generate and process these AI embeddings against hyper-local user behavior signals, and the ultimate impact this personalized shift has had on customer retention and long-term app engagement.
10:25 am - 10:50 am
Morning Break

Add a Title
Add a Title

Add a Title
Add a Title

Add a Title
Add a Title

Add a Title
Add a Title
10:50 am - 11:20 am
How Netflix Uses Versioned Caching Paradigms to Drive User Experience

Sunjeet Singh
Staff Software Engineer, Netflix

Senior Software Engineer, Netflix
William Schor

Add a Title
Add a Title

Add a Title
Add a Title
Delivering a seamless playback experience requires a caching footprint that serves different workloads and access patterns using the best-matching operational model. While traditional demand-filled caches remain a popular pattern, Netflix increasingly relies on versioned datasets, snapshot-published with periodic refreshes, to power the streaming control plane.
This session breaks down four distinct versioned dataset architectures deployed at scale. We will examine how they vary across dimensions of size, read latency, staleness tolerance, and cost, ranging from small in-memory datasets written in near-real-time to distributed petabyte-scale datasets published from offline jobs. Attendees will learn how matching the right caching architecture to specific media workloads (from asset manifests to user state) sidesteps classic failure modes of demand-filled caches such as stampedes and cold-start misses while meeting strict read-performance KPIs.
11:20 am - 11:40 am
The Pixel Economy: Building generative models for taste, speed, and scale

Mathieu Tuli
Machine Learning Engineer, Meta

Add a Title
Add a Title

Add a Title
Add a Title

Add a Title
Add a Title
Building AI features for creators means chasing a moving target where quality and compute are in constant contention.
Training and deploying models is expensive and higher quality usually means bigger models; so how do you iterate quickly and build highly aesthetic, low compute models?
Taking it a step further, how do you build creation features that are aligned with what creators want to make and also what users want to see, and do so on demand? In this talk, we'll peek into how Meta is thinking about this, how we navigate these tradeoffs, how we train aesthetic models and adapt to virality, and where we're going as AI demand grows.
11:40 am - 12:00 pm
MOQ Around and Find Out: Leaving Legacy Latency Behind

Ali C. Begen
Prof., Ozyegin University

Add a Title
Add a Title

Add a Title
Add a Title

Add a Title
Add a Title
For years, the streaming industry has been trapped in a toxic relationship with legacy, segment-based protocols. We've twisted HLS and DASH into every low-latency shape imaginable, only to hit the inherent architectural limits. It's time to move on.
Media over QUIC (MOQ) promises to change the game, offering ultra-low latency without sacrificing the scalability of modern CDN architectures. But what happens when you actually stop looking at the IETF drafts and start writing code?
In this talk, we'll dive deep into the reality of building a real-time streaming pipeline using MOQ and MOQtail (our open-source implementation). We'll skip the introductory fluff and get straight into the engineering trenches.
12:00 pm - 1:00 pm
Lunch

Add a Title
Add a Title

Add a Title
Add a Title

Add a Title
Add a Title

Add a Title
Add a Title
1:00 pm - 2:00 pm
Optimizing Playback SDKs, Client Architecture, and Device Capabilities for Next-Gen Experiences

Jonathan Colwell
Senior Director, Video Engineering, Fox Corporation

Product Director, Red Bull Media House
Joshua Lamb

Stephanie (Schuck) Lone
Global SA Leader, Media, Entertainment, Gaming & Sports, Amazon Web Services

David Van der Voort
Sr Manager, Connected TV, International Engineering, Paramount
Supporting a massive breadth of legacy, modern, and even non-traditional devices while delivering resource-heavy features has made the client application layer a critical engineering bottleneck. Traditional implementations routinely fail under the constraints of embedded hardware, rendering pipelines, and concurrent stream demands.
Moderated by Steph Lone (Global Leader, Solutions Architecture M&E and Sports at AWS), this panel explores the client-side architecture required to deploy premium playback experiences at scale. David Van der Voort (Senior Manager, Connected TV, International Engineering, Paramount+) breaks down the technical realities of building a converged, multi-tenant application with Paramount Lite+, while Jonathan Colwell (Director of Engineering, FOX) details the resource management and stream synchronization strategies required to power dynamic Multiview for the World Cup. Finally, Josh Lamb (Product Director, Red Bull Media House) discusses how to squeeze maximum performance out of diverse playback devices and pushing streaming standards to their limits to unlock cutting-edge features.
2:00 pm - 2:15 pm
Driving the AI Video Enhancement Flywheel

Qi Cai
Research Scientist, Meta

Add a Title
Add a Title

Add a Title
Add a Title

Add a Title
Add a Title
Real-time quality enhancement and metadata parsing do not happen in isolation. This session explores how to scale video quality metrics, enhancements, and metadata processing simultaneously. We will break down how we build a unified data flywheel and use that loop to continuously measure and improve operational performance at scale.
2:15 pm - 2:35 pm
From Bits to Tokens: Applying Video Compression Intelligence to Video AI Token Efficiency

Zoe Liu
Co-Founder & CEO, Visionular

Add a Title
Add a Title

Add a Title
Add a Title

Add a Title
Add a Title
Video is token-hungry by nature, directly driving inference cost across compute, memory, and latency at scale. One minute of video at 30fps could represent 1.8 million words, yet the vast majority is spatially, temporally, or perceptually redundant. We argue that AI video token compression shares a profound connection with conventional video compression engineering. RD optimization, motion-compensated prediction, perceptual salience modeling, and ROI encoding all have direct analogs in efficient video tokenization. Further combined with lightweight domain-specialized pre-analyzers and vertical-specific priors, compression-domain thinking offers a practical path to making large-scale multimodal video inference computationally tractable.
2:35 pm - 2:50 pm
Square Pegs and Round Holes: WebRTC at Broadcast Scale

Kyle Lee
Director of Engineering, Amazon Web Services

Sr Engineering Manager, Amazon Web Services
Alexander Allen

Add a Title
Add a Title

Add a Title
Add a Title
It is a common misconception that WebRTC cannot scale beyond a few hundred (or thousand) concurrent viewers. What is inarguably true is that thanks to its peer-to-peer roots, scaling WebRTC takes special care and effort: because it was never designed with CDN delivery in mind, each WebRTC distribution system is ultimately a customized implementation with its own quirks and tradeoffs.
Learn how Amazon IVS built a WebRTC CDN capable of serving 1MM+ concurrent viewers on a single broadcast with unit economics similar to HLS/DASH. We’ll talk about lessons we learned along the way and core investments we’ve made in improving time-to-video (TTV), quality of service (QoS), and cost. Lastly, we’ll talk about the future, including intentional divergence from the WebRTC standard, first class multicodec and transcoding support, and forays into MoQ (Media over QUIC).
2:50 pm - 3:20 pm
From Zero to Everywhere: Engineering a FAST Linear Streaming Platform at Scale

Spencer Shanson
SVP Architecture, Paramount

Add a Title
Add a Title

Add a Title
Add a Title

Add a Title
Add a Title
Building a Free Ad-Supported Streaming TV platform is deceptively simple at the start — spin up some channels, stitch together a playlist, serve a stream. The hard problems emerge when you're operating thousands of channels, a library of millions of VOD assets, live events that can't fail, and a platform deployed across multiple geographic regions. This talk traces the full arc of that journey: the early architectural decisions that either buy you flexibility or create ceilings you'll hit painfully later, the inflection points where systems that worked at hundreds of channels collapse at thousands, and the operational realities of keeping linear streams reliable at a scale most engineers never encounter.
We'll cover the core engineering challenges — playlist assembly and scheduling at scale, ad insertion complexity, content ingestion pipelines that have to absorb millions of assets without falling over, live event integration into a fundamentally pre-recorded medium, and the regional distribution architecture that makes global delivery economically viable. The talk is grounded in real production experience and designed for engineers and architects who are either building a FAST platform or scaling one that's already outgrown its original design.
3:20 pm - 3:35 pm
Engineering Real-Time Dialogue Isolation for Live Broadcast

Abhay Nadkarni
Technical GTM & AI Audio Deployment, AudioShake

Add a Title
Add a Title

Add a Title
Add a Title

Add a Title
Add a Title
Traditional live audio cleanup relies on noise suppression, which attenuates non-speech frequencies within a mix. This approach sacrifices detail and never yields a truly clean signal. True dialogue isolation takes a different approach: instead of suppressing the mix, it extracts dialogue as an independent stem. Historically, this has only been viable in post-production, because the context a separation model needs for high-quality output is fundamentally at odds with the latency budget of a live feed.
This session covers the engineering required to close that gap and run true dialogue isolation inline with a live broadcast at 11ms. We start with the central trade-off in source separation: better separation requires more signal context, and real-time constraints force hard decisions about how much context you can use. We then break down the latency budget across the broadcast signal chain, examining where time actually accumulates, and why broadcast-grade latency is a combined problem of fast model architecture and sufficient compute. Finally, we look at the downstream payoffs across viewer engagement, transcription and captioning, and compliance.
3:35 pm - 3:55 pm
Afternoon Break

Add a Title
Add a Title

Add a Title
Add a Title

Add a Title
Add a Title

Add a Title
Add a Title
3:55 pm - 4:25 pm
Before AI Can Help, You Have to See Everything: Building the Observability Foundation for Intelligent Streaming Operations

Bradley Burciaga
Software Architect, Paramount

Add a Title
Add a Title

Add a Title
Add a Title

Add a Title
Add a Title
Large-scale video streaming platforms are notoriously difficult to operate — fragmented telemetry, siloed systems, and incident investigations that demand engineers manually correlate data across dozens of services before they can even begin to diagnose a problem.
This talk makes the case that observability isn't just an operational nicety; it's the prerequisite for applying AI to streaming infrastructure. Drawing on a real-world audit of a major streaming platform's video pipeline, we'll walk through how implementing clean, standardized telemetry with OpenTelemetry creates a unified data foundation that unlocks the next generation of intelligent operations — AI-driven root cause analysis, autonomous remediation of common failures, and continuous optimization of pipeline performance, reliability, and cost. We'll cover the engineering trade-offs of instrumenting at scale, what "good" telemetry actually looks like in practice, and why telemetry quality is the critical gating factor that determines whether AI delivers meaningful operational value or just expensive noise.
4:25 pm - 4:40 pm
What Real-World MoQ Deployments Teach Us About Next-Gen Live Media

Luke Curley
CEO @ moq.dev, moq.dev

Add a Title
Add a Title

Add a Title
Add a Title

Add a Title
Add a Title
Media over QUIC (MoQ) is no longer a theoretical next-gen protocol; it is active in production, proving its worth under some of the most unforgiving use cases. Today, MoQ is being used to remotely fly drones, drive boats, fold t-shirts in real-time robotics labs, and is soon set to pilot autonomous vehicles. But what do these telemetry and control use cases mean for the future of traditional media? We will explore how its unique sub-second transport mechanism solves massive pain points in remote live production (REMI), camera-to-cloud workflows, and interactive fan engagement at scale.
4:40 pm - 5:10 pm
Scaling the Screen With Content Protection, Intelligent Operations, and Rapid AI Feature Velocity

Daniela Miao
Co-Founder & CTO, Momento

Add a Title
Add a Title

Add a Title
Add a Title

Add a Title
Add a Title
Engineering and product leaders have always been faced with the challenge of balancing content protection, increasing operational awareness, and accelerating the pace of user experience innovation. In this session, we'll explore ways of solving these challenges at scale and what we learned doing so.
We will detail how next-generation content protection is evolving, showcasing how real-time concurrency tracking acts as a critical line of defense against piracy and unauthorized credential sharing. We'll dive into the latest in operational dashboards and using under-the-hood data to enable natural language inquiries for bespoke business intelligence. Finally, we will explore how AI can be used to not only create personalized user experiences but also increase feature velocity.
4:45 pm - 6:45 pm
Happy hour

Add a Title
Add a Title

Add a Title
Add a Title

Add a Title
Add a Title

Add a Title
Add a Title
Day 2 Agenda
7:30 am - 8:30 am
Breakfast

Add a Title
Add a Title

Add a Title
Add a Title
8:30 am - 9:00 am
Optimizing, Accelerating, and Augmenting Video with NVIDIA

Jeff Kember
Senior Director, Product Management AI, NVIDIA

Rick Champagne
Director, Global M&E Strategy and Marketing, NVIDIA
NVIDIA accelerated computing and AI are improving how video is processed, delivered and experienced today. Join NVIDIA to learn about the latest technologies for increasing encoding efficiency, enhancing video quality, supporting real-time streaming and adding intelligent capabilities to media workflows. The session will connect the latest NVIDIA advances with practical opportunities across broadcast, streaming and video infrastructure.
9:00 am - 9:15 am
Content Steering - a Unified Standard for Multi-CDN Streaming Across HLS and DASH

Yuriy Reznik
Chair of DASH-IF, P&P, & E&P working groups, SVTA

Add a Title
Add a Title
Multi‑CDN delivery is now essential for large‑scale streaming, yet past solutions were proprietary, format‑specific, and difficult to deploy. HLS/DASH Content Steering introduces a unified, standards‑based approach for interoperable multi‑CDN control. This talk reviews the standard's origins, core mechanisms, and benefits, surveys current client and server support, and highlights ongoing work by HLS-interest, IETF, DASH‑IF, and SVTA groups. It concludes with SVTA study results showing clear QoE gains from content steering.
9:15 am - 9:30 am
Architecting Media Infrastructure for MCP and Agentic AI

Igor Oreper
Chief Architect and Chief Strategy Officer , Bitmovin

Add a Title
Add a Title
Model Context Protocol (MCP) is moving agentic AI in media from demo to production, shifting us from fixed, hard-coded API endpoints to context-aware orchestration. This session covers the technical realities of building MCP servers for media infrastructure and using them as a single agentic interface across the full video lifecycle - encoding, playback, automated QA and device testing, and observability - so an agent can drive the whole stack in natural language.
We break down real examples, from configuring encodes and players to provisioning cross-device test beds and interpreting playback telemetry, and why none of it works without an observability foundation underneath. We close on the prerequisites a media company needs before agents can act: structured stream data, first-party session telemetry, and quality signaling that connects frontend playback to backend systems, so LLMs can interface with production data safely and effectively.
9:30 am - 9:45 am
Lessons from delivering the FIFA Worldcup Buffer Free

Megha Kande
Senior Manager - Product Management, Amazon CloudFront, Amazon Web Services

Phil Harrison
Sr Product Manager, AWS Elemental, Amazon Web Services
The 2026 FIFA World Cup exposed every assumption in LL-HLS and ad-insertion pipelines at scale. This talk breaks down what happens when server guided ad insertion and LL-HLS operate together at peak traffic, and shares design patterns for configuring TTLs to minimize origin latency, distributing channels across regions for resilience, and prefetching ad-decisions for increased monetization.
9:45 am - 10:00 am
Beyond the Optimal Split: Operating a Multi-CDN Control Loop

Tom Howe
Director of Insights, Hydrolix

Add a Title
Add a Title
Automatic traffic steering stays feasible under load only when the controller plans over a look-ahead horizon instead of reacting interval-by-interval and because that decision is a microsecond LP, you can forecast centrally and run the optimizer at the edge.
10:00 am - 10:15 am
Morning Break

Add a Title
Add a Title

Add a Title
Add a Title
10:15 am - 12:00 pm
AI Code Boot Camp for Media Leaders

Khawaja Shams
Co-Founder & CEO, Momento

Add a Title
Add a Title
Engineering lifecycles for media applications, spanning streaming infrastructure, back-office operations, and real-time personalization, are accelerating at an unprecedented pace. For leadership, understanding the technology driving this velocity is no longer optional; seeing exactly how these tools anchor into traditional M&E workloads is key to unlocking team insight and protecting operational efficiency.
Designed specifically for media technology executives, CTOs, and product leaders, this intensive 2-hour workshop invites you to step out of the meeting room and into an active, AI-native development environment.
This boot camp completely demystifies the mechanics of agentic AI coding. Rather than reviewing conceptual slides, you will spend high-quality, hands-on time executing actual orchestration and development tasks alongside experts focused entirely on workflows relevant to your media business. Executives will walk away with a practical framework to effectively direct, audit, and maximize the ROI of AI-driven engineering teams, providing the blueprint needed to quickly evaluate new ideas and compress product roadmaps from quarters to weeks.
12:00 pm - 1:00 pm
Lunch

Add a Title
Add a Title
