I am an AI Researcher at Descript, working on generative models that power audio and video editing in production. My current focus spans audio codecs, diffusion models and audio-text forced alignment. A full list of papers is on Google Scholar.

Applied Research

Audio Regenerate

Seamless word-level audio editing via latent inpainting.

Audio in
Latent frames
Text in "… new words …"
Generator
Flow-matching transformer
Inpainted masked span
Audio out

Video Regenerate and Translation

Audio-driven lip-sync for editing and translation.

Video in
Latent frames
Audio in
References a few frames of you
Generator
Flow-matching transformer
Lower-face latents
Video out

Anchored Tree Sampling

A training-free inference scheduler that replaces autoregressive rollout with sparse-to-dense tree sampling, enabling drift-free long-horizon video-to-video generation.

ATS tree structure: root, guidance, and leaf calls
Tree structure
ATS sparse-to-dense filling over the horizon
Sparse-to-dense filling

Publications & Submissions

Other Research