I am an AI Researcher at Descript, working on generative models that power audio and video editing in production. My current focus spans audio codecs, diffusion models and audio-text forced alignment. A full list of papers is on Google Scholar.

Applied Research

Audio Regenerate

Seamless word-level audio editing via latent inpainting.

Audio in
→
Latent frames
→
Text in "… new words …"
Generator
Flow-matching transformer
→
Inpainted masked span
→
Audio out

Video Regenerate and Translation

Audio-driven lip-sync for editing and translation.

Video in
→
Latent frames
→
Audio in
References a few frames of you
Generator
Flow-matching transformer
→
Lower-face latents
→
Video out

Anchored Tree Sampling

A training-free inference scheduler that replaces autoregressive rollout with sparse-to-dense tree sampling, enabling drift-free long-horizon video-to-video generation.

ATS tree structure: root, guidance, and leaf calls
Tree structure
ATS sparse-to-dense filling over the horizon
Sparse-to-dense filling

Publications & Submissions

Other Research