DriftAudio
Marginal Drifting for Distributional Post-Training of Text-to-Audio Generators
Listen side by side
10 prompts · 4 models + Ground TruthThe text shown is the exact caption used in our test877 manifest. Starting one player pauses the others.
| Text prompt | MeanAudio-S-FullReleased checkpoint | DriftAudioFrom MeanAudio · step 4400Final checkpoint | FdAudioOfficial released checkpoint | DriftAudioFrom FdAudio · step 1800Final checkpoint | Ground TruthReal reference recordingAudioCaps · test877 |
|---|
About these examples
Clip IDs follow the ten examples on the FdAudio demo page. Our captions may differ from that page because each AudioCaps clip has multiple captions. All four generated columns use the same local test877 caption for each clip.
All audio is displayed as 10-second, 16 kHz mono, 16-bit lossless FLAC. The first 160,000 PCM samples of each source are unchanged. No resampling, loudness normalization, denoising or lossy compression was applied. These examples are illustrative and do not replace aggregate evaluation.
Sample list & processing detailsGround Truth uses the same clip IDs and evaluation-preprocessed real recordings as test877. It is a real recording, encoded to the same lossless format as the generated audio.