AUDIO COMPARISONS

DriftAudio

Marginal Drifting for Distributional Post-Training of Text-to-Audio Generators

Text-to-audio generationOne-step generationDistributional post-trainingDrifting

Listen side by side

10 prompts · 4 models + Ground Truth

The text shown is the exact caption used in our test877 manifest. Starting one player pauses the others.

Text promptMeanAudio-S-FullReleased checkpointDriftAudioFrom MeanAudio · step 4400Final checkpointFdAudioOfficial released checkpointDriftAudioFrom FdAudio · step 1800Final checkpointGround TruthReal reference recordingAudioCaps · test877

About these examples

Clip IDs follow the ten examples on the FdAudio demo page. Our captions may differ from that page because each AudioCaps clip has multiple captions. All four generated columns use the same local test877 caption for each clip.

All audio is displayed as 10-second, 16 kHz mono, 16-bit lossless FLAC. The first 160,000 PCM samples of each source are unchanged. No resampling, loudness normalization, denoising or lossy compression was applied. These examples are illustrative and do not replace aggregate evaluation.

Sample list & processing details

Ground Truth uses the same clip IDs and evaluation-preprocessed real recordings as test877. It is a real recording, encoded to the same lossless format as the generated audio.