AudioGAR

AudioGAR: Bridging Reconstruction and Generation

in Latent Audio Generative Models

Xianghong Fang1,*· Geeyang Tay1,*· Wentao Ma1· Tim G. J. Rudner1,2· Dehan Kong1

1University of Toronto2Vijil*Equal contribution

Can we obtain generation-relevant latents that remain paired with source audio for decoder fine-tuning?

In our paper, we show that AudioGAR constructs intermediate latents between reconstruction and generation, retaining source correspondence at lower noise levels to support generation-aware codec decoder adaptation.

31.3%lower FAD on MusicCaps
1.60 → 1.10
16.8%lower FAD on AudioCaps
1.61 → 1.34
1.5%of original training audio
508.3 audio hours
0.26%of original training cost
10.3 H100-equivalent hours

One decoder.
Two latent distributions.

Latent audio models learn in two stages. A codec learns to reconstruct audio, then a generative model learns to produce latents in the codec’s space. The decoder trains on encoder-induced latents, but receives generator-produced latents at inference.

For a source waveform x, reconstruction decodes ze = Eθ(x) ∼ Pe. Generation decodes zg ∼ Pg, produced from noise and a text caption. When Pe ≠ Pg, the decoder faces a train-generation mismatch.

Strong reconstruction quality therefore need not transfer to generated audio. AudioGAR adapts the decoder to latents that incorporate the generative process while preserving a source waveform for supervision.

The mismatch
is measurable

We construct encoder and generator latents from matched audio-caption examples in AudioX. Linear Discriminant Analysis (LDA) finds a direction that separates the two collections, both at the decoder input and within its feature spaces.

LDA density plots for AudioX on AudioCaps, MusicCaps, and VGGSound-Omni, showing separation of encoder and generated latents at the input and three decoder stages.
Figure 1. Empirical evidence of distribution mismatch (paper §2.2). Rows show AudioCaps, MusicCaps, and VGGSound-Omni; columns show latent space, Conv In, Block 2, and Block 4. Curves are Gaussian-smoothed, normalized densities of one-dimensional LDA scores for encoder latents ze and generated latents zg; shading marks overlap. AUC is the area under the receiver operating characteristic curve, and d reports standardized separation.

Latent-space AUC reaches 0.962 on AudioCaps, 0.862 on MusicCaps, and 0.845 on VGGSound-Omni. Separation persists after the latents enter the decoder, including Conv In, Block 2, and Block 4.

Reconstruction quality leaves a generation gap

We compare reconstructed and generated audio across 29 model-dataset combinations. FD uses PANNs embeddings; FAD uses VGGish embeddings. The prefixes r and g denote reconstruction and generation. Lower values indicate closer agreement with the reference audio distribution.

Table 1

Reconstruction-generation performance gap

Reconstruction-generation performance gap across AudioCaps, MusicCaps, and VGGSound-Omni. Lower is better for every metric.
DatasetModelrFD ↓gFD ↓rFAD ↓gFAD ↓
AudioCapsAudioLDM-2-Base3.1911.541.312.04
AudioLDM-2-Large3.2011.531.311.86
AudioLDM-S-Full3.0521.851.274.79
AudioX7.5311.833.491.61
Stable Audio Open7.5330.363.493.41
Tango2.9714.471.141.52
Tango Full2.9722.231.144.10
Tango Full (FT AudioCaps + MusicCaps)2.979.771.142.59
Tango Full (FT AudioCaps)2.9710.421.142.39
Tango (AF-AC, FT AudioCaps)2.9712.081.142.55
MusicCapsAudioLDM-2-Large2.2018.840.662.51
AudioLDM-2-Base2.2021.710.663.82
AudioX4.319.560.951.60
Stable Audio Open4.3137.720.953.42
Tango2.1846.800.565.64
Tango Full2.1835.670.565.78
TangoMusic2.0415.200.421.85
Tango Full (FT AudioCaps + MusicCaps)2.1820.430.563.61
VGGSound-OmniAudioLDM-2-Base1.6410.030.491.72
AudioLDM-2-Large1.649.930.491.49
AudioLDM-S-Full1.6318.780.503.44
AudioX3.669.671.911.74
Stable Audio Open3.6626.071.912.68
Tango1.6224.030.502.88
Tango Full1.6219.010.503.69
TangoMusic1.6213.430.502.36
Tango (AF-AC, FT AudioCaps)1.6211.720.501.46
Tango Full (FT AudioCaps + MusicCaps)1.6213.260.502.28
Tango Full (FT AudioCaps)1.6219.120.503.06

All 29 comparisons from the supplied table (paper Table 2). Blue columns show FD; green columns show FAD. FT denotes fine-tuning. Model names link to their checkpoints.

rFD is lower than gFD in all 29 comparisons; rFAD is lower than gFAD in 26. For example, Tango on MusicCaps has rFD 2.18 and gFD 46.80. AudioX on AudioCaps is an FAD exception, with rFAD 3.49 and gFAD 1.61.

The AudioCaps quantile comparison checks whether a few models drive the FD gap. It compares the empirical reconstruction and generation distributions across the evaluated models.

Empirical quantile curves of reconstruction FD and generation FD on AudioCaps, separated across the full quantile range with a median gap of 8.95.
Figure 2. AudioCaps quantiles of rFD and gFD (paper Figure 2). The horizontal axis is FD and the vertical axis is quantile level. The generation median exceeds the reconstruction median by 8.95.

The curves remain separated across the full quantile range. The gap alone does not isolate the decoder’s contribution, because generation also depends on the latent generator. Combined with the LDA evidence, it motivates adapting the decoder to generation-relevant inputs.

Generation-aware
decoder adaptation

Fully generated latents lack a paired source waveform. AudioGAR starts from an encoded waveform, adds controlled noise, and denoises through the frozen generative model. Lower-noise outputs retain the correspondence needed for paired supervision.

We first construct intermediate latents, then use them as decoder inputs while supervising against the original audio. The codec encoder and latent generator remain frozen.

AudioGAR encodes a source waveform, perturbs its latent with Gaussian noise, denoises with the frozen generative model, and decodes the intermediate latent back to audio.
Figure 3. AudioGAR latent construction (paper §3.1). Eθ maps source audio x to ze. Gaussian noise ε ∼ N(0, I) produces znoisyt = atze + btε. The frozen generator denoises this input, conditioned on the source caption, to obtain zgt ∼ Pgt. The decoder reconstructs the waveform from this intermediate latent.

The relative noise level ηt = bt / (at + bt) controls the transition. ηt = 0 recovers reconstruction; ηt = 1 recovers generation from pure noise. Our experiments use at = cos(πt/2) and bt = sin(πt/2).

01 / Encode

Keep the source pair

Encode a waveform with the pretrained, frozen codec encoder.

02 / Perturb

Choose the noise level

Mix the encoder latent with Gaussian noise using the diffusion schedule.

03 / Denoise

Use the generator

Denoise with the frozen model and the caption paired with the source audio.

04 / Adapt

Fine-tune the decoder

Decode the AudioGAR latent and apply the original codec’s decoder-side objective.

LAudioGAR = Lcodec(x, Dϕ(zgt))

We retain the codec’s reconstruction and adversarial objectives and their loss weights. AudioX updates only the codec decoder; TangoMusic jointly updates the VAE decoder and HiFi-GAN vocoder. Both keep the encoder and latent generator frozen.

AudioGAR improves generation by adapting the decoder to generation-relevant latents with paired source supervision.

How do latents shift
as noise increases?

Latent-FD compares empirical means and covariances under a Gaussian approximation, directly in latent space. We measure the distance from each intermediate distribution Pgt to both the encoder endpoint Pe and generator endpoint Pg.

On AudioCaps, Latent-FD to encoder latents increases with noise while Latent-FD to generated latents decreases.(a) AudioCaps
On MusicCaps, intermediate distributions also move away from encoder latents and toward generated latents as noise increases.(b) MusicCaps
Figure 4. AudioGAR bridges the two latent distributions (paper §3.2, Figure 4). Blue curves compare Pgt with Pe; green curves compare Pgt with Pg. The horizontal axis is noise level ηt; the vertical axis is Latent-FD. Both datasets trace the same transition toward the generation endpoint.

On both datasets, increasing noise moves AudioGAR latents farther from Pe and closer to Pg. Lower noise preserves source correspondence; larger noise moves closer to generation while weakening that correspondence.

Better generation.
A small adaptation budget.

AudioGAR improves FD and FAD on AudioX for both audio and music generation. It also transfers to TangoMusic’s decoder and vocoder pipeline. The inference pipeline uses the adapted weights with no additional inference cost.

Table 2

Generation before and after AudioGAR

Reproduced baselines and AudioGAR results from paper Table 3.
DatasetModelgFD ↓gFAD ↓KL ↓IS ↑PC ↑PQ ↑
AudioCapsAudioX11.831.611.3112.473.165.73
AudioX + AudioGAR11.621.341.2912.063.115.73
MusicCapsAudioX9.561.601.003.654.786.61
AudioX + AudioGAR8.331.100.993.594.756.56
TangoMusic15.201.851.092.855.627.22
TangoMusic + AudioGAR14.251.471.092.835.617.16

Subset of paper Table 3 using our reproduced baselines. Shading identifies AudioGAR rows. KL: Kullback-Leibler divergence; IS: Inception Score; PC: Production Complexity; PQ: Production Quality. FD and FAD measure distributional agreement; PC and PQ use Meta Audiobox Aesthetics.

On MusicCaps, AudioGAR reduces AudioX FAD by 31.3% and TangoMusic FAD by 20.5%. IS and PC decline slightly, and PQ stays unchanged or declines slightly; the main gains are in FD and FAD.

Table 3

AudioX adaptation cost and data

AudioGAR adaptation relative to original AudioX training, from paper Table 1.
TrainingH100-equivalent hoursRelative costAudio hoursRelative data
Original AudioX4,000100%33,052.3100%
AudioGAR adaptation10.30.26%508.31.5%

Adaptation runs for 100k steps on one NVIDIA H100 80GB GPU. Cost and data percentages use original AudioX training as the denominator. The 10.3 hours are the adaptation cost, additional to pretrained AudioX.

What makes adaptation work?

Does the choice of decoder latents matter?

We keep the adaptation procedure fixed and replace encoder latents with AudioGAR latents. Fine-tuning on encoder latents already improves FD and FAD, but AudioGAR yields larger gains on both datasets.

Table 4

Decoder input ablation

FD and FAD subset of paper Table 4, comparing pretrained AudioX with two decoder adaptation inputs.
Decoder settingAudioCaps FD ↓AudioCaps FAD ↓MusicCaps FD ↓MusicCaps FAD ↓
Pretrained AudioX11.831.619.561.60
Adapt with encoder latents ze11.721.558.871.38
Adapt with AudioGAR latents zgt11.621.348.331.10

FD and FAD subset of paper Table 4. All values evaluate generated audio; lower is better.

On MusicCaps, replacing encoder latents with AudioGAR latents lowers FAD from 1.38 to 1.10. This comparison isolates the benefit of generation-aware inputs beyond decoder fine-tuning alone.

  1. Encoder and generator latents differ across the evaluated audio models, and strong reconstruction alone does not predict strong generation.
  2. AudioGAR constructs intermediate latents that incorporate the generative process while retaining source correspondence at lower noise levels.
  3. Adapting decoder components improves FD and FAD on AudioX and TangoMusic without adding inference cost.
  4. AudioX adaptation uses 508.3 hours of audio and 10.3 H100-equivalent hours, or 1.5% of its original training audio and 0.26% of its training cost.
@article{fang2026audiogar,
  title   = {AudioGAR: Bridging Reconstruction and Generation in Latent Audio Generative Models},
  author  = {Fang, Xianghong and Tay, Geeyang and Ma, Wentao and Rudner, Tim G. J. and Kong, Dehan},
  journal = {Arxiv},
  year    = {2026}
}