Music generation · Audio demo

Beyond Reconstruction:
Full-Context Generative DiT for Music Generation

FullDiT treats acoustic rendering as full-context generation from an imperfect discrete plan, combining frame-aligned RVQ conditions, independently encoded lyrics and captions, and non-causal full-song context.

Yunjia Li1,2, Menglin Wu2, Junyu Dai2, Xinyue Fan2, Xiangang Li2, Haoxu Wang2, Jianwei Yu2, Huaicheng Zhang2, Han Zhao2, Weiqin Li2, Yufei Shi2, Cheng Wen2, Sitong Zhao2, Qixi Zheng2, Haina Zhu2, Wei Li1

1 College of Computer Science and Artificial Intelligence, Fudan University  ·  2 Alibaba Token Foundry

+0.77ViSQOL under synthetic corruption
69.7%non-tied preference over no EMDC
15 / 18automatic metrics led
Top 3public vocals leaderboard

Method

Rendering beyond reconstruction

Error-Matched Distractor Conditioning (EMDC) exposes the renderer to realistic codec-token mistakes without changing the acoustic target. Four-way classifier-free guidance independently controls codec, lyric, and caption conditioning at inference time.

FullDiT architecture showing caption, lyrics, codec tokens, the full-context generative DiT, and recovered waveform.
FullDiT acoustic renderer and Error-Matched Distractor Conditioning.
01

Full-song context

Non-causal self-attention models the complete acoustic latent sequence rather than isolated local windows.

02

Error-matched training

Codebook-specific replacement rates and cosine-KNN distractors mimic plausible language-model errors.

03

Independent guidance

4-CFG separates codec, lyric, and caption guidance increments for controllable rendering.

Full-length samples

Complete music generations

Listen to the full tracks. Audio loads only after you press play.

01Complete vocal song ·

I Am Free

Prompt. Alternative metal driven by thick, distorted guitars, thunderous drums, gritty male vocals, and an anthemic narrative about breaking free and reclaiming one's identity.

02Complete vocal song ·

Still Waiting

Prompt. An intimate midnight vocal jazz track exploring the theme of lost love through smoky vocals and a sparse acoustic arrangement.

03Complete vocal song ·

New Year Fireworks

Prompt. An upbeat reggae-pop New Year's Eve celebration driven by acoustic guitar upstrokes, a syncopated electric bassline, punchy drums, bright brass, and a cheerful male lead vocal, with an energetic and uplifting mood.

04Instrumental ·

Changing Pace

Prompt. A mellow contemporary R&B instrumental built from syncopated electric bass, laid-back drums, rhythmic electric-piano chords, warm synth pads, and seamless shifts between standard and double-time feels, ending in a slow fade-out.

Controlled comparison

What changes the rendering?

Prompt, lyrics, upstream LM codec tokens, duration, random seed, sampler, and guidance tuple (1, 2, 1) are fixed. Only the FullDiT configuration changes.

Shared prompt

A dreamy Folk Pop track evoking the feeling of falling asleep under the stars, featuring a soft male vocal and gentle acoustic instrumentation.

Duration
3:17
Vocal
English · male
Tempo
75 BPM
Key
G major
M1Complete

FullDiT

Full-song context, renderer-side caption and lyrics, and EMDC.

M2aAblation

Local context

Local 30-second training context; otherwise matched to M1.

M2bAblation

No renderer text

No renderer-side caption or lyric conditions; otherwise matched to M1.

M3Ablation

No EMDC

No Error-Matched Distractor Conditioning; otherwise matched to M1.

Paired reconstruction

Suno V5.5 original vs. FullDiT

Each pair has matching duration. Switch between the two players for a direct comparison.

Pair 01

Instrumental

AReference

Suno V5.5 original

BReconstruction

FullDiT reconstruction

Pair 02

Vocal

AReference

Suno V5.5 original

BReconstruction

FullDiT reconstruction

Results

External evaluation

Automatic comparisons across leading music-generation systems on a shared multilingual set of vocal music.

Automatic comparison of six music-generation systems on a shared multilingual set of vocal music, evaluated with SongBench, SongEval, Audiobox, and CMI-RM. Best result in each row is bold.
Evaluator Metric Ours Suno 5.5 Suno 5 Mureka 8 Lyria 3 Pro MiniMax 2.6
SongBench Melody 7.20016.69396.92027.06607.03776.7245
Arrangement7.38796.84067.08057.26017.15116.8167
Musicality6.36785.82976.00966.25596.18105.8823
Vocal7.62347.09817.31607.62487.45607.2513
Instrument7.35177.00807.14877.25857.05836.8458
Mixing7.30477.03327.11017.14337.00936.7397
Structure6.96506.61586.73637.02377.05656.5637
SongEval Coherence 4.50944.29054.32784.38704.46734.2772
Musicality4.40984.15474.19704.28754.34664.1781
Memorability4.46634.21234.24734.36804.41534.2180
Clarity4.38384.12464.17414.24314.30994.1448
Naturalness4.26004.00464.06794.17674.20213.9919
Audiobox Content Enjoyment 7.70897.52897.55077.59217.69507.6844
Content Usefulness7.99897.95357.80207.86347.89197.9578
Prod. Complexity6.87746.64866.81186.87686.75616.7514
Production Quality8.29868.20838.14118.10988.27808.2923
CMI-RM Alignment 2.17861.92161.91192.30892.03772.1546
Musicality2.76112.36652.37562.75902.52802.6661