TACD: Distilling Efficient Text-to-Motion Models via Terminal Amplification Control
TL;DR. We distill large text-to-motion teachers into few-step students from prompts alone (no real motion data). On a fixed supervision grid, velocity matching repeatedly overweights errors near the end of denoising; TACD bounds that weight by tying the latest teacher query to the student's step size, without changing inference.
Abstract
Recent text-to-motion models have improved motion quality and instruction following, yet many-step denoising and large model components make deployment slow and memory-intensive. We present Terminal-Amplification-Controlled Distillation (TACD), an on-policy approach for training efficient motion generators from text prompts and pretrained teachers, without real-motion training data.
Building on segmented on-policy flow distillation, we supervise clean-motion predictions along student-generated trajectories. We identify a failure mode in which velocity matching on a fixed supervision grid repeatedly overweights errors near the denoising endpoint, degrading few-step generation. TACD ties the latest teacher query to the student's step size, bounding the effective loss weights in clean-motion space without changing inference.
Experiments on HumanML3D and KIT-ML demonstrate improved few-step generation, including a 58% reduction in eight-step HY-Motion student FID relative to distillation without this bound. For diffusion teachers, the endpoint-matching form of TACD yields four-step students with lower FID and matched or improved text–motion retrieval relative to their 50-step teachers on HumanML3D. On HY-Motion and Kimodo, eight-step students with compact components achieve 7.7–11.9× end-to-end speedups and reduce peak GPU memory by 3.8–6.7× relative to their teachers.
Samples
Eight-step samples from the distilled student, one prompt and one noise sample per clip (30 fps).
Method
TACD is a supervision design for segmented on-policy distillation. In each denoising segment the student predicts the complete clean motion once; this prediction fixes the student's states within the segment, and the frozen teacher, queried at those states during training only, supplies the targets that correct it.
Segmented on-policy distillation
Given text and noise, the student generates a complete motion in evaluations, without real-motion training targets. Flow time runs from noise at to clean motion at . The sampling grid fixes the student evaluations and splits denoising into the segments . Following the single-knot DX parameterization of -Flow, one clean-motion prediction defines the student policy over each segment: at the segment start , the student's velocity output gives
The prediction estimates the clean motion at , not the state at the next boundary; holding it fixed within the segment yields the inference update
Supervision along the student rollout. Teacher supervision is obtained on states that the student itself visits. An EMA copy of the student generates the rollout: starting from noise , it produces the segment starts . Within each segment we place supervision times
where the terminal margin keeps the last query below , at which the policy is undefined. This supervision grid is used only during training and is distinct from the sampling grid, which always ends at . The rollout states at the supervision times are
with the stop-gradient. The online student computes its prediction once at each detached segment start and reuses it for the segment's queries; the frozen teacher velocity is queried at these states. Velocity matching penalizes, at each query, the residual
where are segment weights and is the squared error averaged over valid frames and channels. Rollout states and teacher targets are constants for the update, so gradients pass only through the online predictions at the segment starts. This construction fixes where supervision is obtained; the matching space sets how strongly each query counts.
Terminal amplification in velocity matching
The teacher's instantaneous velocity implies a clean-motion estimate, . Substituting it into the velocity residual at a supervision state and time gives
with the clean-motion loss: velocity matching is clean-motion matching with an inverse-square time-dependent weight. The objective is therefore
where the segment weights are a design choice but the coefficients are not: they are induced by matching in velocity space.
The last query of the supervision grid sits at and receives the terminal coefficient . On a uniform sampling grid every other coefficient is at most , so a small numerical margin makes one coefficient exceed all others by orders of magnitude: gives , while at and the other fifteen coefficients are at most .
Because the supervision grid is reused at every update, this weight returns at the same supervision time in every update, instead of being met occasionally as under random time sampling. We call this persistent overweighting terminal amplification.
Terminal amplification control
TACD bounds every coefficient by a value tied to the student's step size. It retains the sampling grid, the inference update, and the training budget, changes only the supervision used during training, and adds no network module or training stage. For a flow teacher it ties the terminal margin to the step of the uniform sampling grid :
The latest teacher query then sits at the midpoint of the final student step, so the teacher is never queried next to the singular time, while the sampling grid still ends at . The terminal coefficient has the same order as the others on the grid, and the rule is closed-form in the step count, with no tuned threshold.
Two other enforcements meet the same bound at the original queries, each with a different loss. The three differ in how they meet the bound:
1 Grid capping TACD
Moves the final-segment queries and keeps the velocity loss unchanged. TACD's rule for every flow result.
2 Endpoint matching
Removes the conversion factor at every query and matches the teacher's clean-motion estimates directly. The form TACD takes for diffusion teachers.
3 Weight capping
Lowers one coefficient at the original queries; with it changes a single term of the objective.
At the bound is , and TACD queries the final segment at and .
One training iteration, flow teacher
- Set and build the supervision times ; sample noise and text .
- Detached rollout. For : the rollout EMA predicts , stores the states for , and steps to .
- Loss. For each : the online student predicts at the detached ; for each , query the frozen teacher's estimate and add to the loss.
- Update ; update the rollout EMA and the separate evaluation EMA.
Only the online student receives gradients. The endpoint-matching control keeps the original supervision grid and uses in place of .
Extension to diffusion teachers
For a diffusion teacher the student again predicts clean motion at each segment start and learns from frozen-teacher targets on detached rollout states, with two changes: the states follow deterministic DDIM sub-trajectories, and the loss is the objective above with . Let be the teacher sampler's -step timestep schedule with cumulative signal level , and the student's clean-motion prediction at the segment start ( is the flow case's under reversed time). Holding the rollout prediction fixed over the segment also fixes the noise it implies, and the DDIM states at supervision times within the segment are
The target is the frozen teacher's clean-motion estimate , computed with the teacher sampler's guidance. Written in clean-motion space from the start, this loss carries no factor; matching noise predictions on the same states would instead weight the clean-motion error by the signal-to-noise ratio, which also grows without bound toward the clean end. Endpoint matching thus carries no time-dependent weight, and for diffusion teachers TACD denotes this endpoint-matching form.
Results
HumanML3D and KIT-ML use dataset-specific Guo evaluators and real-motion references, averaged over 20 evaluator repetitions. The main flow student, TACD (HY-Motion), pairs the 0.46B motion backbone of the released HY-Motion Lite with frozen Qwen3-0.6B conditioning and a trainable projection. Training uses training-split captions as text only, with teacher supervision and no real-motion targets; evaluation uses the test split.
Main comparison
At eight steps with the HY-Motion teacher, TACD is best in R@1, R@3, MM-Dist, and FID.
| Method | NFE | R@1 ↑ | R@3 ↑ | MM-Dist ↓ | FID ↓ | Div → | MMod ↑ |
|---|---|---|---|---|---|---|---|
| Real motions | – | .511 | .797 | 2.974 | .002 | 9.503 | – |
| Our evaluation: HY-Motion references | |||||||
| HY-Motion teacher | 50 | .455 | .743 | 3.471 | .541 | 9.13 | 3.28 |
| HY-Motion Lite (released) | 50 | .435 | .728 | 3.575 | .726 | 8.95 | – |
| Truncated teacher | 8 | .378 | .652 | 4.225 | 3.250 | 8.24 | – |
| Controlled Lite students, three-seed means: Qwen3-0.6B encoder, teacher-only supervision | |||||||
| Uncapped student | 8 | .402 | .708 | 3.703 | 2.486 | 8.01 | 3.994 |
| TACD | 8 | .479 | .782 | 3.192 | 1.041 | 8.92 | 3.541 |
| Non-adversarial ports to the Lite student, three-seed means: teacher-only supervision | |||||||
| DMD port | 8 | .426 | .730 | 3.493 | 1.829 | 8.22 | 3.023 |
| Consistency port | 8 | .394 | .699 | 3.754 | 2.908 | 7.77 | 3.803 |
| Reflow port | 8 | .370 | .660 | 4.063 | 3.341 | 8.07 | 3.886 |
HumanML3D. NFE counts guided denoising steps per motion; Div →: closer to real is better. Bold: best of the eight-step HY-Motion rows.
At the same eight-step Lite budget, TACD reduces mean FID from 2.486 to 1.041 and improves R@3 from .708 to .782. It exceeds the R@3 of the 50-step teacher (.743) and of the released Lite at 50 steps (.728). The ports, our reimplementations of DMD, consistency distillation, and Reflow on the same Lite student, share its architecture, prompts, and generator-update budget. Under the MARDM evaluator, TACD leads the truncated teacher and every eight-step student in each seed and cuts the uncapped student's FID from 7.559 to 3.575.
Terminal control
Bounding the coefficient restores quality; restoring it reverses the gain.
At and , the original supervision grid queries the final segment at .937 and .999: the terminal coefficient is , while the other fifteen coefficients are at most 252. All three enforcements of the bound recover FID to 1.01–1.04. Conversely, restoring on the capped grid, without moving any query, worsens FID in all three seeds, to a mean of 2.669.
| Configuration | Final-segment queries | FID ↓ | seed range | vs. uncapped | |
|---|---|---|---|---|---|
| Uncapped student | 0.937, 0.999 | 2.486 | [2.419–2.563] | – | |
| 1 TACD (grid capping) | 0.90625, 0.9375 | 256 | 1.041 | [1.009–1.074] | −58% |
| 2 Endpoint matching | 0.937, 0.999 | 1 | 1.029 | [0.988–1.091] | −59% |
| 3 Weight capping | 0.937, 0.999 | 256 | 1.011 | [0.980–1.044] | −59% |
| Last-point-only | 0.937, 0.9375 | 256 | 1.039 | [1.032–1.048] | −58% |
| Capped grid, restored to | 0.90625, 0.9375 | 2.669 | [2.645–2.703] | +7% |
Eight-step Lite student, HumanML3D; is the terminal coefficient. Each bound (1–3) cuts FID by 58–59%; restoring (rose) undoes it. Brackets: three-seed range.
A dose–response. FID stays at 1.01–1.07 from to and rises monotonically beyond it, to 1.14, 1.30, 1.70, and 2.49 at , , , and , while the median pre-clip gradient norm rises in step from 6 to 253. sits in a plateau at least wide.
Gradient concentration follows the grid. The original grid puts 96.6–97.6% of the gradient norm on the final segment and the capped grid 3.6–11.8%.
One query carries the effect. Keeping .937 and moving only .999 to .9375 gives FID 1.039, close to full capping's 1.041.
VLM judge
On SSAE, a VLM-judged semantic-alignment score, TACD scores highest.
| Model | SSAE ↑ |
|---|---|
| HY-Motion teacher @50 | .695 |
| HY-Motion Lite (released) @50 | .692 |
| Truncated teacher @8 | .587 |
| Uncapped student @8 | .650 |
| DMD port @8 | .672 |
| Consistency port @8 | .700 |
| Reflow port @8 | .590 |
| TACD @8 | .720 |
SSAE: a VLM judge's yes-rate on questions about the prompted motion, on a fixed set of 640 prompts.
TACD scores highest (.720), above the 50-step teacher (.695), the released Lite (.692), and every other eight-step student (.587–.700).
Human study
In a blind ranking study, annotators rank eight-step TACD above its 50-step teacher in 51% of comparisons, on par with it.
| Model | Above teacher (%) ↑ [95% CI] | Prompts W/T/L |
|---|---|---|
| HY-Motion teacher @50 | 50 (reference) | – |
| MotionLCM-V2 @4 | 41.2 [30.4, 51.7] | 17/0/33 |
| MLCT @4 | 33.9 [24.5, 43.4] | 11/0/39 |
| TACD @8 | 50.8 [41.1, 60.5] | 28/1/21 |
Win rate against the 50-step teacher and prompts won/tied/lost. 50 prompts, four systems rendered from joints, systems hidden and screen positions balanced; 17 annotators contributed 691 rankings.
By the same measure, the four-step MotionLCM-V2 and MLCT reach 41% and 34%.
Diffusion teachers
Four-step TACD has lower FID than its 50-step teacher and the released MotionLCM student.
| Method | Training signal | NFE | R@1 ↑ | R@3 ↑ | MM-Dist ↓ | FID ↓ | Div → | MMod ↑ |
|---|---|---|---|---|---|---|---|---|
| MLD teacher and MotionLCM-V1 | ||||||||
| MLD teacher (DDIM) | Reference | 50 | .502 | .795 | 3.048 | .387 | 9.53 | – |
| Truncated MLD teacher | None | 4 | .459 | .768 | 3.295 | .585 | 9.49 | – |
| MotionLCM-V1 | Real motions | 4 | – | .798 | – | .286 | – | – |
| TACD | Teacher outputs | 4 | .525 | .818 | 2.909 | .222 | 9.80 | – |
| MotionLCM-V2 teacher and student | ||||||||
| MotionLCM-V2 teacher | Reference | 50 | .537 | .824 | 2.838 | .066 | 9.59 | – |
| MotionLCM-V2 student | Real motions | 4 | .554 | .841 | 2.760 | .072 | 9.54 | 1.72 |
| TACD | Teacher outputs | 4 | .538 | .824 | 2.838 | .055 | 9.57 | 1.94 |
HumanML3D, our evaluation; bold: best at four steps in each group. TACD FID is the mean of three training seeds (±.002 for MLD, ±.001 for MotionLCM-V2).
With endpoint matching, four-step TACD improves on the 50-step MLD teacher in both FID (.387 → .222) and R@3 (.795 → .818), and its FID is below the .286 we measure for the released MotionLCM-V1 checkpoint, a student trained on real motions. With the MotionLCM-V2 teacher, four-step TACD reaches FID .055 against the teacher's .066 and the released official student's .072 under the same evaluation, at the teacher's R@3 (.824).
KIT-ML and Kimodo
On KIT-ML (MDM teacher) and Kimodo, TACD with the teacher's encoder beats the truncated teacher at every NFE.
KIT-ML. At the same NFE, TACD reduces FID from the truncated teacher's 1.530, .976, and .745 to 1.038, .852, and .649 at two, four, and eight steps, 32% lower at two steps and 13% at four and eight.
| Method | NFE | R@1 ↑ | R@3 ↑ | MM-Dist ↓ | FID ↓ | Div → |
|---|---|---|---|---|---|---|
| Real motions (GT) | – | .426 | .780 | 2.786 | .027 | 10.99 |
| MDM teacher | 1000 | .403 | .725 | 3.115 | .506 | 10.74 |
| Truncated teacher | 2 | .395 | .728 | 3.193 | 1.530 | 11.24 |
| TACD | 2 | .410 | .756 | 2.996 | 1.038 | 11.06 |
| Truncated teacher | 4 | .410 | .755 | 3.014 | .976 | 11.38 |
| TACD | 4 | .416 | .761 | 2.932 | .852 | 11.16 |
| Truncated teacher | 8 | .416 | .756 | 2.948 | .745 | 11.34 |
| TACD | 8 | .421 | .763 | 2.873 | .649 | 11.08 |
KIT-ML, MDM teacher. TACD rows: three-seed means after 10k updates.
Kimodo. At eight steps, TACD raises R@1 over the truncated teacher from 69.36% to 71.39% and lowers FID from .064 to .054, with the same ordering at four and two steps.
| Method | NFE | R@1 (%) ↑ | R@3 (%) ↑ | FID ↓ | Time (s) ↓ |
|---|---|---|---|---|---|
| Teacher | 100 | 73.15 | 88.63 | 0.045 | 2.96 |
| Truncated teacher | 8 | 69.36 | 84.35 | 0.064 | 0.29 |
| TACD | 8 | 71.39 | 86.55 | 0.054 | 0.29 |
| Truncated teacher | 4 | 67.05 | 82.37 | 0.079 | 0.18 |
| TACD | 4 | 68.42 | 83.25 | 0.067 | 0.18 |
| Truncated teacher | 2 | 60.46 | 77.27 | 0.107 | 0.13 |
| TACD | 2 | 62.36 | 79.57 | 0.095 | 0.13 |
Kimodo teacher, 607-caption TMR evaluation; TACD keeps the teacher's 8B text encoder. Time: end to end, one RTX L40.
Efficiency
With compact backbones and text encoders, eight-step students run 7.7–11.9× faster end to end and use 3.8–6.7× less peak GPU memory than their teachers.
HY-Motion. The Lite backbone and Qwen3-0.6B encoder of TACD (HY-Motion) reduce peak GPU memory from 17.63 to 2.62 GB (6.7×), mostly from the 8B-to-0.6B text encoder, and eight-step sampling cuts end-to-end time on an RTX L40 (bf16, batch 1) from the teacher's 2077 to 268 ms (7.7×).
Kimodo. TACD (Kimodo), with a trained-in 1B encoder, reduces peak memory from 15.3 to 4.0 GB (3.8×) and end-to-end time from 2.96 to .25 s (11.9×) at the R@3 and FID of the 8B-encoder student.
Models
The TACD student we release. It is trained from the frozen HY-Motion teacher and text prompts, with teacher supervision and no real-motion targets.
| Student | Teacher | Benchmark | Steps | Main result |
|---|---|---|---|---|
| TACD (HY-Motion) coming soon0.46B Lite backbone, Qwen3-0.6B encoder | HY-Motion, 50 steps | HumanML3D | 8 | FID 1.041 · R@3 .782 |
Loading with Hugging Face Transformers, once released:
Release candidate: one of the three training seeds, with FID 1.0085, R@3 0.7829 and MM-Dist 3.1811 on HumanML3D; the row above is the three-seed mean. | ||||
HumanML3D: Guo evaluator, 20 repetitions.
BibTeX
@article{huang2026tacd,
title = {TACD: Distilling Efficient Text-to-Motion Models via Terminal Amplification Control},
author = {Huang, Wei-Jin and Li, Yuan-Ming and Lin, Kun-Yu and Luo, Wang and Zhu, Yinlin and
Yu, Yue and Ye, Shenghao and Yuan, Junbin and Hong, Fa-Ting and Zhang, Qing and Zheng, Wei-Shi},
journal = {arXiv preprint},
year = {2026}
}




