TACD: Distilling Efficient Text-to-Motion Models via Terminal Amplification Control

Wei-Jin Huang1 Yuan-Ming Li1 Kun-Yu Lin2 Wang Luo1 Yinlin Zhu1 Yue Yu3 Shenghao Ye4 Junbin Yuan1 Fa-Ting Hong5 Qing Zhang1 Wei-Shi Zheng1,†

1Sun Yat-sen University 2Nanyang Technological University 3Wuhan University 4University of Science and Technology of China 5The Hong Kong University of Science and Technology

†Corresponding author

TL;DR. We distill large text-to-motion teachers into few-step students from prompts alone (no real motion data). On a fixed supervision grid, velocity matching repeatedly overweights errors near the end of denoising; TACD bounds that weight by tying the latest teacher query to the student's step size, without changing inference.

Motion examples from the teacher, the teacher truncated to fewer steps, and the TACD student, for the same prompts and noise.
Teacher, the same teacher truncated to fewer steps, and TACD. Rows share prompt, noise, frames and camera.

Abstract

Recent text-to-motion models have improved motion quality and instruction following, yet many-step denoising and large model components make deployment slow and memory-intensive. We present Terminal-Amplification-Controlled Distillation (TACD), an on-policy approach for training efficient motion generators from text prompts and pretrained teachers, without real-motion training data.

Building on segmented on-policy flow distillation, we supervise clean-motion predictions along student-generated trajectories. We identify a failure mode in which velocity matching on a fixed supervision grid repeatedly overweights errors near the denoising endpoint, degrading few-step generation. TACD ties the latest teacher query to the student's step size, bounding the effective loss weights in clean-motion space without changing inference.

Experiments on HumanML3D and KIT-ML demonstrate improved few-step generation, including a 58% reduction in eight-step HY-Motion student FID relative to distillation without this bound. For diffusion teachers, the endpoint-matching form of TACD yields four-step students with lower FID and matched or improved text–motion retrieval relative to their 50-step teachers on HumanML3D. On HY-Motion and Kimodo, eight-step students with compact components achieve 7.7–11.9× end-to-end speedups and reduce peak GPU memory by 3.8–6.7× relative to their teachers.

Samples

Eight-step samples from the distilled student, one prompt and one noise sample per clip (30 fps).

A person walks forward and then turns around and walks in the opposite direction
A person jumps twice with his arms relaxed at his sides
A person kneels down onto all fours, crawls towards the right, and then stands back up
A person balances on one leg with arms out like an airplane
A person standing upright is energetically performing hip-hop dance
A person paddles a kayak alternating left and right

Method

TACD is a supervision design for segmented on-policy distillation. In each denoising segment the student predicts the complete clean motion once; this prediction fixes the student's states within the segment, and the frozen teacher, queried at those states during training only, supplies the targets that correct it.

Diagram: a K-step student rollout; each segment is trained on detached student states against a frozen teacher; in the final segment the velocity loss weight grows as (1-s)^-2 and TACD caps it.
(a) One detached on-policy training step on a segment of the student's own rollout. (b) In the final segment the velocity loss weight (1−s)−2(1-s)^{-2} grows without bound near the endpoint; TACD caps the latest teacher query at 1−12K1-\tfrac{1}{2K}.

Segmented on-policy distillation

Given text cc and noise, the student generates a complete motion in KK evaluations, without real-motion training targets. Flow time runs from noise at t=0t=0 to clean motion at t=1t=1. The sampling grid 0=t0<⋯<tK=10=t_0<\cdots<t_K=1 fixes the KK student evaluations and splits denoising into the segments [tk,tk+1][t_k,t_{k+1}]. Following the single-knot DX parameterization of π\pi-Flow, one clean-motion prediction defines the student policy over each segment: at the segment start x(k)x^{(k)}, the student's velocity output vθv_\theta gives

x^1,θ(k)=x(k)+(1−tk) vθ(x(k),tk;c),πθ(k)(x,t)=x^1,θ(k)−x1−t,t∈[tk,tk+1].\begin{aligned} \hat x_{1,\theta}^{(k)}&=x^{(k)}+(1-t_k)\,v_\theta(x^{(k)},t_k;c),\\[4pt] \pi_\theta^{(k)}(x,t)&=\frac{\hat x_{1,\theta}^{(k)}-x}{1-t}, \qquad t\in[t_k,t_{k+1}]. \end{aligned}
x^1,θ(k)=x(k)+(1−tk) vθ(x(k),tk;c),πθ(k)(x,t)=x^1,θ(k)−x1−t,t∈[tk,tk+1].\begin{aligned} \hat x_{1,\theta}^{(k)}&=x^{(k)}+(1-t_k)\,v_\theta(x^{(k)},t_k;c),\\[4pt] \pi_\theta^{(k)}(x,t)&=\frac{\hat x_{1,\theta}^{(k)}-x}{1-t}, \\[5pt] &\quad t\in[t_k,t_{k+1}]. \end{aligned}

The prediction estimates the clean motion at t=1t=1, not the state at the next boundary; holding it fixed within the segment yields the inference update

x(k+1)=x(k)+tk+1−tk1−tk(x^1,θ(k)−x(k)).x^{(k+1)}=x^{(k)}+\frac{t_{k+1}-t_k}{1-t_k}\left(\hat x_{1,\theta}^{(k)}-x^{(k)}\right).

Supervision along the student rollout. Teacher supervision is obtained on states that the student itself visits. An EMA copy θˉ\bar\theta of the student generates the rollout: starting from noise xˉ(0)\bar x^{(0)}, it produces the segment starts xˉ(k)\bar x^{(k)}. Within each segment we place MM supervision times

sk,j(δ)=tk+jM[min⁡(tk+1, 1−δ)−tk],j=1,…,M,s_{k,j}(\delta)=t_k+\frac{j}{M}\left[\min(t_{k+1},\,1-\delta)-t_k\right], \qquad j=1,\ldots,M,
sk,j(δ)=tk+jM[min⁡(tk+1, 1−δ)−tk],j=1,…,M,\begin{gathered} s_{k,j}(\delta)=t_k+\frac{j}{M}\left[\min(t_{k+1},\,1-\delta)-t_k\right], \\[7pt] j=1,\ldots,M, \end{gathered}

where the terminal margin δ>0\delta>0 keeps the last query below t=1t=1, at which the policy is undefined. This supervision grid is used only during training and is distinct from the sampling grid, which always ends at tK=1t_K=1. The rollout states at the supervision times are

xˉ(k,j)=sg⁡ ⁣[xˉ(k)+sk,j−tk1−tk(x^1,θˉ(k)−xˉ(k))],\bar x^{(k,j)}=\operatorname{sg}\!\left[\bar x^{(k)}+\frac{s_{k,j}-t_k}{1-t_k} \left(\hat x_{1,\bar\theta}^{(k)}-\bar x^{(k)}\right)\right],

with sg⁡[⋅]\operatorname{sg}[\cdot] the stop-gradient. The online student computes its prediction x^1,θ(k)\hat x_{1,\theta}^{(k)} once at each detached segment start and reuses it for the segment's MM queries; the frozen teacher velocity vteav_{\mathrm{tea}} is queried at these states. Velocity matching penalizes, at each query, the residual

ℓv(sk,j)=∥πθ(k)(xˉ(k,j),sk,j)−vtea(xˉ(k,j),sk,j;c)∥m2,L~v=∑k=0K−1wkM∑j=1Mℓv(sk,j),\ell_v(s_{k,j})=\left\|\pi_\theta^{(k)}(\bar x^{(k,j)},s_{k,j})-v_{\mathrm{tea}}(\bar x^{(k,j)},s_{k,j};c)\right\|_m^2, \qquad \widetilde{\mathcal L}_{v}=\sum_{k=0}^{K-1}\frac{w_k}{M}\sum_{j=1}^{M}\ell_v(s_{k,j}),
ℓv(sk,j)=∥πθ(k)(xˉ(k,j),sk,j)−vtea(xˉ(k,j),sk,j;c)∥m2,L~v=∑k=0K−1wkM∑j=1Mℓv(sk,j),\begin{gathered} \ell_v(s_{k,j})=\left\|\pi_\theta^{(k)}(\bar x^{(k,j)},s_{k,j})-v_{\mathrm{tea}}(\bar x^{(k,j)},s_{k,j};c)\right\|_m^2, \\[7pt] \widetilde{\mathcal L}_{v}=\sum_{k=0}^{K-1}\frac{w_k}{M}\sum_{j=1}^{M}\ell_v(s_{k,j}), \end{gathered}

where wkw_k are segment weights and ∥⋅∥m2\|\cdot\|_m^2 is the squared error averaged over valid frames and channels. Rollout states and teacher targets are constants for the update, so gradients pass only through the online predictions at the segment starts. This construction fixes where supervision is obtained; the matching space sets how strongly each query counts.

Terminal amplification in velocity matching

The teacher's instantaneous velocity implies a clean-motion estimate, x^1tea(x,t;c)=x+(1−t) vtea(x,t;c)\hat x_1^{\mathrm{tea}}(x,t;c)=x+(1-t)\,v_{\mathrm{tea}}(x,t;c). Substituting it into the velocity residual at a supervision state xˉ\bar x and time ss gives

πθ(k)(xˉ,s)−vtea(xˉ,s;c)=x^1,θ(k)−x^1tea(xˉ,s;c)1−s⟹ℓv(s)=(1−s)−2 ℓx(s),\pi_\theta^{(k)}(\bar x,s)-v_{\mathrm{tea}}(\bar x,s;c) =\frac{\hat x_{1,\theta}^{(k)}-\hat x_1^{\mathrm{tea}}(\bar x,s;c)}{1-s} \qquad \Longrightarrow\qquad \ell_v(s)=(1-s)^{-2}\,\ell_x(s),
πθ(k)(xˉ,s)−vtea(xˉ,s;c)=x^1,θ(k)−x^1tea(xˉ,s;c)1−s⟹ℓv(s)=(1−s)−2 ℓx(s),\begin{gathered} \pi_\theta^{(k)}(\bar x,s)-v_{\mathrm{tea}}(\bar x,s;c) =\frac{\hat x_{1,\theta}^{(k)}-\hat x_1^{\mathrm{tea}}(\bar x,s;c)}{1-s} \\[7pt] \Longrightarrow\qquad \ell_v(s)=(1-s)^{-2}\,\ell_x(s), \end{gathered}

with ℓx(s)=∥x^1,θ(k)−x^1tea(xˉ,s;c)∥m2\ell_x(s)=\|\hat x_{1,\theta}^{(k)}-\hat x_1^{\mathrm{tea}}(\bar x,s;c)\|_m^2 the clean-motion loss: velocity matching is clean-motion matching with an inverse-square time-dependent weight. The objective is therefore

LS,b=∑k=0K−1wkM∑j=1Mbk,j∥x^1,θ(k)−sg⁡ ⁣[x^1tea(xˉ(k,j),sk,j;c)]∥m2,bk,j=(1−sk,j)−2,\mathcal L_{\mathcal S,b} =\sum_{k=0}^{K-1}\frac{w_k}{M}\sum_{j=1}^{M}b_{k,j} \left\|\hat x_{1,\theta}^{(k)}-\operatorname{sg}\!\left[\hat x_1^{\mathrm{tea}}(\bar x^{(k,j)},s_{k,j};c)\right]\right\|_m^2, \qquad b_{k,j}=(1-s_{k,j})^{-2},
LS,b=∑k=0K−1wkM∑j=1Mbk,j∥x^1,θ(k)−sg⁡ ⁣[x^1tea(xˉ(k,j),sk,j;c)]∥m2,bk,j=(1−sk,j)−2,\begin{gathered} \mathcal L_{\mathcal S,b} =\sum_{k=0}^{K-1}\frac{w_k}{M}\sum_{j=1}^{M}b_{k,j} \left\|\hat x_{1,\theta}^{(k)}-\operatorname{sg}\!\left[\hat x_1^{\mathrm{tea}}(\bar x^{(k,j)},s_{k,j};c)\right]\right\|_m^2, \\[7pt] b_{k,j}=(1-s_{k,j})^{-2}, \end{gathered}

where the segment weights wkw_k are a design choice but the coefficients bk,jb_{k,j} are not: they are induced by matching in velocity space.

Few-step FID with and without terminal amplification control.
(a) The known weight (1−s)−2(1-s)^{-2} on our supervision grid. (b) Bounding it improves FID; hollow bars enforce the same bound on the original grid (dots: seeds).

The last query of the supervision grid sits at sK−1,M=1−δs_{K-1,M}=1-\delta and receives the terminal coefficient bK−1,M=δ−2b_{K-1,M}=\delta^{-2}. On a uniform sampling grid every other coefficient is at most (MK)2(MK)^2, so a small numerical margin makes one coefficient exceed all others by orders of magnitude: δ=10−3\delta=10^{-3} gives 10610^6, while at K=8K=8 and M=2M=2 the other fifteen coefficients are at most 252252.

Because the supervision grid is reused at every update, this weight returns at the same supervision time in every update, instead of being met occasionally as under random time sampling. We call this persistent overweighting terminal amplification.

Terminal amplification control

TACD bounds every coefficient bk,jb_{k,j} by a value tied to the student's step size. It retains the sampling grid, the inference update, and the training budget, changes only the supervision used during training, and adds no network module or training stage. For a flow teacher it ties the terminal margin to the step Δt=1/K\Delta t=1/K of the uniform sampling grid tk=k/Kt_k=k/K:

δ=Δt2=12K⟹max⁡k,jbk,j=δ−2=4Δt2=4K2.\delta=\frac{\Delta t}{2}=\frac{1}{2K} \qquad \Longrightarrow\qquad \max_{k,j}b_{k,j}=\delta^{-2}=\frac{4}{\Delta t^{2}}=4K^2 .
δ=Δt2=12K⟹max⁡k,jbk,j=δ−2=4Δt2=4K2.\begin{gathered} \delta=\frac{\Delta t}{2}=\frac{1}{2K} \\[7pt] \Longrightarrow\qquad \max_{k,j}b_{k,j}=\delta^{-2}=\frac{4}{\Delta t^{2}}=4K^2 . \end{gathered}

The latest teacher query then sits at the midpoint of the final student step, so the teacher is never queried next to the singular time, while the sampling grid still ends at tK=1t_K=1. The terminal coefficient has the same order as the others on the grid, and the rule is closed-form in the step count, with no tuned threshold.

Two other enforcements meet the same bound at the original queries, each with a different loss. The three differ in how they meet the bound:

1 Grid capping TACD

δ=12K\delta=\dfrac{1}{2K}

Moves the final-segment queries and keeps the velocity loss unchanged. TACD's rule for every flow result.

2 Endpoint matching

bk,j≡1b_{k,j}\equiv 1

Removes the conversion factor at every query and matches the teacher's clean-motion estimates directly. The form TACD takes for diffusion teachers.

3 Weight capping

b=min⁡{(1−s)−2, 4K2}b=\min\{(1-s)^{-2},\,4K^2\}

Lowers one coefficient at the original queries; with M=2M=2 it changes a single term of the objective.

At K=8K=8 the bound is 4K2=2564K^2=256, and TACD queries the final segment at 0.906250.90625 and 0.93750.9375.

One training iteration, flow teacher

  1. Set δ←1/(2K)\delta\leftarrow 1/(2K) and build the supervision times sk,j(δ)s_{k,j}(\delta); sample noise xˉ(0)\bar x^{(0)} and text cc.
  2. Detached rollout. For k=0,…,K−1k=0,\ldots,K-1: the rollout EMA θˉ\bar\theta predicts x^1,θˉ(k)\hat x_{1,\bar\theta}^{(k)}, stores the states xˉ(k,j)\bar x^{(k,j)} for j=1,…,Mj=1,\ldots,M, and steps to xˉ(k+1)\bar x^{(k+1)}.
  3. Loss. For each kk: the online student predicts x^1,θ(k)\hat x_{1,\theta}^{(k)} at the detached xˉ(k)\bar x^{(k)}; for each jj, query the frozen teacher's estimate x^1tea(xˉ(k,j),sk,j;c)\hat x_1^{\mathrm{tea}}(\bar x^{(k,j)},s_{k,j};c) and add wkM ℓv(sk,j)\tfrac{w_k}{M}\,\ell_v(s_{k,j}) to the loss.
  4. Update θ\theta; update the rollout EMA θˉ\bar\theta and the separate evaluation EMA.

Only the online student receives gradients. The endpoint-matching control keeps the original supervision grid and uses ℓx\ell_x in place of ℓv\ell_v.

Extension to diffusion teachers

For a diffusion teacher the student again predicts clean motion at each segment start and learns from frozen-teacher targets on detached rollout states, with two changes: the states follow deterministic DDIM sub-trajectories, and the loss is the objective above with bk,j≡1b_{k,j}\equiv1. Let τ0>⋯>τK\tau_0>\cdots>\tau_K be the teacher sampler's KK-step timestep schedule with cumulative signal level αˉτ\bar\alpha_\tau, and x^0,θ(k)\hat x_{0,\theta}^{(k)} the student's clean-motion prediction at the segment start (x0x_0 is the flow case's x1x_1 under reversed time). Holding the rollout prediction x^0,θˉ(k)\hat x_{0,\bar\theta}^{(k)} fixed over the segment also fixes the noise ϵˉ(k)\bar\epsilon^{(k)} it implies, and the DDIM states at MM supervision times τk,j\tau_{k,j} within the segment are

xˉ(k,j)=sg⁡ ⁣[αˉτk,j  x^0,θˉ(k)+1−αˉτk,j  ϵˉ(k)].\bar x^{(k,j)}=\operatorname{sg}\!\left[\sqrt{\bar\alpha_{\tau_{k,j}}}\;\hat x_{0,\bar\theta}^{(k)} +\sqrt{1-\bar\alpha_{\tau_{k,j}}}\;\bar\epsilon^{(k)}\right].

The target is the frozen teacher's clean-motion estimate x^0tea\hat x_0^{\mathrm{tea}}, computed with the teacher sampler's guidance. Written in clean-motion space from the start, this loss carries no (1−s)−2(1-s)^{-2} factor; matching noise predictions on the same states would instead weight the clean-motion error by the signal-to-noise ratio, which also grows without bound toward the clean end. Endpoint matching thus carries no time-dependent weight, and for diffusion teachers TACD denotes this endpoint-matching form.

Results

HumanML3D and KIT-ML use dataset-specific Guo evaluators and real-motion references, averaged over 20 evaluator repetitions. The main flow student, TACD (HY-Motion), pairs the 0.46B motion backbone of the released HY-Motion Lite with frozen Qwen3-0.6B conditioning and a trainable projection. Training uses training-split captions as text only, with teacher supervision and no real-motion targets; evaluation uses the test split.

Main comparison

At eight steps with the HY-Motion teacher, TACD is best in R@1, R@3, MM-Dist, and FID.

MethodNFER@1 ↑R@3 ↑MM-Dist ↓FID ↓Div →MMod ↑
Real motions–.511.7972.974.0029.503–
Our evaluation: HY-Motion references
HY-Motion teacher50.455.7433.471.5419.133.28
HY-Motion Lite (released)50.435.7283.575.7268.95–
Truncated teacher8.378.6524.2253.2508.24–
Controlled Lite students, three-seed means: Qwen3-0.6B encoder, teacher-only supervision
Uncapped student8.402.7083.7032.4868.013.994
TACD8.479.7823.1921.0418.923.541
Non-adversarial ports to the Lite student, three-seed means: teacher-only supervision
DMD port8.426.7303.4931.8298.223.023
Consistency port8.394.6993.7542.9087.773.803
Reflow port8.370.6604.0633.3418.073.886

HumanML3D. NFE counts guided denoising steps per motion; Div →: closer to real is better. Bold: best of the eight-step HY-Motion rows.

At the same eight-step Lite budget, TACD reduces mean FID from 2.486 to 1.041 and improves R@3 from .708 to .782. It exceeds the R@3 of the 50-step teacher (.743) and of the released Lite at 50 steps (.728). The ports, our reimplementations of DMD, consistency distillation, and Reflow on the same Lite student, share its architecture, prompts, and generator-update budget. Under the MARDM evaluator, TACD leads the truncated teacher and every eight-step student in each seed and cuts the uncapped student's FID from 7.559 to 3.575.

Terminal control

Bounding the coefficient restores quality; restoring it reverses the gain.

At K=8K=8 and M=2M=2, the original supervision grid queries the final segment at .937 and .999: the terminal coefficient is 10610^6, while the other fifteen coefficients are at most 252. All three enforcements of the bound recover FID to 1.01–1.04. Conversely, restoring C=106C=10^6 on the capped grid, without moving any query, worsens FID in all three seeds, to a mean of 2.669.

ConfigurationFinal-segment
queries
CCFID ↓seed rangevs. uncapped
Uncapped student0.937, 0.99910610^62.486[2.419–2.563]–
1 TACD (grid capping)0.90625, 0.93752561.041[1.009–1.074]−58%
2 Endpoint matching0.937, 0.99911.029[0.988–1.091]−59%
3 Weight capping0.937, 0.9992561.011[0.980–1.044]−59%
Last-point-only0.937, 0.93752561.039[1.032–1.048]−58%
Capped grid, CC restored to 10610^60.90625, 0.937510610^62.669[2.645–2.703]+7%

Eight-step Lite student, HumanML3D; CC is the terminal coefficient. Each bound (1–3) cuts FID by 58–59%; restoring C=106C=10^6 (rose) undoes it. Brackets: three-seed range.

Three stacked curves against the terminal coefficient C: FID, R@3 and the median pre-clip gradient norm; all three are flat near the TACD value and degrade as C grows to the uncapped value.
Sweeping CC on the eight-step Lite student: TACD's 4K24K^{2} (dotted) sits on the plateau.

A dose–response. FID stays at 1.01–1.07 from C=256C=256 to 40964096 and rises monotonically beyond it, to 1.14, 1.30, 1.70, and 2.49 at C=16384C=16384, 6553665536, 262144262144, and 10610^6, while the median pre-clip gradient norm rises in step from 6 to 253. C=4K2C=4K^2 sits in a plateau at least 16×16\times wide.

Gradient concentration follows the grid. The original grid puts 96.6–97.6% of the gradient norm on the final segment and the capped grid 3.6–11.8%.

One query carries the effect. Keeping .937 and moving only .999 to .9375 gives FID 1.039, close to full capping's 1.041.

Key finding. In fixed-grid flow distillation, velocity matching can repeatedly overweight errors near the denoising endpoint. TACD controls this amplification with a step-size-aware bound, reducing eight-step HY-Motion student FID by 58% versus uncapped distillation on HumanML3D without changing inference. Restoring the full weight reverses the gain.

VLM judge

On SSAE, a VLM-judged semantic-alignment score, TACD scores highest.

ModelSSAE ↑
HY-Motion teacher @50.695
HY-Motion Lite (released) @50.692
Truncated teacher @8.587
Uncapped student @8.650
DMD port @8.672
Consistency port @8.700
Reflow port @8.590
TACD @8.720

SSAE: a VLM judge's yes-rate on questions about the prompted motion, on a fixed set of 640 prompts.

TACD scores highest (.720), above the 50-step teacher (.695), the released Lite (.692), and every other eight-step student (.587–.700).

Human study

In a blind ranking study, annotators rank eight-step TACD above its 50-step teacher in 51% of comparisons, on par with it.

ModelAbove teacher (%) ↑ [95% CI]Prompts W/T/L
HY-Motion teacher @5050 (reference)–
MotionLCM-V2 @441.2 [30.4, 51.7]17/0/33
MLCT @433.9 [24.5, 43.4]11/0/39
TACD @850.8 [41.1, 60.5]28/1/21

Win rate against the 50-step teacher and prompts won/tied/lost. 50 prompts, four systems rendered from joints, systems hidden and screen positions balanced; 17 annotators contributed 691 rankings.

By the same measure, the four-step MotionLCM-V2 and MLCT reach 41% and 34%.

Diffusion teachers

Four-step TACD has lower FID than its 50-step teacher and the released MotionLCM student.

MethodTraining
signal
NFER@1 ↑R@3 ↑MM-Dist ↓FID ↓Div →MMod ↑
MLD teacher and MotionLCM-V1
MLD teacher (DDIM)Reference50.502.7953.048.3879.53–
Truncated MLD teacherNone4.459.7683.295.5859.49–
MotionLCM-V1Real motions4–.798–.286––
TACDTeacher outputs4.525.8182.909.2229.80–
MotionLCM-V2 teacher and student
MotionLCM-V2 teacherReference50.537.8242.838.0669.59–
MotionLCM-V2 studentReal motions4.554.8412.760.0729.541.72
TACDTeacher outputs4.538.8242.838.0559.571.94

HumanML3D, our evaluation; bold: best at four steps in each group. TACD FID is the mean of three training seeds (±.002 for MLD, ±.001 for MotionLCM-V2).

With endpoint matching, four-step TACD improves on the 50-step MLD teacher in both FID (.387 → .222) and R@3 (.795 → .818), and its FID is below the .286 we measure for the released MotionLCM-V1 checkpoint, a student trained on real motions. With the MotionLCM-V2 teacher, four-step TACD reaches FID .055 against the teacher's .066 and the released official student's .072 under the same evaluation, at the teacher's R@3 (.824).

KIT-ML and Kimodo

On KIT-ML (MDM teacher) and Kimodo, TACD with the teacher's encoder beats the truncated teacher at every NFE.

Four panels against the number of sampling steps: FID and R@3 on KIT-ML, FID and R@1 on Kimodo; the TACD curve is better than the truncated teacher's at every step count.
Few-step curves: TACD (teal) and the truncated teacher (grey) at each NFE; dashed: the full-step teacher.

KIT-ML. At the same NFE, TACD reduces FID from the truncated teacher's 1.530, .976, and .745 to 1.038, .852, and .649 at two, four, and eight steps, 32% lower at two steps and 13% at four and eight.

MethodNFER@1 ↑R@3 ↑MM-Dist ↓FID ↓Div →
Real motions (GT)–.426.7802.786.02710.99
MDM teacher1000.403.7253.115.50610.74
Truncated teacher2.395.7283.1931.53011.24
TACD2.410.7562.9961.03811.06
Truncated teacher4.410.7553.014.97611.38
TACD4.416.7612.932.85211.16
Truncated teacher8.416.7562.948.74511.34
TACD8.421.7632.873.64911.08

KIT-ML, MDM teacher. TACD rows: three-seed means after 10k updates.

Kimodo. At eight steps, TACD raises R@1 over the truncated teacher from 69.36% to 71.39% and lowers FID from .064 to .054, with the same ordering at four and two steps.

MethodNFER@1 (%) ↑R@3 (%) ↑FID ↓Time (s) ↓
Teacher10073.1588.630.0452.96
Truncated teacher869.3684.350.0640.29
TACD871.3986.550.0540.29
Truncated teacher467.0582.370.0790.18
TACD468.4283.250.0670.18
Truncated teacher260.4677.270.1070.13
TACD262.3679.570.0950.13

Kimodo teacher, 607-caption TMR evaluation; TACD keeps the teacher's 8B text encoder. Time: end to end, one RTX L40.

Efficiency

With compact backbones and text encoders, eight-step students run 7.7–11.9× faster end to end and use 3.8–6.7× less peak GPU memory than their teachers.

Bar charts of parameters, peak memory and time per motion for the HY-Motion teacher, HY-Motion Lite, TACD (HY-Motion), the Kimodo teacher and TACD (Kimodo).
Deployment cost, RTX L40, batch 1; ×: vs. the teacher; @NN: NN steps.

HY-Motion. The Lite backbone and Qwen3-0.6B encoder of TACD (HY-Motion) reduce peak GPU memory from 17.63 to 2.62 GB (6.7×), mostly from the 8B-to-0.6B text encoder, and eight-step sampling cuts end-to-end time on an RTX L40 (bf16, batch 1) from the teacher's 2077 to 268 ms (7.7×).

Kimodo. TACD (Kimodo), with a trained-in 1B encoder, reduces peak memory from 15.3 to 4.0 GB (3.8×) and end-to-end time from 2.96 to .25 s (11.9×) at the R@3 and FID of the 8B-encoder student.

Models

The TACD student we release. It is trained from the frozen HY-Motion teacher and text prompts, with teacher supervision and no real-motion targets.

StudentTeacherBenchmarkStepsMain result
TACD (HY-Motion) coming soon0.46B Lite backbone, Qwen3-0.6B encoderHY-Motion, 50 stepsHumanML3D8FID 1.041 · R@3 .782

Loading with Hugging Face Transformers, once released:

import torch
from transformers import AutoModel

model = AutoModel.from_pretrained(
    "<ORG>/TACD-HY-Motion-Lite", trust_remote_code=True
).to("cuda")
out = model.generate(
    ["a person walks forward, turns around and waves"], duration=4.0, seed=0
)
out.joints   # (1, 120, 22, 3) joint positions at 30 fps

Release candidate: one of the three training seeds, with FID 1.0085, R@3 0.7829 and MM-Dist 3.1811 on HumanML3D; the row above is the three-seed mean.

HumanML3D: Guo evaluator, 20 repetitions.

BibTeX

@article{huang2026tacd,
  title   = {TACD: Distilling Efficient Text-to-Motion Models via Terminal Amplification Control},
  author  = {Huang, Wei-Jin and Li, Yuan-Ming and Lin, Kun-Yu and Luo, Wang and Zhu, Yinlin and
             Yu, Yue and Ye, Shenghao and Yuan, Junbin and Hong, Fa-Ting and Zhang, Qing and Zheng, Wei-Shi},
  journal = {arXiv preprint},
  year    = {2026}
}