MethodThree-stage progressive training
Part-aware motion tokenization supplies structured sign representations to a shared language model, which is then adapted for both mapping directions and prompted task use.
01Part-aware tokenization
PHVQ models a body-face stream and two hand streams, using body-to-hand conditioning and bidirectional multiscale temporal encoding.
02Bidirectional joint training
GHMLM reuses PHVQ features as motion embeddings and predicts either text or synchronized part-aware motion tokens from shared hidden states.
03Instruction fine-tuning
Progressive LoRA-based training equips each dataset-specific model to follow prompted sign-to-text and text-to-sign instructions.