SignGPT: Toward LLM-Mediated Sign Language Interaction through Gloss-Free Translation and Generation

A unified pose-based framework for bidirectional sign language understanding, generation, and exploratory sign-to-sign interaction.

About the paper

One model for sign language translation, generation, and conversation

Abstract

Large language models (LLMs) provide limited support for sign language interaction. Unifying sign language translation (SLT) and generation (SLG) to enable sign language as both input and output can reduce switching between models during sign–text interaction. We present SignGPT, a unified, pose-based framework for gloss-free SLT and SLG. SignGPT integrates part-aware hierarchical representations of body, hand, and facial motion into a shared language model and employs asymmetric multi-token prediction and progressive training for bidirectional modeling. We evaluate SignGPT on How2Sign (ASL) and Phoenix-2014T (DGS) through benchmark comparisons, qualitative analyses, and component ablations. An exploratory study with 12 Deaf ASL signers assesses an LLM-mediated sign-to-sign response pipeline, highlighting the potential of unified modeling to support sign language conversation (SLC).

SignGPT translation, generation, and one-turn sign-to-sign response pipeline
SignGPT overview. A central SignGPT model connects translation and generation. On the left, a text instruction is converted into a sequence of signing avatars. In the middle, an input signing sequence is translated into text. On the right, a signed question is translated, answered in English, and converted into a signed response, illustrating the one-turn response pipeline.
Contributions

Connecting representation, bidirectional modeling, and interaction

SignGPT is designed and evaluated as a unified framework rather than as isolated translation and generation systems.

Unified sign-text modeling

We present SignGPT, a unified pose-based framework that connects gloss-free sign-to-text translation and text-to-sign generation through part-aware motion quantization, shared hidden states, and heterogeneous prediction heads.

Benchmarks and ablations

We provide benchmark comparisons and component ablations on ASL and DGS datasets, separately evaluating motion reconstruction quality and task-level translation and generation performance.

Exploratory interaction study

We construct an exploratory LLM-mediated sign-to-sign response pipeline and examine raters' perceptions of response appropriateness and motion smoothness, characterizing its opportunities and current limitations.

Method

Three-stage progressive training

Part-aware motion tokenization supplies structured sign representations to a shared language model, which is then adapted for both mapping directions and prompted task use.

The three training stages of SignGPT: PHVQ tokenization, bidirectional task pre-training, and instruction fine-tuning
Figure 2. Overview of SignGPT's three training stages. (1) PHVQ discretizes continuous body and hand motion, with an optional text-alignment objective; (2) GHMLM is jointly trained on SLG and SLT and uses AMTP to decode text or part-aware motion tokens; and (3) instruction fine-tuning supports prompted translation and generation.
01

Part-aware tokenization

PHVQ models a body-face stream and two hand streams, using body-to-hand conditioning and bidirectional multiscale temporal encoding.

02

Bidirectional joint training

GHMLM reuses PHVQ features as motion embeddings and predicts either text or synchronized part-aware motion tokens from shared hidden states.

03

Instruction fine-tuning

Progressive LoRA-based training equips each dataset-specific model to follow prompted sign-to-text and text-to-sign instructions.

Experiments

Benchmark results on DGS and ASL

Table 1. Sign language generation (SLG) on Phoenix-2014T and How2Sign

SignGPT's part-aware representation substantially reduces aggregate DTW-JPE relative to the matched MotionGPT* adaptation. The optional text-space alignment variant gives small additional gains.

MethodBi-TGlossInputPhoenix-2014T · DTW ↓How2Sign · DTW ↓
PoseRGBBodyHandAvgBodyHandAvg
MotionGPT*8.9710.149.459.4310.969.82
SOKE2.585.894.267.923.075.49
T2S-GPT7.329.868.287.1511.218.49
MoMask3.556.755.38
NSA3.096.805.417.837.337.44
SignGPT+TSA (Ours)3.314.984.324.984.114.76
SignGPT (Ours)3.355.074.365.054.184.82

DTW-JPE is reported for the body, hands, and the aggregate labeled “Avg” in each implementation. Bi-T indicates support for both sign-to-text and text-to-sign directions within one dataset-specific model. SignGPT+TSA uses the optional text-embedding alignment objective with paired sentence-level translations during PHVQ training. MotionGPT* results are reproduced using our matched sign-language adaptation; all other baseline values are transcribed from the source studies cited for the corresponding results. Bold and underline mark the best and second-best values.

Table 2. Sign language translation (SLT) on Phoenix-2014T and How2Sign

SignGPT remains gloss-free and bidirectional while achieving the strongest listed How2Sign translation scores. Phoenix-2014T comparisons include methods with different input modalities and task-specific priors.

MethodBi-TGlossInputPhoenix-2014T · Dev / TestHow2Sign · Test
PoseRGBDev B4 ↑Dev R ↑Test B4 ↑Test R ↑B4 ↑R ↑
MotionGPT*11.5327.1410.9827.058.7228.61
SLTCC11.831.1
SLT20.6945.5420.1745.34
CSGCR15.0838.9615.1838.85
Uni-Sign14.936.0
SignLLM25.2547.2323.4044.49
MixSignGraph24.8751.7124.0251.1410.4128.01
SignGPT+TSA (Ours)24.1742.9623.5642.1316.7738.25
SignGPT (Ours)23.5841.7222.9541.2816.4237.69

We report BLEU-4 (B4) and ROUGE-L (R). Bi-T indicates support for both sign-to-text translation and text-to-sign generation within one dataset-specific model. MotionGPT* results are reproduced using our matched sign-language adaptation; all other baseline values are transcribed from the source studies cited for the corresponding results.

Exploratory User Study

How Deaf ASL signers rated one-turn responses

We separately evaluate whether each response plausibly answers the signed question and whether its frame-to-frame motion is smooth.

12Deaf ASL users, ages 20–30
1,000How2Sign question samples
3Different raters per sample
12,000Total scalar ratings

Table 4. End-to-end response relevance and exploratory subjective ratings for the single-turn ASL response pipeline

System labels were concealed and response order was randomized. Each output was rated independently on response appropriateness and motion smoothness.

MethodResponse relevance
LLM-AR (%) ↑
Subjective pilot · 1–5
Response appropriateness ↑Motion smoothness ↑
MotionGPT*15.72.711.26
SignGPT (Ours)52.23.674.35

LLM-AR is the percentage of 1,000 outputs judged relevant by GPT-4o after shared frozen SLT back-translation. Each subjective mean summarizes 1,000 three-rater sample averages (3,000 raw ratings per system and item).

SignGPT's mean motion-smoothness rating (4.35) is higher than its response-appropriateness rating (3.67), indicating that temporal coherence was perceived more strongly than semantic appropriateness. The results are descriptive rather than population-level estimates.

Participants and allocation

Participants reported at least 10 years of ASL use and completed the remote study during a one-month window. Each participant received 250 samples. Every selected question was assigned to exactly three different participants, yielding 3,000 participant-sample assignments.

Participant-facing instructions

View the signed question (together with its corresponding translation) and two rendered signed responses, then rate each response separately. For response appropriateness, focus on whether the response plausibly answers the question and ignore rendering artifacts. For motion smoothness, focus only on temporal continuity and ignore semantic correctness.

Item-level rating questions

No preference question was used; both responses were rated independently on both items.

Response appropriateness1: no recognizable relation · 2: mostly unrelated · 3: partially related or ambiguous · 4: clearly related and mostly answers · 5: clearly and properly responds.
Motion smoothness1: severe jitter or broken motion · 2: frequent jitter or disrupted transitions · 3: some noticeable jitter but generally continuous · 4: mostly smooth with minor artifacts · 5: coherent with no perceptible jitter.

Each participant completed 250 trials, covering 500 response outputs and 1,000 scalar ratings.

01

Sign Language Generation

Given one shared prompt, compare the ground-truth sign, MotionGPT* output, and SignGPT output. Facial parameters are also included in the visualization. Videos use a common frame clock; a shorter sequence holds on its final frame until the longest sequence finishes.* Adapted model.

Prompt text
Ground Truth
Last frame held
MotionGPT*
Last frame held
SignGPTOurs
Last frame held
0 / 0 frames
02

Sign Language Translation

One sign input is translated by three targets. Meaning-preserving words and paraphrases are shown in green; only meaning-changing errors are highlighted in red.

Semantically consistentMeaning-changing error
Sign motionInput
Ground TruthReference
MotionGPT*Generated text
SignGPTGenerated text
03

Sign Language Conversation

The same question sign branches into two pipelines: sign-to-text question understanding, LLaMA response generation, and text-to-sign answer synthesis. The animation reveals each stage in order.

Ready

Shared input · Question sign

MotionGPT*
Step 1 · Sign → Text

Recognized question

Step 2 · LLaMA

Generated answer

Step 3 · Text → Sign

Back translation

SignGPTOurs
Step 1 · Sign → Text

Recognized question

Step 2 · LLaMA

Generated answer

Step 3 · Text → Sign

Back translation

Step 3 · Generated sign