The invention provides a sign
language translation method and
system based on a pre-training
diffusion large
language model, and belongs to the field of sign language video translation. The method comprises the following steps: preprocessing a video containing sign language actions to obtain a sign language video
frame sequence, inputting the sign language video
frame sequence into a visual
feature extraction network to extract features, and fusing to obtain a
time sequence visual fusion feature sequence; giving a text cue word of a sign
language translation task, constructing an initial
mask sequence for a target translation position, taking the text cue word, the
time sequence visual fusion feature sequence and the initial
mask sequence as guide conditions, injecting the guide conditions into a
diffusion language model, iteratively denoising and predicting lexical elements of a masked position in combination with a
diffusion mask mechanism, and obtaining the sign
language translation task. A
natural language translation sequence is obtained, and sign language translation is completed; wherein when the diffusion
language model is trained, through an internal feature alignment mechanism, the guiding effect of guiding conditions on
text generation is optimized, so that the accuracy, coherence and robustness of long text translation are improved, and the actual requirements of a barrier-free public service scene are better met.