The application discloses a sign
language recognition method and device based on cross-stage focus
distillation and multi-
branch time sequence learning, acquires a sign language video to be recognized, and inputs the sign language video to a trained sign
language recognition model to obtain a recognition result; the method and the device realize complementary deep and shallow features through a cross-stage focus
distillation unit, and simultaneously capture
time sequence dynamics of different granularities with the help of a multi-
branch time sequence learning unit, so that the overall modeling capability for key
semantics and action evolution in a
continuous flow of sign language is improved; the cross-stage focus
distillation unit guides the sign
language recognition model to focus on key discrimination areas in the sign language video through selective reinforcement of channel and spatial dimensions, effectively solving the contradiction between incomplete shallow feature
semantics and lost deep feature fine-grained information; and the multi-
branch time
sequence learning unit models short-time actions at the word level and long-range dependencies at the
sentence level synchronously through a parallel multi-branch
convolution structure, overcoming the deficiency of a traditional time sequence network in capturing multi-
granularity time dynamics.