The invention discloses a skeleton sign
language recognition method of a double-flow space-time dynamic graph convolutional network fused with residual learning, and belongs to the technical field of
artificial intelligence and
gesture recognition. According to the method, an input gesture skeleton sequence relative to a face is divided into two data streams, namely a hand form
data stream and a
wrist track
data stream through double-reference-
system differential homeomorphic mapping; the method comprises the following steps: firstly,
processing hand posture data, and capturing a
hand joint spatial topological relation by combining a spatial-temporal dynamic graph convolutional network (STDGCNN) with a residual convolutional block; meanwhile, a Finsler trajectory dynamics
encoder (FTDE) is adopted to carry out differential
geometric modeling on the
wrist trajectory, and the direction sensitivity characteristic of the trajectory is captured through a multi-scale causal convolutional network. Then, mutual enhancement of double-flow features is realized through a bidirectional cross feature
enhancer (BCFE), and the problem of geometric inconsistency of a heterogeneous feature space is solved through a geometric-driven optimal transmission fusion device (Geometric-OT). The method solves the challenge that a traditional sign
language recognition method processes complex space-
time correlation of gesture forms and motion tracks at the same time, the technical problems of insufficient relation between hands and faces, insufficient feature expression ability and low space-time
feature extraction efficiency, and the problems of geometric inconsistency, single reference
system and the like. And the identification accuracy and the real-time performance are obviously improved. Experiments show that the accuracy rate of the method in complex hundreds of sign language vocabulary recognition tasks reaches 95% or above on average, the reasoning speed is only 17ms on average, high-precision real-time sign
language recognition is achieved, and the method has higher robustness in complex environments such as
noise and shielding.