Continuous sign language recognition method and system based on hierarchical attention feature fusion

By employing a hierarchical attention feature fusion and multi-level regularization framework, this approach addresses the shortcomings of existing continuous sign language recognition methods in multi-scale feature fusion, temporal boundary awareness, and knowledge distillation, thereby improving recognition accuracy and robustness, particularly in performance on small-scale datasets.

CN122135437BActive Publication Date: 2026-07-21TIANJIN POLYTECHNIC UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TIANJIN POLYTECHNIC UNIV
Filing Date
2026-04-30
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing continuous sign language recognition methods have shortcomings in multi-scale visual feature fusion, temporal boundary perception, knowledge distillation, and regularization strategies, resulting in insufficient recognition accuracy and robustness. They are particularly prone to overfitting and inaccurate word boundary segmentation on small datasets.

Method used

We adopt a hierarchical attention feature fusion approach, which uses a cross-scale aggregation module, a dynamic hierarchical attention mechanism, a boundary prediction auxiliary task, and an online self-distillation mechanism of EMA, combined with a multi-level regularization framework, to adaptively fuse multi-scale features, explicitly model sign language vocabulary boundaries, and improve the model’s generalization ability through efficient online self-distillation and systematic regularization.

Benefits of technology

It significantly improves the model's recognition accuracy and robustness in complex sign language scenarios, reduces word error rate, and improves recognition accuracy while maintaining computational efficiency, making it suitable for small-scale datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135437B_ABST
    Figure CN122135437B_ABST
Patent Text Reader

Abstract

The application provides a continuous sign language recognition method and system based on hierarchical attention feature fusion. The sign language video is output as multiple hierarchical multi-scale feature maps through a visual encoder, is sent to a classifier for vocabulary classification through a cross-scale aggregation module containing a hierarchical attention network, and a vocabulary sequence is output by a CTC decoder. Meanwhile, two-layer full connection networks are used for binary classification prediction to determine whether each frame is a boundary frame of a sign language vocabulary. An online self-distillation mechanism based on an exponential moving average (EMA) is established, an exponential moving average version of model parameters is used as an EMA teacher model, a student model continuously learns, and a more stable vocabulary sequence output is obtained. A multi-level regularization framework is constructed to cooperatively optimize from four levels of label space, time sequence space, feature space and parameter space. The application effectively improves the generalization ability of the model under small sample conditions, thereby reducing the word error rate (WER) of continuous sign language recognition.
Need to check novelty before this filing date? Find Prior Art