Sign language translation method and system based on deep learning
By using the lightweight deep learning model MediaPipe and the multi-stream key point attention model, combined with the mBART large language model, the problem of slow sign language translation technology is solved, and efficient and accurate sign language translation is achieved.
Patent Information
- Application Number
- CN202510036051.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-09
AI Technical Summary
The existing sign language translation technology based on computer vision is difficult to guarantee the speed of sign language translation due to the large amount of calculations.
The lightweight deep learning model MediaPipe is used to identify key point information in sign language videos, build a multi-stream key point attention model, analyze the key point characteristics through the attention mechanism, and input it into the mBART large language model for translation.
It improves the speed and accuracy of sign language translation and achieves real-time translation effect.
Smart Images

Figure CN119964239A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning, and in particular to a sign language translation method and system based on deep learning. Background Art
[0002] At present, the existing sign language translation technology is based on a large-scale sign language database. It collects sign language videos from different countries, regions, dialects and personal habits, and shoots sign language videos from multiple angles to capture all-round information of gestures. Through the database structure, it can store and quickly retrieve gesture features and corresponding meanings.
[0003] However, existing sign language translation based on computer vision uses a camera to collect sign language videos, and then uses image processing and computer vision technology to analyze gestures. In recent years, with the development of deep learning technology, more and more methods use 3D convolution to process sign language videos, but the amount of calculation required has increased significantly, making it difficult to ensure the speed of sign language translation. Summary of the invention
[0004] The present invention provides a sign language translation method and system based on deep learning, which is used to solve the above-mentioned problem existing in the prior art, that is, how to improve the sign language translation speed in the prior art. The present invention provides a sign language translation method based on deep learning, which comprises:
[0005] The lightweight deep learning model MediaPipe is used to identify key points of the left hand, right hand, face, and the whole body in the acquired sign language video to obtain key point information;
[0006] Build a multi-stream keypoint attention model, including four keypoint attention networks;
[0007] According to the key point information, the key point sequence is obtained, and the key point sequence is decoupled into left-hand flow, right-hand flow, face flow and whole-body flow. Based on the attention mechanism, four key point attention networks are used to perform correlation analysis on the left-hand flow, right-hand flow, face flow and whole-body flow respectively, and the key point features corresponding to the left-hand flow, right-hand flow, face flow and whole-body flow are determined; the key point features corresponding to the left-hand flow, right-hand flow, face flow and whole-body flow are pooled and fused to generate sign language intermediate annotations;
[0008] Input the sign language intermediate annotations into the trained mBART large language model to obtain the translated text.
[0009] Also includes:
[0010] The acquired dataset for sign language translation is subjected to temporal augmentation, key point rotation, and key point translation operations, and the key point attention network is trained by using the dataset.
[0011] Optionally, the step of performing time sequence enhancement, key point rotation and key point translation operations on the acquired data set for sign language translation specifically includes:
[0012] By adopting a certain range of time series scaling ratios, the key point information in the data set is time-enhanced. By using a rotation matrix to obtain the new coordinates of all points after rotation around the origin, the key point information of the data set is rotated and enhanced. By adding the offset matrix to the original key point information, the key points are translated up, down, left, and right.
[0013] Optionally, the key point features corresponding to the left hand stream, the right hand stream, the face stream and the whole body stream are pooled and fused to generate the sign language intermediate annotation, specifically including:
[0014] Pooling and fusion operations are performed on the key point features corresponding to the left-hand flow, right-hand flow, face flow, and whole-body flow to determine the feature set;
[0015] After determining the feature set, distillation loss operations are performed on the left-hand flow, right-hand flow, face flow, and whole-body flow to obtain the prediction results of each key point flow and generate sign language intermediate annotations.
[0016] It also includes identifying key points of the left hand, right hand, face and the whole human body in the acquired sign language video, and after acquiring the key point information, performing tracking missing detection on the key point information, specifically including:
[0017] Initialize the key point value and determine whether all key point positions of the first frame of the key point data are detected. When the key point position of a certain frame is not detected, detect the key point information of the previous frame corresponding to the frame. If the key point position of the current frame is not 0, the key point information of the previous frame is assigned to the frame until the key point position of the frame is detected.
[0018] The method also includes performing a model pruning operation on the mBART large language model, specifically including:
[0019] The features obtained by fusing the left-hand stream, right-hand stream, facial stream and whole-body stream are matched with the intermediate annotations to adapt to the target language.
[0020] The present invention provides a sign language translation system based on deep learning, comprising:
[0021] The acquisition module is used to use the lightweight deep learning model MediaPipe to identify the key points of the left hand, right hand, face and the whole body in the acquired sign language video to obtain the key point information;
[0022] A building block for building a multi-stream keypoint attention model, including four keypoint attention networks;
[0023] The analysis and fusion module is used to obtain the key point sequence according to the key point information, decouple the key point sequence into left-hand flow, right-hand flow, face flow and whole-body flow, and perform correlation analysis on the left-hand flow, right-hand flow, face flow and whole-body flow respectively through four key point attention networks based on the attention mechanism to determine the key point features corresponding to the left-hand flow, right-hand flow, face flow and whole-body flow; perform pooling and fusion operations on the key point features corresponding to the left-hand flow, right-hand flow, face flow and whole-body flow to generate sign language intermediate annotations;
[0024] The translation module is used to input the sign language intermediate annotations into the trained mBART large language model to obtain the translated text.
[0025] Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention provides a sign language translation method based on deep learning. The method is based on the attention mechanism and constructs a multi-stream key point attention model. The model decouples the key point sequence into four streams, namely the left-hand stream, the right-hand stream, the face stream and the whole-body stream. When performing sign language translation, the continuity of facial expressions, body postures and movements can be obtained, and then by fusing different types of features, the sign language information that the translator wants to convey can be more comprehensively obtained and analyzed, thereby improving the accuracy of sign language translation. When performing sign language translation, the key point information can be directly input into the sign language translation model to obtain the translated text, achieving the effect of real-time translation and effectively improving the translation speed of sign language translation. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0027] Figure 1 A flowchart of a sign language translation method based on deep learning provided by an embodiment of the present invention;
[0028] Figure 2 A diagram of a multi-stream key point attention model provided for an embodiment of the present invention. DETAILED DESCRIPTION
[0029] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0030] The technical solution of the present invention and how the technical solution of the present invention solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below in conjunction with the accompanying drawings.
[0031] Figure 1 is a flow chart of a sign language translation method based on deep learning provided by an embodiment of the present invention. Figure 1 As shown, this embodiment shows a sign language translation method based on deep learning, including:
[0032] S1: Use the lightweight deep learning model MediaPipe to identify key points of the left hand, right hand, face, and the whole body in the acquired sign language video to obtain key point information.
[0033] Optionally, use the MediaPipe model to detect human hands, face, and torso to obtain the coordinates and confidence of each key point in the image.
[0034] In this step, the MediaPipe algorithm can be used to quickly identify key points of the human body, effectively improving the speed of subsequent sign language translation.
[0035] S2: Construct a multi-stream keypoint attention model, including four keypoint attention networks.
[0036] S3: Obtain the key point sequence according to the key point information, decouple the key point sequence into left-hand flow, right-hand flow, face flow and whole-body flow, and based on the attention mechanism, perform correlation analysis on the left-hand flow, right-hand flow, face flow and whole-body flow through four key point attention networks to determine the key point features corresponding to the left-hand flow, right-hand flow, face flow and whole-body flow; perform pooling and fusion operations on the key point features corresponding to the left-hand flow, right-hand flow, face flow and whole-body flow to generate sign language intermediate annotations.
[0037] Exemplarily, the key point sequence is decoupled into four streams, namely the left-hand stream, the right-hand stream, the face stream and the whole-body stream. Each stream focuses on one aspect of the key point sequence. By concatenating and fusing the features of the left-hand, right-hand, face and whole-body streams, the model can better learn the whole-body stream and the local feature stream, and can have a more comprehensive understanding of sign language recognition and translation. The results of each stream can be added together to obtain a set result, and the obtained set is used to guide the distillation operation of each stream to make the predictions of each stream more similar.
[0038] Optionally, the acquired dataset for sign language translation is subjected to temporal augmentation, key point rotation, and key point translation operations, and a key point attention network is trained by using the dataset.
[0039] Optionally, the key point information in the data set is time-enhanced by adopting a certain range of time scaling ratios, the rotation data of the key point information in the data set is enhanced by adopting a rotation matrix to obtain the new coordinates of all points after rotation around the origin, and the key point information is translated up, down, left, and right by adding the offset matrix to the original key point information.
[0040] For example, when enhancing the timing, a timing scaling ratio of 0.5X-1.5X may be selected to allow the input data to have more features.
[0041] Exemplarily, during network training, HRNet is used to extract key point information, and MediaPipe is used to extract key points when the model is deployed in the system. The two models have different key point generation quantities and positions. In order to avoid the above-mentioned problem that leads to the loss of key point tracking, thereby affecting the accuracy of the translation results, the present invention determines whether all key point positions of the first frame of the acquired key point information are detected; if the key point position disappears in a certain frame, the key point information of the previous frame is queried, and if the key point is detected in the previous frame, the data of the previous frame is assigned to the frame until the frame is re-detected, so that the key points generated by MediaPipe can be aligned to the trained multi-stream key point attention model, thereby solving the problem of partial key point recognition loss, thereby enhancing the accuracy of subsequent sign language translation.
[0042] For example, after the key point information is enhanced by translation data by introducing an offset matrix and adding it to the original key point information, the key point information can be dimensionally upgraded through 2D convolution, and then attention calculation is performed, and then subsequent pooling and multi-stream fusion operations are performed.
[0043] like Figure 2 As shown in the figure, after obtaining the key point information, the key point dimension is increased to 64 dimensions through the dimension-up operation, and then the information after dimension-up is decoupled and passed through the whole body stream key point attention network, the right hand stream key point attention network, the left hand stream key point attention network and the face stream key point attention network. After passing through the multi-layer attention network, the key point features obtained are pooled and classified, and the whole body classification network, the right hand classification network, the left hand classification network and the fusion classification network are used for CTC loss supervision, which only contains the time dimension and the feature dimension, and eliminates the key point dimension; the results of each stream of the left hand and face, the right hand and face, the whole, the left hand, the face, the right hand and the whole are fused to obtain the set result, and the obtained set is used to guide each stream to perform the distillation loss operation, in which the gradient calculation is prohibited for the set to make the prediction of each stream more similar.
[0044] Exemplarily, the key point information is input into a multi-stream key point attention model, and a convolution operation is performed on the key point information. Based on the convolved key point information, the key point features are obtained through the attention mechanism. The key point features corresponding to the left-hand stream, right-hand stream, facial stream, and whole-body stream can be pooled and fused to determine the feature set. After determining the feature set, a distillation loss operation is performed on the left-hand stream, right-hand stream, facial stream, and whole-body stream to obtain the prediction results of each key point stream, and generate sign language intermediate annotations.
[0045] S4: Input the sign language intermediate annotation into the trained mBART large language model to obtain the translated text.
[0046] Exemplarily, the intermediate annotations are input into the mBART model, and the encoder and decoder of the mBART model are used to obtain the translated text.
[0047] Exemplarily, the mBART large language model may be subjected to a model pruning operation, specifically including: matching features obtained by fusing the left-hand stream, the right-hand stream, the facial stream, and the whole body stream with intermediate annotations to adapt to the target language.
[0048] The above is a sign language translation method based on deep learning provided by one or more embodiments of this specification. Based on the same idea, this specification also provides a corresponding sign language translation system based on deep learning, including:
[0049] The acquisition module is used to use the lightweight deep learning model MediaPipe to identify the key points of the left hand, right hand, face and the whole body in the acquired sign language video to obtain the key point information;
[0050] A building block for building a multi-stream keypoint attention model, including four keypoint attention networks;
[0051] The analysis and fusion module is used to obtain the key point sequence according to the key point information, decouple the key point sequence into left-hand flow, right-hand flow, face flow and whole-body flow, and perform correlation analysis on the left-hand flow, right-hand flow, face flow and whole-body flow respectively through four key point attention networks based on the attention mechanism to determine the key point features corresponding to the left-hand flow, right-hand flow, face flow and whole-body flow; perform pooling and fusion operations on the key point features corresponding to the left-hand flow, right-hand flow, face flow and whole-body flow to generate sign language intermediate annotations;
[0052] The translation module is used to input the sign language intermediate annotations into the trained mBART large language model to obtain the translated text.
[0053] For the specific limitations of the sign language translation system based on deep learning, please refer to the limitations of the sign language translation method based on deep learning above, which will not be repeated here. Each module in the above-mentioned sign language translation system based on deep learning can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0054] The technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of the present invention.
Claims
1. A sign language translation method based on deep learning, characterized in that: include: The lightweight deep learning model MediaPipe is used to identify key points of the left hand, right hand, face, and the whole body in the acquired sign language video to obtain key point information; Build a multi-stream keypoint attention model, including four keypoint attention networks; According to the key point information, the key point sequence is obtained, and the key point sequence is decoupled into left-hand flow, right-hand flow, face flow and whole-body flow. Based on the attention mechanism, four key point attention networks are used to perform correlation analysis on the left-hand flow, right-hand flow, face flow and whole-body flow respectively, and the key point features corresponding to the left-hand flow, right-hand flow, face flow and whole-body flow are determined; the key point features corresponding to the left-hand flow, right-hand flow, face flow and whole-body flow are pooled and fused to generate sign language intermediate annotations; Input the sign language intermediate annotations into the trained mBART large language model to obtain the translated text.
2. The sign language translation method based on deep learning as claimed in claim 1, characterized in that: The method further comprises: The acquired dataset for sign language translation is subjected to temporal augmentation, key point rotation, and key point translation operations, and the key point attention network is trained by using the dataset.
3. The sign language translation method based on deep learning as claimed in claim 2, characterized in that: The step of performing time sequence enhancement, key point rotation and key point translation operations on the acquired data set for sign language translation specifically includes: By adopting a certain range of time series scaling ratios, the key point information in the data set is time-enhanced. By using a rotation matrix to obtain the new coordinates of all points after rotation around the origin, the key point information of the data set is rotated and enhanced. By adding the offset matrix to the original key point information, the key points are translated up, down, left, and right.
4. The sign language translation method based on deep learning as claimed in claim 1, characterized in that: The key point features corresponding to the left hand stream, right hand stream, face stream and whole body stream are pooled and fused to generate the sign language intermediate annotation, specifically including: Pooling and fusion operations are performed on the key point features corresponding to the left-hand flow, right-hand flow, face flow, and whole-body flow to determine the feature set; After determining the feature set, distillation loss operations are performed on the left-hand flow, right-hand flow, face flow, and whole-body flow to obtain the prediction results of each key point flow and generate sign language intermediate annotations.
5. The sign language translation method based on deep learning as claimed in claim 1, characterized in that: It also includes identifying key points of the left hand, right hand, face and the whole human body in the acquired sign language video, and after acquiring the key point information, performing tracking missing detection on the key point information, specifically including: Initialize the key point value and determine whether all key point positions of the first frame of the key point data are detected. When the key point position of a certain frame is not detected, detect the key point information of the previous frame corresponding to the frame. If the key point position of the current frame is not 0, the key point information of the previous frame is assigned to the frame until the key point position of the frame is detected.
6. The sign language translation method based on deep learning as claimed in claim 1, characterized in that: The method also includes performing a model pruning operation on the mBART large language model, specifically including: For the features obtained by fusing the left-hand stream, right-hand stream, facial stream and whole-body stream, the features and intermediate annotations are matched to adapt to the target language.
7. A sign language translation system based on deep learning, characterized in that: include: The acquisition module is used to use the lightweight deep learning model MediaPipe to identify the key points of the left hand, right hand, face and the whole body in the acquired sign language video to obtain the key point information; A building block for building a multi-stream keypoint attention model, including four keypoint attention networks; The analysis and fusion module is used to obtain the key point sequence according to the key point information, decouple the key point sequence into left-hand flow, right-hand flow, face flow and whole-body flow, and perform correlation analysis on the left-hand flow, right-hand flow, face flow and whole-body flow respectively through four key point attention networks based on the attention mechanism to determine the key point features corresponding to the left-hand flow, right-hand flow, face flow and whole-body flow; perform pooling and fusion operations on the key point features corresponding to the left-hand flow, right-hand flow, face flow and whole-body flow to generate sign language intermediate annotations; The translation module is used to input the sign language intermediate annotations into the trained mBART large language model to obtain the translated text.
Citation Information
Cited By
Multi-modal fusion self-adaptive sign language digital character generation method and multi-modal fusion self-adaptive sign language digital character generation system
CN120876686A
Sign language translation model, system and method based on full-modal alignment
CN120894834A