Sign Language Translation Method, Electronic Device, and Storage Medium
By extracting the keyframes of sign language videos and performing segmentation processing, and generating feature map sequences, the problem of low sign language translation speed and accuracy is solved, and a more efficient sign language translation effect is achieved.
Patent Information
- Application Number
- CN202410309773.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-18
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-03-18
AI Technical Summary
In the prior art, the speed and accuracy of sign language translation are low, mainly due to the redundancy and complex background of sign language videos, the difficulty of keyframe extraction, and the unknown semantic boundaries of continuous sign language, which improves the difficulty of segmentation gestures.
By obtaining the question text and the sign language video to be translated, the keyframes in the video are determined, and the keyframes are segmented to generate a feature chart sequence. Enter the question text and feature chart sequence into the pre-trained sign language translation model to generate sign language translation text.
Effectively reduce fuzzy frames, solve the natural redundancy problem of continuous frames, reduce background interference, and improve the speed and accuracy of sign language translation.
Smart Images

Figure CN118230412B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology. Specifically, it relates to a sign language translation method, an electronic device, and a storage medium. Background Art
[0002] For some special groups, it is very important to translate sign language into text for communication with others. For example, the action gestures in a sign language video can be converted into text information through vision technology and natural language processing technology. Specifically, sign language translation can be based on a sign language translation model.
[0003] Most traditional sign language translation models map sign language videos to corresponding translation texts based on technologies such as convolutional neural networks, recurrent convolutional neural networks, and multi-cue collaborative sign language translation. However, sign language translation faces two major challenges. First, sign language videos are redundant, and there will also be a large number of blank frames and interference from complex backgrounds, which makes it difficult to extract key frames. Second, the semantic boundaries of continuous sign language are unknown. Sign language has a rich vocabulary and many terms have similar gestures and movements, and each person has different speeds of gestures, so it increases the difficulty of segmenting gestures, thereby reducing the speed and accuracy of sign language translation. Summary of the Invention
[0004] The purpose of this application is to provide a sign language translation method, an electronic device, and a storage medium for the deficiencies in the above-mentioned existing technologies, so as to solve the problem of low speed and accuracy of sign language translation in the existing technologies.
[0005] To achieve the above purpose, the technical solutions adopted in the embodiments of this application are as follows:
[0006] In a first aspect, an embodiment of this application provides a sign language translation method, and the method includes:
[0007] Obtain a question text and a sign language video to be translated, where the sign language video is a sign language video formed by the user's gesture answer based on the question text;
[0008] Determine the key frames in the sign language video;
[0009] Perform segmentation processing on the key frames to obtain a segmentation result, and determine a sequence of feature maps based on the segmentation result;
[0010] Input the question text and the sequence of feature maps into a pre-trained sign language translation model to obtain a sign language translation text corresponding to the sign language video.
[0011] As a possible implementation, the determining the key frames in the sign language video includes:
[0012] Perform frame extraction on the sign language video to obtain multiple initial video frames;
[0013] Perform hand detection processing on each of the initial video frames, and screen out multiple target video frames with hand features from the multiple initial video frames according to the results of the hand detection processing;
[0014] According to the frame extraction order corresponding to each target video frame, perform key point detection processing on each target video frame respectively, and determine the key frame according to the key point detection results.
[0015] As a possible implementation, the performing key point detection processing on each target video frame respectively and determining the key frame according to the key point detection results includes:
[0016] Perform key point detection on the area with hand features in each target video frame to obtain key point detection results, where the key point detection results include the change range of the limb angle and the change range of the hand center point coordinates corresponding to each target video frame;
[0017] Determine the key frame according to the change range of the limb angle corresponding to each target video frame and the preset angle change threshold, and the change range of the hand center point coordinates corresponding to each target video frame and the preset center point coordinate change threshold.
[0018] As a possible implementation, the performing key point detection on the area with hand features in each target video frame to obtain key point detection results includes:
[0019] Determine the change range of the limb angle corresponding to the current target video frame according to the limb angle information of the previous target video frame of the current target video frame and the limb angle information of the current target video frame;
[0020] Determine the change range of the hand center point coordinates corresponding to the current target video frame according to the hand center point coordinates in the previous target video frame and the hand center point coordinates in the current target video frame, where the current target video frame is any target video frame other than the first target video frame among the multiple target videos.
[0021] As a possible implementation, the determining the key frame according to the change range of the limb angle corresponding to each target video frame and the preset angle change threshold, and the change range of the hand center point coordinates corresponding to each target video frame and the preset center point coordinate change threshold includes:
[0022] If the change range of the limb angle corresponding to the target video frame is greater than the angle change threshold, and the change range of the hand center point coordinates corresponding to the target video frame is greater than the center point coordinate change threshold, then use the target video frame as the key frame.
[0023] As a possible implementation, the process of segmenting the key frame to obtain a segmentation result and determining a feature map sequence based on the segmentation result includes:
[0024] Perform semantic segmentation on the regions containing facial features and the regions containing hand features in the key frame respectively to obtain a segmentation result, where the segmentation result includes the relative position information corresponding to the facial features and hand features in the key frame;
[0025] Determine the position information of a preset part in the key frame, and determine the feature map corresponding to the key frame according to the position information of the preset part in the key frame and the relative position information;
[0026] Determine the feature map sequence according to the feature map corresponding to the key frame and the frame extraction order corresponding to the key frame.
[0027] As a possible implementation, the training process of the sign language translation model includes:
[0028] Collect training samples, where the training samples include sign language sample videos and the corresponding question sample texts and sign language translation sample texts for the sign language sample videos;
[0029] Determine the key sample frames in the sign language video;
[0030] Perform segmentation processing on the key sample frames to obtain a segmentation result, and determine a feature map sequence based on the segmentation result;
[0031] Train an initial translation model based on the question sample text, the sign language translation sample text, and the feature map sequence to obtain a sign language translation model.
[0032] As a possible implementation, the training of the initial translation model based on the question sample text, the sign language translation sample text, and the feature map sequence to obtain a sign language translation model includes:
[0033] Perform encoding processing on the question sample text and the sign language translation sample text to obtain a question semantic vector and a translation semantic vector, and extract the video features corresponding to each feature map from the feature map sequence;
[0034] Use a predefined special marker to mask a preset proportion of the feature elements in the translation semantic vector;
[0035] Input the question semantic vector, the video features corresponding to each feature map, and the masked translation semantic vector into the initial translation model. The initial translation model predicts the feature elements masked in the translation semantic vector based on the video features corresponding to each feature map and the question semantic vector, and corrects the parameters of the initial translation model according to the prediction results. Repeat the process until the initial translation model stops training when it reaches a preset stop condition, and obtain the sign language translation model.
[0036] In a second aspect, an embodiment of the present application provides a sign language translation device, which includes:
[0037] An acquisition module, configured to acquire a question text and a sign language video to be translated, where the sign language video is a sign language video formed by a user's gesture answer based on the question text;
[0038] A determination module, configured to determine key frames in the sign language video;
[0039] A processing module, configured to perform segmentation processing on the key frames to obtain a segmentation result, and determine a feature map sequence based on the segmentation result;
[0040] A translation module, configured to input the question text and the feature map sequence into a pre-trained sign language translation model to obtain a sign language translation text corresponding to the sign language video.
[0041] As a possible implementation manner, the determination module is specifically configured to:
[0042] Perform frame extraction processing on the sign language video to obtain a plurality of initial video frames;
[0043] Perform hand detection processing on each of the initial video frames, and screen out a plurality of target video frames with hand features from the plurality of initial video frames according to the results of the hand detection processing;
[0044] Perform key point detection processing on each of the target video frames in the order of frame extraction corresponding to each target video frame, and determine the key frames according to the key point detection results.
[0045] As a possible implementation manner, the determination module is further configured to:
[0046] Perform key point detection on the region with hand features in each of the target video frames to obtain key point detection results, where the key point detection results include the change range of the limb angle corresponding to each target video frame and the change range of the hand center point coordinates;
[0047] Determine the key frames according to the limb angle change range corresponding to each target video frame, the preset angle change threshold, the change range of the hand center point coordinates corresponding to each target video frame, and the preset center point coordinate change threshold.
[0048] As a possible implementation manner, the determining module is further configured to:
[0049] Determine the limb angle change range corresponding to the current target video frame according to the limb angle information of the previous target video frame of the current target video frame and the limb angle information of the current target video frame;
[0050] Determine the change range of the hand center point coordinates corresponding to the current target video frame according to the hand center point coordinates in the previous target video frame and the hand center point coordinates in the current target video frame, where the current target video frame is any target video frame other than the first target video frame among the multiple target videos.
[0051] As a possible implementation manner, the determining module is further configured to:
[0052] If the limb angle change range corresponding to the target video frame is greater than the angle change threshold, and the change range of the hand center point coordinates corresponding to the target video frame is greater than the center point coordinate change threshold, then use the target video frame as the key frame.
[0053] As a possible implementation manner, the processing module is specifically configured to:
[0054] Perform semantic segmentation processing on the area containing the face features and the area containing the hand features in the key frame respectively to obtain a segmentation result, where the segmentation result includes the relative position information of the face features and the hand features corresponding in the key frame;
[0055] Determine the position information of the preset part in the key frame, and determine the feature map corresponding to the key frame according to the position information of the preset part in the key frame and the relative position information;
[0056] Determine the feature map sequence according to the feature map corresponding to the key frame and the extraction order of the key frame.
[0057] As a possible implementation manner, the sign language translation device further includes a training module, and the training module is specifically configured to:
[0058] Collect training samples, where the training samples include sign language sample videos, the question sample texts corresponding to the sign language sample videos, and the sign language translation sample texts;
[0059] Determine the key sample frames in the sign language sample video;
[0060] Perform a segmentation process on the key sample frames to obtain a segmentation result, and determine a sequence of feature maps based on the segmentation result;
[0061] Train an initial translation model based on the query sample text, the sign language translation sample text, and the sequence of feature maps to obtain a sign language translation model.
[0062] As a possible implementation, the training module is further configured to:
[0063] Perform an encoding process on the query sample text and the sign language translation sample text to obtain a query semantic vector and a translation semantic vector, and extract video features corresponding to each feature map from the sequence of feature maps;
[0064] Use a predefined special marker to mask a preset proportion of feature elements in the translation semantic vector;
[0065] Input the query semantic vector, the video features corresponding to each feature map, and the masked translation semantic vector into the initial translation model. The initial translation model predicts the masked feature elements in the translation semantic vector based on the video features corresponding to each feature map and the query semantic vector, and corrects the parameters of the initial translation model according to the prediction result. Repeat the process until the initial translation model reaches a preset stop condition and stop training to obtain the sign language translation model.
[0066] In a third aspect, an embodiment of the present application provides an electronic device, including: a processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the storage medium through the bus. The processor executes the machine-readable instructions to perform the steps of the sign language translation method according to any one of the first aspects described above.
[0067] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, it performs the steps of the sign language translation method according to any one of the first aspects described above.
[0068] According to the sign language translation method, electronic device, and storage medium of the embodiments of the present application, obtain a question text and a sign language video to be translated, determine key frames in the sign language video and perform segmentation processing on the key frames to obtain a segmentation result, determine a sequence of feature maps based on the segmentation result, and input the question text and the sequence of feature maps into a pre-trained sign language translation model to obtain a sign language translation text corresponding to the sign language video. According to the embodiments of the present application, by extracting key frames, it is possible to effectively reduce blurred frames and solve the natural redundancy problem of consecutive frames. And by performing segmentation processing on the key frames, further optimizing the feature maps based on the segmentation result, and generating a sequence of feature maps, it is possible to reduce background interference, making the human features such as facial expressions, gestures, and postures prominent in sign language translation, thereby reducing the learning difficulty of the sign language translation model. Moreover, introducing the question text in sign language translation helps the sign language translation model for cross-modal learning and establishes the basis for sign language translation in the context of the sign language video, thus improving the speed and accuracy of sign language translation. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0070] Figure 1 The flowchart showing a sign language translation method provided by the embodiments of the present application;
[0071] Figure 2 The flowchart showing a method for determining key frames provided by the embodiments of the present application;
[0072] Figure 3 The flowchart showing another method for determining key frames provided by the embodiments of the present application;
[0073] Figure 4 The schematic diagram showing a limb angle provided by the embodiments of the present application;
[0074] Figure 5 The flowchart showing a method for determining a sequence of feature maps provided by the embodiments of the present application;
[0075] Figure 6 The schematic diagram showing a feature map provided by the embodiments of the present application;
[0076] Figure 7 The flowchart showing a sign language video translation provided by the embodiments of the present application;
[0077] Figure 8The figure shows a schematic flowchart of a sign language translation model training method provided by an embodiment of the present application;
[0078] Figure 9 The figure shows a schematic flowchart of another sign language translation model training method provided by an embodiment of the present application;
[0079] Figure 10 The figure shows a schematic structural diagram of an optimized training model provided by an embodiment of the present application;
[0080] Figure 11 The figure shows a schematic structural diagram of a sign language translation device provided by an embodiment of the present application;
[0081] Figure 12 The figure shows a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0082] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. It should be understood that the accompanying drawings in the present application are only for the purposes of illustration and description, and are not used to limit the protection scope of the present application. In addition, it should be understood that the schematic drawings are not drawn to actual scale. The flowcharts used in the present application show operations implemented according to some embodiments of the present application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical context relationships may be reversed or implemented simultaneously. In addition, those skilled in the art may add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of the present application.
[0083] In addition, the described embodiments are only some of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application described and illustrated in the accompanying drawings here may be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts fall within the protection scope of the present application.
[0084] To enable those skilled in the art to use the content of the present application, the following implementation manners are given in combination with the specific application scenario of "sign language translation". For those skilled in the art, without departing from the spirit and scope of the present application, the general principles defined here can be applied to other embodiments and application scenarios. Although the present application is mainly described around the sign language translation method, it should be understood that this is only an exemplary embodiment.
[0085] It should be noted that the term "including" will be used in the embodiments of the present application to indicate the presence of the features stated thereafter, but does not exclude the addition of other features.
[0086] A sign language translation method based on video key frames and a sign language translation model provided by an embodiment of the present application is applicable not only to the financial industry, but also to sub - industries such as medical care and education, and can effectively improve the speed and accuracy of sign language translation.
[0087] Figure 1 The flowchart of a sign language translation method provided by an embodiment of the present application is shown. The execution subject of this method is an electronic device. Refer to Figure 1 As shown, the method specifically includes the following steps:
[0088] S101. Obtain a question text and a sign language video to be translated. The sign language video is a sign language video formed by the user's gesture answer based on the question text.
[0089] Optionally, the sign language video to be translated is obtained based on a shooting device, and the question text is a record of the information asked by the questioner. That is to say, the question text contains the question of the questioner, and the sign language video is the answer made by the person being questioned based on the question of the questioner and shot by the shooting device. For example, in the financial field, a financial salesman is the questioner, and a special user is the person being questioned. The financial salesman asks the special user a question and records the question asked to form a question text. Correspondingly, the answer made by the special user based on gestures is stored in video form, that is, the sign language video is obtained.
[0090] S102. Determine the key frames in the sign language video.
[0091] Optionally, the sign language video is composed of many video frames. Since the sign language video includes many gesture actions, and for different users, the speed and amplitude of gesture changes are different. If all the video frames of the sign language video are used as the basic frames for sign language translation, it will not only reduce the accuracy of sign language translation, but also occupy too many resources and reduce the efficiency of sign language translation. Therefore, key frames are screened and determined from the multiple video frames included in the sign language video, and subsequent sign language translation operations are performed based on the key frames.
[0092] S103. Perform segmentation processing on the key frames to obtain a segmentation result, and determine a sequence of feature maps based on the segmentation result.
[0093] Optionally, performing segmentation processing on the key frames, that is, performing semantic segmentation processing on the key frames. For sign language translation, there are special regions such as hand features and facial features in the key frames that affect the sign language translation result. Performing semantic segmentation processing is to perform semantic segmentation processing on the special regions in the key frames, thereby reducing the influence of some invalid information in the key frames on sign language translation. Further, after performing semantic segmentation processing on the key frames, a feature map is determined according to the segmentation result, and a feature map sequence is determined in combination with the frame extraction order corresponding to the key frames, and the feature map sequence is used as the input parameter of the sign language translation model and input into the pre-trained sign language translation model.
[0094] S104. Input the question text and the feature map sequence into the pre-trained sign language translation model to obtain the sign language translation text corresponding to the sign language video.
[0095] Optionally, before performing sign language translation, the initial sign language translation model is pre-trained to obtain a trained sign language translation model. While using the feature map sequence as the input parameter of the pre-trained sign language translation model, the question text is also used as the input parameter of the sign language translation model and input into the pre-trained sign language translation model together with the feature map sequence. The sign language translation model performs sign language translation based on the feature map sequence and the sign language translation basis established based on the input text, thereby obtaining the sign language translation text corresponding to the sign language video.
[0096] Based on this, according to the sign language translation method provided by the embodiments of the present application, the question text and the sign language video to be translated are obtained, the key frames in the sign language video are determined and segmentation processing is performed on the key frames to obtain a segmentation result, a feature map sequence is determined based on the segmentation result, and the question text and the feature map sequence are input into the pre-trained sign language translation model to obtain the sign language translation text corresponding to the sign language video. According to the embodiments of the present application, by extracting key frames, it is possible to effectively reduce blurred frames and solve the natural redundancy problem of consecutive frames. By performing segmentation processing on the key frames and further optimizing the feature map based on the segmentation result to generate a feature map sequence, it is possible to reduce background interference, make human features such as facial expressions, gestures, and postures prominent in sign language translation, thereby reducing the learning difficulty of the sign language translation model. And introducing the question text in sign language translation helps the sign language translation model for cross-modal learning and establishes the basis for sign language translation in the context of the sign language video, thereby improving the speed and accuracy of sign language translation.
[0097] As a possible implementation manner, as Figure 2 shown, the above step S102 of determining the key frames in the sign language video specifically includes the following steps:
[0098] S201. Perform frame extraction processing on the sign language video to obtain a plurality of initial video frames.
[0099] Exemplarily, frame extraction refers to the process of simulating taking a photo every once in a while and combining them to form a video by extracting several frames from a video, similar to slow-motion photography. Before frame extraction is performed on each video, the number of frames is formed first. The number of frames in a video is the amount of pictures transmitted within 1 second, which can also be understood as the refresh rate of the graphics processor per second. Performing frame extraction on the sign language video to obtain multiple initial video frames is also to facilitate the screening process for multiple initial video frames, so as to improve the speed and accuracy of sign language translation. Among them, each initial video frame is a static image, and quickly displaying each initial video frame continuously creates the illusion of motion.
[0100] Exemplarily, usually, video frames can be extracted from a video at a certain frequency. For example, a fixed number of video frames are extracted per second, such as 25 frames per second. However, in the sign language translation scenario, for different sign language gestures, the user's hand speed varies. Therefore, when performing frame extraction on the sign language video, a low-frequency extraction method is adopted for sign language video segments with a slow hand speed, and a high-frequency extraction method is adopted for sign language video segments with a fast hand speed. Based on this, different frame extraction frequency periods correspond to sign language videos with different hand speeds, which can improve the efficiency of obtaining key frames.
[0101] S202. Perform hand detection processing on each initial video frame, and screen out multiple target video frames with hand features from multiple initial video frames according to the results of the hand detection processing.
[0102] Exemplarily, since the semantic boundaries of continuous sign language are unknown, the sign language vocabulary is rich and many terms have similar gestures and actions, and each person has different sign language speeds, sign language segmentation is relatively difficult. Each initial video frame is also obtained by performing frame extraction on the sign language video, which is equivalent to performing a sign language segmentation on the sign language video. However, not all of the multiple initial video frames obtained by segmentation contain clear hand features, and there will also be some blurred frames and blank frames. Therefore, hand detection processing is performed on each initial video frame respectively, and multiple target video frames with clear hand features are screened out from multiple video frames. That is to say, when performing hand detection processing on each initial video frame, only the target video frames that can detect both hands are retained.
[0103] Exemplarily, there are initial video frames A, B, C, D, E, F, and the corresponding frame extraction order of each initial video frame is 1, 2, 3, 4, 5, 6 in sequence. If hand detection processing is performed on each initial video frame, the target video frames A, B, D, F with hand features and clear features are screened out.
[0104] S203. According to the frame extraction order corresponding to each target video frame, perform key point detection processing on each target video frame respectively, and determine the key frames according to the key point detection results.
[0105] Optionally, each initial video frame is obtained by extracting frames from a sign language video. Therefore, each initial video frame has a corresponding frame extraction order. Each target video frame is selected from multiple initial video frames. Based on the frame extraction order of each initial video frame and the position of each target video frame among the multiple initial video frames, the frame extraction order corresponding to each target video frame can be determined.
[0106] Exemplarily, continuing with the above example where there are initial video frames A, B, C, D, E, F, and the corresponding frame extraction orders of each initial video frame are 1, 2, 3, 4, 5, 6 in sequence, and hand detection processing is performed on each initial video frame, and target video frames A, B, D, F with clear hand features are selected. At this time, the frame extraction orders of video frames A, B, D, F will change from the original 1, 2, 4, 6 to 1, 2, 3, 4, that is, the frame extraction order corresponding to each target video frame is determined.
[0107] As a possible implementation, as Figure 3 shown, in step S203 above, key point detection processing is respectively performed on each target video frame, and key frames are determined according to the key point detection results, which specifically include the following steps:
[0108] S301. Perform key point detection on the area with hand features in each target video frame to obtain the key point detection result. The key point detection result includes the change range of the limb angle corresponding to each target video frame and the change range of the hand center point coordinates.
[0109] Optionally, according to the limb angle information of the previous target video frame of the current target video frame and the limb angle information of the current target video frame, determine the change range of the limb angle corresponding to the current target video frame.
[0110] Exemplarily, denote the current target video frame as Cur_Frame hand , and denote the previous target video frame of the current target video frame as Base_Frame hand . For the current target video frame, the previous target video frame of the current target video frame is the reference frame of the current target video frame. Therefore, the reference frame changes in real time.
[0111] Exemplarily, as Figure 4As shown, the limb angle information includes the angle α1 between the right upper arm and the right forearm, the angle α2 between the right upper arm and the left edge of the torso, the angle β1 between the left upper arm and the left forearm, and the angle β2 between the left upper arm and the left edge of the torso. Specifically, the limb angle change range corresponding to the current target video frame is determined comprehensively based on the right limb angle change and the left limb angle change. The left limb angle change is shown in the following formula (1), the right limb angle change is shown in the following formula (2), and the limb angle change range corresponding to the current target video frame is determined according to the following formula (3):
[0112] diff_angle left =sin(|cur_α1 - base_α1|)+sin(|cur_α2 - base_α2|)(1)
[0113] Where diff_angle left represents the left limb angle change, cur_α1 represents the angle between the right upper arm and the right forearm in the current target video frame, base_α1 represents the angle between the right upper arm and the right forearm in the previous target video frame of the current target video frame, cur_α2 represents the angle between the right upper arm and the left edge of the torso in the current target video frame, and base_α2 represents the angle between the right upper arm and the left edge of the torso in the previous target video frame of the current target video frame.
[0114] diff_angle right =sin(|cur_β1 - base_β1|)+sin(|cur_β2 - base_β2|) (2)
[0115] Where ditt_angle right represents the right limb angle change, cur_β1 represents the angle between the left upper arm and the left forearm in the current target video frame, base_β1 represents the angle between the left upper arm and the left forearm in the previous target video frame of the current target video frame, cur_β2 represents the angle between the left upper arm and the left edge of the torso in the current target video frame, and base_α2 represents the angle between the left upper arm and the left edge of the torso in the previous target video frame of the current target video frame.
[0116] diff_angle=max(diff_angle left ,diff_angle right ) (3)
[0117] Where diff_angle represents the limb angle change range corresponding to the current target video frame, diff_angle left represents the left limb angle change, and diff_angle right represents the right limb angle change.
[0118] Optionally, determine the change range of the hand center point coordinates corresponding to the current target video frame according to the hand center point coordinates in the previous target video frame and the hand center point coordinates in the current target video frame, where the current target video frame is any target video frame other than the first target video frame among multiple target videos.
[0119] Exemplarily, determine the change range of the hand center point coordinates corresponding to the current target video frame according to the following formula (4):
[0120] diff_distance = max(|A|, |B|) (4)
[0121] where A = Cur_Lett_Hand pos -Base_Lett_Hand pos , B = Cur_Right_Hand pos -Base_Right_Hand pos .
[0122] where diff_distance represents the change range of the hand center point coordinates corresponding to the current target video frame, Cur_Left_Hand pos represents the left hand center point coordinates in the current target video frame, Base_Left_Hand pos represents the left hand center point coordinates in the previous target video frame, Cur_Right_Hand pos represents the right hand center point coordinates in the current target video frame, Base_Right_Hand pos represents the right hand center point coordinates in the previous target video frame.
[0123] S302. Determine the key frames according to the limb angle change range corresponding to each target video frame and the preset angle change threshold, and the change range of the hand center point coordinates corresponding to each target video frame and the preset center point coordinate change threshold.
[0124] Exemplarily, if the limb angle change range corresponding to the target video frame is greater than the angle change threshold, and the change range of the hand center point coordinates corresponding to the target video frame is greater than the center point coordinate change threshold, then regard the target video frame as a key frame.
[0125] Exemplarily, the target video frames with the limb angle change range greater than the angle change threshold and the change range of the hand center point coordinates greater than the center point coordinate change threshold are used as key frames in order to select video frames containing significant gesture changes from multiple target video frames and ignore the static or minimally changing video frames. Based on this, while reducing the data volume, sufficient information to express gesture changes is retained, which can not only improve the processing efficiency of key frames but also improve the accuracy of gesture recognition.
[0126] Exemplarily, an angle change threshold and a center point coordinate change threshold are preset, and the specific values corresponding to the thresholds are set according to the actual situation and are not specifically limited herein. If the limb angle change range diff_angle corresponding to the current target video frame and the change range diff_distance of the hand center point coordinates corresponding to the current target video frame respectively exceed the corresponding thresholds, then the current target video frame is retained and used as a key frame.
[0127] Based on this, through performing hand detection processing and screening processing on multiple initial video frames obtained by frame extraction of a sign language video, screening target video frames from the multiple initial video frames, and then determining key frames from the target video frames, according to the key frame determination method provided in the embodiments of the present application, it is possible to effectively reduce blurred frames and solve the natural redundancy problem of consecutive frames, thereby improving the inference speed and accuracy of the sign language translation model.
[0128] As a possible implementation manner, as Figure 5 shown, the above step S103 performs segmentation processing on the key frame to obtain a segmentation result, and determines a feature map sequence based on the segmentation result, which specifically includes the following steps:
[0129] S501. Respectively perform semantic segmentation processing on the region containing facial features and the region containing hand features in the key frame to obtain a segmentation result, and the segmentation result includes the relative position information corresponding to the facial features and the hand features in the key frame.
[0130] Exemplarily, semantic segmentation refers to assigning a corresponding class label to each pixel point in an image or video to achieve the segmentation of different class objects. Respectively performing semantic segmentation processing on the region containing facial features and the region containing hand features in the key frame can adopt a semantic segmentation model based on deep learning, such as semantic segmentation models like U-Net, Mask R-CNN, etc., to accurately segment the facial features, hand features, and background part in the key frame, and further determine the relative position information corresponding to the facial features and the hand features in the key frame.
[0131] S502. Determine the position information of the preset part in the key frame, and determine the feature map corresponding to the key frame according to the position information of the preset part in the key frame and the relative position information.
[0132] Exemplarily, determining the position information of the preset part in the key frame means obtaining key points, such as Figure 6 each joint point marked in the figure, which are mainly distributed at the wrist, elbow and shoulder. Drawing the segmentation result and the key points on a black canvas of the same size, and connecting the human key points with white solid lines, the feature map as shown in Figure 6 can be obtained.
[0133] S503. Determine the feature map sequence according to the feature map corresponding to the key frame and the frame extraction order corresponding to the key frame.
[0134] Exemplarily, for each key frame, the corresponding feature map will be obtained based on the above processing method, and then combined with the frame extraction order corresponding to the key frame to obtain the feature map sequence. Among them, the feature map sequence includes feature maps corresponding to multiple key frames arranged in sequence.
[0135] Exemplarily, continue with the above example where there are target video frames A, B, D, F, and the frame extraction orders corresponding to each target video frame are 1, 2, 3, 4. After performing key point detection processing on each target video frame, the key frames that meet the conditions are A, B, D. The frame extraction orders corresponding to the key frames A, B, D are 1, 2, 3, and the key frames A, B, D respectively correspond to feature Figure 1 , feature Figure 2 and feature Figure 3 . Then, identify and sort each key frame according to the frame extraction order corresponding to the key frame, and the feature map sequence can be obtained.
[0136] Exemplarily, Figure 7 shows a schematic flow chart of a sign language video preprocessing provided by an embodiment of the present application. After processing the sign language video through the above process, the question text and the obtained feature map sequence are input into a pre-trained sign language translation model for inference, and then the sign language translation text is obtained.
[0137] Based on this, applying the key frame selection and optimized feature map strategy provided by the embodiments of the present application to various types of sign language translation models will significantly improve the translation accuracy of the sign language translation model.
[0138] As a possible implementation manner, Figure 8 shows a schematic flow chart of a sign language translation model training method provided by an embodiment of the present application. Referring to Figure 8 shown, the training process of the sign language translation model specifically includes the following steps:
[0139] S801. Collect training samples, where the training samples include sign language sample videos, as well as the corresponding question sample texts and sign language translation sample texts for the sign language sample videos.
[0140] Exemplarily, the question sample text contains the question of the questioner. The sign language sample video is the answer made by the questioned person to the question of the questioner based on the shooting device. The sign language translation sample text is the translation result of the sign language video.
[0141] S802. Determine the key sample frames in the sign language sample video.
[0142] Exemplarily, the sign language sample video is composed of many video frames. Similarly, since the sign language sample video includes many gesture actions, and the gesture change speed and amplitude vary for different users, if all the video frames of the sign language sample video are used as the basic frames for sign language translation, it will not only reduce the accuracy of the sign language translation model, but also occupy too much resources and reduce the efficiency of the sign language translation model. Therefore, key frames are screened and determined from the multiple video frames included in the sign language sample video, and subsequent sign language translation operations are then based on the key frames.
[0143] Specifically, frame extraction processing is performed on the sign language sample video to obtain multiple initial sample video frames. Hand detection processing is performed on each initial sample video frame. According to the results of the hand detection processing, multiple target sample video frames with hand features are screened out from the multiple initial sample video frames. Then, according to the frame extraction order corresponding to each target sample video frame, key point detection processing is performed on each target sample video frame respectively, and the key sample frames are determined according to the key point detection results.
[0144] S803. Perform segmentation processing on the key sample frames to obtain a segmentation result, and determine a feature map sequence based on the segmentation result.
[0145] Exemplarily, the specific implementation process of performing segmentation processing on the key sample frames and determining the feature map sequence according to the segmentation result is the same as the processing process in the sign language translation method, that is, semantic segmentation processing is respectively performed on the area containing the facial features and the area containing the hand features in the key sample frames to determine the position information of the preset parts in the key sample frames, and the feature map corresponding to the key sample frames is determined according to the position information and relative position information of the preset parts in the key sample frames. Then, according to the feature map corresponding to the key sample frames and the frame extraction order corresponding to the key sample frames, the feature map sequence is determined.
[0146] S804. Train the initial translation model based on the question sample text, the sign language translation sample text, and the feature map sequence to obtain a sign language translation model.
[0147] Exemplarily, different from the sign language translation method, sign language translation sample texts are introduced during the translation of the initial translation model, the Transformer. That is to say, in the sign language translation method, the input parameters of the pre-trained sign language translation model are the question text and the sequence of feature maps, while in the sign language translation model training method, the input parameters of the initial translation model are the question sample text, the sign language translation sample text, and the sequence of feature maps. Then, the initial translation model is trained based on the question sample text, the sign language translation sample text, and the sequence of feature maps to obtain the sign language translation model.
[0148] Based on this, according to the training method of the sign language translation model provided by the embodiments of the present application, by extracting key frames, it is possible to effectively reduce blurred frames and solve the natural redundancy problem of consecutive frames. And by performing segmentation processing on the key frames and further optimizing the feature maps based on the segmentation results to generate a sequence of feature maps, background interference can be reduced, making the human features such as facial expressions, gestures, and postures prominent in sign language translation, thereby reducing the learning difficulty of the sign language translation model. Moreover, introducing the question text in sign language translation helps the sign language translation model for cross-modal learning and establishes the basis for sign language translation in the sign language video context, thus improving the speed and accuracy of the sign language translation model.
[0149] As a possible implementation, as Figure 9 shown, the above step S804 trains the initial translation model based on the question sample text, the sign language translation sample text, and the sequence of feature maps to obtain the sign language translation model, which specifically includes the following steps:
[0150] S901. Encode the question sample text and the sign language translation sample text to obtain a question semantic vector and a translation semantic vector, and extract the video features corresponding to each feature map from the sequence of feature maps.
[0151] Exemplarily, the question sample text and the sign language translation sample text are respectively encoded into a question semantic vector "question tokens" and a translation semantic vector "translation tokens", and the sequence of feature maps is input into the translation model Swin Transformer to extract video features "video tokens".
[0152] S902. Mask a preset proportion of the feature elements in the translation semantic vector using a predefined special marker.
[0153] Exemplarily, the predefined special token is token[MASK], that is, masking the translation semantic vectors, and replacing a preset proportion of the word tokens in the translation semantic vectors with the predefined special token token[MASK] to obtain the masked translation semantic vectors.
[0154] S903. Input the question semantic vector, the video features corresponding to each feature map, and the masked translation semantic vectors into the initial translation model. The initial translation model predicts the masked feature elements in the translation semantic vectors based on the video features corresponding to each feature map and the question semantic vector, and corrects the parameters of the initial translation model according to the prediction results. Repeat the execution until the initial translation model reaches the preset stop condition and stops training to obtain the sign language translation model.
[0155] Exemplarily, as Figure 10 shown, input the question semantic vector, the video features corresponding to each feature map, and the masked translation semantic vectors into the initial translation model, predict the masked feature elements in the translation semantic vectors. Similar to the encoding and decoding process, through the encoding and decoding of the multi-modal transformer encoder, an output prediction result is obtained. Repeat the above process, and use the output prediction result to iteratively correct the parameters of the initial translation model until the initial translation model reaches the preset stop condition and stops training to obtain the sign language translation model.
[0156] Among them, the initial translation model is a translation model constructed based on the SwinBERT structure. The preset stop condition is that the similarity between the meaning of the predicted masked feature elements and the meaning of the originally masked feature elements reaches a preset threshold, such as more than 95%, etc. No specific limitation is made here, but undoubtedly, the higher the similarity, the more accurate the model prediction result.
[0157] Based on this, by determining the key sample frames and the feature map sequence for model training, the influence of external factors on the prediction result of the sign language translation model is reduced, and the translation speed and accuracy of the sign language translation model are effectively improved.
[0158] Based on the same inventive concept, an embodiment of the present application also provides a sign language translation device corresponding to the sign language translation method. Since the principle of solving problems by the device in the embodiment of the present application is similar to the above sign language translation method in the embodiment of the present application, the implementation of the sign language translation device can refer to the implementation of the method, and the repeated parts will not be elaborated.
[0159] Refer to Figure 11As shown in the figure, it is a schematic structural diagram of a sign language translation device provided by an embodiment of the present application. The sign language translation device 1100 includes: an acquisition module 1101, a determination module 1102, a processing module 1103, and a translation module 1104, where:
[0160] The acquisition module 1101 is used to acquire a question text and a sign language video to be translated. The sign language video is a sign language video formed by the user's gesture answer based on the question text;
[0161] The determination module 1102 is used to determine the key frames in the sign language video;
[0162] The processing module 1103 is used to perform segmentation processing on the key frames to obtain a segmentation result, and determine a feature map sequence based on the segmentation result;
[0163] The translation module 1104 is used to input the question text and the feature map sequence into a pre-trained sign language translation model to obtain a sign language translation text corresponding to the sign language video.
[0164] Thus, according to the sign language translation device provided by the embodiment of the present application, a question text and a sign language video to be translated are acquired, the key frames in the sign language video are determined and segmented, a segmentation result is obtained, a feature map sequence is determined based on the segmentation result, the question text and the feature map sequence are input into a pre-trained sign language translation model, and a sign language translation text corresponding to the sign language video is obtained. According to the embodiment of the present application, by extracting key frames, it is possible to effectively reduce blurred frames and solve the natural redundancy problem of consecutive frames, and by performing segmentation processing on the key frames and further optimizing the feature map based on the segmentation result to generate a feature map sequence, it is possible to reduce background interference, making the human features such as facial expressions, gestures, and postures prominent in sign language translation, thereby reducing the learning difficulty of the sign language translation model. And introducing the question text in sign language translation helps the sign language translation model to perform cross-modal learning and establish a basis for sign language translation in the context of the sign language video, thereby improving the speed and accuracy of sign language translation.
[0165] In a possible implementation manner, the above determination module 1102 is specifically used for:
[0166] Perform frame extraction on the sign language video to obtain a plurality of initial video frames;
[0167] Perform hand detection processing on each initial video frame, and screen out a plurality of target video frames with hand features from the plurality of initial video frames according to the result of the hand detection processing;
[0168] Perform key point detection processing on each target video frame respectively according to the frame extraction order corresponding to each target video frame, and determine the key frames according to the key point detection results.
[0169] In a possible implementation manner, the above-mentioned determining module 1102 is further configured to:
[0170] Perform key point detection on the regions with hand features in each target video frame to obtain a key point detection result, where the key point detection result includes the range of limb angle changes and the range of changes in the hand center point coordinates corresponding to each target video frame;
[0171] Determine key frames according to the range of limb angle changes corresponding to each target video frame and a preset angle change threshold, and the range of changes in the hand center point coordinates corresponding to each target video frame and a preset center point coordinate change threshold.
[0172] In a possible implementation manner, the above-mentioned determining module 1102 is further configured to:
[0173] Determine the range of limb angle changes corresponding to the current target video frame according to the limb angle information of the previous target video frame of the current target video frame and the limb angle information of the current target video frame;
[0174] Determine the range of changes in the hand center point coordinates corresponding to the current target video frame according to the hand center point coordinates in the previous target video frame and the hand center point coordinates in the current target video frame, where the current target video frame is any target video frame other than the first target video frame among the multiple target videos.
[0175] In a possible implementation manner, the above-mentioned determining module 1102 is further configured to:
[0176] If the range of limb angle changes corresponding to the target video frame is greater than the angle change threshold, and the range of changes in the hand center point coordinates corresponding to the target video frame is greater than the center point coordinate change threshold, then use the target video frame as a key frame.
[0177] In a possible implementation manner, the above-mentioned processing module 1103 is specifically configured to:
[0178] Perform semantic segmentation processing on the regions containing facial features and the regions containing hand features in the key frames respectively to obtain a segmentation result, where the segmentation result includes the relative position information of the facial features and the hand features corresponding in the key frames;
[0179] Determine the position information of the preset part in the key frame, and determine the feature map corresponding to the key frame according to the position information of the preset part in the key frame and the relative position information;
[0180] Determine a feature map sequence according to the feature map corresponding to the key frame and the frame extraction order corresponding to the key frame.
[0181] In a possible implementation manner, the above-mentioned sign language translation device 1100 further includes a training module, and the training module is specifically configured to:
[0182] Collect training samples, where the training samples include sign language sample videos, as well as question sample texts and sign language translation sample texts corresponding to the sign language sample videos;
[0183] Determine the key sample frames in the sign language sample videos;
[0184] Perform segmentation processing on the key sample frames to obtain a segmentation result, and determine a feature map sequence based on the segmentation result;
[0185] Train an initial translation model based on the question sample text, the sign language translation sample text, and the feature map sequence to obtain a sign language translation model.
[0186] In a possible implementation manner, the above training module is further configured to:
[0187] Perform encoding processing on the question sample text and the sign language translation sample text to obtain a question semantic vector and a translation semantic vector, and extract video features corresponding to each feature map from the feature map sequence;
[0188] Use a predefined special marker to mask a preset proportion of feature elements in the translation semantic vector;
[0189] Input the question semantic vector, the video features corresponding to each feature map, and the masked translation semantic vector into the initial translation model. The initial translation model predicts the masked feature elements in the translation semantic vector based on the video features corresponding to each feature map and the question semantic vector, and corrects the parameters of the initial translation model according to the prediction result. Repeat the execution until the initial translation model reaches a preset stop condition and stop training to obtain a sign language translation model.
[0190] For the processing flow of each module in the device and the interaction flow between modules, reference can be made to the relevant descriptions in the above method embodiments, which will not be elaborated here.
[0191] This application embodiment also provides an electronic device 1200, as Figure 12 shown, which is a schematic structural diagram of the electronic device 1200 provided by this application embodiment, including: a processor 1201, a memory 1202. Optionally, a bus 1203 may also be included. The memory 1202 stores machine-readable instructions executable by the processor 1201. When the electronic device 1200 runs, the processor 1201 communicates with the memory 1202 through the bus 1203. When the machine-readable instructions are executed by the processor 1201, the steps of the sign language translation method described in any one of the above are executed.
[0192] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the sign language translation method described in any one of the above.
[0193] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the method embodiments, which will not be repeated in this application. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or modules can be in electrical, mechanical or other forms.
[0194] In addition, in each embodiment of this application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical discs that can store program codes.
[0195] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application.
Claims
1. A sign language translation method, characterized in that: include: Obtaining a question text and a sign language video to be translated, wherein the sign language video is a sign language video formed by a gesture answer made by a user based on the question text; Determining key frames in the sign language video; Performing segmentation processing on the key frame to obtain a segmentation result, and determining a feature map sequence based on the segmentation result; Inputting the question text and the feature graph sequence into a pre-trained sign language translation model to obtain a sign language translation text corresponding to the sign language video; The step of segmenting the key frame to obtain a segmentation result, and determining a feature map sequence based on the segmentation result includes: Performing semantic segmentation processing on the area containing facial features and the area containing hand features in the key frame respectively to obtain a segmentation result, wherein the segmentation result includes relative position information corresponding to the facial features and the hand features in the key frame; Determining position information of a preset part in the key frame, and determining a feature map corresponding to the key frame according to the position information of the preset part in the key frame and the relative position information; The feature map sequence is determined according to the feature map corresponding to the key frame and the frame extraction order corresponding to the key frame.
2. The sign language translation method according to claim 1, characterized in that: Determining the key frames in the sign language video includes: Performing frame extraction processing on the sign language video to obtain multiple initial video frames; Performing hand detection processing on each of the initial video frames, and selecting a plurality of target video frames having hand features from the plurality of initial video frames according to the result of the hand detection processing; According to the frame extraction sequence corresponding to each target video frame, key point detection processing is performed on each target video frame respectively, and the key frame is determined according to the key point detection result.
3. The sign language translation method according to claim 2, characterized in that: The step of performing key point detection processing on each target video frame and determining the key frame according to the key point detection result includes: Perform key point detection on the area with hand features in each target video frame to obtain a key point detection result, wherein the key point detection result includes a limb angle change range and a hand center point coordinate change range corresponding to each target video frame; The key frame is determined according to the limb angle change range corresponding to each target video frame and the preset angle change threshold, as well as the hand center point coordinate change range corresponding to each target video frame and the preset center point coordinate change threshold.
4. The sign language translation method according to claim 3, characterized in that: The key point detection is performed on the area with hand features in each target video frame to obtain the key point detection result, including: Determine a limb angle variation range corresponding to the current target video frame according to limb angle information of a previous target video frame of the current target video frame and limb angle information of the current target video frame; According to the hand center point coordinates in the previous target video frame and the hand center point coordinates in the current target video frame, the variation range of the hand center point coordinates corresponding to the current target video frame is determined, wherein the current target video frame is any target video frame among the multiple target videos except the first target video frame.
5. The sign language translation method according to claim 3, characterized in that: The determining of the key frame according to the limb angle change range and the preset angle change threshold corresponding to each target video frame, and the hand center point coordinate change range and the preset center point coordinate change threshold corresponding to each target video frame, comprises: If the limb angle change range corresponding to the target video frame is greater than the angle change threshold, and the hand center point coordinate change range corresponding to the target video frame is greater than the center point coordinate change threshold, the target video frame is used as the key frame.
6. The sign language translation method according to claim 1, characterized in that: The training process of the sign language translation model includes: Collecting training samples, wherein the training samples include a sign language sample video and a question sample text and a sign language translation sample text corresponding to the sign language sample video; Determining key sample frames in the sign language sample video; Performing segmentation processing on the key sample frame to obtain a segmentation result, and determining a feature map sequence based on the segmentation result; The initial translation model is trained based on the question sample text, the sign language translation sample text and the feature graph sequence to obtain a sign language translation model.
7. The sign language translation method according to claim 6, characterized in that: The initial translation model is trained based on the question sample text, the sign language translation sample text and the feature graph sequence to obtain a sign language translation model, including: Encoding the question sample text and the sign language translation sample text to obtain a question semantic vector and a translation semantic vector, and extracting video features corresponding to each feature map from the feature map sequence; Using a predefined special mark to perform masking on a preset proportion of feature elements in the translation semantic vector; The question semantic vector, the video features corresponding to each feature graph, and the translation semantic vector after masking are input into the initial translation model, and the initial translation model predicts the feature elements in the translation semantic vector that are masked based on the video features corresponding to each feature graph and the question semantic vector, and corrects the parameters of the initial translation model according to the prediction result. The training is executed in a loop until the initial translation model reaches a preset stop condition, thereby obtaining the sign language translation model.
8. An electronic device, characterized in that: include: A processor and a memory, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor executes the machine-readable instructions to perform the steps of the sign language translation method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the sign language translation method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Sign language recognition method based on interaction scene
CN113688685A
Dialogue scene sign language translation model based on double-flow attention and method thereof
CN115862140A
Sign language recognition method fusing hand skeleton and facial expression features
CN116486484A