Sign language recognition method and device, electronic equipment and readable storage medium
By splitting the translation model and employing weighted processing and a fixed sequence length normalization scheme, the problem of insufficient learning of discrete word semantic information in sign language recognition is solved, thereby improving the accuracy of sign language recognition and reducing model overfitting.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- VIVO MOBILE COMM CO LTD
- Filing Date
- 2022-10-19
- Publication Date
- 2026-04-10
AI Technical Summary
Existing sign language recognition technologies struggle to fully learn the semantic information of discrete sign language words, especially for small-amplitude sign language movements. Furthermore, traditional methods cannot simultaneously preserve the relative size relationships of key points in both the temporal and spatial dimensions.
The traditional translation model is split into a body feature translation model and a hand feature translation model. Different weights are used to weight the body and hand feature information in the time and space dimensions, and the relative size relationship of the key points in the space and time dimensions is preserved by a fixed sequence length standardization scheme.
It improves the accuracy of sign language recognition, reduces the number of model parameters, lowers the risk of model overfitting, and enables more comprehensive learning of sign language semantic information.
Smart Images

Figure CN115546897B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of artificial intelligence, and particularly relates to a sign language recognition method and device, electronic equipment and readable storage medium. BACKGROUND
[0002] Sign language is a tool for the hearing-impaired to communicate and express their thoughts, and is used to convey their information and express some complex or abstract semantic concepts. Generally, information is conveyed according to specific grammar, with finger movements combined with body language and facial expressions. With the development of electronic device technology, sign language recognition and translation functions have emerged to provide convenience for these hearing-impaired people.
[0003] Generally, sign language recognition schemes all input video stream information into a visual model for classification training to obtain sign language semantic information from a fixed template, or extract skeleton key points and then use a graph network or generate a heat map and then use a convolutional neural network (CNN) model for classification. However, this method cannot fully learn some sign language actions with small amplitudes. Alternatively, a skeleton key point method based on a transformer model is used. This method adds a cnn convolution layer to the transformer model, standardizes the sign language features of the human body input into the transformer, and uses the same attention module to recognize the sign language feature information.
[0004] Therefore, conventional sign language recognition schemes are too single and fixed, which can result in insufficient learning of discrete word sign language semantic information. SUMMARY
[0005] The embodiments of the present application aim to provide a sign language recognition method, device, electronic equipment and readable storage medium, which can solve the problem of how to fully learn discrete word sign language semantic information.
[0006] In a first aspect, the embodiments of the present application provide a sign language recognition method, which comprises: obtaining first human feature information of a target user in a first image frame, the human feature information comprising first body feature information and first hand feature information; inputting the first human feature information into a translation model, and performing weighted processing on the first body feature information and the first hand feature information respectively to obtain second body feature information and second hand feature information; splicing the second body feature information and the second hand feature information to obtain second human feature information; and performing classification processing on the second human feature information to output sign language semantic information of the target user.
[0007] In a second aspect, an embodiment of the present application provides a sign language recognition device, comprising: an acquisition module and a processing module; the acquisition module is configured to acquire first human feature information of a target user in a first image frame, the human feature information comprising first body feature information and first hand feature information; the processing module is configured to input the first human feature information into a translation model, and perform weighted processing on the first body feature information and the first hand feature information respectively to obtain second body feature information and second hand feature information; the processing module is further configured to splice the second body feature information and the second hand feature information to obtain second human feature information; and the processing module is further configured to perform classification processing on the second human feature information, and output sign language semantic information of the target user.
[0008] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions are executed by the processor to implement the steps of the method according to the first aspect.
[0009] In a fourth aspect, an embodiment of the present application provides a readable storage medium, wherein the readable storage medium stores programs or instructions, and the programs or instructions are executed by a processor to implement the steps of the method according to the first aspect.
[0010] In a fifth aspect, an embodiment of the present application provides a chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is configured to run programs or instructions to implement the method according to the first aspect.
[0011] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium, and is executed by at least one processor to implement the method according to the first aspect.
[0012] In this embodiment, the first human body feature information of the target user in the first image frame is obtained. The human body feature information includes first body feature information and first hand feature information. The first human body feature information is input into a translation model, and the first body feature information and the first hand feature information are weighted and processed respectively to obtain second body feature information and second hand feature information. The second body feature information and the second hand feature information are concatenated to obtain second human body feature information. The second human body feature information is classified and processed to output the sign language semantic information of the target user. Thus, by inputting the target user's body feature information and hand feature information into the translation model provided in this application, the body feature information of the current frame is fused with the body feature information of the previous and next frames in the time dimension through weighted processing. Similarly, the hand feature information of the current frame is fused with the hand feature information of the previous and next frames in the time dimension through weighted processing. At the same time, the body feature information and hand feature information are fused in the spatial dimension through weighted processing. This allows the electronic device to not only learn the target user's sign language semantic information more fully based on the fused body feature information and hand feature information, but also reduces the number of parameters between models by splitting the traditional human feature translation model into a body feature translation model and a hand feature translation model, which helps to reduce model overfitting. Attached Figure Description
[0013] Figure 1 This is one of the flowcharts illustrating a sign language recognition method provided in an embodiment of this application;
[0014] Figure 2 This is one of the schematic diagrams of a sign language recognition method provided in the embodiments of this application;
[0015] Figure 3 This is a second schematic diagram of a sign language recognition method provided in an embodiment of this application;
[0016] Figure 4 This is one of the example schematic diagrams of a sign language recognition method provided in the embodiments of this application;
[0017] Figure 5 This is a second example schematic diagram of a sign language recognition method provided in the embodiments of this application;
[0018] Figure 6 This is a second schematic flowchart of a sign language recognition method provided in an embodiment of this application;
[0019] Figure 7 This is a schematic diagram of the structure of a sign language recognition device provided in an embodiment of this application;
[0020] Figure 8Fig. 1 is a schematic diagram of a hardware structure of an electronic device according to an embodiment of the present application.
[0021] Figure 9 Fig. 2 is another schematic diagram of a hardware structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0022] The technical solutions in the embodiments of the present application will be clearly described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art belong to the scope of protection of the present application.
[0023] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of a kind and do not limit the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification and claims indicates at least one of the connected objects, and the character " / ", generally indicates that the front and rear associated objects are in an "or" relationship.
[0024] The sign language recognition provided by the embodiments of the present application will be described in detail below in combination with the drawings, through specific embodiments and their application scenarios.
[0025] First, in the prior art, electronic devices usually use open source toolkits (such as mediapipe, OpenPose, MMPose, etc., hereinafter taking mediapipe as an example) to extract joint information, however, the human joint coordinates extracted by them are only available in x and y axes, and the z axis coordinates are not available (as indicated in the official document), and the z axis coordinates represent depth information, which represents the distance of the palm from the human body in the depth direction. It is an important feature of sign language recognition. If we cannot accurately obtain the palm z axis information, the information entropy of the model input will be lost, resulting in the model being unable to learn sufficiently.
[0026] Secondly, the traditional transformer adopts layer normalization (layer norm) to standardize the features. The layer norm standardization cannot reflect the changes of the same coordinate point in the time dimension, and the batch normalization (batchnorm) cannot reflect the relative size of the node features in the single frame image. In addition, some human joint nodes remain stationary in consecutive frames of data, but due to the detection error of the key point module (such as mediepipe), the coordinates have slight fluctuations. If the data processed by the batch norm cannot retain the information of the stationary state, the current standardization scheme cannot simultaneously retain the relative size relationship of the coordinates in the spatial dimension and the time dimension. However, in sign language recognition, the relative position of the coordinate points in the spatial dimension and the time dimension plays a crucial role.
[0027] In addition, understanding sign language is crucial to understanding hand shapes, followed by the position of the palm relative to the body. The palm joint nodes are very dense and have small changes in distance between each other, while the body joint nodes are very sparse and have large changes in distance between each other. When the joint coordinate is used as the sign language feature input model, the palm joint coordinate changes little but contains the most important sign language information, while the body joint coordinate changes relatively weakly. The traditional Transformer standardizes all features through a layer norm. Although the layer norm can change the data distribution, it cannot change the relative size of the data, and cannot deeply integrate all feature information of the body and the palm. If the body and palm features are directly input into the model for training, the model cannot fully learn the semantic information of the palm.
[0028] In the embodiments of the present application, first, the first human feature information of the target user in the first image frame is obtained, which includes first body feature information and first hand feature information; and the depth information of the body and the palm is included in the human feature information, and the depth information of the body and the palm is spliced respectively for overall standardization to obtain the first human feature information. Then, the first human feature information is input into a translation model. The translation model is different from the traditional translation model. The traditional translation model is divided into three parts in the present application, and different weights are used for the feature information of the hand and the body, respectively, for processing in the time dimension and the spatial dimension. Specifically, the first weight and the second weight are used to process the first body feature information to obtain the second body feature information, and the second weight and the third weight are used to process the first hand feature information to obtain the second hand feature information; the second body feature information and the second hand feature information are spliced to obtain the second human feature information. Finally, the second human feature information is classified and processed to output the sign language semantic information of the target user.
[0029] Thus, after the depth information of the body and the palm in the human feature information is obtained, more accurate body feature information and hand feature information can be extracted according to the distances of the body, the palm and the camera, so that the more accurate body feature information and hand feature information of the target user can be input into the translation model provided in the application, the current frame body feature information and the previous and next frame body feature information are fused in the time dimension by using the first weight, and correspondingly, the current frame hand feature information and the previous and next frame hand feature information are fused in the time dimension by using the third weight, and at the same time, the body feature information and the hand feature information are fused in the spatial dimension by using the second weight, so that the electronic device can not only learn the sign language semantic information of the target user more fully according to the fused body feature information and hand feature information, but also the traditional human feature translation model is split into a body feature translation model and a hand feature translation model in the new translation model, thereby reducing the parameter amount between models and facilitating reducing model overfitting.
[0030] The execution subject of the sign language recognition method provided in the embodiments of the application can be a sign language recognition device, which can be an electronic device or a functional module in the electronic device. Hereinafter, the electronic device will be taken as an example for illustration.
[0031] The embodiments of the application provide a sign language recognition method, Figure 1 A flowchart of a sign language recognition method provided in the embodiments of the application is shown, and the method can be applied to an electronic device. As shown in Figure 1 The sign language recognition method provided in the embodiments of the application can include the following steps 201 to 204.
[0032] Step 201, obtaining first human feature information of a target user in a first image frame.
[0033] In the embodiments of the application, the human feature information includes first body feature information and first hand feature information.
[0034] In the embodiments of the application, the first hand feature information includes first left hand feature information and first right hand feature information.
[0035] It should be noted that the body feature information in the embodiments of the application refers to the head, the trunk and the limb parts except the hand of the target user, such as the elbow joint, the shoulder joint, etc. Meanwhile, the hand feature information in the embodiments of the application refers to the joint features of the hand of the target user, such as the finger joint, the wrist joint, etc.
[0036] In the embodiments of the application, the first image frame is one of all image frames in a sign language video in which the target user participates.
[0037] In the embodiments of the present application, the sign language video can include a recorded sign language video, a video of a user performing a sign language action in a real-time call environment, and the like.
[0038] In step 202, the first human feature information is input into a translation model, and the first body feature information and the first hand feature information are respectively weighted and processed to obtain second body feature information and second hand feature information.
[0039] For example, after the first human feature information is input into the translation model, the first body feature information and the first hand feature information are respectively weighted and processed in different modules by using different weights, so as to obtain the second body feature information and the second hand feature information.
[0040] Optionally, in the embodiments of the present application, in the process of step 202, the first body feature information and the first hand feature information are respectively weighted and processed to obtain the second body feature information and the second hand feature information, the following step 202a is included:
[0041] In step 202a, the first human feature information is input into the translation model, the first body feature information is processed by using a first weight and a second weight to obtain the second body feature information, and the first hand feature information is processed by using the second weight and a third weight to obtain the second hand feature information.
[0042] For example, the first weight is used to represent the relevance between the body feature information in the image frame before the first image frame and the first body feature information. It can be understood that the first weight represents the relevance between the body feature information in the image frame before the first image frame and the first body feature information in the time dimension.
[0043] For example, the second weight is used to represent the relevance between the first body feature information and the first hand feature information. It can be understood that the second weight represents the relevance between the first body feature information and the first hand feature information in the spatial dimension.
[0044] For example, the third weight is used to represent the relevance between the hand feature information in the image frame before the first image frame and the first hand feature information. It can be understood that the third weight represents the relevance between the hand feature information in the image frame before the first image frame and the first hand feature information in the time dimension.
[0045] Optionally, in the embodiments of the present application, the translation model includes a first multi-head attention module, a second multi-head attention module, a three-section multi-head attention module residual and a normalization module, and a front feedback module.
[0046] In a possible embodiment, the step 202a of processing the first body feature information by using the first weight and the second weight to obtain second body feature information includes steps 202a1-202a4.
[0047] The step 202a1 includes processing the first body feature information by using the first weight based on a first multi-head attention module to obtain third body feature information.
[0048] The step 202a2 includes processing the first body feature information by using the second weight based on a three-section multi-head attention module to obtain fourth body feature information.
[0049] For example, in the three-section multi-head attention module, the first body feature information is processed based on a first formula, a second formula and a third formula to obtain the fourth body feature information.
[0050] For example, the first formula is as follows:
[0051] For example, the second formula is as follows:
[0052] For example, the third formula is as follows:
[0053] wherein a=1, 2, 3 respectively represent left hand feature information, right hand feature information and body feature information; b=1, 2, 3 respectively represent left hand feature information, right hand feature information and body feature information.
[0054] z a represents the attention-weighted subpart feature information (i.e., the fourth body feature information); a ab represents the normalized weight of the bth subpart when calculating the ath subpart vector.
[0055] It should be noted that the three-section multi-head attention module is used to calculate the self-attention weight and the mutually weighted feature information of the body, the left hand and the right hand in the same frame of image. Since there is no time sequence relationship when the left hand, the right hand and the body feature are associated with each other, the position information does not need to be preserved, so the relative position encoding does not need to be added in this module, and there is also no absolute position encoding feature.
[0056] The step 202a3 includes calculating the mean and the standard deviation value corresponding to the third body feature information and the fourth body feature information based on the residual and standardization module, and processing the third body feature information and the fourth body feature information based on the mean and the standard deviation value.
[0057] Step 202a4, based on the front feedback module, all feature information in the processed third body feature information and the fourth body feature information is fused to obtain the second body feature information.
[0058] In one possible embodiment, in the process of the step 202a "processing the first hand feature information by using the second weight and the third weight to obtain the second hand feature information", steps 202a5 to 202a8 are included:
[0059] Step 202a5, based on the second multi-head attention module, the first hand feature information is processed by using the third weight to obtain the third hand feature information.
[0060] Step 202a6, based on the three-section multi-head attention module, the first hand feature information is processed by using the second weight to obtain the fourth hand feature information.
[0061] Exemplarily, the process of processing the first hand feature information based on the three-section multi-head attention module is the same as the process of processing the first body feature information based on the three-section multi-head attention module in the step 202a2, which will not be repeated here.
[0062] Step 202a7, based on the residual and standardization module, the mean and standard deviation values corresponding to the third hand feature information and the fourth hand feature information are calculated, and the third hand feature information and the fourth hand feature information are processed based on the mean and standard deviation values.
[0063] Step 202a8, based on the front feedback module, all feature information in the processed third hand feature information and the fourth hand feature information is fused to obtain the second hand feature information.
[0064] Exemplarily, first, the first body feature information and the first hand feature information are spliced to obtain the first human body feature information, which is input into the transformer model. The transformer model provided in the embodiments of the present application is, for example, Figure 2As shown, after input, each sign language gesture is split into three sub-gestures (left-hand gesture, right-hand gesture, and body gesture, i.e., the first left-hand feature information, the first right-hand feature information, and the first body feature information). Then, in the attention module (i.e., the first multi-head attention module and the second multi-head attention module, which can also be called the body-multi-head attention module and the hand-multi-head attention module) and the residual and normalization module (Add&Norm), the three sub-gestures are first subjected to self-attention learning in the temporal dimension, and then mutual attention learning of the three sub-gestures is performed in the same frame. Next, the feedforward network module is input to further fuse these feature information. Finally, residual and normalization processing is performed on the three sub-vectors separately. It should be noted that we use the same self-attention module for the left and right hands because the left and right hand gestures are symmetrical in sign language. By processing the left hand gesture symmetrically, it can be processed uniformly with the right hand gesture, so it can share the same multi-head attention module.
[0065] Generally, traditional transformers ensure the sequential relationship of video frames by adding relative or absolute position encoding to the features of each frame. Combined with... Figure 2 ,like Figure 3 As shown, the body-multi-head attention module and the hand-multi-head attention module (i.e., the first multi-head attention module and the second multi-head attention module) use relative position encoding, consistent with traditional self-attention modules. The model results do not require modification. The first and third weights are used to calculate the self-attention weights and weighted feature information of body features and left and right hand features in the time dimension. Next, the three-segment multi-head attention module uses the second weight to calculate the self-attention weights and weighted feature information of the body, left hand, and right hand in the same image frame. Finally, body feature information and hand feature information that are correlated in both the time and spatial dimensions are obtained. At this point, the features of the body feature information and hand feature information are also fused together.
[0066] Thus, by splitting the traditional transformer, the model parameters are reduced: assuming that the model has L transformer layers in total, and the vector dimension of each sub-pose is dim, then the attention parameter amount before splitting is: L*3*(3*dim)2=27*L*dim2, and the attention parameter amount after splitting is: L*3*2*(dim)2=6*L*dim2. Taking L as about 10 and dim as 100 dimensions as an example, about 2 million parameters can be reduced. For sign language recognition with less training data, reducing the number of parameters is beneficial to reduce model overfitting.
[0067] In step 203, the second body feature information and the second hand feature information are spliced to obtain second human body feature information.
[0068] Exemplarily, the split second body feature information and the second hand feature information of the same image frame are spliced to obtain complete second human body feature information.
[0069] Exemplarily, the second body feature information obtained by processing the first body feature information in the translation module by using the first weight and the second weight, and the second hand feature information obtained by processing the first hand feature information of the same image frame by using the second weight and the third weight, are correspondingly spliced to obtain the second human body feature information in the complete image frame.
[0070] In step 204, the second human body feature information is classified and processed to output target user sign language semantic information.
[0071] Optionally, in the embodiment of the present application, in the process of step 204, "classifying and processing the second human body feature information to output target user sign language semantic information", step 204a and step 204b are included.
[0072] In step 204a, the second human body feature information is input into a semantic analysis model to obtain semantic analysis information that has a mapping relationship with the second human body feature information, and based on the semantic analysis information, target prediction parameters are obtained.
[0073] Exemplarily, the target prediction parameters include the probability that the semantic of the user sign language embodied by the second human body feature information belongs to different preset semantics.
[0074] Exemplarily, the preset semantics are semantics in a preset semantic library of the system.
[0075] In step 204b, target user sign language semantic information is obtained based on the target prediction parameters.
[0076] Exemplarily, the spliced second human feature information is input into a semantic analysis model, passes through a full connection layer and a RELU activation layer, and then passes through another full connection layer to obtain semantic analysis information of the second human feature information existing in a mapping relationship, i.e., an output n-dimensional vector. Finally, the n-dimensional vector passes through a softmax function to obtain a prediction parameter of target user sign language semantic information corresponding to the second human feature information, and a class corresponding to a probability maximum of the prediction parameter is the sign language word class corresponding to the sign language video.
[0077] In addition, when training the semantic analysis model, the n-dimensional vector and the label y can be put into a cross-entropy loss function for learning.
[0078] In the sign language recognition method provided in the embodiment of the application, first human feature information of a target user in a first image frame is obtained, the human feature information including first body feature information and first hand feature information; the first human feature information is input into a translation model, and the first body feature information and the first hand feature information are respectively subjected to weighted processing to obtain second body feature information and second hand feature information; the second body feature information and the second hand feature information are spliced to obtain second human feature information; and the second human feature information is subjected to classification processing to output sign language semantic information of the target user. In this way, the body feature information and the hand feature information of the target user are input into the translation model provided in the application, the current frame body feature information and the previous and next frame body feature information are fused in the time dimension by using weighted processing, the current frame hand feature information and the previous and next frame hand feature information are fused in the time dimension by using weighted processing, and the body feature information and the hand feature information are fused in the space dimension by using weighted processing, so that the electronic device can not only learn the sign language semantic information of the target user more fully according to the fused body feature information and hand feature information, but also the traditional human feature translation model is split into a body feature translation model and a hand feature translation model in the new translation model, thereby reducing the parameter quantity between models and being conducive to reducing model overfitting.
[0079] Optionally, in the embodiment of the application, before the step 201 of obtaining first human feature information of a target user in a first image frame, the sign language recognition method provided in the embodiment of the application further includes steps 301 to 303:
[0080] The step 301 includes obtaining joint information of a human joint of the target user in the first image frame.
[0081] Exemplarily, the human joint includes a body joint and a hand joint, and the hand joint includes a left hand joint and a right hand joint.
[0082] Exemplarily, the mediapipe toolkit can be used to obtain joint information of the body joints of the target user.
[0083] Exemplarily, the joint information can include a feature sequence composed of coordinate information corresponding to the skeletal joint of the target user, and can also include the joint position of the human body joint.
[0084] For example, as shown in the figure, the skeletal joints of the body joints include the head, the torso, and the limb parts except the hands (for example, the head nodes 0-10, the torso nodes 11, 12, 23, and 24, and the limb nodes 13 and 14), and the skeletal joints of the hand joints include the skeletal joints of the left hand and the right hand (for example, the left hand skeletal joints 0-20 and the right hand skeletal joints 0-20). Figure 4
[0085] Step 302: splicing the joint information of the body joints of the target user to obtain first body joint information, and inputting the first body joint information into a fixed sequence length normalization module for feature extraction to obtain the first body feature information.
[0086] Step 303: obtaining first hand joint information based on the joint information of the body joints of the target user, and inputting the first hand joint information into the fixed sequence length normalization module for feature extraction to obtain the first hand feature information.
[0087] Exemplarily, in the case of a fixed length sequence, the relative size relationship of the skeletal joint coordinates in the spatial dimension and the time dimension can be retained.
[0088] Exemplarily, the first hand joint information includes first right hand joint information and first left hand joint information.
[0089] In an example, taking the right hand as an example, it is assumed that the right hand has m skeletal joints, and the x and y coordinates of each skeletal joint plus depth information (i.e., the first right hand joint information) are 2m+1-dimensional feature vectors. The right hand features of the continuous k frames (k, 2m+1) are spliced into k*(2m+1)-dimensional features, and then the fourth formula is used for normalization. The normalized features are restored to the original (k, 2m+1) shape. The parameters in the fourth formula are obtained by the fifth formula and the sixth formula .
[0090] In this way, the improved normalization method in the present application simultaneously contains the advantages of batch normalization (batch norm) and layer normalization (layer norm), thereby simultaneously retaining the relative size relationship of the skeletal joint coordinates in the spatial dimension and the time dimension, and conforming to the normal distribution.
[0091] Optionally, in this embodiment of the application, step 303, "obtaining first hand joint information based on the joint information of the target user's human body joints," includes steps 303a to 303c:
[0092] Step 303a: Calculate the shoulder width information of the target user based on the joint information of the target user's human body joints.
[0093] For example, the human shoulder width information may include the shoulder width length of the target user's body and the position of the target user's shoulder width.
[0094] For example, the seventh formula is used to calculate the target user's shoulder width information based on the joint information of the target user's human body.
[0095] For example, the seventh formula is:
[0096] Among them, L cd x represents the width of the human shoulder, and x and y represent the coordinates of the two sides of the human shoulder.
[0097] Step 303b: Based on the human shoulder width information and the joint information of the target user's hand joints, construct a target coordinate system.
[0098] For example, the target coordinate system is a coordinate system with the target user's shoulder width as the side length and the center of the target user's hand as the center.
[0099] For example, the coordinates of the center point of the target user's hand are calculated using the eighth formula.
[0100] For example, the eighth formula is:
[0101] in, The coordinates of the center point of the hand;
[0102] x i y i Let be the coordinates of the m-th hand joint.
[0103] In one example, taking the right hand joint as an example, such as Figure 5 As shown, with the center point 51 of the hand as the center, and the shoulder width L as the center, cd Draw a square EFGH with sides of length EFGH. Since the coordinates of the hand's center point and shoulder width are already determined, the formula can be used... and calculation Find the coordinates of the four vertices of the square. Then, establish the target coordinate system with vertex E of the square as the origin of the target coordinate system and vertex G of the square as the coordinate point (1,1) of the target coordinate system.
[0104] Step 303c, mapping the joint information of the hand joint of the target user to the target coordinate system to obtain the first hand joint information.
[0105] Exemplarily, the hand joint information is mapped into the target coordinate system by using a ninth formula to obtain new joint information.
[0106] Exemplarily, the ninth formula is: i = (x i -x e ) / (x g -x e ), γ i = (y i -y e ) / (y g -y e )
[0107] wherein χ i , γ i are new coordinate points of the hand joint in the target coordinate system.
[0108] Further exemplarily, after mapping the joint information of the hand joint of the target user to the target coordinate system, a tenth formula is used to calculate the first hand joint information.
[0109] Exemplarily, the tenth formula is:
[0110] wherein A ij is the dispersion of the hand joint in the target coordinate system.
[0111] It should be noted that the new coordinate system is established because the dispersion of the hand joint is not only related to the depth information of the hand from the body, but also related to the distance of the hand from the camera, so in order to obtain accurate depth information of the hand from the body, the influence of the body distance from the camera needs to be eliminated.
[0112] In this way, by establishing a new coordinate system with the shoulder width of the human body as the basis, and scaling the hand key point coordinates according to the distance of the body from the camera, the influence can be offset, and the hand key point dispersion calculated in this way can represent the hand depth information.
[0113] The sign language recognition method provided by the present application will be exemplarily described below with a specific sign language video. Specifically, as shown in Figure 6 , the method can include the following steps 101 to 106:
[0114] Step 101: Extract the skeleton joints of the person in the sign language video, obtain the coordinates of the body joints of the user in the sign language video (i.e., the joint information of the body joints), and calculate the shoulder width of the body based on the coordinates of the body joints.
[0115] It should be noted that the patent splits each sign language posture of the human body into: body posture + left hand posture + right hand posture. The advantage of this is to reduce the number of sign language postures. Assuming that the body, left hand, and right hand each have 100 different postures, if not split, 100*100*100 sign language postures can be formed, which increases the difficulty of the model to learn the relevance and importance of each posture; on the contrary, after splitting, there are a total of 300 postures, and the model is easier to learn the relevance and importance between each posture.
[0116] Step 102: Calculate the depth information of the left and right hands of the user in the sign language video (i.e., the first hand joint information), and splice the depth information to the left and right hand features.
[0117] Step 103: Perform fix-length norm standardization processing on the spliced features of the left and right hands and the body features of the continuous frames (i.e., the first body joint information).
[0118] Step 104: After the standardization processing of the spliced features of the left and right hands and the body features (i.e., the first body feature information), input the transformer model to extract the spatiotemporal attention weighted features of the body, left and right hands (i.e., the second body feature information and the second hand feature information).
[0119] It should be noted that the weighting process can refer to step 202 shown in the above, which will not be repeated here.
[0120] Step 105: Splice the spatiotemporal attention weighted features of the body, left and right hands (i.e., the second body feature information).
[0121] At this time, the body features not only weight and integrate the body features in other time dimensions, but also weight and integrate the left and right hand features in the current frame. The same is true for the left and right hand features.
[0122] Step 106: Put the transformer encoding features into the classifier for classification to obtain the final discrete word sign language semantic information.
[0123] Therefore, the hand sign recognition method provided in the application can obtain hand depth information by using a new method, a new normalization scheme combining the advantages of layer norm and batch norm, and a new input structure and attention mechanism of a transformer, so that the model can not only learn palm gesture information sufficiently, but also greatly reduce the model parameters, thereby reducing the overfitting of the model and improving the accuracy of recognizing semantic information of the hand sign.
[0124] It should be noted that the execution subject of the hand sign recognition method provided in the application can be a hand sign recognition device or an electronic device, and can also be a functional module or an entity in the electronic device. The hand sign recognition method executed by the hand sign recognition device is taken as an example in the application embodiment to illustrate the hand sign recognition device provided in the application embodiment.
[0125] Figure 7 A possible structural schematic diagram of the hand sign recognition device involved in the application embodiment is shown. As shown in the figure, Figure 7 The hand sign recognition device 700 can include an acquisition module 701 and a processing module 702. The acquisition module 701 is configured to acquire first human feature information of a target user in a first image frame, the human feature information including first body feature information and first hand feature information. The processing module 702 is configured to input the first human feature information into a translation model, and perform weighted processing on the first body feature information and the first hand feature information respectively to obtain second body feature information and second hand feature information. The processing module 702 is further configured to splice the second body feature information and the second hand feature information to obtain second human feature information. The processing module 702 is further configured to perform classification processing on the second human feature information, and output semantic information of a hand sign of the target user.
[0126] Optionally, in the application embodiment, the processing module 702 is specifically configured to input the first human feature information into the translation model, process the first body feature information by using a first weight and a second weight to obtain second body feature information, and process the first hand feature information by using the second weight and a third weight to obtain second hand feature information. The first weight is used to represent the relevance between body feature information in an image frame before the first image frame and the first body feature information. The second weight is used to represent the relevance between the first body feature information and the first hand feature information. The third weight is used to represent the relevance between hand feature information in an image frame before the first image frame and the first hand feature information.
[0127] Optionally, in the embodiment of the present application, the processing module 702 is specifically configured to: based on the first multi-head attention module, process the first body feature information by using the first weight to obtain third body feature information; based on the three-section multi-head attention module, process the first body feature information by using the second weight to obtain fourth body feature information; based on the residual and standardization module, calculate the mean and standard deviation values corresponding to the third body feature information and the fourth body feature information, and process the third body feature information and the fourth body feature information based on the mean and standard deviation values; based on the front feedback module, fuse all feature information in the processed third body feature information and the fourth body feature information to obtain the second body feature information.
[0128] Optionally, in the embodiment of the present application, the processing module 702 is specifically configured to: based on the second multi-head attention module, process the first hand feature information by using the third weight to obtain third hand feature information; based on the three-section multi-head attention module, process the first hand feature information by using the second weight to obtain fourth hand feature information; based on the residual and standardization module, calculate the mean and standard deviation values corresponding to the third hand feature information and the fourth hand feature information, and process the third hand feature information and the fourth hand feature information based on the mean and standard deviation values; based on the front feedback module, fuse all feature information in the processed third hand feature information and the fourth hand feature information to obtain the second hand feature information.
[0129] Optionally, in the embodiment of the present application, the acquisition module 701 is further configured to acquire joint information of a human joint of a target user in a first image frame, the human joint including a body joint and a hand joint; the processing module 702 is further configured to splice the joint information of the body joint of the target user to obtain first body joint information, and input the first body joint information into a fixed sequence length standardization module for feature extraction to obtain the first body feature information; the processing module 702 is further configured to obtain first hand joint information based on the joint information of the human joint, and input the first hand joint information into the fixed sequence length standardization module for feature extraction to obtain the first hand feature information.
[0130] Optionally, in the embodiment of the present application, the processing module 702 is specifically configured to: calculate human shoulder width information of the target user based on the joint information of the human joint; construct a target coordinate system based on the human shoulder width information and the joint information of the hand joint of the target user, the target coordinate system being a coordinate system with the human shoulder width of the target user as the side length and the center of the hand of the target user as the center; and map the joint information of the hand joint of the target user to the target coordinate system to obtain the first hand joint information.
[0131] Optionally, in the embodiment of the present application, the processing module 702 is specifically configured to: input the second human feature information into a semantic analysis model to obtain semantic analysis information that has a mapping relationship with the second human feature information, and obtain target prediction parameters based on the semantic analysis information; the target prediction parameters include a probability that a semantic of a sign language of a user embodied by the second human feature information belongs to different preset semantics; and obtain target user sign language semantic information based on the target prediction parameters.
[0132] In the sign language recognition apparatus provided in the embodiment of the present application, the apparatus obtains first human feature information of a target user in a first image frame, the human feature information including first body feature information and first hand feature information; inputs the first human feature information into a translation model to perform weighted processing on the first body feature information and the first hand feature information respectively to obtain second body feature information and second hand feature information; splices the second body feature information and the second hand feature information to obtain second human feature information; and performs classification processing on the second human feature information to output target user sign language semantic information. In this way, the body feature information and the hand feature information of the target user are input into the translation model provided in the present application, the current frame body feature information and the previous and next frames of body feature information are fused in the time dimension by using weighted processing, and correspondingly, the current frame hand feature information and the previous and next frames of hand feature information are fused in the time dimension by using weighted processing, and at the same time, the body feature information and the hand feature information are fused in the space dimension by using weighted processing, so that the electronic device can not only learn the target user sign language semantic information more fully based on the fused body feature information and hand feature information, but also the traditional human feature translation model is split into a body feature translation model and a hand feature translation model in the new translation model, thereby reducing the parameter amount between models and being conducive to reducing model overfitting.
[0133] The sign language recognition apparatus in the embodiments of the present applicationapplicationbe an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic deviceapplicationbe a terminal or other device than a terminal. For example, the electronic deviceapplicationbe a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), or the like, andapplicationbe a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, and the like, and the embodiments of the present application do not make a specific limitation.
[0134] The sign language recognition apparatus in the embodiments of the present applicationapplicationbe a device having an operating system. The operating systemapplicationbe an Android operating system, an ios operating system, or other possible operating system, and the embodiments of the present application do not make a specific limitation.
[0135] The sign language recognition apparatus provided in the embodiments of the present applicationapplicationimplement the method embodiments, and each process of the method embodimentsapplicationbe implemented by the sign language recognition apparatus, and thus will not be repeated here. Figure 7 The method embodimentsapplicationimplement each process of the method embodiments, and thus will not be repeated here.
[0136] Optionally, as shown in Figure 8 The embodiments of the present application further provide an electronic device 800, whichapplicationinclude a processor 801 and a memory 802, and the memory 802applicationstore programs or instructions whichapplicationbe run on the processor 801. The programs or instructionsapplicationbe executed by the processor 801 to implement each step of the sign language recognition method embodiments and achieve the same technical effects, and thus will not be repeated here.
[0137] It should be noted that the electronic device in the embodiments of the present applicationapplicationinclude the mobile electronic device and the non-mobile electronic device.
[0138] Figure 9 A hardware structure schematic diagram of an electronic device for implementing the embodiments of the present application.
[0139] The electronic device 100 includes, but is not limited to, a radio frequency unit 101, a network module 102, an audio output unit 103, an input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, and a processor 110, etc.
[0140] As can be appreciated by those skilled in the art, the electronic device 100 can further include a power supply (such as a battery) that supplies power to each component, and the power supply can be logically connected to the processor 110 through a power management system, so that the power management system can realize functions such as management of charging, discharging, and power consumption management. Figure 9 The electronic device structure shown in the figure does not constitute a limitation on the electronic device, and the electronic device can include more or fewer components than shown, or combine certain components, or different component arrangements, which are not described here.
[0141] The processor 110 is configured to obtain first human feature information of a target user in a first image frame, the human feature information including first body feature information and first hand feature information; the processor 110 is further configured to input the first human feature information into a translation model, and perform weighted processing on the first body feature information and the first hand feature information respectively to obtain second body feature information and second hand feature information; the processor 110 is further configured to splice the second body feature information and the second hand feature information to obtain second human feature information; and the processor 110 is further configured to perform classification processing on the second human feature information and output sign language semantic information of the target user.
[0142] Optionally, in the embodiments of the present application, the processor 110 is specifically configured to input the first human feature information into a translation model, perform processing on the first body feature information using a first weight and a second weight to obtain second body feature information, and perform processing on the first hand feature information using the second weight and a third weight to obtain second hand feature information; wherein the first weight is used to represent the relevance between body feature information in an image frame before the first image frame and the first body feature information; the second weight is used to represent the relevance between the first body feature information and the first hand feature information; and the third weight is used to represent the relevance between hand feature information in an image frame before the first image frame and the first hand feature information.
[0143] Optionally, in the embodiment of the present application, the processor 110 is specifically configured to: based on the first multi-head attention module, process the first body feature information by using the first weight to obtain third body feature information; based on the three-section multi-head attention module, process the first body feature information by using the second weight to obtain fourth body feature information; based on the residual and standardization module, calculate the mean and standard deviation values corresponding to the third body feature information and the fourth body feature information, and process the third body feature information and the fourth body feature information based on the mean and standard deviation values; based on the front feedback module, fuse all feature information in the processed third body feature information and the fourth body feature information to obtain the second body feature information.
[0144] Optionally, in the embodiment of the present application, the processor 110 is specifically configured to: based on the second multi-head attention module, process the first hand feature information by using the third weight to obtain third hand feature information; based on the three-section multi-head attention module, process the first hand feature information by using the second weight to obtain fourth hand feature information; based on the residual and standardization module, calculate the mean and standard deviation values corresponding to the third hand feature information and the fourth hand feature information, and process the third hand feature information and the fourth hand feature information based on the mean and standard deviation values; based on the front feedback module, fuse all feature information in the processed third hand feature information and the fourth hand feature information to obtain the second hand feature information.
[0145] Optionally, in the embodiment of the present application, the processor 110 is further configured to obtain joint information of a human joint of a target user in a first image frame, the human joint including a body joint and a hand joint; the processor 110 is further configured to splice the joint information of the body joint of the target user to obtain first body joint information, and input the first body joint information into a fixed sequence length standardization module for feature extraction to obtain the first body feature information; the processor 110 is further configured to obtain first hand joint information based on the joint information of the human joint, and input the first hand joint information into the fixed sequence length standardization module for feature extraction to obtain the first hand feature information.
[0146] Optionally, in the embodiment of the present application, the processor 110 is specifically configured to: calculate human shoulder width information of the target user based on the joint information of the human joints; construct a target coordinate system based on the human shoulder width information and the joint information of the hand joints of the target user, the target coordinate system being a coordinate system with the human shoulder width of the target user as the side length and the center of the hand of the target user as the center; and map the joint information of the hand joints of the target user to the target coordinate system to obtain the first hand joint information.
[0147] Optionally, in the embodiment of the present application, the processor 110 is specifically configured to: input the second human feature information into a semantic analysis model to obtain semantic analysis information that has a mapping relationship with the second human feature information, and obtain target prediction parameters based on the semantic analysis information; the target prediction parameters include a probability that a semantic of a sign language of a user embodied by the second human feature information belongs to different preset semantics; and obtain target user sign language semantic information based on the target prediction parameters.
[0148] In the electronic device provided in the embodiment of the present application, the electronic device obtains first human feature information of a target user in a first image frame, the human feature information including first body feature information and first hand feature information; inputs the first human feature information into a translation model to perform weighted processing on the first body feature information and the first hand feature information respectively to obtain second body feature information and second hand feature information; splices the second body feature information and the second hand feature information to obtain second human feature information; and performs classification processing on the second human feature information to output target user sign language semantic information. In this way, the body feature information and the hand feature information of the target user are input into the translation model provided in the present application, the current frame body feature information and the previous and next frames of body feature information are fused in the time dimension by using weighted processing, and the current frame hand feature information and the previous and next frames of hand feature information are fused in the time dimension by using weighted processing, and at the same time, the body feature information and the hand feature information are fused in the space dimension by using weighted processing, so that the electronic device can not only learn the target user sign language semantic information more fully based on the fused body feature information and hand feature information, but also the traditional human feature translation model is split into a body feature translation model and a hand feature translation model in the new translation model, thereby reducing the parameter quantity between models and being conducive to reducing model overfitting.
[0149] It should be understood that in the embodiments of the present application, the input unit 104 can include a graphics processing unit (GPU) 1041 and a microphone 1042. The graphics processing unit 1041 processes image data of a still picture or a video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 106 can include a display panel 1061, which can be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 can include two parts of a touch detection device and a touch controller. The other input devices 1072 can include, but are not limited to, a physical keyboard, function keys (such as volume control keys, on-off keys, etc.), a trackball, a mouse, a joystick, and the like, which will not be described here.
[0150] The memory 109 can be used to store software programs and various data. The memory 109 can mainly include a first storage area storing programs or instructions and a second storage area storing data, wherein the first storage area can store an operating system, application programs or instructions required by at least one function (such as a sound playing function, an image playing function, etc.), and the like. In addition, the memory 109 can include a volatile memory or a non-volatile memory, or the memory 109 can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synch link DRAM (SLDRAM), and a direct memory bus random access memory (Direct Rambus RAM, DRRAM). The memory 109 in the embodiments of the present application includes but is not limited to these and any other suitable types of memory.
[0151] The processor 110 can include one or more processing units; optionally, the processor 110 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and an application program, and the modem processor mainly processes a wireless communication signal, such as a baseband processor. It can be understood that the modem processor can also not be integrated into the processor 110.
[0152] The embodiment of the present application further provides a readable storage medium, and the readable storage medium stores a program or instructions, which are executed by a processor to implement various processes of the sign language recognition method embodiment and achieve the same technical effects. To avoid repetition, details are not described herein.
[0153] The processor is the processor in the electronic device in the embodiment. The readable storage medium includes a computer readable storage medium, such as a computer readable only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and the like.
[0154] The embodiment of the present application further provides a chip, and the chip includes a processor and a communication interface. The communication interface is coupled with the processor, and the processor is configured to run a program or instructions to implement various processes of the sign language recognition method embodiment and achieve the same technical effects. To avoid repetition, details are not described herein.
[0155] It should be understood that the chip mentioned in the embodiment of the present application can also be referred to as a system level chip, a system chip, a chip system, or a system on chip, and the like.
[0156] The embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement various processes of the sign language recognition method embodiment and achieve the same technical effects. To avoid repetition, details are not described herein.
[0157] It should be noted that, in the present document, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element. Furthermore, it is to be understood that the method and apparatus of the present application can be carried out by more than one process, method, article, or apparatus either simultaneously, concurrently, or with intervening action that are carried out at the same time, either in a simultaneous fashion or in a fashion that is staggered in time. For example, the described methods can be carried out in a different order than described, and / or various steps can be combined or omitted, and / or additional steps can be added, without departing from the scope of the present application. Also, features described with respect to certain examples can be combined in other examples.
[0158] From the above description of the embodiments, it is apparent that the method of the embodiments can be implemented by means of software and the requisite general purpose hardware platform, of course, can also be implemented by hardware, but in many cases the former is the preferred implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product stored in a storage medium (such as ROM / RAM, magnetic disc, optical disc), including a number of instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the method described in the various embodiments of the present application.
[0159] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the specific embodiments described, which are merely illustrative and not restrictive, and a person of ordinary skill in the art can make many forms under the inspiration of the present application without departing from the scope of the present application and the protection scope of the claims.
Claims
1. A sign language recognition method, characterized by, The method comprises: obtaining first human feature information of a target user in a first image frame, wherein the human feature information comprises first body feature information and first hand feature information; inputting the first human feature information into a translation model, processing the first body feature information by using a first weight based on a first multi-head attention module to obtain third body feature information; processing the first body feature information by using a second weight based on a three-section multi-head attention module to obtain fourth body feature information; calculating mean and standard deviation values corresponding to the third body feature information and the fourth body feature information based on a residual and standardization module, and processing the third body feature information and the fourth body feature information based on the mean and standard deviation values; fusing all feature information in the processed third body feature information and the fourth body feature information based on a front feedback module to obtain second body feature information, and processing the first hand feature information by using the second weight and a third weight to obtain second hand feature information; and splicing the second body feature information and the second hand feature information to obtain second human feature information; classifying and processing the second human feature information to output sign language semantic information of the target user; wherein the translation model comprises a first multi-head attention module, a three-section multi-head attention module, a residual and standardization module, and a front feedback module; the first weight is used to represent the relevance between body feature information in image frames before the first image frame and the first body feature information; the second weight is used to represent the relevance between the first body feature information and the first hand feature information; and the third weight is used to represent the relevance between hand feature information in image frames before the first image frame and the first hand feature information.
2. The method of claim 1, wherein, The translation model comprises a second multi-head attention module, a three-section multi-head attention module, a residual and standardization module, and a front feedback module. The processing of the first hand feature information by using the second weight and the third weight to obtain second hand feature information comprises: processing the first hand feature information by using the third weight based on the second multi-head attention module to obtain third hand feature information; processing the first hand feature information by using the second weight based on the three-section multi-head attention module to obtain fourth hand feature information; calculating mean and standard deviation values corresponding to the third hand feature information and the fourth hand feature information based on the residual and standardization module, and processing the third hand feature information and the fourth hand feature information based on the mean and standard deviation values; fusing all feature information in the processed third hand feature information and the fourth hand feature information based on the front feedback module to obtain the second hand feature information.
3. The method of claim 1, wherein, Before the obtaining of the first human feature information of the target user in the first image frame, the method further comprises: Obtaining joint information of human joints of the target user in the first image frame, the human joints including body joints and hand joints; Splicing joint information of body joints of the target user to obtain first body joint information, and inputting the first body joint information into a fixed sequence length normalization module for feature extraction to obtain first body feature information; Based on the joint information of the human joints, obtaining first hand joint information, and inputting the first hand joint information into the fixed sequence length normalization module for feature extraction to obtain first hand feature information.
4. The method of claim 3, wherein, The first hand joint information is obtained based on the joint information of the human joints, including: Based on the joint information of the human joints, calculating human shoulder width information of the target user; Based on the human shoulder width information and the joint information of the hand joints of the target user, constructing a target coordinate system, the target coordinate system being a coordinate system with the human shoulder width of the target user as the side length and the center of the hand of the target user as the center; Mapping the joint information of the hand joints of the target user to the target coordinate system to obtain the first hand joint information.
5. The method of claim 1, wherein, The second body feature information is classified and processed to output the sign language semantic information of the target user, including: Inputting the second body feature information into a semantic analysis model to obtain semantic analysis information that has a mapping relationship with the second body feature information, and obtaining target prediction parameters based on the semantic analysis information; the target prediction parameters include the probability that the semantic of the sign language of the user embodied by the second body feature information belongs to different preset semantics; Based on the target prediction parameters, obtaining the sign language semantic information of the target user.
6. A sign language recognition apparatus characterized by comprising: The sign language recognition device includes an acquisition module and a processing module; The acquisition module is configured to acquire first body feature information of a target user in a first image frame, the body feature information including first body feature information and first hand feature information; The processing module is configured to input the first body feature information acquired by the acquisition module into a translation model, process the first body feature information based on a first multi-head attention module using a first weight to obtain third body feature information; The processing module is further configured to process the first body feature information based on a three-section multi-head attention module using a second weight to obtain fourth body feature information; The processing module is further configured to calculate mean and standard deviation values corresponding to the third body feature information and the fourth body feature information based on a residual and normalization module, and process the third body feature information and the fourth body feature information based on the mean and standard deviation values; The processing module is further configured to fuse all feature information in the processed third body feature information and the fourth body feature information based on a front feedback module to obtain second body feature information, and process the first hand feature information using the second weight and a third weight to obtain second hand feature information; The processing module is further configured to splice the second body feature information and the second hand feature information to obtain second human body feature information. The processing module is further configured to perform classification processing on the second human body feature information, and output the target user sign language semantic information. The translation model comprises a first multi-head attention module, a three-section multi-head attention module residual and normalization module, and a front feedback module; the first weight is used to represent the relevance between the body feature information in the image frame before the first image frame and the first body feature information; the second weight is used to represent the relevance between the first body feature information and the first hand feature information; and the third weight is used to represent the relevance between the hand feature information in the image frame before the first image frame and the first hand feature information.
7. The apparatus of claim 6, wherein, The translation model comprises a second multi-head attention module, a three-section multi-head attention module residual and normalization module, and a front feedback module. The processing module is specifically configured to: based on the second multi-head attention module, process the first hand feature information by using the third weight to obtain third hand feature information; based on the three-section multi-head attention module, process the first hand feature information by using the second weight to obtain fourth hand feature information; based on the residual and normalization module, calculate the mean and standard deviation values corresponding to the third hand feature information and the fourth hand feature information, and process the third hand feature information and the fourth hand feature information based on the mean and standard deviation values; based on the front feedback module, fuse all feature information in the processed third hand feature information and the fourth hand feature information to obtain the second hand feature information.
8. The apparatus of claim 6, wherein The acquisition module is further configured to acquire joint information of human body joints of the target user in a first image frame, the human body joints comprising body joints and hand joints; The processing module is further configured to splice the joint information of the body joints of the target user acquired by the acquisition module to obtain first body joint information, and input the first body joint information into a fixed sequence length normalization module for feature extraction to obtain the first body feature information; The processing module is further configured to acquire first hand joint information based on the joint information of the human body joints, and input the first hand joint information into the fixed sequence length normalization module for feature extraction to obtain the first hand feature information.
9. The apparatus of claim 8, wherein The processing module is specifically configured to: calculate human shoulder width information of the target user based on the joint information of the human body joints; construct a target coordinate system based on the human shoulder width information and the joint information of the hand joints of the target user, the target coordinate system being a coordinate system with the human shoulder width of the target user as the side length and the center of the hand of the target user as the center. Map joint information of a hand joint of the target user to the target coordinate system to obtain first hand joint information.
10. The apparatus of claim 6, wherein, The processing module is specifically configured to: input the second human feature information into a semantic analysis model, acquire semantic analysis information in a mapping relationship with the second human feature information, and obtain target prediction parameters based on the semantic analysis information; the target prediction parameters include a probability that a semantic of a sign language of a user embodied by the second human feature information belongs to different preset semantics; obtain the target user sign language semantic information based on the target prediction parameters.
11. An electronic device, comprising: A processor, a memory, and a program or instructions stored on the memory and executable on the processor, the program or instructions being executed by the processor to implement the steps of the sign language recognition method of any one of claims 1 to 5.
12. A readable storage medium, characterized by, A program or instructions stored on the readable storage medium, the program or instructions being executed by the processor to implement the steps of the sign language recognition method of any one of claims 1 to 5.
Citation Information
Patent Citations
Sign language recognition method and system based on double-flow space-time diagram convolutional neural network
CN111325099A