A Real-time Sign Language Recognition Method for Video Streams Based on Human Key Points
Through the real-time video stream sign language recognition method based on human key points, the human body posture estimation and space-time graph convolution network combined with the encoder-decoder network are used to solve the real-time and accuracy problems of sign language recognition in the video stream, and the statement-level sign language recognition of real-time video stream is realized, reducing the influence of environmental factors.
Patent Information
- Application Number
- CN202211054559.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-31
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-08-31
AI Technical Summary
The existing continuous sign language recognition technology is difficult to recognize in real time in video streaming information, and is susceptible to factors such as character clothing and ambient lighting, resulting in confusion in translation semantics.
The real-time video stream sign language recognition method based on human body key points is adopted, and the video stream is read frame by frame, and the key points are extracted using the human body pose estimation network, combined with the space-time graph convolution network and the encoder-decoder network, the segmentation and feature extraction of sign language actions are realized, and the complete sentence is output.
The statement-level sign language recognition of real-time video streams is realized, which reduces the impact of clothing and ambient lighting and improves the accuracy of continuous sign language recognition.
Smart Images

Figure CN115457654B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computing, reckoning or counting, and particularly relates to a real-time video stream sign language recognition method based on human key points in the field of image processing and pattern recognition. Background Art
[0002] Hearing-impaired people cannot conveniently obtain information and express their wishes, and often face many difficulties in social interaction, education, employment, etc. This is because most hearing-impaired people usually use sign language for communication, but very few hearing people can understand sign language. As a visual language, sign language has differences in grammar and expression compared with the auditory language used by ordinary people. In different countries and regions, sign languages often also vary. Sign language recognition technology aims to translate sign languages in different regions and countries into corresponding written languages to solve the communication problems of hearing-impaired people.
[0003] Sign language recognition technology usually takes sign language images or videos as input, extracts features and classifies different sign language actions, and finally outputs text sentences. At present, sign language recognition is divided into isolated word recognition and continuous sentence recognition. The former is the recognition of a single sign language word, and the latter is the recognition of a complete sentence composed of a series of sign language words. Obviously, the continuous sentence sign language recognition is more practical. The current continuous sentence sign language recognition only focuses on the recognition of a single sentence, and there are often restrictions on the video length. For a video containing multiple sign language sentences, it needs to be artificially segmented. However, in practical applications, what is often faced is video stream information, and the usual continuous sign language recognition methods are difficult to perform real-time sign language recognition end-to-end.
[0004] The paper "Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition" introduced the spatio-temporal graph convolution method into the field of action recognition, which is of great significance for the research in the field of sign language recognition and is currently one of the commonly used methods for sign language recognition.
[0005] The encoder-decoder model is often used for sequence-to-sequence conversion problems. Continuous sign language recognition can also be regarded as a problem of converting a video sequence to a word sequence. Therefore, the encoder-decoder model is very effective for solving sign language recognition problems.
[0006] The Chinese patent with the application number CN202010301154.X discloses a sign language recognition method and system. This method first extracts feature frames from the collected video frames through a convolutional neural network, then inputs the feature frames into a preset hierarchical long short-term memory network to extract effective frames, and finally inputs the effective frames into a preset sign language recognition model to output the target sentence text aligned with the sign language video. This method performs feature extraction based on RGB images, and the recognition effect may be affected by factors such as the environment, and it is only applicable to the recognition of sign language videos within a certain length, and it is difficult to process video stream information.
[0007] The Chinese patent with the application number CN202010648991.X discloses a sign language recognition system and method based on spatio-temporal semantic features. This method first performs data preprocessing and frame segmentation on the input sign language video data, then extracts features from a series of video segments after frame segmentation through a spatio-temporal feature module, and then through semantic mining and decoding processing of the feature sequence, finally outputs the corresponding text information. This method uses a fixed-length frame segmentation strategy and is only applicable to the recognition scenario of a single sentence. When processing a video stream, it cannot well distinguish between the previous and subsequent sentences, easily leading to semantic confusion in translation. Summary of the Invention
[0008] The present invention solves the problems existing in the prior art and provides a real-time video stream sign language recognition method based on human key points. For real-time video streams, it solves the problem of sign language sentence segmentation in video streams. Based on human key points, it effectively reduces the influence of factors such as task clothing and environmental lighting on the algorithm; through the present invention, sentence-level sign language recognition can be performed on longer sign language video streams or real-time sign language video streams.
[0009] The technical solution adopted by the present invention is a real-time video stream sign language recognition method based on human key points, and the method includes the following steps:
[0010] Step 1: Read the input sign language video stream frame by frame;
[0011] Step 2: Use a human pose estimation network to extract human key points in any frame image read in Step 1, including but not limited to the nodes of the head, torso, and both hands, for identifying the human pose;
[0012] Step 3: Calculate the action difference degree between the current frame and the previous frame frame by frame and accumulate it;
[0013] Step 4: When the accumulated difference degree within t time is higher than the threshold T1, it is determined that the sign language action starts; when the difference degree is lower than the threshold T2, it is determined that the action is stationary, and a convolutional neural network is used to determine whether the current frame action is an end action; T1>T2>0;
[0014] Step 5: Save the human key point data of all frames within the time period from the start to the end of the sign language action to obtain a human key point sequence;
[0015] Step 6: Use a spatio-temporal graph convolutional network to extract features from the human key point sequence in Step 5 to obtain a feature sequence X;
[0016] Step 7: Use an encoder-decoder network, take the feature sequence X in Step 6 as the input, and output a complete sentence to achieve continuous sign language recognition.
[0017] Preferably, Step 2 includes the following steps:
[0018] Step 2.1: Input any frame image into a human pose estimation network and output key point information v,
[0019] v = {(x1, y1, c1), (x2, y2, c2), …, (x M , y M , c M )}
[0020] where M represents the number of output key points, and x i , yi, c i respectively represent the x coordinate, y coordinate and prediction confidence of the i-th key point, M ≥ 1, and i is the index of the key point;
[0021] Step 2.2: Screen the key points for sign language recognition, denoted as
[0022] v′ = {(x1, y1, c1), (x2, y2, c2), …, (x N , y N , c N )}
[0023] where N is the number of key points, 1 ≤ N ≤ M.
[0024] Preferably, Step 3 includes the following steps:
[0025] Step 3.1: Read the key point coordinates of the current frame,
[0026] P = {(x1, y1), (x2, y2), …, (x N , y N )}
[0027] Take P cur as the set of key point coordinates of the current frame, take P pre as the set of key point coordinates of the previous frame. If the current frame is the first frame of the video stream, then let P cur = P, P pre = P, otherwise, let P cur= P;
[0028] Step 3.2: Calculate the spatial difference degree δ of the key points in adjacent frames using the Euclidean distance between the key points corresponding to the current frame and the previous frame.
[0029]
[0030] Among them, x_cur i and y_cur i respectively represent the x - coordinate and y - coordinate of the i - th key point in the set P of key point coordinates of the current frame; x_pre cur and y_pre i respectively represent the x - coordinate and y - coordinate of the i - th key point in the set P of key point coordinates of the previous frame. i pre
[0031] Step 3.3: Repeat Step 3.2 and save the difference degree δ in a queue; the length of the queue is L, and L = t×fps, where t represents the time threshold. When a hearing - impaired person performs sign language, a short pause is used to indicate the end of a complete sentence. t is set according to the pause time and usually t = 0.3s can be taken, and fps represents the number of frames transmitted per second in the video stream.
[0032] Preferably, in Step 4, for the accumulated difference degree S within t time, when S > T1, it indicates the start of a sign - language action; when S < T2, it indicates a static action. Input the current frame image into the convolutional neural network. If it is judged as an invalid sign - language action, it indicates the end of the sign - language action. Appropriate values of T1 and T2 can be selected according to the number of nodes, and T1 > T2 can prevent frequent jumps.
[0033] In the present invention, the convolutional neural network is an image - classification network for binary classification of valid sign - language actions and invalid sign - language actions; among them, the invalid sign - language actions can be that both hands are naturally lowered or both hands are folded across the abdomen, etc.
[0034] Preferably, Step 5 includes the following steps:
[0035] Step 5.1: Save the sequence V of human key - point information of all frames of a sign - language action.
[0036] V = {v1, v2,..., v L}
[0037] Among them, L represents the number of frames of this sign - language action.
[0038] Step 5.2: Based on the input dimension of the spatio - temporal graph convolutional model, adjust the length of the key - point information sequence in Step 3.3 to T using the method of frame extraction or frame padding. in , T in Determined by the input dimension of the spatio-temporal graph convolution model.
[0039] Preferably, in step 6, the spatio-temporal graph convolution network performs spatial graph convolution and temporal graph convolution on the input data; the key points selected in step 2 are used as the nodes of the graph, connected into edges according to the human body structure to form spatial edges, and the same nodes in adjacent frames are connected into edges to form temporal edges.
[0040] Preferably, the spatio-temporal graph convolution network is composed of 10 spatio-temporal graph convolution units, and a global pooling layer is added to obtain a feature sequence.
[0041] X = (x1, x2,..., x T′ )
[0042] where T3 = 1, 2,... 4 in ;
[0043] Any spatio-temporal graph convolution unit includes a spatial graph convolution network and a temporal graph convolution network.
[0044] Preferably, the graph partitioning strategy of the spatial graph convolution adopts spatial configuration partitioning. The first subset connects the neighbor nodes that are further away from the entire skeleton in spatial position than the root node, representing the centrifugal movement in sign language. The second subset connects the neighbor nodes closer to the center, representing the centripetal movement in sign language. The third subset is the root node itself, representing the static movement in sign language.
[0045] Preferably, the spatial graph convolution formula is
[0046]
[0047] where f in represents the input feature sequence, with a dimension of where C in represents the dimension of the node information data, T in represents the number of input frames, and N represents the number of nodes; f out represents the output feature sequence, with a dimension of where C out represents the output feature dimension, and Tout represents the number of output frames; A j represents the normalized adjacency matrix formed according to the graph partitioning strategy; W j represents the weight matrix.
[0048] In the present invention, the temporal graph convolution learns the change features of the same nodes in the kernel_size key frames adjacent to the current frame with a convolution kernel of size (kernel_size, 1).
[0049] Preferably, step 7 includes the following steps:
[0050] Step 7.1: Input the feature sequence X obtained in Step 6 into the recurrent layer of the encoder to get the output o of the i-th recurrent unit i ,
[0051] o i = Encoder(x i , o i-1 )
[0052] where x i is the i-th feature vector in X, i is a positive integer, and o0 is a zero vector;
[0053] Step 7.2: Through the hidden state h j-1 at the previous moment and the word embedding g j-1 output at the previous moment, the decoder generates the output y j at the next moment, updates the hidden state h j ,
[0054] y j , h j = Decoder(g j-1 , h j-1 )
[0055] g j-1 = wordEmbedding(y j-1 )
[0056] where the initial hidden state h0 is the hidden state corresponding to the last encoding unit o [ in Step 7.1, and the identifier of the initial output y0 is set as the identifier of the start of the sequence;
[0057] Step 7.3: When the sequence end identifier appears, the output is completed to obtain the output word sequence Y = {y1, y2,..., y p}, and the output word sequence is concatenated to obtain a sentence.
[0058] The present invention relates to a real-time video stream sign language recognition method based on human key points. After reading the input sign language video stream frame by frame, a human pose estimation network is used to extract the human key points in any frame image read; the action difference degree between the current frame and the previous frame is calculated frame by frame and accumulated, and based on the accumulated difference degree within t time, it is judged whether the sign language action starts and the action is static, and a convolutional neural network is used to judge whether the current frame action is an end action; after the end, the human key point data of all frames within the time period from the start to the end of the sign language action is saved, a spatio-temporal graph convolutional network is used to extract features from the human key point sequence, and the obtained feature sequence X is input into the encoder-decoder network to output a complete sentence, realizing continuous sign language recognition.
[0059] The beneficial effects of the present invention are as follows:
[0060] (1) By segmenting the real-time sign language video stream, sentence-level sign language video segments are obtained, enabling continuous sign language recognition of the real-time video stream.
[0061] (2) Based on human key points, factors such as the clothing of the person and the environmental lighting can be avoided from affecting the algorithm.
[0062] (3) By adopting the method of combining spatio-temporal graph convolutional network with encoder-decoder structure, the accuracy of continuous sign language recognition at the sentence level can be effectively improved. Description of the Drawings
[0063] Figure 1 is the overall process schematic diagram of the present invention;
[0064] Figure 2 is the schematic diagram of human key points related to sign language recognition in the present invention;
[0065] Figure 3 is the flowchart of step 4 in the embodiment of the present invention. Detailed Embodiment
[0066] The present invention will be further described in detail below in conjunction with embodiments, but the protection scope of the present invention is not limited thereto.
[0067] As Figure 1 shown, the present invention relates to a flowchart of a real-time video stream sign language recognition method based on human key points, which includes the following steps:
[0068] Step 1: Read the input sign language video stream frame by frame;
[0069] Step 2: Use a human pose estimation algorithm to extract the human key points in the image frame read in step 1;
[0070] The said step 2 includes:
[0071] Step 2.1: Input the current image frame into the human pose estimation network and output the key point information:
[0072] v = {(x1, y1, c1), (x2, y2, c2),..., (x M , y M , c M )}
[0073] where M represents the number of output key points, and x i , y i , c i respectively represent the x coordinate, y coordinate and predicted confidence of the i-th key point, M ≥ 1, and i is the index of the key point;
[0074] In this embodiment, OpenPose is used for human pose estimation.
[0075] Step 2.2: Screen the human key points with relatively high influence weights on sign language recognition. In this embodiment, generally, 7 key points of the nose, both shoulders, both elbows, and both wrists, as well as 20 key points for each of the left and right hands are retained, including the palm root, 2 joints of the thumb and the tip of the thumb, 3 joints of the index finger and the tip of the index finger, 3 joints of the middle finger and the tip of the middle finger, 3 joints of the ring finger and the tip of the ring finger, and 3 joints of the little finger and the tip of the little finger, a total of 47 key points. As Figure 2 shown, denoted as:
[0076] v′ = {(x1, y1, c1), (x2, y2, c2),..., (x N , y N , c N )}
[0077] where N represents the number of key points, and 1 ≤ N ≤ M.
[0078] Step 3: Calculate the action difference degree between the current frame and the previous frame frame by frame and accumulate it;
[0079] The said Step 3 includes:
[0080] Step 3.1: Read the key point coordinates of the current frame:
[0081] P = {(x1, y1), (x2, y2),..., (x N , y N )}
[0082] Taking P cur as the key point coordinate set of the current frame, and P pre as the key point coordinate set of the previous frame. If the current frame is the first frame of the video stream, then let P cur = P, P pre = P, otherwise, let P cur = P;
[0083] Step 3.2: Calculate the spatial difference degree δ of the key points between adjacent frames using the Euclidean distance between the corresponding key points of the current frame and the previous frame.
[0084]
[0085] where x_cur i and y_cur i respectively represent the x coordinate and y coordinate of the i-th key point in the key point coordinate set P cur of the current frame; x_pre i and y_pre irespectively represent the set of key point coordinates P of the previous frame pre the x - coordinate and y - coordinate of the i - th key point in
[0086] Step 3.3: Repeat Step 3.2, and save the difference degree δ in a queue; the length of the queue is L, and L = t×fps, where t represents the time threshold and fps represents the number of frames transmitted per second in the video stream.
[0087] Step 4: When the accumulated difference degree within t time is higher than the threshold T1, it is determined that the sign language action starts; when the difference degree is lower than the threshold T2, it is determined that the action is static, and a convolutional neural network is used to determine whether the current frame action is an end action; T1>T2>0;
[0088] The program flowcharts matching Step 3.3 and Step 4 are as Figure 3 shown, including:
[0089] The length of the queue is L, and L = t×fps, where t represents the time threshold. When a hearing - impaired person performs sign language, a short pause is used to indicate the end of a complete sentence. t is set according to the pause time and usually t = 0.3s can be taken; fps represents the number of frames transmitted per second in the video stream;
[0090] In the said Step 4, for the accumulated difference degree S within t time, when S>T1, it indicates that the sign language action starts, when S<T2, it indicates that the action is static. Input the current frame image into the convolutional neural network. If it is judged as an invalid sign language action, it indicates that the sign language action ends.
[0091] Step 4.2 can further include:
[0092] Step 4.2.1: Obtain the status of the start flag flag. When flag = 0, it indicates that sign language performance has not been carried out; when flag = 1, it indicates that sign language performance is in progress; the initial value of flag is 0;
[0093] Step 4.2.2: If the queue is not full, then directly add the error sum to δ in Step 3.2, S = S + δ, and insert δ into the tail of the queue; if the queue is full, then pop δ at the head of the queue head , and the error sum S = S + δ - δ head , and then insert δ into the tail of the queue;
[0094] Step 4.2.3: When S > T1 and flag = 0, it indicates that the posture has changed significantly and no sign language performance has been carried out before, so it is determined as the start of a sign language performance, and flag is set to 1; when S < T2 and flag = 1, it indicates that the posture remains basically unchanged and a sign language performance is in progress, so it is determined as a static action. The current image frame is input into the convolutional neural network. If the output category is an invalid sign language action, it indicates the end of the sign language performance, and flag is set to 0; if the output category is a valid sign language action, it indicates that this static state is not the end of a sign language sentence.
[0095] In this embodiment, the convolutional neural network described in step 4.2 is an image classification network, which performs binary classification on valid sign language actions and invalid sign language actions; among them, invalid sign language actions can be natural lowering of both hands or folding both hands across the abdomen, etc.; image classification using a convolutional neural network is a well-known technology in the art, and those skilled in the art can select a suitable convolutional neural network based on their needs. For example, ResNet is used as the convolutional neural network in step 4.2.3 to achieve image classification; in this embodiment, picture samples of a sign language performer with natural lowering of both hands or folding both hands across the abdomen are labeled as invalid sign language actions, and other samples are labeled as valid sign language actions for model training, and the network model with the best accuracy on the validation set is selected.
[0096] In this embodiment, appropriate values of T1 and T2 can be selected according to the number of nodes, where T1 > T2 to prevent frequent jumps. For example, T1 = 15 and T2 = 20 are taken.
[0097] Step 5: Save the human key point data of all frames within the time period from the start to the end of the sign language action to obtain a human key point sequence.
[0098] The said step 5 includes:
[0099] Step 5.1: Save the sequence V of human key point information of all frames of a sign language action.
[0100] V = {v1, v2,..., v L}
[0101] where L represents the number of frames of this sign language action.
[0102] Step 5.2: Based on the input dimension of the spatio-temporal graph convolutional model, adjust the length of the key point information sequence in step 3.3 to T in T in determined by the input dimension of the spatio-temporal graph convolutional model.
[0103] In the present invention, the frame extraction method in step 5.2 is to delete one frame of key point information every L / d frames, where d is the number of frames to be deleted; the frame filling method in step 5.2 is to supplement vectors with the same dimension and a value of 0 at the end of the key point information sequence to make the sequence length supplemented to T in 。
[0104] Step 6: Use a spatio-temporal graph convolutional network to extract features from the human key point sequence in step 5 to obtain a feature sequence X;
[0105] In the said step 6, the spatio-temporal graph convolutional network performs spatial graph convolution and temporal graph convolution processing on the input data; using the key points selected in step 2 as the nodes of the graph, connecting them into edges according to the human body structure to form spatial edges, and connecting the same nodes in adjacent frames into edges to form temporal edges.
[0106] The graph partitioning strategy of the said spatial graph convolution adopts spatial configuration partitioning. In this embodiment, as Figure 2 shown, taking the nose node C as the skeleton center (root node), and taking the node of the elbow joint S of the hand as an example, the first subset connects the neighbor nodes O (wrists) whose spatial positions are farther from the entire skeleton than the root node (nose), indicating the centrifugal movement in sign language, the second subset connects the neighbor nodes I (shoulders) closer to the center, indicating the centripetal movement in sign language, and the third subset is the root node S itself, indicating the static movement in sign language. At this time, for the graph convolution of a node, its weight matrix contains three weight vectors, so as to better focus on the important parts in the centrifugal movement, centripetal movement and static movement.
[0107] The said spatio-temporal graph convolutional network is composed of 10 spatio-temporal graph convolutional units, and a global pooling layer is added. In this embodiment, the global pooling layer calculates the feature mean for each node, compresses the dimension of the output matrix, and obtains the feature sequence,
[0108] X=(x1,x2,...,x T′ )
[0109] where T′ = 1,2,...T in ;
[0110] Any spatio-temporal graph convolutional unit includes a spatial graph convolutional network and a temporal graph convolutional network.
[0111] In this embodiment, the feature dimension of the sequence is 256 and the length of the sequence is 38.
[0112] The formula for the said spatial graph convolution is:
[0113]
[0114] where f in represents the input feature sequence, and the dimension is Among them, C in represents the dimension of the node information data, T in represents the number of input frames, and N represents the number of nodes; f out represents the output feature sequence, and the dimension is Among them, C out represents the output feature dimension, and Tout represents the number of output frames; A j represents the normalized adjacency matrix formed according to the graph partitioning strategy; W j represents the weight matrix.
[0115] The time graph convolution uses a convolution kernel of size (kernel_size, 1) to learn the change features of the same nodes in the kernel_size key frames adjacent to the current frame. In this embodiment, kernel_size = 9 is taken.
[0116] Step 7: Use an encoder-decoder network, take the feature sequence X in Step 6 as the input, and output a complete sentence to achieve continuous sign language recognition.
[0117] In the said Step 7, the encoder-decoder network consists of two parts, an encoder and a decoder, including:
[0118] Step 7.1: Pass the feature sequence X obtained in Step 6 into the recurrent layer of the encoder to obtain the output o of the i-th recurrent unit i ,
[0119] o i = Encoder(x i , o i-1 )
[0120] Among them, x i is the i-th feature vector in X, i is a positive integer, and o0 is a zero vector;
[0121] Step 7.2: Through the encoder hidden state (Encoder Hidden States) h j-1 at the previous moment and the word embedding vector (Word Embedding) g j-1 output at the previous moment, the decoder generates the output y j at the next moment, updates the hidden state h j ,
[0122] y j , h j = Decoder(g j-1 , h j-1 )
[0123] g j-1 = wordEmbedding(yj-1 )
[0124] Among them, the initial hidden state h0 is the last encoding unit o in step 7.1 T The corresponding hidden state, set the identifier of the initial output y0, such as <sos>As an identifier for the start of a sequence.
[0125] In the present invention, the word embedding method in step 7.2 uses a fully connected-based linear mapping to convert the one-hot vector corresponding to a word into a representation g in a dense space. j .
[0126] In the present invention, an attention mechanism is added to the encoder-decoder network to provide additional state information for the decoder, thereby ensuring the consistency between the sign language video and the generated sentence.
[0127] In the present invention, a context vector is constructed to assist decoding. For each decoding, the context vector is obtained by a weighted sum of the encoded outputs.
[0128]
[0129] Wherein, represents the attention weight, describing the correlation between the encoder input x i and the generated word y j ; finally, the attention vector A j is calculated from the context vector c j and the hidden state h j .
[0130] A j = tanh(W c [c j ; h j )
[0131] After adding the attention mechanism to the decoder formula in step 7.2, it can be expressed as
[0132] y j , h j = Decoder'(g j-1 , h j-1 , A j-1 ).
[0133] Step 7.3: When the sequence end identifier, such as <eos>When it appears, the output is completed, and the obtained word sequence Y = {y1, y2,..., y p}, and the output word sequence is concatenated to obtain a sentence.
[0134] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0135] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0136] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0137] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0138] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.
[0139] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.< / eos> < / sos>
Claims
1. A real-time video stream sign language recognition method based on human key points, characterized in that: The method includes the following steps: Step 1: Read the input sign language video stream frame by frame; Step 2: Use a human pose estimation network to extract human key points in any frame image read; Step 3: Calculate the action difference degree between the current frame and the previous frame frame by frame and accumulate it; Step 4: When the accumulated difference degree within t time is higher than the threshold T1, it is determined that the sign language action starts; when the difference degree is lower than the threshold T2, it is determined that the action is stationary, and a convolutional neural network is used to judge whether the current frame action is an end action; T1 > T2 > 0; Step 5: Save the human key point data of all frames within the time period from the start to the end of the sign language action to obtain a human key point sequence; Step 6: Use a spatio-temporal graph convolutional network to extract features from the human key point sequence in Step 5 to obtain a feature sequence X; Step 7: Use an encoder-decoder network, take the feature sequence X in Step 6 as the input, and output a complete sentence to realize continuous sign language recognition.
2. The real-time video stream sign language recognition method based on human key points according to claim 1, characterized in that: The said Step 2 includes the following steps: Step 2.1: Input any frame image into the human pose estimation network and output key point information v, v = {(x1, y1, c1), (x2, y2, c2),..., (x M , y M , c M )} where M represents the number of output key points, and x i , y i , c i represent the x - coordinate, y - coordinate and predicted confidence of the i - th key point respectively, M ≥ 1, and i is the index of the key point; Step 2.2: Screen the key points for sign language recognition, denoted as v′ = {(x1, y1, c1), (x2, y2, c2),..., (x N , y N , c N )} where N is the number of key points, 1 ≤ N ≤ M.
3. A real-time video stream sign language recognition method based on human key points according to claim 1, characterized in that: The said Step 3 includes the following steps: Step 3.1: Read the key point coordinates of the current frame, P = {(x1, y1), (x2, y2),..., (x N , y N )} Let \(P\) cur be the set of key point coordinates of the current frame, and let \(P\) pre be the set of key point coordinates of the previous frame. If the current frame is the first frame of the video stream, then let \(P\) cur = \(P\), \(P\) pre = \(P\), otherwise, let \(P\) cur = \(P\); Step 3.2: Use the Euclidean distance between the corresponding key points of the current frame and the previous frame to calculate the spatial difference degree δ of the key points between adjacent frames, Among them, x_cur i and y_cur i respectively represent the x - coordinate and y - coordinate of the i - th key point in the set of key - point coordinates P cur of the current frame; x_pre i and y_pre i respectively represent the x - coordinate and y - coordinate of the i - th key point in the set of key - point coordinates P pre of the previous frame; Step 3.3: Repeat Step 3.2 and save the difference degree δ in a queue; the queue length is L, L = t × fps, where t represents the time threshold and fps represents the number of frames transmitted per second of the video stream.
4. A real-time video stream sign language recognition method based on human key points according to claim 3, characterized in that: In the said Step 4, for the accumulated difference degree S within t time, when S > T1, it means the sign language action starts, when S < T2, it means the action is stationary, input the current frame image into the convolutional neural network, if it is judged as an invalid sign language action, it means the sign language action ends.
5. A real-time video stream sign language recognition method based on human key points according to claim 3, characterized in that: The said Step 5 includes the following steps: Step 5.1: Save the human key point information sequence V of all frames of a sign language action, V = {v1, v2, ..., v L} where L represents the number of frames of this sign language action; Step 5.2: Based on the input dimension of the spatio-temporal graph convolutional model, adjust the length of the key point information sequence in Step 3.3 to T by using the method of frame extraction or empty frame filling in .
6. A real-time video stream sign language recognition method based on human key points according to claim 1, characterized in that: In the said Step 6, the spatio-temporal graph convolutional network performs spatial graph convolution and temporal graph convolution processing on the input data; use the key points selected in Step 2 as the nodes of the graph, connect them into edges according to the human body structure to form spatial edges, and connect the same nodes of adjacent frames into edges to form temporal edges.
7. A real-time video stream sign language recognition method based on human key points according to claim 6, characterized in that: The spatio-temporal graph convolutional network consists of 10 spatio-temporal graph convolutional units, and a global pooling layer is added to obtain a feature sequence, X = (x1, x2,..., x T′ ) where T′ = 1, 2,... T in ; Any spatio-temporal graph convolutional unit includes a spatial graph convolutional network and a temporal graph convolutional network.
8. A real-time video stream sign language recognition method based on human key points according to claim 6, characterized in that: The graph partitioning strategy of the said spatial graph convolution adopts spatial configuration partitioning. The first subset connects the neighbor nodes whose spatial positions are farther from the entire skeleton than the root node, representing the centrifugal movement in sign language, the second subset connects the neighbor nodes closer to the center, representing the centripetal movement in sign language, and the third subset is the root node itself, representing the stationary movement in sign language.
9. A real-time video stream sign language recognition method based on human key points according to claim 8, characterized in that: The formula of the said spatial graph convolution is, where f in represents the input feature sequence, with a dimension of where C in represents the dimension of the node information data, T in represents the number of input frames, and N represents the number of nodes; f out represents the output feature sequence, with a dimension of where C out represents the output feature dimension, T out represents the number of output frames; A j represents the normalized adjacency matrix formed according to the graph partitioning strategy; W j represents the weight matrix.
10. A real-time video stream sign language recognition method based on human key points according to claim 7, characterized in that: The said Step 7 includes the following steps: Step 7.1: Input the feature sequence X obtained in Step 6 into the recurrent layer of the encoder to obtain the output o of the i-th recurrent unit i , o i = Encoder(x i , o i-1 ) where x i is the i-th eigenvector in X, i is a positive integer, and o0 is a zero vector; Step 7.2: Using the hidden state h of the encoder at the previous moment j-1 and the word embedding g output at the previous moment j-1 , the decoder generates the output y at the next moment j , updates the hidden state h j , y j ,h j = Decoder(g j-1 ,h j-1 ) g j-1 = wordEmbedding(y j-1 ) Among them, the initial hidden state h0 is the output o of the last encoding unit in step 7.1 T For the corresponding hidden state, set the identifier of the initial output y0 as the identifier indicating the start of the sequence; Step 7.3: When the sequence end flag appears, complete the output to obtain the output word sequence Y = {y1, y2,..., y p}, concatenate the output word sequence to obtain a sentence.
Citation Information
Patent Citations
A sign language recognition method and system
CN111340005B
A Sign Language Recognition System and Method Based on Spatiotemporal Semantic Features
CN111797777B
Sign language recognition method and system based on double-flow space-time diagram convolutional neural network
CN111325099A
Sign language recognition method and system
CN111340006A