Gesture Recognition Method, Apparatus and Electronic Device
By using deep neural network based on similarity scores and confidence queue processing technology in the gesture recognition system, the inaccuracy problem of gesture recognition when the acquisition distance is far away or the scene is blurred is solved, and the accuracy and robustness of gesture recognition are improved.
Patent Information
- Application Number
- CN202510396669.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-01
AI Technical Summary
Existing gesture recognition technology cannot obtain clear hand key points when the acquisition distance is long or the scene is blurred, resulting in inaccurate gesture recognition. In addition, when optical sensing devices collect hand images, gesture changes lead to similar gesture matching errors, reducing the accuracy of gesture recognition.
By obtaining a sequence of video frames, traversing the video frames and inputting them into a deep neural network constructed based on similarity scores for gesture recognition, generating gesture detection results. A confidence queue is generated based on the original confidence level, a standard confidence and similarity are calculated, and a video frame is determined whether the video frame is a target video frame, and a gesture parameter is calculated based on the sequence value and confidence of the target video frame, and finally a gesture recognition result is generated.
This method can reduce the probability of gesture error detection during gesture transformation, improve the accuracy of gesture recognition, and enhance the robustness of gesture changes.
Smart Images

Figure CN119920013B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of recognition technologies, and in particular, to a gesture recognition method, apparatus, and electronic device. Background Art
[0002] Gesture recognition is a human-computer interaction technology that captures, analyzes, and understands the movements and postures of the human hand through technical means such as computer vision and sensors, thereby realizing the interaction between the user and the device. Gesture recognition can be applied to technical fields such as smart homes, entertainment games, medical assistance, and vehicle control.
[0003] By collecting hand images through optical sensing devices such as cameras, and extracting gesture-related features from the collected images based on image processing and deep learning, and matching the extracted features with predefined gesture models, the meaning represented by the current gesture can be determined.
[0004] However, when the above method extracts gesture-related features, using hand key point data as the basic data, it is easy to encounter the situation where clear hand key points cannot be obtained when the acquisition distance is far or the acquisition scene is blurred, which may lead to inaccurate gesture recognition. Moreover, since the gesture changes during the process of collecting hand images by the optical sensing device, similar gestures may appear during the gesture change process, resulting in incorrect matching between the extracted gesture features and the predefined gesture models, and thus reducing the accuracy rate of gesture recognition. Summary of the Invention
[0005] To solve the above problems, this application provides a gesture recognition method, apparatus, and electronic device, which can reduce the probability of gesture misdetection during the gesture transformation process, and thus improve the accuracy rate of gesture recognition.
[0006] In a first aspect, this application provides a gesture recognition method, including: obtaining a video frame sequence, where the video frame sequence includes video frames and sequence values corresponding to the video frames, and the sequence values are positive integers greater than or equal to 1;
[0007] Traversing the video frames in the video frame sequence, and inputting the video frames into a gesture recognition network to obtain gesture detection results corresponding to the video frames. The gesture recognition network is a deep neural network constructed based on similarity scores, and the gesture detection results include gesture coordinates, gesture categories, original confidence levels, and similarities;
[0008] Generating a confidence queue based on the original confidence levels, and generating a standard confidence level of the gesture detection results according to the confidence queue;
[0009] Determine the video frames with the standard confidence greater than or equal to the preset confidence threshold and the similarity greater than or equal to the preset similarity threshold as target video frames; and determine the video frames with the standard confidence less than the preset confidence threshold or the similarity less than the preset similarity threshold as non-target video frames;
[0010] When the video frame is the target video frame, calculate the first gesture parameter according to the sequence value corresponding to the target video frame; and calculate the second gesture parameter according to the standard confidence and the similarity of the target video frame;
[0011] Generate a gesture recognition result based on the standard confidence, the first gesture parameter, and the second gesture parameter, where the gesture recognition result is used to characterize the gesture category and gesture position of the target video frame.
[0012] In an alternative embodiment, before inputting the video frame into a gesture recognition network to obtain a gesture detection result corresponding to the video frame, it further includes: obtaining a case data set, where the case data set includes at least one piece of training gesture data, and the training gesture data is provided with a gesture annotation area and a gesture annotation category;
[0013] Generate a standard gesture label according to the training gesture data, and calculate a similarity score between the training gesture data and the standard gesture label, where the similarity score is used to characterize the difference between the training gesture data and the standard gesture label, and the similarity score is in an associated relationship with the training gesture data;
[0014] Input the training gesture data and the similarity score into the gesture recognition network, and control the gesture recognition network to output regression parameters, where the regression parameters include coordinate information, dimension information, confidence information, classification result information, and similarity information;
[0015] Train the gesture recognition network based on the regression parameters, the gesture annotation category, and a loss function to optimize the model parameters of the gesture recognition network.
[0016] In an alternative embodiment, generating a standard gesture label according to the training gesture data includes: grouping the training gesture data according to the gesture annotation category to obtain at least one group of gesture classification groups;
[0017] Obtain the target gesture area of the training gesture data according to the gesture annotation area, and input the target gesture area into a feature extraction model to obtain gesture features, where the gesture features are in an associated relationship with the training gesture data;
[0018] Generate a standard gesture label according to the gesture feature, where the standard gesture label is used to represent a reference gesture of the gesture standard category.
[0019] In an optional implementation manner, the training gesture data includes continuous frame data and discontinuous frame data; when the training gesture data is the continuous frame data, mark the gesture feature of the training gesture data as a continuous frame gesture feature;
[0020] When the training gesture data is the discontinuous frame data, multiply a preset weight by the gesture feature of the training gesture data to obtain a discontinuous frame gesture feature.
[0021] In an optional implementation manner, the generating the standard gesture label according to the gesture feature includes: obtaining a reference gesture feature of the training gesture data in each gesture category group, where the reference gesture feature includes the continuous frame gesture feature and / or the discontinuous frame gesture feature;
[0022] Calculate an average value of the reference gesture feature to obtain the standard gesture label.
[0023] In an optional implementation manner, calculating a similarity score between the training gesture data and the standard gesture label includes: obtaining reference gesture data, where the reference gesture data is the standard gesture label of the gesture classification group where the training gesture data is located;
[0024] Calculate a cosine distance between the training gesture data and the reference gesture data to obtain the similarity score.
[0025] In an optional implementation manner, the loss function includes a class loss, a target loss, a position loss, and a similarity loss. The similarity loss is a quotient of a class edit distance and the number of classes. The class edit distance is used to represent a difference between the classification result information and the gesture annotation category, and the number of classes is used to represent the number of gesture classification groups;
[0026] Training the gesture recognition network based on the regression parameter, the gesture annotation category, and the loss function to optimize the model parameters of the gesture recognition network includes:
[0027] Optimizing the classification parameters of the gesture recognition network based on the class loss, where the classification parameters are used to calculate the classification result information;
[0028] Optimizing the confidence parameters of the gesture recognition network based on the target loss, where the confidence parameters are used to calculate the confidence information;
[0029] Optimize the position parameters of the gesture recognition network based on the position loss, where the position parameters are used to calculate the coordinate information, width information, and height information;
[0030] Optimize the similarity parameters of the gesture recognition network based on the similarity loss, where the similarity parameters are used to calculate the similarity information.
[0031] In an alternative embodiment, generate a confidence queue based on the original confidence, and generate the standard confidence of the gesture detection result according to the confidence queue, including: when the sequence value of the video frame is equal to 1, the standard confidence of the gesture detection result is the original confidence;
[0032] When the sequence value of the video frame is greater than 1, obtain a reference confidence, where the reference confidence is the original confidence of the reference frame, and the sequence value corresponding to the reference frame is less than the sequence value corresponding to the video frame; and generate the confidence queue according to the reference confidence and the original confidence, and calculate the average value of the reference confidence and the original confidence to obtain the standard confidence of the gesture detection result.
[0033] In an alternative embodiment, the gesture recognition method further includes: when the video frame is the non-target video frame, input the backward video frame into the gesture recognition network to perform gesture recognition on the backward video frame, where the backward video frame is the next frame of the non-target video frame.
[0034] In an alternative embodiment, calculate a first gesture parameter according to the sequence value corresponding to the target video frame, including: generate a state index according to the sequence value corresponding to the target video frame, where the state index is used to represent the number of target video frames between the video frame with a sequence value of 1 and the current video frame in the video frame sequence;
[0035] Calculate the first gesture parameter using a weighted average formula, where the weighted average formula is:
[0036] ,
[0037] where, represents the state index, represents the first gesture parameter, represents an intensity factor, and calculate the intensity factor using an intensity formula, where the intensity formula is:
[0038] ,
[0039] where, represents a preset number of frames, represents a preset step size.
[0040] In an alternative embodiment, calculating a second gesture parameter according to the standard confidence and the similarity includes: when the sequence value of the target video frame is equal to 1, setting the second gesture parameter to a first preset value;
[0041] When the sequence value of the target video frame is greater than 1, obtaining a forward confidence and a forward similarity, where the forward confidence is the standard confidence of a forward video frame, the forward similarity is the similarity of the forward video frame, and the forward video frame is the previous frame of the target video frame;
[0042] When the forward confidence is less than the standard confidence of the target video frame and the forward similarity is less than the similarity of the target video frame, setting the second gesture parameter to a first preset value;
[0043] When the forward confidence is greater than or equal to the standard confidence of the target video frame and the forward similarity is greater than or equal to the similarity of the target video frame, setting the second gesture parameter to a second preset value.
[0044] In an alternative embodiment, generating a gesture recognition result based on the standard confidence, the first gesture parameter, and the second gesture parameter includes: multiplying the standard confidence, the first gesture parameter, and the second gesture parameter to obtain a first detection result;
[0045] Performing non-maximum suppression algorithm processing on the first detection result to obtain a second detection result;
[0046] Decoding the second detection result to generate the gesture recognition result.
[0047] In a second aspect, the present application provides a gesture recognition device, including: a video frame acquisition module configured to: obtain a video frame sequence, where the video frame sequence includes video frames and sequence values corresponding to the video frames, and the sequence values are positive integers greater than or equal to 1;
[0048] A gesture detection module configured to: traverse the video frames in the video frame sequence, input the video frames into a gesture recognition network to obtain gesture detection results corresponding to the video frames, where the gesture recognition network is a deep neural network constructed based on similarity scores, and the gesture detection results include gesture coordinates, gesture categories, original confidence, and similarity;
[0049] A confidence processing module configured to: generate a confidence queue based on the original confidence and generate a standard confidence of the gesture detection result according to the confidence queue;
[0050] A video frame determination module, configured to: determine a video frame whose standard confidence is greater than or equal to a preset confidence threshold and whose similarity is greater than or equal to a preset similarity threshold as a target video frame; and determine a video frame whose standard confidence is less than the preset confidence threshold or whose similarity is less than the preset similarity threshold as a non-target video frame;
[0051] A result output module, configured to: when the video frame is the target video frame, calculate a first gesture parameter according to a sequence value corresponding to the target video frame; and calculate a second gesture parameter according to the standard confidence and the similarity of the target video frame; generate a gesture recognition result based on the standard confidence, the first gesture parameter, and the second gesture parameter, where the gesture recognition result is used to represent a gesture category and a gesture position of the target video frame;
[0052] When the video frame is the non-target video frame, input a backward video frame into the gesture recognition network to perform gesture recognition on the backward video frame, where the backward video frame is the next frame of the non-target video frame.
[0053] In a third aspect, the present application provides an electronic device, including: a memory, and one or more processors; the memory is configured to: store one or more programs; where, when the one or more programs are executed by the one or more processors, the one or more processors implement the gesture recognition method of the first aspect.
[0054] As can be seen from the above technical solutions, the present application provides a gesture recognition method, apparatus, and electronic device. The gesture recognition method includes: obtaining a video frame sequence, traversing the video frames in the video frame sequence, and inputting the video frames into a gesture recognition network to obtain gesture detection results, including gesture coordinates, gesture categories, original confidence, and similarity; generating a confidence queue based on the original confidence, and generating a standard confidence of the gesture detection result according to the confidence queue; determining whether a video frame is a target video frame according to the standard confidence and the similarity. When the video frame is the target video frame, calculate a first gesture parameter according to a sequence value corresponding to the target video frame; and calculate a second gesture parameter according to the standard confidence and the similarity of the target video frame; generate a gesture recognition result based on the standard confidence, the first gesture parameter, and the second gesture parameter. The above method can improve the accuracy of gesture recognition and reduce the situation of false detection of gestures caused by gesture transformation. Description of the Drawings
[0055] In order to more clearly illustrate the technical solutions of the present application, the drawings required for the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0056] Figure 1 Schematic flowchart of a gesture recognition method provided in this embodiment;
[0057] Figure 2 Schematic flowchart of the training process of a gesture recognition network provided in this embodiment;
[0058] Figure 3 Schematic diagram of the regression parameters of a gesture recognition network provided in this embodiment;
[0059] Figure 4 Schematic diagram of a method for generating a standard confidence level provided in this embodiment;
[0060] Figure 5 Schematic structural diagram of a gesture recognition device provided in this embodiment;
[0061] Figure 6 Schematic diagram of an electronic device provided in this embodiment. Detailed implementation manners
[0062] In order to enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0063] It should be understood that the "multiple" mentioned herein refers to two or more. In the description of the embodiments of this application, unless otherwise specified, " / " means "or". For example, A / B may represent A or B; the "and / or" herein is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in order to clearly describe the technical solutions of the embodiments of this application, in the embodiments of this application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that the terms "first" and "second" do not limit the quantity and execution order, and the terms "first" and "second" do not necessarily limit being different.
[0064] In addition, the terms "comprising", "having", and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0065] Gesture recognition is a human-computer interaction technology that captures, analyzes, and understands the movements and postures of the human hand through technical means such as computer vision and sensors, thereby enabling interaction between the user and the device. Gesture recognition can be applied to technical fields such as smart home, entertainment games, medical assistance, and vehicle control.
[0066] In some embodiments, the user wears a wearable device such as a glove or a wristband that carries a sensor and makes a gesture. The wearable device captures the change in the electrical signal, converts the electrical signal into a digital signal, and then performs analysis and recognition on the gesture. This method can overcome the influence of the external environment and accurately capture the hand movement and posture, but it requires the support of the wearable device.
[0067] In other embodiments, an optical sensing device such as a camera is used to collect the hand image of the user, and image processing and deep learning are used to distinguish the gesture. This method does not require the support of external devices, can improve the user experience, and is suitable for gesture recognition of users in scenarios with high requirements for real-time performance and freedom.
[0068] However, in the process of using an optical sensing device such as a camera to collect the hand image of the user, it is easy to encounter the situation where clear hand key points cannot be obtained when the collection distance is far or the collection scene is blurred, which may lead to inaccurate gesture recognition. Moreover, since the gesture changes during the process of the optical sensing device collecting the hand image, similar gestures may appear during the gesture change process, resulting in incorrect matching between the extracted gesture features and the predefined gesture model, and thus reducing the accuracy rate of gesture recognition.
[0069] In order to reduce the probability of gesture misdetection during the gesture transformation process and improve the accuracy rate of gesture recognition, a first aspect of the present application provides a gesture recognition method. Figure 1 It is a schematic flowchart of a gesture recognition method provided for this embodiment. Next, according to Figure 1 The gesture recognition method of the embodiments of the present application will be described in detail.
[0070] Step S100: Obtain a video frame sequence.
[0071] In the embodiments of the present application, a video containing gestures can be obtained through a camera device. It should be understood that a video is essentially a visual effect formed by rapidly playing a series of consecutive images, and each individual image in these consecutive images is called a video frame. The video is processed by video processing software to generate a video frame sequence. Among them, the video frame sequence includes video frames and sequence values corresponding to the video frames, and the sequence values are positive integers greater than or equal to 1.
[0072] The sequence value is used to represent the order of the video frame in the video frame sequence. For example, the sequence value of the first video frame is 1, the sequence value of the second video frame is 2, and the sequence value of the nth video frame is n.
[0073] In some embodiments, a sliding window can be used to process the video containing gestures frame by frame to obtain video frames. The sliding window maintains a window of a specific size, slides in the video at a set step size, and performs image processing on the images within the window, such as object detection, feature extraction, and image classification, etc., to obtain video frames.
[0074] Step S200: Traverse the video frames in the video frame sequence, and input the video frames into the gesture recognition network to obtain the gesture detection results corresponding to the video frames.
[0075] In the embodiments of the present application, the gesture recognition network is a deep neural network constructed based on similarity scores. Specifically, it is a neural network model used for detecting specific targets or objects in the fields of computer vision and artificial intelligence. After inputting the video frame into the gesture recognition network, the gesture recognition network outputs the gesture detection results corresponding to the video frame after feature extraction, target classification, and post-processing. Among them, the gesture detection results include gesture coordinates, gesture categories, original confidence levels, and similarities.
[0076] During the process of the gesture recognition network processing the video frame, the gesture recognition network will first use a feature extractor constructed by convolutional layers, pooling layers, etc. to process the input image or video frame to convert the original image data into feature maps, which contain various semantic information and spatial information in the image. After obtaining the feature maps, the detection network will predict the position and category of the target through a series of prediction heads or branches. For the prediction of the target position, the coordinates of the target's bounding box will be output, such as the coordinates of the upper left corner and the lower right corner, or the center coordinates and width and height information. For the prediction of the target category, the probability or score of each possible category will be output to represent the possibility that a target of a certain category exists at the current position. Finally, the results obtained through network prediction usually need to be post-processed to obtain the final detection results. Common post-processing operations include Non-Maximum Suppression (NMS), which is used to remove redundant detection boxes with high overlap degrees and retain the detection box with the highest score as the final result.
[0077] In some embodiments, in order to improve the accuracy of gesture detection results, it is necessary to build a trainable gesture recognition network based on the similarity score. Figure 2 As shown, the step of constructing a gesture recognition network includes steps S210 to S240.
[0078] Step S210: Obtain case data set.
[0079] In some embodiments, the electronic device obtains the case data set by means of a web crawler, querying a professional data set, etc. The case data set includes at least one training gesture data, and the training gesture data is used to characterize different gesture videos. For example, the training gesture data may be different gesture types, different shooting angles, different lighting conditions, and different shooting backgrounds.
[0080] The electronic device annotates the collected training gesture data to mark the category to which each gesture belongs and the area where the gesture is located. The training gesture data is preprocessed, such as normalization, cropping, scaling, data enhancement, etc. Among them, normalization can scale the pixel values of the training gesture data to a preset range, cropping and scaling can adjust the image to a uniform size, and data enhancement can expand the data set by rotating, flipping, and adding noise, thereby increasing the diversity of the training gesture data.
[0081] Each training gesture data is provided with a gesture annotation area and a gesture annotation category. The gesture annotation area is used to characterize the coordinates of the gesture in the training gesture data, and the gesture annotation category is used to characterize the type of gesture in the training gesture data, such as: indicating gesture, thumbs-up gesture, and number gesture.
[0082] In this way, the gesture recognition network can train the gesture annotation area and gesture annotation category of the gesture data in the case data set, identify the shape, texture and other characteristics of different gestures, and improve the accuracy of gesture recognition.
[0083] Step S220: Generate a standard gesture label according to the training gesture data, and calculate a similarity score between the training gesture data and the standard gesture label.
[0084] In some embodiments, the electronic device groups the training gesture data in the case data set according to the gesture annotation category set in the training gesture data to obtain at least one group of gesture classification groups, such as an indication gesture group, a thumbs-up gesture group, and a quantity gesture group.
[0085] Next, the electronic device obtains the target gesture area for training gesture data according to the gesture annotation area, and inputs the target gesture area into the feature extraction model to obtain gesture features. Among them, the gesture features and the training gesture data are in an associated relationship, that is, each piece of training gesture data has its corresponding gesture features. Finally, a standard gesture label is generated according to the gesture features, and the standard gesture label is used to represent the reference gesture of the gesture annotation category.
[0086] It should be noted that the feature extraction model is a convolutional neural network. The convolutional layer of the feature extraction model performs convolution operations by sliding the convolution kernel on the image, and automatically extracts local features of the gesture image, such as low-level features like edges, corners, and textures. As the number of network layers increases, the higher-level convolutional layers can combine these low-level features into more abstract and representative high-level features, such as the combination pattern of fingers and the shape structure of the overall gesture. The pooling layer can compress and downsample the features extracted by the convolutional layer, reducing the data volume while retaining the main features, and improving the computational efficiency and generalization ability of the model.
[0087] In some embodiments, the training gesture data includes continuous frame data and discontinuous frame data. Among them, the continuous frame data is used to represent that the gesture action is stored in a continuous video frame, and the discontinuous frame data is used to represent that the gesture action is stored in multiple video frames, and the multiple video frames form the discontinuous frame data.
[0088] When the training gesture data is continuous frame data, the gesture features of the training gesture data are marked as continuous frame gesture features; when the training gesture data is discontinuous frame data, the preset weight is multiplied by the gesture features of the training gesture data to obtain discontinuous frame gesture features.
[0089] In some embodiments, the method for generating the standard gesture label according to the gesture features is: obtaining the reference gesture features of the training gesture data in each gesture category group, where the reference gesture features include continuous frame gesture features and / or discontinuous frame gesture features; and calculating the average value of the reference gesture features to obtain the standard gesture label.
[0090] Exemplarily, the first gesture group includes first continuous frame data and first discontinuous frame data. Then, the first continuous frame gesture features and the first discontinuous frame gesture features are obtained, where the first discontinuous frame gesture features are the features processed by the Exponential Moving Average (EMA). The first continuous frame gesture features and the first discontinuous frame gesture features are averaged to generate the standard gesture label of the first gesture group. In the above process, the exponential moving average is a weighted moving average, which processes the data according to the preset weight to smooth the data and reduce the short-term noise and fluctuations in the data.
[0091] In this way, by calculating the average value of the reference gesture features, a standard gesture label is obtained, which can make the feature result not affected by a single outlier and improve the accuracy of the standard gesture label.
[0092] Furthermore, the electronic device calculates the similarity score between the training gesture data and the standard gesture label. Among them, the similarity score is used to characterize the difference between the training gesture data and the standard gesture label, and the similarity score and the training gesture data are in an associated relationship.
[0093] In some embodiments, the cosine similarity is used to calculate the similarity score, including: obtaining the reference gesture data, where the reference gesture data is the standard gesture label of the gesture classification group where the training gesture data is located; and calculating the cosine distance between the training gesture data and the reference gesture data, and then obtaining the similarity score.
[0094] It should be noted that the method for calculating the similarity score also includes: using the Euclidean distance to calculate the distance between the training gesture data and the reference gesture data, where the smaller the Euclidean distance, the more similar the training gesture data and the reference gesture data are. This application does not specifically limit the method for calculating the similarity score.
[0095] By calculating the similarity between the training gesture data and the reference gesture data, it can provide a data basis for gesture recognition for the gesture recognition network and improve the reliability of the gesture recognition network for gesture recognition.
[0096] Step S230: Input the training gesture data and the similarity score into the gesture recognition network, and control the gesture recognition network to output regression parameters.
[0097] The regression parameters include coordinate information, size information, confidence information, classification result information, and similarity information.
[0098] In some embodiments, in order to ensure the real-time performance of the gesture recognition function, the gesture recognition network adopts YOLO (You Only Look Once) as the basic network architecture. Among them, YOLO is a real-time object detection algorithm widely used in the field of computer vision. It adopts a single-stage test framework, unifies the object detection task into a regression problem, and directly predicts the position and category of the object in one forward propagation process. The detection speed is relatively fast and can meet the application scenarios with high real-time requirements.
[0099] The structure of YOLO includes a backbone network, a neck network, and a head network. Among them, the backbone network is used to extract the features of the image. The backbone network usually contains multiple convolutional layers and pooling layers, and extracts different levels of features of the image through continuous convolutional operations, such as from low-level features such as edges and textures to more abstract semantic features.
[0100] The neck network is between the backbone network and the detection head and is used for feature fusion and processing. For example, the spatial pyramid pooling layer can fuse features of different scales, and the feature pyramid network can transfer and fuse information between feature maps of different levels, enhancing the model's detection ability for targets of different scales.
[0101] The head network is used to output detection results, including the category and location information of the target. It usually includes a convolutional layer and a regression layer. By performing convolutional operations on the feature map, it predicts the coordinates of the target's bounding box, class probability, etc.
[0102] To improve the scalability and flexibility of the gesture recognition network, the embodiment of the present application uses YOLOV5n as the basic network architecture, so that the gesture recognition network can adjust the network depth and width according to requirements and support data augmentation techniques and training strategies, thereby improving the performance and robustness of the gesture recognition network.
[0103] Furthermore, the electronic device inputs the training gesture data and similarity scores into the gesture recognition network and controls the gesture recognition network to output regression parameters, which include coordinate information, size information, confidence information, classification result information, and similarity information.
[0104] In some embodiments, the gesture recognition network applies a convolutional layer to extract features from the input gesture image and outputs three branches. The parameters of each branch include coordinate information, size information, confidence information conf, and classification result information class. As Figure 3 shown, the coordinate information includes the abscissa x and the ordinate y, and the size information includes the width w and the height h. Then, the parameters output by each branch are adjusted to add similarity information m, thereby generating regression parameters, which include coordinate information, size information, confidence information conf, classification result information class, and similarity information m. Similarly, the coordinate information includes the abscissa x and the ordinate y, and the size information includes the width w and the height h. In this way, by adding similarity information m, the accuracy of gesture recognition can be constrained, thereby improving the accuracy of the output results of the gesture recognition network.
[0105] Step S240: Train the gesture recognition network based on the regression parameters, the gesture standard category, and the loss function to optimize the model parameters of the gesture recognition network.
[0106] The loss function is a function used in machine learning and deep learning to measure the difference between the predicted results of a model and the true results. The loss function can convert the difference between the predicted results (i.e., regression parameters) output by the gesture recognition network and the true results (i.e., gesture standard categories) into a numerical value, which can be used as a feedback signal for training the gesture recognition network. The gesture recognition network continuously optimizes the model parameters through an optimization algorithm to reduce the value of the loss function.
[0107] In some embodiments, the loss function includes class loss, object loss, location loss, and similarity loss. The gesture recognition network is trained based on the regression parameters, gesture annotation categories, and the loss function to optimize the model parameters of the gesture recognition network, including: optimizing the classification parameters of the gesture recognition network based on the class loss, where the classification parameters are used to calculate classification result information; optimizing the confidence parameters of the gesture recognition network based on the object loss, where the confidence parameters are used to calculate confidence information; optimizing the location parameters of the gesture recognition network based on the location loss, where the location parameters are used to calculate coordinate information, width information, and height information; optimizing the similarity parameters of the gesture recognition network based on the similarity loss, where the similarity parameters are used to calculate similarity information.
[0108] Exemplarily, the loss function can be expressed by the following formula:
[0109] ; (1)
[0110] Where, represents the loss function, represents the class loss, represents the weight of the class loss, represents the object loss, represents the weight of the object loss, represents the location loss, represents the weight of the location loss, represents the similarity loss, represents the weight of the similarity loss.
[0111] In some embodiments, the gesture recognition network can calculate the intersection over union (IoU) of the predicted gesture region and the true gesture region based on the object loss, and then evaluate the accuracy and effectiveness of the gesture recognition network to optimize the confidence parameters. Among them, the predicted gesture region is the predicted region output by the gesture recognition network, and the true gesture region is a pre-set prior box (anchor box), and the prior box is used to represent the width and height of the common gesture region. Exemplarily, dividing the area of the overlapping part of the predicted gesture region and the true gesture region by the combined area of the predicted gesture region and the true gesture region can calculate the intersection over union of the predicted gesture region and the true gesture region.
[0112] In some embodiments, the similarity loss is the quotient of the class edit distance and the number of classes, where the class edit distance is used to characterize the difference between the classification result information and the gesture annotation class, and the number of classes is used to characterize the number of gesture classification groups. The similarity loss can be calculated using the following formula:
[0113] ; (2)
[0114] Wherein, represents the similarity loss, represents the gesture standard class, represents the classification result information, represents the number of classes, represents the class edit distance.
[0115] Generating a loss function based on the class edit distance and training a gesture recognition network based on the loss function can enable the gesture recognition network to perform real-time processing on the input video frame sequence. If there is a gesture target in the video frame, the gesture recognition network outputs the coordinate information, width information, size information, confidence information, classification result information, and similarity information of the gesture.
[0116] Exemplarily, in the process of calculating the class edit distance between the gesture standard class and the classification result information, the gesture standard class is initialized as the first string a, and the classification result information is initialized as the second string b. Use to represent the class edit distance between the first m characters of the first string a and the first n characters of the second string b. Set , which is used to characterize that the class edit distance from an empty string to the second string b of length n is n insertion operations, and set , which is used to characterize that the class edit distance from the first string a of length m to an empty string is m deletion operations.
[0117] For m > 0 and n > 0, the class edit distance is calculated by determining whether the characters of the first string a and the second string b are equal, including: if (that is, the (m - 1)-th character in the first string a is the same as the (n - 1)-th character in the second string b), then ; if is not equal to (that is, the -th character in the first string a is different from the -th character in the second string b), then obtain the minimum result value among the deletion operation, insertion operation, and replacement operation as the class edit distance. Among them, the result value of the deletion operation is , and the result value of the deletion operation is used to characterize deleting the m-th character of the first string a; the result value of the insertion operation is , the result value of the insertion operation is used to represent inserting the nth character of the second string b at the mth position of the first string a; the result value of the replacement operation is , the result value of the replacement operation is used to represent replacing the mth character of the first string a with the nth character of the second string b. By traversing the characters in the first string a and the characters in the second string b, the category edit distance between the gesture standard category and the classification result information is calculated.
[0118] Step S300: Generate a confidence queue based on the original confidence, and generate the standard confidence of the gesture detection result according to the confidence queue.
[0119] In some embodiments, in order to reduce the sudden jump of the gesture confidence caused by the change of the shooting environment, it is necessary to perform smoothing processing on the gesture detection result to update the original confidence of the gesture detection result to the standard confidence. As Figure 4 shown, when the sequence value of the video frame is equal to 1, the standard confidence of the gesture detection result is the original confidence; when the sequence value of the video frame is greater than 1, obtain the reference confidence, and generate a confidence queue according to the reference confidence and the original confidence. Then calculate the average value of the reference confidence and the original confidence to obtain the standard confidence of the gesture detection result. Among them, the reference confidence is the original confidence of the reference frame, and the sequence value corresponding to the reference frame is less than the sequence value corresponding to the video frame.
[0120] Exemplarily, when generating the standard confidence of the video frame with a sequence value of 5, sequentially obtain the original confidences of the video frames with sequence values from 1 to 5, generate a confidence queue, calculate the average value of all the data in the confidence queue, and update the above average value to the standard confidence of the video frame with a sequence value of 5.
[0121] Exemplarily, when generating the standard confidence of the video frame with a sequence value of 6, sequentially obtain the original confidences of the video frames with sequence values from 2 to 6, generate a confidence queue, calculate the average value of all the data in the confidence queue, and update the above average value to the standard confidence of the video frame with a sequence value of 6.
[0122] It should be noted that the length of the confidence queue is a parameter that can be adjusted, and this parameter is determined according to the actual application scenario. In addition, the standard confidence of the gesture detection result can be calculated by means of weighted average or median filtering. The present application does not limit the calculation method of the standard confidence, as long as the confidence can be smoothed.
[0123] Step S400: Determine the video frames with a standard confidence greater than or equal to the preset confidence threshold and a similarity greater than or equal to the preset similarity threshold as target video frames; and determine the video frames with a standard confidence less than the preset confidence threshold or a similarity less than the preset similarity threshold as non-target video frames.
[0124] In some embodiments, video frames are labeled as target video frames or non-target video frames according to standard confidence and similarity. Among them, target video frames are used to represent that there are gestures in the video frames, and gesture recognition needs to be performed on the gestures. Non-target video frames are used to represent that there are no gestures in the video frames, and no other processing needs to be performed on the video frames.
[0125] Step S5011: When the video frame is a target video frame, calculate the first gesture parameter according to the sequence value corresponding to the target video frame.
[0126] The video frame sequence is used to describe dynamic gestures. In some embodiments, dynamic gestures include three stages, namely the preliminary preparation stage, the mid-term completion stage, and the late retraction stage. Among them, the preliminary preparation stage is used to represent the starting position, speed, amplitude, etc. of the gesture. The mid-term completion stage is used to represent the execution of the gesture action. The mid-term completion stage includes a complete and smooth gesture action. The late retraction stage is used to represent the process of the gesture recovering after the gesture expression is completed. It can be seen that the gesture information in the video frames corresponding to the preliminary preparation stage is relatively fuzzy. When the video frame corresponding to the preliminary preparation stage is a target video frame, it is necessary to calculate the first gesture parameter according to the sequence value corresponding to the target video frame, so as to improve the credibility of gesture recognition.
[0127] In the process of calculating the first gesture parameter, first generate a state index according to the sequence value corresponding to the target video frame, where the state index is used to represent the number of target video frames between the video frame with a sequence value of 1 and the current video frame in the video frame sequence; then calculate the first gesture parameter using the weighted average formula, and the weighted average formula is as follows:
[0128] ; (3)
[0129] Among them, represents the first gesture parameter, represents the state index, represents the intensity factor. The intensity factor is used to represent the intensity of the standard confidence corresponding to the target video frame. The intensity factor can be calculated using the intensity formula, and the intensity formula is as follows:
[0130] ; (4)
[0131] Among them, represents the preset number of frames. The preset number of frames is determined according to prior information and can be the number of frames in which the gesture appears in the preset gesture sequence. represents the preset step size, which is used to represent the number of frames between the video frames selected for two adjacent operations when operating on the video frames.
[0132] In this way, by calculating the first gesture parameter through a weighted average strategy, the ambiguity of the gesture in the preliminary preparation stage can be reduced, and the accuracy of gesture recognition can be improved.
[0133] Step S5012: Calculate the second gesture parameter according to the standard confidence level and similarity of the target video frame.
[0134] In some embodiments, when the sequence value of the target video frame is equal to 1, the second gesture parameter is set to a first preset value; when the sequence value of the target video frame is greater than 1, the second gesture parameter is determined according to the forward confidence level and forward similarity, where the forward confidence level is the standard confidence level of the forward video frame, the forward similarity is the similarity of the forward video frame, and the forward video frame is the previous frame of the target video frame.
[0135] When the sequence value of the target video frame is greater than 1, the process of calculating the second gesture parameter includes: when the forward confidence level is less than the standard confidence level of the target video frame and the forward similarity is less than the similarity of the target video frame, setting the second gesture parameter to the first preset value; when the forward confidence level is greater than or equal to the standard confidence level of the target video frame and the forward similarity is greater than or equal to the similarity of the target video frame, setting the second gesture parameter to the second preset value.
[0136] In some embodiments, the sequence value of the target video frame is , is an integer greater than 1, the standard confidence level of the target video frame is , the similarity of the target video frame is . The sequence value of the forward video frame is , the forward confidence level is , the forward similarity is , then the second gesture parameter is calculated using the following formula:
[0137] ; (5)
[0138] where, represents the second gesture parameter, represents the first preset value, represents the second preset value.
[0139] Exemplarily, the first preset value is 1 and the second preset value is 0.
[0140] In some embodiments, the second gesture parameter can be calculated by setting the state of a single activation model. The single activation model refers to a model that, during operation, performs one activation operation based on the input data to generate an output result, and can complete the mapping from input to output through one calculation step, thereby quickly processing and responding to the input information.
[0141] Exemplarily, the standard confidence, similarity, forward confidence, and forward similarity of the target video frame are input into a single activation model. After calculation, the single activation model outputs a model state, which can be an active state and an end state. Among them, when the forward confidence is less than the standard confidence of the target video frame and the forward similarity is less than the similarity of the target video frame, the model state output by the single activation model is the active state. When the forward confidence is greater than or equal to the standard confidence of the target video frame and the forward similarity is greater than or equal to the similarity of the target video frame, the model state output by the single activation model is the end state.
[0142] Then, the second gesture parameter is calculated according to the model state output by the single activation model. When the model state is the active state, the second gesture parameter is set to a first preset value; when the model state is the end state, the second gesture parameter is set to a second preset value. For example, when the model state is the active state, the second gesture parameter is set to 1, and when the model state is the end state, the second gesture parameter is set to 0.
[0143] In this way, the second gesture parameter is used to further characterize the similarity and confidence of the target video frame, providing a data basis for the generation of subsequent gesture recognition results.
[0144] Step S5013: Generate a gesture recognition result based on the standard confidence, the first gesture parameter, and the second gesture parameter.
[0145] In some embodiments, the gesture recognition result is used to characterize the gesture category and gesture position of the target video frame. In the process of generating the gesture result, first multiply the standard confidence, the first gesture parameter, and the second gesture parameter to obtain a first detection result, then perform non-maximum suppression algorithm processing on the first detection result to obtain a second detection result, and finally decode the second detection result to generate a gesture recognition result. Among them, the non-maximum suppression algorithm is used to screen out the final target detection result from a large number of candidate detection frames corresponding to the first detection result as the second detection result, thereby improving the accuracy and efficiency of gesture recognition.
[0146] Step S5021: When the video frame is a non-target video frame, input the backward video frame into the gesture recognition network to perform gesture recognition on the backward video frame.
[0147] In some embodiments, the non-target video frame is used to indicate that the video frame does not include a gesture and no other processing needs to be performed on this video frame. Input the next frame of the non-target video frame into the gesture recognition network to continue performing gesture recognition on the subsequent video frames in the video frame sequence until all video frames in the video frame sequence are detected to obtain a gesture recognition result.
[0148] As can be seen from the above technical solution, the present application provides a gesture recognition method, including: obtaining a video frame sequence, traversing the video frames in the video frame sequence, inputting the video frames into a gesture recognition network to obtain gesture detection results, including gesture coordinates, gesture categories, original confidence levels, and similarity degrees; generating a confidence level queue based on the original confidence levels, and generating a standard confidence level of the gesture detection results according to the confidence level queue; determining whether a video frame is a target video frame according to the standard confidence level and the similarity degree, and calculating a first gesture parameter according to the sequence value corresponding to the target video frame when the video frame is a target video frame; and calculating a second gesture parameter according to the standard confidence level and the similarity degree of the target video frame; generating a gesture recognition result based on the standard confidence level, the first gesture parameter, and the second gesture parameter. The above method can improve the accuracy of gesture recognition and reduce the situation of false detection of gestures caused by gesture transformation.
[0149] Based on the gesture recognition method provided in the above embodiments, some embodiments of the present application further provide a gesture recognition device. Refer to Figure 5 , the gesture recognition device includes: a video frame acquisition module 501, a gesture detection module 502, a confidence level processing module 503, a video frame determination module 504, and a result output module 505.
[0150] The video frame acquisition module 501 is configured to: obtain a video frame sequence, where the video frame sequence includes video frames and sequence values corresponding to the video frames, and the sequence values are positive integers greater than or equal to 1.
[0151] The gesture detection module 502 is configured to: traverse the video frames in the video frame sequence, input the video frames into a gesture recognition network to obtain gesture detection results corresponding to the video frames, the gesture recognition network is a deep neural network constructed based on similarity scores, and the gesture detection results include gesture coordinates, gesture categories, original confidence levels, and similarity degrees.
[0152] The confidence level processing module 503 is configured to: generate a confidence level queue based on the original confidence levels, and generate a standard confidence level of the gesture detection results according to the confidence level queue.
[0153] The video frame determination module 504 is configured to: determine a video frame with a standard confidence level greater than or equal to a preset confidence level threshold and a similarity degree greater than or equal to a preset similarity degree threshold as a target video frame; and determine a video frame with a standard confidence level less than the preset confidence level threshold or a similarity degree less than the preset similarity degree threshold as a non-target video frame.
[0154] The result output module 505 is configured to: when the video frame is a target video frame, calculate a first gesture parameter according to the sequence value corresponding to the target video frame; calculate a second gesture parameter according to the standard confidence level and similarity of the target video frame; generate a gesture recognition result based on the standard confidence level, the first gesture parameter, and the second gesture parameter, where the gesture recognition result is used to characterize the gesture category and gesture position of the target video frame; when the video frame is a non-target video frame, input the backward video frame into the gesture recognition network to perform gesture recognition on the backward video frame, where the backward video frame is the next frame of the non-target video frame.
[0155] Some embodiments of the present application further provide an electronic device, including: a memory, and one or more processors. The memory is configured to store one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the gesture recognition method.
[0156] Figure 6 It is a schematic structural diagram of the electronic device provided by the embodiments of the present application. As Figure 6 shown, the electronic device includes a processor 601, at least one communication bus 602, a user interface 603, at least one external communication interface 604, and a memory 605. Among them, the communication bus 602 is configured to implement connection communication between these components. The user interface 603 may include a display screen, and the external communication interface 604 may include a standard wired interface and a wireless interface. A computer program is stored in the memory 605. The processor 601 is used to execute the computer program stored in the memory 605.
[0157] For the similar parts between the embodiments provided in the present application, reference can be made to each other. The specific embodiments provided above are only several examples under the general concept of the present application, and do not constitute a limitation on the protection scope of the present application. For those skilled in the art, any other implementation manner extended based on the solution of the present application without creative efforts belongs to the protection scope of the present application.
Claims
1. A gesture recognition method, characterized in that: include: Obtain a video frame sequence, wherein the video frame sequence includes video frames and sequence values corresponding to the video frames, and the sequence value is a positive integer greater than or equal to 1; Traversing the video frames in the video frame sequence, and inputting the video frames into a gesture recognition network to obtain gesture detection results corresponding to the video frames, wherein the gesture recognition network is a deep neural network constructed based on a similarity score, and the gesture detection results include gesture coordinates, gesture categories, original confidence, and similarity; Generate a confidence queue based on the original confidence, and generate a standard confidence of the gesture detection result according to the confidence queue; Determine the video frame whose standard confidence is greater than or equal to a preset confidence threshold and whose similarity is greater than or equal to a preset similarity threshold as a target video frame; and determine the video frame whose standard confidence is less than the preset confidence threshold or whose similarity is less than the preset similarity threshold as a non-target video frame; In the case where the video frame is the target video frame, a first gesture parameter is calculated according to a sequence value corresponding to the target video frame; and a second gesture parameter is calculated according to the standard confidence and the similarity of the target video frame; A gesture recognition result is generated based on the standard confidence, the first gesture parameter, and the second gesture parameter, where the gesture recognition result is used to characterize the gesture category and gesture position of the target video frame.
2. The gesture recognition method according to claim 1, characterized in that: Before inputting the video frame into a gesture recognition network to obtain a gesture detection result corresponding to the video frame, the method further includes: Acquire a case data set, wherein the case data set includes at least one training gesture data, and the training gesture data is provided with a gesture annotation area and a gesture annotation category; Generate a standard gesture label according to the training gesture data, and calculate a similarity score between the training gesture data and the standard gesture label, wherein the similarity score is used to characterize the difference between the training gesture data and the standard gesture label, and the similarity score is associated with the training gesture data; Inputting the training gesture data and the similarity score into the gesture recognition network, and controlling the gesture recognition network to output regression parameters, wherein the regression parameters include coordinate information, size information, confidence information, classification result information, and similarity information; The gesture recognition network is trained based on the regression parameters, the gesture annotation categories and the loss function to optimize the model parameters of the gesture recognition network.
3. The gesture recognition method according to claim 2, characterized in that: The step of generating a standard gesture label according to the training gesture data comprises: Grouping the training gesture data according to the gesture annotation categories to obtain at least one set of gesture classification groups; Acquire a target gesture area of the training gesture data according to the gesture annotation area, and input the target gesture area into a feature extraction model to obtain gesture features, wherein the gesture features are associated with the training gesture data; A standard gesture label is generated according to the gesture feature, and the standard gesture label is used to represent a reference gesture of the gesture annotation category.
4. The gesture recognition method according to claim 3, characterized in that: The training gesture data includes continuous frame data and discontinuous frame data; In a case where the training gesture data is the continuous frame data, marking the gesture features of the training gesture data as continuous frame gesture features; In the case where the training gesture data is the non-continuous frame data, a preset weight is multiplied by the gesture feature of the training gesture data to obtain a non-continuous frame gesture feature.
5. The gesture recognition method according to claim 4, characterized in that: The step of generating a standard gesture label according to the gesture feature comprises: Acquire reference gesture features of the training gesture data in each of the gesture category groups, the reference gesture features including the continuous frame gesture features and / or the non-continuous frame gesture features; The average value of the reference gesture features is calculated to obtain the standard gesture label.
6. The gesture recognition method according to claim 3, characterized in that: The calculating the similarity score between the training gesture data and the standard gesture label comprises: Acquire reference gesture data, where the reference gesture data is a standard gesture label of the gesture classification group where the training gesture data belongs; A cosine distance between the training gesture data and the reference gesture data is calculated to obtain the similarity score.
7. The gesture recognition method according to claim 3, characterized in that: The loss function includes class loss, target loss, position loss and similarity loss, wherein the similarity loss is the quotient of the class edit distance and the number of classes, wherein the class edit distance is used to characterize the difference between the classification result information and the gesture annotation class, and the number of classes is used to characterize the number of the gesture classification groups; The step of training the gesture recognition network based on the regression parameter, the gesture annotation category and the loss function to optimize the model parameters of the gesture recognition network includes: Optimizing classification parameters of the gesture recognition network based on the class loss, wherein the classification parameters are used to calculate the classification result information; Optimizing a confidence parameter of the gesture recognition network based on the target loss, wherein the confidence parameter is used to calculate the confidence information; Optimizing position parameters of the gesture recognition network based on the position loss, wherein the position parameters are used to calculate the coordinate information, width information, and height information; A similarity parameter of the gesture recognition network is optimized based on the similarity loss, and the similarity parameter is used to calculate the similarity information.
8. The gesture recognition method according to claim 1, characterized in that: Generating a confidence queue based on the original confidence, and generating a standard confidence of the gesture detection result according to the confidence queue, including: When the sequence value of the video frame is equal to 1, the standard confidence of the gesture detection result is the original confidence; When the sequence value of the video frame is greater than 1, a reference confidence is obtained, where the reference confidence is the original confidence of the reference frame, and the sequence value corresponding to the reference frame is less than the sequence value corresponding to the video frame; and the confidence queue is generated according to the reference confidence and the original confidence, and the average value of the reference confidence and the original confidence is calculated to obtain the standard confidence of the gesture detection result.
9. The gesture recognition method according to claim 1, characterized in that: Also includes: In a case where the video frame is the non-target video frame, a backward video frame is input into the gesture recognition network to perform gesture recognition on the backward video frame, where the backward video frame is a next frame of the non-target video frame.
10. The gesture recognition method according to claim 1, characterized in that: The calculating the first gesture parameter according to the sequence value corresponding to the target video frame includes: Generate a state index according to the sequence value corresponding to the target video frame, where the state index is used to represent the number of target video frames between the video frame with the sequence value of 1 and the current video frame in the video frame sequence; The first gesture parameter is calculated using a weighted average formula, where the weighted average formula is: , in, represents the state index, represents the first gesture parameter, represents the intensity factor, which is calculated using the intensity formula, which is: , in, Indicates the preset frame number. Indicates the preset step size.
11. The gesture recognition method according to claim 1, characterized in that: The calculating the second gesture parameter according to the standard confidence and the similarity comprises: When the sequence value of the target video frame is equal to 1, setting the second gesture parameter to a first preset value; When the sequence value of the target video frame is greater than 1, a forward confidence and a forward similarity are obtained, wherein the forward confidence is a standard confidence of the forward video frame, the forward similarity is a similarity of the forward video frame, and the forward video frame is a previous frame of the target video frame; When the forward confidence is less than the standard confidence of the target video frame, and the forward similarity is less than the similarity of the target video frame, setting the second gesture parameter to a first preset value; When the forward confidence is greater than or equal to the standard confidence of the target video frame, and the forward similarity is greater than or equal to the similarity of the target video frame, the second gesture parameter is set to a second preset value.
12. The gesture recognition method according to claim 1, characterized in that: The generating a gesture recognition result based on the standard confidence, the first gesture parameter and the second gesture parameter includes: multiplying the standard confidence, the first gesture parameter, and the second gesture parameter to obtain a first detection result; Performing non-maximum suppression algorithm processing on the first detection result to obtain a second detection result; The second detection result is decoded to generate the gesture recognition result.
13. A gesture recognition device, characterized in that: include: A video frame acquisition module is configured to: obtain a video frame sequence, wherein the video frame sequence includes video frames and sequence values corresponding to the video frames, and the sequence value is a positive integer greater than or equal to 1; A gesture detection module is configured to: traverse the video frames in the video frame sequence, input the video frames into a gesture recognition network to obtain a gesture detection result corresponding to the video frame, wherein the gesture recognition network is a deep neural network constructed based on a similarity score, and the gesture detection result includes gesture coordinates, gesture category, original confidence and similarity; A confidence processing module is configured to: generate a confidence queue based on the original confidence, and generate a standard confidence of the gesture detection result according to the confidence queue; The video frame determination module is configured to: determine the video frame whose standard confidence is greater than or equal to a preset confidence threshold and whose similarity is greater than or equal to a preset similarity threshold as a target video frame; and determine the video frame whose standard confidence is less than the preset confidence threshold or whose similarity is less than the preset similarity threshold as a non-target video frame; The result output module is configured to: when the video frame is the target video frame, calculate the first gesture parameter according to the sequence value corresponding to the target video frame; and calculate the second gesture parameter according to the standard confidence and the similarity of the target video frame; generating a gesture recognition result based on the standard confidence, the first gesture parameter, and the second gesture parameter, wherein the gesture recognition result is used to characterize the gesture category and gesture position of the target video frame; In a case where the video frame is the non-target video frame, a backward video frame is input into the gesture recognition network to perform gesture recognition on the backward video frame, where the backward video frame is a next frame of the non-target video frame.
14. An electronic device, characterized in that: include: memory, one or more processors; The memory is configured to: store one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the gesture recognition method according to any one of claims 1-12.
Citation Information
Patent Citations
Self-weight fitness auxiliary coach system, method and terminal based on human body posture recognition
CN113762133A
Methods and Systems for Applications for Z-numbers
US20130330008A1