Gesture recognition method and device, electronic device, and storage medium

By using a pre-trained gesture recognition model and a perceptual coding network, the positional information of gestures in a preset space is represented, which solves the problem of insufficient accuracy in gesture recognition in existing technologies and achieves accurate recognition of gestures with subtle differences.

CN115311683BActive Publication Date: 2026-04-10NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-04
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing gesture recognition solutions cannot perceive subtle differences between different gestures, resulting in poor accuracy of recognition results.

Method used

A pre-trained gesture recognition model is used. The position information of the gesture in the preset space is represented by the perceptual encoding. The perceptual encoding of the gesture image to be recognized is extracted as feature information, and gesture recognition is performed based on the perceptual encoding. The gesture perceptual encoding network is trained to narrow the distance between similar gestures and widen the distance between dissimilar gestures.

Benefits of technology

It improves the accuracy of gesture recognition, enabling the identification of gestures with subtle differences and enhancing the perceptual ability of gesture image recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115311683B_ABST
    Figure CN115311683B_ABST
Patent Text Reader

Abstract

The application provides a gesture recognition method and device, electronic equipment and storage medium, and relates to the technical field of deep learning. The method comprises: acquiring at least one frame of to-be-recognized gesture image; using a pre-trained gesture recognition model to recognize the perception code of each to-be-recognized gesture image, and recognizing the gesture recognition result of at least one frame of to-be-recognized gesture image based on the perception code, wherein the perception code is used to represent the position information of the gesture in the to-be-recognized gesture image in a preset space. The method extracts the perception code of the to-be-recognized gesture image as the feature information of the to-be-recognized gesture image by using the trained gesture recognition model. Since the perception code represents the position information of the gesture in the preset space, the position information corresponding to different gestures is different, that is, each gesture has a unique corresponding perception code. By introducing the perception code as the feature information of the gesture image, the perception ability of similar gestures can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of deep learning, in particular to a gesture recognition method and device, electronic equipment and a storage medium. BACKGROUND

[0002] With the rapid development of deep learning and neural network technology, gesture recognition is widely used in intelligent home appliances, game interaction, AR(Augmented Reality) / VR(Virtual Reality) interaction, smart phone manipulation and other scenarios due to its convenience. The user experience is largely dependent on the accuracy of gesture recognition.

[0003] Most of the current gesture recognition schemes use a network model trained by sample gesture images pre-labeled with gesture categories to perform gesture recognition.

[0004] However, the above method cannot perceive the subtle differences between different gestures, resulting in poor accuracy of gesture recognition results. SUMMARY

[0005] The present application aims to solve the problem of poor accuracy of gesture recognition results due to the inability to perceive the differences between different gestures in the prior art.

[0006] To achieve the above purpose, the technical solutions adopted by the embodiments of the present application are as follows:

[0007] In a first aspect, the embodiments of the present application provide a gesture recognition method, comprising:

[0008] obtaining at least one frame of gesture image to be recognized;

[0009] using a pre-trained gesture recognition model to recognize the perception code of each gesture image to be recognized, and obtaining the gesture recognition result of the at least one frame of gesture image to be recognized based on the perception code, wherein the perception code is used to represent the position information of the gesture in the pre-set space in the gesture image to be recognized.

[0010] In a second aspect, the embodiments of the present application also provide a gesture recognition device, comprising an acquisition module and a recognition module.

[0011] The acquisition module is configured to obtain at least one frame of gesture image to be recognized.

[0012] The identification module is configured to identify a perception code of each of the to-be-identified gesture images by using the pre-trained gesture identification model, and identify a gesture identification result of the at least one frame of to-be-identified gesture images based on the perception code, where the perception code is used to represent position information of a gesture in the to-be-identified gesture images in a preset space.

[0013] In a third aspect, an electronic device is provided, including: a processor, a storage medium, and a bus. The storage medium stores machine readable instructions executable by the processor. When the electronic device is running, the processor communicates with the storage medium through the bus. The processor executes the machine readable instructions to perform the steps of the method provided in the first aspect.

[0014] In a fourth aspect, a storage medium is provided. The storage medium stores a computer program. When the computer program is executed by a processor, the steps of the method provided in the first aspect are performed.

[0015] The present application has the following beneficial effects:

[0016] The present application provides a gesture identification method and device, an electronic device, and a storage medium. The method includes: obtaining at least one frame of to-be-identified gesture images; identifying a perception code of each of the to-be-identified gesture images by using a pre-trained gesture identification model, and identifying a gesture identification result of the at least one frame of to-be-identified gesture images based on the perception code. The perception code is used to represent position information of a gesture in the to-be-identified gesture images in a preset space. The present application extracts the perception code of the to-be-identified gesture images as feature information of the to-be-identified gesture images by using the trained gesture identification model. Since the perception code represents the position information of the gesture in the preset space, different gestures correspond to different position information, that is, each gesture has a unique perception code. Even if the gestures have slight differences, the perception codes of the gestures can be obtained respectively. Therefore, the gesture images can be accurately identified based on the perception codes. By introducing the perception code as the feature information of the gesture images, the perception ability of the gesture identification can be effectively improved, the slight differences between different gestures can be identified, and the accuracy of the gesture image identification result is improved.

[0017] The perception code of the gesture image can be extracted by using a gesture perception code network in the trained gesture identification model. The gesture perception code network is trained based on the constructed gesture triplets. In the training process, the distance between the perception codes of similar gestures in the continuous representation space is shortened, and the distance between the perception codes of dissimilar gestures in the continuous representation space is lengthened based on the modification principle that the distance between similar gestures is smaller than the distance between dissimilar gestures. The trained gesture perception code network can perceive the slight differences between similar gestures, thereby improving the accuracy of the gesture image identification result. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 A flowchart illustrating a gesture recognition method provided in an embodiment of this application;

[0020] Figure 2 A flowchart illustrating another gesture recognition method provided in an embodiment of this application;

[0021] Figure 3 This is a schematic diagram of the architecture of a gesture recognition system provided in an embodiment of this application;

[0022] Figure 4 A flowchart illustrating yet another gesture recognition method provided in an embodiment of this application;

[0023] Figure 5 A flowchart illustrating another gesture recognition method provided in an embodiment of this application;

[0024] Figure 6 A flowchart illustrating another gesture recognition method provided in an embodiment of this application;

[0025] Figure 7 A flowchart illustrating yet another gesture recognition method provided in an embodiment of this application;

[0026] Figure 8 A flowchart illustrating another gesture recognition method provided in an embodiment of this application;

[0027] Figure 9 This is a schematic diagram of a gesture recognition result provided in an embodiment of this application;

[0028] Figure 10 A schematic diagram of a gesture recognition device provided in an embodiment of this application;

[0029] Figure 11 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0030] In order to make the purposes, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. It should be understood that the drawings in the present application serve only the purpose of description and illustration, and do not serve to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn according to the actual proportions. The flowcharts used in the present application show the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowcharts can not be implemented in sequence, and the steps without logical context relationship can be reversed in sequence or implemented simultaneously. In addition, one or more other operations can be added to the flowcharts or one or more operations can be removed from the flowcharts under the guidance of the content of the present application.

[0031] In addition, the described embodiments are only some of the embodiments of the present application, not all the embodiments. The components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0032] In order to enable those skilled in the art to use the content of the present application, the following implementation is given in combination with a specific application scenario "gesture recognition". For those skilled in the art, the general principles defined herein can be applied to other embodiments and application scenarios without departing from the spirit and scope of the present application. Although the present application is mainly described in relation to gesture recognition, it should be understood that this is only an exemplary embodiment. The present application can be applied to any other object recognition scenario. For example, the present application can be applied to some recognition scenarios such as expression recognition, face recognition, etc.

[0033] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.

[0034] Figure 1 A flowchart of a gesture recognition method provided by the embodiments of the present application is shown in FIG. 1. The execution subject of the method can be a computer device. As shown in FIG. 1, the method can include: Figure 1

[0035] S101, acquiring at least one frame of to-be-recognized gesture image.

[0036] ​The method can be used for recognizing a single frame of gesture image, and can also be used for recognizing continuous multiple frames of gesture image, which can be understood as recognizing a dynamic gesture image.

[0037] The at least one frame of gesture image to be recognized can be obtained in real time from a video, or can be directly obtained from a gesture image pre-shot and stored in a database.

[0038] In S102, a pre-trained gesture recognition model is used to recognize the perceptual code of each gesture image to be recognized, and a gesture recognition result of the at least one frame of gesture image to be recognized is obtained based on the perceptual code. The perceptual code is used to represent the position information of the gesture in the preset space in the gesture image to be recognized.

[0039] In this embodiment, the gesture recognition model can be used to recognize the at least one frame of gesture image to be recognized, specifically to recognize the gesture category in the gesture image to be recognized, so as to determine which gesture the gesture in the gesture image to be recognized belongs to, for example, a fist, a V sign, a hand waving, etc.

[0040] The gesture recognition model can extract the perceptual code of each gesture image to be recognized as the feature information of each gesture image to be recognized based on the input of each gesture image to be recognized, and perform gesture category recognition according to the perceptual code of each gesture image to be recognized, so as to obtain the gesture recognition result of the at least one frame of gesture image to be recognized.

[0041] The perceptual code is used to represent the position information of the gesture in the preset space in the gesture image to be recognized. The position information herein can refer to the absolute position information of a single gesture in space, or can refer to the relative position information of different gestures in space. The position information of different gestures in the preset space is unique.

[0042] In some application scenarios, when the at least one frame of gesture image to be recognized is directly subjected to gesture recognition, the absolute position information of the gesture in the preset space in each gesture image to be recognized can be determined based on the extracted perceptual code of each gesture image to be recognized, so as to determine the category of the gesture in each gesture image to be recognized according to the absolute position information.

[0043] In another application scenario, if a gesture image containing a target gesture or an image video containing the target gesture is given, and it is needed to find a candidate gesture same as or similar to the target gesture from another given video or from a database, the absolute position information of the perceptual code of the target gesture and the perceptual code of each candidate gesture can be determined based on the extracted perceptual code of the target gesture and the extracted perceptual code of each candidate gesture, and the candidate gesture closest to the absolute position of the target gesture is taken as the final candidate gesture to be extracted.

[0044] Since the positions of different gestures in the preset space are unique, the perceptual encodings of different gestures are also unique, that is, even if the gestures are similar, they also have unique perceptual encodings, so that based on the perceptual encodings of the extracted gesture images to be recognized, the categories of the gestures in the gesture images can be uniquely recognized, and even the gestures with high similarity can also be accurately recognized according to the perceptual encodings.

[0045] Compared with the prior art, by classifying different gesture images, training a recognition model, and using the recognition model to recognize gestures, in the recognition process, the specific gesture in the extracted gesture image is used as the recognition result. When two gestures are similar, the same recognition result is easily obtained, that is, the gesture images with subtle differences cannot be accurately recognized.

[0046] Since the perceptual encodings corresponding to different gestures are unique, the gesture images can be accurately recognized based on the perceptual encodings, and even the gestures with subtle differences can be accurately recognized, thereby effectively improving the accuracy of the gesture recognition result.

[0047] In summary, the gesture recognition method provided in the embodiment includes: obtaining at least one frame of gesture image to be recognized; using a pre-trained gesture recognition model to recognize the perceptual encoding of each gesture image to be recognized, and obtaining the gesture recognition result of at least one frame of gesture image to be recognized based on the perceptual encoding. The perceptual encoding is used to represent the position information of the gesture in the preset space in the gesture image to be recognized. The method extracts the perceptual encoding of the gesture image to be recognized as the feature information of the gesture image to be recognized by using the trained gesture recognition model. Since the perceptual encoding represents the position information of the gesture in the preset space, the position information corresponding to different gestures is different, that is, different gestures have unique perceptual encodings, and even the gestures with subtle differences can also obtain the perceptual encoding of each gesture, so that accurate gesture image recognition can be performed based on the perceptual encoding. By introducing the perceptual encoding as the feature information of the gesture image, the perceptual ability of gesture recognition can be effectively improved, the subtle differences between different gestures can be recognized, and the accuracy of the gesture image recognition result is improved.

[0048] Figure 2 Another flowchart of a gesture recognition method provided in the embodiment of the application is shown in the figure. Optionally, the gesture recognition model can include a gesture perceptual encoding network and a gesture discriminator. As shown in the figure, Figure 2 In step S102, the pre-trained gesture recognition model is used to recognize the perceptual encoding of each gesture image to be recognized, and the gesture recognition result of at least one frame of gesture image to be recognized is obtained based on the perceptual encoding. It can include:

[0049] S201, input at least one frame of to-be-recognized gesture image into a gesture perception coding network, and extract the perception coding of each to-be-recognized gesture image.

[0050] In the embodiment, the gesture recognition module can be composed of two parts. The first part is the gesture perception coding network, which is used for feature extraction of the to-be-recognized gesture image, so as to extract the perception coding of the to-be-recognized gesture image. The second part can be the gesture discriminator, which can perform gesture recognition according to the perception coding of the to-be-recognized gesture image.

[0051] Figure 3 The gesture recognition system provided in the embodiment of the present application is shown in the schematic diagram of the architecture. The gesture perception coding network in the gesture recognition model can be placed in the front end of the gesture discriminator. At least one frame of to-be-recognized gesture image is input into the gesture perception coding network as input data. The gesture perception coding network extracts the perception coding of each to-be-recognized gesture image as the feature information of each to-be-recognized gesture image.

[0052] S202, input the perception coding of each to-be-recognized gesture image into the gesture discriminator, and sequentially identify the gesture recognition result of each to-be-recognized gesture image.

[0053] The perception coding of each to-be-recognized gesture image output by the gesture perception coding network is input into the gesture discriminator connected in the rear end as input data. The gesture discriminator can obtain the gesture recognition result of each to-be-recognized gesture image based on the perception coding of each to-be-recognized gesture image, that is, obtain the gesture category of each gesture in each to-be-recognized gesture image, and output the gesture recognition result.

[0054] Optionally, the gesture recognition model can only include the gesture discriminator. The above gesture perception coding network can be a network independent of the gesture recognition model. After the perception coding of the to-be-recognized gesture image is extracted, the to-be-recognized gesture image and the perception coding of each to-be-recognized gesture image are input into the gesture recognition model for gesture recognition.

[0055] In the embodiment, the gesture perception coding network is used as a part of the gesture recognition model. In actual application, even if the gesture perception coding network is a network independent of the gesture recognition model, the gesture recognition method of the present application is still applicable.

[0056] Figure 4 The flowchart of another gesture recognition method provided in the embodiment of the present application is shown. Optionally, the above gesture recognition model can be obtained by training in the following manner:

[0057] S401, a sample training set is collected, the sample training set includes a plurality of target gesture triplets, each target gesture triplet is composed of a first gesture image, a second gesture image and a third gesture image, the similarity of the second gesture image and the first gesture image is greater than a first preset threshold, the similarity of the third gesture image and the first gesture image is less than a second preset threshold, the first preset threshold is greater than the second preset threshold; each target gesture triplet has annotation information, the annotation information includes: image similarity indication information and gesture category of each gesture image.

[0058] In this embodiment, the pre-constructed target gesture triplet is taken as a training sample, and the sample training set is obtained by combining a plurality of target gesture triplets.

[0059] Each target gesture triplet can include three gesture images, wherein the similarity of the first gesture image and the second gesture image is greater than a first preset threshold, the similarity of the third gesture image and the first gesture image is less than a second preset threshold, and the first preset threshold is greater than the second preset threshold.

[0060] It can be understood that the first gesture image is taken as a target gesture image, the second gesture image has high similarity with the first gesture image, and the second gesture image is taken as an image similar to the target gesture image; the third gesture image has low similarity with the first gesture image, and the third gesture image is taken as an image dissimilar to the target gesture image.

[0061] Each target gesture triplet has annotation information, the annotation information can include: image similarity indication information, the image similarity indication information can indicate the identification of the gesture image dissimilar to the target gesture image, or can indicate the identification of the gesture image similar to the target gesture image, that is, the identification of the first gesture image and the third gesture image in the target gesture triplet, or the identification of the first gesture image and the second gesture image in the target gesture triplet, so that in the training process, the first gesture image, the second gesture image and the third gesture image in the target gesture triplet can be accurately determined according to the annotated image similarity indication information, for training the gesture perception coding network.

[0062] In addition, the annotation information can also include: gesture category of each gesture image, that is, the gesture category of the first gesture image, the second gesture image and the third gesture image in the target gesture triplet can also be annotated respectively, for training the gesture discriminator.

[0063] S402, inputting the sample training set as input data into the initial perception coding network to train and obtain the gesture perception coding network.

[0064] Based on the above collected sample training set, it can be input into the initial perception coding network, and the gesture perception coding network is trained.

[0065] The initial perception coding network can be understood as having the same network architecture as the final gesture perception coding network to be trained, but the network parameters of the initial perception coding network are initial default parameters. After learning the sample training set data, the initial default parameters can be continuously adjusted to obtain target network parameters. The initial perception coding network with the target network parameters is used as the trained gesture perception coding network.

[0066] S403, input the sample training set as input data into the gesture perception coding network, and obtain the perception coding of each gesture image in each target gesture triple output by the gesture perception coding network.

[0067] Based on the trained gesture perception coding network, the sample training set can be re-input into the gesture perception coding network to obtain the perception coding of each gesture image in each target gesture triple extracted by the gesture perception coding network.

[0068] S404, input the sample training set and the perception coding of each gesture image in each target gesture triple output by the gesture perception coding network into the initial discriminator as input data, and train to obtain the gesture discriminator.

[0069] In some embodiments, the sample training set and the perception coding of each gesture image in each target gesture triple extracted by the gesture perception coding network can be input into the initial discriminator as input data to train the gesture discriminator.

[0070] Similarly, the initial discriminator here can also be understood as having the same network architecture as the final gesture discriminator to be trained, but the network parameters of the initial discriminator are initial default parameters. After learning the input data, the initial default parameters can be continuously adjusted to obtain target network parameters. The initial discriminator with the target network parameters is used as the trained gesture discriminator.

[0071] Figure 5 Another flowchart of a gesture recognition method provided by the embodiment of the present application is provided. Optionally, in step S401, collecting the sample training set can include:

[0072] S501, collect a plurality of sample gesture images, and construct a plurality of candidate gesture triples based on the plurality of sample gesture images.

[0073] Optionally, a large number of sample gesture images can be obtained from the gesture database in real time, and based on the sample gesture images, the matching data meeting the image condition in the target gesture triplets can be constructed, that is, the matching data of the target gesture image, the similar gesture image and the dissimilar gesture image are constructed, so as to obtain a plurality of candidate gesture triplets.

[0074] Herein, the candidate gesture triplets are called because not all of the constructed gesture triplets are required, and there may be invalid data, which need to be screened to determine the target gesture triplets from the candidate gesture triplets.

[0075] S502, validity verification is performed on each candidate gesture triplet, and the candidate gesture triplet meeting the preset condition is taken as a valid gesture triplet.

[0076] For each candidate gesture triplet, a plurality of annotators can perform information annotation, and if the consistency of the plurality of annotation results corresponding to any candidate gesture triplet reaches a threshold, it can be considered as a valid gesture triplet, otherwise it is considered as an invalid gesture triplet.

[0077] For example, for the candidate gesture triplet 1, six annotators perform information annotation to obtain six annotation results, and it is set that when four of the six annotation results are the same, it is considered that the consistency of the plurality of annotation results of the candidate gesture triplet 1 reaches the threshold, and then the candidate gesture triplet 1 can be taken as a valid gesture triplet.

[0078] For the candidate gesture triplet 2, six annotators perform information annotation to obtain six annotation results, and only three of the six annotation results are the same, so it is considered that the consistency of the plurality of annotation results of the candidate gesture triplet 2 does not reach the threshold, and then the candidate gesture triplet 2 can be taken as an invalid gesture triplet and screened out from the candidate gesture triplets.

[0079] S503, according to the plurality of annotation information of each valid gesture triplet, the target annotation information of each valid gesture triplet is determined.

[0080] For the determined valid gesture triplet, the target annotation information of the valid gesture triplet can be determined from the plurality of annotation information of the valid gesture triplet.

[0081] For example, four of the six annotation results of the valid gesture triplet 1 are the same, and then according to the voting mode of minority to majority, the annotation information of the same four annotation results is taken as the target annotation information of the valid gesture triplet.

[0082] The embodiment herein is only one implementable gesture triplet screening mode and annotation information determination mode.

[0083] S504, each valid gesture triplet and the target label information of each valid gesture triplet are taken as a target gesture triplet to obtain a sample training set.

[0084] Then, the sample training set can be composed of the valid gesture triplets determined above, and each valid gesture triplet has target label information.

[0085] Figure 6 Another flowchart of a gesture recognition method provided by the embodiment of the application is shown in FIG. 6. Optionally, in step S402, the sample training set is input into the initial perception coding network as input data to train the gesture perception coding network, which can include the following steps.

[0086] S601, the sample training set is input into the initial perception coding network as input data to obtain the predicted perception codes of each gesture image in each target gesture triplet output by the initial perception coding network.

[0087] Optionally, the sample training set obtained above can be input into the initial perception coding network to train the gesture perception coding network as a feature extractor.

[0088] During training, for each target gesture triplet <A, P, N> input, the initial perception coding network can output a corresponding result <f(A), f(P), f(N)> where A represents the first gesture image, P represents the second gesture image, and N represents the third gesture image, f(A) represents the predicted perception code of the first gesture image, f(P) represents the predicted perception code of the second gesture image, and f(N) represents the predicted perception code of the third gesture image.

[0089] S602, a first loss parameter of the initial perception coding network is calculated according to the predicted perception codes of each gesture image in each target gesture triplet.

[0090] For each target gesture triplet, a sub-loss function of each target gesture triplet can be obtained according to the predicted perception codes of each gesture image in each target gesture triplet, and the first loss parameter of the initial perception coding network can be obtained by averaging or summing the sub-loss functions of each target gesture triplet.

[0091] S603, the network parameters of the initial perception coding network are corrected according to the first loss parameter, and the correction is iteratively performed until the first loss parameter meets a third preset threshold, the correction is stopped, and the current initial perception coding network is taken as the gesture perception coding network.

[0092] Based on the first loss parameter of the current obtained initial perceptual coding network, it is determined whether the first loss parameter meets a first preset threshold. If yes, the current initial perceptual network is taken as the gesture perceptual coding network. If no, the network parameter of the current initial perceptual coding network is corrected to obtain a new initial perceptual coding network, and steps S601-S602 are repeatedly executed based on the new initial perceptual coding network to calculate the first loss parameter of the new initial perceptual coding network, and it is continuously determined whether the first loss parameter of the initial perceptual coding network meets the first preset threshold until the first loss parameter meets the third preset threshold, and the execution is stopped.

[0093] Figure 7 A flowchart of another gesture recognition method provided by an embodiment of the present application is shown in FIG. 6. Optionally, in step S602, the first loss parameter of the initial perceptual coding network is calculated according to the predicted perceptual codes of the gesture images in each target gesture triple. The calculation can include:

[0094] S701, the first gesture image, the second gesture image and the third gesture image in each target gesture triple are determined according to the image similarity indication information in the annotation information of each target gesture triple.

[0095] For each target gesture triple, the first gesture image, the second gesture image and the third gesture image in the target gesture triple can be determined according to the image similarity indication information contained in the annotation information of the target gesture triple.

[0096] Suppose the image similarity indication information indicates the identity of the first gesture image and the identity of the second gesture image, then the first gesture image and the second gesture image can be determined according to the identity of the first gesture image and the identity of the second gesture image, so that the remaining gesture image is taken as the third gesture image.

[0097] S702, the first distance corresponding to each target gesture triple is calculated according to the predicted perceptual code of the first gesture image and the predicted perceptual code of the second gesture image in each target gesture triple.

[0098] The first loss parameter of the initial perceptual coding network can be calculated by the following formula:

[0099] L TRI (A,P,N)=max(‖f(A)-f(P)‖ 2 -‖f(A)-f(N)‖ 2 +α,0)

[0100] wherein, ‖f(A)-f(P)‖ 2The distance between the predicted perceptual encoding of the gesture in the first gesture image and the predicted perceptual encoding of the gesture in the second gesture image in the continuous representation space. That is, the distance between the predicted perceptual encodings of the gestures in two similar gesture images in the continuous representation space; corresponding to the first distance above.

[0101] S703, calculating a second distance corresponding to each target gesture triple according to the predicted perceptual encoding of the first gesture image and the predicted perceptual encoding of the third gesture image in each target gesture triple.

[0102] ‖f(A)-f(N)‖ 2 The distance between the predicted perceptual encoding of the gesture in the first gesture image and the predicted perceptual encoding of the gesture in the third gesture image in the continuous representation space. That is, the distance between the predicted perceptual encodings of the gestures in two dissimilar gesture images in the continuous representation space; corresponding to the second distance above.

[0103] And a can represent a preset distance threshold, which can be set as an acceptable distance limit.

[0104] S704, determining a first loss parameter of the initial perceptual encoding network according to the first distance corresponding to each target gesture triple, the second distance corresponding to each target gesture triple, and the preset distance threshold.

[0105] Based on the first distance and the second distance corresponding to the target gesture triple obtained above and the preset distance threshold a, the calculation formula of the first loss function above can be used to obtain the first loss function of the initial perceptual encoding network in the current round.

[0106] If the first loss function of the initial perceptual encoding network in the current round does not satisfy the first preset threshold, the network parameters of the current initial perceptual encoding network can be corrected.

[0107] When the network parameters are corrected, it is required that the first distance calculated based on the predicted perceptual encoding of the first gesture image and the predicted perceptual encoding of the second gesture image obtained by the corrected initial perceptual encoding network is smaller than the first distance calculated in the last round before the correction, and the second distance calculated based on the predicted perceptual encoding of the first gesture image and the predicted perceptual encoding of the third gesture image obtained by the corrected initial perceptual encoding network is larger than the second distance calculated in the last round before the correction. That is, the distance between the perceptual encodings of similar gestures in the continuous representation space is shortened, and the distance between the perceptual encodings of dissimilar gestures in the continuous representation space is lengthened, so that the distance between similar gestures is smaller than the distance between dissimilar gestures.

[0108] Based on the above correction principle, the network parameter correction of the initial perceptual coding network can be iteratively performed. When the correction is completed, the current initial perceptual coding network is taken as the gesture perceptual coding network.

[0109] In this embodiment, when the gesture perceptual coding network is trained, the network learning is performed by using the constructed gesture triplets, the perceptual coding of the gesture image is predicted, the first loss parameter is calculated, and when the network parameter correction is performed based on the first loss parameter, the principle of continuously narrowing the distance between similar gestures and continuously widening the distance between dissimilar gestures in the target gesture triplet is used, so that the gesture perceptual coding network obtained by training can accurately distinguish the perceptual coding of similar gestures, and can accurately extract the perceptual coding corresponding to any gesture image. When the gesture recognition is performed based on the perceptual coding, the perceptual ability of gesture recognition is improved, the subtle differences between gestures can be recognized, and the purpose of improving the accuracy of gesture recognition result is achieved.

[0110] Figure 8 Another flowchart of a gesture recognition method provided by the embodiment of the application is shown in the figure. Optionally, in step S403, the sample training set and the perceptual coding of each gesture image in each target gesture triplet output by the gesture perceptual coding network are input into the initial discriminator as input data, and the gesture discriminator is trained and obtained, which can include the following steps.

[0111] S801, input the sample training set and the perceptual coding of each gesture image in each target gesture triplet into the initial discriminator as input data, and obtain the predicted gesture category of each gesture image in each target gesture triplet output by the initial discriminator.

[0112] For the training of the gesture discriminator, the sample training set collected and the perceptual coding of each gesture image in each target gesture triplet obtained by processing the sample training set by the trained gesture perceptual coding network can be input into the initial discriminator as input data.

[0113] For the initial discriminator, the input data received can be multiple sample data, and each sample data includes: a target gesture triplet and the perceptual coding of each gesture image in the target gesture triplet.

[0114] S802, according to the predicted gesture category of each gesture image in each target gesture triplet output by the gesture discriminator and the gesture category of each gesture image in the annotation information of each target gesture triplet, a second loss function of the initial discriminator is calculated.

[0115] The initial discriminator can process the predicted gesture category of each gesture image in each target gesture triplet, and calculate a second loss parameter of the initial discriminator based on the gesture category of each gesture image labeled in the annotation information of each target gesture triplet in the input data and the predicted gesture category of each gesture image.

[0116] S803, correct the network parameter of the initial discriminator according to the second loss parameter, and iteratively execute until the second loss parameter meets a fourth preset threshold value, stop correction, and take the current initial discriminator as the gesture discriminator.

[0117] For each target gesture triplet, a sub-loss function of each target gesture triplet can be calculated, and by averaging or accumulating the sum of the sub-loss functions of each target gesture triplet, the second loss parameter of the initial discriminator can be obtained.

[0118] Based on the current obtained second loss parameter of the initial discriminator, it can be judged whether the second loss parameter meets the fourth preset threshold value. If it meets, the current initial discriminator can be taken as the gesture discriminator, and if it does not meet, the network parameter of the current initial discriminator is corrected to obtain a new initial discriminator, and steps S801-S802 are repeatedly executed based on the new initial discriminator to calculate the second loss parameter of the new initial discriminator, and it is continuously judged whether the second loss parameter of the new initial discriminator meets the fourth preset threshold value, until the second loss parameter meets the fourth preset threshold value, and the execution is stopped.

[0119] Alternatively, in step S802, the second loss function of the initial discriminator can be calculated according to the predicted gesture category of each gesture image in each target gesture triplet output by the gesture discriminator and the gesture category of each gesture image in the annotation information of each target gesture triplet, which can include: performing cross-entropy calculation according to the predicted gesture category of each gesture image in each target gesture triplet output by the gesture discriminator, the gesture category of each gesture image in the annotation information of each target gesture triplet, and the number of gesture images in the sample training set, to obtain the second loss function of the initial discriminator.

[0120] Alternatively, the second loss function can be calculated by the following formula:

[0121]

[0122] Wherein, N refers to the number of samples used in training, that is, the total number of gesture images in each target triplet in the sample training set. Assuming that the sample training set contains 5 target gesture triplets, since one target gesture triplet contains three gesture images, N is 15. i y i represents the actual value of the i th gesture image (that is, the gesture category of the i th gesture image labeled in the annotation information), a predicted value of the i-th gesture image (i.e., a predicted gesture class of the i-th gesture image output by the initial discriminator).

[0123] After the second loss parameter of the initial discriminator is calculated according to the above formula in each round, it can be judged whether the second loss parameter meets the fourth preset threshold. If it meets, the current initial discriminator is taken as the gesture discriminator. If it does not meet, the network parameter of the initial discriminator is modified until the second loss parameter meets the fourth preset threshold.

[0124] Optionally, in step S102, the pre-trained gesture recognition model is used to identify the perceptual code of each to-be-identified gesture image, and the gesture recognition result of at least one frame of to-be-identified gesture image is obtained based on the perceptual code. It can include:

[0125] If the at least one frame of to-be-identified gesture image includes multiple frames of to-be-identified gesture image, the pre-trained gesture recognition model is used to identify the perceptual code of each frame of to-be-identified gesture image in turn, and a gesture recognition result sequence of the at least one frame of to-be-identified gesture image is obtained based on the perceptual code of each frame of to-be-identified gesture image. The gesture recognition result sequence is composed of the gesture recognition results of each frame of to-be-identified gesture image in turn.

[0126] In the process of applying the gesture recognition model, if the input to-be-identified gesture image is a gesture image sequence, that is, it contains multiple frames of to-be-identified gesture image, then the perceptual code of each frame of to-be-identified gesture image can be extracted based on the gesture perceptual encoder first, and the perceptual codes of each frame of to-be-identified gesture image are combined to obtain a perceptual code sequence. The gesture discriminator can recognize the gesture recognition result sequence according to the perceptual code sequence, and the gesture recognition result sequence is arranged in frame order according to the gesture recognition result of each to-be-identified gesture image. The gesture recognition result of each to-be-identified gesture image can be the gesture class of the gesture in each to-be-identified gesture image.

[0127] Figure 9 A gesture recognition result schematic diagram provided by an embodiment of the present application. For the input to-be-identified gesture image sequence, after being processed by the gesture perceptual coding network, a perceptual code sequence can be output. The perceptual code sequence is input into the gesture discriminator, and after being processed by the gesture discriminator, a gesture recognition result sequence can be obtained. The gesture recognition result sequence is arranged from top to bottom, the first gesture class corresponds to the first frame of to-be-identified gesture image in the to-be-identified gesture image sequence, the second gesture class corresponds to the second frame of to-be-identified gesture image in the to-be-identified gesture image sequence, and so on.

[0128] In summary, the gesture recognition method provided in this embodiment includes: acquiring at least one frame of a gesture image to be recognized; using a pre-trained gesture recognition model to identify the perceptual codes of each gesture image to be recognized, and acquiring the gesture recognition result of at least one frame of the gesture image to be recognized based on the perceptual codes. The perceptual codes are used to represent the positional information of the gesture in the gesture image to be recognized in a preset space. This method extracts the perceptual codes of the gesture images to be recognized as feature information of the gesture images through the trained gesture recognition model. Since the perceptual codes represent the positional information of the gesture in the preset space, different gestures correspond to different positional information. That is, different gestures have unique corresponding perceptual codes. Even gestures with subtle differences can obtain the perceptual codes of each gesture, thereby enabling accurate gesture image recognition based on the perceptual codes. By introducing perceptual codes as feature information of gesture images, the perceptual ability of gesture recognition can be effectively improved, and subtle differences between different gestures can be identified, thereby improving the accuracy of gesture image recognition results.

[0129] The gesture recognition model can extract the perceptual encoding of gesture images through a gesture perception encoding network. The gesture perception encoding network is trained from the constructed gesture triples. During training, based on the correction principle of narrowing the distance between the perceptual encodings of similar gestures in the continuous representation space and widening the distance between the perceptual encodings of dissimilar gestures in the continuous representation space, so that the distance between similar gestures is smaller than the distance between dissimilar gestures, the trained gesture perception encoding network can perceive subtle differences between similar gestures, thereby improving the accuracy of gesture image recognition results.

[0130] The following describes the apparatus, device, and storage medium used to execute the gesture recognition method provided in this application. The specific implementation process and technical effects are described above and will not be repeated below.

[0131] Figure 10 This is a schematic diagram of a gesture recognition device provided in an embodiment of this application. The function implemented by this gesture recognition device corresponds to the steps performed by the method described above. This device can be understood as the aforementioned electronic device or computer device, such as... Figure 10 As shown, the device may include: an acquisition module 110 and an identification module 120;

[0132] The acquisition module 110 is used to acquire at least one frame of the gesture image to be recognized;

[0133] The recognition module 120 is used to identify the perceptual encoding of each gesture image to be recognized using a pre-trained gesture recognition model, and to obtain the gesture recognition result of at least one frame of gesture image to be recognized based on the perceptual encoding. The perceptual encoding is used to characterize the position information of the gesture in the gesture image to be recognized in a preset space.

[0134] Optionally, the gesture recognition model comprises: a gesture perception encoding network and a gesture discriminator.

[0135] The recognition module 120 is specifically configured to input at least one frame of the to-be-recognized gesture image into the gesture perception encoding network, and extract a perception encoding of each to-be-recognized gesture image.

[0136] The perception encoding of each to-be-recognized gesture image is input into the gesture discriminator, and a gesture recognition result of each to-be-recognized gesture image is sequentially recognized.

[0137] Optionally, the device further comprises a training module.

[0138] The training module is configured to collect a sample training set, the sample training set comprising a plurality of target gesture triplets, each target gesture triplet being composed of a first gesture image, a second gesture image and a third gesture image, the similarity of the second gesture image to the first gesture image being greater than a first preset threshold, the similarity of the third gesture image to the first gesture image being less than a second preset threshold, the first preset threshold being greater than the second preset threshold; each target gesture triplet having annotation information, the annotation information comprising: image similarity indication information and a gesture category of each gesture image.

[0139] The sample training set is input as input data into an initial perception encoding network to train and obtain the gesture perception encoding network.

[0140] The sample training set is input as input data into the gesture perception encoding network to obtain the perception encoding of each gesture image in each target gesture triplet output by the gesture perception encoding network.

[0141] The sample training set and the perception encoding of each gesture image in each target gesture triplet output by the gesture perception encoding network are input as input data into an initial discriminator to train and obtain the gesture discriminator.

[0142] Optionally, the training module is specifically configured to collect a plurality of sample gesture images, and construct a plurality of candidate gesture triplets based on the plurality of sample gesture images.

[0143] Each candidate gesture triplet is subjected to validity verification, and a candidate gesture triplet satisfying a preset condition is taken as a valid gesture triplet.

[0144] The target annotation information of each valid gesture triplet is determined according to a plurality of pieces of annotation information of each valid gesture triplet.

[0145] Each valid gesture triplet and the target annotation information of each valid gesture triplet are taken as a target gesture triplet to obtain the sample training set.

[0146] Optionally, the training module is specifically configured to input the sample training set as input data into the initial perceptual coding network, and obtain the predicted perceptual codes of the gesture images in each target gesture triplet output by the initial perceptual coding network.

[0147] According to the predicted perceptual codes of the gesture images in each target gesture triplet, a first loss parameter of the initial perceptual coding network is calculated.

[0148] According to the first loss parameter, the network parameters of the initial perceptual coding network are corrected, and the correction is iteratively performed until the first loss parameter meets a third preset threshold, the correction is stopped, and the current initial perceptual coding network is taken as the gesture perceptual coding network.

[0149] Optionally, the training module is specifically configured to determine the first gesture image, the second gesture image, and the third gesture image in each target gesture triplet according to the image similarity indication information in the annotation information of each target gesture triplet.

[0150] According to the predicted perceptual codes of the first gesture image and the second gesture image in each target gesture triplet, a first distance corresponding to each target gesture triplet is calculated.

[0151] According to the predicted perceptual codes of the first gesture image and the third gesture image in each target gesture triplet, a second distance corresponding to each target gesture triplet is calculated.

[0152] According to the first distance corresponding to each target gesture triplet, the second distance corresponding to each target gesture triplet, and a preset distance threshold, the first loss parameter of the initial perceptual coding network is determined.

[0153] Optionally, the training module is specifically configured to input the sample training set and the perceptual codes of the gesture images in each target gesture triplet as input data into the initial discriminator, and obtain the predicted gesture categories of the gesture images in each target gesture triplet output by the initial discriminator.

[0154] According to the predicted gesture categories of the gesture images in each target gesture triplet output by the gesture discriminator and the gesture categories of the gesture images in the annotation information of each target gesture triplet, a second loss function of the initial discriminator is calculated.

[0155] According to the second loss parameter, the network parameters of the initial discriminator are corrected, and the correction is iteratively performed until the second loss parameter meets a fourth preset threshold, the correction is stopped, and the current initial discriminator is taken as the gesture discriminator.

[0156] Optionally, the training module is specifically configured to perform cross-entropy calculation according to the predicted gesture category of each gesture image in each target gesture triple output by the gesture discriminator, the gesture category of each gesture image in the annotation information of each target gesture triple, and the number of gesture images in the sample training set, to obtain the second loss function of the initial discriminator.

[0157] Optionally, the training module is specifically configured to, if the at least one frame of to-be-recognized gesture image includes multiple frames of to-be-recognized gesture image, sequentially recognize the perceptual encoding of each frame of to-be-recognized gesture image by using the pre-trained gesture recognition model, and identify the gesture recognition result sequence of the at least one frame of to-be-recognized gesture image based on the perceptual encoding of each frame of to-be-recognized gesture image, the gesture recognition result sequence being composed of gesture recognition results of each frame of to-be-recognized gesture image arranged in sequence.

[0158] In the foregoing manner, when the electronic device executes the gesture recognition method, the acquisition module can acquire at least one frame of to-be-recognized gesture image, and the recognition model uses the pre-trained gesture recognition model to recognize the perceptual encoding of each to-be-recognized gesture image, and identifies the gesture recognition result of the at least one frame of to-be-recognized gesture image based on the perceptual encoding, the perceptual encoding being used to represent the position information of the gesture in the to-be-recognized gesture image in the preset space. In this method, the perceptual encoding of the to-be-recognized gesture image is extracted as the feature information of the to-be-recognized gesture image by using the trained gesture recognition model. Since the perceptual encoding represents the position information of the gesture in the preset space, different gestures correspond to different position information, that is, each gesture has a unique perceptual encoding, and even if the gestures have slight differences, the perceptual encoding of each gesture can be obtained, respectively, so that accurate gesture image recognition can be performed based on the perceptual encoding. By introducing the perceptual encoding as the feature information of the gesture image, the perceptual ability of gesture recognition can be effectively improved, and the slight differences between different gestures can be recognized, thereby improving the accuracy of the gesture image recognition result.

[0159] The perceptual encoding of the gesture image can be extracted by using the gesture perceptual encoding network in the trained gesture recognition model, the gesture perceptual encoding network being trained by using the constructed gesture triple. In the training process, the distance between the perceptual encodings of similar gestures in the continuous representation space is narrowed, and the distance between the perceptual encodings of dissimilar gestures in the continuous representation space is widened, according to the modification principle that the distance between similar gestures is less than the distance between dissimilar gestures, so that the trained gesture perceptual encoding network can perceive the slight differences between similar gestures, thereby improving the accuracy of the gesture image recognition result.

[0160] The above modules can be one or more integrated circuits configured to implement the above methods, for example, one or more application specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs), etc. For another example, when a certain module above is implemented in the form of a processing element scheduling code, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor that can invoke program code. For another example, the modules can be integrated together to implement in the form of a system on a chip (SOC).

[0161] The above modules can be connected or communicated with each other via wired connection or wireless connection. The wired connection can include metal cable, optical cable, hybrid cable, etc., or any combination thereof. The wireless connection can include connection in the form of LAN, WAN, Bluetooth, ZigBee, or NFC, etc., or any combination thereof. Two or more modules can be combined into a single module, and any one module can be divided into two or more units. It can be clearly understood by those skilled in the art that, for the convenience and brevity of description, the specific working process of the system and device described above can refer to the corresponding process in the method embodiment, which will not be described herein.

[0162] Figure 11 A structural schematic diagram of an electronic device provided in an embodiment of the present application includes a processor 801, a storage medium 802, and a bus 803. The storage medium 802 stores machine readable instructions executable by the processor 801. When the electronic device runs a gesture recognition method as in an embodiment, the processor 801 and the storage medium 802 communicate through the bus 803. The processor 801 executes the machine readable instructions to perform the following steps:

[0163] Obtain at least one frame of to-be-recognized gesture image;

[0164] Use a pre-trained gesture recognition model to recognize the perceptual code of each to-be-recognized gesture image, and obtain a gesture recognition result of the at least one frame of to-be-recognized gesture image based on the perceptual code. The perceptual code is used to represent the position information of the gesture in the to-be-recognized gesture image in a preset space.

[0165] In an implementable implementation, the gesture recognition model comprises a gesture perception encoding network and a gesture discriminator, and the processor 801, when executing the pre-trained gesture recognition model to recognize the perception encodings of the to-be-recognized gesture images and to recognize the gesture recognition results of the at least one frame of to-be-recognized gesture images based on the perception encodings, is specifically configured to:

[0166] input the at least one frame of to-be-recognized gesture images into the gesture perception encoding network to extract the perception encodings of the to-be-recognized gesture images;

[0167] input the perception encodings of the to-be-recognized gesture images into the gesture discriminator to sequentially recognize the gesture recognition results of the to-be-recognized gesture images.

[0168] In an implementable implementation, the processor 801, when executing the gesture recognition model training, is specifically configured to:

[0169] collect a sample training set, the sample training set comprising a plurality of target gesture triplets, each target gesture triplet being composed of a first gesture image, a second gesture image and a third gesture image, the similarity of the second gesture image to the first gesture image being greater than a first preset threshold, the similarity of the third gesture image to the first gesture image being less than a second preset threshold, the first preset threshold being greater than the second preset threshold, each target gesture triplet having annotation information, the annotation information comprising image similarity indication information and gesture categories of the gesture images;

[0170] input the sample training set as input data into an initial perception encoding network to train and obtain the gesture perception encoding network;

[0171] input the sample training set as input data into the gesture perception encoding network to obtain the perception encodings of the gesture images in each target gesture triplet output by the gesture perception encoding network;

[0172] input the sample training set and the perception encodings of the gesture images in each target gesture triplet output by the gesture perception encoding network as input data into an initial discriminator to train and obtain the gesture discriminator.

[0173] In an implementable implementation, the processor 801, when executing the collection of the sample training set, is specifically configured to:

[0174] collect a plurality of sample gesture images and construct a plurality of candidate gesture triplets based on the plurality of sample gesture images;

[0175] perform validity verification on each candidate gesture triplet, and take the candidate gesture triplets satisfying a preset condition as valid gesture triplets;

[0176] determine target annotation information of each valid gesture triplet according to votes on a plurality of pieces of annotation information of each valid gesture triplet.

[0177] The effective gesture triplets and the target annotation information of each effective gesture triplet are taken as target gesture triplets to obtain a sample training set.

[0178] In a feasible implementation, when the processor 801 executes the inputting of the sample training set as input data into the initial perception encoding network to train the gesture perception encoding network, the processor 801 is specifically configured to:

[0179] inputting the sample training set as input data into the initial perception encoding network to obtain the predicted perception encodings of the gesture images in each target gesture triplet output by the initial perception encoding network;

[0180] calculating a first loss parameter of the initial perception encoding network according to the predicted perception encodings of the gesture images in each target gesture triplet;

[0181] correcting the network parameters of the initial perception encoding network according to the first loss parameter, and iteratively performing until the first loss parameter meets a third preset threshold value, stopping the correction, and taking the current initial perception encoding network as the gesture perception encoding network.

[0182] In a feasible implementation, when the processor 801 executes the calculation of the first loss parameter of the initial perception encoding network according to the predicted perception encodings of the gesture images in each target gesture triplet, the processor 801 is specifically configured to:

[0183] determining a first gesture image, a second gesture image, and a third gesture image in each target gesture triplet according to the image similarity indication information in the annotation information of each target gesture triplet;

[0184] calculating a first distance corresponding to each target gesture triplet according to the predicted perception encoding of the first gesture image and the predicted perception encoding of the second gesture image in the target gesture triplet;

[0185] calculating a second distance corresponding to each target gesture triplet according to the predicted perception encoding of the first gesture image and the predicted perception encoding of the third gesture image in the target gesture triplet;

[0186] determining the first loss parameter of the initial perception encoding network according to the first distance corresponding to each target gesture triplet, the second distance corresponding to each target gesture triplet, and a preset distance threshold value.

[0187] In a feasible implementation, when the processor 801 executes the inputting of the sample training set and the perception encodings of the gesture images in each target gesture triplet output by the gesture perception encoding network as input data into the initial discriminator to train the gesture discriminator, the processor 801 is specifically configured to:

[0188] inputting the sample training set and the perceptual codes of the gesture images in each target gesture triple as input data into the initial discriminator, obtaining the predicted gesture categories of the gesture images in each target gesture triple output by the initial discriminator;

[0189] calculating a second loss function of the initial discriminator according to the predicted gesture categories of the gesture images in each target gesture triple output by the gesture discriminator and the gesture categories of the gesture images in the annotation information of each target gesture triple;

[0190] modifying the network parameters of the initial discriminator according to the second loss parameter, iteratively performing until the second loss parameter meets a fourth preset threshold, stopping the modification, and taking the current initial discriminator as the gesture discriminator.

[0191] In a feasible implementation, when the processor 801 performs the calculation of the second loss function of the initial discriminator according to the predicted gesture categories of the gesture images in each target gesture triple output by the gesture discriminator and the gesture categories of the gesture images in the annotation information of each target gesture triple, the processor 801 is specifically configured to:

[0192] performing cross-entropy calculation according to the predicted gesture categories of the gesture images in each target gesture triple output by the gesture discriminator, the gesture categories of the gesture images in the annotation information of each target gesture triple, and the number of gesture images in the sample training set, to obtain the second loss function of the initial discriminator.

[0193] In a feasible implementation, when the processor 801 performs the calculation of the second loss function of the initial discriminator according to the predicted gesture categories of the gesture images in each target gesture triple output by the gesture discriminator and the gesture categories of the gesture images in the annotation information of each target gesture triple, the processor 801 is specifically configured to:

[0194] If the at least one frame of gesture image to be recognized includes a plurality of frames of gesture image to be recognized, the pre-trained gesture recognition model is used to sequentially identify the perceptual codes of the frames of gesture image to be recognized, and a gesture recognition result sequence of the at least one frame of gesture image to be recognized is obtained based on the perceptual codes of the frames of gesture image to be recognized, the gesture recognition result sequence being composed of the gesture recognition results of the frames of gesture image to be recognized in sequence.

[0195] By the above manner, when the electronic device executes the gesture recognition method, the processor can acquire at least one frame of to-be-recognized gesture image, and adopts the pre-trained gesture recognition model to recognize the perception code of each to-be-recognized gesture image, and recognizes the gesture recognition result of at least one frame of to-be-recognized gesture image based on the perception code, and the perception code is used to represent the position information of the gesture in the preset space in the to-be-recognized gesture image. Wherein, the method extracts the perception code of the to-be-recognized gesture image as the feature information of the to-be-recognized gesture image by the trained gesture recognition model. Since the perception code represents the position information of the gesture in the preset space, the position information corresponding to different gestures is different, that is, each gesture has a unique corresponding perception code. Even if the gestures have slight differences, the perception codes of each gesture can be obtained respectively, so that the gesture image recognition can be accurately performed based on the perception code. By introducing the perception code as the feature information of the gesture image, the perception ability of the gesture recognition can be effectively improved, the slight differences between different gestures are recognized, and the accuracy of the gesture image recognition result is improved.

[0196] Wherein, the extraction of the perception code of the gesture image can be performed by the gesture perception code network in the trained gesture recognition model, the gesture perception code network is trained by the constructed gesture triplets, and in the training process, the distance between the perception codes of similar gestures in the continuous representation space is narrowed, and the distance between the perception codes of dissimilar gestures in the continuous representation space is widened, so that the distance between similar gestures is smaller than the distance between dissimilar gestures under the correction principle, so that the trained gesture perception code network can perceive the slight differences between similar gestures, thereby improving the accuracy of the gesture image recognition result.

[0197] Wherein, the storage medium 802 stores program codes, when the program codes are executed by the processor 801, the processor 801 executes various steps in the gesture recognition method according to various exemplary embodiments of the present application described in the above "exemplary method" part of the specification.

[0198] The processor 801 can be a general processor, such as a central processing unit (CPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, and can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the present application can be directly embodied as completed by a hardware processor, or completed by a combination of hardware and software modules in the processor.

[0199] The storage medium 802 is a non-volatile computer readable storage medium, and can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The storage medium can include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic storage, magnetic disk, optical disk, etc. The storage medium is any other medium capable of carrying or storing desired program code in the form of instructions or data structures and capable of being accessed by a computer, but is not limited to this. The storage medium 802 in the embodiments of the present application can also be a circuit or any other device capable of realizing a storage function, used to store program instructions and / or data.

[0200] Optionally, the embodiments of the present application further provide a computer readable storage medium, and the computer readable storage medium stores a computer program. When the computer program is run by a processor, the processor executes the following steps:

[0201] Obtaining at least one frame of to-be-identified gesture image;

[0202] The pre-trained gesture recognition model is used to recognize the perceptual codes of each to-be-recognized gesture image, and a gesture recognition result of at least one frame of to-be-recognized gesture image is obtained based on the perceptual codes. The perceptual code is used to represent position information of a gesture in a preset space in the to-be-recognized gesture image.

[0203] In a feasible implementation, the gesture recognition model includes a gesture perceptual coding network and a gesture discriminator. When the processor 801 executes the pre-trained gesture recognition model to recognize the perceptual codes of each to-be-recognized gesture image, and obtains the gesture recognition result of at least one frame of to-be-recognized gesture image based on the perceptual codes, the processor 801 is specifically configured to:

[0204] The at least one frame of to-be-recognized gesture image is input into the gesture perceptual coding network to extract the perceptual codes of each to-be-recognized gesture image;

[0205] The perceptual codes of each to-be-recognized gesture image are input into the gesture discriminator to sequentially recognize and obtain the gesture recognition result of each to-be-recognized gesture image.

[0206] In a feasible implementation, when the processor 801 executes the gesture recognition model training, the processor 801 is specifically configured to:

[0207] A sample training set is collected. The sample training set includes a plurality of target gesture triplets. Each target gesture triplet is composed of a first gesture image, a second gesture image, and a third gesture image. The similarity of the second gesture image to the first gesture image is greater than a first preset threshold, and the similarity of the third gesture image to the first gesture image is less than a second preset threshold. The first preset threshold is greater than the second preset threshold. Each target gesture triplet has annotation information. The annotation information includes image similarity indication information and a gesture category of each gesture image.

[0208] The sample training set is input into the initial perceptual coding network as input data to train and obtain the gesture perceptual coding network.

[0209] The sample training set is input into the gesture perceptual coding network as input data to obtain the perceptual codes of each gesture image in each target gesture triplet output by the gesture perceptual coding network.

[0210] The sample training set and the perceptual codes of each gesture image in each target gesture triplet output by the gesture perceptual coding network are input into the initial discriminator as input data to train and obtain the gesture discriminator.

[0211] In a feasible implementation, when the processor 801 executes the collection of the sample training set, the processor 801 is specifically configured to:

[0212] A plurality of sample gesture images are collected, and a plurality of candidate gesture triplets are constructed based on the plurality of sample gesture images.

[0213] The validity of each candidate gesture triplet is checked, and a candidate gesture triplet satisfying a preset condition is taken as a valid gesture triplet;

[0214] The target annotation information of each valid gesture triplet is determined according to the voting of the multiple pieces of annotation information of each valid gesture triplet;

[0215] Each valid gesture triplet and the target annotation information of each valid gesture triplet are taken as a target gesture triplet, and a sample training set is obtained.

[0216] In a feasible implementation, when the processor 801 executes the inputting of the sample training set as input data into the initial perception coding network and the training of the gesture perception coding network, the processor 801 is specifically configured to:

[0217] The sample training set is inputted as input data into the initial perception coding network, and the predicted perception codes of the gesture images in each target gesture triplet output by the initial perception coding network are obtained.

[0218] The first loss parameter of the initial perception coding network is calculated according to the predicted perception codes of the gesture images in each target gesture triplet.

[0219] The network parameters of the initial perception coding network are corrected according to the first loss parameter, and the correction is iteratively executed until the first loss parameter satisfies a third preset threshold value, the correction is stopped, and the current initial perception coding network is taken as the gesture perception coding network.

[0220] In a feasible implementation, when the processor 801 executes the calculation of the first loss parameter of the initial perception coding network according to the predicted perception codes of the gesture images in each target gesture triplet, the processor 801 is specifically configured to:

[0221] The first gesture image, the second gesture image, and the third gesture image in each target gesture triplet are respectively determined according to the image similarity indication information in the annotation information of each target gesture triplet.

[0222] The first distance corresponding to each target gesture triplet is calculated according to the predicted perception code of the first gesture image and the predicted perception code of the second gesture image in the target gesture triplet.

[0223] The second distance corresponding to each target gesture triplet is calculated according to the predicted perception code of the first gesture image and the predicted perception code of the third gesture image in the target gesture triplet.

[0224] The first loss parameter of the initial perception coding network is determined according to the first distance corresponding to each target gesture triplet, the second distance corresponding to each target gesture triplet, and a preset distance threshold value.

[0225] In an implementable implementation, when performing the following steps, the processor 801 is specifically configured to:

[0226] inputting the sample training set and the perception codes of the gesture images in each target gesture triplet as input data into the initial discriminator, and obtaining the predicted gesture categories of the gesture images in each target gesture triplet output by the initial discriminator;

[0227] calculating a second loss function of the initial discriminator according to the predicted gesture categories of the gesture images in each target gesture triplet output by the gesture discriminator and the gesture categories of the gesture images in the annotation information of each target gesture triplet;

[0228] correcting the network parameters of the initial discriminator according to the second loss parameter, and iteratively performing until the second loss parameter meets the third preset threshold or the second loss parameter meets the fourth preset threshold, stopping the correction, and taking the current initial discriminator as the gesture discriminator.

[0229] In an implementable implementation, when performing the following steps, the processor 801 is specifically configured to:

[0230] performing cross-entropy calculation according to the predicted gesture categories of the gesture images in each target gesture triplet output by the gesture discriminator, the gesture categories of the gesture images in the annotation information of each target gesture triplet, and the number of gesture images in the sample training set, to obtain the second loss function of the initial discriminator.

[0231] In an implementable implementation, when performing the following steps, the processor 801 is specifically configured to:

[0232] if the at least one frame of gesture image to be recognized includes a plurality of frames of gesture image to be recognized, using the pre-trained gesture recognition model to sequentially recognize the perception codes of the frames of gesture image to be recognized, and based on the perception codes of the frames of gesture image to be recognized, recognizing a gesture recognition result sequence of the at least one frame of gesture image to be recognized, the gesture recognition result sequence being composed of the gesture recognition results of the frames of gesture image to be recognized in sequence.

[0233] By the above manner, when the electronic device executes the gesture recognition method, the processor can acquire at least one frame of the to-be-recognized gesture image, and adopts the pre-trained gesture recognition model to recognize the perception code of each to-be-recognized gesture image, and recognizes the gesture recognition result of at least one frame of the to-be-recognized gesture image based on the perception code, and the perception code is used to represent the position information of the gesture in the preset space in the to-be-recognized gesture image. In the method, the perception code of the to-be-recognized gesture image is extracted as the feature information of the to-be-recognized gesture image by the trained gesture recognition model. Since the perception code represents the position information of the gesture in the preset space, the position information corresponding to different gestures is different, that is, each gesture has a unique corresponding perception code. Even if the gestures have slight differences, the perception codes of the gestures can be obtained respectively, so that the gesture image recognition can be accurately performed based on the perception code. By introducing the perception code as the feature information of the gesture image, the perception ability of the gesture recognition can be effectively improved, the slight differences between different gestures are recognized, and the accuracy of the gesture image recognition result is improved.

[0234] In the gesture recognition method, the perception code of the gesture image can be extracted by the gesture perception code network in the trained gesture recognition model. The gesture perception code network is trained by the constructed gesture triplets. In the training process, the distance between the perception codes of similar gestures in the continuous representation space is narrowed, and the distance between the perception codes of dissimilar gestures in the continuous representation space is widened, so that the distance between similar gestures is smaller than the distance between dissimilar gestures according to the correction principle. The trained gesture perception code network can perceive the slight differences between similar gestures, so that the accuracy of the gesture image recognition result is improved.

[0235] In the embodiments of the present application, the computer program can also execute other machine readable instructions when executed by the processor to perform the methods described in other embodiments. For specific method steps and principles, refer to the description of the embodiments, which will not be described in detail here.

[0236] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented by other manners. For example, the apparatus embodiments described above are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units or components shown or discussed can be indirect coupling or communication connection through some interfaces, apparatuses or units, and can be electrical, mechanical or other forms.

[0237] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, may be located in one place, or may be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0238] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of hardware plus software functional unit.

[0239] The integrated unit realized in the form of software functional unit can be stored in a computer readable storage medium. The software functional unit stored in a storage medium includes a plurality of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (English: processor) to execute part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (English: Read-Only Memory, for short: ROM), random access memory (English: Random Access Memory, for short: RAM), magnetic disk or optical disk and various program code storage media.

Claims

1. A gesture recognition method, characterized in that, include: Acquire at least one frame of the gesture image to be recognized; A pre-trained gesture recognition model is used to identify the perceptual encoding of each gesture image to be recognized, and the gesture recognition result of at least one frame of gesture image to be recognized is obtained based on the perceptual encoding. The perceptual encoding is used to characterize the position information of the gesture in the gesture image to be recognized in a preset space. The perceptual encoding of different gestures is unique; The gesture recognition model includes: a gesture perception coding network and a gesture discriminator; The gesture recognition model was trained in the following manner: A sample training set is collected, which includes multiple target gesture triples. Each target gesture triple consists of a first gesture image, a second gesture image, and a third gesture image. The similarity between the second gesture image and the first gesture image is greater than a first preset threshold, and the similarity between the third gesture image and the first gesture image is less than a second preset threshold. The first preset threshold is greater than the second preset threshold. Each target gesture triple has annotation information, which includes: image similarity indication information and the gesture category of each gesture image. The sample training set is used as input data to the initial perceptual coding network to train and obtain the gesture perceptual coding network. The sample training set is used as input data to the gesture perception coding network to obtain the perception coding of each gesture image in each target gesture triple output by the gesture perception coding network. The sample training set and the perceptual encoding of each gesture image in each target gesture triple output by the gesture perception coding network are used as input data into the initial discriminator to train and obtain the gesture discriminator.

2. The method according to claim 1, characterized in that, A pre-trained gesture recognition model is used to identify the perceptual encoding of each gesture image to be recognized, and the gesture recognition result of at least one frame of the gesture image to be recognized is obtained based on the perceptual encoding, including: The at least one frame of the gesture image to be recognized is input into the gesture perception coding network to extract the perception coding of each gesture image to be recognized; The perceptual encoding of each of the gesture images to be recognized is input into the gesture discriminator, and the gesture recognition results of each of the gesture images to be recognized are obtained sequentially.

3. The method according to claim 1, characterized in that, The collected sample training set includes: Multiple sample gesture images are acquired, and multiple candidate gesture triples are constructed based on the multiple sample gesture images; The validity of each candidate gesture triplet is verified, and the candidate gesture triplet that meets the preset conditions is taken as the valid gesture triplet. The target annotation information for each valid gesture triplet is determined by voting on multiple annotation information of each valid gesture triplet. The sample training set is obtained by taking each valid gesture triplet and the target annotation information of each valid gesture triplet as the target gesture triplet.

4. The method according to claim 1, characterized in that, The step of using the sample training set as input data to the initial perceptual coding network and training the gesture perceptual coding network includes: The sample training set is used as input data to the initial perceptual coding network to obtain the predicted perceptual code of each gesture image in each target gesture triplet output by the initial perceptual coding network. The first loss parameter of the initial perceptual coding network is calculated based on the predicted perceptual coding of each gesture image in each target gesture triplet. The network parameters of the initial perception coding network are corrected according to the first loss parameter, and the process is repeated iteratively until the first loss parameter meets the third preset threshold. The correction is then stopped, and the current initial perception coding network is used as the gesture perception coding network.

5. The method according to claim 4, characterized in that, The step of calculating the first loss parameter of the initial perceptual coding network based on the predicted perceptual coding of each gesture image in each target gesture triplet includes: Based on the image similarity indication information in the annotation information of each target gesture triplet, the first gesture image, the second gesture image, and the third gesture image in each target gesture triplet are determined respectively. The first distance corresponding to each target gesture triplet is calculated based on the predicted perceptual encoding of the first gesture image and the predicted perceptual encoding of the second gesture image in each target gesture triplet. The second distance corresponding to each target gesture triplet is calculated based on the predicted perceptual encoding of the first gesture image and the predicted perceptual encoding of the third gesture image in each target gesture triplet. The first loss parameter of the initial perceptual coding network is determined based on the first distance corresponding to each target gesture triplet, the second distance corresponding to each target gesture triplet, and a preset distance threshold.

6. The method according to claim 1, characterized in that, The step of inputting the perceptual encodings of each gesture image in each target gesture triplet output by the gesture perception coding network as input data into the initial discriminator to train and obtain the gesture discriminator includes: The sample training set and the perceptual encoding of each gesture image in each target gesture triplet are used as input data to the initial discriminator to obtain the predicted gesture category of each gesture image in each target gesture triplet output by the initial discriminator. The second loss function of the initial discriminator is calculated based on the predicted gesture category of each gesture image in each target gesture triplet output by the gesture discriminator and the gesture category of each gesture image in the annotation information of each target gesture triplet. The network parameters of the initial discriminator are corrected according to the second loss parameter, and the process is repeated iteratively until the second loss parameter meets the fourth preset threshold. The correction is then stopped, and the current initial discriminator is used as the gesture discriminator.

7. The method according to claim 6, characterized in that, The step of calculating the second loss function of the initial discriminator based on the predicted gesture category of each gesture image in each target gesture triplet output by the gesture discriminator, and the gesture category of each gesture image in the annotation information of each target gesture triplet, includes: Based on the predicted gesture category of each gesture image in each target gesture triple output by the gesture discriminator, the gesture category of each gesture image in the annotation information of each target gesture triple, and the number of gesture images in the sample training set, cross-entropy is calculated to obtain the second loss function of the initial discriminator.

8. The method according to claim 1, characterized in that, The step of employing a pre-trained gesture recognition model to identify the perceptual encoding of each gesture image to be recognized, and obtaining the gesture recognition result of at least one frame of the gesture image to be recognized based on the perceptual encoding, includes: If at least one frame of the gesture image to be recognized includes multiple frames of gesture images to be recognized, a pre-trained gesture recognition model is used to sequentially recognize the perceptual codes of each frame of the gesture image to be recognized, and the gesture recognition result sequence of the at least one frame of the gesture image to be recognized is recognized based on the perceptual codes of each frame of the gesture image to be recognized. The gesture recognition result sequence is composed of the gesture recognition results of each frame of the gesture image to be recognized arranged sequentially.

9. A gesture recognition device, characterized in that, include: Acquisition module, recognition module, training module; The acquisition module is used to acquire at least one frame of the gesture image to be recognized; The recognition module is used to identify the perceptual encoding of each of the gesture images to be recognized using a pre-trained gesture recognition model, and to obtain the gesture recognition result of at least one frame of the gesture image to be recognized based on the perceptual encoding. The perceptual encoding is used to characterize the position information of the gesture in the gesture image to be recognized in a preset space. The perceptual encoding of different gestures is unique; The gesture recognition model includes: a gesture perception coding network and a gesture discriminator; The training module is used to collect a sample training set, which includes multiple target gesture triples. Each target gesture triple consists of a first gesture image, a second gesture image, and a third gesture image. The similarity between the second gesture image and the first gesture image is greater than a first preset threshold, and the similarity between the third gesture image and the first gesture image is less than a second preset threshold. The first preset threshold is greater than the second preset threshold. Each target gesture triple has annotation information, which includes: image similarity indication information and the gesture category of each gesture image. The sample training set is used as input data to the initial perceptual coding network to train and obtain the gesture perceptual coding network. The sample training set is used as input data to the gesture perception coding network to obtain the perception coding of each gesture image in each target gesture triple output by the gesture perception coding network. The sample training set and the perceptual encoding of each gesture image in each target gesture triple output by the gesture perception coding network are used as input data into the initial discriminator to train and obtain the gesture discriminator.

10. An electronic device, characterized in that, include: The device includes a processor, a storage medium, and a bus, wherein the storage medium stores program instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the program instructions to perform the steps of the method as described in any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, performs the steps of the method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Gesture recognition method and device and computer readable storage medium

    CN111722717A

  • Gesture recognition method and device, computer readable storage medium and terminal equipment

    CN113536864A