Gesture recognition method, model construction method, device, and storage medium
By extracting and fusing features from the gesture recognition model, the problem of inaccurate gesture recognition in existing technologies has been solved, achieving high accuracy recognition and low memory usage under multimodal data, thus improving the user interaction experience.
Patent Information
- Application Number
- CN202210957515.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-10
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2042-08-10
AI Technical Summary
Existing gesture recognition technologies cannot accurately identify the type of user gestures, resulting in a poor interactive experience. Traditional methods are limited by the transmission distance of sensor signals and the modality of neural network models, leading to low recognition accuracy.
A gesture recognition model is adopted, including a first recognition model and a second recognition model. Through a feature extraction module, a feature fusion module and a recognition layer, the positional features and image features of the gesture are recognized. During the testing phase, it adapts to multiple modal data inputs and reduces memory usage through parameter sharing.
It improves the accuracy of gesture recognition, adapts to multiple modal data inputs, reduces memory usage, and provides multiple recognition modes to enhance the user experience.
Smart Images

Figure CN115457650B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of machine learning, in particular to a gesture recognition method, a model construction method, a device and a storage medium. BACKGROUND
[0002] Gesture interaction is an efficient and convenient interaction method, but the interaction experience of gesture interaction technology or products on the market is poor at present. The reason is that the user gesture type cannot be accurately recognized.
[0003] Traditional gesture recognition includes perceiving environmental changes based on specific sensors and giving feedback. This method can only support small-range and short-distance interaction, and is easily triggered by mistake. This gesture recognition method has been gradually abandoned.
[0004] With the development of artificial intelligence, methods of using neural network models to determine gesture types have gradually attracted attention, mainly direct and indirect methods. The direct method is to read the RGB image of the camera and input it into a pre-trained neural network model for recognition. The indirect method first detects the hand key points in the image, and then inputs the key point position information into a pre-trained neural network model for recognition. However, neither the direct method nor the indirect method can fuse the hand detection result with the image information, so the accuracy of using a neural network model to recognize the gesture type in the image is not very high. SUMMARY
[0005] The present application provides a gesture recognition method, a model construction method, a device and a storage medium, which can improve the accuracy of recognizing gesture types.
[0006] In a first aspect, the present application provides a gesture recognition method, which comprises:
[0007] obtaining a to-be-recognized image containing a user gesture, inputting the to-be-recognized image into a pre-constructed gesture recognition model, and obtaining a first recognition result;
[0008] The gesture recognition model comprises a first recognition model, which comprises a first feature extraction module, a second feature extraction module, a feature fusion module and a first recognition layer. The first feature extraction module is used to extract image features of the to-be-recognized image layer by layer to obtain shallow layer features. The feature fusion module is used to determine the position features of the user gesture according to the shallow layer features, fuse the shallow layer features with the position features, and input them into the second feature extraction module. The second feature extraction module is used for feature extraction. The first recognition layer is used to recognize the image features extracted by the second feature extraction module to obtain the first recognition result.
[0009] In a second aspect, the present application further provides a gesture recognition model construction method, which comprises:
[0010] obtaining an image sample containing a user gesture;
[0011] inputting the image sample into a gesture recognition model to train the gesture recognition model;
[0012] The gesture recognition model comprises a first recognition model, and the first recognition model comprises a first feature extraction module, a second feature extraction module, a feature fusion module and a first recognition layer. The first feature extraction module is configured to extract image features of the image sample layer by layer to obtain shallow layer features. The feature fusion module is configured to determine position features of the user gesture according to the shallow layer features, and fuse the shallow layer features and the position features and input them into the second feature extraction module. The second feature extraction module is configured to extract features. The first recognition layer is configured to recognize image features extracted by the second feature extraction module to obtain a gesture prediction result, so as to train the gesture recognition model.
[0013] In a third aspect, the present application further provides a computer device, which comprises:
[0014] a memory and a processor;
[0015] The memory is connected with the processor and is configured to store programs.
[0016] The processor is configured to realize steps of the gesture recognition method provided in any of the embodiments of the present application by running the programs stored in the memory, or realize steps of the gesture recognition model construction method provided in any of the embodiments of the present application.
[0017] In a fourth aspect, the present application further provides a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to make the processor realize steps of the gesture recognition method provided in any of the embodiments of the present application, or realize steps of the gesture recognition model construction method provided in any of the embodiments of the present application.
[0018] The gesture recognition method, the model construction method, the device and the storage medium disclosed in the present application can strengthen the hand region features by fusing the key point features of the image to be recognized and the shallow layer features, so that the recognition result is more accurate.
[0019] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort.
[0021] Figure 1 is a block schematic diagram of a first recognition model in a gesture recognition method provided by an embodiment of the present application;
[0022] Figure 2 is a structural schematic diagram of a stage module in a gesture recognition method provided by an embodiment of the present application;
[0023] Figure 3a is a structural schematic diagram of a first recognition model in a gesture recognition model provided by an embodiment of the present application;
[0024] Figure 3b is a structural schematic diagram of a second recognition model in a gesture recognition model provided by an embodiment of the present application;
[0025] Figure 4 is a structural schematic diagram of a gesture recognition model provided by an embodiment of the present application;
[0026] Figure 5 is a block schematic diagram of model recombination provided by an embodiment of the present application;
[0027] Figure 6 is a step schematic diagram of a gesture recognition method provided by an embodiment of the present application;
[0028] Figure 7 is a step schematic diagram of another gesture recognition method provided by an embodiment of the present application;
[0029] Figure 8 is a block schematic diagram of calculating a loss function in a gesture recognition model provided by an embodiment of the present application;
[0030] Figure 9 is a schematic block diagram of a computer device provided by an embodiment of the present application.
[0031] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. DETAILED DESCRIPTION
[0032] With reference to the drawings, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts are within the scope of the present application.
[0033] The flowcharts shown in the drawings are only illustrative, and do not necessarily include all the contents and operations / steps, nor are they necessarily executed in the order described. For example, some operations / steps can be further decomposed, combined or partially merged, so the actual execution order can be changed according to the actual situation.
[0034] It should be understood that the terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and the appended claims of the present application, unless otherwise clear from the context, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0035] It should be understood that, in order to clearly describe the technical solutions of the embodiments of the present application, in the embodiments of the present application, the terms "first", "second", etc. are used to distinguish the same or similar items with basically the same function and role. For example, the first identification model and the second identification model are only used to distinguish different callback functions, and do not limit the order. Those skilled in the art can understand that the terms "first", "second", etc. do not limit the quantity and execution order, and the terms "first", "second", etc. also do not necessarily mean different.
[0036] It should also be understood that the term "and / or" used in the specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.
[0037] The current gesture interaction technology or product on the market has poor interaction experience, and the reason is that the user gesture type cannot be accurately recognized, including left and right waving, up and down waving, covering the face, etc. In the prior art, the gesture recognition method can be roughly divided into two schemes: a scheme without using a neural network and a scheme using a neural network.
[0038] The scheme without using a neural network, that is, a traditional gesture recognition, mainly judges the change of the front environment through a sensor (such as a photosensitive element, a depth sensor, etc.), and feeds back the environmental change to a server or other equipment, so as to facilitate the user to view the environmental change. However, due to the limitation of the signal transmission distance of the sensor, the recognition method can only support small-range and short-distance interaction; moreover, the sensor cannot distinguish whether the instruction is issued by the hand or not, for example, the target object walks, the light is turned on, and other non-hand actions may also trigger the sensor; therefore, the recognition method may also have a false triggering condition.
[0039] The scheme using a neural network can be divided into two kinds, (1) a direct method: reading a to-be-recognized image or a video picture containing a target object, inputting the to-be-recognized image or the video picture into a pre-trained neural network model for recognition, and directly predicting a gesture type. (2) an indirect method: using a key point detection model to predict the coordinate positions of 21 key points of a hand in a picture, and judging a gesture type in the to-be-recognized image or the video picture based on a series of pre-set rules or machine learning methods.
[0040] However, the pre-trained neural network model does not fuse the hand detection result with the image information to assist the neural network model in paying attention to and recognizing the hand region in the image, which is why the gesture recognition result is not accurate enough.
[0041] In summary, the current gesture interaction has a large room for improvement in the application scene and the recognition effect.
[0042] Therefore, embodiments of the present application provide a gesture recognition method, a model construction method, a computer device and a storage medium. In the test phase, the gesture recognition model can adapt to any modal data input, can extract the position features and image features of the target object gesture from the to-be-recognized image of the target object, and fuse them, thereby improving the accuracy of gesture recognition.
[0043] Some embodiments of the present application will be described in detail below with reference to the accompanying drawings. In the case of no conflict, the embodiments described below and the features in the embodiments can be combined with each other.
[0044] Please refer to Figure 1 , Figure 1 is a block schematic diagram of a first recognition model in a gesture recognition method provided by embodiments of the present application. The gesture recognition method can be applied to a computer device.
[0045] The gesture recognition method comprises: acquiring a to-be-recognized image containing a user gesture, inputting the to-be-recognized image into a pre-constructed gesture recognition model, and obtaining a recognition result.
[0046] The gesture recognition model comprises a first recognition model, and the first recognition model comprises a first feature extraction module, a second feature extraction module, a feature fusion module and a first recognition layer. The first feature extraction module is configured to extract image features of the to-be-recognized image layer by layer to obtain shallow layer features. The feature fusion module is configured to determine position features of a user gesture according to the shallow layer features, fuse the shallow layer features and the position features, and input the fused features to the second feature extraction module. The second feature extraction module is configured to extract features. The first recognition layer is configured to recognize image features extracted by the second feature extraction module to obtain the recognition result.
[0047] In the above scheme, the gesture recognition model can recognize the position features and shallow layer features of the hand key points from the to-be-recognized image, fuse the position features and the shallow layer features of the hand key points, thereby improving the accuracy of gesture type recognition, and the recognition result is more reliable.
[0048] Further, in the embodiment of the present application, the first feature extraction module and the second feature extraction module can each comprise one or more sub-feature extraction modules connected in sequence.
[0049] The sub-feature extraction module can comprise a Stage module of a Backbone network. Referring to FIG. 2, Figure 2 Figure 2 FIG. 2 is a structural diagram of a stage module in a gesture recognition method according to an embodiment of the present application; the Stage module can comprise at least one down-sampling module, a plurality of residual modules and a plurality of TSM modules, and the down-sampling module, the plurality of residual modules and the plurality of TSM modules are connected in sequence.
[0050] In the embodiment of the present application, the Backbone network can be composed of a plurality of Stage modules, and the Backbone is used to extract video frame features, that is, to extract features of the to-be-recognized image. The down-sampling module and the residual module in each Stage module are commonly used modules in a convolutional neural network, and the down-sampling module and the residual module can be used to extract spatial features of each image frame. The TSM (Temporal Shift Module) module can be used to construct the temporal features between video frames.
[0051] It should be noted that Backbone is a general term for a feature extraction network, which can be one of ResNet, MobileNet, ShuffleNet, etc. In the embodiment of the present application, the backbone is preferably a MobileNet network.
[0052] Preferably, in the embodiment of the present application, referring to FIG. 3, Figure 3a Figure 3a is a structural diagram of a first recognition model in a gesture recognition model in a gesture recognition method provided by an embodiment of the present application. In the first feature module extraction module of the Backbone network, stage1…stage i in Figure 3a , the second feature extraction module can include stage i+1…stage5 in Figure 3b . That is, the Backbone network in the embodiment of the present application can include five stage modules, which are respectively referred to as a first stage module, a second stage module, a third stage module, a fourth stage module, and a fifth stage module, where i is a positive integer. In the first recognition model, the first stage module, the second stage module, the feature fusion module, the third stage module, the fourth stage module, and the fifth stage module can be sequentially connected.
[0053] Figure 3a In the feature fusion module, the SimDR module is used to extract the position feature of the user gesture from the shallow feature extracted from the previous stage i module, and then the fusion module is used to multiply and fuse the shallow feature extracted from the previous stage i module with the position feature to strengthen the hand region feature.
[0054] Referring to Figure 4 , Figure 4 is a structural diagram of a gesture recognition model in a gesture recognition method provided by an embodiment of the present application. Figure 4 The solid line part is the process of predicting the gesture type by the first recognition model. After the image to be recognized is input into the first recognition model in the gesture recognition model, the first stage module and the second stage module (i.e., MobileNet stage1-2 in Figure 4 ) sequentially extract shallow features from the image to be recognized; then the SimDR module detects the key point position information of the left and right hands from the shallow features, generates position features according to the position information of the key points, and the fusion module (i.e., element product in Figure 4 ) multiplies and fuses the position features with the shallow features; then the third stage module, the fourth stage module, and the fifth stage module (i.e., MobileNet stage3-5 in Figure 4 ) further extract features from the fused feature images, and output (i.e., the FC layer in the solid line part in the figure) the gesture recognition result through the first recognition layer.
[0055] It can be understood that in the above content, the position feature is actually also an image, and in the image, the positions corresponding to the left and right hands are highlighted to improve the recognition accuracy of the gesture.
[0056] The inventors of this application have discovered that while using neural network models to recognize gesture types can improve the accuracy and stability of gesture recognition, neural network models often require training data to train them and enable them to predict the type of input data during the testing phase (i.e., the usage phase). If the training data is only unimodal (e.g., RGB images), while the test data is of other modalities (e.g., near-infrared images, depth images, optical flow, etc.), the recognition performance will decrease due to the data domain mismatch problem.
[0057] A common approach is to train a neural network model using multimodal training data. However, this requires the neural network to be trained across multiple data domains, and its performance on a single modality may not be optimal during the testing phase.
[0058] For example, existing direct and indirect methods have a single recognition modality, supporting only RGB image input. If the camera switches to near-infrared input images in insufficient lighting conditions, it will not match the input of the pre-trained neural network model, so the recognition performance of the pre-trained neural network model may decrease, or it may even fail to recognize gesture types.
[0059] In view of this, in the embodiments of this application, the number of first recognition models can be expanded according to the data type of the image to be recognized, and the expanded first recognition model has the same structure and shares parameters with the original first recognition model, thereby ensuring that during the testing phase, the first recognition model can adapt to data input of any modality and achieve accurate gesture type recognition. Since the parameters of each first recognition model are shared, it will not occupy too much memory.
[0060] Specifically, see Figure 5 , Figure 6 As shown, Figure 5 This is a block diagram illustrating model reconstruction provided in an embodiment of this application. Figure 6 This is a schematic diagram illustrating the steps of a gesture recognition method provided in an embodiment of this application. In one embodiment of this application, in order to enable the first recognition model to adapt to data input of any modality, the gesture recognition method of this application further includes:
[0061] Step S1011: Obtain the image types of the multiple images to be identified.
[0062] Image type, which can be understood as image format, can be used to determine the modality of the image to be identified. Typically, the image to be identified can include multiple modalities such as RGB images, depth images, optical flow images, and near-infrared images.
[0063] Step S1012, determine the number of modalities according to the image type, and replicate the first recognition model to make the number of first recognition models equal to the number of modalities; and only one modality of the to-be-recognized image is input into each first recognition model. That is, the model structures of the first recognition models corresponding to different image types or different modalities are the same, and the model parameters are shared.
[0064] Each image type corresponds to one modality. Taking the number of modalities as four, which are RGB images, depth images, optical flow images, and near-infrared images. Then the number of first recognition models can be made equal to the number of image types by replicating the first recognition model. The RGB image, the depth image, the optical flow image, and the near-infrared image each correspond to a first recognition model, which can be denoted as the first RGB recognition model, the first depth recognition model, the first optical flow recognition model, and the first near-infrared recognition model.
[0065] Step S1013, input the to-be-recognized images of different image types into the corresponding first recognition models respectively to obtain initial recognition results output by each first recognition model.
[0066] That is, the RGB image, the depth image, the optical flow image, and the near-infrared image can be input into the first RGB recognition model, the first depth recognition model, the first optical flow recognition model, and the first near-infrared recognition model respectively to ensure that each first recognition model only includes one modality of to-be-recognized image. The first RGB recognition model, the first depth recognition model, the first optical flow recognition model, and the first near-infrared recognition model output the prediction probability y i of the gesture type. i That is, the initial result.
[0067] Step S1014, determine the recognition result of the user gesture according to a plurality of initial recognition results.
[0068] According to the prediction probability y i , the overall prediction probability y of the gesture type is obtained by weighted summation.
[0069] The specific overall prediction probability can be calculated by the following formula:
[0070]
[0071] Wherein, N is the number of modalities of the to-be-recognized image, a i represents the weight of the i-th modality, y i represents the prediction probability output by the i-th first recognition model; which can be adjusted according to actual effect. Since the model parameters are shared, the memory occupation of all parameter quantities is low.
[0072] Once the overall prediction probability y is obtained and the entire gesture prediction process is completed, only one first recognition model in the gesture recognition model can be retained to reduce memory usage.
[0073] Furthermore, in another embodiment of this application, in order to enable the first recognition model to adapt to data input of any modality, the gesture recognition model may include multiple gesture recognition models of different modalities. The gesture recognition method of this application may further include: obtaining the image type of the image to be recognized; obtaining the corresponding modality according to the image type; determining the gesture recognition model corresponding to the image to be recognized according to the correspondence between the modality and the gesture recognition model, wherein different image types correspond to different modal gesture recognition models; inputting the image to be recognized into the gesture recognition model corresponding to the image type of the image to be recognized; and repeating steps S1013 and S1014 to obtain the overall prediction probability y of the gesture type.
[0074] In summary, even in practical applications where the input image to be recognized has multiple image types or modalities, this gesture recognition model can accurately predict the gesture type. Furthermore, the first recognition models for multiple different modalities share the same structure and parameters, thus minimizing memory consumption.
[0075] It should be noted that parameter sharing can be implemented by: replicating the gesture recognition model and the same number of parameters based on the number of images to be recognized for different modalities; calculating the loss function of each gesture recognition model and averaging the results; and then updating each gesture recognition model based on the average value to achieve parameter sharing.
[0076] The inventors of this application have also discovered that when recognizing the gesture type of a target object in a video, there may be situations where it is necessary to input all video frames into the gesture recognition model for recognition.
[0077] In view of this, in the embodiments of this application, see Figure 3b , Figure 4 The dashed line portion shown in the diagram, Figure 3b This is a schematic diagram illustrating the structure of the second recognition model in a gesture recognition model provided in an embodiment of this application. The gesture recognition model may further include the second recognition model. In this case, the method may further include:
[0078] All video frames containing the target object are used as images to be recognized and input into the second recognition model for recognition, resulting in a second recognition result.
[0079] The second recognition model may include a first feature extraction module, a third feature extraction module, and a feature enhancement module (i.e., Figure 3bThe feature enhancement module can be used to enhance the position feature of the user gesture in the shallow feature of the first feature extraction module, and input the enhanced shallow feature to the third feature extraction module; the third feature extraction module is used for feature extraction; and the second recognition layer is used for recognizing the image feature extracted by the third feature extraction module to obtain a second recognition result.
[0080] It should be emphasized that the first recognition model and the second recognition model can share the first feature extraction module.
[0081] In an embodiment of the present application, the third feature extraction module can be the same as the second feature extraction module. The feature enhancement module can provide spatio-temporal attention enhanced features, so that the input image to be recognized can include all video frames, and the second recognition model can have better recognition effect under the action of the feature enhancement module.
[0082] The present inventors have also found that in actual use of the gesture recognition model, in most cases, it is necessary to accurately recognize the gesture type of the target object using the gesture recognition model, but there can be some cases where it is not necessary to accurately recognize the gesture type of the target object.
[0083] Therefore, in order to improve the user experience, in an embodiment of the present application, the following can also be included:
[0084] The recognition mode is determined according to the recognition mode.
[0085] The recognition mode can include a first recognition mode, a second recognition mode and a third recognition mode, the gesture recognition model corresponding to the first recognition mode includes the first recognition model or the second recognition model, the gesture recognition model corresponding to the second recognition mode includes the first recognition model and the second recognition model, and the gesture recognition model corresponding to the third recognition mode includes a plurality of different modal first recognition models.
[0086] When the user selects the first mode, the image to be recognized containing the user gesture is obtained, and the image to be recognized is input into the pre-constructed gesture recognition model to obtain a recognition result, which can include: obtaining an image to be recognized containing a user gesture, and inputting the image to be recognized into the first recognition model or the second recognition model; or inputting the image to be recognized for all video frames into the second recognition model at one time to recognize the user gesture type.
[0087] When the user selects the second mode, the image to be recognized containing the user gesture is acquired, the image to be recognized is input into the pre-constructed gesture recognition model, and a recognition result is obtained, which can include: the image to be recognized can be input into the first recognition network frame by frame; or the image to be recognized for all video frames can be input into the second recognition model at one time to identify the user gesture type.
[0088] It should be noted that the first recognition model and the second recognition model can be used to identify the gesture type, but the input manner of the image to be recognized is different. Moreover, the SimDR module in the first recognition model can be used to predict the positions of the left and right hands in the horizontal and vertical directions, and the fusion module can multiply and fuse the key point features and the shallow features to strengthen the hand region features. The feature enhancement module of the second recognition model focuses on enhancing the region and time features corresponding to the gesture in the image to be recognized.
[0089] When the user selects the third mode, the image to be recognized containing the user gesture is acquired, the image to be recognized is input into the pre-constructed gesture recognition model, and a recognition result is obtained, which can include: the image to be recognized containing the user gesture is acquired, the image to be recognized is processed in a mode, and the processed image to be recognized is input into the first recognition model corresponding to the mode. The combination of multiple first recognition models can make the prediction result of the gesture type more accurate. Of course, in the case of the third mode, if the image to be recognized has only one mode, the first recognition model corresponding to the mode can be called for recognition according to the mode.
[0090] Among them, the mode processing can be to judge the mode of the current image to be recognized through the type or format of the image to be recognized.
[0091] Specifically, in the third mode, the gesture recognition model can include multiple first recognition models expanded, and the number of first recognition models is the same as the number of modes. When recognition is needed, the image to be recognized is input into the corresponding first recognition model.
[0092] It should be emphasized that when the image to be recognized is saved offline in the form of data, the user can separately call the first recognition model or the second recognition model.
[0093] When the image to be recognized is online data, the user can call the first recognition model for recognition.
[0094] When online recognition is needed and the image to be recognized has multiple modalities, the user can select a third recognition mode. The third recognition mode has multiple gesture recognition models corresponding to the modalities. At this time, the server can determine the corresponding modality according to the image type of the image to be recognized. Then the image to be recognized of different modalities is input into the corresponding gesture recognition model to realize combined recognition of multiple models and obtain more accurate recognition results.
[0095] Through the above scheme, the gesture recognition method provided by the embodiment of the present application can identify the position feature and the shallow feature of the hand key point from the image to be recognized, and multiply and fuse the position feature and the shallow feature of the hand key point, thereby improving the accuracy of gesture type recognition. Meanwhile, the first recognition model in the gesture recognition model can be expanded in quantity, and the parameters of each first recognition model are shared after expansion, which can input the image to be recognized of multiple modalities into the gesture recognition model and improve the accuracy, and also does not occupy too much memory.
[0096] The embodiment of the present application also provides a method for constructing a gesture recognition model. Referring to Figure 7 Figure 7 is a step schematic diagram of another gesture recognition method provided by the embodiment of the present application; and the construction method can include steps S201-S202.
[0097] It should be emphasized that the following content is described for the first recognition model and the second recognition model in the training stage; and the first feature extraction module, the second feature extraction module, the feature fusion module and the first recognition layer included in the first recognition model are also modules in the training stage.
[0098] Step S201, acquiring an image sample containing a user gesture.
[0099] Step S202, inputting the image sample into a gesture recognition model to train the gesture recognition model.
[0100] The gesture recognition model includes a first recognition model, and the first recognition model includes a first feature extraction module, a second feature extraction module, a feature fusion module and a first recognition layer; the first feature extraction module is used to extract image features of the image sample layer by layer to obtain shallow features; the feature fusion module is used to determine the position feature of the user gesture according to the shallow features, fuse the shallow features with the position feature and input them into the second feature extraction module; the second feature extraction module is used for feature extraction; and the first recognition layer is used to recognize the image features extracted by the second feature extraction module to obtain the gesture prediction result, so as to train the gesture recognition model.
[0101] Further, the gesture recognition model can further include a second recognition model, and a training step of the second recognition model is similar to steps S201-S202, and thus will not be described in detail.
[0102] Further, referring to Figure 8 as shown, Figure 8 is a block diagram for calculating a loss function of a gesture recognition model provided by an embodiment of the present application. In the embodiment of the present application, the gesture recognition model can include a plurality of first recognition models corresponding to different modalities, and a second recognition model, and the image sample can include sub-image samples of different modalities. The method for constructing the gesture recognition model can further include:
[0103] inputting the sub-image sample corresponding to the RGB modality into the second recognition model to obtain a first loss function of a second recognition model prediction result; inputting the sub-image sample into the first recognition model corresponding to the modality of the sub-image sample to obtain a second loss function of each first recognition model prediction result; wherein the number of first recognition models is the same as the number of second recognition models; obtaining a third loss function of adjacent two prediction results according to the loss function of the second recognition model and the loss function of the first recognition model; updating the gesture recognition model through the first loss function, the second loss function and the third loss function, and retaining one first recognition model.
[0104] For the convenience of understanding, taking an image sample including four modalities of RGB image samples, depth image samples, optical flow image samples and near-infrared image samples as an example; in the training stage, the first recognition model can be copied to make the number of first recognition models also four, and the parameters of each first recognition model are the same. The training steps are described as follows:
[0105] 1) pre-process the image sample, respectively:
[0106] 11) frame a series of RGB image samples from a video containing a target object, and pre-process the RGB image samples, which can include: scaling the RGB image samples to a certain size (the size can be determined according to actual conditions), and then subtracting the mean value and dividing by the standard deviation; the mean value and the standard deviation can be obtained by statistics on the data in advance.
[0107] 12) the optical flow modality can be calculated from the image samples of the RGB modality. The specific optical flow modality calculation method is prior art, and thus will not be described in detail. The pre-processing method of the optical flow modality can include scaling the image samples of the optical flow modality to a certain size (the size can be determined according to actual conditions), and then subtracting the mean value and dividing by the standard deviation (the mean value and the standard deviation can be obtained by statistics on the data), and the processing result is used as the input of the Flow online recognition model.
[0108] 13) The image samples of the near-infrared modal, the image samples of the depth image can be obtained by a specific camera, and the preprocessing thereof is similar to that of the RGB image samples, which will not be elaborated here.
[0109] 2) The RGB image sample is input into the second recognition model, and the second recognition model is trained; the RGB image sample, the depth image sample, the optical flow image sample, and the near-infrared image sample are respectively input into four first recognition models, and the first recognition models are trained; and the loss functions Loss1-Loss5 of each prediction result are calculated. In order to ensure that the prediction results of the first recognition models corresponding to each modal are consistent, the loss functions Loss6-Loss9 of the prediction results of adjacent two recognition models can also be calculated.
[0110] Specifically, the loss function of the second recognition model is Loss1, and the loss functions of the first recognition networks corresponding to the RGB, optical flow, depth, and near-infrared four modes are Loss2, Loss3, Loss4, and Loss5 respectively. Then, the loss function of the prediction results between the adjacent second recognition model and the first recognition model corresponding to the RGB modal is Loss6; the loss function of the prediction results between the adjacent first recognition model corresponding to the RGB modal and the first recognition model corresponding to the depth modal is Loss7; the loss function of the prediction results between the first recognition model corresponding to the depth modal and the first recognition model corresponding to the optical flow modal is Loss8; and the loss function of the prediction results between the first recognition model corresponding to the optical flow modal and the first recognition model corresponding to the near-infrared modal is Loss9.
[0111] Finally, all the loss functions are summed up, and the formula of the total loss function Loss of the model is as follows:
[0112]
[0113] wherein, λ i is the weight representing the i-th loss function.
[0114] In summary, the gesture recognition model is calculated by Loss. However, considering that the structures of all the first recognition neural network models are consistent and the parameters are shared, the parameters of all the first recognition models are the same after the training is completed, and only one first recognition model parameter needs to be retained, which greatly reduces the memory occupation.
[0115] In summary, the gesture recognition model construction method provided in the embodiments of the present application trains multiple first recognition models through multiple modal image samples. Because the parameters of each first recognition model are shared, less memory can be occupied. Meanwhile, only one first recognition model can be finally retained, further reducing the occupied memory. The trained second recognition model can be used to input all video frames at one time. The second recognition model and the first recognition model cooperate with each other, and can provide multiple recognition modes for user selection.
[0116] Please refer to Figure 9 , Figure 9 is a schematic block diagram of a computer device provided in the embodiments of the present application. As shown in Figure 9 , the computer device 600 includes one or more processors 601 and a memory 602, and the processor 601 and the memory 602 are connected through a bus, such as an I2C (Inter-integrated Circuit) bus.
[0117] The one or more processors 601 work alone or jointly to execute the steps of the gesture recognition method provided in the above embodiments, or to execute the steps of the gesture recognition model construction method provided in the above embodiments.
[0118] Specifically, the processor 601 can be a micro-control unit (MCU), a central processing unit (CPU) or a digital signal processor (DSP), etc.
[0119] Specifically, the memory 602 can be a flash chip, a read-only memory (ROM) disk, an optical disk, a U disk or a mobile hard disk, etc.
[0120] The processor 601 is configured to run a computer program stored in the memory 602, and to implement the steps of the gesture recognition method provided in the above embodiments when executing the computer program.
[0121] For example, the processor 601 is configured to run a computer program stored in the memory 602, and to implement the following steps when executing the computer program:
[0122] Obtain an image to be recognized containing a user gesture, and input the image to be recognized into a pre-constructed gesture recognition model to obtain a first recognition result.
[0123] The gesture recognition model comprises a first recognition model, and the first recognition model comprises a first feature extraction module, a second feature extraction module, a feature fusion module and a first recognition layer; the first feature extraction module is configured to extract image features of the to-be-recognized image layer by layer to obtain shallow layer features; the feature fusion module is configured to determine position features of a user gesture according to the shallow layer features, fuse the shallow layer features and the position features, and input the fused features to the second feature extraction module; the second feature extraction module is configured to extract features; and the first recognition layer is configured to recognize image features extracted by the second feature extraction module to obtain the first recognition result.
[0124] In some embodiments, the first feature extraction module and the second feature extraction module comprise one or more sub-feature extraction modules connected in sequence; wherein the sub-feature extraction module comprises a Stage module of a Backbone network, the Stage module comprises at least one down-sampling module and a plurality of residual modules, and a plurality of TSM modules, and the down-sampling module, the plurality of residual modules and the plurality of TSM modules are connected in sequence.
[0125] In some embodiments, the feature fusion module comprises a SimDR module and a fusion module, the SimDR module is configured to extract position features of a user gesture from the shallow layer features, and the fusion module is configured to fuse the shallow layer features and the position features.
[0126] In some embodiments, before the processor implements the inputting of the to-be-recognized image into the pre-constructed gesture recognition model, the processor is further configured to implement:
[0127] acquiring an image type of the to-be-recognized image, determining a number of modalities according to the image type, and duplicating the first recognition model to make the number of the first recognition models equal to the number of modalities.
[0128] In some embodiments, each of the first recognition models has the same model structure and shares model parameters.
[0129] In some embodiments, before the processor implements the inputting of the to-be-recognized image into the pre-constructed gesture recognition model, the processor is further configured to implement:
[0130] acquiring a plurality of to-be-recognized images containing a user gesture, the modalities of the plurality of to-be-recognized images being different; inputting to-be-recognized images of the same modality into the first recognition model in sequence to obtain initial recognition results output by each of the first recognition models; wherein only to-be-recognized images of one modality are input into each of the first recognition models; and determining a recognition result of the user gesture according to the plurality of initial recognition results.
[0131] In some embodiments, the gesture recognition model further comprises a second recognition model; and the processor, when implementing the gesture recognition method, is further configured to implement:
[0132] input all video frames comprising the target object video as to-be-recognized images to the second recognition model for recognition to obtain a second recognition result;
[0133] The second recognition model comprises a first feature extraction module, a third feature extraction module, a feature enhancement module and a second recognition layer; the feature enhancement module is configured to enhance a position feature of a user gesture in a shallow feature of the first feature extraction module, and input the enhanced shallow feature to the third feature extraction module; the third feature extraction module is configured to perform feature extraction; and the second recognition layer is configured to recognize image features extracted by the third feature extraction module to obtain the second recognition result.
[0134] In some embodiments, the first recognition model and the second recognition model share the first feature extraction module.
[0135] In some embodiments, before the step of inputting the to-be-recognized image to the pre-constructed gesture recognition model, the processor is further configured to implement:
[0136] obtain a user-selected recognition mode, and determine a corresponding gesture recognition model according to the recognition mode; wherein the recognition mode comprises a first recognition mode, a second recognition mode and a third recognition mode, the gesture recognition model corresponding to the first recognition mode comprises the first recognition model or the second recognition model, the gesture recognition model corresponding to the second recognition mode comprises the first recognition model and the second recognition model, and the gesture recognition model corresponding to the third recognition mode comprises a plurality of first recognition models of different modalities.
[0137] For example, the processor 601 is configured to run a computer program stored in the memory 602, and when the computer program is executed, the following is implemented:
[0138] An image sample containing a user gesture is acquired; the image sample is input into a gesture recognition model to train the gesture recognition model. The gesture recognition model includes a first recognition model, and the first recognition model includes a first feature extraction module, a second feature extraction module, a feature fusion module, and a first recognition layer. The first feature extraction module is configured to extract image features of the image sample layer by layer to obtain shallow layer features. The feature fusion module is configured to determine a position feature of the user gesture according to the shallow layer features, fuse the shallow layer features and the position feature, and input the fused features into the second feature extraction module. The second feature extraction module is configured to extract features. The first recognition layer is configured to recognize the image features extracted by the second feature extraction module to obtain a gesture prediction result, so as to train the gesture recognition model.
[0139] The embodiment of the present application further provides a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to make the processor implement the steps of the gesture recognition method provided by the above embodiment, or implement the steps of the construction method of the gesture recognition model provided by the above embodiment.
[0140] The computer readable storage medium can be an internal storage unit of the computer device, for example, a hard disk or a memory of the terminal device. The computer readable storage medium can also be an external storage device of the terminal device, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.
[0141] The above merely describes the specific implementation of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A gesture recognition method, characterized by, The method comprises the following steps: acquiring an image to be recognized containing a user gesture, and acquiring image types of a plurality of the image to be recognized; determining a number of modalities according to the image types, and duplicating a first recognition model in a pre-constructed gesture recognition model, so that the number of the first recognition model is equal to the number of the modalities; inputting the image to be recognized into the gesture recognition model to obtain a first recognition result; wherein the first recognition model comprises a first feature extraction module, a second feature extraction module, a SimDR module, a fusion module, and a first recognition layer; the first feature extraction module is configured to extract image features of the image to be recognized layer by layer to obtain shallow layer features; the SimDR module is configured to extract position features of the user gesture from the shallow layer features; the fusion module is configured to fuse the shallow layer features and the position features and input them into the second feature extraction module; the second feature extraction module is configured to extract features; and the first recognition layer is configured to recognize image features extracted by the second feature extraction module to obtain the first recognition result.
2. The gesture recognition method of claim 1, wherein, The first feature extraction module and the second feature extraction module comprise one or more sub-feature extraction modules connected in sequence; wherein the sub-feature extraction module comprises a Stage module of a Backbone network, the Stage module comprises at least one down-sampling module, a plurality of residual modules, and a plurality of time transfer modules, and the down-sampling module, the plurality of residual modules, and the plurality of time transfer modules are connected in sequence.
3. The gesture recognition method of claim 1, wherein, The model structures of each of the first recognition models are the same, and the model parameters are shared.
4. The gesture recognition method of claim 1, wherein, The step of acquiring the image to be recognized containing the user gesture and inputting the image to be recognized into the pre-constructed gesture recognition model comprises the following steps: acquiring a plurality of images to be recognized containing a user gesture, the modalities of the plurality of images to be recognized being different; inputting the images to be recognized of the same modality into the same first recognition model in sequence to obtain initial recognition results output by each of the first recognition models; wherein one modality of the image to be recognized is input into each of the first recognition models; and determining a recognition result of the user gesture according to the plurality of initial recognition results.
5. The gesture recognition method according to any one of claims 1-4, characterized in that, The gesture recognition model further comprises a second recognition model; and the method comprises the following steps: inputting all video frames of a target object video as images to be recognized into the second recognition model to obtain a second recognition result; wherein the second recognition model comprises a first feature extraction module, a third feature extraction module, a feature enhancement module, and a second recognition layer; the feature enhancement module is configured to enhance position features of a user gesture in shallow layer features of the first feature extraction module, and input the enhanced shallow layer features into the third feature extraction module; the third feature extraction module is configured to extract features; and the second recognition layer is configured to recognize image features extracted by the third feature extraction module to obtain the second recognition result.
6. The gesture recognition method of claim 5, wherein, The first feature extraction module is shared by the first recognition model and the second recognition model.
7. The gesture recognition method of claim 5, wherein, Before the step of inputting the image to be recognized into the pre-constructed gesture recognition model, the method further comprises: obtaining a user-selected recognition mode, and determining a corresponding gesture recognition model according to the recognition mode; wherein the recognition mode comprises a first recognition mode, a second recognition mode and a third recognition mode, the gesture recognition model corresponding to the first recognition mode comprises a first recognition model or a second recognition model, the gesture recognition model corresponding to the second recognition mode comprises the first recognition model and the second recognition model, and the gesture recognition model corresponding to the third recognition mode comprises a plurality of first recognition models of different modalities. 8.A method for constructing a gesture recognition model, comprising: The construction method comprises: obtaining image samples containing user gestures, and obtaining image types of a plurality of the image samples; determining a number of modalities according to the image types, and copying a first recognition model in a gesture recognition model, so that the number of the first recognition models is equal to the number of the modalities; inputting the image samples into the gesture recognition model to train the gesture recognition model; wherein the first recognition model comprises a first feature extraction module, a second feature extraction module, a SimDR module, a fusion module and a first recognition layer; the first feature extraction module is used to extract image features of the image samples layer by layer to obtain shallow layer features; the SimDR module is used to extract position features of user gestures from the shallow layer features; the fusion module is used to fuse the shallow layer features and the position features and input them into the second feature extraction module; the second feature extraction module is used for feature extraction; and the first recognition layer is used to recognize image features extracted by the second feature extraction module to obtain a gesture prediction result, so as to train the gesture recognition model.
9. The construction method of claim 8, wherein, The gesture recognition model comprises a first recognition model and a second recognition model; the image samples comprise sub-image samples of different modalities; and the method further comprises: inputting sub-image samples corresponding to an RGB modality into the second recognition model to obtain a first loss function of a second recognition model prediction result; inputting the sub-image samples into first recognition models corresponding to the modalities of the sub-image samples to obtain a second loss function of a prediction result of each first recognition model; wherein the number of the first recognition models is the same as that of the second recognition models; obtaining a third loss function of adjacent two prediction results according to the loss function of the second recognition model and the loss function of the first recognition model; updating the gesture recognition model through the first loss function, the second loss function and the third loss function, and retaining one first recognition model.
10. The construction method of claim 9, wherein, After the step of obtaining image samples containing user gestures, the method further comprises: copying the first recognition models according to the number of the image samples of different modalities, so that the number of the first recognition models is equal to the number of the image samples of different modalities, and the parameters of the first recognition models are shared.
11. A computer device, comprising: The computer device comprises: a memory and a processor; wherein the memory is connected with the processor and is used to store programs. The processor is configured to implement the steps of the gesture recognition method of any one of claims 1-7 by running a program stored in the memory.
12. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program which, when executed by a processor, causes the processor to implement the steps of the gesture recognition method of any one of claims 1-7.
Citation Information
Patent Citations
Gesture recognition method and device, electronic equipment, readable storage medium and chip
CN112699849A
Gesture recognition method and device, storage medium and computer program product
CN113537169A