Method, device, equipment and storage medium for gesture recognition
Through the preprocessing and correction processing of multiple data sets, the gesture recognition model is optimized, and the accuracy problem of the prior art when identifying variable or micro gestures is solved, achieving higher recognition accuracy and reliability.
Patent Information
- Application Number
- CN202510186011.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-20
AI Technical Summary
Existing deep learning-based gesture recognition technologies are difficult to accurately identify variable or tiny gestures under environmental changes, hand differences and occlusion.
By acquiring multiple data sets (including public data, 3D generated data and real data), pre-processing and correction processing, the comprehensiveness and quality of the training data are improved, thereby optimizing the accuracy and robustness of the gesture recognition model.
Significantly improves the accuracy and reliability of gesture recognition, especially when identifying variable or tiny gestures.
Smart Images

Figure CN119672767B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition technology, and in particular to a method, device, equipment and storage medium for gesture recognition. Background Art
[0002] With the widespread application of gesture recognition, people's requirements for recognition accuracy are also increasing. Although gesture recognition based on deep learning has made great progress, due to environmental changes, differences in hands, occlusion of the hands themselves, etc., the existing gesture recognition based on deep learning still faces great challenges, especially for some changeable or small gestures that are difficult to detect and recognize. Summary of the invention
[0003] In view of the shortcomings of existing gesture recognition, a method, device, equipment and storage medium for gesture recognition are provided, which can improve the comprehensiveness of training data, thereby optimizing the accuracy and robustness of the gesture recognition model, and effectively improve the accuracy of gesture recognition.
[0004] In order to achieve the above-mentioned purpose, in a first aspect, a method for gesture recognition is provided, comprising: obtaining a plurality of training data on gesture recognition, the training data comprising a plurality of data sets, each data set comprising a plurality of images; preprocessing the training data according to the types of the plurality of data sets; performing correction processing on the plurality of data sets after preprocessing respectively to obtain correction results, and obtaining correction data through the correction results; training a gesture recognition model through the correction data, and using the trained gesture recognition model for gesture recognition.
[0005] In some optional embodiments, the multiple data sets include public data, 3D generated data and real data, wherein the 3D generated data and the real data include images without hands.
[0006] In some optional embodiments, the preprocessing of the training data according to the types of the multiple data sets includes: in response to the data set being the real data, using an open source hand detection model to mark the position of the hand in the image of the real data to preprocess the training data.
[0007] In some optional embodiments, the use of an open source hand detection model to mark the hand position in the image of the real data to preprocess the training data includes: marking the hand position in the image of the real data using multiple open source hand detection models to construct a hand position sample; determining the hand position marked in each open source hand detection model for each image in the hand position sample, and discarding images with inconsistent hand position marking results to complete the screening of the real data in the preprocessing process.
[0008] In some optional embodiments, the preprocessing of the training data according to the types of the multiple data sets includes: in response to the types of the multiple data sets being public data or real data, determining the camera intrinsic parameters and distortion parameters corresponding to each image in the data set; transforming the coordinates of the hand position of each image according to the camera intrinsic parameters and distortion parameters, wherein the hand position of each image is determined by an open source hand detection model; and cropping the image according to the hand position in the transformed image.
[0009] In some optional implementations, the 3D generated data and the real data include images that do not contain hands as negative samples.
[0010] In some optional embodiments, the multiple pre-processed data sets are respectively corrected to obtain correction results, and correction data is obtained through the correction results, including: detecting and correcting the hand positions of the images in each data set through multiple open source hand detection models to obtain correction results; determining the labeling accuracy of each data set and the accuracy of each open source model; and performing weighted average calculation on the correction results according to the labeling accuracy of each data set and the accuracy of each open source model to obtain correction data.
[0011] In some optional embodiments, the training of the gesture recognition model using the correction data further includes: after the convergence of the gesture recognition model reaches a preset requirement after training, the correction data is subjected to random blurring and noise addition processing for use in training the gesture recognition model.
[0012] In some optional implementations, the training of the gesture recognition model using the correction data further includes: inputting an image of the correction data into the gesture recognition model, and outputting a hand position and a hand confidence.
[0013] In some optional embodiments, using the trained gesture recognition model for gesture recognition also includes: converting the gesture recognition model into a running format of the device, recognizing an image captured by the device, and when the hand confidence is greater than a first threshold, performing optical flow tracking on the hand area of the image.
[0014] In a second aspect, a gesture recognition device is provided, comprising: a first module, used to obtain multiple training data about gesture recognition, the training data comprising multiple data sets, each data set comprising multiple images; a second module, used to preprocess the training data according to the types of the multiple data sets; a third module, used to perform correction processing on the multiple data sets after preprocessing, respectively, to obtain correction processing, and obtain correction data through the correction results; a fourth module, used to train a gesture recognition model through the correction data, and use the trained gesture recognition model for gesture recognition.
[0015] According to a third aspect, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described above when executing the computer program.
[0016] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the method as described above is implemented.
[0017] A gesture recognition method of the present application can effectively improve the quantity and quality of training data by preprocessing and correcting training data, further enhance the reliability and practicality of training data, so that the recognition effect of the gesture recognition model can be greatly enhanced when training the model, optimize the gesture recognition model required for machine learning, greatly improve the accuracy and reliability of gesture recognition, and especially can carry out targeted enhancement of the training data of various variable or small gestures, effectively improving the recognition ability of the corresponding gestures. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0019] Figure 1 is a flowchart of the method for gesture recognition described in the embodiment;
[0020] Figure 2 is a schematic diagram of the structure of the device for gesture recognition described in the embodiment;
[0021] Figure 3 It is a schematic diagram of the structure of the electronic device described in the embodiment. DETAILED DESCRIPTION
[0022] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0023] Gesture recognition technology mainly relies on computer vision, machine learning and image processing technology. It captures the user's hand images or movements through cameras or other sensors, and then uses algorithms to extract features and classify and identify them, ultimately achieving understanding and response to user gestures. At present, gesture recognition technology has made significant progress, with recognition accuracy and stability continuously improving, and application scenarios becoming increasingly rich, especially in the application of gesture recognition on virtual glasses devices.
[0024] The present application is described below with reference to specific embodiments in conjunction with the accompanying drawings.
[0025] like Figure 1 As shown, a method for gesture recognition includes the following steps:
[0026] Step 101 : Acquire a plurality of training data on gesture recognition, where the training data includes a plurality of data sets, and each data set includes a plurality of images.
[0027] The multiple training data may be jointly constituted by public data, 3D generated data, and real device data or used in selective combination, which is not limited here. In some optional embodiments, the public data, 3D generated data, and real device data are all used as data sources of training data, where:
[0028] Public data, specifically public data of the required scenarios. For example, for scenarios for eyewear devices, you can collect data sets that include the lens angle of view, the distance between the palm and the lens, and the angle of view of the hand seen by the camera in the glasses, which is similar to the lens distance, such as the InterHand2.6M dataset, epic-kichens dataset, 3DPose dataset, etc.
[0029] 3D data generation, specifically, uses 3D modeling software such as Maya, 3ds Max, Blender, etc. to create 3D hand models of various shapes, and applies 3D texture mapping to improve the simulation effect of the 3D hand model. Then, use public hand joint motion sequence data, such as the Dexter 1 dataset, to drive the 3D hand model to simulate various gestures. After the joint-driven 3D hand model is realized, the 3D hand model is projected onto a 2D plane according to the camera intrinsic parameters, and finally fused with a large number of real images that do not contain hands. The rectangular position and joint position information of the hand are extracted from the fused image to generate the final hand sample dataset.
[0030] Real data, specifically, uses the target device to collect a large number of samples with and without hands, including hand images under different lighting, backgrounds, and gestures, to ensure that the gesture recognition model can adapt to various actual situations. For the data set formed by collecting a large number of samples, open source gesture recognition models, such as the YOLOv8 (You Only Look Once version 8) model, can be used to perform hand detection in advance, which can provide a preliminary estimate of the hand position and improve the effectiveness of the data set for subsequent gesture recognition model training.
[0031] Step 102: pre-process the training data according to the types of various data sets.
[0032] In some optional embodiments, in response to the data set being real data, an open source hand detection model is used to mark the position of the hand in the image of the real data to perform preprocessing of the training data.
[0033] Furthermore, the hand positions in the images of real data can be marked separately through multiple open source hand detection models to construct hand position samples; the hand positions marked in each open source hand detection model for each image in the hand position sample are determined, and images with inconsistent hand position marking results are discarded to complete the screening of real data in the preprocessing process.
[0034] Specifically, multiple open source gesture recognition models are used to perform hand detection, and a multi-model voting method is adopted to combine the detection results of multiple open source gesture recognition models to determine the position of the hand. For example, hard voting or soft voting is selected as the voting strategy. Hard voting determines the final result based on the judgment of the majority model, while soft voting is a weighted average based on the probability of the model output. The use of multiple gesture recognition models and corresponding judgment methods can effectively improve the quality of the data set.
[0035] Furthermore, samples with inconsistent voting results from multiple models are discarded, thereby ensuring the quality of the dataset and the effectiveness of model training.
[0036] In some optional embodiments, ensuring the diversity of hand positions when collecting data can effectively prevent overfitting of positions, which is particularly suitable for scenarios where the hand is in the central area of the field of vision, such as glasses scenarios.
[0037] In some optional embodiments, the 3D generated data and the real data are supplemented with images that do not contain hands, and the images that do not contain hands are further used as negative samples, which helps the model learn the characteristics of "no hands" and thus improves the model's discrimination ability and robustness.
[0038] In addition, data enhancement techniques, such as mosaic enhancement and random cropping, can be used to further increase the diversity and quantity of samples and reduce the demand for training data.
[0039] In some optional embodiments, in response to the types of the multiple data sets being public data or real data, the camera intrinsic parameters and distortion parameters corresponding to each image in the data set are determined; the hand position of each image is transformed into coordinates according to the camera intrinsic parameters and distortion parameters, wherein the hand position of each image is determined by an open source hand detection model; and the image is cropped according to the hand position in the transformed image.
[0040] Specifically, the images of the training data can be subjected to internal reference distortion normalization processing and quality normalization processing.
[0041] Image intrinsic distortion normalization processing is for data sets from multiple different sources, such as data sets formed by collecting real data, or related data sets based on public data collection. The fisheye image of the data set can be stretched into a plane image according to the corresponding camera intrinsic parameters and distortion parameters. The size of the plane image is the same as the size of the input image during model training. For example, a 160*160 image size can be selected, and then the coordinates of the hand position are transformed accordingly, and the transformed position is cropped according to the new image boundary.
[0042] Image quality normalization processing, specifically, random blur processing and noise addition processing are performed on multiple data sets. For example, Gaussian blur, motion blur, Gaussian noise, and salt and pepper noise can be selected. Samples can also be randomly rotated and perspective transformed to increase the richness of samples. The data set that has undergone image quality normalization processing can be used for gesture model training to improve the recognition ability of the gesture recognition model, especially for cameras with poor imaging quality on devices, which effectively improves the recognition of their captured images.
[0043] Step 103, respectively calibrate the various pre-processed data sets to obtain calibration results, and obtain calibration data through the calibration results.
[0044] In some optional embodiments, the hand positions of the images in each data set are detected and corrected using multiple open source hand detection models to obtain correction results; the labeling accuracy of each data set and the accuracy of each open source model are determined; and the correction results are weighted averaged according to the labeling accuracy of each data set and the accuracy of each open source model to obtain corrected data.
[0045] Since datasets from multiple different sources may have different annotation standards and accuracies, these differences may be due to the subjectivity of the annotators, the differences in annotation tools, or the differences in the dataset collection environment. Therefore, an open source model of gesture recognition is used for unified detection and correction to improve the quality of training data. At the same time, in order to overcome the limitations of a single open source model, multiple open source models of gesture recognition can be used to re-detect samples, making full use of the respective advantages of different open source models in hand detection and key point positioning.
[0046] Furthermore, after obtaining multiple detection results based on multiple open source models of gesture recognition, correction is performed using the annotation results generated by the circumscribed rectangle of the annotated joint positions, that is, for each detected hand, a circumscribed rectangle can be calculated, which can contain all the annotated joint positions. By comparing the circumscribed rectangles generated by different models, the hand position can be corrected, and finally the required correction data can be obtained, that is, the correction data is the processing result after the training data is re-detected by multiple open source models and the multiple detection results are compared; for example, if the hand position detected by one model is significantly different from that of another model, a more accurate hand position can be determined by calculating the intersection or average position of these rectangles.
[0047] After obtaining the correction results based on the above correction processing, the correction results are finally weighted averaged according to the accuracy of multiple data sets and the accuracy of multiple open source models.
[0048] That is, after obtaining the correction results, the correction data can be weighted averaged according to the accuracy of multiple data sets and multiple open source models. That is, the accuracy of each data set and open source model may be different. These accuracies can be evaluated, and a weight can be assigned to each data set and open source model. The hand detection results can be weighted averaged according to these weights to obtain a comprehensive and more accurate hand position estimate to further optimize the correction data. For example, if one open source model shows an accuracy of 80% in an independent test, while another model has an accuracy of only 60%, then the detection result of the former will be given a greater weight when the correction data is weighted averaged.
[0049] Step 104: train a gesture recognition model using the correction data, and use the trained gesture recognition model for gesture recognition.
[0050] In some optional embodiments, when the correction data is used for training a gesture recognition model, the following steps are also included: inputting an image of the correction data to the gesture recognition model, and outputting the position of the hand and the confidence of the hand.
[0051] Model training requires the definition of a gesture recognition model. The input of the gesture recognition model is a grayscale image, whose size can be 160*160. The output is the position of the hand and the confidence of the hand. It can usually be set to output the position of two hands and the confidence of each hand. The output position information includes center_x, center_y, width, and height, which respectively represent the x-coordinate, y-coordinate, width, and height of the center point of the rectangular bounding box of the hand; the output confidence represents the confidence level of the gesture recognition model in its detection results, which is used to quantify the certainty of the model in its prediction results. For example, in some application scenarios of glasses, at most two hands appear in the field of view, so the confidence of the two hands is defined separately. If two hands appear in the training data, the two confidences are 1. If only one hand appears, the first confidence is 1 and the second confidence is 0. If no hand appears, the two confidences are 0. Correspondingly, in the actual gesture recognition process, if a hand appears, the corresponding confidence will be close to 1, otherwise it will be close to 0. Therefore, a threshold can be set to determine whether to accept the hand detection result.
[0052] In some optional embodiments, training the gesture recognition model through correction data further includes: after the convergence of the gesture recognition model reaches a preset requirement after training, randomly blurring and noising the correction data for use in training the gesture recognition model.
[0053] In order to ensure the convergence of the gesture recognition model, in the initial training stage, random blurring and noise addition processing can be chosen not to be performed on the training data. After the model converges and gradually stabilizes, that is, after the required preset requirements are met, for example, during the training process, the loss function value continues to decrease and finally tends to a stable value, or the accuracy of the model continues to improve and tends to a stable value, random blurring and noise addition processing can be chosen to be performed, that is, the image quality normalization processing of the training data is continued to be completed. By performing image quality normalization processing on the training data in stages, the training of the gesture recognition model can be efficiently completed.
[0054] In some optional embodiments, training a gesture recognition model through correction data also includes the following steps: converting the gesture recognition model into a running format of the device, recognizing an image captured by the device, and performing optical flow tracking on the hand area of the image when the hand confidence is greater than a first threshold.
[0055] For the specific application of the gesture recognition model, it is necessary to first convert the gesture recognition model into a running format supported by the device and deploy it on the device. Then, when performing recognition, the device's camera can be used to capture images, and based on the camera calibration knowledge base, the obtained camera intrinsic parameter matrix and distortion coefficients can be used for distortion correction. That is, the camera intrinsic parameters and distortion parameters are used to flatten the image, and then the hand position can be detected for the captured image sequence.
[0056] Furthermore, when the confidence of the hand output by the gesture recognition model is greater than a first threshold, the detection is considered successful, wherein the first threshold can be set according to the scene, such as the first threshold can be set to 0.95. When the confidence of the hand is greater than 0.95, optical flow tracking is performed on the hand area of the image. That is, when it is confirmed that the hand is detected for the first time, feature point tracking of the hand area is started to obtain the position of the hand in the next frame image.
[0057] By combining model detection and optical flow tracking methods, the consumption of limited hardware resources by the device can be effectively reduced, saving the hardware overhead of hand detection.
[0058] The gesture recognition method disclosed in the present application effectively improves the quantity and quality of training data through in-depth processing of training data, including using training data from multiple sources and utilizing multiple open source gesture recognition models to verify training data. Targeted processing is also performed on training data from different sources, and finally a unified verification process is performed. The correlation and complementarity between training data from multiple sources are fully utilized, and the respective advantages of different open source models are brought into play, and the reliability and practicality of training data are further enhanced, so that the recognition effect of the gesture recognition model can be greatly enhanced when training the model, the gesture recognition model required for machine learning is optimized, and the accuracy and reliability of gesture recognition are greatly improved. The method is particularly suitable for gesture recognition through a panoramic camera in an eyewear device.
[0059] Based on the above method embodiments, an electronic device is provided for implementing the embodiments and preferred embodiments of the above method. The details that have been explained will not be repeated here. The terms "module", "unit", "sub-unit", etc. used below refer to a combination of software and / or hardware that can implement the predetermined functions. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware, is also possible and conceivable.
[0060] like Figure 2 As shown, the electronic device includes:
[0061] The first module 210 is used to obtain a plurality of training data related to gesture recognition, where the training data includes a plurality of data sets, and each data set includes a plurality of images.
[0062] The second module 220 is used to pre-process the training data according to the types of various data sets.
[0063] The third module 230 is used to perform correction processing on the various pre-processed data sets respectively, and obtain correction data through the correction results.
[0064] The fourth module 240 is used to train a gesture recognition model using the correction data, and use the trained gesture recognition model for gesture recognition.
[0065] In some optional embodiments, the first module 210 obtains training data from public data, 3D generated data, and real data.
[0066] In some optional embodiments, the first module 210 obtains 3D generated data and real data, and adds images that do not contain hands as negative samples.
[0067] In some optional embodiments, after acquiring the real data, the second module 220 utilizes an open source gesture recognition model and adopts a multi-model voting method for processing.
[0068] In some optional embodiments, when the fourth module 240 is used to use the correction data for training the gesture recognition model, the image of the correction data is input to the gesture recognition model, and the position of the hand and the confidence of the hand are output.
[0069] In some optional embodiments, when the fourth module 240 is used to use the correction data for training the gesture recognition model, after the gesture recognition model achieves stable convergence, the correction data will be randomly blurred and denoised, and used to complete the training of the gesture recognition model.
[0070] In some optional embodiments, when the fourth module 240 uses the trained gesture recognition model for gesture recognition, it converts the gesture recognition model into the running format of the device, recognizes the image captured by the device, and performs optical flow tracking on the hand area of the image when the hand confidence is greater than the first threshold.
[0071] Based on the above method implementation, please refer to Figure 3 , Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. An electronic device 300 provided in an embodiment of the present application includes: a processor 301 and a memory 302, the memory 302 stores machine-readable instructions executable by the processor 301, and the machine-readable instructions execute the above method when executed by the processor 301; the electronic device can be a physical device, such as a mobile phone, a PC, a wearable device, a household appliance, a monitoring device, a game console, a drone, a robot, etc., or a virtual device, such as a virtual machine, a container, etc.
[0072] Based on the above method embodiments, a computer-readable storage medium is also provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method in any of the above embodiments are implemented.
[0073] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0074] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0075] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0076] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0077] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0078] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0079] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information, which can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0080] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0081] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0082] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A method for gesture recognition, characterized in that: include: Acquire a plurality of training data on gesture recognition, wherein the training data includes a plurality of data sets, each data set includes a plurality of images, and the plurality of data sets include public data, 3D generated data, and real data, wherein the 3D generated data and the real data include images without hands; Preprocessing the training data according to the types of the multiple data sets; The pre-processed data sets are respectively calibrated to obtain calibration results, and calibration data is obtained through the calibration results, wherein: The hand positions of the images in each data set are detected and corrected by using multiple open source hand detection models to obtain correction results, including: after obtaining multiple detection results based on the multiple open source hand detection models, correction is performed using the annotation results generated by the circumscribed rectangles of the annotated joint positions, the circumscribed rectangles generated by different models are compared, the intersection or average position of the circumscribed rectangles is calculated, and the position of the hand is corrected to obtain the correction result; The gesture recognition model is trained using the correction data, and the trained gesture recognition model is used for gesture recognition.
2. The method for gesture recognition according to claim 1, characterized in that: The preprocessing of the training data according to the types of the multiple data sets includes: In response to the data set being the real data, an open source hand detection model is used to mark the position of a hand in an image of the real data to perform preprocessing of the training data.
3. The method for gesture recognition according to claim 2, characterized in that: The method of marking the position of the hand in the image of the real data by using the open source hand detection model to preprocess the training data includes: Using multiple open source hand detection models, respectively mark the hand positions in the image of the real data to construct a hand position sample; The hand position marked in each open source hand detection model for each image in the hand position sample is determined, and images with inconsistent hand position marking results are discarded to complete the screening of the real data in the preprocessing process.
4. The method for gesture recognition according to claim 1, characterized in that: The preprocessing of the training data according to the types of the multiple data sets includes: In response to the types of the multiple data sets being public data or real data, determining a camera intrinsic parameter and a distortion parameter corresponding to each image in the data set; Converting the coordinates of the hand position of each image according to the camera intrinsic parameters and the distortion parameters, wherein the hand position of each image is determined by an open source hand detection model; The image is cropped based on the hand position in the transformed image.
5. The method for gesture recognition according to claim 1, characterized in that: The 3D generated data and the real data use images that do not contain hands as negative samples.
6. The method for gesture recognition according to claim 1, characterized in that: The method of performing correction processing on the preprocessed data sets to obtain correction results and obtaining correction data through the correction results includes: Use multiple open source hand detection models to detect and correct the hand positions of images in each dataset and obtain the correction results; Determine the annotation accuracy of each dataset and the accuracy of each open source model; The correction results are weighted averaged according to the annotation accuracy of each data set and the accuracy of each open source model to obtain the correction data.
7. A gesture recognition device, characterized in that: include: A first module is used to obtain a plurality of training data on gesture recognition, wherein the training data includes a plurality of data sets, each data set includes a plurality of images, and the plurality of data sets include public data, 3D generated data and real data, wherein the 3D generated data and the real data include images without hands; A second module is used to pre-process the training data according to the types of the multiple data sets; The third module is used to perform correction processing on the pre-processed data sets respectively, and obtain correction data through the correction results, wherein: The hand positions of the images in each data set are detected and corrected by using multiple open source hand detection models to obtain correction results, including: after obtaining multiple detection results based on the multiple open source hand detection models, correction is performed using the annotation results generated by the circumscribed rectangles of the annotated joint positions, the circumscribed rectangles generated by different models are compared, the intersection or average position of the circumscribed rectangles is calculated, and the position of the hand is corrected to obtain the correction result; The fourth module is used to train a gesture recognition model using the correction data, and use the trained gesture recognition model for gesture recognition.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for gesture recognition according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for gesture recognition according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Gesture recognition model training method and gesture recognition method and device
CN111428639A
Image detection method and device, computer readable storage medium and computer equipment
CN113284142A
Target detection method and device, and vehicle
CN117422760A