Gesture angle recognition model optimization method and device, equipment and medium

By dividing and optimizing the gesture image sample set, and using the key point prediction model to update the gesture angle recognition model, the problem of decreased generalization ability caused by manual annotation errors is solved, and the accuracy and efficiency of the model are improved.

CN120997533APending Publication Date: 2025-11-21RUIMO INTELLIGENT TECH (HANGZHOU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511094845.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

In existing technologies, errors in manually labeled data lead to a decrease in the generalization ability of neural networks during training, and current solutions are inefficient and have low accuracy.

Method used

By dividing the gesture image sample set into an ideal sample set and a sample set to be optimized, the gesture key point prediction model is trained using the ideal sample set, and the key point and angle prediction values ​​of the sample set to be optimized are updated to optimize the gesture angle recognition model.

Benefits of technology

It improves the accuracy of the gesture angle recognition model, reduces errors in manually labeled data, and enhances the training effect of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997533A_ABST
    Figure CN120997533A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an optimization method and device for a gesture angle recognition model, equipment and a medium, and the method comprises the steps: dividing each gesture image sample into an ideal sample set and a to-be-optimized sample set according to a first gesture angle prediction value and a gesture angle labeling value of each gesture image sample; training by using the ideal sample set to obtain a gesture key point prediction model, predicting a first prediction key point and a second prediction key point of a gesture in each to-be-optimized sample by using the gesture key point prediction model, and determining a second gesture angle prediction value of each to-be-optimized sample according to the two prediction key points; and according to the first gesture angle prediction value, the second gesture angle prediction value and the gesture angle labeling value of each to-be-optimized sample, updating the gesture angle labeling value of each to-be-optimized sample to obtain a target optimized sample set. According to the technical scheme provided by the invention, the accuracy of the gesture angle recognition model can be improved, and the accuracy of labeling in the gesture image sample set is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of deep learning, and in particular to a gesture angle recognition model optimization method, device, equipment and medium. BACKGROUND

[0002] In deep learning, the error of manually labeled data can significantly affect the training of a neural network. Subjective differences, fatigue, different understanding of standards, and the complexity of the data itself can all lead to inaccurate labeling results. When these error-labeled data are used for training, the model will learn incorrect feature associations and decision rules, resulting in decreased generalization ability and increased training difficulty and uncertainty.

[0003] Current solutions to the above problems generally involve multiple people labeling the same sample and then selecting the relatively optimal labeling result. However, the current method is essentially a summary of human experience, which is inefficient and has low accuracy. SUMMARY

[0004] The present application provides a gesture angle recognition model optimization method, device, equipment and medium. The method of the present application embodiment can optimize the training set of the gesture angle recognition model, and then train the gesture angle recognition model using the optimized training set, thereby improving the accuracy of the gesture angle recognition model.

[0005] In a first aspect, the present application embodiment provides a gesture angle recognition model optimization method, comprising:

[0006] In the gesture angle recognition model, input a gesture image sample set to obtain a first gesture angle prediction value corresponding to each gesture image sample; wherein the gesture angle recognition model is trained using a gesture image sample set, the gesture image sample set is manually labeled by multiple labeling personnel, and the gesture image sample includes a first labeled key point, a second labeled key point of a gesture, and a gesture angle label value determined by the two labeled key points;

[0007] According to the first gesture angle prediction value and the gesture angle label value of each gesture image sample, each gesture image sample is divided into an ideal sample set and a to-be-optimized sample set;

[0008] After training a gesture key point prediction model using the ideal sample set, the gesture key point prediction model is used to predict the first predicted key point and the second predicted key point of the gesture in each to-be-optimized sample, and the second gesture angle prediction value of each to-be-optimized sample is determined according to the two predicted key points;

[0009] According to the first gesture angle prediction value, the second gesture angle prediction value and the gesture angle label value of each to-be-optimized sample, the gesture angle label value of each to-be-optimized sample is updated to obtain a target optimization sample set;

[0010] The target optimization sample set is input to the gesture angle recognition model again, and the gesture angle recognition model is optimized.

[0011] In a second aspect, an embodiment of the present application provides an optimization device of a gesture angle recognition model, including:

[0012] A first prediction module is configured to input a gesture image sample set into the gesture angle recognition model to obtain a first gesture angle prediction value corresponding to each gesture image sample, wherein the gesture angle recognition model is trained by using a gesture image sample set, the gesture image sample set is manually labeled by multiple labelers, and the gesture image sample includes a first labeled key point, a second labeled key point of a gesture, and a gesture angle label value determined by the two labeled key points;

[0013] A division module is configured to divide each gesture image sample into an ideal sample set and a to-be-optimized sample set according to the first gesture angle prediction value and the gesture angle label value of each gesture image sample;

[0014] A second prediction module is configured to predict a first prediction key point and a second prediction key point of a gesture in each to-be-optimized sample by using a gesture key point prediction model after the gesture key point prediction model is trained by using the ideal sample set, and determine a second gesture angle prediction value of each to-be-optimized sample according to the two prediction key points;

[0015] An update module is configured to update the gesture angle label value of each to-be-optimized sample according to the first gesture angle prediction value, the second gesture angle prediction value and the gesture angle label value of each to-be-optimized sample to obtain a target optimization sample set;

[0016] A training module is configured to input the target optimization sample set to the gesture angle recognition model again to optimize the gesture angle recognition model.

[0017] In a third aspect, an embodiment of the present application provides an electronic device, including:

[0018] at least one processor; and

[0019] a memory connected with the at least one processor; wherein

[0020] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the optimization method of the gesture angle recognition model in any one of the embodiments of the present application.

[0021] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, which stores computer instructions for causing a processor to implement the gesture angle recognition model optimization method according to any one of the embodiments of the present application when executed.

[0022] The embodiment of the present application provides a gesture angle recognition model optimization method, device, equipment and medium. The method comprises the following steps: inputting a gesture image sample set into a gesture angle recognition model to obtain first gesture angle prediction values corresponding to the gesture image samples respectively; the gesture angle recognition model is obtained by training using the gesture image sample set; the gesture image sample set is obtained by manual labeling by multiple labelers; the gesture image sample comprises first labeled key points, second labeled key points of a gesture and a gesture angle label value determined by the two labeled key points; dividing the gesture image samples into an ideal sample set and a to-be-optimized sample set according to the first gesture angle prediction values and the gesture angle label values of the gesture image samples; after obtaining a gesture key point prediction model by training using the ideal sample set, predicting first prediction key points and second prediction key points of the gesture in each to-be-optimized sample by using the gesture key point prediction model, and determining second gesture angle prediction values of the to-be-optimized samples according to the two prediction key points; updating the gesture angle label values of the to-be-optimized samples according to the first gesture angle prediction values, the second gesture angle prediction values and the gesture angle label values of the to-be-optimized samples to obtain a target optimized sample set; and inputting the target optimized sample set into the gesture angle recognition model again to optimize the gesture angle recognition model. Specifically, for the samples in the to-be-optimized sample set, the first prediction key points and the second prediction key points can be determined by using the gesture key point prediction model, and the second gesture angle prediction values of the to-be-optimized samples are recalculated. Then, the gesture angle label values of the to-be-optimized samples are updated according to the first gesture angle prediction values, the second gesture angle prediction values and the gesture angle label values of the to-be-optimized samples to obtain the target optimized sample set, and the gesture angle recognition model can be retrained by using the target optimized sample set to improve the accuracy of the gesture angle recognition model. By using the method of the embodiment of the present application, the accuracy of the gesture angle recognition model can be improved, and the error of the gesture angle in the manually labeled data set can be reduced. BRIEF DESCRIPTION OF DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort.

[0024] Figure 1 A gesture angle recognition model optimization method is provided for the first embodiment of the present application.

[0025] Figure 2 A flow chart of an optimization method of a gesture angle recognition model provided for the second embodiment of the present application;

[0026] Figure 3 A structural schematic diagram of an optimization device of a gesture angle recognition model provided for the third embodiment of the present application;

[0027] Figure 4 A structural schematic diagram of an electronic device provided for the fourth embodiment of the present application. DETAILED DESCRIPTION

[0028] In order to make the personnel in the technical field better understand the present application scheme, the technical scheme in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by the person skilled in the art without creative labor should belong to the scope of protection of the present application.

[0029] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0030] It should be noted that in the technical scheme of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the technical scheme comply with the relevant legal regulations and do not violate public order and good customs.

[0031] Embodiment one

[0032] Figure 1 The optimization method of a gesture angle recognition model provided for the first embodiment of the present application, which is particularly suitable for optimizing the training set of a gesture angle recognition model, and can be particularly suitable for solving the problem of large error of manually labeled training set. The embodiment of the present application can be executed by an optimization device of a gesture angle recognition model, which can be composed of software and / or hardware and configured in a computer or a server.

[0033] As Figure 1 shown, comprising:

[0034] Step 110, input a gesture image sample set in a gesture angle recognition model, to obtain a first gesture angle prediction value corresponding to each gesture image sample respectively; wherein the gesture angle recognition model is obtained by training using a gesture image sample set, and the gesture image sample set is obtained by manual labeling by multiple labelers, and the gesture image sample includes a first labeled key point, a second labeled key point of a gesture, and a gesture angle label value determined by the two labeled key points.

[0035] Wherein, the gesture angle recognition model is used to determine the gesture angle of the gesture image, such as the first gesture angle prediction value of the like gesture, the one gesture, or the victory gesture. In some embodiments, the like gesture can estimate the angle using the line connecting the root of the index finger (first labeled key point) and the root of the little finger (second labeled key point), and can also estimate the angle using the line connecting the tip of the thumb (first labeled key point) and the root of the thumb (second labeled key point); the one gesture can estimate the angle using the line connecting the tip of the index finger (first labeled key point) and the root of the index finger (second labeled key point), can also estimate the angle using the line connecting the first node of the index finger (first labeled key point) and the root of the index finger (second labeled key point), or can estimate the angle using the line connecting the first node of the index finger (first labeled key point) and the second node (second labeled key point); the victory gesture can estimate the angle using the line connecting the tip of the middle finger (first labeled key point) and the root of the middle finger (second labeled key point), etc. Further, due to the current technical limitations, for the samples of the gesture image sample set, they are all manually labeled by multiple labelers, i.e. according to the joint type of the fingers in the gesture, the first labeled key point and the second labeled key point of the gesture are manually labeled, and then according to the first labeled key point and the second labeled key point of the gesture, the gesture angle label value is determined. However, due to the error of manual labeling, the positions of the first labeled key point and the second labeled key point may have some deviation from the actual target key point. Therefore, the gesture angle label value calculated will also have deviation. Further, the gesture angle recognition model trained by the gesture image sample set with partial deviation will also have the problem of low prediction accuracy. Therefore, it is necessary to adjust the samples with large labeling errors in the gesture image sample set by the method of the embodiment of the present application, generate new accurate standard values, and then retrain the gesture angle recognition model by using the updated gesture image sample set, so as to improve the accuracy of the gesture angle recognition model.

[0036] Step 120, according to the first gesture angle prediction value and the gesture angle label value of each gesture image sample, divide each gesture image sample into an ideal sample set and a to-be-optimized sample set.

[0037] Optionally, if a first error determined by the first gesture angle prediction value and the gesture angle label value of the gesture image sample is greater than a preset threshold, the gesture image sample is added to the to-be-optimized sample set, otherwise, the gesture image sample is added to the ideal sample set.

[0038] Specifically, the gesture angle label value of the gesture image sample in the ideal sample set has a smaller error with the first gesture angle prediction value, meets the training requirement, and does not need to update the gesture angle label value. The gesture angle label value of the gesture image sample in the to-be-optimized sample set has a larger error with the first gesture angle prediction value, does not meet the training requirement, and needs to update the gesture angle label value according to the method of the embodiment of the present application.

[0039] Step 130, after the gesture key point prediction model is trained using the ideal sample set, the gesture key point prediction model is used to predict the first predicted key point and the second predicted key point of the gesture in each to-be-optimized sample, and the second gesture angle prediction value of each to-be-optimized sample is determined according to the two predicted key points.

[0040] Further, the samples in the ideal sample set have smaller errors and meet the training requirement, so the positions of the first labeled key point and the second labeled key point in the ideal sample set can be considered accurate, and thus can be used to train the gesture key point prediction model. The gesture key point prediction model that can be used to predict the position of the gesture key point is obtained. Further, since the positions of the first labeled key point and the second labeled key point in each to-be-optimized sample are not accurate, the first predicted key point and the second predicted key point of the gesture in each to-be-optimized sample can be determined again through the gesture key point prediction model.

[0041] Optionally, the gesture key point prediction model is trained using the ideal sample set, including:

[0042] The position coordinates of the first labeled key point and the second labeled key point in the ideal sample set are determined as the labels of the samples in the ideal sample set. The gesture key point prediction model is trained through the ideal sample set until a preset iteration condition is met. The preset iteration condition can be that the number of iterations reaches a preset number or the training error is less than a preset threshold.

[0043] Step 140, updating the gesture angle label value of each to-be-optimized sample according to the first gesture angle prediction value, the second gesture angle prediction value and the gesture angle label value of each to-be-optimized sample, to obtain a target optimized sample set.

[0044] Specifically, the prediction value with smaller error can be determined as the target gesture angle prediction value according to the errors of the first gesture angle prediction value and the second gesture angle prediction value with the gesture angle label value, and then the gesture angle label value of each to-be-optimized sample is updated according to the target gesture angle prediction value to obtain the target optimized sample set.

[0045] Further, for the determination method of determining the target optimization sample set, the following method can also be used:

[0046] According to the first angle gesture prediction value and the gesture angle label value, the image sample is rotated to a specific angle, such as 90 degrees, and then according to the rotated image sample, the angle difference between the gesture vector of the image sample and the straight line of 90 degrees is determined by artificial judgment to determine whether to update the gesture angle label value. Specifically, if the angle difference between the gesture vector of the rotated image sample corresponding to the first angle prediction value and 90 degrees is less than the angle difference between the gesture vector of the rotated image sample corresponding to the gesture angle label value and 90 degrees, it indicates that the label error of the first angle prediction value is smaller, and the gesture angle label value of the image sample can be updated to the first angle prediction value, and the updated image sample is determined as the target optimization sample.

[0047] Step 150, re-input the target optimization sample set to the gesture angle recognition model to optimize the gesture angle recognition model.

[0048] Specifically, since the gesture angle label value of the target optimization sample set has been more accurately updated, the gesture angle recognition model can be retrained by the target optimization sample set to make the prediction accuracy of the gesture angle recognition model better.

[0049] The embodiment of the application provides a gesture angle recognition model optimization method, which comprises the following steps: inputting a gesture image sample set into a gesture angle recognition model to obtain first gesture angle prediction values corresponding to each gesture image sample; wherein the gesture angle recognition model is trained by using the gesture image sample set, the gesture image sample set is manually labeled by a plurality of labelers, the gesture image sample comprises first labeled key points, second labeled key points of a gesture, and a gesture angle label value determined by the two labeled key points; each gesture image sample is divided into an ideal sample set and a to-be-optimized sample set according to the first gesture angle prediction value and the gesture angle label value of each gesture image sample; after a gesture key point prediction model is trained by using the ideal sample set, the first prediction key point and the second prediction key point of the gesture in each to-be-optimized sample are predicted by using the gesture key point prediction model, and the second gesture angle prediction value of each to-be-optimized sample is determined according to the two prediction key points; the gesture angle label value of each to-be-optimized sample is updated according to the first gesture angle prediction value, the second gesture angle prediction value and the gesture angle label value of each to-be-optimized sample, and a target optimization sample set is obtained; the gesture angle recognition model is optimized by re-inputting the target optimization sample set into the gesture angle recognition model. Specifically, for the samples in the to-be-optimized sample set, the first prediction key point and the second prediction key point can be determined by using the gesture key point prediction model, and the second gesture angle prediction value of the to-be-optimized sample is recalculated. Then, the gesture angle label value of each to-be-optimized sample is updated according to the first gesture angle prediction value, the second gesture angle prediction value and the gesture angle label value of each to-be-optimized sample, and a target optimization sample set is obtained, and then the gesture angle recognition model can be retrained by using the target optimization sample set to improve the accuracy of the gesture angle recognition model. By using the method of the embodiment of the application, the accuracy of the gesture angle recognition model can be improved, and the error of the gesture angle in the manually labeled data set can be reduced.

[0050] Embodiment two

[0051] Figure 2 The flowchart of the gesture angle recognition model optimization method provided for the embodiment two of the application is based on the above-mentioned embodiment, and further limits the determination method of the second gesture angle prediction value and the target optimization sample set.

[0052] As shown in Figure 2 , it comprises:

[0053] In step 210, a gesture image sample set is input into a gesture angle recognition model to obtain first gesture angle prediction values corresponding to each gesture image sample; wherein the gesture angle recognition model is trained by using the gesture image sample set, the gesture image sample set is manually labeled by a plurality of labelers, the gesture image sample comprises first labeled key points, second labeled key points of a gesture, and a gesture angle label value determined by the two labeled key points;

[0054] Step 220, according to the first gesture angle prediction value and the gesture angle label value of each gesture image sample, dividing each gesture image sample into an ideal sample set and a to-be-optimized sample set.

[0055] Step 230, after obtaining the gesture key point prediction model by training the ideal sample set, using the gesture key point prediction model to predict the first predicted key point and the second predicted key point of the gesture in each to-be-optimized sample.

[0056] Step 240, determining the second gesture angle prediction value of each to-be-optimized sample according to the two predicted key points.

[0057] Optionally, step 240 comprises:

[0058] determining a prediction vector according to the position coordinates of the first predicted key point and the second predicted key point; obtaining a standard vector corresponding to a horizontal axis in a preset coordinate system; determining the vector modulus of the prediction vector and the standard vector, and the vector product of the prediction vector and the standard vector; determining the included angle between the preset vector and the standard vector according to the vector modulus of the prediction vector and the standard vector and the vector product of the prediction vector and the standard vector, and determining the included angle as the second gesture angle prediction value.

[0059] Specifically, the second gesture angle prediction value can be determined according to the following formula: wherein, is a prediction vector composed of the first predicted key point A and the second predicted key point B; is a standard vector parallel to the positive half axis of the horizontal axis, which can be (0, 1), and are the vector modulus of the prediction vector and the standard vector respectively, is the vector product of the prediction vector and the standard vector.

[0060] Step 250, obtaining the first error between the first gesture angle prediction value of each to-be-optimized sample and the matching gesture angle label value, and the second error between the second gesture angle prediction value of each to-be-optimized sample and the matching gesture angle label value.

[0061] Specifically, the first error and the second error can be the absolute error between two numbers: absolute error = |A-B|, and the relative error |A-B|÷B. Herein, no limitation is made.

[0062] Step 260, updating the gesture angle label value of each to-be-optimized sample according to the first error and the second error corresponding to each to-be-optimized sample respectively, to obtain a target optimized sample.

[0063] Specifically, if the first error is greater than the second error, and the second error is less than an error threshold, the second gesture angle prediction value is determined as a target gesture angle prediction value; if the second error is greater than the first error, and the first error is less than the error threshold, the first gesture angle prediction value is determined as the target gesture angle prediction value; and the gesture angle label value of the to-be-optimized sample is updated according to the target gesture angle prediction value to obtain a target optimization sample.

[0064] In order to avoid that the relatively smaller one of the first error and the second error is still too large and cannot meet the use requirement, an error threshold is introduced to avoid that the target optimization sample still has a large error and further affects the accuracy of training.

[0065] Step 270: re-inputting the target optimization sample set into the gesture angle recognition model to perform model optimization on the gesture angle recognition model.

[0066] Optionally, after the target optimization sample set is re-inputted into the gesture angle recognition model to perform model optimization on the gesture angle recognition model, the method further includes:

[0067] After the target optimization sample set is used to update the gesture image sample set, the operation of inputting the gesture image sample set into the gesture angle recognition model to obtain the first gesture angle prediction value corresponding to each gesture image sample is performed until a preset model optimization iteration condition is met.

[0068] Specifically, a single iteration may still be unable to completely eliminate the samples with large errors in the gesture image sample set, and therefore, the operation of the embodiment of the present application can be performed multiple times, and the label in the gesture image sample set is finally made more accurate through multiple iterations, and the gesture angle recognition model obtained through training is made to predict more accurately.

[0069] The embodiment of the present application provides a gesture angle recognition model optimization method, and through the method of the embodiment of the present application, the label in the gesture image sample set can be made more accurate, and the gesture angle recognition model obtained through training can be made to predict more accurately.

[0070] Embodiment three

[0071] Figure 3 FIG. 1 is a structural schematic diagram of a gesture angle recognition model optimization device provided by the embodiment three of the present application. As shown in the figure, the device includes: Figure 3

[0072] ​The first prediction module 310 is configured to input a gesture image sample set into a gesture angle recognition model to obtain a first gesture angle prediction value corresponding to each gesture image sample, respectively; wherein the gesture angle recognition model is obtained by training using a gesture image sample set, and the gesture image sample set is obtained by manual labeling by multiple labelers; the gesture image sample includes a first labeled key point, a second labeled key point of a gesture, and a gesture angle labeled value determined by the two labeled key points;

[0073] The division module 320 is configured to divide each gesture image sample into an ideal sample set and a to-be-optimized sample set according to the first gesture angle prediction value and the gesture angle labeled value of each gesture image sample.

[0074] The second prediction module 330 is configured to, after training a gesture key point prediction model using the ideal sample set, predict the first prediction key point and the second prediction key point of the gesture in each to-be-optimized sample by using the gesture key point prediction model, and determine the second gesture angle prediction value of each to-be-optimized sample according to the two prediction key points.

[0075] The update module 340 is configured to update the gesture angle labeled value of each to-be-optimized sample according to the first gesture angle prediction value, the second gesture angle prediction value and the gesture angle labeled value of each to-be-optimized sample, to obtain a target optimized sample set.

[0076] The training module 350 is configured to re-input the target optimized sample set into the gesture angle recognition model to perform model optimization on the gesture angle recognition model.

[0077] The embodiment of the present application provides a gesture angle recognition model optimization device, which comprises the following steps: inputting gesture image sample sets in a gesture angle recognition model to obtain first gesture angle prediction values corresponding to each gesture image sample; wherein the gesture angle recognition model is obtained by training using gesture image sample sets, the gesture image sample sets are obtained by manual labeling by multiple labelers, the gesture image sample sets comprise first labeled key points, second labeled key points of gestures, and gesture angle labeled values determined by the two labeled key points; each gesture image sample is divided into an ideal sample set and a to-be-optimized sample set according to the first gesture angle prediction value and the gesture angle labeled value of each gesture image sample; after a gesture key point prediction model is obtained by training using the ideal sample set, the first prediction key point and the second prediction key point of the gesture in each to-be-optimized sample are predicted by using the gesture key point prediction model, and the second gesture angle prediction value of each to-be-optimized sample is determined according to the two prediction key points; the gesture angle labeled value of each to-be-optimized sample is updated according to the first gesture angle prediction value, the second gesture angle prediction value and the gesture angle labeled value of each to-be-optimized sample, and a target optimization sample set is obtained; the gesture angle recognition model is optimized by inputting the target optimization sample set to the gesture angle recognition model again. Specifically, for the samples in the to-be-optimized sample set, the first prediction key point and the second prediction key point can be determined by using the gesture key point prediction model, and the second gesture angle prediction value of the to-be-optimized sample is recalculated. Then, the gesture angle labeled value of each to-be-optimized sample is updated according to the first gesture angle prediction value, the second gesture angle prediction value and the gesture angle labeled value of each to-be-optimized sample, and a target optimization sample set is obtained, and then the gesture angle recognition model can be retrained by using the target optimization sample set to improve the accuracy of the gesture angle recognition model. By using the device of the embodiment of the present application, the accuracy of the gesture angle recognition model can be improved, and the error of the gesture angle in the manually labeled data set can be reduced.

[0078] Optionally, the dividing module 320 is specifically used for:

[0079] If the first error determined by the first gesture angle prediction value and the gesture angle labeled value of the gesture image sample is greater than a preset threshold, the gesture image sample is added to the to-be-optimized sample set, otherwise, the gesture image sample is added to the ideal sample set.

[0080] Optionally, the second prediction module 330 comprises:

[0081] The training unit is used for determining the position coordinates of the first labeled key point and the second labeled key point in the ideal sample set as the label of the sample in the ideal sample set; the gesture key point prediction model is trained by using the ideal sample set until a preset iteration condition is met.

[0082] The second prediction module 330 comprises:

[0083] The computing unit is configured to determine a second gesture angle prediction value of each sample to be optimized according to the two predicted key points.

[0084] The computing unit comprises:

[0085] The vector determining subunit is configured to determine a prediction vector according to the position coordinates of the first predicted key point and the second predicted key point, and obtain a standard vector corresponding to a horizontal axis in a preset coordinate system.

[0086] The vector product determining subunit is configured to determine a vector module of the prediction vector and the standard vector, and a vector product of the prediction vector and the standard vector.

[0087] The computing subunit is configured to determine an included angle between the preset vector and the standard vector according to the vector module of the prediction vector and the standard vector and the vector product of the prediction vector and the standard vector, and determine the included angle as the second gesture angle prediction value.

[0088] The updating module 340 comprises:

[0089] The error calculating unit is configured to obtain a first error between the first gesture angle prediction value of each sample to be optimized and a matching gesture angle label value, and a second error between the second gesture angle prediction value of each sample to be optimized and the matching gesture angle label value.

[0090] The updating unit is configured to update the gesture angle label value of each sample to be optimized according to the first error and the second error corresponding to each sample to be optimized respectively, to obtain a target optimized sample.

[0091] Optionally, the updating unit comprises:

[0092] The judging subunit is configured to determine the second gesture angle prediction value as a target gesture angle prediction value if the first error is greater than the second error and the second error is less than an error threshold, and determine the first gesture angle prediction value as the target gesture angle prediction value if the second error is greater than the first error and the first error is less than the error threshold.

[0093] The updating subunit is configured to update the gesture angle label value of the sample to be optimized according to the target gesture angle prediction value, to obtain the target optimized sample.

[0094] Optionally, the device further comprises an iteration module configured to return to perform the operation of inputting the gesture image sample set into the gesture angle recognition model to obtain the first gesture angle prediction value corresponding to each gesture image sample after updating the gesture image sample set using the target optimized sample set, until a preset model optimization iteration condition is met.

[0095] The gesture angle recognition model optimization device provided by the embodiment of the present application can execute the gesture angle recognition model optimization method provided by any embodiment of the present application, and has the corresponding function modules and beneficial effects of the execution method.

[0096] Embodiment Four

[0097] Figure 4 A structural schematic diagram of an electronic device 10 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit implementations of the applications described and / or claimed in this document.

[0098] As Figure 4 shown, the electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11, wherein the memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0099] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunications networks.

[0100] The processor 11 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, and the like. The processor 11 performs various methods and processes described above, such as the optimization method of the gesture angle recognition model.

[0101] In some embodiments, the optimization method of the gesture angle recognition model can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded onto the RAM 13 and executed by the processor 11, one or more steps of the optimization method of the gesture angle recognition model described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the optimization method of the gesture angle recognition model by any other suitable means, such as by means of firmware.

[0102] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0103] Computer programs used to implement the methods of the application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed, implements the functions / acts specified in the flowcharts and / or block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, or entirely on a remote machine or server.

[0104] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0105] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0106] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0107] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.

[0108] It should be understood that the various forms of flow shown above can be used to reorder, add or delete steps. For example, each step described in the present application can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions of the present application can be achieved, which is not limited herein.

[0109] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. An optimization method for a gesture angle recognition model, characterized in that, include: In the gesture angle recognition model, a set of gesture image samples is input, and the first gesture angle prediction value corresponding to each gesture image sample is obtained. The gesture angle recognition model is trained using the gesture image sample set, which is manually annotated by multiple annotators. The gesture image sample includes the first annotation key point, the second annotation key point, and the gesture angle annotation value determined by the two annotation key points. Based on the predicted value of the first gesture angle and the labeled value of the gesture angle of each gesture image sample, each gesture image sample is divided into an ideal sample set and a sample set to be optimized. After training the gesture keypoint prediction model using the ideal sample set, the gesture keypoint prediction model is used to predict the first and second predicted keypoints of the gesture in each sample to be optimized, and the second gesture angle prediction value of each sample to be optimized is determined based on the two predicted keypoints. Based on the predicted values ​​of the first and second gesture angles and the labeled gesture angles of each sample to be optimized, update the labeled gesture angle values ​​of each sample to be optimized to obtain the target optimized sample set. The target optimized sample set is re-input into the gesture angle recognition model to optimize the model.

2. The method according to claim 1, characterized in that, Based on the predicted first gesture angle and the labeled gesture angle of each gesture image sample, each gesture image sample is divided into an ideal sample set and a sample set to be optimized, including: If the first error determined by the first gesture angle prediction value and the gesture angle annotation value of the gesture image sample is greater than a preset threshold, then the gesture image sample is added to the sample set to be optimized; otherwise, it is added to the ideal sample set.

3. The method according to claim 1, characterized in that, A gesture keypoint prediction model is trained using an ideal sample set, including: The position coordinates of the first and second labeled key points in the ideal sample set are determined as the labels of the samples in the ideal sample set; The gesture key point prediction model is trained using an ideal sample set until a preset iteration condition is met.

4. The method according to claim 1, characterized in that, The predicted value of the second gesture angle for each sample to be optimized is determined based on two key prediction points, including: The prediction vector is determined based on the position coordinates of the first and second prediction key points. Obtain the standard vector corresponding to the horizontal axis in the preset coordinate system; Determine the vector magnitudes of the predicted vector and the standard vector, and the vector product of the predicted vector and the standard vector; Based on the vector magnitude of the predicted vector and the standard vector and the vector product of the predicted vector and the standard vector, the angle between the preset vector and the standard vector is determined, and the angle is determined as the second gesture angle prediction value.

5. The method according to claim 1, characterized in that, Based on the predicted first gesture angle, the predicted second gesture angle, and the gesture angle annotation value of each sample to be optimized, the gesture angle annotation value of each sample to be optimized is updated to obtain the target optimized sample, including: Obtain the first error between the first gesture angle prediction value and the matching gesture angle annotation value for each sample to be optimized, and the second error between the second gesture angle prediction value and the matching gesture angle annotation value for each sample to be optimized. Based on the first error and the second error corresponding to each sample to be optimized, the gesture angle annotation value of each sample to be optimized is updated to obtain the target optimized sample.

6. The method according to claim 5, characterized in that, Based on the first error and second error corresponding to each sample to be optimized, the gesture angle annotation value of each sample to be optimized is updated to obtain the target optimized sample, including: If the first error is greater than the second error and the second error is less than the error threshold, then the second gesture angle prediction value is determined as the target gesture angle prediction value; if the second error is greater than the first error and the first error is less than the error threshold, then the first gesture angle prediction value is determined as the target gesture angle prediction value. The gesture angle annotation value of the sample to be optimized is updated based on the predicted value of the target gesture angle to obtain the target optimized sample.

7. The method according to any one of claims 1-6, characterized in that, After re-inputting the target optimized sample set into the gesture angle recognition model and optimizing the model, the following steps are also included: After updating the gesture image sample set using the target optimization sample set, the process returns to the gesture angle recognition model, where the input gesture image sample set is used to obtain the first gesture angle prediction value corresponding to each gesture image sample, until the preset model optimization iteration conditions are met.

8. An optimization device for a gesture angle recognition model, characterized in that, include: The first prediction module is used to input a set of gesture image samples into the gesture angle recognition model and obtain the first gesture angle prediction value corresponding to each gesture image sample. The gesture angle recognition model is trained using the set of gesture image samples, which is manually annotated by multiple annotators. The gesture image samples include the first annotation key point, the second annotation key point, and the gesture angle annotation value determined by the two annotation key points. The partitioning module is used to divide each gesture image sample into an ideal sample set and a sample set to be optimized based on the first gesture angle prediction value and gesture angle annotation value of each gesture image sample; The second prediction module is used to train a gesture key point prediction model using an ideal sample set, and then use the gesture key point prediction model to predict the first and second predicted key points of the gesture in each sample to be optimized, and determine the second gesture angle prediction value of each sample to be optimized based on the two predicted key points. The update module is used to update the gesture angle annotation value of each sample to be optimized based on the first gesture angle prediction value, the second gesture angle prediction value, and the gesture angle annotation value, so as to obtain the target optimized sample set. The training module is used to re-input the target optimized sample set into the gesture angle recognition model to optimize the model.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the optimization method of the gesture angle recognition model according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the optimization method of the gesture angle recognition model according to any one of claims 1-7.