Hand key feature point detection model training method and device

By constructing the target loss function, considering the position information and weights of the key feature points and related feature points in the hand, training the deep learning network model, solving the problem of low detection accuracy in the existing technology, and achieving higher accuracy of key feature points in the hand.

CN120340107APending Publication Date: 2025-07-18BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410064588.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing deep learning network model fails to effectively consider the mutual influence between key feature points in the hand during training, resulting in low detection accuracy.

Method used

By determining the relevant feature points corresponding to the key feature points of the hand image in the training dataset, the target loss function is constructed, the position information and weights of the key feature points and the related feature points are considered, and the model is trained until the target loss function converges, and the target network model is obtained.

Benefits of technology

The detection accuracy of key feature points in the hand is improved, unreasonable prediction results such as finger distortion and fracture are reduced, and the detection stability and accuracy of the model are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340107A_ABST
    Figure CN120340107A_ABST
Patent Text Reader

Abstract

The invention provides a hand key feature point detection model training method and device, and relates to the technical field of artificial intelligence, and the method comprises the steps: determining at least one related feature point corresponding to each key feature point of a hand image in a training data set, the at least one related feature point comprises one or more of the remaining key feature points except the key feature points in the hand image; obtaining the key feature point and the position information of at least one related feature point corresponding to the key feature point in a pixel coordinate system; respectively determining corresponding weights for the position information of the key feature point and the position information of the at least one related feature point, and obtaining a target loss function; and training the hand key feature point detection model by taking the target loss function as a loss function of the hand key feature point detection model until the target loss function converges, thereby obtaining a target network model. The method is used for improving the detection precision of a hand key feature point detection model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a training method and device for a hand key feature point detection model. Background Art

[0002] Gesture recognition has very important applications in human-computer interaction scenarios, and the position detection of hand key feature points is a key means for gesture recognition. Usually, the joints of fingers are used as hand key feature points.

[0003] Currently, common methods include: training a deep learning network model using hand images, and predicting the positions of hand key feature points based on the trained deep learning network model. However, since the positions of hand key feature points are usually affected by adjacent key feature points, and the mutual influence between key feature points is not considered during the training of the deep learning network model currently, the detection accuracy of the trained deep learning network model is insufficient, and the detection accuracy of hand key feature points is not high.

[0004] Therefore, how to improve the detection accuracy of hand key feature points is an urgent problem to be solved. Summary of the Invention

[0005] To solve the problem of low detection accuracy of hand key feature points, an embodiment of this application provides a training method and device for a hand key feature point detection model to improve the detection accuracy of hand key feature points.

[0006] In a first aspect, an embodiment of this application provides a training method for a hand key feature point detection model, including:

[0007] Determine at least one relevant feature point corresponding to each key feature point in the hand image of the training dataset, where the at least one relevant feature point includes one or more of the remaining key feature points in the hand image except the key feature point;

[0008] Obtain the position information of the key feature point and the at least one relevant feature point corresponding to the key feature point in the pixel coordinate system;

[0009] Determine corresponding weights for the position information of the key feature point and the position information of the at least one relevant feature point respectively to obtain a target loss function;

[0010] Use the target loss function as the loss function of the hand key feature point detection model, and train the hand key feature point detection model with the training dataset until the target loss function converges to obtain a target network model.

[0011] As an optional implementation manner of an embodiment of the present application, the position information of the key feature points includes: a first offset of the predicted position of the key feature points in the X-axis direction and a second offset in the Y-axis direction, and the position information of the at least one related feature point includes: a third offset of each related feature point among the at least one related feature point in the X-axis direction relative to the key feature point and a fourth offset in the Y-axis direction.

[0012] The obtaining of the position information of the key feature points and the at least one related feature point corresponding to the key feature points in the pixel coordinate system includes:

[0013] Inputting the hand image into the hand key feature point detection model for feature extraction, and outputting the predicted position of the key feature points, the first offset of the predicted position in the X-axis direction and the second offset in the Y-axis direction, and the third offset of each related feature point among the at least one related feature point in the X-axis direction relative to the key feature point and the fourth offset in the Y-axis direction.

[0014] As an optional implementation manner of an embodiment of the present application, the obtaining of the target loss function by respectively determining corresponding weights for the position information of the key feature points and the position information of the at least one related feature point includes:

[0015] Determining the weights corresponding to the predicted position, the first offset, the second offset, the third offset, and the fourth offset respectively;

[0016] Performing weighted summation on the predicted position, the first offset, the second offset, the third offset, and the fourth offset through the weights to obtain the target loss function.

[0017] As an optional implementation manner of an embodiment of the present application, the inputting of the hand image into the hand key feature point detection model for feature extraction, and outputting the predicted position of the key feature points, the first offset of the predicted position in the X-axis direction and the second offset in the Y-axis direction, and the third offset of each related feature point among the at least one related feature point in the X-axis direction relative to the key feature point and the fourth offset in the Y-axis direction includes:

[0018] Inputting the hand image into the hand key feature point detection model, and extracting the global features of the hand image through the residual network in the hand key feature point detection model;

[0019] Input the global features of the hand image into the fully connected layer in the hand key feature point detection model for a fully connected operation, and output the predicted positions of the key feature points, the first offset, the second offset, the third offset, and the fourth offset.

[0020] As an optional implementation manner of the embodiment of the present application, determining at least one relevant feature point corresponding to each key feature point in the hand image in the training dataset includes:

[0021] Group the key feature points to obtain multiple key feature point groups, and each key feature point group corresponds to a finger in the hand image;

[0022] For the key feature points on non-thumb fingers, determine at least one relevant feature point corresponding to the key feature point group in the key feature point group where the key feature point is located and the key feature point group corresponding to the thumb;

[0023] For the key feature points on the thumb, determine at least one relevant feature point corresponding to the key feature point group in the key feature point group corresponding to the thumb.

[0024] As an optional implementation manner of the embodiment of the present application, for the key feature points on non-thumb fingers, determining at least one relevant feature point corresponding to the key feature point group in the key feature point group where the key feature point is located and the key feature point group corresponding to the thumb includes:

[0025] For the key feature points on non-thumb fingers, if the key feature point is the root joint of the finger, determine at least one relevant feature point corresponding to the key feature point group in the key feature point group where the key feature point is located and the key feature point group corresponding to the thumb;

[0026] If the key feature point is not the root joint of the finger, determine at least one relevant feature point corresponding to the key feature point group in the key feature point group where the key feature point is located.

[0027] As an optional implementation manner of the embodiment of the present application, after obtaining the target network model, the method further includes:

[0028] Obtain a hand image to be detected, and input the hand image to be detected into the target network model;

[0029] Detect each key feature point in the hand image to be detected through the target network model, and output the position information of each key feature point in the hand image to be detected.

[0030] Second aspect, an embodiment of the present application provides a training device for a hand key feature point detection model, including:

[0031] A determination module, configured to determine at least one associated feature point corresponding to each key feature point of the hand images in the training dataset, where the at least one associated feature point includes one or more of the remaining key feature points in the hand images other than the key feature point;

[0032] An acquisition module, configured to acquire the position information of the key feature point and at least one associated feature point corresponding to the key feature point in the pixel coordinate system;

[0033] A processing module, configured to obtain a target loss function by respectively determining corresponding weights for the position information of the key feature point and the position information of the at least one associated feature point.

[0034] A training module, configured to use the target loss function as the loss function of the hand key feature point detection model, and train the hand key feature point detection model through the training dataset until the target loss function converges to obtain a target network model.

[0035] As an optional implementation manner of an embodiment of the present application, the position information of the key feature point includes: a first offset of the predicted position of the key feature point in the X-axis direction and a second offset in the Y-axis direction, and the position information of the at least one associated feature point includes: a third offset of each associated feature point in the at least one associated feature point relative to the key feature point in the X-axis direction and a fourth offset in the Y-axis direction;

[0036] The acquisition module is specifically configured to input the hand image into the hand key feature point detection model for feature extraction, and output the predicted position of the key feature point, the first offset of the predicted position in the X-axis direction and the second offset in the Y-axis direction, and the third offset of each associated feature point in the at least one associated feature point relative to the key feature point in the X-axis direction and the fourth offset in the Y-axis direction.

[0037] As an optional implementation manner of an embodiment of the present application, the processing module is specifically configured to determine the weights corresponding to the predicted position, the first offset, the second offset, the third offset, and the fourth offset respectively;

[0038] Perform weighted summation on the predicted position, the first offset, the second offset, the third offset, and the fourth offset through the weights to obtain a target loss function.

[0039] As an optional implementation manner of the embodiment of the present application, the obtaining module is specifically configured to input the hand image into the hand key feature point detection model, and extract the global feature of the hand image through the residual network in the hand key feature point detection model;

[0040] Input the global feature of the hand image into the fully connected layer in the hand key feature point detection model for a fully connected operation, and output the predicted positions of the key feature points, the first offset, the second offset, the third offset, and the fourth offset.

[0041] As an optional implementation manner of the embodiment of the present application, the determining module is specifically configured to group the key feature points to obtain a plurality of key feature point groups, and each key feature point group corresponds to a finger in the hand image;

[0042] For the key feature points on non-thumb fingers, determine at least one relevant feature point corresponding to the key feature point group in the key feature point group where the key feature point is located and the key feature point group corresponding to the thumb;

[0043] For the key feature points on the thumb, determine at least one relevant feature point corresponding to the key feature point group in the key feature point group corresponding to the thumb.

[0044] As an optional implementation manner of the embodiment of the present application, the determining module is specifically configured to, for the key feature points on non-thumb fingers, if the key feature point is the root joint of a finger, determine at least one relevant feature point corresponding to the key feature point group in the key feature point group where the key feature point is located and the key feature point group corresponding to the thumb;

[0045] If the key feature point is not the root joint of a finger, determine at least one relevant feature point corresponding to the key feature point group in the key feature point group where the key feature point is located.

[0046] As an optional implementation manner of the embodiment of the present application, the device further includes: a detection module;

[0047] After obtaining the target network model, it is configured to obtain a hand image to be detected, and input the hand image to be detected into the target network model;

[0048] Detect each key feature point in the hand image to be detected through the target network model, and output the position information of each key feature point in the hand image to be detected.

[0049] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory and a processor, where the memory is used to store a computer program, and the processor is used to execute the training method of the hand key feature point detection model according to the first aspect or any optional implementation manner of the first aspect when calling the computer program.

[0050] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the training method of the hand key feature point detection model according to the first aspect or any optional implementation manner of the first aspect.

[0051] The technical solution provided by the embodiment of the present application has the following advantages compared with the prior art:

[0052] An embodiment of the present application provides a training method and device for a hand key feature point detection model. The method includes: determining at least one relevant feature point corresponding to the key feature point of the hand image in the training dataset, where the at least one relevant feature point includes one or more of the remaining key feature points in the hand image except the key feature point; constructing a target loss function based on the key feature point and the at least one relevant feature point; using the target loss function as the loss function of the hand key feature point detection model, and training the hand key feature point detection model through the training dataset until the target loss function converges to obtain a target network model; obtaining a hand image to be detected, inputting the hand image to be detected into the target network model, and obtaining the position of the hand key feature point corresponding to the hand image to be detected. The target loss function in the embodiment of the present application is constructed based on the key feature point and the relevant feature point, so that the hand key feature point detection model can fully learn to predict the position of the key feature point based on the key feature point and the relevant feature point during the training process. Since the influence of the relevant feature point on the key feature point to be detected is considered, the detection accuracy of the trained target network model is higher, and thus the detection accuracy of the hand key feature point can be improved. Description of the Drawings

[0053] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0055] Figure 1Flowchart of a method for training a hand key feature point detection model provided according to one or more embodiments of the present application;

[0056] Figure 2 Schematic diagram of hand key feature points provided according to one or more embodiments of the present application;

[0057] Figure 3 Flowchart of a method for training a hand key feature point detection model provided according to one or more embodiments of the present application;

[0058] Figure 4 Schematic diagram of the processing process of a hand key feature point detection model for a hand image provided according to one or more embodiments of the present application;

[0059] Figure 5 Structural block diagram of a training device for a hand key feature point detection model provided according to one or more embodiments of the present application;

[0060] Figure 6 Internal structure diagram of an electronic device provided according to one or more embodiments of the present application. Detailed implementation manners

[0061] To make the objectives, implementation manners, and advantages of the present application clearer, the following will clearly and completely describe the exemplary implementation manners of the present application with reference to the accompanying drawings in the exemplary embodiments of the present application. Apparently, the described exemplary embodiments are only a part of the embodiments of the present application, rather than all the embodiments.

[0062] Based on the exemplary embodiments described in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the claims of the present application. In addition, although the disclosed content in the present application is introduced according to exemplary one or several examples, it should be understood that each aspect of these disclosed contents can also be independently constituted as a complete implementation manner. It should be noted that the brief description of the terms in the present application is only for the convenience of understanding the following described implementation manners, rather than intending to limit the implementation manners of the present application. Unless otherwise specified, these terms should be understood according to their ordinary and general meanings.

[0063] The embodiments of the present application provide a method and a device for training a hand key feature point detection model. Among them, the method determines at least one related feature point corresponding to the key feature point of the hand image in the training dataset, and constructs a target loss function based on the key feature point and the related feature point. Therefore, the influence of the related feature point on the key feature point to be detected is considered during the model training process, so that the detection accuracy of the trained target network model is higher, and further the detection accuracy of the hand key feature point is improved.

[0064] The training method of the hand key feature point detection model provided in the embodiment of the present application can be executed by the electronic device provided in the embodiment of the present application, and can also be implemented by the training device of the hand key feature point detection model provided in the embodiment of the present application, and can also be implemented by one or more functional entities on the vehicle, which is not specifically limited in the embodiment of the present application.

[0065] The following is a detailed description of the training method of the hand key feature point detection model provided in the embodiments of the present application through several specific examples.

[0066] Figure 1 The flowchart of the training method of the hand key feature point detection model provided in the embodiment of the present application is shown in FIG. Figure 1 As shown, the training method of the hand key feature point detection model provided in this embodiment includes the following steps:

[0067] S11, determining at least one relevant feature point corresponding to each key feature point of the hand image in the training data set.

[0068] Among them, the at least one related feature point includes one or more of the remaining key feature points in the hand image except the key feature point. The embodiment of the present application needs to obtain each key feature point and at least one related feature point corresponding to each key feature point, that is, the key feature point in step S11 is any key feature point, and the key feature point can be a finger joint, a fingertip, etc. The embodiment of the present application takes the hand including 21 key feature points as an example. For example, refer to Figure 2 As shown, Figure 2 A schematic diagram of key feature points of a hand provided for an embodiment of the present application includes 16 joint points, namely 0 to 3, 5 to 7, 9 to 11, 13 to 15, 17 to 19, and 5 fingertips, namely 4, 8, 12, 16, and 20.

[0069] The at least one relevant feature point may include two key feature points, three key feature points, or four key feature points, and the number of relevant feature points corresponding to each key feature point may be determined according to actual conditions. When determining at least one relevant feature point corresponding to a key feature point, the correlation between the key feature point in the hand image and the remaining key feature points in the hand image may be first determined; and according to the strength of the correlation, a preset number of remaining key feature points with strong correlations may be determined as relevant feature points corresponding to the key feature point.

[0070] When determining the correlation between key feature points in a hand image and the remaining key feature points in the hand image, determine the positional relationship and distance between the joints of the key feature points in the hand image and the joints of the remaining key feature points in the hand image; based on the positional relationship and the distance, determine the correlation between the key feature points and the remaining key feature points. Exemplarily, the correlation between key feature points located on the same finger is relatively strong. On the same finger, the closer the distance between two key feature points, the stronger the correlation. On each finger, except for the fingertip, each key feature point corresponds to a joint that can be adjusted within a certain range, and the degrees of freedom that can be controlled are relatively large, that is, the joint can move within a certain range. The correlation between key feature points located on the same finger is relatively large, and the correlation between key feature points on two fingers that are far apart is relatively weak. For example, the correlation between key feature point 4 and key feature point 20 is very small, and the correlation between key feature point 0 and key feature points 4, 8, 12, 16, and 20 is also very small. For the fingertip, other key feature points on the corresponding finger can be used as the relevant feature points of the fingertip.

[0071] Exemplarily, group the key feature points based on the positional relationship and the joint movement range to obtain multiple key feature point groups. For example, group the key feature points located on the same finger into one group. For the joint points on the palm, they can be grouped into the key feature point group corresponding to the thumb. Group the key feature points to obtain multiple key feature point groups, and each key feature point group corresponds one-to-one to a finger in the hand image; for the key feature points on fingers other than the thumb, determine at least one relevant feature point corresponding to the key feature point group in the key feature point group where the key feature point is located and the key feature point group corresponding to the thumb; for the key feature points on the thumb, determine at least one relevant feature point corresponding to the key feature point group in the key feature point group corresponding to the thumb. The key feature point group corresponding to the thumb includes the key feature points on the thumb and the key feature points on the palm.

[0072] For the key feature points on fingers other than the thumb, if the key feature point is the root joint of the finger, determine at least one relevant feature point corresponding to the key feature point group in the key feature point group where the key feature point is located and the key feature point group corresponding to the thumb. For example, for key feature point 5, determine key feature point 6 in the key feature point group corresponding to the index finger and key feature point 0 in the key feature point group corresponding to the thumb as the relevant feature points of key feature point 5; another example, for key feature point 13, determine key feature point 14 in the key feature point group corresponding to the ring finger and key feature point 0 in the key feature point group corresponding to the thumb as the relevant feature points of key feature point 13.

[0073] For the key feature points on fingers other than the thumb, if the key feature point is not the root joint of the finger, at least one relevant feature point corresponding to the key feature point group where the key feature point is located is determined in the key feature point group. For example, for key feature point 6, key feature points 5, 7, and 8 in the key feature point group corresponding to the index finger are determined as the relevant feature points of key feature point 6; for another example, for key feature point 15, key feature points 13, 14, and 16 in the key feature point group corresponding to the ring finger are determined as the relevant feature points of key feature point 15.

[0074] For the key feature points on the thumb, at least one relevant feature point corresponding to the key feature point group where the key feature point is located is determined in the key feature point group corresponding to the thumb. For example, for key feature point 2, any two or more of key feature points 0, 1, 3, and 4 are determined as the relevant feature points of key feature point 2; for key feature point 0, key feature points 1 and 2 are determined as the relevant feature points of key feature point 0.

[0075] Taking the determination of the relevant feature points of key feature point 8 as an example, during the movement of the finger, key feature points 5 to 7 have the greatest influence on key feature point 8, and the correlation between key feature points 5 to 7 and key feature point 8 is stronger than that between other key feature points 0 to 4, 9 to 20 and key feature point 8. Then, key feature points 5 to 7 are determined as the relevant feature points of key feature point 8, that is, the position of key feature point 8 can be corrected by key feature points 5 to 7. Taking the determination of the relevant feature points of key feature point 6 as an example, during the movement of the finger, key feature points 5 and 7 have the greatest influence on key feature point 6, and the correlation between key feature points 5 and 7 and key feature point 6 is stronger than that between other key feature points 0 to 4, 8 to 20 and key feature point 6. Then, key feature points 5 and 7 are determined as the relevant feature points of key feature point 6, that is, the position of key feature point 6 can be corrected by key feature points 5 to 7.

[0076] S12. Obtain the position information of the key feature point and at least one relevant feature point corresponding to the key feature point in the pixel coordinate system.

[0077] Among them, the position information of the key feature point includes: the first offset of the predicted position of the key feature point in the X-axis direction and the second offset in the Y-axis direction, and the position information of the at least one relevant feature point includes: the third offset of each relevant feature point in the at least one relevant feature point in the X-axis direction compared with the key feature point and the fourth offset in the Y-axis direction.

[0078] Exemplarily, the position information of the key feature points and at least one associated feature point corresponding to the key feature points in the pixel coordinate system can be determined by a hand key feature point detection model. For example, the hand image is input into the hand key feature point detection model for feature extraction, and the predicted positions of the key feature points, the first offset in the X-axis direction and the second offset in the Y-axis direction of the predicted positions, and the third offset in the X-axis direction and the fourth offset in the Y-axis direction of each associated feature point in the at least one associated feature point compared to the key feature point are output.

[0079] Among them, the first offset is the offset between the predicted position of the key feature point in the X-axis direction in the pixel coordinate system and the actual position of the key feature point, the second offset is the offset between the predicted position of the key feature point in the Y-axis direction and the actual position of the key feature point, the third offset is the offset between the associated feature point predicted by the hand key feature point detection model and the key feature point in the X-axis direction, and the fourth offset is the offset between the associated feature point predicted by the hand key feature point detection model and the key feature point in the Y-axis direction. In the hand image, there is an offset between the position of the key feature point and the positions of the associated feature points, and the third offset and the fourth offset are the predicted values corresponding to this offset.

[0080] The hand key feature point detection model includes a residual network and fully connected layers. There can be multiple fully connected layers, and the parameters of each fully connected layer are different. Determining the predicted position of the key feature point, the first offset, the second offset, the third offset, and the fourth offset through the hand key feature point detection model can be achieved in the following way: the hand image is input into the hand key feature point detection model, and the global features of the hand image are extracted through the residual network in the hand key feature point detection model; the global features of the hand image are input into the fully connected layers in the hand key feature point detection model for fully connected operations, and the predicted position of the key feature point, the first offset, the second offset, the third offset, and the fourth offset are output.

[0081] Refer to Figure 3 as shown Figure 3 is a schematic diagram of the processing process of the hand key feature point detection model for the input hand image. For example, a single-channel grayscale image of 192x192x1 is input into the hand key feature point detection model, and 32-fold downsampling is performed. A global feature map of 6x6x512 is obtained through the residual network, and the global feature map is subjected to fully connected processing through multiple fully connected layers, and the predicted position of the key feature point, the first offset, the second offset, the third offset, and the fourth offset are output.

[0082] S13. Determine the corresponding weights for the position information of the key feature points and the position information of the at least one related feature point respectively, and obtain the target loss function.

[0083] Exemplarily, by assigning different weights to the predicted position and each offset, and performing weighted summation, the target loss function is obtained.

[0084] Since the target loss function is constructed based on the predicted position of the key feature points, the offset between the predicted position and the true position of the key feature points, and the predicted offset between the related feature points and the position of the key feature points, the bone characteristics of the hand are fully considered, and the target loss function plays a role of bone constraint.

[0085] S14. Use the target loss function as the loss function of the hand key feature point detection model, and train the hand key feature point detection model with the training data set until the target loss function converges, to obtain the target network model.

[0086] Among them, the hand key feature point detection model is a model with a residual network model as the backbone network. For example, it can be a model with resNet18 as the backbone network. The target network model is the hand key feature point detection model trained based on the target loss function.

[0087] During the model training process, due to the existence of the target loss function, the hand key feature point detection model can not only learn how to predict the position of a single key feature point, but also learn the position relationship between the key feature points and the corresponding related feature points and the movement range of the joints. Therefore, it is possible to avoid large prediction deviations for the position of a certain key feature point, resulting in unreasonable prediction results such as finger distortion, backward folding, and fracture in the overall hand key feature point prediction.

[0088] An embodiment of the present application provides a method for training a hand key feature point detection model, including: determining at least one relevant feature point corresponding to each key feature point in a hand image in a training dataset, where the at least one relevant feature point includes one or more of the remaining key feature points in the hand image other than the key feature point; obtaining the position information of the key feature point and at least one relevant feature point corresponding to the key feature point in a pixel coordinate system; determining corresponding weights for the position information of the key feature point and the position information of the at least one relevant feature point respectively to obtain a target loss function; using the target loss function as the loss function of the hand key feature point detection model, and training the hand key feature point detection model through the training dataset until the target loss function converges to obtain a target network model. The target loss function in the embodiment of the present application is constructed based on key feature points and relevant feature points, enabling the hand key feature point detection model to fully learn to predict the position of key feature points based on key feature points and relevant feature points during the training process. Since the influence of relevant feature points on the key feature points to be detected is considered, the detection accuracy of the trained target network model is higher, thereby improving the detection accuracy of hand key feature points.

[0089] Referring to Figure 4 as shown, Figure 4 is a flowchart of a method for training a hand key feature point detection model provided by another embodiment of the present application. In this embodiment, the same or similar steps as those in Figure 1 the embodiment shown will not be specifically described and explained. For details, refer to Figure 1 the description and explanation in the embodiment shown.

[0090] S41. Determine at least one relevant feature point corresponding to each key feature point in a hand image in a training dataset.

[0091] S42. Input the hand image into the hand key feature point detection model, and extract the global feature of the hand image through the residual network in the hand key feature point detection model.

[0092] S43. Input the global feature of the hand image into the fully connected layer in the hand key feature point detection model for a fully connected operation, and output the predicted position of the key feature point, the first offset, the second offset, the third offset, and the fourth offset.

[0093] S44. Determine the weights corresponding to the predicted position, the first offset, the second offset, the third offset, and the fourth offset respectively.

[0094] S45. Weighted sum of the predicted position, the first offset, the second offset, the third offset, and the fourth offset is performed using the weights to obtain an objective loss function.

[0095] Exemplarily, if the predicted position is denoted as Out Map, the first offset as Offset X, the second offset as Offset Y, the third offset as Neighbor Offset X, and the fourth offset as Neighbor Offset Y, the loss of Out Map is Loss(Map), the corresponding weight is λ1, the loss of Offset X is Loss(Offset X), the corresponding weight is λ2, the loss of Offset Y is Loss(Offset Y), the corresponding weight is λ3, the loss of Neighbor Offset X is Loss(N Offset X), the corresponding weight is λ4, and the loss of Neighbor Offset Y is Loss(N Offset Y), the corresponding weight is λ5, then the constructed objective loss function All Loss = λ1 * Loss(Map) + λ2 * Loss(Offset X) + λ3 * Loss(Offset Y) + λ4 * Loss(N Offset X) + λ5 * Loss(N Offset Y).

[0096] The construction of the objective loss function fully considers the predicted position of the key feature points, the offsets between the predicted position and the true position of the key feature points, and the predicted offsets between the positions of the related feature points and the key feature points, enabling the objective loss function to play a role in skeletal constraint, making the target network model also have the role of skeletal constraint, and reducing the deviation between the position of the predicted key feature points and the true position.

[0097] S46. Use the objective loss function as the loss function of the hand key feature point detection model, and train the hand key feature point detection model using the training dataset until the objective loss function converges to obtain a target network model.

[0098] After obtaining the target network model, it may further include: obtaining a hand image to be detected, inputting the hand image to be detected into the target network model, and obtaining the position information of the hand key feature points corresponding to the hand image to be detected.

[0099] The target network model is a model trained considering the influence between key feature points. Therefore, inputting the hand image to be detected into the target network model can make the prediction result of the target network model more accurate.

[0100] Based on any of the above embodiments, in order to verify the influence of relevant feature points on the predicted positions of key feature points, during the process of training a hand key feature point detection model to obtain a target network model, the hand key feature point detection model is obtained using the same training dataset to obtain a comparison network model. Among them, the loss function of the comparison network model is constructed based on the predicted position of the key feature point, the first offset in the X-axis direction of the predicted position, and the second offset in the Y-axis direction. For example, the weights corresponding to the predicted position, the first offset, and the second offset are determined, and the predicted position, the first offset, and the second offset are weighted and summed through the weights to obtain a loss function for training the comparison network model.

[0101] After obtaining the target network model and the comparison network model, the data in the test dataset is input into the target network model to obtain the first predicted position of the key feature point. The data in the test dataset is input into the comparison network model to obtain the second predicted position of the key feature point. The first predicted position and the second predicted position are respectively compared with the true position of the key feature point, and the prediction error of the target network model and the prediction error of the comparison network model can be obtained. If the prediction error of the target network model is less than the prediction error of the comparison network model, it proves that the relevant feature points will affect the position of the key feature point, and the detection accuracy of the target network model is higher.

[0102] The verification results show that the convergence speed of the target network model is faster than that of the comparison network model, the prediction error decreases, the application of the target loss function makes the target network model more stable, and the detection accuracy is higher.

[0103] After obtaining the target network model, it also includes the application process of the target network model, including: obtaining a hand image to be detected, and inputting the hand image to be detected into the target network model; detecting each key feature point in the hand image to be detected through the target network model, and outputting the position information of each key feature point in the hand image to be detected. Exemplarily, the hand image to be detected is input into the target network model, and at least one relevant feature point corresponding to each key feature point in the hand image to be detected is determined based on the target network model. The at least one relevant feature point includes one or more of the remaining key feature points in the hand image to be detected except the key feature point; predicting the first position information of the key feature point in the pixel coordinate system and the second position information of at least one relevant feature point corresponding to the key feature point in the pixel coordinate system; correcting the first position information through the second position information to obtain the position information of the key feature point, and outputting the position information corresponding to each key feature point in the hand image to be detected based on the target network model.

[0104] Among them, the method of correcting the first position information by the second position information can refer to the training process of the hand key feature point detection model, which will not be elaborated here.

[0105] Based on the same inventive concept, as an implementation of the above method, an embodiment of the present application further provides a training device for the hand key feature point detection model provided in the above embodiment, and a hand key feature point detection device for executing the hand key feature point detection method provided in the above embodiment. The device embodiment corresponds to the foregoing method embodiment. For the convenience of reading, the details in the foregoing method embodiment will not be elaborated one by one in this device embodiment. However, it should be clear that the detection of the hand key feature points in this embodiment can correspondingly implement all the contents in the foregoing method embodiment.

[0106] Figure 5 The structural schematic diagram of the training device for the hand key feature point detection model provided by an embodiment of the present application is as Figure 5 shown. The training device 500 for the hand key feature point detection model provided in this embodiment includes:

[0107] A determination module 510, configured to determine at least one associated feature point corresponding to each key feature point of the hand image in the training data set, where the at least one associated feature point includes one or more of the remaining key feature points other than the key feature point in the hand image;

[0108] An acquisition module 520, configured to acquire the position information of the key feature point and at least one associated feature point corresponding to the key feature point in the pixel coordinate system;

[0109] A processing module 530, configured to obtain a target loss function by respectively determining corresponding weights for the position information of the key feature point and the position information of the at least one associated feature point.

[0110] A training module 540, configured to use the target loss function as the loss function of the hand key feature point detection model, and train the hand key feature point detection model through the training data set until the target loss function converges, to obtain a target network model.

[0111] As an optional implementation manner of an embodiment of the present application, the position information of the key feature point includes: a first offset of the predicted position of the key feature point in the X-axis direction and a second offset in the Y-axis direction, and the position information of the at least one associated feature point includes: a third offset of each associated feature point in the at least one associated feature point relative to the key feature point in the X-axis direction and a fourth offset in the Y-axis direction;

[0112] The obtaining module 520 is specifically configured to input the hand image into the hand key feature point detection model for feature extraction, and output the predicted positions of the key feature points, a first offset in the X-axis direction and a second offset in the Y-axis direction of the predicted positions, and a third offset in the X-axis direction and a fourth offset in the Y-axis direction of each of the at least one associated feature point relative to the key feature point.

[0113] As an optional implementation manner of the embodiment of the present application, the processing module 530 is specifically configured to determine weights corresponding to the predicted position, the first offset, the second offset, the third offset, and the fourth offset respectively; perform weighted summation on the predicted position, the first offset, the second offset, the third offset, and the fourth offset through the weights to obtain an objective loss function.

[0114] As an optional implementation manner of the embodiment of the present application, the obtaining module 520 is specifically configured to input the hand image into the hand key feature point detection model, and extract the global features of the hand image through the residual network in the hand key feature point detection model; input the global features of the hand image into the fully connected layer in the hand key feature point detection model for a fully connected operation, and output the predicted positions of the key feature points, the first offset, the second offset, the third offset, and the fourth offset.

[0115] As an optional implementation manner of the embodiment of the present application, the determining module 510 is specifically configured to group the key feature points to obtain a plurality of key feature point groups, and each key feature point group corresponds to a finger in the hand image; for the key feature points on non-thumb fingers, determine at least one associated feature point corresponding to the key feature point group in the key feature point group where the key feature point is located and the key feature point group corresponding to the thumb; for the key feature points on the thumb, determine at least one associated feature point corresponding to the key feature point group in the key feature point group corresponding to the thumb.

[0116] As an optional implementation manner of the embodiment of the present application, the determining module 510 is specifically configured to, for the key feature points on non-thumb fingers, if the key feature point is the root joint of a finger, determine at least one associated feature point corresponding to the key feature point group in the key feature point group where the key feature point is located and the key feature point group corresponding to the thumb; if the key feature point is not the root joint of a finger, determine at least one associated feature point corresponding to the key feature point group in the key feature point group where the key feature point is located.

[0117] As an optional implementation manner of the embodiment of the present application, the device further includes: a detection module, configured to obtain a hand image to be detected after obtaining the target network model, and input the hand image to be detected into the target network model; detect each key feature point in the hand image to be detected through the target network model, and output the position information of each key feature point in the hand image to be detected.

[0118] In one embodiment, an electronic device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of any one of the hand key feature point detection model training methods described in the above method embodiments.

[0119] Exemplarily, Figure 6 is a schematic structural diagram of the electronic device provided by the embodiment of the present application. As Figure 6 shown, the electronic device provided in this embodiment includes: a memory 61 and a processor 62. The memory 61 is used to store a computer program; the processor 62 is used to execute the steps in the hand key feature point detection model training method provided in the above method embodiment when calling the computer program. The implementation principle and technical effect are similar and will not be elaborated here. Those skilled in the art can understand that Figure 6 the structure shown in

[0120] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of any one of the hand key feature point detection model training methods described in the above method embodiments.

[0121] Those of ordinary skill in the art can understand that all or part of the processes in the above-described embodiment methods can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-described methods. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static random access memory (SRAM) and dynamic random access memory (DRAM), etc.

[0122] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0123] For the sake of convenience of explanation, the above description has been made in conjunction with specific embodiments. However, the above discussion in some embodiments is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. According to the above teachings, various modifications and variations can be obtained. The selection and description of the above embodiments are for the purpose of better explaining the principles and actual applications, so that those skilled in the art can better use the embodiments and various different modified embodiments suitable for specific use considerations.

Claims

1. A training method for a hand key feature point detection model, characterized in that Including: Determine at least one associated feature point corresponding to each key feature point in the training dataset of hand images, where the at least one associated feature point includes one or more of the remaining key feature points in the hand image other than the key feature point; Obtain the position information of the key feature point and the at least one associated feature point corresponding to the key feature point in the pixel coordinate system; Obtain the target loss function by respectively determining the corresponding weights for the position information of the key feature point and the position information of the at least one associated feature point; Use the target loss function as the loss function of the hand key feature point detection model, and train the hand key feature point detection model with the training dataset until the target loss function converges to obtain the target network model.

2. The method according to claim 1, wherein The position information of the key feature point includes: the first offset in the X-axis direction and the second offset in the Y-axis direction of the predicted position of the key feature point, and the position information of the at least one associated feature point includes: the third offset in the X-axis direction and the fourth offset in the Y-axis direction of each associated feature point in the at least one associated feature point compared to the key feature point; The obtaining the position information of the key feature point and the at least one associated feature point corresponding to the key feature point in the pixel coordinate system includes: Input the hand image into the hand key feature point detection model for feature extraction, and output the predicted position of the key feature point, the first offset in the X-axis direction and the second offset in the Y-axis direction of the predicted position, and the third offset in the X-axis direction and the fourth offset in the Y-axis direction of each associated feature point in the at least one associated feature point compared to the key feature point.

3. The method according to claim 2, wherein The obtaining the target loss function by respectively determining the corresponding weights for the position information of the key feature point and the position information of the at least one associated feature point includes: Determine the weights corresponding to the predicted position, the first offset, the second offset, the third offset, and the fourth offset respectively; Perform weighted summation on the predicted position, the first offset, the second offset, the third offset, and the fourth offset through the weights to obtain the target loss function.

4. The method according to claim 2, characterized in that, The inputting the hand image into the hand key feature point detection model for feature extraction, and outputting the predicted position of the key feature point, the first offset in the X-axis direction and the second offset in the Y-axis direction of the predicted position, and the third offset in the X-axis direction and the fourth offset in the Y-axis direction of each associated feature point in the at least one associated feature point compared to the key feature point includes: Input the hand image into the hand key feature point detection model, and extract the global features of the hand image through the residual network in the hand key feature point detection model; Input the global features of the hand image into the fully connected layer in the hand key feature point detection model for a fully connected operation, and output the predicted positions of the key feature points, the first offset, the second offset, the third offset, and the fourth offset.

5. The method according to claim 1, wherein The determination of at least one associated feature point corresponding to each key feature point in the hand images of the training dataset includes: Group the key feature points to obtain a plurality of key feature point groups, where each key feature point group corresponds one-to-one to a finger in the hand image; For the key feature points on non-thumb fingers, determine at least one associated feature point corresponding to the key feature point group in the key feature point group where the key feature point is located and the key feature point group corresponding to the thumb; For the key feature points on the thumb, determine at least one associated feature point corresponding to the key feature point group in the key feature point group corresponding to the thumb.

6. The method according to claim 5, wherein The step of, for the key feature points on non-thumb fingers, determining at least one associated feature point corresponding to the key feature point group in the key feature point group where the key feature point is located and the key feature point group corresponding to the thumb includes: For the key feature points on non-thumb fingers, if the key feature point is the root joint of a finger, determine at least one associated feature point corresponding to the key feature point group in the key feature point group where the key feature point is located and the key feature point group corresponding to the thumb; If the key feature point is not the root joint of a finger, determine at least one associated feature point corresponding to the key feature point group in the key feature point group where the key feature point is located.

7. The method according to any one of claims 1-6, characterized in that, After obtaining the target network model, the method further includes: Obtain a hand image to be detected, and input the hand image to be detected into the target network model; Detect each key feature point in the hand image to be detected through the target network model, and output the position information of each key feature point in the hand image to be detected.

8. A training device for a hand key feature point detection model, characterized in that, It includes: A determination module, configured to determine at least one associated feature point corresponding to each key feature point in the hand images of the training dataset, where the at least one associated feature point includes one or more of the remaining key feature points in the hand image other than the key feature point; An acquisition module, configured to acquire the position information of the key feature point and at least one associated feature point corresponding to the key feature point in the pixel coordinate system; A processing module, configured to obtain a target loss function by respectively determining corresponding weights for the position information of the key feature point and the position information of the at least one associated feature point; A training module, configured to use the target loss function as the loss function of the hand key feature point detection model, and train the hand key feature point detection model through the training dataset until the target loss function converges to obtain a target network model.

9. An electronic device, comprising: A memory and a processor, the memory storing a computer program, characterized in that when the processor executes the computer program, it implements the training method of the hand key feature point detection model according to any one of claims 1 to 6, or when the processor executes the computer program, it implements the hand key feature point detection method according to claim 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the training method of the hand key feature point detection model according to any one of claims 1 to 6, or when the computer program is executed by the processor, it implements the hand key feature point detection method according to claim 7.

11. A vehicle, characterized in that, The vehicle is configured with the training device of the hand key feature point detection model according to claim 8, or the electronic device according to claim 9, or the storage medium according to claim 10.