Hand key point detection method and device
By acquiring and combining the continuous sequence image features of the hand image, using the combination of residual network and fully connected network, the problem of inaccurate detection of hand key points in the prior art is solved, and the accuracy of gesture recognition is improved.
Patent Information
- Application Number
- CN202410064521.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-16
- Publication Date
- 2025-07-18
AI Technical Summary
In the deep learning-based network model, the prior art has failed to effectively solve the problem of hand key point detection accuracy on continuous sequence images, resulting in low accuracy of gesture recognition.
By acquiring multiple consecutive sequence images of the hand image to be predicted, the first key point feature with timing information is extracted, and combined with the second key point feature of the key point detection model, the fully connected network is input to obtain the position information of the hand key point position, and the residual network is used for feature extraction and fusion.
The accuracy of hand key points detection is improved, the motion characteristics and jitter characteristics of hand key points movement changes are retained, and the accuracy of detection is enhanced.
Smart Images

Figure CN120340106A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of object detection, and more particularly, to a method and device for detecting hand key points. Background Art
[0002] Gesture recognition has very important applications in human-computer interaction scenarios, and the detection of the positions of hand key points is a key means for gesture recognition.
[0003] Currently, hand key points are usually detected based on a deep learning network model. For example, after using a convolutional network to extract the global features of a hand image, a feature heat map of each hand key point is obtained through downsampling and then upsampling, so as to determine the position of the hand key point. However, this method does not consider the jitter problem when the hand image is a continuous sequence of images. Therefore, when detecting hand key points on a continuous sequence of images based on the existing method, the accuracy of the predicted hand key points is not high. Summary of the Invention
[0004] In order to improve the accuracy of the predicted hand key points, embodiments of this application provide a method, device, electronic device, and computer-readable storage medium for detecting hand key points.
[0005] In a first aspect, embodiments of this application provide a method for detecting hand key points, including:
[0006] Obtain a set of hand sequence images corresponding to the hand image to be predicted, where the set of hand sequence images includes multiple consecutive hand sequence images before the hand image to be predicted;
[0007] Extract features from multiple hand sequence images in the set of hand sequence images to obtain first key point features of the hand key points with temporal information;
[0008] Input the hand image to be predicted into the residual network in the key point detection model to obtain second key point features of the hand key points in the hand image to be predicted;
[0009] Merge the first key point features and the second key point features to obtain a fusion feature of the hand image to be predicted, and input the fusion feature into the fully connected network in the key point detection model to obtain position information of multiple hand key points in the hand image to be predicted.
[0010] As an optional implementation manner of the embodiments of this application, the extracting features from multiple hand sequence images in the set of hand sequence images to obtain first key point features of the hand key points with temporal information includes:
[0011] Determine the actual positions of the hand key points in the pixel coordinate system in each hand sequence image, the first offset feature of the hand key points in the X-axis direction, and the second offset feature of the hand key points in the Y-axis direction;
[0012] Perform a convolution operation on the actual positions of the hand key points, the first offset feature, and the second offset feature in each hand sequence image to obtain the first key point feature of the hand key points based on time series.
[0013] As an optional implementation manner of an embodiment of the present application, before inputting the hand image to be predicted into the residual network in the key point detection model to obtain the second key point feature of the hand key points in the hand image to be predicted, the method further includes:
[0014] Obtain a training data set, and train the key point detection model based on the training data set to obtain a trained key point detection model;
[0015] The obtaining of the training data set includes:
[0016] Obtain multiple training images and a set of sequence images respectively corresponding to each training image in the multiple training images as the training data set, where the set of sequence images corresponding to each training image includes multiple consecutive sequence images before the training image.
[0017] As an optional implementation manner of an embodiment of the present application, the training the key point detection model based on the training data set to obtain a trained key point detection model includes:
[0018] For each training image, determine at least one related key point corresponding to each hand key point in the training image, and the at least one related key point includes one or more of the remaining hand key points other than the hand key points in the training image;
[0019] Input the training image and the set of sequence images corresponding to the training image into the key point detection model to obtain the first position information of the hand key points and the second position information of at least one related key point corresponding to the hand key points;
[0020] Adjust the parameters of the key point detection model based on the first position information and the second position information until the training end condition is met, and obtain a trained key point detection model.
[0021] As an optional implementation manner of an embodiment of the present application, the first position information includes the predicted positions of each hand key point in the training image in the pixel coordinate system, the first offset of the hand key point in the X-axis direction, and the second offset of the hand key point in the Y-axis direction; the second position information includes the predicted positions of at least one related key point corresponding to the hand key point in the pixel coordinate system, the third offset of the at least one related key point in the X-axis direction, and the fourth offset of the at least one related key point in the Y-axis direction;
[0022] Adjusting the parameters of the key point detection model based on the first position information and the second position information until the training end condition is met to obtain a trained key point detection model includes:
[0023] Determining the weights corresponding to the predicted position, the first offset, the second offset, the third offset, and the fourth offset of the hand key point in the pixel coordinate system respectively;
[0024] Performing weighted summation on the predicted position, the first offset, the second offset, the third offset, and the fourth offset through the weights to obtain an objective loss function;
[0025] Training the objective loss function as the loss function of the key point detection model until the training end condition is met to obtain a trained key point detection model.
[0026] As an optional implementation manner of an embodiment of the present application, for each training image, determining at least one related key point corresponding to each hand key point in the training image includes:
[0027] For each training image, grouping the hand key points in the training image to obtain a plurality of hand key point groups, and each hand key point group corresponds to a finger in the training image;
[0028] For the hand key points on non-thumb fingers, determining at least one related key point corresponding to the hand key point in the hand key point group where the hand key point is located and the hand key point group corresponding to the thumb;
[0029] For the hand key points on the thumb, determining at least one related key point corresponding to the hand key point in the hand key point group corresponding to the thumb.
[0030] As an optional implementation manner of an embodiment of the present application, for the hand key points on non-thumb fingers, determining at least one relevant key point corresponding to the hand key point in the hand key point group where the hand key point is located and the hand key point group corresponding to the thumb includes:
[0031] For a hand key point on a non-thumb finger, if the hand key point is the root joint of a finger, determining at least one relevant key point corresponding to the hand key point in the hand key point group where the hand key point is located and the hand key point group corresponding to the thumb;
[0032] If the hand key point is other joints except the root joint of a finger, determining at least one relevant key point corresponding to the hand key point in the hand key point group where the hand key point is located.
[0033] In a second aspect, an embodiment of the present application provides a detection device for hand key points, including
[0034] An acquisition module, configured to acquire a set of hand sequence images corresponding to a hand image to be predicted, where the set of hand sequence images includes multiple consecutive hand sequence images before the hand image to be predicted;
[0035] An extraction module, configured to perform feature extraction on multiple hand sequence images in the set of hand sequence images to obtain first key point features of the hand key points with temporal information;
[0036] A processing module, configured to input the hand image to be predicted into a residual network in a key point detection model to obtain second key point features of hand key points in the hand image to be predicted;
[0037] A prediction module, configured to merge the first key point features and the second key point features to obtain a fusion feature of the hand image to be predicted, and input the fusion feature into a fully connected network in the key point detection model to obtain position information of multiple hand key points in the hand image to be predicted.
[0038] As an optional implementation manner of an embodiment of the present application, the extraction module is specifically configured to determine the actual position of the hand key point in a pixel coordinate system in each hand sequence image, a first offset feature of the hand key point in the X-axis direction, and a second offset feature of the hand key point in the Y-axis direction;
[0039] Perform a convolution operation on the actual position of the hand key point, the first offset feature, and the second offset feature in each hand sequence image to obtain first key point features of the hand key points based on time series.
[0040] As an optional implementation manner of an embodiment of the present application, the device further includes:
[0041] A training module, configured to obtain a training data set, and train the key point detection model based on the training data set to obtain a trained key point detection model;
[0042] Specifically, the training module is configured to obtain multiple training images and a sequence image set corresponding to each training image in the multiple training images as the training data set, where the sequence image set corresponding to each training image includes multiple consecutive sequence images before the training image.
[0043] As an optional implementation manner of an embodiment of the present application, specifically, the training module is configured to, for each training image, determine at least one relevant key point corresponding to each hand key point in the training image, and the at least one relevant key point includes one or more of the remaining hand key points in the training image except the hand key points;
[0044] Input the training image and the sequence image set corresponding to the training image into the key point detection model, and obtain first position information of the hand key points and second position information of at least one relevant key point corresponding to the hand key points;
[0045] Adjust parameters of the key point detection model based on the first position information and the second position information until a training end condition is met, and obtain a trained key point detection model.
[0046] As an optional implementation manner of an embodiment of the present application, the first position information includes the predicted positions of each hand key point in the training image in a pixel coordinate system, a first offset of the hand key point in the X-axis direction, and a second offset of the hand key point in the Y-axis direction; the second position information includes the predicted positions of at least one relevant key point corresponding to the hand key point in the pixel coordinate system, a third offset of the at least one relevant key point in the X-axis direction, and a fourth offset of the at least one relevant key point in the Y-axis direction;
[0047] Specifically, the training module is configured to determine weights corresponding to the predicted position, the first offset, the second offset, the third offset, and the fourth offset of the hand key point in the pixel coordinate system respectively;
[0048] Perform weighted summation on the predicted position, the first offset, the second offset, the third offset, and the fourth offset through the weights to obtain an objective loss function;
[0049] Train the target loss function as the loss function of the key point detection model until the training end condition is met, and obtain the trained key point detection model.
[0050] As an optional implementation manner of the embodiment of the present application, the training module is specifically configured to group the hand key points in each training image for each training image to obtain a plurality of hand key point groups, and each hand key point group corresponds to a finger in the training image one by one;
[0051] For the hand key points on non-thumb fingers, determine at least one relevant key point corresponding to the hand key point in the hand key point group where the hand key point is located and the hand key point group corresponding to the thumb;
[0052] For the hand key points on the thumb, determine at least one relevant key point corresponding to the hand key point in the hand key point group corresponding to the thumb.
[0053] As an optional implementation manner of the embodiment of the present application, the training module is specifically configured to, for the hand key points on non-thumb fingers, if the hand key point is the root joint of a finger, determine at least one relevant key point corresponding to the hand key point in the hand key point group where the hand key point is located and the hand key point group corresponding to the thumb;
[0054] If the hand key point is other joints except the root joint of the finger, determine at least one relevant key point corresponding to the hand key point in the hand key point group where the hand key point is located.
[0055] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory and a processor, where the memory is used to store a computer program, and the processor is used to execute the hand key point detection method according to the first aspect or any optional implementation manner of the first aspect when calling the computer program.
[0056] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and the computer program, when executed by a processor, implements the hand key point detection method according to the first aspect or any optional implementation manner of the first aspect.
[0057] The technical solution provided by the embodiment of the present application has the following advantages compared with the prior art:
[0058] The embodiments of the present application provide a method, an apparatus, an electronic device, and a computer-readable storage medium for detecting hand key points. The method includes: obtaining a set of hand sequence images corresponding to a hand image to be predicted, where the set of hand sequence images includes multiple consecutive hand sequence images before the hand image to be predicted; extracting features from the multiple hand sequence images in the set of hand sequence images to obtain first key point features of the hand key points with temporal information; inputting the hand image to be predicted into a residual network in a key point detection model to obtain second key point features of the hand key points in the hand image to be predicted; merging the first key point features and the second key point features to obtain a fused feature of the hand image to be predicted, and inputting the fused feature into a fully connected network in the key point detection model to obtain position information of multiple hand key points in the hand image to be predicted. In the embodiments of the present application, the introduction of multiple consecutive hand sequence images takes into account the continuous movement and jitter of hand key points in a temporal frame, so that the obtained first key point features of the hand key points have temporal information. The fused feature obtained by merging the first key point features containing temporal information and the second key point features of the hand key points predicted by the key point detection model is equivalent to retaining the motion features and jitter features of the action changes of the hand key points, thereby improving the accuracy of the hand key points obtained based on the fused feature. Description of the Drawings
[0059] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0060] To more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0061] Figure 1 It is a flowchart of a method for detecting hand key points according to one or more embodiments of the present application;
[0062] Figure 2 It is a schematic diagram of the process of obtaining the position information of hand key points according to one or more embodiments of the present application;
[0063] Figure 3 It is a flowchart of a method for detecting hand key points according to one or more embodiments of the present application;
[0064] Figure 4 It is a schematic diagram of hand key points according to one or more embodiments of the present application;
[0065] Figure 5 Block diagram of a detection device for hand key points provided for one or more embodiments of the present application;
[0066] Figure 6 Block diagram of a detection device for hand key points provided for one or more embodiments of the present application;
[0067] Figure 7 Internal structure diagram of an electronic device provided according to one or more embodiments of the present application. Detailed implementation manners
[0068] To make the objectives, implementation manners and advantages of the present application clearer, the following will clearly and completely describe the exemplary implementation manners of the present application with reference to the accompanying drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only a part rather than all of the embodiments of the present application.
[0069] Based on the exemplary embodiments described in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the appended claims of the present application. In addition, although the disclosed content in the present application is introduced according to one or several exemplary instances, it should be understood that each aspect of these disclosed contents can also constitute a complete implementation manner alone. It should be noted that the brief description of the terms in the present application is only for the convenience of understanding the subsequent described implementation manners, rather than intending to limit the implementation manners of the present application. Unless otherwise specified, these terms should be understood according to their ordinary and common meanings.
[0070] The embodiments of the present application provide a method and a device for detecting hand key points. Among them, the method obtains multiple consecutive hand sequence images before the hand image to be predicted, obtains the first key point features based on time series of each key point based on the multiple consecutive hand sequence images, and obtains the position information of multiple key points in the hand image to be predicted by combining and processing the first key point features and the second key point features predicted by the key point detection model. Since the time series features of each key point are integrated, the spatio-temporal features of the action changes of the hand key points are retained, thereby improving the detection accuracy of the hand key points.
[0071] The method for detecting hand key points provided by the embodiments of the present application can be executed by the electronic device provided by the embodiments of the present application, can also be implemented by the detection device for hand key points provided by the embodiments of the present application, and can also be implemented by one or more functional entities on a vehicle. The embodiments of the present application do not make specific limitations.
[0072] The following elaborates in detail on the method for detecting hand key points provided by the embodiments of the present application through several specific embodiments.
[0073] Figure 1 This is a flowchart of the method for detecting hand key points provided by the embodiments of the present application. Referring to Figure 1 as shown, the method for detecting hand key points provided in this embodiment includes the following steps:
[0074] S11. Obtain a set of hand sequence images corresponding to the hand image to be predicted.
[0075] The set of hand sequence images includes multiple consecutive hand sequence images before the hand image to be predicted. The number of consecutive hand sequence images can be determined according to actual needs. For example, it can be 10, 12, etc.
[0076] Exemplarily, taking consecutive sequence images as an example, when the set of hand sequence images included in the set of hand sequence images is 10 hand images, when the hand image to be predicted is the 11th frame image, the 1st to 10th frames in the consecutive sequence images are used as the set of hand sequence images corresponding to the hand image to be predicted; when the hand image to be predicted is the 15th frame image, the 5th to 14th frames in the consecutive sequence images are used as the set of hand sequence images corresponding to the hand image to be predicted.
[0077] S12. Extract features from multiple hand sequence images in the set of hand sequence images to obtain first key point features of the hand key points with temporal information.
[0078] As an implementable way to extract features from the hand sequence images, it includes: determining the actual positions of the hand key points in the pixel coordinate system in each hand sequence image, the first offset feature of the hand key points in the X-axis direction, and the second offset feature of the hand key points in the Y-axis direction; performing a convolution operation on the actual positions of the hand key points, the first offset feature, and the second offset feature in each hand sequence image to obtain the first key point features of the hand key points based on time series.
[0079] Exemplarily, taking a hand sequence image as an example for illustration, the hand sequence image is divided into a preset number of pixel regions, and the target pixel region where the hand key point is located is determined. The position of the target pixel region in the pixel coordinate system is the actual position of the hand key point in the pixel coordinate system. In this target pixel region, the offset X and offset Y of the hand key point are determined. Offset X is the first offset feature of the hand key point in the X-axis direction, and offset Y is the second offset feature of the hand key point in the Y-axis direction.
[0080] Exemplarily, the obtained hand sequence image is processed into a grayscale image with a size of 192x192, and the 192x192 hand sequence image is re-divided according to the downsampling factor of the key point detection model. For example, if the downsampling factor of the key point detection model is 32, the 192x192 hand sequence image is re-divided into a 6x6 global feature map. Taking 10 hand sequence images as an example, if each hand sequence image includes 21 hand key points, a feature map of 21x10x6x6 is obtained after re-division, and the actual position of each hand key point in the feature map is determined. If the position information of a hand key point in the 192x192 image is (110, 170), in the 6x6 global feature map, the value at the position (3, 5) is 1, and the values at other positions from (0, 0) to (5, 5) are all 0. Then, for this hand key point in the 6x6 global feature map, the corresponding value of the first offset feature is 0.4375, and the calculation method is (110–3*32) / 32. The calculation method of the corresponding value of the second offset feature is similar and will not be elaborated here.
[0081] The sequential sorting of the feature maps in space directly represents the temporal features. The global feature maps, the first bias features, and the second bias features of each hand sequence image are concatenated in chronological order, and then feature fusion is performed through a trained convolutional network. For example, three-layer convolutional operations are performed to obtain the first key point feature containing temporal features of the hand key points.
[0082] S13. Input the hand image to be predicted into the residual network in the key point detection model to obtain the second key point feature of the hand key points in the hand image to be predicted.
[0083] The key point detection model is a network model with a residual network as the backbone network, and the residual network can be resNet18. That is, the hand image to be predicted is downsampled according to the downsampling factor through the residual network in the key point detection model, and feature extraction is performed with resNet18 as the backbone network to obtain the global feature map corresponding to the hand image to be predicted. For example, a 192x192 input image is processed into a 6x6 global feature map to obtain the second key point feature of the hand key points in the hand image to be predicted.
[0084] S14. Merge the first key point feature and the second key point feature to obtain the fusion feature of the hand image to be predicted, and input the fusion feature into the fully connected network in the key point detection model to obtain the position information of multiple hand key points in the hand image to be predicted.
[0085] The "concat" here includes any one of addition, subtraction, multiplication, and division operations, which combines the first key point feature with the second key point feature to obtain the fused feature of the hand image to be predicted, and then inputs the fused feature into the fully connected network in the key point detection model for a fully connected operation to obtain the position information of multiple hand key points in the hand image to be predicted.
[0086] The method for detecting hand key points provided by the embodiments of the present application includes: obtaining a set of hand sequence images corresponding to the hand image to be predicted, where the set of hand sequence images includes multiple consecutive hand sequence images before the hand image to be predicted; extracting features from the multiple hand sequence images in the set of hand sequence images to obtain the first key point feature of the hand key points with temporal information; inputting the hand image to be predicted into the residual network in the key point detection model to obtain the second key point feature of the hand key points in the hand image to be predicted; combining the first key point feature with the second key point feature to obtain the fused feature of the hand image to be predicted, and inputting the fused feature into the fully connected network in the key point detection model to obtain the position information of multiple hand key points in the hand image to be predicted. In the embodiments of the present application, the introduction of multiple consecutive hand sequence images takes into account the continuous movement and jitter of hand key points in the temporal frame, so that the obtained first key point feature of the hand key points has temporal information. The fused feature obtained by combining the first key point feature containing temporal information with the second key point feature of the hand key points predicted by the key point detection model is equivalent to retaining the motion feature and jitter feature of the action change of the hand key points, thereby improving the accuracy of the hand key points obtained based on the fused feature.
[0087] On the basis of the above steps, it also includes the training process of the key point detection model. Exemplarily, refer to Figure 2 as shown. Figure 2 FIG. 10 is a schematic diagram of the process of obtaining the hand image to be predicted provided by an embodiment of the present application. The key point detection model is trained based on the training data set to obtain a trained key point detection model. Exemplarily, in combination with Figure 3 as shown. Figure 3 FIG. 14 is a flowchart of the steps of the method for detecting hand key points provided by another embodiment of the present application. Before step S11, the following steps S31 to S34 are further included. In this embodiment, the steps that are the same or similar to the embodiment Figure 1 shown are not explained and described again. Specifically, reference can be made to the explanation and description in the embodiment Figure 1 shown.
[0088] S31. Obtain multiple training images and a set of sequence images corresponding to each training image in the multiple training images as a training data set.
[0089] Among them, the sequence image set corresponding to each training image includes multiple consecutive sequence images before the training image.
[0090] S32. For each training image, determine at least one relevant key point corresponding to each hand key point in the training image.
[0091] Among them, the at least one relevant key point includes one or more of the remaining hand key points in the training image except the hand key point.
[0092] In the embodiments of the present application, it is necessary to obtain each hand key point and at least one relevant key point corresponding to each hand key point. That is, the hand key point in step S11 is any hand key point. The hand key point can be a finger joint, a fingertip, etc. In the embodiments of the present application, an example is given with 21 hand key points included in the hand. Exemplarily, referring to Figure 4 as shown Figure 4 is a schematic diagram of hand key points provided by an embodiment of the present application, including 16 joint points such as 0 to 3, 5 to 7, 9 to 11, 13 to 15, 17 to 19, and 5 fingertips such as 4, 8, 12, 16, 20.
[0093] The at least one relevant key point may include two hand key points, may also include three hand key points, and may further include four hand key points. The number of relevant key points corresponding to each hand key point can be determined according to the actual situation. When determining at least one relevant key point corresponding to a hand key point, the correlation between the hand key point in the hand image and the remaining hand key points in the hand image can be determined first; according to the strength of the correlation, a preset number of the remaining hand key points with strong correlation are determined as the relevant key points corresponding to the hand key point.
[0094] When determining the correlation between the hand key points in the hand image and the remaining hand key points in the hand image, determine the positional relationship and distance between the joints of the hand key points in the hand image and the joints of the remaining hand key points in the hand image; based on the positional relationship and the distance, determine the correlation between the hand key points and the remaining hand key points. Exemplarily, the correlation between the hand key points located on the same finger is relatively strong. On the same finger, the closer the two hand key points are, the stronger the correlation between them. On each finger, except for the fingertip, each hand key point corresponds to a joint that can be adjusted within a certain range, and the degree of freedom that can be controlled is relatively large, that is, the joint can move within a certain range. The correlation between the hand key points located on the same finger is relatively large, and the correlation between the hand key points on two fingers that are far apart is relatively weak. For example, the correlation between hand key point 4 and key point feature point 20 is very small, and the correlation between key point feature point 0 and key point feature points 4, 8, 12, 16, 20 is also very small. For the fingertip, other hand key points on the corresponding finger can be used as the relevant key points of the fingertip.
[0095] Exemplarily, for each training image, group the hand key points based on the positional relationship and the joint movement range to obtain multiple hand key point groups. For example, group the hand key points located on the same finger into one group. For the joint points on the palm, they can be grouped into the hand key point group corresponding to the thumb. Group the hand key points to obtain multiple hand key point groups, and each hand key point group corresponds one-to-one to the finger in the hand image; for the hand key points on the non-thumb fingers, determine at least one relevant key point corresponding to the hand key point in the hand key point group where the hand key point is located and the hand key point group corresponding to the thumb; for the hand key points on the thumb, determine at least one relevant key point corresponding to the hand key point in the hand key point group corresponding to the thumb. The hand key point group corresponding to the thumb includes the hand key points on the thumb and the hand key points on the palm.
[0096] For the hand key points on the non-thumb fingers, if the hand key point is the root joint of the finger, determine at least one relevant key point corresponding to the hand key point in the hand key point group where the hand key point is located and the hand key point group corresponding to the thumb. For example, for hand key point 5, determine hand key point 6 in the hand key point group corresponding to the index finger and hand key point 0 in the hand key point group corresponding to the thumb as the relevant key points of hand key point 5; for another example, for hand key point 13, determine hand key point 14 in the hand key point group corresponding to the ring finger and hand key point 0 in the hand key point group corresponding to the thumb as the relevant key points of hand key point 13.
[0097] For the hand key points on fingers other than the thumb, if the hand key point is not the root joint of a finger, at least one related key point corresponding to the hand key point is determined in the group of hand key points where the hand key point is located. For example, for hand key point 6, hand key points 5, 7, and 8 in the group of hand key points corresponding to the index finger are determined as the related key points of hand key point 6; for another example, for hand key point 15, hand key points 13, 14, and 16 in the group of hand key points corresponding to the ring finger are determined as the related key points of hand key point 15.
[0098] For the hand key points on the thumb, at least one related key point corresponding to the group of hand key points is determined in the group of hand key points corresponding to the thumb. For example, for hand key point 2, any two or more of hand key points 0, 1, 3, and 4 are determined as the related key points of hand key point 2; for hand key point 0, hand key points 1 and 2 are determined as the related key points of hand key point 0.
[0099] Taking the determination of the related key points of hand key point 8 as an example, during the movement of the finger, hand key points 5 to 7 have the greatest influence on hand key point 8, and the correlation between hand key points 5 to 7 and hand key point 8 is stronger than that between other hand key points 0 to 4, 9 to 20 and hand key point 8. Then, hand key points 5 to 7 are determined as the related key points of hand key point 8, that is, the position of hand key point 8 can be corrected through hand key points 5 to 7. Taking the determination of the related key points of hand key point 6 as an example, during the movement of the finger, hand key points 5 and 7 have the greatest influence on hand key point 6, and the correlation between hand key points 5 and 7 and hand key point 6 is stronger than that between other hand key points 0 to 4, 8 to 20 and hand key point 6. Then, hand key points 5 and 7 are determined as the related key points of hand key point 6, that is, the position of hand key point 6 can be corrected through hand key points 5 to 7.
[0100] S33. Input the training image and the corresponding sequence image set of the training image into the key point detection model to obtain the first position information of the hand key points and the second position information of at least one related key point corresponding to the hand key points.
[0101] Among them, the first position information includes the predicted positions of the hand key points in the training image in the pixel coordinate system, the first offset of the hand key point in the X-axis direction, and the second offset of the hand key point in the Y-axis direction. The second position information includes the predicted positions of at least one related key point corresponding to the hand key point in the pixel coordinate system, the third offset of the at least one related key point in the X-axis direction, and the fourth offset of the at least one related key point in the Y-axis direction;
[0102] The key point detection model may include a residual network and multiple fully-connected layers. During the training process of the key point detection model, the training image is input into the residual network to extract the corresponding global feature map. The sequential images corresponding to the training image are processed according to the method described in step S12 above to obtain a global feature map with temporal information. The two global feature maps are combined as the input of the fully-connected layer to obtain the predicted position of the hand key points in the pixel coordinate system, the first offset of the hand key points in the X-axis direction, and the second offset of the hand key points in the Y-axis direction. The second position information includes the predicted positions of at least one related key point corresponding to the hand key points in the pixel coordinate system, the third offset of the at least one related key point in the X-axis direction, and the fourth offset of the at least one related key point in the Y-axis direction.
[0103] The introduction of the sequential images corresponding to the training image is equivalent to introducing temporal information during the training process of the key point detection model, improving the accuracy and stability of the key point detection model.
[0104] S34. Adjust the parameters of the key point detection model based on the first position information and the second position information until the training end condition is met, and obtain the trained key point detection model.
[0105] Exemplarily, determine the weights corresponding to the predicted position of the hand key points, the first offset, the second offset, the third offset, and the fourth offset in the pixel coordinate system; perform weighted summation on the predicted position, the first offset, the second offset, the third offset, and the fourth offset through the weights to obtain the target loss function; use the target loss function as the loss function of the key point detection model for training until the training end condition is met, and obtain the trained key point detection model.
[0106] Exemplarily, if the predicted position is denoted as Out Map, the first offset as Offset X, the second offset as Offset Y, the third offset as Neighbor Offset X, the fourth offset as Neighbor Offset Y, the loss of Out Map as Loss(Map), the corresponding weight as λ1, the loss of Offset X as Loss(Offset X), the corresponding weight as λ2, the loss of Offset Y as Loss(Offset Y), the corresponding weight as λ3, the loss of Neighbor Offset X as Loss(N Offset X), the corresponding weight as λ4, and the loss of Neighbor Offset Y as Loss(N Offset Y), the corresponding weight as λ5, then the constructed objective loss function All Loss = λ1 * Loss(Map) + λ2 * Loss(Offset X) + λ3 * Loss(Offset Y) + λ4 * Loss(N Offset X) + λ5 * Loss(N Offset Y).
[0107] During the model training process, due to the existence of the objective loss function, the key point detection model can not only learn how to predict the positions of individual hand key points, but also learn the positional relationships between hand key points and their corresponding related key points and the movement ranges of joints. Therefore, it can avoid large prediction deviations in the positions of certain hand key points, which may lead to unreasonable prediction results such as finger distortion, backward folding, and fracture in the overall hand key point prediction, and play a role in skeletal constraint.
[0108] To improve the robustness of the network, in the embodiments of the present application, the sequence images in the sequence image set corresponding to the training images are subjected to interference processing. The interference processing includes copying, filling with blanks, and randomly swapping adjacent two-frame sequence images to form an interference sequence image set, and the key point detection model is trained.
[0109] Based on the same inventive concept, as an implementation of the above method, the embodiments of the present application also provide a detection device for hand key points that executes the method provided in the above embodiments. The device embodiments correspond to the foregoing method embodiments. For the convenience of reading, the device embodiments of the present application will not repeat the detailed content in the foregoing method embodiments one by one. However, it should be clear that the detection device for hand key points in this embodiment can correspondingly implement all the content in the foregoing method embodiments.
[0110] Figure 5 The structural schematic diagram of the detection device for hand key points provided in an embodiment of the present application is as Figure 5 shown. The detection device 500 for hand key points provided in this embodiment includes:
[0111] An acquisition module 510, configured to acquire a set of hand sequence images corresponding to a hand image to be predicted, where the set of hand sequence images includes multiple consecutive hand sequence images before the hand image to be predicted;
[0112] An extraction module 520, configured to extract features from multiple hand sequence images in the set of hand sequence images to obtain first key point features of the hand key points with temporal information;
[0113] A processing module 520, configured to input the hand image to be predicted into a residual network in a key point detection model to obtain second key point features of hand key points in the hand image to be predicted;
[0114] A prediction module 530, configured to merge the first key point features and the second key point features to obtain fused features of the hand image to be predicted, and input the fused features into a fully connected network in the key point detection model to obtain position information of multiple hand key points in the hand image to be predicted.
[0115] As an optional implementation manner of an embodiment of this application, the extraction module 520 is specifically configured to determine the actual positions of the hand key points in each hand sequence image in a pixel coordinate system, a first offset feature of the hand key points in the X-axis direction, and a second offset feature of the hand key points in the Y-axis direction; perform a convolution operation on the actual positions of the hand key points, the first offset feature, and the second offset feature to obtain first key point features of the hand key points based on time series.
[0116] Figure 6 A structural schematic diagram of a detection device for hand key points provided in another embodiment of this application, as Figure 6 shown, on the basis of the detection device 500 for hand key points shown in Figure 5 shown, further includes:
[0117] A training module 610, configured to acquire a training data set, and train the key point detection model based on the training data set to obtain a trained key point detection model.
[0118] The training module is specifically configured to acquire multiple training images and a set of sequence images respectively corresponding to each of the multiple training images as a training data set, where the set of sequence images includes multiple consecutive sequence images before the training image.
[0119] As an optional implementation manner of the embodiment of the present application, the training module 610 is specifically configured to, for each training image, determine at least one relevant key point corresponding to each hand key point in the training image, where the at least one relevant key point includes one or more of the remaining hand key points in the training image other than the hand key points; input the training image and the sequence image set corresponding to the training image into the key point detection model, and obtain first position information of the hand key points and second position information of at least one relevant key point corresponding to the hand key points; adjust the parameters of the key point detection model based on the first position information and the second position information until the training end condition is met, so as to obtain a trained key point detection model.
[0120] As an optional implementation manner of the embodiment of the present application, the first position information includes the predicted positions of the hand key points in the training image in the pixel coordinate system, a first offset of the hand key points in the X-axis direction, and a second offset of the hand key points in the Y-axis direction; the second position information includes the predicted positions of at least one relevant key point corresponding to the hand key points in the pixel coordinate system, a third offset of the at least one relevant key point in the X-axis direction, and a fourth offset of the at least one relevant key point in the Y-axis direction; the training module 610 is specifically configured to determine weights corresponding to the predicted positions, the first offset, the second offset, the third offset, and the fourth offset of the hand key points in the pixel coordinate system; perform weighted summation on the predicted positions, the first offset, the second offset, the third offset, and the fourth offset through the weights to obtain an objective loss function; use the objective loss function as the loss function of the key point detection model for training until the training end condition is met, so as to obtain a trained key point detection model.
[0121] As an optional implementation manner of the embodiment of the present application, the training module 610 is specifically configured to, for each training image, group the hand key points in the training image to obtain a plurality of hand key point groups, and each hand key point group corresponds to a finger in the training image; for the hand key points on non-thumb fingers, determine at least one relevant key point corresponding to the hand key points in the hand key point group where the hand key points are located and the hand key point group corresponding to the thumb; for the hand key points on the thumb, determine at least one relevant key point corresponding to the hand key points in the hand key point group corresponding to the thumb.
[0122] As an optional implementation manner of the embodiment of the present application, the training module 610 is specifically configured to, for the hand key points on non-thumb fingers, if the hand key point is the root joint of a finger, determine at least one related key point corresponding to the hand key point in the hand key point group where the hand key point is located and the hand key point group corresponding to the thumb; if the hand key point is other joints except the root joint of a finger, determine at least one related key point corresponding to the hand key point in the hand key point group where the hand key point is located.
[0123] In one embodiment, an electronic device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of any one of the hand key point detection methods described in the above method embodiments are implemented.
[0124] Exemplarily, Figure 7 is a schematic structural diagram of the electronic device provided by the embodiment of the present application. As Figure 7 shown, the electronic device provided in this embodiment includes: a memory 71 and a processor 72. The memory 71 is used to store a computer program; the processor 72 is used to execute the steps in the hand key point detection method provided by the above method embodiment when calling the computer program. The implementation principle and technical effect are similar and will not be elaborated here. Those skilled in the art can understand that Figure 7 the structure shown in
[0125] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the hand key point detection methods described in the above method embodiments are implemented.
[0126] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static random access memory (SRAM) and dynamic random access memory (DRAM), etc.
[0127] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0128] For the sake of convenience of explanation, the above description has been made in conjunction with specific embodiments. However, the above discussion in some embodiments is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. According to the above teachings, various modifications and variations can be obtained. The selection and description of the above embodiments are for better explaining the principles and practical applications, so that those skilled in the art can better use the embodiments and various different modified embodiments suitable for specific use considerations.
Claims
1. A method for detecting hand key points, characterized in that, Including: Obtain a set of hand sequence images corresponding to the hand image to be predicted, where the set of hand sequence images includes multiple consecutive hand sequence images before the hand image to be predicted; Extract features from the multiple hand sequence images in the set of hand sequence images to obtain first key-point features of the hand key points with temporal information; Input the hand image to be predicted into the residual network in the key-point detection model to obtain second key-point features of the hand key points in the hand image to be predicted; Merge the first key-point features and the second key-point features to obtain fused features of the hand image to be predicted, and input the fused features into the fully-connected network in the key-point detection model to obtain position information of multiple hand key points in the hand image to be predicted.
2. The method according to claim 1, wherein The step of extracting features from the multiple hand sequence images in the set of hand sequence images to obtain first key-point features of the hand key points with temporal information includes: Determine the actual positions of the hand key points in each hand sequence image in the pixel coordinate system, a first offset feature of the hand key points in the X-axis direction, and a second offset feature of the hand key points in the Y-axis direction; Perform a convolution operation on the actual positions, the first offset feature, and the second offset feature of the hand key points in each hand sequence image to obtain first key-point features of the hand key points based on time series.
3. The method according to claim 1, characterized in that Before inputting the hand image to be predicted into the residual network in the key-point detection model to obtain second key-point features of the hand key points in the hand image to be predicted, the method further includes: Obtain a training data set, and train the key-point detection model based on the training data set to obtain a trained key-point detection model; The step of obtaining the training data set includes: Obtain multiple training images and a set of sequence images corresponding to each of the multiple training images as the training data set. Among them, the set of sequence images corresponding to each training image includes multiple consecutive sequence images before the training image.
4. The method according to claim 3, wherein The step of training the key-point detection model based on the training data set to obtain a trained key-point detection model includes: For each training image, determine at least one relevant key point corresponding to each hand key point in the training image, where the at least one relevant key point includes one or more of the remaining hand key points in the training image other than the hand key points; Input the training image and the set of sequence images corresponding to the training image into the key-point detection model to obtain first position information of the hand key points and second position information of at least one relevant key point corresponding to the hand key points; Adjust the parameters of the key-point detection model based on the first position information and the second position information until the training end condition is met, and obtain a trained key-point detection model.
5. The method according to claim 4, wherein The first position information includes the predicted positions of each hand key point in the training image in the pixel coordinate system, the first offset of the hand key point in the X-axis direction, and the second offset of the hand key point in the Y-axis direction; the second position information includes the predicted positions of at least one associated key point corresponding to the hand key point in the pixel coordinate system, the third offset of the at least one associated key point in the X-axis direction, and the fourth offset of the at least one associated key point in the Y-axis direction; Adjusting the parameters of the key point detection model based on the first position information and the second position information until the training end condition is met to obtain a trained key point detection model includes: Determining the weights corresponding to the predicted position, the first offset, the second offset, the third offset, and the fourth offset of the hand key point in the pixel coordinate system respectively; Performing weighted summation on the predicted position, the first offset, the second offset, the third offset, and the fourth offset through the weights to obtain a target loss function; Training the target loss function as the loss function of the key point detection model until the training end condition is met to obtain a trained key point detection model.
6. The method according to claim 4, characterized in that, For each training image, determining at least one associated key point corresponding to each hand key point in the training image includes: For each training image, grouping the hand key points in the training image to obtain a plurality of hand key point groups, and each hand key point group corresponds to a finger in the training image; For the hand key points on non-thumb fingers, determining at least one associated key point corresponding to the hand key point in the hand key point group where the hand key point is located and the hand key point group corresponding to the thumb; For the hand key points on the thumb, determining at least one associated key point corresponding to the hand key point in the hand key point group corresponding to the thumb.
7. The method according to claim 6, characterized in that, For the hand key points on non-thumb fingers, determining at least one associated key point corresponding to the hand key point in the hand key point group where the hand key point is located and the hand key point group corresponding to the thumb includes: For the hand key points on non-thumb fingers, if the hand key point is the root joint of a finger, determining at least one associated key point corresponding to the hand key point in the hand key point group where the hand key point is located and the hand key point group corresponding to the thumb; If the hand key point is other joints except the root joint of a finger, determining at least one associated key point corresponding to the hand key point in the hand key point group where the hand key point is located.
8. A detection device for hand key points, characterized in that, Includes: An acquisition module for acquiring a set of hand sequence images corresponding to a hand image to be predicted, where the set of hand sequence images includes a plurality of consecutive hand sequence images before the hand image to be predicted; An extraction module for extracting features from the plurality of hand sequence images in the set of hand sequence images to obtain first key point features of the hand key points with temporal information; A processing module, configured to input the hand image to be predicted into a residual network in a key point detection model, and obtain second key point features of hand key points in the hand image to be predicted; A prediction module, configured to merge the first key point features and the second key point features to obtain fused features of the hand image to be predicted, and input the fused features into a fully connected network in the key point detection model to obtain position information of multiple hand key points in the hand image to be predicted.
9. An electronic device, comprising: A memory and a processor, the memory stores a computer program, wherein the processor, when executing the computer program, implements the method for detecting hand key points according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the method for detecting hand key points according to any one of claims 1 to 7 is implemented.
11. A vehicle, characterized in that, The vehicle is configured with the hand key point detection device according to claim 8, or the electronic device according to claim 9, or the storage medium according to claim 10.