Key point detection model training method, device, electronic device and storage medium

By using a training data set containing the original and interfering sample images, the key point detection model is trained, which solves the problem of low recognition accuracy of the model in complex interfering environments, and achieves more stable and accurate key point detection, improving the user experience.

CN113139564BActive Publication Date: 2025-05-23TSINGHUA UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010066238.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-01-20
Publication Date
2025-05-23
Estimated Expiration
2040-01-20

AI Technical Summary

Technical Problem

When facing complex interference environments, the existing key point detection model has low recognition accuracy, resulting in severe jitter in the target recognition image and affecting the user experience.

Method used

The key point detection model is trained by obtaining a collection of training sample images containing the original sample image and the interfering sample image. Specific methods include: detecting predicted key points for each training sample image and generating a predicted heat map; adjusting the model weight parameters when the predicted heat map does not match the real heat map; and outputting the trained model when the convergence conditions are met.

Benefits of technology

It improves the stability and recognition accuracy of the key point detection model, reduces the jitter phenomenon of the recognition image, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113139564B_ABST
    Figure CN113139564B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a training method, device, electronic device and storage medium for a key point detection model. The method includes: each time a training sample image is read, a key point recognition model is used to obtain corresponding predicted key points, and corresponding predicted heat maps are generated respectively, wherein each time a predicted key point is read, a corresponding layer is generated, and based on the relative distance between each pixel point on the layer and the predicted key point, the predicted response value of each pixel point is recorded on the layer to obtain a corresponding predicted heat map, and when it is determined that the predicted heat map corresponding to at least one predicted key point does not match the corresponding real heat map, the corresponding weight parameter in the key point detection model is adjusted; when it is determined that the preset convergence condition is met, the trained key point detection model is output. The model is trained by using the method of determining the predicted response value of the pixel point by relative distance to improve the stability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a training method, device, electronic device and storage medium for a key point detection model. Background Art

[0002] With the development of science and technology, key point detection models are applied in real life to provide people with corresponding application services. For example, when a user shoots a video, a gesture recognition model is used to determine the user's target gesture image, and special effects are added to the target gesture image to improve playability; for another example, an object recognition model is used in a search website to identify the image to be detected input by the user and determine the object image in the image to be detected.

[0003] Since the trained key point detection model is very sensitive to small changes in the input image, the output target recognition image will have serious jitter, which greatly reduces the user experience. In order to solve the problem of low model stability, Gaussian noise operation is performed on the original sample image set, and then the key point detection model is trained based on the original sample image set after the Gaussian noise operation.

[0004] However, when the above-trained key point detection model is actually applied, the following problems may arise: Gaussian noise has limited interference on the image, and the original sample image set after the Gaussian noise operation cannot fully reflect the complex interference environment in actual applications. Therefore, the key point detection model obtained after training is still unable to accurately identify images with multiple interference factors added, and the target recognition image will still experience serious jitter, resulting in a poor user experience.

[0005] In view of this, it is necessary to design a new training method, device, electronic device and storage medium for a key point detection model to overcome the above-mentioned defects. Summary of the invention

[0006] The present invention provides a training method, device, electronic device and storage medium for a key point detection model, which are used to improve the stability of the model.

[0007] The specific technical solutions provided by this disclosure are as follows:

[0008] According to a first aspect of an embodiment of the present disclosure, a method for training a key point detection model is provided, comprising:

[0009] Acquire a preset training sample image set, wherein the training sample image set includes an original sample image and an interference sample image corresponding to the original sample image;

[0010] Each training sample image is read from the training sample image set in sequence, wherein each time a training sample image is read, the following operations are performed:

[0011] Using the key point detection model, each predicted key point is detected from the one training sample image, and corresponding predicted heat maps are generated based on the predicted key points, wherein each time a predicted key point is read, a corresponding layer is generated, based on the relative distance between each pixel point in the one layer and the one predicted key point, the predicted response value corresponding to each pixel point is determined and recorded in the one layer, and the processed one layer is determined as a predicted heat map;

[0012] Determining that a predicted heat map corresponding to at least one predicted key point does not match a corresponding real heat map, adjusting a weight parameter corresponding to the at least one predicted key point in the key point detection model accordingly;

[0013] When it is determined that the key point detection model meets the preset convergence condition, a trained key point detection model is obtained.

[0014] Optionally, further including:

[0015] An interference sample image is an image obtained by adding a set interference factor to a corresponding original sample image, and an original sample image corresponds to at least two interference sample images.

[0016] Optionally, determining a predicted response value corresponding to a pixel point in the layer based on a relative distance between the pixel point and the predicted key point includes:

[0017] Calculating the relative distance between the one pixel point and the one predicted key point;

[0018] When the relative distance does not reach the set distance threshold, determining a predicted response value of the one pixel based on the relative distance;

[0019] When it is determined that the relative distance is greater than the set distance threshold, the predicted response value of the one pixel is determined to be zero.

[0020] Optionally, calculating the relative distance between the one pixel point and the one predicted key point includes:

[0021] Obtaining the coordinates of the one pixel point and the coordinates of the one predicted key point;

[0022] Using a preset chessboard distance function, based on each of the obtained coordinates, calculate the relative distance between the one pixel point and the one predicted key point;

[0023] The step of determining the predicted response value of the pixel point based on the relative distance includes:

[0024] Based on the relative distance and a preset attenuation coefficient, a predicted response value of the one pixel is calculated.

[0025] Optionally, determining that a predicted heat map corresponding to a predicted key point does not match the corresponding true heat map includes:

[0026] Determine the error between the predicted key point and the corresponding true key point based on the predicted response value of each pixel point in the predicted heat map and the corresponding true response value;

[0027] When the error is higher than a set threshold, it is determined that the predicted heat map corresponding to the one predicted key point does not match the corresponding true heat map.

[0028] Optionally, determining the error between the predicted key point and the corresponding true key point based on the predicted response value of each pixel point in the predicted heat map and the corresponding true response value includes:

[0029] Using a multi-label classification algorithm, determine the weight of each pixel in the original sample image corresponding to the predicted heat map;

[0030] A weighted cross entropy loss function is used to determine the error between the predicted key point and the corresponding true key point by combining the predicted response value of each pixel point in the predicted heat map, the true response value of each pixel point in the original sample image corresponding to the predicted heat map, and the weight.

[0031] Optionally, when it is determined that the key point detection model satisfies a preset convergence condition, outputting the trained key point detection model includes:

[0032] When all training sample images are read, it is determined that the key point detection model training is completed; or,

[0033] When the recognition accuracy of the key point detection model reaches a set threshold value, it is determined that the training of the key point detection model is completed.

[0034] According to a second aspect of an embodiment of the present disclosure, a method for key point detection is provided, including:

[0035] Inputting the acquired image to be detected into the key point detection model generated by the method of the first aspect to generate a plurality of heat maps of key points to be detected;

[0036] In each of the obtained key point heat maps to be detected, the pixel point corresponding to the maximum predicted response value is determined as the predicted key point output by the corresponding key point heat map to be detected, wherein the predicted response value of a pixel point represents the probability value of the pixel point being the predicted key point.

[0037] According to a third aspect of an embodiment of the present disclosure, a training device for a key point detection model is provided, comprising:

[0038] An acquisition unit is configured to acquire a preset training sample image set, wherein the training sample image set includes an original sample image and an interference sample image corresponding to the original sample image;

[0039] The processing unit is configured to read each training sample image from the training sample image set in sequence, wherein each time a training sample image is read, the following operations are performed:

[0040] Using the key point detection model, each predicted key point is detected from the one training sample image, and corresponding predicted heat maps are generated based on the predicted key points, wherein each time a predicted key point is read, a corresponding layer is generated, based on the relative distance between each pixel point in the one layer and the one predicted key point, the predicted response value corresponding to each pixel point is determined and recorded in the one layer, and the processed one layer is determined as a predicted heat map;

[0041] Determining that a predicted heat map corresponding to at least one predicted key point does not match a corresponding real heat map, adjusting a weight parameter corresponding to the at least one predicted key point in the key point detection model accordingly;

[0042] The output unit is configured to output the trained key point detection model when it is determined that the key point detection model meets the preset convergence condition.

[0043] Optionally, it is further configured as:

[0044] An interference sample image is an image obtained by adding a set interference factor to a corresponding original sample image, and an original sample image corresponds to at least two interference sample images.

[0045] Optionally, based on a relative distance between a pixel point in the layer and the prediction key point, the prediction response value corresponding to the pixel point is determined, and the processing unit is configured to:

[0046] Calculating the relative distance between the one pixel point and the one predicted key point;

[0047] When the relative distance does not reach the set distance threshold, determining a predicted response value of the one pixel based on the relative distance;

[0048] When it is determined that the relative distance is greater than the set distance threshold, the predicted response value of the one pixel is determined to be zero.

[0049] Optionally, the calculating the relative distance between the one pixel point and the one predicted key point, the processing unit is configured to:

[0050] Obtaining the coordinates of the one pixel point and the coordinates of the one predicted key point;

[0051] Using a preset chessboard distance function, based on each of the obtained coordinates, calculate the relative distance between the one pixel point and the one predicted key point;

[0052] The step of determining the predicted response value of the one pixel based on the relative distance, wherein the processing unit is configured to:

[0053] Based on the relative distance and a preset attenuation coefficient, a predicted response value of the one pixel is calculated.

[0054] Optionally, it is determined that a predicted heat map corresponding to a predicted key point does not match a corresponding real heat map, and the processing unit is configured to:

[0055] Determine the error between the predicted key point and the corresponding true key point based on the predicted response value of each pixel point in the predicted heat map and the corresponding true response value;

[0056] When the error is higher than a set threshold, it is determined that the predicted heat map corresponding to the one predicted key point does not match the corresponding true heat map.

[0057] Optionally, based on the predicted response values ​​of each pixel point in the predicted heat map and the corresponding true response values, the error between the predicted key point and the corresponding true key point is determined, and the processing unit is configured to:

[0058] Using a multi-label classification algorithm, determine the weight of each pixel in the original sample image corresponding to the predicted heat map;

[0059] A weighted cross entropy loss function is used to determine the error between the predicted key point and the corresponding true key point by combining the predicted response value of each pixel point in the predicted heat map, the true response value of each pixel point in the original sample image corresponding to the predicted heat map, and the weight.

[0060] Optionally, when it is determined that the key point detection model satisfies a preset convergence condition, the trained key point detection model is output, and the output unit is configured to:

[0061] When all training sample images are read, it is determined that the key point detection model training is completed; or,

[0062] When the recognition accuracy of the key point detection model reaches a set threshold value, it is determined that the training of the key point detection model is completed.

[0063] According to a fourth aspect of an embodiment of the present disclosure, there is provided a key point detection device, including:

[0064] A generating unit is configured to input the acquired image to be detected into a key point detection model generated by the method of the first aspect to generate a plurality of heat maps of key points to be detected;

[0065] The detection unit is configured to determine the pixel point corresponding to the maximum predicted response value in each obtained key point heat map to be detected as the predicted key point output by the corresponding key point heat map to be detected, wherein the predicted response value of a pixel point represents the probability value of the pixel point being the predicted key point.

[0066] According to a fifth aspect of an embodiment of the present disclosure, there is provided an electronic device, including:

[0067] A memory for storing executable instructions;

[0068] A processor is used to read and execute the executable instructions stored in the memory to implement any of the above methods.

[0069] According to a sixth aspect of an embodiment of the present disclosure, a storage medium is provided, which enables the execution of steps of any one of the above methods when instructions in the storage medium are executed by a processor.

[0070] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects:

[0071] In the embodiment of the present disclosure, each training sample image is read in sequence, and each time a training sample image is read, the corresponding predicted key points are obtained by using the key point recognition model, and the corresponding predicted heat maps are generated respectively, wherein each time a predicted key point is read, a corresponding layer is generated, and based on the relative distance between each pixel point on the layer and the predicted key point, the predicted response value of each pixel point is recorded on the layer to obtain the corresponding predicted heat map, and when it is determined that the predicted heat map corresponding to at least one predicted key point does not match the corresponding real heat map, the corresponding weight parameter in the key point detection model is adjusted; when it is determined that the preset convergence condition is met, the trained key point detection model is output. The embodiment of the present disclosure uses relative distance to determine the predicted response value of the pixel point. When the relative distance between other pixel points around the predicted key point and the predicted key point is close, the predicted response value of the other pixel points obtained by calculation is also greatly different from the predicted response value of the predicted key point. Therefore, the heat map to be detected output by the embodiment of the present disclosure has a high discrimination characteristic compared with the traditional predicted heat map. Using the above method to train the key point detection model can improve the model stability, thereby improving the detection accuracy and user experience.

[0072] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0074] Figure 1 is a schematic diagram of a process of training a key point detection model according to an exemplary embodiment;

[0075] Figure 2a is a diagram showing an original gesture image according to an exemplary embodiment;

[0076] Figure 2b A first type of layer is shown according to an exemplary embodiment;

[0077] Figure 3a is a diagram showing an interfering clover image according to an exemplary embodiment;

[0078] Figure 3b A second type of layer is shown according to an exemplary embodiment;

[0079] Figure 4a is a diagram showing a conventional prediction heat map according to an exemplary embodiment;

[0080] Figure 4b is a prediction heat map showing high discrimination according to an exemplary embodiment;

[0081] Figure 5 is a prediction heat map showing a predicted gesture key point according to an exemplary embodiment;

[0082] Figure 6 is a prediction heat map showing a predicted clover key point according to an exemplary embodiment;

[0083] Figure 7 is a schematic diagram showing a process of key point detection according to an exemplary embodiment;

[0084] Figure 8 is a block diagram of a training device for a key point detection model according to an exemplary embodiment;

[0085] Fig. 9 is a block diagram of a key point detection device according to an exemplary embodiment;

[0086] Fig.10 The diagram is a schematic structural diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0087] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.

[0088] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0089] In order to improve the stability of the model, a solution is provided in an embodiment of the present disclosure, which is: each time a training sample image is read, a key point recognition model is used to obtain corresponding predicted key points, and corresponding predicted heat maps are generated respectively, wherein each time a predicted key point is read, a corresponding layer is generated, and based on the relative distance between each pixel point on the layer and the predicted key point, the predicted response value of each pixel point is recorded on the layer to obtain a corresponding predicted heat map, and when it is determined that the predicted heat map corresponding to at least one predicted key point does not match the corresponding real heat map, the corresponding weight parameters in the key point detection model are adjusted; when it is determined that the preset convergence conditions are met, the trained key point detection model is output.

[0090] The preferred embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.

[0091] Before using the key point detection model to perform recognition operations, the key point detection model needs to be trained first. Figure 1 As shown, in the embodiment of the present disclosure, the training process of the key point detection model is as follows:

[0092] S101: Obtain a preset training sample image set, wherein the training sample image set includes an original sample image and an interference sample image corresponding to the original sample image.

[0093] Optionally, an interference sample image is an image obtained by adding a set interference factor to a corresponding original sample image, and an original sample image corresponds to at least two interference sample images. The set interference factor can be a combination of any two or more of the following interference factors: adding Gaussian noise, increasing / decreasing brightness, increasing / decreasing contrast, increasing / decreasing saturation, increasing / decreasing grayscale value, model, shadow, sharpening. Therefore, the original sample image and the corresponding interference sample image have the same size, that is, the two sample images contain the same number of pixels.

[0094] For example, Gaussian noise is added to the original sample image 1 to obtain the corresponding interference sample image 1; Gaussian noise is added to the original sample image 1 and the gray value is reduced at the same time to obtain the corresponding interference sample image 2.

[0095] One original sample image corresponds to at least two interference sample images, which can fully simulate the complex interference environment in practical applications. Using the above training sample image set to train the key point detection model is conducive to improving the anti-interference ability of the model, improving the stability of the model, and reducing the jitter phenomenon of the output recognition image.

[0096] S102: Read a training sample image from the training sample image set.

[0097] In the disclosed embodiment, when reading each training sample image from the training sample image set in sequence, they can be read in the order of original sample image + interference sample image, that is, after each original sample image is read to train the key point detection model, the associated interference sample image will continue to be read to further train the key point detection model. In this way, the key point detection model can be targeted and concentratedly trained based on the original scene + interference scene, thereby effectively improving the training accuracy and training efficiency.

[0098] In the subsequent embodiments, the training sample images are used as the operation objects for unified description, and the original sample images and the interference sample images are distinguished when giving examples.

[0099] S103: Using a key point detection model, detect and obtain various predicted key points from a training sample image.

[0100] For example, if an original sample image is read, the key point detection model is used to detect and obtain the corresponding first-category predicted key points;

[0101] For another example, if an interference sample image is read, the key point detection model is used to perform detection to obtain the corresponding second-category predicted key points.

[0102] S104: Read a predicted key point and generate a corresponding layer.

[0103] Specifically, after reading a predicted key point from a training sample image, a layer separation method can be used to generate a layer corresponding to the predicted key point based on the training sample image and the predicted key point, wherein all pixel points in the layer include the above-mentioned predicted key point.

[0104] On the other hand, the purpose of generating layers by layer separation is to ensure that the size of the layers is consistent with the training sample image, and to ensure that the number of pixels contained in the layers is consistent with the number of pixels contained in the training sample image.

[0105] If the original sample image is read, the first type of predicted key points are read, and the generated layer is also called the first type of layer. Figure 2a The original gesture image shown, Figure 2a A black dot in the image represents a predicted gesture key point of the original gesture image, and a predicted gesture key point is read to generate a first-class layer such as Figure 2b As shown;

[0106] If the interference sample image is read, the second type of prediction key points are read, and the generated layer is also called the second type of layer. Figure 3a The interfering clover image shown, Figure 3a A black dot in the image is represented as a predicted clover key point that interferes with the clover image, and a black snowflake intersection is represented as a set interference factor. Then a predicted clover key point is read, and a second type layer is generated as follows Figure 3b shown.

[0107] S105: Based on the relative distance between each pixel point in the one layer and the one prediction key point, determine the prediction response value corresponding to each pixel point and record it in the one layer, and determine the processed one layer as a prediction heat map.

[0108] Optionally, any pixel point (hereinafter referred to as pixel point P 1 ) as an example, the calculation process of the predicted response value is introduced as follows:

[0109] A1. Calculate pixel point P 1 The relative distance between the predicted key point and the

[0110] In the embodiments of the present disclosure, a chessboard distance function, a Euclidean distance function, a block distance function or other types of distance functions may be used to calculate the above relative distance. For ease of understanding, the following embodiments use the chessboard distance function as an example to calculate the above relative distance, but are not limited to this implementation method.

[0111] First, get a pixel point P 1 The coordinates (x 1 ,y 1 ), and a predicted keypoint M 1 The coordinates (x 2 ,y 2 );

[0112] Then, the preset chessboard distance function formula (1) is used to calculate the pixel point P based on the obtained coordinates. 1 and predicted key points M 1 relative distance.

[0113] Distance(P 1 , M 1 )=max(|x 1 -x 2 |,|y 1 -y 2 |) Formula (1);

[0114] Finally, using formula (2), based on the relative distance and the preset attenuation coefficient, the pixel point P is calculated 1 The predicted response value of .

[0115] T=α n Formula (2);

[0116] In formula (2), α is the preset attenuation coefficient, α∈[0,1]; n is the pixel point P 1 And the predicted key point M 1 The relative distance between them.

[0117] A2. When the relative distance does not reach the set distance threshold, the pixel point P can be determined based on the relative distance. 1 The predicted response value, accordingly, when the relative distance is greater than the set distance threshold, it can be determined that the pixel point P 1 The predicted response value is zero.

[0118] Specifically, pixel P 1 And the predicted key point M 1 The farther the relative distance between them, the closer the pixel point P is. 1 The lower the probability value of the target key point, that is, the pixel point P 1 The lower the predicted response value, the more redundant the data is. Therefore, in order to discard redundant data, when n is greater than the set distance threshold, the corresponding predicted response value is determined to be 0.

[0119] The above only takes an arbitrary pixel point as an example to introduce the generation process of the predicted response value. In the embodiment of the present disclosure, the corresponding predicted response value can be calculated in the same way for each pixel point in a training sample image, which will not be repeated here.

[0120] And each pixel point (including the above-mentioned predicted key point M 1 ) are recorded in the one layer, the processed layer can be determined as a prediction heat map.

[0121] In the traditional prediction heat map, the difference between the predicted response value of the prediction key point and the predicted response values ​​of other pixels around it is very small, resulting in many pixels in the prediction heat map that meet the set response value threshold. Figure 4a shown.

[0122] In the embodiment of the present disclosure, the pixel point P 1 The predicted response value is the nth power of the attenuation coefficient, where n is the pixel point P 1 And the predicted key point M 1 The relative distance between them, since the attenuation coefficient ranges from [0,1], when predicting the key point M 1Other surrounding pixels and predicted key points M 1 When the relative distance between them is close, the predicted response value obtained by calculation is also close to the predicted key point M 1 The predicted response values ​​of 1 The predicted response value of is much higher than the predicted response values ​​of other pixels, generating a prediction heat map with high discrimination. Figure 4b In this way, when the trained key point detection model detects images with slight changes, it can reduce the jitter of the predicted key points and improve the model stability and user experience.

[0123] For example, when reading each first-class prediction key point, a corresponding first-class prediction heat map is generated for each prediction key point. Figure 2a The original gesture image shown in the figure uses the layer separation method to obtain a predicted heat map of one of the predicted gesture key points. For details, see Figure 5 shown.

[0124] For another example, when reading each second-category prediction key point, the above method can also be used to generate a corresponding second-category prediction heat map for each second-category prediction key point. Figure 3a The interference clover image shown in the figure uses the layer separation method to obtain a predicted heat map of one of the predicted clover key points. For details, see Figure 6 shown.

[0125] S106: Determine whether all predicted key points have been read, if so, execute step 107; otherwise, return to step 104.

[0126] S107: Determine whether the predicted heat maps corresponding to each predicted key point match the corresponding real heat maps. If so, execute step 109; otherwise, execute step 108.

[0127] Before executing step 107, it is necessary to annotate each original sample image, determine each real key point, and use a layer separation method to generate a real heat map corresponding to each real key point, wherein a real heat map records the real response value of each identified pixel point, and each pixel point contains a real key point.

[0128] Optionally, taking any prediction key point as an example, when determining whether the prediction heat map corresponding to any prediction key point matches the corresponding real heat map, the specific calculation process is as follows:

[0129] B1. Based on the predicted response value of each pixel in the predicted heat map and the corresponding true response value, determine the error between a predicted key point and the corresponding true key point.

[0130] Optionally, the error between a predicted key point and the corresponding true key point is determined, and the specific calculation process is as follows:

[0131] First, a multi-label classification algorithm is used to determine the weight of each pixel in the original sample image corresponding to the predicted heat map.

[0132] In order to reduce the difficulty of training and facilitate model convergence, a multi-label classification method is adopted in the embodiment of the present disclosure to assign a weight to each pixel in the original sample image. Since the true response value of the pixel is negatively correlated with the relative distance between the pixel and the true key point, in order to discard redundant data, when the relative distance does not exceed the set distance threshold, the weight assigned to the pixel is 1; when the relative distance is greater than the set distance threshold, the weight assigned to the pixel is 0.

[0133] Secondly, the weighted cross entropy loss function is used to determine the error between a predicted key point and the corresponding true key point by combining the predicted response value of each pixel in the predicted heat map and the true response value and weight of each pixel in the original sample image corresponding to the predicted heat map.

[0134] Specifically, the embodiment of the present disclosure uses the weighted cross entropy loss function shown in formula (3) to calculate the error between a predicted key point and the corresponding true key point.

[0135] Loss w =-∑ (i,j)∈I w i,i y i,j logp i,j Formula (3);

[0136] Where I represents the predicted heat map, (i, j) represents the pixel point in the i-th row and j-th column on the predicted heat map; p i,j Characterizes the predicted response value of the pixel; w i,j It represents the true response value of the pixel point at the i-th row and j-th column recorded in the original sample image. The true response value represents the probability value of the pixel point at the i-th row and j-th column being the true key point.

[0137] B2. When the error is higher than a set threshold, it is determined that the predicted heat map corresponding to the one predicted key point does not match the corresponding real heat map; when the error does not exceed the set threshold, it is determined that the predicted heat map corresponding to the one predicted key point matches the corresponding real heat map.

[0138] Therefore, the above method can be used to determine whether each first-class predicted heat map matches the corresponding real heat map; similarly, the above method can also be used to determine whether each second-class predicted heat map matches the corresponding real heat map.

[0139] S108: Adjust the weight parameter corresponding to at least one predicted key point in the key point detection model accordingly, and execute step 109.

[0140] Specifically, when adjusting the weight parameters, it can be divided into the following two cases:

[0141] Case 1: When the errors corresponding to the prediction key points are higher than the set threshold, the weight parameters corresponding to the prediction key points are adjusted accordingly.

[0142] In case 2, when the partial errors corresponding to several prediction key points are higher than the set threshold, the weight parameters corresponding to the corresponding prediction key points are adjusted accordingly.

[0143] For example, among 21 errors, only the 3rd to 5th errors are higher than the set threshold, so only the weight parameters corresponding to the 3rd to 5th prediction key points are adjusted accordingly.

[0144] Specifically, taking the error corresponding to any one of the prediction key points as an example, which is higher than the set threshold, the specific process of adjusting the weight parameter corresponding to any one of the prediction key points is as follows:

[0145] Firstly, the chain rule of composite functions is used to derive each neuron corresponding to a predicted key point in the key point detection model based on the error between the predicted key point and the corresponding true key point;

[0146] Secondly, the derivatives of each neuron are negated respectively, and each negated derivative is multiplied by the set step size respectively, and each processed derivative is added to the initial weight parameter corresponding to the one prediction key point to obtain the adjusted weight parameter.

[0147] S109: Determine whether the key point detection model meets the preset convergence condition. If so, execute step 110; otherwise, return to step 102.

[0148] The training of the key point detection model can be stopped if any of the following conditions are met:

[0149] When all training sample images have been read, it is determined that the key point detection model training is completed;

[0150] When the recognition accuracy of the key point detection model reaches the set threshold value, it is determined that the key point detection model training is completed.

[0151] S110: Output the trained key point detection model.

[0152] The embodiment of the present disclosure determines the predicted response value of a pixel point based on relative distance. During the calculation process, since the preset attenuation coefficient has a value range of [0,1], when the relative distance between other pixel points around the predicted key point and the predicted key point is close, the predicted response value obtained by calculation is also greatly different from the predicted response value of the predicted key point. Therefore, the heat map to be detected output by the embodiment of the present disclosure has a high discrimination characteristic compared to the traditional prediction heat map. Using the above method to train the key point detection model can improve the model stability.

[0153] Furthermore, a training sample image set including original sample images and interference sample images is used to train the key point detection model. In this way, the trained model can not only detect the predicted key points in images without adding interference factors, but also identify the predicted key points in images with adding multiple interference factors, thereby improving the detection accuracy.

[0154] See also Figure 7 As shown, a method for key point detection is provided in an embodiment of the present disclosure, and the specific process is as follows:

[0155] Step 701: input the acquired image to be detected into the key point detection model trained by the above method to generate multiple heat maps of key points to be detected.

[0156] Step 702: In each of the obtained key point heat maps to be detected, the pixel point corresponding to the maximum predicted response value is determined as the predicted key point output by the corresponding key point heat map to be detected, wherein the predicted response value of a pixel point represents the probability value of the pixel point being the predicted key point.

[0157] The disclosed embodiment determines the predicted response value of a pixel point based on relative distance. During the calculation process, since the preset attenuation coefficient has a value range of [0,1], when the relative distance between other pixel points around the predicted key point and the predicted key point is close, the predicted response value obtained by calculation is also greatly different from the predicted response value of the predicted key point. Therefore, the heat map to be detected output by the disclosed embodiment has a high discrimination characteristic compared to the traditional prediction heat map, and the pixel point corresponding to the maximum predicted response value can be directly selected as the predicted key point. In this way, not only the model stability is improved, but also the detection efficiency and user experience are improved.

[0158] Based on the above embodiments, see Figure 8 As shown, in an embodiment of the present disclosure, a training device 800 for a key point detection model is provided, which at least includes an acquisition unit 801, a processing unit 802 and an output unit 803, wherein:

[0159] The acquisition unit 801 is configured to acquire a preset training sample image set, wherein the training sample image set includes an original sample image and an interference sample image corresponding to the original sample image;

[0160] The processing unit 802 is configured to read each training sample image from the training sample image set in sequence, wherein each time a training sample image is read, the following operations are performed:

[0161] Using the key point detection model, each predicted key point is identified from the one training sample image, and corresponding predicted heat maps are generated based on the predicted key points, wherein each time a predicted key point is read, a corresponding layer is generated, based on the relative distance between each pixel point in the one layer and the one predicted key point, the predicted response value corresponding to each pixel point is determined and recorded in the one layer, and the one layer is determined as a predicted heat map;

[0162] Determining that a predicted heat map corresponding to at least one predicted key point does not match a corresponding real heat map, adjusting a weight parameter corresponding to the at least one predicted key point in the key point detection model accordingly;

[0163] The output unit 803 is configured to output the trained key point detection model when it is determined that the key point detection model meets the preset condition for stopping training.

[0164] Optionally, it is further configured as:

[0165] An interference sample image is an image obtained by adding a set interference factor to a corresponding original sample image, and an original sample image corresponds to at least two interference sample images.

[0166] Optionally, based on a relative distance between a pixel point in the layer and the prediction key point, the prediction response value corresponding to the pixel point is determined, and the processing unit 802 is configured to:

[0167] Calculating the relative distance between the one pixel point and the one predicted key point;

[0168] When it is determined that the relative distance does not reach the set distance threshold, determining a predicted response value of the one pixel based on the relative distance;

[0169] When it is determined that the relative distance is greater than the set distance threshold, the predicted response value of the one pixel is determined to be zero.

[0170] Optionally, the calculating the relative distance between the one pixel point and the one predicted key point, the processing unit 802 is configured to:

[0171] Obtaining the coordinates of the one pixel point and the coordinates of the one predicted key point;

[0172] Using a preset chessboard distance function, based on each of the obtained coordinates, calculate the relative distance between the one pixel point and the one predicted key point;

[0173] The step of determining the predicted response value of the pixel point based on the relative distance, the processing unit 802 is configured to:

[0174] Based on the relative distance and a preset attenuation coefficient, a predicted response value of the one pixel is calculated.

[0175] Optionally, it is determined that a predicted heat map corresponding to a predicted key point does not match a corresponding real heat map, and the processing unit 802 is configured to:

[0176] Determine the error between the predicted key point and the corresponding true key point based on the predicted response value of each pixel point in the predicted heat map and the corresponding true response value;

[0177] When it is determined that the error is higher than a set threshold, it is determined that the predicted heat map corresponding to the one predicted key point does not match the corresponding true heat map.

[0178] Optionally, based on the predicted response value of each pixel point in the predicted heat map and the corresponding true response value, the error between the predicted key point and the corresponding true key point is determined, and the processing unit 802 is configured to:

[0179] Using a multi-label classification algorithm, determine the weight of each pixel in the original sample image corresponding to the predicted heat map;

[0180] A weighted cross entropy loss function is used to determine the error between the predicted key point and the corresponding true key point by combining the predicted response value of each pixel point in the predicted heat map, the true response value of each pixel point in the original sample image corresponding to the predicted heat map, and the weight.

[0181] Optionally, when it is determined that the key point detection model satisfies a preset convergence condition, the trained key point detection model is output, and the output unit 803 is configured to:

[0182] When all training sample images are read, it is determined that the key point detection model training is completed; or,

[0183] When the recognition accuracy of the key point detection model based on the training sample image set reaches a set threshold value, it is determined that the training of the key point detection model is completed.

[0184] Based on the above embodiments, see Fig. 9 As shown, in an embodiment of the present disclosure, a key point detection device 900 is provided, which at least includes a generation unit 901 and a detection unit 902, wherein:

[0185] A generating unit 901 is configured to input the acquired image to be detected into a key point detection model generated by the method of the first aspect to generate a plurality of heat maps of key points to be detected;

[0186] The detection unit 902 is configured to determine the pixel point corresponding to the maximum predicted response value in each obtained heat map of the key point to be detected as the predicted key point output by the corresponding heat map of the key point to be detected, wherein the predicted response value of a pixel point represents the probability value of the pixel point being the predicted key point.

[0187] Based on the above embodiments, see Fig.10 As shown, in an embodiment of the present disclosure, an electronic device is provided, which includes at least a memory 1001 and a processor 1002, wherein:

[0188] Memory 1001, used for storing executable instructions;

[0189] The processor 1002 is configured to obtain a preset training sample image set, wherein the training sample image set includes an original sample image and an interference sample image set corresponding to the original sample image;

[0190] Each training sample image is read from the training sample image set in sequence, wherein each time a training sample image is read, the following operations are performed:

[0191] Using the key point detection model, each predicted key point is identified from the one training sample image, and corresponding predicted heat maps are generated based on the predicted key points, wherein each time a predicted key point is read, a corresponding layer is generated, based on the relative distance between each pixel point in the one layer and the one predicted key point, the predicted response value corresponding to each pixel point is determined and recorded in the one layer, and the processed one layer is determined as a predicted heat map;

[0192] Determining that a predicted heat map corresponding to at least one predicted key point does not match a corresponding real heat map, adjusting a weight parameter corresponding to the at least one predicted key point in the key point detection model accordingly;

[0193] When it is determined that the key point detection model meets the preset convergence condition, outputting the trained key point detection model;

[0194] A power supply component 1003, used to provide electrical energy;

[0195] The communication component 1004 is used to implement the communication function.

[0196] Optionally, the processor 1002 is further configured to:

[0197] An interference sample image is an image obtained by adding a set interference factor to a corresponding original sample image, and an original sample image corresponds to at least two interference sample images.

[0198] Optionally, based on a relative distance between a pixel point in the layer and the prediction key point, determining a prediction response value corresponding to the pixel point, the processor 1002 is configured to:

[0199] Calculating the relative distance between the one pixel point and the one predicted key point;

[0200] When the relative distance does not reach the set distance threshold, determining a predicted response value of the one pixel based on the relative distance;

[0201] When it is determined that the relative distance is greater than the set distance threshold, the predicted response value of the one pixel is determined to be zero.

[0202] Optionally, the calculating the relative distance between the one pixel point and the one predicted key point, the processor 1002 is configured to:

[0203] Obtaining the coordinates of the one pixel point and the coordinates of the one predicted key point;

[0204] Using a preset chessboard distance function, based on each of the obtained coordinates, calculate the relative distance between the one pixel point and the one predicted key point;

[0205] The processor 1002 is used to determine the predicted response value of the pixel point based on the relative distance:

[0206] Based on the relative distance and a preset attenuation coefficient, a predicted response value of the one pixel is calculated.

[0207] Optionally, it is determined that a predicted heat map corresponding to a predicted key point does not match a corresponding real heat map, and the processor 1002 is configured to:

[0208] Determine the error between the predicted key point and the corresponding true key point based on the predicted response value of each pixel point in the predicted heat map and the corresponding true response value;

[0209] When it is determined that the error is higher than a set threshold, it is determined that the predicted heat map corresponding to the one predicted key point does not match the corresponding true heat map.

[0210] Optionally, based on the predicted response values ​​of each pixel in the predicted heat map and the corresponding true response values, the error between the predicted key point and the corresponding true key point is determined, and the processor 1002 is used to:

[0211] Using a multi-label classification algorithm, determine the weight of each pixel in the original sample image corresponding to the predicted heat map;

[0212] A weighted cross entropy loss function is used to determine the error between the predicted key point and the corresponding true key point by combining the predicted response value of each pixel point in the predicted heat map, the true response value of each pixel point in the original sample image corresponding to the predicted heat map, and the weight.

[0213] Optionally, when it is determined that the key point detection model satisfies a preset convergence condition, the trained key point detection model is outputted, and the processor 1002 is used to:

[0214] When all training sample images are read, it is determined that the key point detection model training is completed; or,

[0215] When the recognition accuracy of the key point detection model based on the training sample image set reaches a set threshold value, it is determined that the training of the key point detection model is completed.

[0216] The processor 1002 is used to input the acquired image to be detected into the key point detection model generated by the method of the first aspect to generate a plurality of heat maps of key points to be detected;

[0217] In each of the obtained key point heat maps to be detected, the pixel point corresponding to the maximum predicted response value is determined as the predicted key point output by the corresponding key point heat map to be detected, wherein the predicted response value of a pixel point represents the probability value of the pixel point being the predicted key point.

[0218] Based on the above embodiment, a storage medium is provided, which at least includes: when the instructions in the storage medium are executed by a processor, the steps of any one of the above methods can be executed.

[0219] In summary, each training sample image is read in sequence, and each time a training sample image is read, the key point detection model is used to obtain the corresponding predicted key points, wherein each time a predicted key point is read, a corresponding layer is generated, and based on the relative distance between each pixel point on the layer and the predicted key point, the predicted response value of each pixel point is recorded on the layer, and the layer recording the predicted response value is determined as a predicted heat map; when the predicted heat map corresponding to at least one predicted key point does not match the corresponding real heat map, the corresponding weight parameter in the key point detection model is adjusted; when the key point detection model meets the preset conditions, the trained key point detection model is output.

[0220] The predicted response value of the pixel point is determined based on the relative distance. When the relative distance between other pixel points around the predicted key point and the predicted key point is close, the predicted response value obtained by calculation is also greatly different from the predicted response value of the predicted key point. Therefore, the heat map to be detected output by the embodiment of the present disclosure has a high discrimination characteristic compared with the traditional prediction heat map. The above method is used to train the key point detection model to improve the model stability, thereby improving the detection accuracy and user experience.

[0221] Specifically, in the process of calculating the predicted response value of a pixel point, the predicted response value of the pixel point is calculated based on the relative distance and a preset attenuation coefficient. The relative distance will increase the gap between the predicted response value of the predicted key point and the predicted response values ​​of other pixel points around it. Since the attenuation coefficient has a value range of [0,1], the gap will be further strengthened. Therefore, the embodiment of the present disclosure can generate a thermal map to be detected with high discrimination.

[0222] In addition, a training sample image set containing original sample images and interference sample images is used to train the key point detection model. In this way, the trained model can not only detect the predicted key points in images without adding interference factors, but also detect the predicted key points in images with multiple interference factors added, thereby improving the model stability, recognition accuracy and user experience.

[0223] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0224] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A training method for a key point detection model, It is characterized in that include: Acquire a preset training sample image set, wherein the training sample image set includes an original sample image and an interference sample image corresponding to the original sample image; Each training sample image is read from the training sample image set in sequence, wherein each time a training sample image is read, the following operations are performed: The key point detection model is used to detect and obtain each predicted key point from the one training sample image, and corresponding predicted heat maps are generated based on the predicted key points, wherein each time a predicted key point is read, a corresponding layer is generated, and the relative distance between each pixel point in the one layer and the one predicted key point is obtained respectively; for each pixel point, the following is performed respectively: when the relative distance does not reach the set distance threshold, the predicted response value of the corresponding pixel point is set to zero, otherwise, the relative distance is used as the parameter n of the attenuation coefficient to determine the predicted response value of the corresponding pixel point; the layer recording the predicted response values ​​of each pixel point is determined as a predicted heat map; When it is determined that the predicted heat map corresponding to at least one predicted key point does not match the corresponding real heat map, the weight parameter corresponding to the at least one predicted key point in the key point detection model is adjusted accordingly, and when it is determined that the key point detection model meets the preset convergence conditions, the trained key point detection model is obtained.

2. The method according to claim 1, It is characterized in that Further including: An interference sample image is an image obtained by adding a set interference factor to a corresponding original sample image, and an original sample image corresponds to at least two interference sample images.

3. The method according to claim 1, It is characterized in that The relative distance between each pixel point in the layer and the predicted key point is obtained respectively, wherein for each pixel point, the following steps are performed respectively: Obtaining the coordinates of a pixel point and the coordinates of the predicted key point; A preset chessboard distance function is used to calculate the relative distance between the pixel point and the predicted key point based on the obtained coordinates.

4. The method according to claim 1, It is characterized in that Determine that the predicted heat map corresponding to a predicted key point does not match the corresponding real heat map, including: Using a multi-label classification algorithm, determine the weight of each pixel in the original sample image corresponding to the predicted heat map; Using a weighted cross entropy loss function, combining the predicted response value of each pixel in the predicted heat map, the true response value of each pixel in the original sample image corresponding to the predicted heat map and the weight, to determine the error between the predicted key point and the corresponding true key point; When the error is higher than a set threshold, it is determined that the predicted heat map corresponding to the one predicted key point does not match the corresponding true heat map.

5. The method according to any one of claims 1 to 4, It is characterized in that When it is determined that the key point detection model meets the preset convergence condition, outputting the trained key point detection model includes: When all training sample images are read, it is determined that the key point detection model training is completed; or, When the recognition accuracy of the key point detection model reaches a set threshold value, it is determined that the training of the key point detection model is completed.

6. A method for key point detection, It is characterized in that include: Inputting the acquired image to be detected into a key point detection model trained by the method according to any one of claims 1 to 4 to generate a plurality of heat maps of key points to be detected; In each of the obtained key point heat maps to be detected, the pixel point corresponding to the maximum predicted response value is determined as the predicted key point output by the corresponding key point heat map to be detected, wherein the predicted response value of a pixel point represents the probability value of the pixel point being the predicted key point.

7. A training device for a key point detection model, It is characterized in that include: An acquisition unit is configured to acquire a preset training sample image set, wherein the training sample image set includes an original sample image and an interference sample image corresponding to the original sample image; The processing unit is configured to read each training sample image from the training sample image set in sequence, wherein each time a training sample image is read, the following operations are performed: The key point detection model is used to detect and obtain each predicted key point from the one training sample image, and corresponding predicted heat maps are generated based on the predicted key points, wherein each time a predicted key point is read, a corresponding layer is generated, and the relative distance between each pixel point in the one layer and the one predicted key point is obtained respectively; for each pixel point, the following is performed respectively: when the relative distance does not reach the set distance threshold, the predicted response value of the corresponding pixel point is set to zero, otherwise, the relative distance is used as the parameter n of the attenuation coefficient to determine the predicted response value of the corresponding pixel point; the layer recording the predicted response values ​​of each pixel point is determined as a predicted heat map; Determining that a predicted heat map corresponding to at least one predicted key point does not match a corresponding real heat map, adjusting a weight parameter corresponding to the at least one predicted key point in the key point detection model accordingly; The output unit is configured to output the trained key point detection model when it is determined that the key point detection model meets the preset convergence condition.

8. The device according to claim 7, It is characterized in that It is further configured as: An interference sample image is an image obtained by adding a set interference factor to a corresponding original sample image, and an original sample image corresponds to at least two interference sample images.

9. The device according to claim 7, It is characterized in that The relative distance between each pixel point in the layer and the predicted key point is obtained respectively, and the processing unit performs the following steps for each pixel point respectively: Obtaining the coordinates of the one pixel point and the coordinates of the one predicted key point; Using a preset chessboard distance function, based on each of the obtained coordinates, calculate the relative distance between the one pixel point and the one predicted key point; The step of determining the predicted response value of the one pixel based on the relative distance, wherein the processing unit is configured to: Based on the relative distance and a preset attenuation coefficient, a predicted response value of the one pixel is calculated.

10. The device according to claim 7, It is characterized in that Determining that a predicted heat map corresponding to a predicted key point does not match a corresponding real heat map, the processing unit is configured to: Using a multi-label classification algorithm, determine the weight of each pixel in the original sample image corresponding to the predicted heat map; Using a weighted cross entropy loss function, combining the predicted response value of each pixel in the predicted heat map, the true response value of each pixel in the original sample image corresponding to the predicted heat map and the weight, to determine the error between the predicted key point and the corresponding true key point; When the error is higher than a set threshold, it is determined that the predicted heat map corresponding to the one predicted key point does not match the corresponding true heat map.

11. The device according to any one of claims 7 to 10, It is characterized in that When it is determined that the key point detection model meets the preset convergence condition, the trained key point detection model is output, and the output unit is configured as follows: When all training sample images are read, it is determined that the key point detection model training is completed; or, When the recognition accuracy of the key point detection model reaches a set threshold value, it is determined that the training of the key point detection model is completed.

12. A device for key point detection, It is characterized in that include: A generating unit, configured to input the acquired image to be detected into a key point detection model trained by the method according to any one of claims 1 to 4, to generate a plurality of heat maps of key points to be detected; The detection unit is configured to determine the pixel point corresponding to the maximum predicted response value in each obtained key point heat map to be detected as the predicted key point output by the corresponding key point heat map to be detected, wherein the predicted response value of a pixel point represents the probability value of the pixel point being the predicted key point.

13. An electronic device, It is characterized in that include: A memory for storing executable instructions; A processor, configured to read and execute executable instructions stored in the memory to implement the training method of the key point detection model as described in any one of claims 1 to 5, or to implement the key point detection method as described in claim 6.

14. A storage medium, It is characterized in that When the instructions in the storage medium are executed by an electronic device, the electronic device is enabled to execute the training method of the key point detection model as described in any one of claims 1 to 5, or to execute the key point detection method as described in claim 6.

Citation Information

Patent Citations

  • A method and apparatus for generating a human key point detection model

    CN109508681A

  • Human body posture index prediction method and device, electronic device and storage medium

    CN110188633A