Key point detection model optimization method and device, equipment, medium and product
By constructing a generative sub-model and a discriminative sub-model of a GAN model and training a key point detection model, the problem of cumbersome and inefficient optimization process in existing technologies is solved, and the optimization efficiency and accuracy of the key point detection model are improved without changing the network structure.
Patent Information
- Application Number
- CN202410627048.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-20
- Publication Date
- 2025-11-21
AI Technical Summary
The optimization process of existing keypoint detection models is cumbersome and inefficient, making it difficult to improve detection capabilities without changing the network structure.
By constructing a Generative Adversarial Network (GAN) model, the GAN model is trained using the feature information of the generative sub-model and the discriminative sub-model, and the key point detection model is optimized. The generative sub-model serves as the optimized key point detection model.
Without changing the network structure of the keypoint detection model, the optimization efficiency and detection accuracy of the model are improved, and the keypoint distribution generated by the sub-model is close to the real distribution.
Smart Images

Figure CN120997527A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of image processing technology, and in particular relates to a method, apparatus, equipment, medium and product for optimizing a key point detection model. Background Technology
[0002] Keypoint detection, also known as keypoint localization or keypoint alignment, typically involves taking an image containing a target (such as a face, body, or hand) as input and outputting a set of predefined keypoint locations, such as facial features and contours, joints of the body, or joints of the hand.
[0003] In related technologies, to improve the detection capability of keypoint detection models and make the data distribution of keypoints detected by the model closer to the data distribution of real keypoints, the keypoint detection model is usually optimized by improving its network structure. However, optimizing the keypoint detection model by improving its network structure is a cumbersome process and has low efficiency. Summary of the Invention
[0004] This application provides a method, apparatus, device, medium, and product for optimizing a key point detection model, which can improve the efficiency of optimizing the key point detection model without changing the network structure of the key point detection model.
[0005] In a first aspect, embodiments of this application provide a key point detection model optimization method, including:
[0006] The feature information of the Generative Adversarial Network (GAN) model, multiple sample images, and the real key points corresponding to each sample image is obtained. The GAN model includes a generator sub-model and a discriminator sub-model. The generator sub-model is constructed based on the key point detection model to be optimized. The feature information includes a confidence image representing the confidence that a pixel in the image is a key point, an X-axis offset image representing the offset of the key point in the X-axis direction, and a Y-axis offset image representing the offset of the key point in the Y-axis direction.
[0007] By using the feature information of multiple sample images and the real key points corresponding to each sample image, a GAN model is trained to obtain the trained GAN model.
[0008] The generated sub-model of the trained GAN model is used as the optimized keypoint detection model.
[0009] Secondly, embodiments of this application provide a key point detection model optimization device, comprising:
[0010] The acquisition module is used to acquire the feature information of the GAN model, multiple sample images, and the real key points corresponding to each sample image. The GAN model includes a generator sub-model and a discriminator sub-model. The generator sub-model is constructed based on the key point detection model to be optimized. The feature information includes a confidence image representing the confidence that a pixel in the image is a key point, an X-axis offset image representing the offset of the key point in the X-axis direction, and a Y-axis offset image representing the offset of the key point in the Y-axis direction.
[0011] The training module is used to train the GAN model using the feature information of multiple sample images and the real key points corresponding to each sample image, and obtain the trained GAN model.
[0012] The determination module is used to use the generated sub-models of the trained GAN model as the optimized keypoint detection model.
[0013] Thirdly, embodiments of this application provide an electronic device, the electronic device including: a processor and a memory storing computer program instructions; the processor reads and executes the computer program instructions to implement the key point detection model optimization method provided in embodiments of this application.
[0014] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the steps of the key point detection model optimization method provided in embodiments of this application.
[0015] Fifthly, embodiments of this application provide a computer program product, the computer program product including computer program instructions, which, when executed by a processor, implement the steps of the key point detection model optimization method provided in embodiments of this application.
[0016] In this embodiment, feature information of a GAN model, multiple sample images, and the corresponding real keypoints for each sample image is obtained. The GAN model includes a generator sub-model and a discriminator sub-model, with the generator sub-model constructed based on the keypoint detection model to be optimized. The GAN model is trained using the feature information of the multiple sample images and the corresponding real keypoints for each sample image, resulting in a trained GAN model. The generator sub-model of the trained GAN model is then used as the optimized keypoint detection model. Thus, by constructing a generator sub-model of the GAN model based on the keypoint detection model to be optimized, the GAN model is trained, and then the trained generator sub-model is used as the optimized keypoint detection model. In the process of optimizing the keypoint detection model, only the generator sub-model of the GAN model is constructed based on the keypoint detection model to be optimized; the network structure of the keypoint detection model is not changed. This allows for optimization of the keypoint detection model without altering its network structure, improving the optimization efficiency of the keypoint detection model. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating the key point detection model optimization method provided in the embodiments of this application;
[0019] Figure 2 This is a schematic diagram of the structure of the GAN model provided in the embodiments of this application;
[0020] Figure 3 This is a schematic diagram of three feature images provided in an embodiment of this application;
[0021] Figure 4 This is a schematic diagram of the key point detection model optimization device provided in the embodiments of this application;
[0022] Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0023] To better understand the above-mentioned objectives, features, and advantages of this application, the solution of this application will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.
[0024] Many specific details are set forth in the following description in order to provide a full understanding of this application, but this application may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some embodiments of this application, and not all embodiments.
[0025] The key point detection model optimization method, apparatus, equipment, medium, and product provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.
[0026] In some possible implementations of the embodiments of this application, the key point detection model optimization method and apparatus provided in the embodiments of this application can be applied to electronic devices. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, handheld computer, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM or self-service machine, etc. The embodiments of this application do not specifically limit it.
[0027] Figure 1 This is a flowchart illustrating the key point detection model optimization method provided in an embodiment of this application. Figure 1 As shown, keypoint detection model optimization methods may include:
[0028] Step 101: Obtain the feature information of the GAN model, multiple sample images, and the real key points corresponding to each sample image. The GAN model includes a generator sub-model and a discriminator sub-model. The generator sub-model is constructed based on the key point detection model to be optimized. The feature information includes a confidence image representing the confidence that a pixel in the image is a key point, an X-axis offset image representing the offset of the key point in the X-axis direction, and a Y-axis offset image representing the offset of the key point in the Y-axis direction.
[0029] In some possible implementations of the embodiments of this application, the main optimization goal of the GAN model is to make the distribution of the data generated by the generative sub-model as close as possible to the distribution of the real data. Based on this, when optimizing the key point detection model, the GAN model can be used to optimize the key point detection model to be optimized. When using the GAN model to optimize the key point detection model to be optimized, the generative sub-model of the GAN model can be constructed according to the key point detection model to be optimized. For example, the key point detection model to be optimized can be used as the generative sub-model of the GAN model, so that the distribution of key points detected by the key point detection model is close to the distribution of real key points.
[0030] Figure 2 This is a schematic diagram of the structure of the GAN model provided in an embodiment of this application. Figure 2 In a GAN model, there are two components: a generator sub-model G and a discriminator sub-model D. The generator sub-model G is constructed from the keypoint detection model to be optimized; that is, the keypoint detection model to be optimized serves as the generator sub-model G of the GAN model. The feature information of real keypoints (Real data) and the feature information of predicted keypoints generated by the generator sub-model G (Fake data) are input into the discriminator sub-model D. When the discriminator sub-model D is input with the feature information of real keypoints, it outputs a larger value; when the discriminator sub-model D is input with the feature information of predicted keypoints, it outputs a smaller value, enabling it to distinguish between real keypoints and the predicted keypoints generated by the generator sub-model G.
[0031] Step 102: Use the feature information of multiple sample images and the real key points corresponding to each sample image to train the GAN model and obtain the trained GAN model.
[0032] In some possible implementations of the embodiments of this application, in step 102, the parameters of the generation sub-model and the discriminator sub-model can be initialized. Based on the feature information of multiple sample images and the real key points corresponding to each sample image, the generation sub-model and the discriminator sub-model are trained simultaneously. The trained discriminator sub-model and the trained generation sub-model are used to train the GAN model and obtain the trained GAN model.
[0033] However, training both the generator and discriminator sub-models simultaneously requires frequent adjustments to their parameters, leading to slower training times for GAN models and consequently lower optimization efficiency for keypoint detection models. Therefore, alternatively, we can first keep the parameters of the generator sub-model constant, and then train the discriminator sub-model using multiple sample images and the feature information of the corresponding real keypoints in each sample image. This results in a well-trained discriminator sub-model. Then, keeping the parameters of the trained discriminator sub-model constant, we train the generator sub-model using multiple sample images. This process is repeated until the training termination condition is met.
[0034] In some possible implementations of the embodiments of this application, the training termination conditions of the embodiments of this application include, but are not limited to: the number of training times reaching a preset training time threshold, and the loss values of the loss function of the discriminant sub-model and the loss function of the generator sub-model both reaching their corresponding minimum values.
[0035] In the embodiments of this application, by keeping the parameters of the generator sub-model unchanged when training the discriminator sub-model and keeping the parameters of the discriminator sub-model unchanged when training the generator sub-model, the training efficiency of the GAN model can be improved, thereby improving the optimization efficiency of the key point detection model.
[0036] In some possible implementations of this application, when training a discriminant sub-model based on multiple sample images and feature information of the real keypoints corresponding to each sample image, and obtaining a trained discriminant sub-model, the sample images can be input into the generator sub-model to obtain feature information of the predicted keypoints; the feature information of the real keypoints and the feature information of the predicted keypoints are input into the discriminant sub-model so that the discriminant sub-model can discriminate the feature information of the real keypoints and the feature information of the predicted keypoints respectively, and obtain the first discrimination results corresponding to the feature information of the real keypoints and the feature information of the predicted keypoints respectively; the first discrimination results are substituted into the loss function of the discriminant sub-model to obtain the first loss value of the loss function of the discriminant sub-model; the parameters of the discriminant sub-model are adjusted according to the first loss value; the process continues until the first loss value satisfies the first preset training stopping condition, and a trained discriminant sub-model is obtained.
[0037] In some possible implementations of the embodiments of this application, the GAN model in the embodiments of this application can be a generative adversarial network model based on Wasserstein distance (i.e., WGAN model). The WGAN model can alleviate the problem of model collapse to a certain extent and improve the stability of GAN model training, avoiding training failure. Here, model collapse refers to the phenomenon that the data generated by the model is monotonous and has very poor diversity.
[0038] In some possible implementations of the embodiments of this application, the GAN model in the embodiments of this application can be a WGAN model with gradient penalty (GP) (i.e., WGAN-GP model). The WGAN-GP model can solve the gradient vanishing and gradient exploding problems of the WGAN model, making the model training more stable and making the distribution of data generated by the sub-model more likely to be close to the distribution of real data.
[0039] In some possible implementations of this application's embodiments, the loss function of the discriminant sub-model can be determined based on the discriminant sub-model's discrimination result for the feature information of the predicted keypoints, the discriminant sub-model's discrimination result for the feature information of the real keypoints, and a gradient penalty term, before substituting the first discrimination result into the discriminant sub-model's loss function to obtain the first loss value of the discriminant sub-model's loss function. Based on this, before substituting the first discrimination result into the discriminant sub-model's loss function to obtain the first loss value of the discriminant sub-model's loss function, the keypoint detection model optimization method provided in this application's embodiments may further include: determining the loss function of the discriminant sub-model by summing the difference between the discriminant sub-model's discrimination result for the feature information of the predicted keypoints and the discriminant sub-model's discrimination result for the feature information of the real keypoints, and the gradient penalty term.
[0040] The loss function of the discriminant sub-model can be expressed as shown in formula (1):
[0041] D loss =D(G(x))-D(x) data )+GP (1)
[0042] Wherein, formula (1) represents the loss function of the discriminant sub-model; in formula (1), D loss D(G)x) represents the loss value of the discriminant sub-model's loss function, and D(x) represents the discriminant sub-model's judgment result on the feature information of the predicted key points. data ) represents the discrimination result of the sub-model on the feature information of the real keypoints, GP is the gradient penalty term, G(x) is the feature information of the predicted keypoints, x is the sample image, and x is the feature information of the predicted keypoints. data The feature information of the real key points. The gradient penalty term GP is shown in the following formula (2):
[0043]
[0044] In formula (2), λ is a parameter. For gradient operators, To distinguish sub-model pairs The judgment result, As shown in formula (3) below:
[0045]
[0046] In formula (3), ε follows a (0,1) distribution, G(x) represents the feature information of the predicted keypoints, and x data This refers to the feature information of real key points.
[0047] In some possible implementations of the embodiments of this application, the first preset stopping condition in the embodiments of this application can be that the first loss value reaches the minimum value while the parameters of the generated sub-model remain unchanged.
[0048] In some possible implementations of the embodiments of this application, when training a generative sub-model based on multiple sample images to obtain a trained generative sub-model, the sample images can be input into the generative sub-model to obtain the feature information of the predicted key points; the feature information of the predicted key points can be input into a discriminative sub-model so that the discriminative sub-model can discriminate the feature information of the predicted key points to obtain a second discrimination result corresponding to the feature information of the predicted key points; the second discrimination result can be substituted into the loss function of the generative sub-model to obtain a second loss value of the loss function of the generative sub-model; the parameters of the generative sub-model can be adjusted according to the second loss value; the process of inputting sample images into the generative sub-model to obtain the feature information of the predicted key points continues until the second loss value meets the second preset training stopping condition, thus obtaining a trained generative sub-model.
[0049] In some possible implementations of this application's embodiments, since the generative sub-model of the GAN model is constructed from the keypoint detection model to be optimized, and the training of the generative sub-model is affected by the discrimination result of the discriminative sub-model of the GAN model, the loss function of the keypoint detection model to be optimized cannot be directly used as the loss function of the generative sub-model. Instead, the loss function of the generative sub-model must be determined based on the loss function of the keypoint detection model to be optimized and the discrimination result of the discriminative sub-model. Therefore, before substituting the second discrimination result into the loss function of the generative sub-model to obtain the second loss value of the generative sub-model's loss function, the keypoint detection model optimization method provided in this application's embodiments may further include: determining the loss function of the generative sub-model as the sum of the loss function of the keypoint detection model to be optimized and the discrimination result of the discriminative sub-model on the feature information of the predicted keypoints.
[0050] The loss function for generating the sub-model can be expressed as shown in formula (4):
[0051] G loss =Net loss +D(G(x)) (4)
[0052] Wherein, formula (4) represents the loss function for generating the sub-model; in formula (4), G loss To generate the loss value of the loss function for the sub-model, Net... lossLet be the loss function of the keypoint detection model to be optimized, D(G(x)) be the discrimination result of the discriminant sub-model on the feature information of the predicted keypoint, G(x) be the feature information of the predicted keypoint, and x be the sample image.
[0053] The embodiments of this application do not elaborate on the process of adjusting model parameters based on the loss value; for details, please refer to the descriptions in related technologies.
[0054] In some possible implementations of the embodiments of this application, before step 101, the key point detection model optimization method provided in the embodiments of this application may further include: determining the confidence level of each pixel in the image as a key point; determining the offset of the pixel as a key point in the image relative to the top-left corner vertex position of the image in the X-axis direction and the offset of the pixel as a key point in the image in the Y-axis direction based on the coordinates of the pixel as a key point in the image; generating a confidence image based on the confidence level of each pixel in the image; generating an X-axis offset image based on the offset of the pixel as a key point in the image relative to the top-left corner vertex position of the image in the X-axis direction; and generating a Y-axis offset image based on the offset of the pixel as a key point in the image relative to the top-left corner vertex position of the image in the Y-axis direction.
[0055] In some possible implementations of the embodiments of this application, for a pixel that is a key point in the image, the confidence level of the pixel is 1, and for a pixel that is not a key point in the image, the confidence level of the pixel is 0.
[0056] When determining the offset of a key pixel relative to the top-left corner vertex of an image based on its coordinates, the X-axis component of the key pixel's coordinates can be used as the offset relative to the top-left corner vertex; and the Y-axis component of the key pixel's coordinates can be used as the offset relative to the top-left corner vertex.
[0057] When generating a confidence image based on the confidence score of each pixel in the image, the coordinate position value corresponding to each pixel in the confidence image is set to the confidence score of that pixel; a confidence image is generated based on the confidence score of each pixel in the image; when generating an X-axis offset image based on the X-axis offset of a keypoint pixel relative to the top-left corner of the image, the coordinate position value corresponding to the keypoint in the X-axis offset image is set to the X-axis offset of the keypoint relative to the top-left corner of the image, and the coordinate positions value corresponding to other pixels are set to 0; when generating a Y-axis offset image based on the Y-axis offset of a keypoint pixel relative to the top-left corner of the image, the coordinate position value corresponding to the keypoint in the Y-axis offset image is set to the Y-axis offset of the keypoint relative to the top-left corner of the image, and the coordinate positions value corresponding to other pixels are set to 0.
[0058] For a sample image with width W, height H, and N key points, after inputting it into the generator sub-model G (i.e., the key point detection model), the generator sub-model G outputs three feature images of W*H*N. Among them, one feature image is the confidence image of W*H*N, one image is the X-axis offset image of W*H*N, and one image is the Y-axis offset image of W*H*N.
[0059] For example, consider a target image with W = 4, H = 4, N = 1, and the pixels in the third row and second column as keypoints. After this target image is input into the generator sub-model G, the generated three feature images are as follows: Figure 3 As shown, Figure 3 This is a schematic diagram of three feature images provided in an embodiment of this application. Figure 3 In the image, the middle image is the X-axis offset image, the image to the left of the middle image is the confidence score image, and the image to the right of the middle image is the Y-axis offset image. Figure 3 In the confidence image, the value in the third row and second column is 1, meaning the confidence score for the pixel in the third row and second column of the target image as a keypoint is 1. The values in the corresponding regions of other rows and columns in this confidence image are 0, indicating that the confidence score for other pixels as keypoints is 0. Figure 3 In the X-axis offset image, the value of 1 in the third row and second column indicates that the keypoint's offset relative to the top-left corner of the target image is 1 in the X-axis direction. The values in the corresponding regions of other rows and columns in this X-axis offset image are 0. Figure 3 In the Y-axis offset image, the value of the third row and second column is 2, which means that the key point is offset by 2 in the Y-axis direction relative to the top left corner of the target image. The values of the corresponding areas in other rows and columns of this Y-axis offset image are 0.
[0060] Step 103: Use the generated sub-model of the trained generative adversarial network model as the optimized key point detection model.
[0061] Once the GAN model is trained, the data distribution of predicted keypoints generated by the generative sub-models included in the GAN model is close to the data distribution of real keypoints. At this point, the generative sub-models of the trained generative adversarial network model can be used as the optimized keypoint detection model.
[0062] In this embodiment, feature information of a GAN model, multiple sample images, and the corresponding real keypoints for each sample image is obtained. The GAN model includes a generator sub-model and a discriminator sub-model, with the generator sub-model constructed based on the keypoint detection model to be optimized. The GAN model is trained using the feature information of the multiple sample images and the corresponding real keypoints for each sample image, resulting in a trained GAN model. The generator sub-model of the trained GAN model is then used as the optimized keypoint detection model. Thus, by constructing a generator sub-model of the GAN model based on the keypoint detection model to be optimized, the GAN model is trained, and then the trained generator sub-model is used as the optimized keypoint detection model. In the process of optimizing the keypoint detection model, only the generator sub-model of the GAN model is constructed based on the keypoint detection model to be optimized; the network structure of the keypoint detection model is not changed. This allows for optimization of the keypoint detection model without altering its network structure, improving the optimization efficiency of the keypoint detection model.
[0063] Corresponding to the above method embodiments, this application also provides a key point detection model optimization device. For example... Figure 4 As shown, Figure 4 This is a schematic diagram of the keypoint detection model optimization device provided in an embodiment of this application. The keypoint detection model optimization device 400 may include:
[0064] The acquisition module 401 is used to acquire the feature information of the GAN model, multiple sample images, and the real key points corresponding to each sample image. The GAN model includes a generator sub-model and a discriminator sub-model. The generator sub-model is constructed based on the key point detection model to be optimized. The feature information includes a confidence image representing the confidence that a pixel in the image is a key point, an X-axis offset image representing the offset of the key point in the X-axis direction, and a Y-axis offset image representing the offset of the key point in the Y-axis direction.
[0065] Training module 402 is used to train the GAN model using the feature information of multiple sample images and the real key points corresponding to each sample image, so as to obtain the trained GAN model.
[0066] The determination module 403 is used to use the generated sub-model of the trained GAN model as the optimized key point detection model.
[0067] In this embodiment, feature information of a GAN model, multiple sample images, and the corresponding real keypoints for each sample image is obtained. The GAN model includes a generator sub-model and a discriminator sub-model, with the generator sub-model constructed based on the keypoint detection model to be optimized. The GAN model is trained using the feature information of the multiple sample images and the corresponding real keypoints for each sample image, resulting in a trained GAN model. The generator sub-model of the trained GAN model is then used as the optimized keypoint detection model. Thus, by constructing a generator sub-model of the GAN model based on the keypoint detection model to be optimized, the GAN model is trained, and then the trained generator sub-model is used as the optimized keypoint detection model. In the process of optimizing the keypoint detection model, only the generator sub-model of the GAN model is constructed based on the keypoint detection model to be optimized; the network structure of the keypoint detection model is not changed. This allows for optimization of the keypoint detection model without altering its network structure, improving the optimization efficiency of the keypoint detection model.
[0068] In some possible implementations of embodiments of this application, the training module 402 may include:
[0069] The initialization submodule is used to initialize the parameters of the generation submodel and the discrimination submodel;
[0070] The discriminant sub-model training sub-module is used to keep the parameters of the generated sub-model unchanged, and train the discriminant sub-model based on the feature information of multiple sample images and the real key points corresponding to each sample image, so as to obtain the trained discriminant sub-model.
[0071] The generator sub-model training sub-module is used to train the generator sub-model based on multiple sample images while keeping the parameters of the trained discriminator sub-model unchanged, thus obtaining the trained generator sub-model.
[0072] In the embodiments of this application, by keeping the parameters of the generator sub-model unchanged when training the discriminator sub-model and keeping the parameters of the discriminator sub-model unchanged when training the generator sub-model, the training efficiency of the GAN model can be improved, thereby improving the optimization efficiency of the key point detection model.
[0073] In some possible implementations of the embodiments of this application, the discriminant sub-model training sub-module is specifically used for:
[0074] The sample image is input into the generator sub-model to obtain the feature information of the predicted key points;
[0075] The feature information of the real key points and the feature information of the predicted key points are input into the discriminant sub-model so that the discriminant sub-model can discriminate the feature information of the real key points and the feature information of the predicted key points respectively, and obtain the first discrimination result corresponding to the feature information of the real key points and the feature information of the predicted key points respectively.
[0076] Substitute the first discrimination result into the loss function of the discriminant sub-model to obtain the first loss value of the loss function of the discriminant sub-model;
[0077] Adjust the parameters of the discriminant sub-model based on the first loss value;
[0078] Return to the input of the sample image into the generator sub-model to obtain the feature information of the predicted key points, until the first loss value meets the first preset training stopping condition, and the trained discriminant sub-model is obtained.
[0079] In some possible implementations of the embodiments of this application, the discriminant sub-model training sub-module can also be used for:
[0080] The difference between the discrimination result of the discriminant sub-model on the feature information of the predicted key points and the discrimination result of the discriminant sub-model on the feature information of the real key points, plus the sum of the gradient penalty term, is determined as the loss function of the discriminant sub-model.
[0081] In some possible implementations of the embodiments of this application, the sub-model training submodule is specifically used for:
[0082] The sample image is input into the generator sub-model to obtain the feature information of the predicted key points;
[0083] The feature information of the predicted key points is input into the discriminant sub-model so that the discriminant sub-model can discriminate the feature information of the predicted key points and obtain the second discrimination result corresponding to the feature information of the predicted key points.
[0084] Substituting the second discrimination result into the loss function of the generator sub-model, we obtain the second loss value of the loss function of the generator sub-model;
[0085] Adjust the parameters of the generated sub-model based on the second loss value;
[0086] Return to the input of the sample image into the generator sub-model to obtain the feature information of the predicted key points, until the second loss value meets the second preset training stopping condition, and the trained generator sub-model is obtained.
[0087] In some possible implementations of the embodiments of this application, the sub-model training submodule can also be used for:
[0088] The loss function of the generation sub-model is determined by summing the loss function of the key point detection model to be optimized and the discrimination result of the discriminant sub-model on the feature information of the predicted key points.
[0089] In some possible implementations of the embodiments of this application, the key point detection model optimization device 400 provided in the embodiments of this application may further include:
[0090] The confidence determination module is used to determine the confidence level of each pixel in the image as a key point;
[0091] The offset determination module is used to determine the offset of a key pixel in the image relative to the top-left corner vertex of the image along the X-axis and the offset of the key pixel relative to the top-left corner vertex of the image, based on the coordinates of the key pixel in the image.
[0092] The generation module is used to generate a confidence image based on the confidence level of each pixel in the image; to generate an X-axis offset image based on the X-axis offset of the key pixels in the image relative to the top-left vertex of the image; and to generate a Y-axis offset image based on the Y-axis offset of the key pixels in the image relative to the top-left vertex of the image.
[0093] Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application.
[0094] The electronic device may include a processor 501 and a memory 502 storing computer program instructions.
[0095] Specifically, the processor 501 may include a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0096] Memory 502 may include mass storage for data or instructions. For example, and not limitingly, memory 502 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where suitable, memory 502 may include removable or non-removable (or fixed) media. Where suitable, memory 502 may be internal or external to an electronic device. In some specific embodiments, memory 502 is a non-volatile solid-state memory.
[0097] In some specific embodiments, the memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Therefore, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the keypoint detection model optimization method according to this application.
[0098] The processor 501 reads and executes computer program instructions stored in the memory 502 to implement the steps of the key point detection model optimization method provided in the embodiments of this application.
[0099] In some examples, the electronic device may also include a communication interface 503 and a bus 510. For example, Figure 5 As shown, the processor 501, memory 502, and communication interface 503 are connected through bus 510 and complete communication with each other.
[0100] The communication interface 503 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0101] Bus 510 includes hardware, software, or both, that couples components of an electronic device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, bus 510 may include one or more buses. Although specific buses are described and illustrated in the embodiments of this application, this application considers any suitable bus or interconnection.
[0102] The electronic device can execute the key point detection model optimization method provided in the embodiments of this application, thereby achieving the corresponding technical effects of the key point detection model optimization method provided in the embodiments of this application.
[0103] In addition, in conjunction with the key point detection model optimization method in the above embodiments, this application also provides a computer-readable storage medium for implementation. This computer-readable storage medium stores computer program instructions; when executed by a processor, these computer program instructions implement the steps of the key point detection model optimization method provided in this application. Examples of computer-readable storage media include non-transitory computer-readable media, such as ROM, RAM, magnetic disks, or optical disks.
[0104] This application also provides a computer program product, which includes computer program instructions. When the computer program instructions are executed by a processor, they implement the steps of the key point detection model optimization method provided in this application and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0105] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the term "comprising" or any other variations thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
Claims
1. A method for optimizing a key point detection model, characterized in that, The method includes: The method involves acquiring feature information of a generative adversarial network (GAN) model, multiple sample images, and the real keypoints corresponding to each sample image. The GAN model includes a generator sub-model and a discriminator sub-model. The generator sub-model is constructed based on the keypoint detection model to be optimized. The feature information includes a confidence image representing the confidence that a pixel in the image is a keypoint, an X-axis offset image representing the offset of the keypoint in the X-axis direction, and a Y-axis offset image representing the offset of the keypoint in the Y-axis direction. The generative adversarial network model is trained using the feature information of the multiple sample images and the real key points corresponding to each sample image, resulting in the trained generative adversarial network model. The generated sub-model of the trained generative adversarial network model is used as the optimized keypoint detection model.
2. The method as described in claim 1, characterized in that, The step of training the generative adversarial network model using the feature information of the sample images and the real key points includes: Initialize the parameters of the generating sub-model and the discriminating sub-model; Keeping the parameters of the generated sub-model unchanged, the discriminative sub-model is trained based on the feature information of the multiple sample images and the real key points corresponding to each sample image, to obtain the trained discriminative sub-model; Keeping the parameters of the trained discriminative sub-model unchanged, the generative sub-model is trained based on the multiple sample images to obtain the trained generative sub-model.
3. The method as described in claim 2, characterized in that, The step of training the discriminant sub-model based on the feature information of the plurality of sample images and the real key points corresponding to each sample image to obtain the trained discriminant sub-model includes: The sample image is input into the generation sub-model to obtain the feature information of the predicted key points; The feature information of the real key points and the feature information of the predicted key points are input into the discrimination sub-model, so that the discrimination sub-model can discriminate the feature information of the real key points and the feature information of the predicted key points respectively, and obtain the first discrimination result corresponding to the feature information of the real key points and the feature information of the predicted key points respectively; Substitute the first discrimination result into the loss function of the discrimination sub-model to obtain the first loss value of the loss function of the discrimination sub-model; Based on the first loss value, adjust the parameters of the discriminant sub-model; The sample image is then input into the generating sub-model to obtain the feature information of the predicted key points, until the first loss value meets the first preset training stopping condition, thus obtaining the trained discriminative sub-model.
4. The method as described in claim 3, characterized in that, Before substituting the first discrimination result into the loss function of the discrimination sub-model to obtain the first loss value of the loss function of the discrimination sub-model, the following steps are included: The difference between the discrimination result of the discriminant sub-model on the feature information of the predicted key points and the discrimination result of the discriminant sub-model on the feature information of the real key points, plus the sum of the gradient penalty term, is determined as the loss function of the discriminant sub-model.
5. The method as described in claim 2, characterized in that, The step of training the generator sub-model based on the multiple sample images to obtain the trained generator sub-model includes: The sample image is input into the generation sub-model to obtain the feature information of the predicted key points; The feature information of the predicted key points is input into the discrimination sub-model so that the discrimination sub-model can discriminate the feature information of the predicted key points and obtain a second discrimination result corresponding to the feature information of the predicted key points. Substitute the second discrimination result into the loss function of the generated sub-model to obtain the second loss value of the loss function of the generated sub-model; Based on the second loss value, adjust the parameters of the generated sub-model; The sample image is then input into the generator sub-model to obtain the feature information of the predicted key points, until the second loss value meets the second preset training stopping condition, thus obtaining the trained generator sub-model.
6. The method as described in claim 5, characterized in that, Before substituting the second discrimination result into the loss function of the generated sub-model to obtain the second loss value of the loss function of the generated sub-model, the following steps are included: The loss function of the generating sub-model is determined by summing the loss function of the key point detection model to be optimized and the discrimination result of the discriminant sub-model on the feature information of the predicted key points.
7. The method according to any one of claims 1 to 6, characterized in that, Before acquiring the feature information of the generative adversarial network model, multiple sample images, and the real keypoints corresponding to each sample image, the process includes: Determine the confidence level of each pixel in the image as a keypoint; Based on the coordinates of the key pixels in the image, determine the offset of the key pixels in the image relative to the top left corner vertex in the X-axis direction and the offset of the key pixels in the Y-axis direction relative to the top left corner vertex in the image. The confidence image is generated based on the confidence level of each pixel in the image; Based on the offset of the key pixels in the image relative to the top left corner vertex of the image along the X-axis, the X-axis offset image is generated. The Y-axis offset image is generated based on the offset of the key pixels in the image relative to the top left corner vertex of the image along the Y-axis.
8. A key point detection model optimization device, characterized in that, The device includes: The acquisition module is used to acquire the feature information of the generative adversarial network model, multiple sample images, and the real key points corresponding to each sample image. The generative adversarial network model includes a generator sub-model and a discriminator sub-model. The generator sub-model is constructed based on the key point detection model to be optimized. The feature information includes a confidence image representing the confidence that a pixel in the image is a key point, an X-axis offset image representing the offset of the key point in the X-axis direction, and a Y-axis offset image representing the offset of the key point in the Y-axis direction. The training module is used to train the generative adversarial network model using the feature information of the multiple sample images and the real key points corresponding to each sample image, so as to obtain the trained generative adversarial network model. The determination module is used to use the generated sub-model of the trained generative adversarial network model as the optimized key point detection model.
9. An electronic device, characterized in that, The electronic device includes: a processor and a memory storing computer program instructions; The processor reads and executes the computer program instructions to implement the steps of the key point detection model optimization method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the steps of the key point detection model optimization method as described in any one of claims 1-7.
11. A computer program product, characterized in that, The computer program product includes computer program instructions, which, when executed by a processor, implement the steps of the key point detection model optimization method as described in any one of claims 1-7.