Image Processing Method, Apparatus, Storage Medium, and Electronic Device

By expanding the sample data set containing the human key point labels and using the label to calibrate the model, the problem of insufficient training samples of the human part segmentation model is solved, and the accuracy of the segmentation model is improved.

CN114913181BActive Publication Date: 2025-07-01BEIJING YUDA ORIENTAL SOFTWARE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210476252.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-29
Publication Date
2025-07-01
Estimated Expiration
2042-04-29

AI Technical Summary

Technical Problem

In the prior art, training of human body part segmentation models requires a large amount of labeled sample data, which leads to high labeling costs and is difficult to obtain sufficient sample data, which affects the accuracy of the model segmentation.

Method used

By using the first pending sample data set containing the human key point label, the target sample data set is expanded, the target human image segmentation model is obtained, and the segmentation accuracy is improved by combining the label calibration model.

Benefits of technology

It improves the efficiency of obtaining training sample data, expands the sample data, and improves the accuracy of the human body part segmentation model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114913181B_ABST
    Figure CN114913181B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an image processing method, apparatus, storage medium, and electronic device. The method includes: obtaining a to-be-processed image including a target human body; inputting the to-be-processed image into a target human body image segmentation model to perform part segmentation on the target human body to obtain a target image after part segmentation; wherein, the target human body image segmentation model is trained according to a target sample data set, the target sample data set includes a plurality of target sample images and target part labels corresponding to each target sample image, the target sample images include first sample images in a first to-be-determined sample data set, and the target part labels are obtained according to the first to-be-determined sample data set and a second to-be-determined sample data set; the first to-be-determined sample data set includes first sample images and first human body key point labels corresponding thereto; the second to-be-determined sample data set includes second sample images and second human body key point labels and second part labels corresponding thereto.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, and in particular, to an image processing method, apparatus, storage medium, and electronic device. Background Art

[0002] Human body part segmentation is to retrieve more fine-grained semantic labels in an image containing a human body, such as head labels, limb labels, and torso labels, etc. The application scope of human body part segmentation is extensive. For example, it can be applied in scenarios such as behavior understanding, robot operation, heuristic reasoning, and identifying human-computer interaction.

[0003] In the related art, human body part segmentation can be achieved through a neural network model. However, in order to train the neural network model, a large amount of labeled training sample data is required, and the cost of labeling the sample data is extremely high, resulting in difficulty in obtaining sufficient sample data to train the neural network model. The small sample size will cause the trained model to be difficult to accurately achieve human body part segmentation. Summary of the Invention

[0004] The purpose of the present disclosure is to provide an image processing method, apparatus, storage medium, and electronic device to partially solve the above problems existing in the related art.

[0005] To achieve the above purpose, a first aspect of the present disclosure provides an image processing method, the method comprising:

[0006] Obtain a to-be-processed image containing a target human body;

[0007] Input the to-be-processed image into a target human body image segmentation model to perform part segmentation on the target human body, so as to obtain a target image after part segmentation;

[0008] Wherein, the target human body image segmentation model is trained according to a target sample data set, the target sample data set includes a plurality of target sample images, and a target part label corresponding to each target sample image. The target sample image includes a first sample image in a first to-be-determined sample data set, and the target part label is obtained according to the first to-be-determined sample data set and a second to-be-determined sample data set; the first to-be-determined sample data set includes a plurality of first sample images, and a first human body key point label corresponding to each first sample image; the second to-be-determined sample data set includes a plurality of second sample images, and a second human body key point label and a second part label corresponding to each second sample image.

[0009] Optionally, the target human body image segmentation model is trained in the following manner:

[0010] Obtain the first to-be-determined sample dataset and the second to-be-determined sample dataset;

[0011] Obtain the first part label corresponding to the first sample image according to the second to-be-determined sample dataset;

[0012] Determine the target sample dataset according to the first sample image and the first part label;

[0013] Train a first preset neural network model according to the target sample dataset to obtain the target human image segmentation model.

[0014] Optionally, the obtaining the first part label corresponding to the first sample image according to the second to-be-determined sample dataset includes:

[0015] Obtain the first human pose corresponding to each first sample image in the first to-be-determined sample dataset according to the first human key point label;

[0016] Obtain the second human pose corresponding to each second sample image in the second to-be-determined sample dataset according to the second human key point label;

[0017] Determine the first part label corresponding to the first sample image according to the first human pose and the second human pose.

[0018] Optionally, the determining the first part label corresponding to the first sample image according to the first human pose and the second human pose includes:

[0019] Obtain the first candidate part label corresponding to the first sample image according to the label obtaining step;

[0020] Determine the first part label corresponding to the first sample image according to the first candidate part label;

[0021] Wherein, the label obtaining step includes:

[0022] Determine one or more target pose sample images corresponding to the to-be-determined sample image from the second to-be-determined sample dataset according to the to-be-determined human pose corresponding to the to-be-determined sample image; the second human pose corresponding to the target pose sample image has a similarity greater than or equal to a preset pose similarity threshold with the to-be-determined human pose;

[0023] Obtain the to-be-determined part label corresponding to the to-be-determined sample image according to the second part label corresponding to the target pose sample image; the to-be-determined sample image includes the first sample image, and the to-be-determined part label includes the first candidate part label corresponding to the first sample image.

[0024] Optionally, obtaining the to-be-determined part label corresponding to the to-be-determined sample image according to the second part label corresponding to the target pose sample image includes:

[0025] Taking the average value of the second part labels of one or more of the target pose sample images as the to-be-determined part label corresponding to the to-be-determined sample image.

[0026] Optionally, determining the first part label corresponding to the first sample image according to the first candidate part label includes:

[0027] Inputting the first sample image, the first human key point label, and the first candidate part label into a pre-trained label calibration model to obtain the first part label corresponding to the first sample image; wherein, the label calibration model is trained according to a second to-be-determined sample data set.

[0028] Optionally, the to-be-determined sample data includes the second sample data, and the to-be-determined part label includes the second candidate part label corresponding to the second sample data; the label calibration model is trained in the following manner:

[0029] According to the label obtaining step, obtaining the second candidate part label corresponding to each second sample data in the second to-be-determined sample data set;

[0030] Training a second preset neural network model according to the second candidate part label and the second to-be-determined sample data set to obtain the label calibration model.

[0031] In a second aspect, the present disclosure provides an image processing device, and the device includes:

[0032] An image acquisition module, configured to acquire a to-be-processed image including a target human body;

[0033] An image processing module, configured to input the to-be-processed image into a target human body image segmentation model to perform part segmentation on the target human body to obtain a target image after part segmentation;

[0034] Among them, the target human body image segmentation model is trained according to a target sample dataset. The target sample dataset includes a plurality of target sample images and target part labels corresponding to each target sample image. The target sample images include the first sample images in the first to-be-determined sample dataset, and the target part labels are obtained according to the first to-be-determined sample dataset and the second to-be-determined sample dataset. The first to-be-determined sample dataset includes a plurality of first sample images and first human key point labels corresponding to each first sample image. The second to-be-determined sample dataset includes a plurality of second sample images and second human key point labels and second part labels corresponding to each second sample image.

[0035] Optionally, the device further includes a first model training module;

[0036] The first model training module is configured to obtain the first to-be-determined sample dataset and the second to-be-determined sample dataset; obtain the first part label corresponding to the first sample image according to the second to-be-determined sample dataset; determine the target sample dataset according to the first sample image and the first part label; and train a first preset neural network model according to the target sample dataset to obtain the target human body image segmentation model.

[0037] Optionally, the first model training module is configured to obtain the first human body pose corresponding to each first sample image in the first to-be-determined sample dataset according to the first human key point label; obtain the second human body pose corresponding to each second sample image in the second to-be-determined sample dataset according to the second human key point label; and determine the first part label corresponding to the first sample image according to the first human body pose and the second human body pose.

[0038] Optionally, the first model training module is configured to obtain the first candidate part label corresponding to the first sample image according to a label obtaining step; and determine the first part label corresponding to the first sample image according to the first candidate part label;

[0039] Among them, the label obtaining step includes:

[0040] Determine one or more target pose sample images corresponding to the to-be-determined sample image from the second to-be-determined sample dataset according to the to-be-determined human body pose corresponding to the to-be-determined sample image; the second human body pose corresponding to the target pose sample image has a similarity greater than or equal to a preset pose similarity threshold with the to-be-determined human body pose;

[0041] Obtain the to-be-determined part label corresponding to the to-be-determined sample image according to the second part label corresponding to the target pose sample image; the to-be-determined sample image includes the first sample image, and the to-be-determined part label includes the first candidate part label corresponding to the first sample image.

[0042] Optionally, the first model training module is configured to use the average value of the second part labels of one or more of the target pose sample images as the to-be-determined part label corresponding to the to-be-determined sample image.

[0043] Optionally, the first model training module is configured to input the first sample image, the first human key point label, and the first candidate part label into a pre-trained label calibration model to obtain the first part label corresponding to the first sample image; wherein, the label calibration model is trained according to a second to-be-determined sample data set.

[0044] Optionally, the to-be-determined sample data includes the second sample data, and the to-be-determined part label includes the second candidate part label corresponding to the second sample data; the apparatus further includes a second model training module;

[0045] The second model training module is configured to obtain the second candidate part label corresponding to each second sample data in the second to-be-determined sample data set according to the label obtaining step; and train a second preset neural network model according to the second candidate part label and the second to-be-determined sample data set to obtain the label calibration model.

[0046] In a third aspect, the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in the first aspect of the present disclosure are implemented.

[0047] In a fourth aspect, the present disclosure provides an electronic device, including: a memory, on which a computer program is stored; a processor, configured to execute the computer program in the memory to implement the steps of the method described in the first aspect of the present disclosure.

[0048] Adopt the above technical solution to obtain a to-be-processed image including a target human body; input the to-be-processed image into a target human body image segmentation model to perform part segmentation on the target human body to obtain a target image after part segmentation; wherein, the target human body image segmentation model is trained according to a target sample data set, the target sample data set includes a plurality of target sample images and a target part label corresponding to each target sample image, the target sample image includes a first sample image in a first to-be-determined sample data set, and the target part label is obtained according to the first to-be-determined sample data set and a second to-be-determined sample data set; the first to-be-determined sample data set includes a plurality of first sample images and a first human body key point label corresponding to each first sample image; the second to-be-determined sample data set includes a plurality of second sample images and a second human body key point label and a second part label corresponding to each second sample image. In this way, it can be expanded based on the sample (the first to-be-determined sample) including human body key points, and on the basis of the first to-be-determined sample data set including human body key point labels, expand to obtain a target sample data set including target part labels, thereby improving the acquisition efficiency of training sample data, efficiently expanding sample data, and improving the accuracy of the trained target human body image segmentation model for human body part segmentation.

[0049] Other features and advantages of the present disclosure will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The drawings are used to provide a further understanding of the present disclosure, and constitute a part of the specification, and are used to explain the present disclosure together with the following specific implementation manners, but do not constitute a limitation to the present disclosure. In the drawings:

[0051] Figure 1 is a flowchart of an image processing method provided by an embodiment of the present disclosure.

[0052] Figure 2 is a flowchart of a training method of a target human body image segmentation model provided by an embodiment of the present disclosure.

[0053] Figure 3 is according to Figure 2 shown in the embodiment is a flowchart of an S202 step.

[0054] Figure 4 is a flowchart of a training method of a label calibration model provided by an embodiment of the present disclosure.

[0055] Figure 5 is a schematic diagram of an image processing device provided by an embodiment of the present disclosure.

[0056] Figure 6 is a schematic diagram of another image processing device provided by an embodiment of the present disclosure.

[0057] Figure 7 It is a schematic diagram of another image processing apparatus provided by an embodiment of the present disclosure.

[0058] Figure 8 It is a block diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0059] The following will describe the detailed implementation manners of the present disclosure with reference to the accompanying drawings. It should be understood that the detailed implementation manners described herein are only for explaining and illustrating the present disclosure, and are not used to limit the present disclosure.

[0060] It should be noted that all actions of obtaining signals, information, or data in the present disclosure are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where it is located, and with the authorization given by the owner of the corresponding device.

[0061] In the present disclosure, terms such as "first" and "second" are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying an order; terms such as "S101", "S102", "S201", "S202", etc. are used to distinguish steps, and do not have to be understood as executing method steps in a specific order or sequence; when the following description involves the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements.

[0062] First, the application scenarios of the present disclosure will be described. The present disclosure can be applied to image processing, especially in the scenario of segmenting human body parts of a target human body in an image. In the related art, in order to implement human body part segmentation through a neural network model, a large number of labeled training sample data are required to train the neural network model. However, the part labels of the sample images need to be manually labeled, which is a large workload and low efficiency, and it is difficult to obtain a large amount of sample data, resulting in the difficulty of accurately implementing human body part segmentation by the trained model.

[0063] To solve the above problems, the present disclosure provides an image processing method, apparatus, storage medium, and electronic device, which can expand a first to-be-determined sample data set including a first sample image and a first human key point label to obtain a target sample data set including a target sample image and a target part label, train a target human image segmentation model according to the target sample data set, and input the to-be-processed image into the target human image segmentation model to segment the parts of the target human body to obtain a target image after part segmentation. Thereby, the acquisition efficiency of the training sample data is improved, the sample data is efficiently expanded, and the accuracy of the trained target human image segmentation model for human body part segmentation is improved.

[0064] The following will describe in detail the specific embodiments of the present disclosure with reference to the accompanying drawings.

[0065] Figure 1 is a flowchart of an image processing method provided by an embodiment of the present disclosure. This method can be applied to an electronic device, which can include a terminal device, such as a smart phone, a smart wearable device, a smart speaker, a smart tablet, a PDA (Personal Digital Assistant), a CPE (Customer Premise Equipment), a personal computer, etc., or can also include a server, such as a local server or a cloud server. As Figure 1 shown, this method includes:

[0066] S101. Obtain a to-be-processed image including a target human body.

[0067] Among them, the to-be-processed image may include the entire target human body, or may only include a part of the target human body, such as the upper body, the lower body, the head, the hand, or the foot, etc. The target human body therein may be frontal, back, or side. The to-be-processed image may be a picture including the target human body, or may be a video including the target human body. The present disclosure does not limit the type of the to-be-processed image.

[0068] In this step, the electronic device may acquire the to-be-processed image in real time through a camera device, or may obtain the pre-stored to-be-processed image, or may also receive the to-be-processed image sent by other devices. The present disclosure does not limit the acquisition method of the to-be-processed image.

[0069] S102. Input the to-be-processed image into a target human body image segmentation model to perform part segmentation on the target human body, so as to obtain a target image after part segmentation.

[0070] Among them, the target human body image segmentation model is trained according to a target sample data set, and the target sample data set includes a plurality of target sample images and target part labels corresponding to each target sample image. The target sample image includes a first sample image in a first to-be-determined sample data set, and the target part label is obtained according to the first to-be-determined sample data set and a second to-be-determined sample data set; the first to-be-determined sample data set includes a plurality of first sample images and first human body key point labels corresponding to each first sample image; the second to-be-determined sample data set includes a plurality of second sample images and second human body key point labels and second part labels corresponding to each second sample image.

[0071] The above-mentioned first body part label and second body part label are both human body part labels. The human body part label can include a head label, limb labels, and a torso label, etc. The limb labels can further include a hand label, an arm label, a leg label, and a foot label, etc. The above-mentioned human body part label needs to perform pixel-level annotation on the sample image, and the cost of pixel-level human body part annotation is very high. Therefore, it is difficult to obtain a sufficient amount of data for training the network by relying on manual annotation. The above-mentioned first human key point label and second human key point label are both human key point labels. The human key point label can be an annotation for a specific key part, which is sparse and does not require pixel-level annotation. Therefore, it is relatively easy to obtain.

[0072] Exemplarily, the above-mentioned first to-be-determined sample data set can include the first sample images annotated with the first human key point labels obtained from an existing database, or the first human key point labels can be obtained by annotating the obtained first sample images through a key point annotation model in related technologies. The above-mentioned second to-be-determined sample data set can be the second sample images partially annotated with the second human key point labels, and further annotated with the second body part labels. For example, the second body part labels can be manually annotated, or can be first annotated by a model and then corrected manually to improve the accuracy of the second body part labels in the second to-be-determined sample data set.

[0073] By adopting the above method, a to-be-processed image containing a target human body is obtained; the to-be-processed image is input into a target human body image segmentation model to perform part segmentation on the target human body to obtain a target image after part segmentation; wherein, the target human body image segmentation model is trained according to a target sample data set, the target sample data set includes a plurality of target sample images, and the target part label corresponding to each target sample image. The target sample image includes the first sample images in the first to-be-determined sample data set, and the target part label is obtained according to the first to-be-determined sample data set and the second to-be-determined sample data set; the first to-be-determined sample data set includes a plurality of first sample images, and the first human key point label corresponding to each first sample image; the second to-be-determined sample data set includes a plurality of second sample images, and the second human key point label and the second body part label corresponding to each second sample image. In this way, it can be expanded based on the sample (the first to-be-determined sample) containing human key points. On the basis of the first to-be-determined sample data set including human key point labels, a target sample data set including target part labels can be expanded, thereby improving the acquisition efficiency of training sample data, efficiently expanding sample data, and improving the accuracy of the target human body image segmentation model for human body part segmentation after training.

[0074] Figure 2 It is a flowchart of a training method for a target human body image segmentation model provided by an embodiment of the present disclosure. As Figure 2 shown, the training method can include:

[0075] S201. Obtain a first to-be-determined sample data set and a second to-be-determined sample data set.

[0076] Among them, the first to-be-determined sample data set includes a plurality of first sample images and first human key point labels corresponding to each first sample image. The second to-be-determined sample data set includes a plurality of second sample images, and second human key point labels and second part labels corresponding to each second sample image.

[0077] Exemplarily, a plurality of first sample images labeled with first human key point labels can be obtained from an existing database as the above-mentioned first to-be-determined sample data set; alternatively, after obtaining a plurality of first sample images, the first sample images obtained can be labeled by a key point labeling model in related technologies to obtain the first human key point labels corresponding to each first sample image.

[0078] It should be noted that the key point labeling model can be a human key point detection model in related technologies. For example, a 2D key point detection model (such as OpenPose, Hourglass, or Cascaded Pyramid Network) can be used to detect the human key point labels; alternatively, a 3D key point detection model (such as VideoPose3D, PoseDRL, or MovNect) can be used to detect the human key point labels.

[0079] Furthermore, the second part labels in the above-mentioned second to-be-determined sample data set can be manually labeled, or can be first labeled by a model and then manually corrected to improve the accuracy of the second part labels in the second to-be-determined sample data set.

[0080] S202. Obtain the first part label corresponding to the first sample image according to the second to-be-determined sample data set.

[0081] Exemplarily, a plurality of second to-be-determined sample images with the same first key point labels as the first sample image can be obtained from the second to-be-determined sample data set, the second part labels of the plurality of second to-be-determined sample images can be classified and counted, and the second part label with the largest proportion among them can be used as the first part label corresponding to the first sample image.

[0082] S203. Determine a target sample data set according to the first sample image and the first part label.

[0083] Exemplarily, the first sample image can be used as the target sample image in the target sample data set, and the first part label corresponding to the first sample image can be used as the target part label corresponding to the target sample image.

[0084] S204. Train the first preset neural network model according to the target sample dataset to obtain a target human body image segmentation model.

[0085] Exemplarily, the first preset neural network model may include a fully convolutional network model. The objective function for training the first preset neural network model may include the following formula (1):

[0086]

[0087] where I i represents the i-th target sample image in the target sample dataset, j represents the coordinate of the target pixel of the target sample image I i , S i (j) represents the target part label of the target pixel j, represents the training part label obtained by classifying the target pixel j through the first fully convolutional network, Φ represents the first model parameter of the first fully convolutional network, represents the loss function, which is used to calculate the single loss value (such as the difference) between the training part label and the target part label of a single target sample data, and ε(Φ) represents the classification result loss value of the objective function, which is used to characterize the comprehensive loss value (such as the average difference) between the training part labels and the target part labels of multiple target sample data in the target sample dataset.

[0088] Through the above method, the training samples can be expanded, so as to train a more accurate target human body image segmentation model.

[0089] In another embodiment of the present disclosure, the target sample dataset may further include the target human body key point label corresponding to the target sample image. For example, the first human body key point label corresponding to the first sample image may be used as the target human body key point label corresponding to the target sample image. In this way, during training, the target human body key point label can be used as an intermediate variable and a constraint item to assist in human body part recognition, which can improve the accuracy of the model in recognizing human body parts, and thus further improve the accuracy of the trained target human body image segmentation model in segmenting human body parts.

[0090] Figure 3 is a flowchart of a step S202 shown according to Figure 2 the embodiment shown. As Figure 2 shown, the above step S202 may include the following steps:

[0091] S2021. Determine the first human body posture corresponding to the first sample image and the second human body posture corresponding to the second sample image.

[0092] In this step, according to the above-mentioned first human key point labels, the first human pose corresponding to each first sample image in the above-mentioned first to-be-determined sample dataset can be obtained; according to the above-mentioned second human key point labels, the second human pose corresponding to each second sample image in the above-mentioned second to-be-determined sample dataset can be obtained.

[0093] Exemplarily, the above-mentioned first human key point labels can be used as the first human pose, and the above-mentioned second human key point labels can be used as the second human pose.

[0094] Furthermore, in order to calculate the similarity between different human poses, the above-mentioned human poses (the first human pose and the second human pose) can be regularized and corrected. Regularization is achieved by transforming the torso of the human body corresponding to the first sample image and the second sample image to the same size. The human key points of the torso part can be selected to calculate the torso length, and then regularization can be performed according to the torso length; then the hip key point can be selected as the reference coordinate to correct the above-mentioned human pose. In this way, it is convenient to perform operations on the above-mentioned first human pose and the second human pose.

[0095] S2022. Determine the first part label corresponding to the first sample image according to the first human pose and the second human pose.

[0096] Exemplarily, the first candidate part label corresponding to the first sample image can be obtained first according to the label acquisition step; then the first part label corresponding to the first sample image can be determined according to the first candidate part label.

[0097] In some embodiments, the label acquisition step may include the following steps:

[0098] S11. Determine one or more target pose sample images corresponding to the to-be-determined sample image from the second to-be-determined sample dataset according to the to-be-determined human pose corresponding to the to-be-determined sample image.

[0099] S12. Obtain the to-be-determined part label corresponding to the to-be-determined sample image according to the second part label corresponding to the target pose sample image.

[0100] Wherein, the similarity between the second human pose corresponding to the above-mentioned target pose sample image and the to-be-determined human pose is greater than or equal to a preset pose similarity threshold; the above-mentioned to-be-determined sample image may include the first sample image or the second sample image. In the case where the to-be-determined sample image includes the first sample image, the to-be-determined part label includes the first candidate part label corresponding to the first sample image; in the case where the to-be-determined sample image includes the second sample image, the to-be-determined part label includes the second candidate part label corresponding to the second sample image.

[0101] It should be noted that the above similarity can include Euclidean Distance, Standardized Euclidean Distance, Mahalanobis Distance or Cosine Distance. For example, taking the similarity including Euclidean Distance as an example, the similarity being greater than or equal to the preset similarity threshold can indicate that the Euclidean Distance is less than or equal to the preset distance threshold.

[0102] Taking the above to-be-determined sample image as the first sample image as an example, it is illustrated as follows:

[0103] If the first to-be-determined sample data set is The second to-be-determined sample data set is Among them, I i1 represents the first sample image, K i1 represents the first human body key point label corresponding to the first sample image I i1 I i2 represents the second sample image, S i2 represents the second part label (such as head label, torso label, limb label or other finer divisions) corresponding to the second sample image I i2 K i2 represents the second human body key point label corresponding to the second sample image I i2 M represents the number of first sample images in the first to-be-determined sample data set, N represents the number of second sample images in the second to-be-determined sample data set, M > N, M can be much larger than N. For example, N is 100 or 1000, and M can be 10000 or 100000.

[0104] In this way, the first human body pose Z i1 can be obtained according to the first human body key point label K i1 , and the second human body pose Z i2 can be obtained according to the second human body key point label K i2 , and the similarity (such as Euclidean Distance) between Z i1 and Z i2 is calculated. One or more second sample images corresponding to the second human body pose with the similarity greater than or equal to the preset pose similarity threshold are used as the target pose sample images, or a preset number of second sample images can be randomly selected from one or more second sample images corresponding to the second human body pose with the similarity greater than or equal to the preset pose similarity threshold as the target pose sample images.

[0105] In some embodiments, in step S12 above, the second part labels corresponding to one or more target pose sample images may be classified and counted, and the second part label with the largest proportion may be used as the to-be-determined part label corresponding to the to-be-determined sample image; alternatively, the average value of the second part labels of one or more target pose sample images may be used as the to-be-determined part label corresponding to the to-be-determined sample image.

[0106] In other embodiments, when the to-be-determined sample image is the first sample image, in step S12 above, the second part label corresponding to the target pose sample image may be subjected to an affine transformation according to the second human key point label corresponding to the target pose sample image and the first human key point label corresponding to the to-be-determined sample image, to obtain the target affine part label corresponding to the target pose sample image; and the to-be-determined part label corresponding to the to-be-determined sample image may be obtained according to the target affine part label.

[0107] Exemplarily, the target affine part label may be calculated according to the following formula (2):

[0108] SF j =T θ (S j2 ; θ j ) (2)

[0109] where SF j represents the target affine part label corresponding to the jth target pose sample image, S j2 represents the second part label corresponding to the target pose sample image, T θ (·) is an affine transformation function, θ j represents the target affine transformation parameter, and the target affine transformation parameter θ j can be calculated according to the first human key point label K t corresponding to the to-be-determined sample image and the second human key point label K j2 corresponding to the target pose sample image. Exemplarily, the target affine transformation parameter can be calculated by the following formula (3):

[0110] θ j =g(K t ,K j2 ) (3)

[0111] where θ j represents the target affine transformation parameter, K t represents the first human key point label corresponding to the to-be-determined sample image, K j2 represents the second human key point label corresponding to the target pose sample image, and g() represents a pre-designed calculation function. Exemplarily, the pre-designed calculation function can be difference calculation or ratio calculation.

[0112] In this way, the target affine part label corresponding to each target pose sample image can be calculated and obtained.

[0113] Furthermore, the average value of multiple target affine part labels can be used as the to-be-determined part label corresponding to this to-be-determined sample image. Exemplarily, the to-be-determined part label corresponding to this to-be-determined sample image can be calculated according to the following formula (4):

[0114]

[0115] where P t represents the to-be-determined part label corresponding to the to-be-determined sample image, SF j represents the target affine part label corresponding to the j-th target pose sample image, n represents the number of target pose sample images corresponding to this to-be-determined sample image, and j is less than or equal to n.

[0116] In this way, in the case where the above to-be-determined sample image includes the first sample image, through this label acquisition step, the first candidate part label corresponding to the first sample image can be obtained;

[0117] In some embodiments, the above manner of determining the first part label corresponding to the first sample image based on the first candidate part label may include any one of the following:

[0118] Method 1: Use the first candidate part label as the first part label corresponding to the first sample image.

[0119] Method 2: Calibrate the first candidate part label to obtain the first part label.

[0120] Exemplarily, the first sample image, the first human key point label, and the first candidate part label can be input into a pre-trained label calibration model to obtain the first part label corresponding to the first sample image.

[0121] Among them, the label calibration model is trained according to a second to-be-determined sample dataset.

[0122] Figure 4 is a flowchart of a training method for a label calibration model provided by an embodiment of the present disclosure. As Figure 4 shown, the training method may include:

[0123] S401. According to the label acquisition step, obtain the second candidate part label corresponding to each second sample data in the second to-be-determined sample dataset.

[0124] It should be noted that the to-be-determined sample data may include the second sample data, and the to-be-determined part label includes the second candidate part label corresponding to the second sample data. The description of this label acquisition step can refer to the description in the above embodiments and will not be elaborated here.

[0125] S402. Train the second preset neural network model according to the second candidate part label and the second to-be-determined sample data set to obtain a label calibration model.

[0126] The second preset neural network model may include a second fully convolutional network. The second loss function for training the second preset neural network model may include the following formula (5):

[0127]

[0128] where I m represents the m-th second sample image in the second sample data set, j represents the coordinate of the second pixel of the second sample image I m , S m (j) represents the second part label of the target pixel j of the m-th second sample image, represents the second training part label obtained by classifying the second pixel j through the second fully convolutional network, Ψ represents the second model parameter of the second fully convolutional network, and ε(Ψ) represents the loss value of the loss function.

[0129] In this way, the label calibration model obtained through training can calibrate the first candidate part label corresponding to the first sample image to obtain the first part label corresponding to the first sample image, thereby further improving the accuracy of the obtained first part label.

[0130] Figure 5 is a schematic structural diagram of an image processing device provided by an embodiment of the present disclosure. As Figure 5 shown, the device includes:

[0131] An image acquisition module 501, configured to acquire a to-be-processed image including a target human body;

[0132] An image processing module 502, configured to input the to-be-processed image into a target human body image segmentation model to perform part segmentation on the target human body to obtain a target image after part segmentation;

[0133] Among them, the target human body image segmentation model is trained according to a target sample data set, the target sample data set includes a plurality of target sample images, and a target part label corresponding to each target sample image, the target sample image includes a first sample image in a first to-be-determined sample data set, and the target part label is obtained according to the first to-be-determined sample data set and a second to-be-determined sample data set; the first to-be-determined sample data set includes a plurality of first sample images, and a first human body key point label corresponding to each first sample image; the second to-be-determined sample data set includes a plurality of second sample images, and a second human body key point label and a second part label corresponding to each second sample image.

[0134] Figure 6 is a schematic structural diagram of another image processing device provided by an embodiment of the present disclosure, as Figure 6 shown, the device further includes a first model training module 601:

[0135] The first model training module 601 is configured to obtain the first to-be-determined sample data set and the second to-be-determined sample data set; obtain a first part label corresponding to the first sample image according to the second to-be-determined sample data set; determine the target sample data set according to the first sample image and the first part label; and train a first preset neural network model according to the target sample data set to obtain the target human body image segmentation model.

[0136] Optionally, the first model training module 601 is configured to obtain a first human body posture corresponding to each first sample image in the first to-be-determined sample data set according to the first human body key point label; obtain a second human body posture corresponding to each second sample image in the second to-be-determined sample data set according to the second human body key point label; and determine a first part label corresponding to the first sample image according to the first human body posture and the second human body posture.

[0137] Optionally, the first model training module 601 is configured to obtain a first candidate part label corresponding to the first sample image according to a label obtaining step; and determine a first part label corresponding to the first sample image according to the first candidate part label;

[0138] Among them, the label obtaining step includes:

[0139] According to a to-be-determined human body posture corresponding to a to-be-determined sample image, determine one or more target posture sample images corresponding to the to-be-determined sample image from the second to-be-determined sample data set; a second human body posture corresponding to the target posture sample image has a similarity greater than or equal to a preset posture similarity threshold;

[0140] Obtain the to-be-determined part label corresponding to the to-be-determined sample image according to the second part label corresponding to the target pose sample image; the to-be-determined sample image includes the first sample image, and the to-be-determined part label includes the first candidate part label corresponding to the first sample image.

[0141] Optionally, the first model training module 601 is configured to use the average value of the second part labels of one or more of the target pose sample images as the to-be-determined part label corresponding to the to-be-determined sample image.

[0142] Optionally, the first model training module 601 is configured to input the first sample image, the first human key point label, and the first candidate part label into a pre-trained label calibration model to obtain the first part label corresponding to the first sample image; wherein, the label calibration model is trained according to a second to-be-determined sample data set.

[0143] Figure 7 It is a schematic structural diagram of another image processing device provided by an embodiment of the present disclosure. As Figure 7 shown, the device further includes a second model training module 701:

[0144] The to-be-determined sample data includes the second sample data, and the to-be-determined part label includes the second candidate part label corresponding to the second sample data;

[0145] The second model training module 701 is configured to obtain the second candidate part label corresponding to each second sample data in the second to-be-determined sample data set according to the label obtaining step; and train a second preset neural network model according to the second candidate part label and the second to-be-determined sample data set to obtain the label calibration model.

[0146] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0147] Figure 8 It is a block diagram of an electronic device 900 shown according to an exemplary embodiment. As Figure 8 shown, the electronic device 900 may include: a processor 901, a memory 902. The electronic device 900 may further include one or more of a multimedia component 903, an input / output (I / O) interface 904, and a communication component 905.

[0148] Among them, the processor 901 is used to control the overall operation of the electronic device 900 to complete all or part of the steps in the above image processing method. The memory 902 is used to store various types of data to support the operation of the electronic device 900. These data may include, for example, instructions for any application or method operating on the electronic device 900, as well as application-related data, such as contact data, received and sent messages, pictures, audio, video, etc. The memory 902 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The multimedia component 903 may include a screen and an audio component. Among them, the screen may be a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signal may be further stored in the memory 902 or sent through the communication component 905. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 904 provides an interface between the processor 901 and other interface modules. The above other interface modules may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 905 is used for wired or wireless communication between the electronic device 900 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, near field communication (NFC), 2G, 3G, 4G, 5G, NB-IOT, eMTC, or other 6G, etc., or a combination of one or more of them, is not limited here. Therefore, the corresponding communication component 905 may include: a Wi-Fi module, a Bluetooth module, an NFC module, etc.

[0149] In an exemplary embodiment, the electronic device 900 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described image processing method.

[0150] In another exemplary embodiment, there is also provided a computer-readable storage medium including program instructions, which implement the steps of the above-described image processing method when executed by a processor. For example, the computer-readable storage medium may be the above-described memory 902 including program instructions, and the above program instructions may be executed by the processor 901 of the electronic device 900 to complete the above-described image processing method. Exemplarily, the computer-readable storage medium may be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium may be ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage devices, etc.

[0151] In another exemplary embodiment, there is also provided a computer program product, which includes a computer program executable by a programmable device, and the computer program has a code portion for performing the above-described image processing method when executed by the programmable device.

[0152] The preferred embodiments of the present disclosure have been described in detail above in conjunction with the accompanying drawings. However, the present disclosure is not limited to the specific details in the above embodiments. Within the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the protection scope of the present disclosure.

[0153] In addition, it should be noted that, in the above specific embodiments, the various specific technical features described can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the present disclosure does not separately describe various possible combination manners.

[0154] Furthermore, any combination can be made between various different embodiments of the present disclosure as long as it does not violate the idea of the present disclosure, and it should also be regarded as the content disclosed by the present disclosure.

Claims

1. An image processing method, characterized in that, The method includes: Obtaining a to-be-processed image including a target human body; Inputting the to-be-processed image into a target human body image segmentation model to perform part segmentation on the target human body, so as to obtain a target image after part segmentation; Wherein, the target human body image segmentation model is trained according to a target sample data set, the target sample data set includes a plurality of target sample images and target part labels corresponding to each target sample image, the target sample images include first sample images in a first to-be-determined sample data set, and the target part labels are obtained according to the first to-be-determined sample data set and a second to-be-determined sample data set; the first to-be-determined sample data set includes a plurality of first sample images and first human body key point labels corresponding to each first sample image; the second to-be-determined sample data set includes a plurality of second sample images and second human body key point labels and second part labels corresponding to each second sample image; The target human body image segmentation model is trained through the following manner: Obtaining the first to-be-determined sample data set and the second to-be-determined sample data set; Obtaining first part labels corresponding to the first sample images according to the second to-be-determined sample data set; Determining the target sample data set according to the first sample images and the first part labels; Training a first preset neural network model according to the target sample data set to obtain the target human body image segmentation model; The obtaining of the first part labels corresponding to the first sample images according to the second to-be-determined sample data set includes: Obtaining a first human body pose corresponding to each first sample image in the first to-be-determined sample data set according to the first human body key point labels; Obtaining a second human body pose corresponding to each second sample image in the second to-be-determined sample data set according to the second human body key point labels; Determining the first part labels corresponding to the first sample images according to the first human body pose and the second human body pose.

2. The method according to claim 1, characterized in that, The determining of the first part labels corresponding to the first sample images according to the first human body pose and the second human body pose includes: Obtaining first candidate part labels corresponding to the first sample images according to a label obtaining step; Determining the first part labels corresponding to the first sample images according to the first candidate part labels; Wherein, the label obtaining step includes: Determining one or more target pose sample images corresponding to the to-be-determined sample image from the second to-be-determined sample data set according to the to-be-determined human body pose corresponding to the to-be-determined sample image; the second human body pose corresponding to the target pose sample image has a similarity greater than or equal to a preset pose similarity threshold with the to-be-determined human body pose; Obtaining to-be-determined part labels corresponding to the to-be-determined sample image according to the second part labels corresponding to the target pose sample images; the to-be-determined sample image includes the first sample images, and the to-be-determined part labels include the first candidate part labels corresponding to the first sample images.

3. The method according to claim 2, characterized in that, Obtaining the to-be-determined part label corresponding to the to-be-determined sample image according to the second part label corresponding to the target pose sample image includes: Taking the average value of the second part labels of one or more of the target pose sample images as the to-be-determined part label corresponding to the to-be-determined sample image.

4. The method according to claim 2, wherein Determining the first part label corresponding to the first sample image according to the first candidate part label includes: Inputting the first sample image, the first human key point label, and the first candidate part label into a pre-trained label calibration model to obtain the first part label corresponding to the first sample image; wherein, the label calibration model is trained according to a second to-be-determined sample data set.

5. The method according to claim 4, characterized in that, The to-be-determined sample image includes the second sample image, and the to-be-determined part label includes the second candidate part label corresponding to the second sample image; the label calibration model is trained in the following manner: Obtaining the second candidate part label corresponding to each second sample image in the second to-be-determined sample data set according to the label obtaining step. Training a second preset neural network model according to the second candidate part label and the second to-be-determined sample data set to obtain the label calibration model.

6. An image processing apparatus, characterized in that, The device includes: An image acquisition module, configured to acquire a to-be-processed image including a target human body. An image processing module, configured to input the to-be-processed image into a target human body image segmentation model to perform part segmentation on the target human body to obtain a target image after part segmentation. Wherein, the target human body image segmentation model is trained according to a target sample data set, the target sample data set includes a plurality of target sample images, and a target part label corresponding to each target sample image, the target sample image includes a first sample image in a first to-be-determined sample data set, and the target part label is obtained according to the first to-be-determined sample data set and a second to-be-determined sample data set; the first to-be-determined sample data set includes a plurality of first sample images, and a first human key point label corresponding to each first sample image; the second to-be-determined sample data set includes a plurality of second sample images, and a second human key point label and a second part label corresponding to each second sample image. The target human body image segmentation model is trained in the following manner: Obtaining the first to-be-determined sample data set and the second to-be-determined sample data set; obtaining the first part label corresponding to the first sample image according to the second to-be-determined sample data set; determining the target sample data set according to the first sample image and the first part label; training a first preset neural network model according to the target sample data set to obtain the target human body image segmentation model. Obtaining the first part label corresponding to the first sample image according to the second to-be-determined sample data set includes: Obtain the first human body pose corresponding to each of the first sample images in the first to-be-determined sample dataset according to the first human body key point labels; obtain the second human body pose corresponding to each of the second sample images in the second to-be-determined sample dataset according to the second human body key point labels; determine the first part label corresponding to the first sample image according to the first human body pose and the second human body pose.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method according to any one of claims 1 to 5.

8. An electronic device, characterized in that, Comprising: A memory storing a computer program thereon; A processor for executing the computer program in the memory to implement the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Semi-supervised machine learning optimization method, device and equipment and storage medium

    CN111222648A

  • Human body image segmentation method, and training method and device of human body image segmentation model

    CN111932568A