Self-Supervised Human Pose Transformation Method and System, Readable Storage Medium
Feature decoupling and fusion are performed through the convolutional neural network's pose feature encoder and decoupling style encoder, combining spatial correlation learning and U-Net image converter, the feature alignment problem in self-supervised human posture conversion is solved, and efficient pose conversion of single human image is achieved, improving the robustness and generalization of the algorithm.
Patent Information
- Application Number
- CN202211665771.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-23
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-12-23
AI Technical Summary
The existing self-supervised human posture conversion algorithm has alignment problems in feature decoupling and fusion methods, resulting in poor large-scale posture conversion effect and lack prior knowledge of invisible areas, which reduces the robustness and generalization of the algorithm.
The pose feature encoder and decoupling style encoder based on convolutional neural network are used for feature extraction and decoupling. The feature correlation is calculated through the cross-channel fusion module and the correlation mining module of spatial correlation learning, and the dense spatial correlation field is constructed for non-rigid deformation, combined with the U-Net image converter for reconstruction, and the image discriminator and loss function are used for model training.
Self-supervised attitude conversion for a single human image is realized, reducing the cost of data acquisition and model training, and improving the effect and robustness of attitude conversion.
Smart Images

Figure CN116416677B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and in particular, to a self-supervised human pose conversion method and system, and a readable storage medium. Background Art
[0002] In recent years, with the continuous development of artificial intelligence technologies, significant progress has been made in image content generation based on deep learning. Among them, the generation of human images with multiple poses (i.e., human pose conversion) has been widely applied in many fields, such as film production, multimedia entertainment, pedestrian re-identification, etc. The purpose of the human pose conversion algorithm is to change the pose of a human image given a target pose. This can be regarded as a non-aligned image-to-image conversion problem, which requires a non-rigid transformation of the original image in the feature space to achieve pose conversion. Currently, human pose conversion algorithms can be divided into supervised methods and self-supervised methods according to different training data. Supervised human pose conversion algorithms need paired data for training. However, in a real deployment environment, obtaining paired data requires collecting images of the same person in different poses, and the data collection cost is relatively high, which is not conducive to the actual implementation of the algorithm. And designing an effective self-supervised pose conversion method mainly lies in designing effective feature decoupling and feature fusion methods. The decoupled features obtained by existing feature decoupling methods still have a certain degree of alignment, and cannot provide sufficient supervision information for self-supervised algorithms to train, resulting in poor effects on large-scale pose conversion. Existing feature fusion methods, such as feature splicing, statistic transfer, etc., are all global operations on feature maps, and it is difficult to achieve non-rigid transformation of features. At the same time, due to the lack of prior knowledge of invisible regions in the self-supervised algorithm for the pose conversion from half body to full body, it is difficult to complete the invisible human regions, reducing the robustness and generalization of the algorithm. Summary of the Invention
[0003] This application aims to solve or improve the above technical problems.
[0004] To this end, the first object of this application is to provide a self-supervised human pose conversion method.
[0005] The second object of this application is to provide a self-supervised human pose conversion system.
[0006] The third object of this application is to provide a self-supervised human pose conversion system.
[0007] The fourth object of this application is to provide a readable storage medium.
[0008] To achieve the first object of the present application, the technical solution of the first aspect of the present application provides a self-supervised human pose conversion method, including: obtaining an input human body image, and obtaining a human pose skeleton map and a human body parsing map according to the input human body image; extracting features from the human pose skeleton map through a pose feature encoder to obtain human pose features, and the pose feature encoder includes a pose feature encoder based on a convolutional neural network; extracting features from the human body parsing map through a decoupled style encoder to obtain human body part style features, and the decoupled style encoder includes a decoupled style encoder based on a convolutional neural network; fusing the human body part style features through a cross-channel fusion module; calculating the feature correlation between different positions according to the human pose features and the human body part style features through a correlation mining module based on spatial correlation learning; constructing a dense spatial correlation field according to the feature correlation; performing non-rigid deformation on the human body part style features based on the dense spatial correlation field to obtain recombined style features; reconstructing the input human body image through an image converter according to the recombined style features to obtain a reconstructed human body image.
[0009] According to the self-supervised human pose conversion method provided by the present application, first, an input human body image is obtained, and a human pose skeleton map and a human body parsing map are obtained according to the input human body image. The pose feature encoder based on a convolutional neural network extracts human pose features from the human pose skeleton map. The decoupled style encoder based on a convolutional neural network extracts human body part style features from each individual human body part image. The cross-channel fusion module based on 1×1 convolution fuses the decoupled human body part style features. The correlation mining module based on spatial correlation learning calculates the feature correlation between different positions according to the decoupled human pose features and the human body part style features. According to the feature correlation, a dense spatial correlation field is constructed. Based on the dense spatial correlation field, non-rigid deformation is performed on the human body style features to obtain recombined style features. The image converter based on U-Net reconstructs the input human body image according to the recombined style features. The self-supervised human pose conversion ability is realized through the pose encoder, the decoupled style encoder, the correlation mining module and the image converter. It does not rely on paired multi-pose human body data sets for supervised training, but reconstructs a single human body image by extracting and fusing non-aligned features, so as to realize the supervision process of model training, effectively reducing the costs of data collection and model training.
[0010] Specifically, the pose encoder and the decoupled style encoder enable non-aligned decoupled expression of the human body image in the feature space. The correlation mining module calculates the correlation between non-aligned decoupled features and recombines the style features, so as to realize the fusion of non-aligned features. The image converter converts the fused features into a real human body image to complete the reconstruction of the human body image.
[0011] In addition, the technical solution provided by this application may also have the following additional technical features:
[0012] In the above technical solution, the self-supervised human pose conversion method further includes: using an image discriminator to judge a real human body image and a reconstructed human body image, and calculating a loss function value.
[0013] In this technical solution, the self-supervised human pose conversion method further includes using an image discriminator to judge a real human body image and a reconstructed human body image, and calculating a loss function value, which can effectively train the model through the constraint of the loss function.
[0014] In the above technical solution, using an image discriminator to judge a real human body image and a reconstructed human body image, and calculating a loss function value specifically includes: using a pose discriminator based on a convolutional neural network to judge a real human body image and its pose skeleton image as a positive example pair, and judging a reconstructed human body image and its pose skeleton image as a negative example pair, and calculating a loss function value; using a style discriminator based on a convolutional neural network to judge a real human body image as a positive example and a reconstructed human body image as a negative example, and calculating a loss function value.
[0015] In this technical solution, using an image discriminator to judge a real human body image and a reconstructed human body image, and calculating a loss function value specifically means using a pose discriminator based on a convolutional neural network to judge a real human body image and its pose skeleton image as a positive example pair, and judging a reconstructed human body image and its pose skeleton image as a negative example pair, and calculating a loss function value. Using a style discriminator based on a convolutional neural network to judge a real human body image as a positive example and a reconstructed human body image as a negative example, and calculating a loss function value.
[0016] In the above technical solution, the self-supervised human pose conversion method further includes: a human body graph generator based on a pre-trained VGG network and a regional average pooling layer, obtaining a human body structure graph representing the human body structure according to the input human body image and the human body parsing graph; calculating a loss function value according to the input human body image, the reconstructed human body image and the human body structure graph.
[0017] In this technical solution, the self-supervised human pose conversion method further includes obtaining a human body structure diagram based on a human body feature generator. The loss function value is calculated based on the input human body image, the reconstructed human body image, and the human body structure diagram, and the effective training of the model is achieved through the constraints of multiple loss functions. Specifically, the human body feature generator consists of a pre-trained VGG network and a regional average pooling layer. The pre-trained VGG network extracts the perceptual feature map of the human body image, and the regional average pooling layer separates the perceptual feature map into multiple regions according to the human body parsing diagram, and extracts the feature vectors of each region through average pooling operations. Among them, each human body region feature vector serves as a node of the human body structure diagram, and the cosine similarity between two feature vectors serves as the edge of the human body structure diagram.
[0018] In the above technical solution, the loss function includes one or a combination of the following: human body image reconstruction loss, human body image perception loss, human body image style loss, human body structure preservation loss based on graph features, and adversarial training loss.
[0019] In this technical solution, the loss function includes human body image reconstruction loss, human body image perception loss, human body image style loss, human body structure preservation loss based on graph features, and adversarial training loss. Among them, the adversarial training loss is divided into pose discrimination loss and style discrimination loss, which are calculated by a pose discriminator and a style discriminator respectively.
[0020] In the above technical solution, the self-supervised human pose conversion method further includes: according to the loss function value, iteratively adjusting the weights of the pose feature encoder, the decoupled style encoder, the image converter, and the image discriminator through the loss gradient backpropagation algorithm until convergence.
[0021] In this technical solution, the self-supervised human pose conversion method further includes iteratively adjusting the weights of the pose feature encoder, the decoupled style encoder, the image converter, and the image discriminator according to the loss function value by using the loss gradient backpropagation algorithm until convergence.
[0022] In the above technical solution, an input human body image is obtained, and a human body pose skeleton diagram and a human body parsing diagram are obtained according to the input human body image. Specifically, it includes: obtaining an input human body image; obtaining a human body pose skeleton diagram according to the input human body image through a preset pose estimation method; obtaining a human body parsing diagram according to the input human body image through a preset human body parsing method.
[0023] In this technical solution, obtaining a human body pose skeleton diagram and a human body parsing diagram specifically means obtaining an input human body image, obtaining the human body pose skeleton diagram of the input human body image based on a pre-constructed pose estimation method, and obtaining the human body parsing diagram of the input human body image based on a pre-constructed human body parsing method.
[0024] To achieve the second object of the present application, the technical solution of the second aspect of the present application provides a self-supervised human pose conversion system, including: an acquisition module, configured to acquire an input human body image and obtain a human pose skeleton map and a human body parsing map according to the input human body image; a pose feature extraction module, configured to extract features from the human pose skeleton map through a pose feature encoder to obtain human pose features, and the pose feature encoder includes a pose feature encoder based on a convolutional neural network; a style feature extraction module, configured to extract features from the human body parsing map through a decoupled style encoder to obtain human body part style features, and the decoupled style encoder includes a decoupled style encoder based on a convolutional neural network; a style feature fusion module, configured to fuse the human body part style features through a cross-channel fusion module; a correlation calculation module, configured to calculate the feature correlation between different positions according to the human pose features and the human body part style features through a correlation mining module based on spatial correlation learning; a correlation field construction module, configured to construct a dense spatial correlation field according to the feature correlation; a recombination module, configured to perform non-rigid deformation on the human body part style features based on the dense spatial correlation field to obtain recombined style features; a reconstruction module, configured to reconstruct the input human body image through an image converter according to the recombined style features to obtain a reconstructed human body image.
[0025] The self-supervised human pose conversion system provided by the present application includes an acquisition module, a pose feature extraction module, a style feature extraction module, a style feature fusion module, a correlation calculation module, a correlation field construction module, a recombination module, and a reconstruction module. Among them, the acquisition module is used to acquire an input human image and obtain a human pose skeleton map and a human parsing map based on the input human image. The pose feature extraction module is used to extract features from the human pose skeleton map through a pose feature encoder to obtain human pose features. The pose feature encoder includes a pose feature encoder based on a convolutional neural network. The style feature extraction module is used to extract features from the human parsing map through a decoupled style encoder to obtain human part style features. The decoupled style encoder includes a decoupled style encoder based on a convolutional neural network. The style feature fusion module is used to fuse the human part style features through a cross-channel fusion module. The correlation calculation module is used to calculate the feature correlation between different positions according to the human pose features and the human part style features through a correlation mining module based on spatial correlation learning. The correlation field construction module is used to construct a dense spatial correlation field according to the feature correlation. The recombination module is used to perform non-rigid deformation on the human part style features based on the dense spatial correlation field to obtain recombined style features. The reconstruction module is used to reconstruct the input human image through an image converter according to the recombined style features to obtain a reconstructed human image. Through the pose encoder, the decoupled style encoder, the correlation mining module, and the image converter, the ability of self-supervised human pose conversion is realized. It does not rely on paired multi-pose human datasets for supervised training, but reconstructs a single human image by extracting and fusing unaligned features, thereby realizing the supervision process of model training and effectively reducing the costs of data acquisition and model training.
[0026] To achieve the third object of the present application, the technical solution of the third aspect of the present application provides a self-supervised human pose conversion system, including: a memory and a processor. Among them, a program or instruction that can run on the processor is stored on the memory. When the processor executes the program or instruction, it implements the self-supervised human pose conversion method in any one of the technical solutions of the first aspect. Therefore, it has the technical effects of any one of the technical solutions in the first aspect, which will not be elaborated here.
[0027] To achieve the fourth object of the present application, the technical solution of the fourth aspect of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements the steps of the self-supervised human pose conversion method in any one of the technical solutions of the first aspect. Therefore, it has the technical effects of any one of the technical solutions in the first aspect, which will not be elaborated here.
[0028] The additional aspects and advantages of the present application will become obvious in the following description section, or be learned through the practice of the present application. Description of the Drawings
[0029] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of embodiments in conjunction with the accompanying drawings, where:
[0030] Figure 1 It is a schematic flowchart of the steps of a self-supervised human pose conversion method according to an embodiment of the present application;
[0031] Figure 2 It is a schematic flowchart of the steps of a self-supervised human pose conversion method according to an embodiment of the present application;
[0032] Figure 3 It is a schematic flowchart of the steps of a self-supervised human pose conversion method according to an embodiment of the present application;
[0033] Figure 4 It is a schematic flowchart of the steps of a self-supervised human pose conversion method according to an embodiment of the present application;
[0034] Figure 5 It is a schematic flowchart of the steps of a self-supervised human pose conversion method according to an embodiment of the present application;
[0035] Figure 6 It is a schematic flowchart of the steps of a self-supervised human pose conversion method according to an embodiment of the present application;
[0036] Figure 7 It is a schematic block diagram of the structure of a self-supervised human pose conversion system according to an embodiment of the present application;
[0037] Figure 8 It is a schematic block diagram of the structure of a self-supervised human pose conversion system according to another embodiment of the present application;
[0038] Figure 9 It is a schematic flowchart of the steps of a self-supervised human pose conversion method according to an embodiment of the present application.
[0039] Wherein, Figure 7 and Figure 8 The corresponding relationship between the reference numerals and component names in is:
[0040] 10: Self-supervised human pose conversion system; 110: Acquisition module; 120: Pose feature extraction module; 130: Style feature extraction module; 140: Style feature fusion module; 150: Correlation calculation module; 160: Correlation field construction module; 170: Recombination module; 180: Reconstruction module; 20: Self-supervised human pose conversion system; 300: Memory; 400: Processor. Detailed implementation manners
[0041] To better understand the above objects, features, and advantages of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other.
[0042] In the following description, many specific details are set forth to facilitate a thorough understanding of the present application. However, the present application may be implemented in other ways different from those described herein. Therefore, the scope of protection of the present application is not limited by the specific embodiments disclosed below.
[0043] The following refers to Figures 1 to 9 Describe a self-supervised human pose conversion method, system, and readable storage medium according to some embodiments of the present application.
[0044] As Figure 1 shown, an embodiment of the first aspect of the present application provides a self-supervised human pose conversion method, including the following steps:
[0045] Step S102: Obtain an input human image, and obtain a human pose skeleton diagram and a human parsing diagram according to the input human image;
[0046] Step S104: Extract features from the human pose skeleton diagram through a pose feature encoder to obtain human pose features. The pose feature encoder includes a pose feature encoder based on a convolutional neural network;
[0047] Step S106: Extract features from the human parsing diagram through a decoupled style encoder to obtain human part style features. The decoupled style encoder includes a decoupled style encoder based on a convolutional neural network;
[0048] Step S108: Fuse the human part style features through a cross-channel fusion module;
[0049] Step S110: Calculate the feature correlation between different positions according to the human pose features and the human part style features through a correlation mining module based on spatial correlation learning;
[0050] Step S112: Construct a dense spatial correlation field according to the feature correlation;
[0051] Step S114: Perform non-rigid deformation on the human part style features based on the dense spatial correlation field to obtain reorganized style features;
[0052] Step S116: Reconstruct the input human image through an image converter according to the reorganized style features to obtain a reconstructed human image.
[0053] According to the self-supervised human pose conversion method provided by this embodiment, first, an input human image is obtained, and a human pose skeleton map and a human parsing map are obtained based on the input human image. Based on the pose feature encoder of the convolutional neural network, human pose features are extracted from the human pose skeleton map. Based on the decoupled style encoder of the convolutional neural network, human part style features are extracted from each individual human part image. Based on the cross-channel fusion module of 1×1 convolution, the decoupled human part style features are fused. Based on the correlation mining module of spatial correlation learning, according to the decoupled human pose features and human part style features, the feature correlations between different positions are calculated. According to the feature correlations, a dense spatial correlation field is constructed. Based on the dense spatial correlation field, non-rigid deformation is performed on the human style features to obtain recombined style features. Based on the U-Net image converter, the input human image is reconstructed according to the recombined style features. Through the pose encoder, the decoupled style encoder, the correlation mining module, and the image converter, the ability of self-supervised human pose conversion is realized. It does not rely on paired multi-pose human datasets for supervised training, but reconstructs a single human image by extracting and fusing unaligned features, thereby realizing the supervision process of model training, effectively reducing the costs of data acquisition and model training.
[0054] Specifically, the pose encoder and the decoupled style encoder enable the unaligned decoupled expression of the human image in the feature space. The correlation mining module calculates the correlations between the unaligned decoupled features and recombines the style features, thereby realizing the fusion of unaligned features. The image converter converts the fused features into a real human image to complete the reconstruction of the human image.
[0055] As Figure 2 shown, according to a self-supervised human pose conversion method of an embodiment proposed by this application, the following steps are further included:
[0056] Step S202: Judge the real human image and the reconstructed human image through an image discriminator, and calculate the loss function value.
[0057] In this embodiment, the self-supervised human pose conversion method further includes judging the real human image and the reconstructed human image through an image discriminator and calculating the loss function value, which can realize the effective training of the model through the constraint of the loss function.
[0058] As Figure 3 shown, according to a self-supervised human pose conversion method of an embodiment proposed by this application, judging the real human image and the reconstructed human image through an image discriminator and calculating the loss function value specifically include the following steps:
[0059] Step S302: Use a pose discriminator based on a convolutional neural network to determine the real human body image and its pose skeleton diagram as a positive example pair, and the reconstructed human body image and its pose skeleton diagram as a negative example pair, and calculate the loss function value;
[0060] Step S304: Use a style discriminator based on a convolutional neural network to determine the real human body image as a positive example and the reconstructed human body image as a negative example, and calculate the loss function value.
[0061] In this embodiment, an image discriminator is used to judge the real human body image and the reconstructed human body image, and calculate the loss function value. Specifically, a pose discriminator based on a convolutional neural network is used to determine the real human body image and its pose skeleton diagram as a positive example pair, and the reconstructed human body image and its pose skeleton diagram as a negative example pair, and calculate the loss function value. A style discriminator based on a convolutional neural network is used to determine the real human body image as a positive example and the reconstructed human body image as a negative example, and calculate the loss function value.
[0062] As Figure 4 shown, the self-supervised human pose conversion method according to an embodiment proposed by the present application further includes the following steps:
[0063] Step S402: Based on a human body feature generator of a pre-trained VGG network and a regional average pooling layer, obtain a human body structure diagram representing the human body structure according to the input human body image and the human body parsing diagram;
[0064] Step S404: Calculate the loss function value according to the input human body image, the reconstructed human body image, and the human body structure diagram.
[0065] In this embodiment, the self-supervised human pose conversion method further includes obtaining a human body structure diagram based on a human body feature generator. Calculate the loss function value based on the input human body image, the reconstructed human body image, and the human body structure diagram, and effectively train the model through the constraints of multiple loss functions. Specifically, the human body feature generator consists of a pre-trained VGG network and a regional average pooling layer. The pre-trained VGG network extracts the perceptual feature map of the human body image, and the regional average pooling layer separates the perceptual feature map into multiple regions according to the human body parsing diagram, and extracts the feature vectors of each region through average pooling operations. Among them, each human body region feature vector serves as a node of the human body structure diagram, and the cosine similarity between two feature vectors serves as the edge of the human body structure diagram.
[0066] In the above embodiment, the loss function includes a human body image reconstruction loss, a human body image perception loss, a human body image style loss, a human body structure preservation loss based on graph representation, and an adversarial training loss. Among them, the adversarial training loss is divided into a pose discrimination loss and a style discrimination loss, which are calculated by a pose discriminator and a style discriminator respectively.
[0067] As shown Figure 5 in the figure, the self-supervised human pose conversion method according to an embodiment of the present application further includes the following steps:
[0068] Step S502: According to the loss function value, iteratively adjust the weights of the pose feature encoder, decoupled style encoder, image converter, and image discriminator through the loss gradient backpropagation algorithm until convergence.
[0069] In this embodiment, the self-supervised human pose conversion method further includes iteratively adjusting the weights of the pose feature encoder, decoupled style encoder, image converter, and image discriminator according to the loss function value through the loss gradient backpropagation algorithm until convergence.
[0070] As shown Figure 6 in the figure, the self-supervised human pose conversion method according to an embodiment of the present application acquires an input human image and obtains a human pose skeleton map and a human parsing map according to the input human image, which specifically includes the following steps:
[0071] Step S602: Acquire the input human image;
[0072] Step S604: Obtain a human pose skeleton map according to the input human image through a preset pose estimation method;
[0073] Step S606: Obtain a human parsing map according to the input human image through a preset human parsing method.
[0074] In this embodiment, to obtain the human pose skeleton map and the human parsing map, specifically, an input human image is acquired, the human pose skeleton map of the input human image is obtained based on a pre-constructed pose estimation method, and the human parsing map of the input human image is obtained based on a pre-constructed human parsing method.
[0075] As shown Figure 7As shown in the figure, an embodiment of the second aspect of the present application provides a self-supervised human pose conversion system 10, including: an acquisition module 110, configured to acquire an input human image and obtain a human pose skeleton map and a human parsing map according to the input human image; a pose feature extraction module 120, configured to extract features from the human pose skeleton map through a pose feature encoder to obtain human pose features, and the pose feature encoder includes a pose feature encoder based on a convolutional neural network; a style feature extraction module 130, configured to extract features from the human parsing map through a decoupled style encoder to obtain human part style features, and the decoupled style encoder includes a decoupled style encoder based on a convolutional neural network; a style feature fusion module 140, configured to fuse the human part style features through a cross-channel fusion module; a correlation calculation module 150, configured to calculate the feature correlation between different positions according to the human pose features and the human part style features through a correlation mining module based on spatial correlation learning; a correlation field construction module 160, configured to construct a dense spatial correlation field according to the feature correlation; a recombination module 170, configured to perform non-rigid deformation on the human part style features based on the dense spatial correlation field to obtain recombined style features; and a reconstruction module 180, configured to reconstruct the input human image through an image converter according to the recombined style features to obtain a reconstructed human image.
[0076] According to the self-supervised human pose conversion system 10 provided in this embodiment, it includes an acquisition module 110, a pose feature extraction module 120, a style feature extraction module 130, a style feature fusion module 140, a correlation calculation module 150, a correlation field construction module 160, a recombination module 170, and a reconstruction module 180. Among them, the acquisition module 110 is used to acquire an input human body image, and obtain a human body pose skeleton map and a human body parsing map according to the input human body image. The pose feature extraction module 120 is used to extract features from the human body pose skeleton map through a pose feature encoder to obtain human body pose features. The pose feature encoder includes a pose feature encoder based on a convolutional neural network. The style feature extraction module 130 is used to extract features from the human body parsing map through a decoupled style encoder to obtain human body part style features. The decoupled style encoder includes a decoupled style encoder based on a convolutional neural network. The style feature fusion module 140 is used to fuse the human body part style features through a cross-channel fusion module. The correlation calculation module 150 is used to calculate the feature correlation between different positions according to the human body pose features and the human body part style features through a correlation mining module based on spatial correlation learning. The correlation field construction module 160 is used to construct a dense spatial correlation field according to the feature correlation. The recombination module 170 is used to perform non-rigid deformation on the human body part style features based on the dense spatial correlation field to obtain recombined style features. The reconstruction module 180 is used to reconstruct the input human body image through an image converter according to the recombined style features to obtain a reconstructed human body image. Through the pose encoder, the decoupled style encoder, the correlation mining module, and the image converter, the ability of self-supervised human pose conversion is realized. It does not rely on a paired multi-pose human body data set for supervised training, but reconstructs a single human body image by extracting and fusing unaligned features, so as to realize the supervision process of model training, effectively reducing the cost of data acquisition and model training.
[0077] As Figure 8 shown, an embodiment of the third aspect of the present application provides a self-supervised human pose conversion system 20, including: a memory 300 and a processor 400. Among them, a program or instruction that can run on the processor 400 is stored on the memory 300. When the processor 400 executes the program or instruction, it implements the steps of the self-supervised human pose conversion method in any one of the embodiments of the first aspect. Therefore, it has the technical effects of any one of the above-mentioned first aspect embodiments and will not be elaborated here.
[0078] An embodiment of the fourth aspect of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements the steps of the self-supervised human pose conversion method in any one of the embodiments of the first aspect. Therefore, it has the technical effects of any one of the above-mentioned first aspect embodiments and will not be elaborated here.
[0079] As Figure 9 shown, according to a self-supervised human pose conversion method provided by the present application, the overall architecture consists of an image generator and an image discriminator. The image generator includes a pose feature encoder based on a convolutional neural network, a decoupled style encoder based on a convolutional neural network, an image transformer based on U-Net, and a human body representation generator. The image discriminator includes a pose discriminator and a style discriminator.
[0080] The self-supervised human pose conversion method includes:
[0081] Obtaining a human pose skeleton map of the input human image based on a pre-constructed pose estimation method;
[0082] Extracting human pose features from the human pose skeleton map by a pose feature encoder based on a convolutional neural network;
[0083] Obtaining a human body parsing map of the input human image based on a pre-constructed human body parsing method;
[0084] Extracting style features from each individual human body part image by a decoupled style encoder based on a convolutional neural network;
[0085] Fusing the decoupled style features by a cross-channel fusion module based on 1×1 convolution;
[0086] Calculating the feature correlation between different positions according to the decoupled pose features and style features by a correlation mining module based on spatial correlation learning;
[0087] Constructing a dense spatial correlation field according to the feature correlation;
[0088] Non-rigidly deforming the human body style features based on the dense spatial correlation field to obtain reorganized style features;
[0089] Reconstructing the input human image according to the reorganized style features by an image transformer based on U-Net;
[0090] Judging the real human image and its pose skeleton map as a positive example pair, judging the reconstructed human image and its pose skeleton map as a negative example pair, and calculating the loss function value by a pose discriminator based on a convolutional neural network;
[0091] Judging the real human image as a positive example and the reconstructed human image as a negative example, and calculating the loss function value by a style discriminator based on a convolutional neural network;
[0092] Obtaining a human body structure diagram representing the human body structure according to the human image and the human body parsing map by a human body representation generator based on a pre-trained VGG network and a regional average pooling layer;
[0093] Calculate the loss function value based on the real human body image and the reconstructed human body image;
[0094] Optionally, the loss function includes:
[0095] Human body image reconstruction loss, human body image perception loss, human body image style loss, human body structure preservation loss based on graph representation, adversarial training loss;
[0096] Human body image reconstruction loss
[0097] Human body image perception loss
[0098] Human body image style loss
[0099] Human body structure preservation loss based on graph representation
[0100] Adversarial training loss L adv = E I,P [log(D s (I)·D p (I,P))]+ E I,P [log((1 - D s (G(I,P)))·(1 - D p (G(I,P),P))))];
[0101] Where, I represents the input human body image, represents the reconstructed human body image, ‖·‖1 represents calculating the Euclidean distance, v l (·) represents the l-th layer of the pre-trained VGG network, G(·) represents the Gram matrix, M(·) represents the human body structure diagram, P represents the human body pose skeleton diagram, S represents the human body pose skeleton diagram, D s represents the style discriminator, D p represents the pose discriminator;
[0102] Calculate the total loss function value L according to the following formula total ;
[0103] L total = β adv L adv + β rec L rec + β perc L perc + β style L style + β graph L graph ;
[0104] where β adv ,β rec ,β perc ,β style ,β graph are the weights of the loss function;
[0105] In this embodiment, the self-supervised human pose conversion ability is realized through a pose encoder, a decoupled style encoder, a correlation mining module, and an image converter, and the effective training of the model is achieved through the constraints of multiple loss functions. Specifically, the pose encoder and the decoupled style encoder enable the decoupled expression of the human body image in the feature space; the correlation mining module calculates the correlation between the non-aligned decoupled features and reorganizes the style features, thereby realizing the fusion of the non-aligned features; the image converter converts the fused features into a real human body image to complete the reconstruction of the human body image.
[0106] Specifically, the self-supervised human pose conversion method includes:
[0107] Step 702: Obtain a certain number of human body images for training;
[0108] Step 704: Obtain the pose skeleton diagram of the human body image based on the pre-trained human pose estimation method;
[0109] Step 706: Obtain the human body parsing diagram of the human body image based on the pre-trained human body parsing method;
[0110] Step 708: Based on the pose extractor, obtain the pose features of the human body according to the pose skeleton diagram of the human body;
[0111] Specifically, in step 708, the pose extractor is composed of a downsampled convolutional neural network, and the network structure is composed of several 3×3 convolutional layers, batch normalization layers, and non-linear activation function layers; the pose extractor extracts features from the entire pose skeleton diagram to obtain a pose feature map with reduced resolution and increased number of channels;
[0112] Step 710: Based on the style feature extractor, obtain the decoupled style features of the human body according to the human body image and its corresponding human body parsing diagram;
[0113] Specifically, in step 710, the style extractor is composed of a downsampled convolutional neural network, and the network structure is composed of several 3×3 convolutional layers, batch normalization layers, and non-linear activation function layers. The style extractor separates the human body parsing diagram into multiple binary mask diagrams according to the human body part labels; and through the Hadamard Product, obtains the decoupled human body part images. The style extractor independently extracts the features of each human body part image and concatenates the features along the channel dimension;
[0114] In step 710, based on the cross-channel fusion module, the human body features of each part are fused in the channel dimension;
[0115] In step 710, the style features of each human body part are independently encoded, and the resulting overall style feature map is misaligned with the pose feature map obtained in step 708;
[0116] Step 712: Based on the spatial correlation mining module, establish a dense spatial correlation field, and reorganize the style features according to the dense spatial correlation field;
[0117] Specifically, in step 712, the spatial correlation mining module calculates the cosine similarity between the feature vectors at each position of the pose feature map and the style feature map, and uses the softmax function for probability normalization. This similarity is saved to the dense spatial correlation field. Based on this dense spatial correlation field, the style features are reorganized in a weighted combination manner, where the weights are read from the dense spatial correlation field;
[0118] Step 714: Based on the image transformer and the reorganized style features, obtain the reconstructed human body image;
[0119] Step 716: Based on the human body graph generator, obtain the human body structure graph;
[0120] Specifically, in step 716, the human body graph generator consists of a pre-trained VGG network and a regional average pooling layer. The pre-trained VGG network extracts the perceptual feature map of the human body image, and the regional average pooling layer separates the perceptual feature map into multiple regions according to the human body parsing graph, and extracts the feature vectors of each region through average pooling operations;
[0121] Furthermore, in step 716, each of the human body region feature vectors serves as a node of the human body structure graph, and the cosine similarity between two feature vectors serves as the edge of the human body structure graph;
[0122] Step 718: Calculate the loss function value based on the input human body image, the reconstructed human body image, and the human body structure graph;
[0123] Optionally, in step 718, the loss function value includes: human body image reconstruction loss, human body image perception loss, human body image style loss, human body structure preservation loss based on graph representation, adversarial training loss;
[0124] In step 718, the adversarial training loss is divided into pose discrimination loss and style discrimination loss, which are calculated by a pose discriminator and a style discriminator respectively;
[0125] Specifically, in step 718, the loss function value is calculated by the following formula:
[0126] Human body image reconstruction loss
[0127] Human body image perception loss
[0128] Human body image style loss
[0129] Human body structure preservation loss based on chart features
[0130] Adversarial training loss L adv = E I,P [log(D s (I)·D p (I,P))]+ E I,P [log((1 - D s (G(I,P)))·(1 - D p (G(I,P),P))))];
[0131] Where I represents the input human body image, represents the reconstructed human body image, ‖·‖1 represents calculating the Euclidean distance, φ l (·) represents the l-th layer of the pre-trained VGG network, G(·) represents the Gram matrix, M(·) represents the human body structure diagram, P represents the human body pose skeleton diagram, S represents the human body pose skeleton diagram, D s represents the style discriminator, D p represents the pose discriminator;
[0132] According to the following formula, calculate the total loss function value L total ;
[0133] L total = β adv L adv + β rec L rec + β perc L perc + β style L style + β graph L graph ;
[0134] Where β adv , β rec , β perc , β style , β graph are the weights of the loss function;
[0135] Step 720: According to the loss function value, use the loss gradient backpropagation algorithm to iteratively adjust the weights of the pose feature encoder, decoupled style encoder, image transformer, and image discriminator until convergence;
[0136] The model designed in this embodiment can be summarized into three parts: a feature decoupling part composed of a pose encoder and a style encoder; a feature fusion part composed of a spatial correlation mining module;
[0137] a feature transformation part composed of an image transformer; the three parts realize the supervision process of model training through the reconstruction of the input human body image;
[0138] The feature decoupling part takes a human body image as input and decouples it into misaligned pose features and style features;
[0139] The feature fusion part calculates the position correlation between feature maps, so that the misaligned style features are embedded into the pose features, realizing the fusion of decoupled features;
[0140] The feature transformation part takes a low-resolution feature map as input and generates a high-resolution real human body image;
[0141] To supervise the model training process, this embodiment uses a human body image reconstruction loss, a human body image perception loss, a human body image style loss, a human body structure preservation loss based on graph representation, a style adversarial training loss, and a pose adversarial training loss to constrain the model training, realizing the training of a self-supervised human pose conversion model.
[0142] It is also possible to perform data augmentation on a single human body image in various forms of spatial transformation, so as to create misaligned image pairs to provide supervision information for model training.
[0143] In summary, the beneficial effects of the embodiments of the present application are as follows:
[0144] 1. It does not rely on paired multi-pose human body datasets for supervised training. Instead, it reconstructs a single human body image by extracting and fusing misaligned features, thereby realizing the supervision process of model training, effectively reducing the costs of data collection and model training.
[0145] In this application, the terms "first", "second", and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance; the term "plural" means two or more, unless otherwise clearly defined. Terms such as "installed", "connected", "joined", "fixed", etc. should be understood in a broad sense. For example, "connected" can be a fixed connection, a detachable connection, or an integral connection; "joined" can be a direct connection or an indirect connection through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0146] In the description of this application, it should be understood that the orientation or positional relationship indicated by terms such as "upper", "lower", "front", "rear", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing this application and simplifying the description, rather than indicating or implying that the device or module referred to must have a specific direction, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to this application.
[0147] In the description of this specification, the description of terms such as "one embodiment", "some embodiments", "specific embodiments", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or instance. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0148] The above are only the preferred embodiments of this application and are not used to limit this application. For those skilled in the art, various changes and modifications can be made to this application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the protection scope of this application.
Claims
1. A self-supervised human pose conversion method, characterized in that Comprising: Obtain an input human body image, and obtain a human body pose skeleton map and a human body parsing map according to the input human body image; Extract features from the human body pose skeleton map through a pose feature encoder to obtain human body pose features, and the pose feature encoder includes a pose feature encoder based on a convolutional neural network; Extract features from the human body parsing map through a decoupled style encoder to obtain human body part style features, and the decoupled style encoder includes a decoupled style encoder based on a convolutional neural network; Fuse the human body part style features through a cross-channel fusion module; According to the human body pose features and the human body part style features, calculate the feature correlation between different positions through a correlation mining module based on spatial correlation learning; Construct a dense spatial correlation field according to the feature correlation; Based on the dense spatial correlation field, perform non-rigid deformation on the human body part style features to obtain recombined style features; Reconstruct the input human body image according to the recombined style features through an image converter to obtain a reconstructed human body image.
2. The self-supervised human pose conversion method according to claim 1, wherein The self-supervised human body pose conversion method further includes: Judge the real human body image and the reconstructed human body image through an image discriminator, and calculate the loss function value.
3. The self-supervised human pose conversion method according to claim 2, wherein The judging the real human body image and the reconstructed human body image through an image discriminator and calculating the loss function value specifically includes: Through a pose discriminator based on a convolutional neural network, judge the real human body image and its pose skeleton map as a positive example pair, judge the reconstructed human body image and its pose skeleton map as a negative example pair, and calculate the loss function value; Through a style discriminator based on a convolutional neural network, judge the real human body image as a positive example, judge the reconstructed human body image as a negative example, and calculate the loss function value.
4. The self-supervised human pose conversion method according to claim 2, wherein The self-supervised human body pose conversion method further includes: Based on a human body graph feature generator of a pre-trained VGG network and a regional average pooling layer, obtain a human body structure graph representing the human body structure according to the input human body image and the human body parsing map; Calculate the loss function value according to the input human body image, the reconstructed human body image and the human body structure graph.
5. The self-supervised human pose conversion method according to claim 4, characterized in that, The loss function includes one or a combination of the following: human body image reconstruction loss, human body image perception loss, human body image style loss, human body structure preservation loss based on graph feature, and adversarial training loss.
6. The self-supervised human pose conversion method according to claim 5, wherein The self-supervised human body pose conversion method further includes: According to the loss function value, iteratively adjust the weights of the pose feature encoder, the decoupled style encoder, the image converter and the image discriminator through the loss gradient backpropagation algorithm until convergence.
7. The self-supervised human pose conversion method according to any one of claims 1 to 6, characterized in that The obtaining the input human body image and obtaining the human body pose skeleton map and the human body parsing map according to the input human body image specifically includes: Obtain an input human body image; Obtain a human body pose skeleton map according to the input human body image through a preset pose estimation method; Obtain a human body parsing map according to the input human body image through a preset human body parsing method.
8. A self-supervised human pose conversion system, characterized in that, Comprising: An acquisition module (110) for acquiring an input human body image and obtaining a human body pose skeleton diagram and a human body parsing diagram according to the input human body image; A pose feature extraction module (120) for extracting features from the human body pose skeleton diagram through a pose feature encoder to obtain human body pose features, where the pose feature encoder includes a pose feature encoder based on a convolutional neural network; A style feature extraction module (130) for extracting features from the human body parsing diagram through a decoupled style encoder to obtain human body part style features, where the decoupled style encoder includes a decoupled style encoder based on a convolutional neural network; A style feature fusion module (140) for fusing the human body part style features through a cross-channel fusion module; A correlation calculation module (150) for calculating the feature correlation between different positions according to the human body pose features and the human body part style features through a correlation mining module based on spatial correlation learning; A correlation field construction module (160) for constructing a dense spatial correlation field according to the feature correlation; A recombination module (170) for performing non-rigid deformation on the human body part style features based on the dense spatial correlation field to obtain recombined style features; A reconstruction module (180) for reconstructing the input human body image through an image converter according to the recombined style features to obtain a reconstructed human body image.
9. A self-supervised human pose conversion system, characterized in that, Comprising: A memory (300) and a processor (400), wherein the memory (300) stores a program or instruction that can run on the processor (400), and when the processor (400) executes the program or the instruction, the steps of the self-supervised human body pose conversion method according to any one of claims 1 to 7 are implemented.
10. A readable storage medium, on which a program or instructions are stored, characterized in that, When the program or the instruction is executed by the processor, the steps of the self-supervised human body pose conversion method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Multi-person posture estimation method based on adversarial learning
CN110598554A
Low-illuminance indoor human body target visible light feature reconstruction method and network based on infrared camera
CN113609893A