Virtual fitting method and device, equipment, storage medium and program product

By extracting and fusing high- and low-level features in human analytical expression in virtual fitting methods, a more accurate human analytical map is generated, and combined with clothing deformation maps, the problem of low authenticity of virtual fitting images is solved, and a more accurate and natural fitting effect image is achieved.

CN120070631APending Publication Date: 2025-05-30CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510125436.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing virtual fitting method based on thin plate spline transformation can easily lead to false obstruction of clothing in body parts during the virtual fitting process, reducing the authenticity of the virtual fitting image.

Method used

By obtaining the human analytical expression, extracting high-level and low-level features, and fusion is performed, the fusion features are obtained to generate more accurate human analytical maps. At the same time, a deformation diagram of the target clothing that meets the target posture is obtained, and analytical diagrams and deformation diagrams are synthesized to generate fitting effect images.

Benefits of technology

It improves the authenticity of virtual fitting images, reduces the probability of wrong occlusion of clothing, and the generated fitting effect image is more accurate and natural.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070631A_ABST
    Figure CN120070631A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a virtual fitting method, device and equipment, a storage medium and a program product, and the method comprises the steps: obtaining a human body analysis expression which at least comprises an analysis graph used for representing human body features of a given figure image and a posture heat graph used for representing a target posture; extracting high-level features and low-level features of the human body analysis expression, and fusing at least part of the high-level features and the low-level features corresponding to the at least part of the features to obtain fused features; based on the fused features, obtaining an analytic graph of the human body under the target posture; obtaining a deformation graph of the target garment conforming to the target posture; and synthesizing the analytic graph and the deformation graph to obtain a fitting effect image of the target garment worn by the human body in the target posture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of artificial intelligence, and particularly relates to a virtual fitting method, device, equipment, storage medium and program product. Background Art

[0002] In related technologies, the virtual fitting method based on thin plate spline transformation is a commonly used virtual fitting method. In the virtual fitting method based on thin plate spline transformation, operations such as human parsing and pose recognition need to be performed. When the accuracy of human parsing is low, it is easy to cause incorrect occlusion of clothing on body parts during the virtual fitting process, reducing the authenticity of the finally generated virtual fitting image. Summary of the Invention

[0003] To solve the problem of low authenticity of virtual fitting images in related technologies, embodiments of this application propose a virtual fitting method, device, equipment, storage medium and program product.

[0004] Embodiments of this application provide a virtual fitting method, and the method includes:

[0005] Obtain a human parsing expression, where the human parsing expression at least includes: a parsing map for representing the human characteristics of a given person image and a pose heat map for representing a target pose;

[0006] Extract the high-level features and low-level features of the human parsing expression, fuse at least some of the features in the high-level features and the corresponding low-level features to obtain the fused features; based on the fused features, obtain a parsing map of the human body in the target pose;

[0007] Obtain a deformation map of a target clothing that conforms to the target pose;

[0008] By synthesizing the parsing map and the deformation map, obtain a fitting effect image of the human body wearing the target clothing in the target pose.

[0009] In some embodiments, the high-level features and the low-level features are represented by feature maps; the fusing at least some of the features in the high-level features and the corresponding low-level features to obtain the fused features includes: determining a first feature in the high-level features whose confidence score is lower than a confidence threshold; determining a target region of the first feature in the corresponding feature map, and determining a second feature in the low-level features corresponding to the target region; fusing the first feature and the second feature to obtain the fused features.

[0010] In some embodiments, obtaining the parsing graph of the human body in the target pose based on the fusion feature includes: obtaining the parsing graph of the human body in the target pose according to the fusion feature and the high-level feature.

[0011] In some embodiments, the human body parsing expression further includes a target clothing mask, and the target clothing mask represents the position correspondence between the target clothing and the human body; the parsing graph of the human body in the target pose can characterize the position correspondence between the target clothing and the human body.

[0012] In some embodiments, obtaining the deformation graph of the target clothing that conforms to the target pose includes: extracting the features of the parsing graph and the features of the image of the target clothing; determining the spatial transformation parameters of the thin plate spline transformation method according to the features of the parsing graph and the features of the image of the target clothing; and processing the spatial transformation parameters by the thin plate spline transformation method under the condition of introducing a preset constraint term to obtain the deformation graph of the target clothing, where the constraint term is used to constrain the deformation range of the target clothing within a set range.

[0013] In some embodiments, obtaining the virtual try-on effect image of the human body wearing the target clothing in the target pose by synthesizing the parsing graph and the deformation graph includes: obtaining a roughly rendered composite image by synthesizing the parsing graph and the deformation graph; and optimizing the facial features of the roughly rendered composite image according to the given person image to obtain the virtual try-on effect image.

[0014] In some embodiments, obtaining a roughly rendered composite image by synthesizing the parsing graph and the deformation graph includes: cropping the clothing part of the given person image to obtain a human body representation omitting clothing information; and synthesizing the parsing graph, the deformation graph and the human body representation omitting clothing information to obtain the roughly rendered composite image.

[0015] The embodiments of the present application further provide a virtual try-on device, and the device includes:

[0016] An obtaining module, configured to obtain a human body parsing expression, where the human body parsing expression at least includes: a parsing graph for representing the human body features of a given person image and a pose heat map for representing a target pose.

[0017] A processing module is configured to extract high-level features and low-level features of the human body parsing expression, fuse at least some of the high-level features and the corresponding low-level features of the at least some of the high-level features to obtain the fused features; based on the fused features, obtain a parsing map of the human body in a target pose; obtain a deformation map of a target garment that conforms to the target pose; and synthesize the parsing map and the deformation map to obtain a virtual try-on effect image of the human body wearing the target garment in the target pose.

[0018] An embodiment of the present application further provides an electronic device, which includes a processor and a memory for storing a computer program that can run on the processor; wherein, the processor is configured to run the computer program to execute any one of the above virtual try-on methods.

[0019] An embodiment of the present application further provides a computer storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements any one of the above virtual try-on methods.

[0020] An embodiment of the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements any one of the above virtual try-on methods.

[0021] It can be seen that the embodiments of the present application can fuse high-level features and low-level features in the human body parsing expression to better capture the global information in the human body parsing expression, so as to more accurately obtain a parsing map of the human body in a target pose, which is beneficial to obtaining a more realistic virtual try-on effect image based on the parsing map. Description of the Drawings

[0022] Figure 1 It is a flowchart of a virtual try-on method according to an embodiment of the present application;

[0023] Figure 2 It is another flowchart of a virtual try-on method according to an embodiment of the present application;

[0024] Figure 3 It is an internal data processing flowchart of an information enhancement module according to an embodiment of the present application;

[0025] Figure 4 It is a structural schematic diagram of a virtual try-on device according to an embodiment of the present application;

[0026] Figure 5 It is a composition structural schematic diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments

[0027] In the related art, two-dimensional virtual try-on methods mainly include a virtual try-on method based on appearance flow and a virtual try-on method based on thin plate spline transformation.

[0028] In the virtual fitting method based on appearance flow, the teacher network can be used to extract the appearance flow between the person image and the clothing image, so as to find the dense correspondence between the two, that is, a set of 2D coordinate vectors. This correspondence is then applied to the student network, which imitates the teacher network to achieve clothing deformation and complete virtual fitting. This method does not require the student network to parse and process the human body image and estimate the posture, but since the teacher network is still trained based on the parser, if the image generated by the teacher network has artifacts, the imitation result of the student network may not be ideal.

[0029] In the virtual fitting method based on thin plate spline transformation, the mesh control points can be interpolated to deform the mesh to obtain a thin plate spline transformation model, and the clothing can be transformed into an image that conforms to the human body posture through the trained thin plate spline transformation model, thereby realizing the virtual fitting function. This method does not require a teacher network, but requires interpolation calculation of the mesh, which is computationally intensive.

[0030] The virtual fitting method based on thin plate spline transformation is a relatively flexible method, but it also has some limitations. First, this method requires operations such as human body analysis and posture recognition, and the accuracy of these operations will directly affect the subsequent fitting effect. Secondly, excessive deviation is prone to occur during the training process, resulting in the deformed clothing shape not being close to the human body posture. When the posture amplitude of the human body's front and back transformation is large, clothing deformation will also produce undesirable effects. In addition, the accuracy of the early human body analysis will also affect the subsequent human body fitting process. If the analysis results are wrong, they will be superimposed layer by layer during the fitting process, resulting in incorrect occlusion of body parts during fitting, such as clothing incorrectly covering the position of the human hand.

[0031] In view of the technical problems existing in the related technologies, a technical solution of the embodiment of the present application is proposed. The embodiment of the present application proposes an overall solution for multi-stage virtual fitting under posture guidance.

[0032] The following is a further detailed description of the embodiments of the present application in conjunction with the accompanying drawings and examples. It should be understood that the embodiments provided herein are only used to explain the embodiments of the present application and are not intended to limit the embodiments of the present application. In addition, the embodiments provided below are partial embodiments for implementing the present application, rather than providing all embodiments for implementing the present application. In the absence of conflict, the technical solutions recorded in the embodiments of the present application can be implemented in any combination.

[0033] It should be noted that in the embodiments of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a method or device including a series of elements not only includes the elements clearly recited, but also includes other elements not explicitly listed, or further includes elements inherent in the implementation of the method or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of other related elements in the method or device including the element (such as steps in a method or units in a device, and the unit can be part of a circuit, part of a processor, part of a program or software, etc.).

[0034] The virtual fitting method provided by the embodiments of the present application includes a series of steps, but the virtual fitting method provided by the embodiments of the present application is not limited to the recited steps. Similarly, the virtual fitting device provided by the embodiments of the present application includes a series of modules, but the device provided by the embodiments of the present application is not limited to including the explicitly recited modules, and may further include modules required for obtaining relevant information or processing based on the information.

[0035] Figure 1 is a flowchart of the virtual fitting method according to the embodiments of the present application, as Figure 1 shown, the process includes:

[0036] Step 101: Obtain a human body parsing expression, where the human body parsing expression at least includes: a parsing diagram for representing the human body characteristics of a given person image and a pose heat map for representing a target pose.

[0037] The given person image can be a pre-acquired image. For example, the given person image can be obtained through the shooting operation of a shooting device, or the given person image pre-stored in an electronic device can be read out, or the given person image can be obtained through a network. The embodiments of the present application do not limit the source of the given person image.

[0038] In the embodiments of the present application, the human body characteristics of the given person image may include the characteristics of various parts of the human body. For example, the human body characteristics may include the characteristics of the hair, face, neck, and body of the human body. In one example, a three-channel parsing diagram can be used to represent the human body characteristic part of the given person image.

[0039] The pose heat map can be obtained according to the conversion of the key points of the human body. Exemplarily, the OpenPose network or other key point extraction models can be used to extract 18 human body key points from the human body image conforming to the target pose, and then the 18 human body key points can be converted into an 18-channel pose heat map.

[0040] OpenPose adopts a multi-stage structure, and each stage is divided into two branches. First, each branch extracts features from the image to be analyzed to obtain the basic feature representation of the image; then, the extracted features are further processed and merged to perform multi-scale feature fusion to obtain the output result of this stage. Finally, OpenPose will visually present the output result in the form of a two-dimensional image so that users can more clearly understand the posture of human actions.

[0041] Step 102: Extract the high-level features and low-level features of the human body parsing expression, fuse at least some of the features in the high-level features and the low-level features corresponding to at least some of the features to obtain the fused features; based on the fused features, obtain the parsing map of the human body in the target pose.

[0042] In the embodiments of the present application, a Unet network or other networks can be used to extract the high-level features and low-level features of the human body parsing expression. The Unet network is mainly composed of a convolutional layer, a max pooling layer, and a transposed convolutional layer. The input image of the Unet network first passes through four modules composed of a convolutional layer and a max pooling layer to perform four downsampling operations; then passes through four modules composed of a transposed convolutional layer and a convolutional layer to perform four upsampling operations.

[0043] The convolutional layer formula, max pooling layer formula, and transposed convolutional layer formula used in the Unet network are formula (1), formula (2), and formula (3), respectively

[0044]

[0045]

[0046] N = S * (X - 1) + F - 2 * S (3)

[0047] Among them, N is the size of the output image of each layer, X is the size of the input image of each layer, F is the size of the convolutional kernel used in the convolutional layer, P is the padding value size, and S is the stride size.

[0048] In some embodiments, the human body parsing expression further includes a target clothing mask, and the target clothing mask represents the position correspondence between the target clothing and the human body; the parsing map of the human body in the target pose can characterize the position correspondence between the target clothing and the human body. Here, the target clothing refers to the clothing for which virtual fitting is to be performed. It can be seen that since the parsing map can characterize the position correspondence between the target clothing and the human body, it is beneficial to obtain a more realistic fitting effect image according to the parsing map.

[0049] Exemplarily, the target clothing mask is a single-channel binary mask. In this binary mask, the part of the human body wearing the target clothing is marked as 1, and the part of the human body not wearing the target clothing (i.e., the part of the human body not covered by the target clothing) is marked as 0.

[0050] In some embodiments, referring to Figure 2 , after obtaining the human body parsing expression based on a given person image, etc., a parsing conversion step can be performed on the human body parsing expression to obtain the parsing map of the human body in the target pose. The input of the parsing conversion step can be a 22-channel human body parsing expression, that is, the human body parsing expression includes a three-channel parsing map, an 18-channel pose heat map, and a single-channel binary mask.

[0051] Step 103: Obtain the deformed map of the target clothing that conforms to the target pose.

[0052] In the embodiments of the present application, referring to Figure 2 , the human body clothing expression can be obtained based on a given person image, etc. The human body clothing expression includes the human body feature representation and the image of the target clothing. The human body feature representation includes a three-channel parsing map for representing the features of the human body head retention area (including hair and face), and also includes the above-mentioned pose heat map and the above-mentioned target clothing mask.

[0053] Referring to Figure 2 , a clothing deformation step can be performed on the human body clothing expression to obtain the deformed map of the target clothing that conforms to the target pose. Clothing deformation is the process of extracting the human body clothing expression and the target clothing features and generating the target clothing that conforms to the target pose. Exemplarily, a Spatial Transformer Network (STN) and a thin plate spline transformation method can be used to complete clothing deformation.

[0054] Exemplarily, the human body feature representation and the target clothing image are first respectively passed through two feature extraction layers to extract high-level features. Then, these two features are combined into a tensor using matrix multiplication. This tensor is used as the input of the regression layer to predict the spatial transformation parameter theta of the thin plate spline transformation method. Next, by inputting the spatial transformation parameter theta into the thin plate spline transformation network, the deformed clothing is generated. Finally, the deformed clothing should match the target pose.

[0055] Step 104: By synthesizing the parsing map and the deformed map, obtain the fitting effect image of the human body wearing the target clothing in the target pose.

[0056] In practical applications, steps 101 to 104 can be implemented based on a processor, and the above-mentioned processor can be at least one of an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Digital Signal Processing Device (DSPD), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a Central Processing Unit (CPU), a controller, a microcontroller, and a microprocessor.

[0057] It can be seen that the embodiments of the present application can fuse high-level features and low-level features in the human body parsing expression, better capture the global information in the human body parsing expression, so as to more accurately obtain the parsing diagram of the human body in the target pose, which is beneficial to obtaining a more realistic virtual fitting effect image based on the parsing diagram.

[0058] In some embodiments of the present application, the above-mentioned high-level features and low-level features can be represented by feature maps. Correspondingly, the process of fusing at least some features in the high-level features and the low-level features corresponding to at least some features to obtain the fused features may include: determining a first feature in the high-level features whose confidence score is lower than the confidence threshold; determining the target region of the first feature in the corresponding feature map, and determining a second feature in the low-level features corresponding to the target region; fusing the first feature and the second feature to obtain the fused features.

[0059] The network used in the parsing conversion step can be the Unet network described above. However, the Unet network uses skip connections to directly splice high-level features and low-level features, but it ignores the complementarity between these two types of features. Using the Unet network to directly splice features will cause the low-level features to not well distinguish the uncertain regions in the high-level features. Therefore, in the embodiments of the present application, an information enhancement module can be introduced, and this module can fuse high-level and low-level features to improve the accuracy and robustness of human body parsing.

[0060] The input of the information enhancement module is the human body parsing expression. The information enhancement module is used to introduce low-level feature information into the high-level feature information of the Unet network and can suppress the influence of introduced noise. In the information enhancement module, the high-level feature map restored by upsampling can be measured to determine its confidence score. Exemplarily, the information enhancement module can use the Softmax function to process the upsampled high-level feature map, calculate the probability of each pixel belonging to each category, and use the maximum value among all the probabilities of each pixel belonging to each category as its confidence score. The confidence score of the high-level feature can be represented by a confidence feature map. According to the confidence feature map, some features (first features) in the high-level feature and the low-level feature can be added element by element of the matrix to obtain a fused feature.

[0061] The Softmax function can be represented by formula (4).

[0062]

[0063] Where x i represents the channel value of each pixel point of the image input to the formula.

[0064] It can be seen that in the embodiments of the present application, by fusing the features with relatively low confidence scores in the high-level features and the low-level features, more information is obtained from the low-level features, which is beneficial to obtaining a more accurate parsing map based on the fused feature.

[0065] In some embodiments of the present application, the process of obtaining the parsing map of the human body in the target pose based on the fused feature includes: obtaining the parsing map of the human body in the target pose according to the fused feature and the high-level feature. Here, the fused feature and the high-level feature can be fused to obtain the parsing map of the human body in the target pose.

[0066] It can be seen that the embodiments of the present application can more accurately obtain the parsing map of the human body in the target pose on the basis of comprehensively considering the fused feature and the high-level feature.

[0067] Referring to Figure 3 , the information enhancement module can extract the low-level features in the input image and obtain the high-level features by processing the input image through an encoder-decoder; then, obtain the confidence feature map according to the confidence score in the high-level feature, complete the fusion of some features in the high-level feature and the low-level feature according to the confidence feature map to obtain the fused feature; finally, fuse the fused feature and the high-level feature to obtain the output result. The output result of the information enhancement module is the above-mentioned parsing map of the human body in the target pose.

[0068] Exemplarily, the relationship between the input image and the output result of the information enhancement module can be represented by formula (5).

[0069]

[0070] Among them, Out represents the output result of the information enhancement module, X represents the input image, and Y represents the output image of X passing through the Unet network. Represents the element-wise addition operation of matrices.

[0071] In the embodiments of the present application, the high-level features and low-level features are fused through the information enhancement module, so that the lost global information can be captured, redundant information can be avoided from being introduced, thereby improving the accuracy of the parsing graph and reducing the probability of incorrect occlusion in the fitting effect image.

[0072] In some embodiments of the present application, the process of obtaining the deformed image of the target clothing that conforms to the target pose includes: extracting the features of the parsing graph and the features of the image of the target clothing; determining the spatial transformation parameters of the thin plate spline transformation method according to the features of the parsing graph and the features of the image of the target clothing; and processing the spatial transformation parameters through the thin plate spline transformation method under the introduction of a preset constraint term to obtain the deformed image of the target clothing, where the constraint term is used to constrain the deformation range of the target clothing within a set range.

[0073] Here, the set range can be set according to actual needs. The thin plate spline transformation method realizes the deformation of the grid by performing interpolation calculations on the grid control points. However, due to the high flexibility of the thin plate spline transformation, excessive deformation may cause distortion of the clothing pattern, especially when the pose conversion amplitude is large. To solve this problem, in the embodiments of the present application, by introducing a preset constraint term, the parameters of the thin plate spline transformation can be restricted, and to a certain extent, the excessive deformation difference of the grid can be avoided. The constraint term is defined on the grid of the thin plate spline transformation to ensure that the deformation of the clothing pattern remains within a certain range.

[0074] To solve the problem of excessive clothing deformation, the embodiments of the present application introduce a preset constraint term in the clothing deformation step. The constraint term acts on the deformed grid obtained through the thin plate spline transformation to avoid excessive deformation differences between the front and back two grids, thereby reducing the irregular deformation of the clothing and better retaining the clothing features.

[0075] In some embodiments of the present application, by synthesizing the parsing graph and the deformed image, the fitting effect image of the human body wearing the target clothing in the target pose is obtained, including: synthesizing the parsing graph and the deformed image to obtain a roughly rendered composite image; and optimizing the facial features of the roughly rendered composite image according to the given person image to obtain the fitting effect image.

[0076] It can be seen that in the embodiments of the present application, after obtaining the coarsely rendered composite image, the facial features of the coarsely rendered composite image can also be optimized based on the given person image, thereby improving the authenticity of the final obtained fitting effect image.

[0077] In some embodiments of the present application, by synthesizing the parsing map and the deformed map, a coarsely rendered composite image is obtained, including: cropping the clothing part of the given person image to obtain a human body representation excluding clothing information; synthesizing the parsing map, the deformed map, and the human body representation excluding clothing information to obtain the coarsely rendered composite image.

[0078] Here, by cropping the clothing part of the given person image, an image of the human body without clothing coverage can be obtained, and the image of the human body without clothing coverage is the above-mentioned human body representation excluding clothing information.

[0079] Referring to Fig. 2, after obtaining the parsing map of the human body in the target pose, the deformed map of the target clothing conforming to the target pose, and the human body representation excluding clothing information, a fitting synthesis step can be performed on the parsing map of the human body in the target pose, the deformed map of the target clothing conforming to the target pose, and the human body representation excluding clothing information to obtain a coarsely rendered composite image, and the coarsely rendered composite image can include a coarsely rendered fitting image and a synthesis mask.

[0080] In some embodiments, a Unet network can be used to perform the fitting synthesis step; since the information enhancement module can focus the high-level feature map on the features of the low-confidence region and achieve the information enhancement effect on the low-confidence region, therefore, the information enhancement module can also be introduced in the fitting synthesis step. By processing the parsing map, the deformed map, and the human body representation excluding clothing information through the Unet network with the information enhancement module introduced, a coarsely rendered composite image can be generated.

[0081] Referring to Figure 2 , after obtaining the coarsely rendered composite image, a facial refinement step can be performed on the coarsely rendered composite image and the given person image to obtain a fitting effect image. Exemplarily, by processing the coarsely rendered composite image and the given person image through the Unet network with the information enhancement module introduced, a fitting effect image can be obtained.

[0082] In the embodiments of the present application, in order to solve the problem of insufficient human body parsing, an information enhancement module can be introduced in the parsing conversion step. This module can better capture the lost global information, help the network extract features more comprehensively, and retain the semantic information of the person, providing better guidance for the feature fusion of the subsequent network. In addition, since the information enhancement module can better fuse high-level features and low-level features, it can also play the role of information enhancement in the try-on synthesis step and the face refinement step, helping to improve the accuracy of the network output result and reduce the wrong occlusion of the try-on result.

[0083] The technical solution of the embodiments of the present application has at least the following advantages compared with the related technologies: 1) Compared with the virtual fitting method based on the appearance flow, the fitting effect image generated in the embodiments of the present application is more realistic and natural, and the face features are also more complete. 2) Compared with the virtual fitting method based on the thin plate spline transformation, the embodiments of the present application reduce the irregular deformation of the clothing and better retain the clothing features; at the same time, reduce the occlusion during the fitting process, and the synthesized fitting result is also more accurate and realistic.

[0084] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order that constitutes any limitation to the implementation process, and the specific execution order of each step should be determined according to its function and possible internal logic.

[0085] Figure 4 It is a schematic structural diagram of the virtual fitting device for the embodiments of the present application, as Figure 4 shown, the device includes:

[0086] An acquisition module 401, configured to acquire a human body parsing expression, where the human body parsing expression at least includes: a parsing map for representing the human body features of a given person image and a pose heat map for representing a target pose;

[0087] A processing module 402, configured to extract high-level features and low-level features of the human body parsing expression, fuse at least some of the high-level features and the corresponding low-level features to obtain the fused features; based on the fused features, obtain a parsing map of the human body in the target pose; obtain a deformation map of the target clothing that conforms to the target pose; and synthesize the parsing map and the deformation map to obtain a fitting effect image of the human body wearing the target clothing in the target pose.

[0088] In some embodiments, the high-level features and the low-level features are represented by feature maps;

[0089] The processing module 402, configured to fuse at least some of the high-level features and the corresponding low-level features to obtain the fused features, includes:

[0090] Determine a first feature in the high-level features with a confidence score lower than a confidence threshold;

[0091] Determine a target region of the first feature in the corresponding feature map, and determine a second feature in the low-level features corresponding to the target region;

[0092] Fuse the first feature and the second feature to obtain the fused feature.

[0093] In some embodiments, the processing module 402 is configured to obtain an analysis map of the human body in the target pose based on the fused feature, including:

[0094] Obtain the analysis map of the human body in the target pose according to the fused feature and the high-level features.

[0095] In some embodiments, the human body analysis expression further includes a target clothing mask, and the target clothing mask represents the positional correspondence between the target clothing and the human body; the analysis map of the human body in the target pose can characterize the positional correspondence between the target clothing and the human body.

[0096] In some embodiments, the processing module 402 is configured to obtain a deformed map of the target clothing that conforms to the target pose, including:

[0097] Extract the features of the analysis map and the features of the image of the target clothing;

[0098] Determine the spatial transformation parameters of the thin plate spline transformation method according to the features of the analysis map and the features of the image of the target clothing;

[0099] Process the spatial transformation parameters by the thin plate spline transformation method with a preset constraint term introduced to obtain the deformed map of the target clothing, where the constraint term is used to constrain the deformation range of the target clothing within a set range.

[0100] In some embodiments, the processing module 402 is configured to obtain a try-on effect image of the human body wearing the target clothing in the target pose by synthesizing the analysis map and the deformed map, including:

[0101] Synthesize the analysis map and the deformed map to obtain a roughly rendered synthesized image;

[0102] Optimize the facial features of the roughly rendered synthesized image according to the given person image to obtain the try-on effect image.

[0103] In some embodiments, the processing module 402 is configured to obtain a roughly rendered composite image by synthesizing the parsed graph and the deformed graph, including:

[0104] Cropping the clothing part of the given person image to obtain a human body representation excluding clothing information;

[0105] Synthesizing the parsed graph, the deformed graph, and the human body representation excluding clothing information to obtain the roughly rendered composite image.

[0106] In practical applications, the acquisition module 401 and the processing module 402 can be implemented based on a processor and a communication device.

[0107] It should be noted that the description of the above device embodiments is similar to that of the above method embodiments and has similar beneficial effects to the method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0108] It should be noted that in the embodiments of the present application, if the above method is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application essentially or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a terminal, a server, etc.) to execute all or part of the methods described in the embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs that can store program codes. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0109] Correspondingly, the embodiments of the present application further provide a computer program product, where the computer program product includes computer-executable instructions for implementing any virtual fitting method provided by the embodiments of the present application.

[0110] Correspondingly, the embodiments of the present application further provide a computer storage medium, where computer-executable instructions are stored on the computer storage medium for implementing any virtual fitting method provided by the above embodiments.

[0111] The embodiments of the present application further provide an electronic device. Figure 5 Shown in the following is a schematic structural diagram of an electronic device provided by the embodiments of the present application. As Figure 5 shown, the electronic device 50 may include:

[0112] A memory 501 for storing executable instructions;

[0113] A processor 502, when executing the executable instructions stored in the memory 501, implements any of the above virtual fitting methods.

[0114] The above-mentioned processor 502 can be at least one of ASIC, DSP, DSPD, PLD, FPGA, CPU, controller, microcontroller, and microprocessor.

[0115] The above-mentioned computer-readable storage medium and memory 502 can be a read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; it can also be various terminals including one or any combination of the above memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0116] In some embodiments, the functions or modules included in the device provided by the embodiments of the present application can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0117] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. Their similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated in this article.

[0118] The methods disclosed in the method embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments.

[0119] The features disclosed in the product embodiments provided by the present application can be arbitrarily combined without conflict to obtain new product embodiments.

[0120] The features disclosed in each method or device embodiment provided by the present application can be arbitrarily combined without conflict to obtain a new method embodiment or device embodiment.

[0121] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0122] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims. All of these are within the protection scope of the present application.

Claims

1. A virtual fitting method, characterized in that: The method comprises: Acquire a human body analytical expression, wherein the human body analytical expression comprises at least: a analytical graph for representing human body features of a given person image and a posture heat map for representing a target posture; Extracting high-level features and low-level features of the analytical expression of the human body, fusing at least part of the high-level features and the low-level features corresponding to the at least part of the features to obtain the fused features; and obtaining an analytical graph of the human body in a target posture based on the fused features; Obtaining a deformation map of a target garment that matches the target posture; By synthesizing the analytical graph and the deformation graph, a fitting effect image of a human body wearing the target garment in the target posture is obtained.

2. The method according to claim 1, characterized in that The high-level features and the low-level features are represented by feature maps; The fusing at least part of the high-level features and the low-level features corresponding to the at least part of the features to obtain the fused features includes: Determining a first feature among the high-level features having a confidence score below a confidence threshold; Determine a target area of ​​the first feature in a corresponding feature map, and determine a second feature corresponding to the target area in the low-level features; The first feature and the second feature are fused to obtain the fused feature.

3. The method according to claim 2, characterized in that The step of obtaining a parsing diagram of a human body in a target posture based on the fusion features includes: According to the fused features and the high-level features, an analytical graph of the human body in the target posture is obtained.

4. The method according to claim 1, characterized in that The human body analysis table The method further includes a target clothing mask, which represents the position correspondence between the target clothing and the human body; the analytical graph of the human body in the target posture can represent the position correspondence between the target clothing and the human body.

5. The method according to claim 1, characterized in that The step of obtaining a deformation map of the target clothing that matches the target posture includes: Extracting features of the analytical graph and features of the image of the target clothing; Determining spatial transformation parameters of a thin plate spline transformation method according to the characteristics of the analytical graph and the characteristics of the image of the target garment; In the case of introducing preset constraint items, the spatial transformation parameters are processed by a thin plate spline transformation method to obtain a deformation map of the target garment, wherein the constraint items are used to constrain the deformation range of the target garment within a set range.

6. The method according to any one of claims 1 to 5, characterized in that: The step of synthesizing the analytical graph and the deformation graph to obtain a fitting effect image of a human body wearing the target garment in the target posture includes: By synthesizing the analytical image and the deformation image, a roughly rendered synthetic image is obtained; The facial features of the roughly rendered composite image are optimized according to the given character image to obtain the fitting effect image.

7. The method according to claim 6, characterized in that The step of synthesizing the analytical graph and the deformation graph to obtain a roughly rendered synthetic image includes: cropping the clothing part of the given person image to obtain a human body representation omitting clothing information; The analytical graph, the deformation graph and the human body representation without clothing information are synthesized to obtain the roughly rendered synthetic image.

8. A virtual fitting device, characterized in that: The device comprises: An acquisition module, used for acquiring a human body analytical expression, wherein the human body analytical expression comprises at least: a analytical graph for representing human body features of a given person image and a posture heat map for representing a target posture; A processing module is used to extract high-level features and low-level features of the analytical expression of the human body, fuse at least part of the high-level features and the low-level features corresponding to the at least part of the features to obtain the fused features; based on the fused features, obtain a analytical graph of the human body in a target posture; obtain a deformation graph of a target garment that conforms to the target posture; and obtain a fitting effect image of a human body wearing the target garment in the target posture by synthesizing the analytical graph and the deformation graph.

9. An electronic device, characterized in that: The electronic device comprises a processor and a memory for storing a computer program that can be run on the processor; wherein, The processor is configured to run the computer program to perform the method according to any one of claims 1 to 7.

10. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

11. A computer program product, comprising a computer program, characterized in that The computer program implements the method according to any one of claims 1 to 7 when executed by a processor.