Virtual fitting model training method, virtual fitting method, device and equipment based on real person images

Through the virtual fitting model training method based on real person images, using two-dimensional images for virtual fitting, the problems of high cost and poor effect in the existing technology are solved, and a more realistic and natural virtual fitting effect is achieved.

CN114926324BActive Publication Date: 2025-06-03SHENZHEN HAOLI SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210593210.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-27
Publication Date
2025-06-03
Estimated Expiration
2042-05-27

AI Technical Summary

Technical Problem

The existing virtual fitting technology is costly and the fitting effect is unreal and unnatural. In particular, 3D virtual fitting requires the construction of realistic 3D models, while 2D virtual fitting is limited by hardware, resulting in poor results.

Method used

A virtual fitting model training method based on real person images is adopted. By obtaining two-dimensional original clothing images and real person images, posture estimation, feature extraction and appearance flow estimation are performed, and a virtual fitting image is generated by combining the generation of adversarial neural networks, and a virtual fitting model is generated through iterative training.

Benefits of technology

It reduces the development and maintenance costs of virtual fittings, improves the authenticity and nature of fitting effects, and avoids the need to generate and use three-dimensional models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114926324B_ABST
    Figure CN114926324B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a method, a virtual fitting method, a device and a device for training a virtual fitting model based on real human images. The present invention relates to the technical field of virtual fitting. The method includes: training a virtual fitting model based on a two-dimensional original clothing image and a two-dimensional real human image, wherein the virtual fitting model includes a feature pyramid network, an appearance flow estimation network and a generative adversarial neural network. After the model is trained, a two-dimensional real human image and a two-dimensional real clothing image selected by a user are obtained; performing pose estimation and image semantic segmentation on the two-dimensional real human image respectively to obtain a human pose estimation and a human clothing area; inputting the human pose estimation, the human clothing area and the two-dimensional real clothing image into the trained virtual fitting model, and generating a virtual fitting image through feature extraction, appearance flow estimation and synthesis processing in sequence. The embodiment of the present invention can not only reduce the fitting cost, but also has a real and natural fitting effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of virtual fitting, and particularly to a method for training a virtual fitting model based on real human images, a virtual fitting method, a device and equipment. Background Art

[0002] With the development of computer technology and online shopping platforms, virtual fitting has received wide attention. Existing virtual fitting mainly falls into two modes: 2D and 3D. Among them, for 3D virtual fitting, in order to achieve the fitting effect, it is generally necessary to construct a realistic 3D model. However, a clothing model needs to be constructed for each piece of clothing, and at the same time, different customer appearances also require an adaptive human body 3D model, resulting in high development and maintenance costs. For 2D virtual fitting, it relies on a 3D-like user image, and on this basis, clothing texture mapping is performed. The method of first generating a three-dimensional model and then forming a two-dimensional image not only requires too much offline processing but also is limited by hardware, often making the fitting effect unrealistic and unnatural. Summary of the Invention

[0003] Embodiments of the present invention provide a method for training a virtual fitting model based on real human images, a virtual fitting method, a device and equipment, aiming to solve the problems of high cost and unrealistic and unnatural fitting effect in existing virtual fitting.

[0004] In a first aspect, embodiments of the present invention provide a method for training a virtual fitting model based on real human images, which includes:

[0005] Obtain original training data, where the original training data includes two-dimensional original clothing images and two-dimensional real human images wearing corresponding clothes;

[0006] Perform pose estimation and marking on the two-dimensional real human image to obtain human pose estimation and human clothing regions;

[0007] Input the human pose estimation, the human clothing regions, and the two-dimensional original clothing images into a preset feature extraction network for feature extraction and fusion to obtain human pose features, human semantic features, and clothing features;

[0008] Input the two-dimensional original clothing images, the human pose features, the human semantic features, and the clothing features into an appearance flow estimation network to obtain an appearance flow estimation map;

[0009] Input the appearance flow estimation map, the human pose estimation, and the human clothing regions into a generative adversarial neural network to obtain a virtual fitting image;

[0010] Iteratively train the preset feature extraction network, the appearance flow estimation network, and the generative adversarial neural network according to the virtual try-on image and the two-dimensional real person image to generate a virtual try-on model.

[0011] In a second aspect, an embodiment of the present invention provides a virtual try-on method, which includes:

[0012] Obtain a two-dimensional real person image and a two-dimensional real clothing image selected by a user;

[0013] Perform pose estimation and image semantic segmentation on the two-dimensional real person image respectively to obtain a human pose estimation and a human clothing area;

[0014] Input the human pose estimation, the human clothing area, and the two-dimensional real clothing image into the virtual try-on model described in the first aspect above, and generate a virtual try-on image through feature extraction, appearance flow estimation, and synthesis processing in sequence.

[0015] In a third aspect, an embodiment of the present invention further provides a virtual try-on model training device based on a real person image, which includes:

[0016] A first acquisition unit, configured to acquire original training data, where the original training data includes a two-dimensional original clothing image and a two-dimensional real person image wearing the corresponding clothing;

[0017] An estimation and marking unit, configured to perform pose estimation and marking on the two-dimensional real person image to obtain a human pose estimation and a human clothing area;

[0018] A feature extraction and fusion unit, configured to input the human pose estimation, the human clothing area, and the two-dimensional original clothing image into a preset feature extraction network for feature extraction and fusion to obtain a human pose feature, a human semantic feature, and a clothing feature;

[0019] An appearance flow estimation unit, configured to input the two-dimensional original clothing image, the human pose feature, the human semantic feature, and the clothing feature into an appearance flow estimation network to obtain an appearance flow estimation map;

[0020] A first generation unit, configured to input the appearance flow estimation map, the human pose estimation, and the human clothing area into a generative adversarial neural network to obtain a virtual try-on image;

[0021] A training unit, configured to iteratively train the preset feature extraction network, the appearance flow estimation network, and the generative adversarial neural network according to the virtual try-on image and the two-dimensional real person image to generate a virtual try-on model.

[0022] Fourthly, an embodiment of the present invention further provides a virtual fitting device, which includes:

[0023] A second acquisition unit, configured to acquire a two-dimensional real human image and a two-dimensional real clothing image selected by a user;

[0024] An estimation and segmentation unit, configured to perform pose estimation and image semantic segmentation on the two-dimensional real human image respectively to obtain a human pose estimation and a human clothing area;

[0025] A second generation unit, configured to input the human pose estimation, the human clothing area, and the two-dimensional real clothing image into the virtual fitting model according to any one of claims 1-5, and generate a virtual fitting image after feature extraction, appearance flow estimation, and synthesis processing in sequence.

[0026] Fifthly, an embodiment of the present invention further provides a computer device, which includes a memory and a processor. A computer program is stored on the memory, and when the processor executes the computer program, the methods of the first aspect and the second aspect are implemented.

[0027] An embodiment of the present invention provides a method for training a virtual fitting model based on a real human image, a virtual fitting method, a device, and a device. The method includes: first, training a virtual fitting model based on a two-dimensional original clothing image and a two-dimensional real human image. The virtual fitting model includes a feature pyramid network, an appearance flow estimation network, and a generative adversarial neural network. After the model is trained, a two-dimensional real human image and a two-dimensional real clothing image selected by a user are acquired; pose estimation and image semantic segmentation are performed on the two-dimensional real human image respectively to obtain a human pose estimation and a human clothing area; the human pose estimation, the human clothing area, and the two-dimensional real clothing image are input into the trained virtual fitting model, and a virtual fitting image is generated after feature extraction, appearance flow estimation, and synthesis processing in sequence. The technical solution of the embodiment of the present invention trains a virtual fitting model based on a two-dimensional original clothing image and a two-dimensional real human image during the model training stage. The virtual fitting model includes a feature pyramid network, an appearance flow estimation network, and a generative adversarial neural network. During the model application stage, based on the acquired two-dimensional real human image and two-dimensional real clothing image, a virtual fitting image is generated through the trained virtual fitting model. The entire virtual fitting process only generates two-dimensional images and does not require generating and using three-dimensional models, which can not only reduce the fitting cost, but also make the fitting effect real and natural. Description of the Drawings

[0028] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0029] Figure 1 It is a schematic flowchart of a virtual fitting model training method based on real person images provided by an embodiment of the present invention;

[0030] Figure 2 It is a diagram showing the effect of a virtual fitting model training method based on real person images provided by an embodiment of the present invention;

[0031] Figure 3 It is a schematic flowchart of a virtual fitting method provided by an embodiment of the present invention;

[0032] Figure 4 It is a schematic block diagram of a virtual fitting model training device based on real person images provided by an embodiment of the present invention;

[0033] Figure 5 It is a schematic block diagram of a virtual fitting device provided by an embodiment of the present invention;

[0034] Figure 6 It is a schematic block diagram of a computer device provided by an embodiment of the present invention. Detailed implementation manners

[0035] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0036] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprise" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0037] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0038] It should also be further understood that the term "and / or" used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0039] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "once" or "in response to determining" or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]" depending on the context.

[0040] Please refer to Figure 1 , Figure 1 FIG. is a schematic flow chart of a virtual fitting model training method based on real person images provided by an embodiment of the present invention. The virtual fitting model training method based on real person images according to the embodiment of the present invention can be applied to a server, for example, the virtual fitting model training method based on real person images can be implemented by a software program configured on the server. The virtual fitting model training method based on real person images will be described in detail below. As Figure 1 shown, the method includes the following steps S100-S150.

[0041] S100. Obtain original training data, where the original training data includes two-dimensional original clothing images and two-dimensional real person images wearing corresponding clothes.

[0042] In the embodiment of the present invention, when training a virtual fitting model, it is first necessary to obtain an original training data set required for training the virtual fitting model. The original training data includes two-dimensional original clothing images and two-dimensional real person images wearing corresponding clothes. Understandably, the two-dimensional original clothing images include various clothing images. For example, long-sleeved images, short-sleeved images, and sleeveless images, and the long-sleeved images, short-sleeved images, and sleeveless images respectively include long-length images, short-length images, and medium-length images, etc.; the two-dimensional real person images include images of models wearing various clothes.

[0043] S110. Perform pose estimation and marking on the two-dimensional real person images to obtain human pose estimation and human clothing areas.

[0044] In the embodiment of the present invention, after obtaining several two-dimensional real human images, pose estimation of the two-dimensional real human images is performed through a pose detection model to obtain human pose estimation, where the pose detection model is an existing pose detection model, for example, the Pr-VIPE pose detection model or the SOTA pose detection model; the hair, face, and lower body clothing regions in the two-dimensional real human images are marked to obtain the human clothing region, that is, the clothing region is marked. It should be noted that in the embodiment of the present invention, the human pose estimation includes data of human key points such as the nose, left eye, right eye, left ear, right ear, left shoulder, and right shoulder.

[0045] S120. Input the human pose estimation, the human clothing region, and the two-dimensional original clothing image into a preset feature extraction network for feature extraction and fusion to obtain human pose features, human semantic features, and clothing features.

[0046] In the embodiment of the present invention, the human pose estimation, the human clothing region, and the two-dimensional original clothing image are input into a preset feature extraction network, a feature pyramid network, for feature extraction to obtain multiple pose features, multiple human features, and multiple clothing features. Among them, the preset feature extraction network is a feature pyramid network, and the feature pyramid network is a five-layer recursive pyramid network, and operations are performed through convolutions with a stride of 2 between adjacent network layers; the multiple pose features, the multiple human features, and the multiple clothing features are respectively fused to obtain human pose features, human semantic features, and clothing features. It should be noted that in the embodiment of the present invention, the feature pyramid network further includes 2 residual modules. It should also be noted that in the embodiment of the present invention, the reason for using the feature pyramid network to extract features is that the feature pyramid network extracts features more accurately.

[0047] S130. Input the two-dimensional original clothing image, the human pose features, the human semantic features, and the clothing features into an appearance flow estimation network to obtain an appearance flow estimation map.

[0048] In the embodiment of the present invention, the appearance flow estimation network first estimates the appearance flow based on the human body pose feature, the human body semantic feature, and the clothing feature; then deforms the two-dimensional original clothing image according to the appearance flow to obtain an appearance flow estimation map. Understandably, in the process of estimating the two-dimensional original clothing image, in order to avoid the appearance of jump discontinuities between adjacent pixel points in the two-dimensional original clothing image, that is, in order to better retain the features of the clothing, a smoothing constraint is introduced to limit the collinearity of the appearance estimation. It should be noted that in the embodiment of the present invention, a second-order smoothing constraint is used to limit the collinearity of the appearance estimation. Understandably, in other embodiments, other-order smoothing constraints can also be used to limit the collinearity of the appearance estimation, for example, a first-order smoothing constraint. It should also be noted that in the embodiment of the present invention, the appearance flow estimation network includes 5 flow network modules, and each flow network module includes 4 convolutions.

[0049] S140. Input the appearance flow estimation map, the human body pose estimation, and the human body clothing area into a generative adversarial neural network to obtain a virtual fitting image.

[0050] In the embodiment of the present invention, the appearance flow estimation map, the human body pose estimation, and the human body clothing area are input into a generative adversarial neural network to obtain a virtual fitting image, where the generative adversarial neural network is pix2pix, which includes a generator and a discriminator, and the virtual fitting image is generated through the generator and the discriminator. Understandably, in other embodiments, the generative adversarial neural network can also be other adversarial networks, such as CGAN.

[0051] S150. Iteratively train the preset feature extraction network, the appearance flow estimation network, and the generative adversarial neural network according to the virtual fitting image and the two-dimensional real person image to generate a virtual fitting model.

[0052] In an embodiment of the present invention, the constructed virtual fitting model includes a Feature Pyramid Network (FPN), an Appearance Flow Estimation Network, and a Generative Adversarial Network (GAN). After generating the virtual fitting image, a loss value is calculated according to the virtual fitting image and the two-dimensional real person image through a preset loss function. If the loss value is not less than the preset loss value and the number of training times is less than the preset number of training times, the model is iteratively trained. Understandably, if the loss value is less than the preset loss value or the number of training times is not less than the preset number of training times, the trained model is used as the virtual fitting model. It should be noted that, in an embodiment of the present invention, the loss value is calculated according to the loss value in the smoothing process, the visual similarity between the fitting image and the reference image, and the logarithm of the element-wise difference, that is, the loss function is composed of a smoothed second-order loss function, a visual similarity loss function, and an element-wise loss function. Among them, the element-wise loss function may include functions such as activation functions, absolute values, and square roots.

[0053] It should also be noted that, during the model training process, as Figure 2 shown, Figure 2 in the first figure from left to right is the two-dimensional real person image, the third figure is the two-dimensional original clothing image, the first figure and the third figure constitute the original training data. The first figure is processed to obtain the second figure and the fourth figure. The second figure is the human body dressing area, and the fourth figure is the figure synthesized by human body pose estimation and the human body dressing area. The fifth figure is the appearance flow estimation figure, that is, the figure generated after passing through the Feature Pyramid Network, the Appearance Flow Estimation Network, and the Generative Adversarial Network.

[0054] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of a virtual fitting method provided by an embodiment of the present invention. The virtual fitting method of the embodiment of the present invention can be applied to a terminal. For example, the virtual fitting method can be implemented through a software program configured on the terminal, such as a software named virtual fitting software, which can not only reduce the fitting cost, but also make the fitting effect real and natural. The virtual fitting method will be described in detail below. As Figure 3 shown, the method includes the following steps S200 - S220.

[0055] S200. Obtain a two-dimensional real person image and a two-dimensional real clothing image selected by the user;

[0056] S210. Perform pose estimation and image semantic segmentation on the two-dimensional real person image respectively to obtain human body pose estimation and the human body dressing area;

[0057] S220. Input the human body pose estimation, the human body dressing area, and the two-dimensional real clothing image into the virtual fitting model, and generate a virtual fitting image through feature extraction, appearance flow estimation, and synthesis processing in sequence.

[0058] In an embodiment of the present invention, before virtual fitting, the user first needs to select a two-dimensional real human image and a two-dimensional real clothing image. In practical applications, the virtual fitting software presents a clothing selection interface including at least one clothing picture to the user for the user to select their favorite clothing as the two-dimensional real clothing image. After the user selects the two-dimensional real clothing image, that is, after the virtual fitting software receives the trigger instruction for the user to select the two-dimensional real clothing image, it will present photographing guidance information to the user so that the user can take a photo according to the photographing guidance information. After the photographing is completed, the virtual fitting software will present a photo selection interface including at least one user picture to the user for the user to select a suitable user picture as the two-dimensional real human image. After the two-dimensional real human image and the two-dimensional real clothing image are confirmed, the virtual fitting software will obtain the two-dimensional real human image and the two-dimensional real clothing image, and perform pose estimation on the two-dimensional real human image through a pose estimation model to obtain a human pose estimation, wherein the pose estimation model is an existing pose estimation model, such as the Openpose pose estimation model. After obtaining the human pose estimation, perform image semantic segmentation on the two-dimensional real human image through a human body image segmentation model to obtain a human body clothing area, wherein the human body image segmentation model is an existing human body image segmentation model. For example, the Bodypix human body image segmentation model. After obtaining the human body pose data and the human body clothing area, input the human body pose estimation, the human body clothing area, and the two-dimensional real clothing image into the virtual fitting model, and generate a virtual fitting image through feature extraction, appearance flow estimation, and synthesis processing in sequence. Understandably, after generating the virtual fitting image, display the virtual fitting image for the user to view.

[0059] Figure 4 FIG. is a schematic block diagram of a virtual fitting model training device 200 based on real human images provided by an embodiment of the present invention. As Figure 4 shown, corresponding to the above virtual fitting model training method based on real human images, the present invention also provides a virtual fitting model training device 200 based on real human images. The virtual fitting model training device 200 based on real human images includes units for executing the above virtual fitting model training method based on real human images, and this device can be configured in a server. Specifically, please refer to Figure 4 , the virtual fitting model training device 200 based on real human images includes a first acquisition unit 201, an estimation and marking unit 202, a feature extraction and fusion unit 203, an appearance flow estimation unit 204, a first generation unit 205, and a training unit 206.

[0060] Among them, the first acquisition unit 201 is used to acquire original training data, and the original training data includes a two-dimensional original clothing image and a two-dimensional real human image wearing the corresponding clothing; the estimation and marking unit 202 is used to perform pose estimation and marking on the two-dimensional real human image to obtain a human pose estimation and a human clothing area; the feature extraction and fusion unit 203 is used to input the human pose estimation, the human clothing area, and the two-dimensional original clothing image into a preset feature extraction network for feature extraction and fusion to obtain a human pose feature, a human semantic feature, and a clothing feature; the appearance flow estimation unit 204 is used to input the two-dimensional original clothing image, the human pose feature, the human semantic feature, and the clothing feature into an appearance flow estimation network to obtain an appearance flow estimation map; the first generation unit 205 is used to input the appearance flow estimation map, the human pose estimation, and the human clothing area into a generative adversarial neural network to obtain a virtual fitting image; the training unit 206 is used to iteratively train the preset feature extraction network, the appearance flow estimation network, and the generative adversarial neural network according to the virtual fitting image and the two-dimensional real human image to generate a virtual fitting model.

[0061] In some embodiments, such as this embodiment, the estimation and marking unit includes a pose estimation unit and a marking unit.

[0062] Among them, the pose estimation unit is used to perform pose estimation on the two-dimensional real human image through a pose detection model to obtain a human pose estimation; the marking unit is used to mark the hair, face, and lower body clothing area in the two-dimensional real human image to obtain a human clothing area.

[0063] In some embodiments, such as this embodiment, the feature extraction and fusion unit 203 includes an extraction unit and a fusion unit.

[0064] Among them, the extraction unit is used to input the human pose estimation, the human clothing area, and the two-dimensional original clothing image into a feature pyramid network for feature extraction to obtain a plurality of pose features, a plurality of human features, and a plurality of clothing features; the fusion unit is used to fuse the plurality of pose features, the plurality of human features, and the plurality of clothing features respectively to obtain a human pose feature, a human semantic feature, and a clothing feature.

[0065] In some embodiments, such as this embodiment, the training unit 206 includes a calculation unit, a training subunit, and a generation subunit.

[0066] Wherein, the calculation unit is configured to calculate a loss value according to the virtual fitting image and the two-dimensional real person image through a preset loss function; the training subunit is configured to iteratively train the preset feature extraction network, the appearance flow estimation network, and the generative adversarial neural network if the loss value is not less than a preset loss value and the number of training times is less than a preset number of training times; the generation subunit is configured to use the trained preset feature extraction network, the appearance flow estimation network, and the generative adversarial neural network as a virtual fitting model if the loss value is less than the preset loss value or the number of training times is not less than the preset number of training times.

[0067] Figure 5 FIG. 4 is a schematic block diagram of a virtual fitting device 300 provided by an embodiment of the present invention. As Figure 5 shown, corresponding to the above virtual fitting method, the present invention also provides a virtual fitting device 300. The virtual fitting device 300 includes units for performing the above virtual fitting method, and the device can be configured in a terminal. Specifically, please refer to Figure 5 FIG. 4, the virtual fitting device 300 includes a second acquisition unit 301, an estimation and segmentation unit 302, and a second generation unit 303.

[0068] Wherein, the second acquisition unit 301 is configured to acquire a two-dimensional real person image and a two-dimensional real clothing image selected by a user; the estimation and segmentation unit 302 is configured to perform pose estimation and image semantic segmentation on the two-dimensional real person image respectively to obtain a human pose estimation and a human clothing area; the second generation unit 303 is configured to input the human pose estimation, the human clothing area, and the two-dimensional real clothing image into a virtual fitting model, and generate a virtual fitting image through feature extraction, appearance flow estimation, and synthesis processing in sequence.

[0069] In some embodiments, such as this embodiment, the virtual fitting device 300 further includes a first display unit, a second display unit, a third display unit, and an acting as unit.

[0070] Wherein, the first display unit is configured to display a clothing selection interface including at least one clothing picture to the user; the second display unit is configured to display photographing guiding information to the user after receiving a trigger instruction for selecting a two-dimensional real clothing image, so that the user can take a photo according to the photographing guiding information; the third display unit is configured to display a photo selection interface including at least one user picture to the user; the acting as unit is configured to receive the user picture selected by the user and use the user picture as a two-dimensional real person image.

[0071] The above virtual fitting model training and virtual fitting device can be implemented in the form of a computer program, and the computer program can be stored in a storage medium such as Figure 6running on the computer device shown.

[0072] Please refer to Figure 6 , Figure 6 which is a schematic block diagram of a computer device provided by an embodiment of the present invention. The computer device 300 is a server or a terminal. Specifically, the server can be an independent server or a server cluster composed of multiple servers.

[0073] Refer to Figure 6 , the computer device 300 includes a processor 302, a memory, and a network interface 305 connected through a system bus 301. Among them, the memory can include a storage medium 303 and an internal memory 304.

[0074] The storage medium 303 can store an operating system 3031 and a computer program 3032. When the computer program 3032 is executed, it can cause the processor 302 to execute a virtual fitting model training method based on real person images. The action pose model trained based on the virtual fitting model training method based on real person images can also cause the processor 302 to execute a virtual fitting method.

[0075] The processor 302 is used to provide computing and control capabilities to support the operation of the entire computer device 300.

[0076] The internal memory 304 provides an environment for the operation of the computer program 3032 in the storage medium 303. When the computer program 3032 is executed by the processor 302, it can cause the processor 302 to execute a virtual fitting model training method based on real person images. The action pose model trained based on the virtual fitting model training method based on real person images can also cause the processor 302 to execute a virtual fitting method.

[0077] The network interface 305 is used for network communication with other devices. Those skilled in the art can understand that Figure 6 the structure shown in

[0078] Among them, the processor 302 is used to run the computer program 3032 stored in the memory to implement the following steps: obtaining original training data, where the original training data includes a two-dimensional original clothing image and a two-dimensional real person image wearing the corresponding clothing; performing pose estimation and marking on the two-dimensional real person image to obtain a human pose estimation and a human clothing area; inputting the human pose estimation, the human clothing area, and the two-dimensional original clothing image into a preset feature extraction network for feature extraction and fusion to obtain a human pose feature, a human semantic feature, and a clothing feature; inputting the two-dimensional original clothing image, the human pose feature, the human semantic feature, and the clothing feature into an appearance flow estimation network to obtain an appearance flow estimation map; inputting the appearance flow estimation map, the human pose estimation, and the human clothing area into a generative adversarial neural network to obtain a virtual try-on image; and iteratively training the preset feature extraction network, the appearance flow estimation network, and the generative adversarial neural network according to the virtual try-on image and the two-dimensional real person image to generate a virtual try-on model.

[0079] In some embodiments, such as this embodiment, when the processor 302 implements the step of performing pose estimation and marking on the two-dimensional real person image to obtain a human pose estimation and a human clothing area, the specific implementation steps are as follows: performing pose estimation on the two-dimensional real person image through a pose detection model to obtain a human pose estimation; marking the hair, face, and lower body clothing area in the two-dimensional real person image to obtain a human clothing area.

[0080] In some embodiments, such as this embodiment, when the processor 302 implements the step of inputting the human pose estimation, the human clothing area, and the two-dimensional original clothing image into a preset feature extraction network for feature extraction and fusion to obtain a human pose feature, a human semantic feature, and a clothing feature, the specific implementation steps are as follows: inputting the human pose estimation, the human clothing area, and the two-dimensional original clothing image into a feature pyramid network for feature extraction to obtain a plurality of pose features, a plurality of human features, and a plurality of clothing features; respectively fusing the plurality of pose features, the plurality of human features, and the plurality of clothing features to obtain a human pose feature, a human semantic feature, and a clothing feature.

[0081] In some embodiments, such as this embodiment, when the processor 302 implements the step of inputting the two-dimensional original clothing image, the human pose feature, the human semantic feature, and the clothing feature into an appearance flow estimation network to obtain an appearance flow estimation map, the specific implementation steps are as follows: inputting the human pose feature, the human semantic feature, and the clothing feature into an appearance flow estimation network to obtain an appearance flow; deforming the two-dimensional original clothing image according to the appearance flow to obtain an appearance flow estimation map.

[0082] In some embodiments, such as this embodiment, when the processor 302 implements the step of iteratively training the preset feature extraction network, the appearance flow estimation network, and the generative adversarial neural network according to the virtual fitting image and the two-dimensional real person image to generate a virtual fitting model, the specific implementation is as follows: calculating a loss value according to the virtual fitting image and the two-dimensional real person image through a preset loss function; if the loss value is not less than a preset loss value and the number of training times is less than a preset number of training times, then iteratively training the preset feature extraction network, the appearance flow estimation network, and the generative adversarial neural network; if the loss value is less than the preset loss value or the number of training times is not less than the preset number of training times, then using the trained preset feature extraction network, appearance flow estimation network, and generative adversarial neural network as a virtual fitting model.

[0083] Among them, the processor 302 is used to run the computer program 3032 stored in the memory to implement the following steps: obtaining a two-dimensional real person image and a two-dimensional real clothing image selected by the user; respectively performing pose estimation and image semantic segmentation on the two-dimensional real person image to obtain a human pose estimation and a human clothing area; inputting the human pose estimation, the human clothing area, and the two-dimensional real clothing image into the virtual fitting model, and generating a virtual fitting image through feature extraction, appearance flow estimation, and synthesis processing in sequence.

[0084] In some embodiments, such as this embodiment, before the processor 302 implements the step of obtaining the two-dimensional real person image and the two-dimensional real clothing image selected by the user, the specific implementation further includes the following steps: displaying a clothing selection interface including at least one clothing picture to the user; after receiving a trigger instruction for the user to select a two-dimensional real clothing image, displaying photographing guidance information to the user so that the user can take a photo according to the photographing guidance information; displaying a photo selection interface including at least one of the user's pictures to the user; receiving the user's selected user picture and using the user picture as a two-dimensional real person image.

[0085] It should be understood that in the embodiments of the present invention, the processor 302 may be a central processing unit (CPU), and the processor 302 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0086] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a storage medium, and the storage medium is a computer-readable storage medium. The computer program is executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.

[0087] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program. When the computer program is executed by a processor, the processor executes any of the embodiments of the above virtual fitting model training method and virtual fitting method based on real person images.

[0088] The storage medium may be various computer-readable storage media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes.

[0089] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0090] In several embodiments provided by the present invention, it should be understood that the disclosed apparatus and method can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0091] The steps in the method embodiments of the present invention can be adjusted, combined, and deleted according to actual needs. The units in the apparatus embodiments of the present invention can be combined, divided, and deleted according to actual needs. In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0092] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0093] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0094] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, provided that these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.

[0095] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or substitutions within the technical scope disclosed by the present invention, and these modifications or substitutions should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for training a virtual fitting model based on real - person images, characterized in that, it includes: Obtain original training data, where the original training data includes two - dimensional original clothing images and two - dimensional real - person images wearing corresponding clothes; Perform pose estimation and marking on the two - dimensional real - person images to obtain human pose estimation and human clothing regions; Input the human pose estimation, the human clothing region, and the two - dimensional original clothing image into a preset feature extraction network for feature extraction and fusion to obtain human pose features, human semantic features, and clothing features; Input the two - dimensional original clothing image, the human pose features, the human semantic features, and the clothing features into an appearance flow estimation network to obtain an appearance flow estimation map; Input the appearance flow estimation map, the human pose estimation, and the human clothing region into a generative adversarial neural network to obtain virtual fitting images; Iteratively train the preset feature extraction network, the appearance flow estimation network, and the generative adversarial neural network according to the virtual fitting images and the two - dimensional real - person images to generate a virtual fitting model; The step of iteratively training the preset feature extraction network, the appearance flow estimation network, and the generative adversarial neural network according to the virtual fitting images and the two - dimensional real - person images to generate a virtual fitting model includes: Calculate a loss value through a preset loss function according to the virtual fitting images and the two - dimensional real - person images; If the loss value is not less than a preset loss value and the number of training times is less than a preset number of training times, then iteratively train the preset feature extraction network, the appearance flow estimation network, and the generative adversarial neural network; If the loss value is less than the preset loss value or the number of training times is not less than the preset number of training times, then use the trained preset feature extraction network, appearance flow estimation network, and generative adversarial neural network as the virtual fitting model.

2. The method for training a virtual fitting model based on real - person images according to claim 1, characterized in that, The step of performing pose estimation and marking on the two - dimensional real - person images to obtain human pose estimation and human clothing regions includes: Perform pose estimation on the two - dimensional real - person images through a pose detection model to obtain human pose estimation; Mark the hair, face, and lower - body clothing regions in the two - dimensional real - person images to obtain human clothing regions.

3. The method for training a virtual fitting model based on real - person images according to claim 1, characterized in that, The preset feature extraction network is a feature pyramid network. The step of inputting the human pose estimation, the human clothing region, and the two - dimensional original clothing image into the preset feature extraction network for feature extraction and fusion to obtain human pose features, human semantic features, and clothing features includes: Input the human pose estimation, the human clothing region, and the two - dimensional original clothing image into the feature pyramid network for feature extraction to obtain multiple pose features, multiple human features, and multiple clothing features; Fuse the multiple pose features, the multiple human body features, and the multiple clothing features respectively to obtain human body pose features, human body semantic features, and clothing features.

4. The method according to claim 1, wherein, inputting the two-dimensional original clothing image, the human body pose features, the human body semantic features, and the clothing features into an appearance flow estimation network to obtain an appearance flow estimation map includes: inputting the human body pose features, the human body semantic features, and the clothing features into an appearance flow estimation network to obtain an appearance flow; deforming the two-dimensional original clothing image according to the appearance flow to obtain an appearance flow estimation map.

5. A virtual fitting method, wherein, it includes: obtaining a two-dimensional real human image and a two-dimensional real clothing image selected by a user; performing pose estimation and image semantic segmentation on the two-dimensional real human image respectively to obtain a human body pose estimation and a human body dressing area; inputting the human body pose estimation, the human body dressing area, and the two-dimensional real clothing image into the virtual fitting model according to any one of claims 1-4, and generating a virtual fitting image through feature extraction, appearance flow estimation, and synthesis processing in sequence.

6. The method according to claim 5, wherein, before obtaining the two-dimensional real human image and the two-dimensional real clothing image selected by the user, it further includes: displaying a clothing selection interface including at least one clothing picture to the user; after receiving a trigger instruction for the user to select a two-dimensional real clothing image, displaying photographing guidance information to the user so that the user can take a photo according to the photographing guidance information; displaying a photo selection interface including at least one user picture to the user; receiving the user picture selected by the user and using the user picture as the two-dimensional real human image.

7. A virtual fitting model training device based on real human images, wherein, it includes: a first acquisition unit for acquiring original training data, where the original training data includes a two-dimensional original clothing image and a two-dimensional real human image wearing the corresponding clothing; an estimation and marking unit for performing pose estimation and marking on the two-dimensional real human image to obtain a human body pose estimation and a human body dressing area; a feature extraction and fusion unit for inputting the human body pose estimation, the human body dressing area, and the two-dimensional original clothing image into a preset feature extraction network for feature extraction and fusion to obtain human body pose features, human body semantic features, and clothing features; an appearance flow estimation unit for inputting the two-dimensional original clothing image, the human body pose features, the human body semantic features, and the clothing features into an appearance flow estimation network to obtain an appearance flow estimation map; a first generation unit for inputting the appearance flow estimation map, the human body pose estimation, and the human body dressing area into a generative adversarial neural network to obtain a virtual fitting image; a training unit for iteratively training the preset feature extraction network, the appearance flow estimation network, and the generative adversarial neural network according to the virtual fitting image and the two-dimensional real human image to generate a virtual fitting model; the training unit includes: A computing unit, configured to calculate a loss value according to the virtual try-on image and the two-dimensional real person image through a preset loss function; A training subunit, configured to, if the loss value is not less than a preset loss value and the number of training times is less than a preset number of training times, perform iterative training on the preset feature extraction network, the appearance flow estimation network, and the generative adversarial neural network; A generation subunit, configured to, if the loss value is less than the preset loss value or the number of training times is not less than the preset number of training times, use the trained preset feature extraction network, the appearance flow estimation network, and the generative adversarial neural network as a virtual try-on model.

8. A virtual try-on device, characterized in that, it includes: A second acquisition unit, configured to acquire a two-dimensional real person image and a two-dimensional real clothing image selected by a user; An estimation and segmentation unit, configured to perform pose estimation and image semantic segmentation on the two-dimensional real person image respectively to obtain a human pose estimation and a human body clothing area; A second generation unit, configured to input the human pose estimation, the human body clothing area, and the two-dimensional real clothing image into the virtual try-on model according to any one of claims 1-4, and generate a virtual try-on image through feature extraction, appearance flow estimation, and synthesis processing in sequence.

9. A computer device, characterized in that, the computer device includes a memory and a processor, a computer program is stored on the memory, and when the processor executes the computer program, the method for training a virtual try-on model based on a real person image according to any one of claims 1-4 is implemented.

Citation Information

Patent Citations

  • Novel virtual fitting network

    CN112613439A

  • Semantic-based multi-posture virtual fitting method

    CN113361560A