Virtual fitting method, device, server and storage medium
By scaling and deforming the high-resolution human body and clothing images, combined with knowledge distillation and adversarial training of the appearance flow deformation module and the generation module, the problem of large resource occupation and insufficient fusion effect of the PF-AFN algorithm is solved, and efficient fusion effect is achieved.
Patent Information
- Application Number
- CN202111404834.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-24
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2041-11-24
AI Technical Summary
The existing PF-AFN algorithm takes up a lot of resources when processing high-resolution input images, and the clothing fusion effect is insufficient.
By obtaining high-resolution target human body and clothing images, scaling processing is performed to generate fusion images, and using the appearance flow deformation module and the generation module for image deformation and fusion, including knowledge distillation and adversarial training to improve the fusion effect of clothes.
High-resolution target dressing images are generated, which improves the fusion effect of clothes and reduces resource usage.
Smart Images

Figure CN114170403B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of image processing technology, and in particular to a virtual fitting method, device, server, and storage medium. Background Art
[0002] Virtual fitting refers to the use of virtual technology to automatically display a three-dimensional image of what a new garment will look like after trying it on. This technology allows customers to put on new clothes and see how they will look without having to remove their clothes.
[0003] Currently, the Parser-Free Appearance Flow Network (PF-AFN) algorithm is a highly effective virtual fitting algorithm. This algorithm applies a distillation algorithm to the appearance flow. A parser-based 2D fitting algorithm generates fitting images, which serve as a teacher network. The fitting images are then used as a subnetwork to learn the appearance flow generation of the teacher network.
[0004] In the process of implementing the embodiments of the present application, the inventors found that the current technical solution has at least the following technical problems: the current PF-AFN algorithm consumes a lot of resources for high-resolution input images, and the fusion effect of clothes is insufficient. Summary of the Invention
[0005] The main technical problem solved by the embodiments of the present application is to provide a virtual fitting method, device, server and storage medium to improve the fusion effect of clothes.
[0006] In a first aspect, an embodiment of the present application provides a virtual fitting method, comprising:
[0007] Acquire a first target human body image and a first target clothing image, wherein the first target human body image and the first target clothing image are both high-resolution images;
[0008] Scaling the first target human body image and the first target clothing image to generate a second target human body image and a second target clothing image;
[0009] generating a first target fused image according to the second target human body image and the second target clothing image;
[0010] generating a first target processed image according to the first target human body image, wherein the first target processed image does not include clothing information and arm information;
[0011] The first target fusion image, the first target clothing image and the first target processed image are input into the first clothing model to generate a target clothing image.
[0012] In some embodiments, the method further includes: pre-training the first clothing model, specifically including:
[0013] Acquire a first data set, wherein the first data set includes a first human body image and a first clothing image, wherein the first human body image and the first clothing image are both high-resolution images;
[0014] Scaling the first human body image and the first clothing image in the first data set to generate a second data set, wherein the second data set includes the scaled first human body image and the scaled first clothing image;
[0015] generating a third data set according to the scaled first human body image and the scaled first clothing image in the second data set, wherein the third data set includes the first fused image;
[0016] generating a first processed image according to the first human body image, wherein the first processed image does not include clothing information and arm information;
[0017] Scaling the first fused image to generate a second fused image, wherein the second fused image has the same resolution as the first human body image;
[0018] A first human body image, a first processed image, and a second fused image are input to train a first clothing model.
[0019] In some embodiments, generating a first target fused image based on the second target human body image and the second target clothing image includes:
[0020] The second target human body image and the second target clothing image are input into the appearance flow deformation module and the generation module to generate the first target fusion image.
[0021] In some embodiments, inputting the second target human body image and the second target clothing image into the appearance flow deformation module and the generation module to generate the first target fused image includes:
[0022] Inputting the second target human body image and the second target clothing image into the appearance flow deformation module to generate second deformation information;
[0023] deforming the second target clothing image using the second deformation information to generate a third target clothing image;
[0024] The second target human body image and the third target clothing image are input into the generation module to generate a first target fusion image.
[0025] In some embodiments, the appearance flow deformation module includes: a first appearance flow deformation module and a second appearance flow deformation module, wherein the second appearance flow deformation module is used to generate the second deformation information, and the method further includes:
[0026] Training the second appearance flow deformation module includes:
[0027] Acquire a first human body image and a first clothing image;
[0028] Acquiring human body image analysis information according to the first human body image;
[0029] Processing the human body image analysis information and the first clothing image through a first appearance flow deformation module to generate first deformation information;
[0030] Acquire a second clothing image, and deform the second clothing image according to the first deformation information to generate a third clothing image;
[0031] Fusing the third clothing image and the human body image analysis information to generate a first fitting image;
[0032] A second appearance flow deformation module is trained based on the first clothing image and the first try-on image.
[0033] In some embodiments, the method further comprises:
[0034] During the training process of the second appearance manifold deformation module, knowledge distillation is performed on the second appearance manifold deformation module through the first appearance manifold deformation module.
[0035] In some embodiments, the first appearance flow deformation module includes a parser-based appearance flow deformation module, and the second appearance flow deformation module includes a parser-free appearance flow deformation module.
[0036] In a second aspect, an embodiment of the present application provides a virtual fitting device, comprising:
[0037] an acquiring unit, configured to acquire a first target human body image and a first target clothing image, wherein the first target human body image and the first target clothing image are both high-resolution images;
[0038] a scaling unit, configured to scale the first target human body image and the first target clothing image to generate a second target human body image and a second target clothing image;
[0039] a generating unit, configured to generate a first target fused image based on the second target human body image and the second target clothing image; and generate a first target processed image based on the first target human body image, wherein the first target processed image does not include clothing information and arm information;
[0040] The fusion unit is used to input the first target fusion image, the first target clothing image and the first target processed image into the first clothing model to generate a target clothing image.
[0041] In a third aspect, an embodiment of the present application provides a server, including:
[0042] A memory and one or more processors, the one or more processors are used to execute one or more computer programs stored in the memory, and when the one or more processors execute the one or more computer programs, the electronic device implements the method of the first aspect.
[0043] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the method of the first aspect.
[0044] Beneficial effects of the embodiments of the present application: Different from the prior art, the embodiments of the present application provide a virtual fitting method, including: acquiring a first target human body image and a first target clothing image, wherein the first target human body image and the first target clothing image are both high-resolution images; scaling the first target human body image and the first target clothing image to generate a second target human body image and a second target clothing image; generating a first target fused image based on the second target human body image and the second target clothing image; generating a first target processed image based on the first target human body image, wherein the first target processed image does not include clothing information and arm information; inputting the first target fused image, the first target clothing image and the first target processed image into a first clothing model to generate a target dressing image.
[0045] By acquiring a high-resolution first target human body image and a second target clothing image and scaling them, a second target human body image and a second target clothing image are obtained to generate a first target fused image, and a first target processed image excluding clothing information and arm information is generated based on the first target human body image, and then the first target fused image, the first target clothing image and the first target processed image are input into the first clothing model to generate a target clothing image. The embodiment of the present application can generate a high-resolution target clothing image and improve the fusion effect of clothing. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0047] Figure 1 This is a schematic diagram of an application environment of a virtual fitting method provided in an embodiment of the present application;
[0048] Figure 2 This is a flowchart of a virtual fitting method provided in Example 1 of the present application;
[0049] Figure 3 This is a schematic diagram of the principle of a virtual fitting method provided in Example 1 of the present application;
[0050] Figure 4 is a schematic diagram of an appearance flow deformation module provided in Example 1 of the present application;
[0051] Figure 5 This is a structural diagram of a virtual fitting device provided in Example 1 of the present application;
[0052] Figure 6 This is a flowchart of a virtual fitting method provided in Example 2 of the present application;
[0053] Figure 7 This is a schematic diagram of the principle of a virtual fitting method provided in Example 2 of the present application;
[0054] Figure 8 This is a structural diagram of a virtual fitting device provided in Example 2 of the present application;
[0055] Figure 9 This is a flowchart of a virtual fitting method provided in Example 3 of the present application;
[0056] Figure 10 This is a schematic diagram of the principle of generating a first target fused image provided in Example 3 of the present application;
[0057] Figure 11 Schematic diagram of the principle of training a second appearance flow deformation module provided in the third embodiment of the present application;
[0058] Figure 12 This is a schematic diagram of the principle of training a first clothing model provided in Example 3 of the present application;
[0059] Figure 13 This is a structural diagram of a virtual fitting device provided in Example 3 of the present application;
[0060] Figure 14 A schematic diagram of the hardware structure of a server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0061] The present application is described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present application, but are not intended to limit the present application in any form. It should be noted that those skilled in the art may make several variations and improvements without departing from the scope of the present application. These all fall within the scope of protection of the present application.
[0062] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0063] It should be noted that, if there is no conflict, the various features in the embodiments of the present application can be combined with each other and are all within the scope of protection of the present application. In addition, although the functional modules are divided in the device schematic and the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in a different order than the module division in the device or the order in the flow chart. In addition, the words "first", "second", "third", etc. used herein do not limit the data and execution order, but only distinguish between the same items or similar items with basically the same functions and effects.
[0064] Unless otherwise defined, all technical and scientific terms used in this specification have the same meanings as those commonly understood by those skilled in the art to which this application belongs. The terms used in this specification and in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The term "and / or" as used in this specification includes any and all combinations of one or more of the relevant listed items.
[0065] In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as they do not conflict with each other.
[0066] Before explaining the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations:
[0067] (1) Generative Adversarial Networks (GANs) use adversarial training to make the samples generated by the generative network conform to the real data distribution. In a generative adversarial network, two networks are trained in adversarial training. One is the discriminant network, whose goal is to judge as accurately as possible whether a sample comes from real data or is generated by the generative network; the other is the generative network, whose goal is to generate samples whose sources cannot be distinguished by the discriminant network. These two networks with opposite goals are continuously trained alternately. When they finally converge, if the discriminant network can no longer determine the source of a sample, then it is equivalent to the generative network being able to generate samples that conform to the real data distribution.
[0068] (2) Appearance flow refers to which pixels in the source can be used to synthesize the two-dimensional coordinate vector of the target.
[0069] (3) Knowledge Distillation (KD) refers to a model compression method and a training method based on the "teacher-student network concept". It extracts the knowledge contained in the trained model by distilling it into another model. By introducing soft-targets related to the teacher network (complex but with excellent reasoning performance) as part of the total loss, it can induce the training of the student network (simplified and low-complexity) to achieve knowledge transfer. That is, first train another more complex teacher network (generally an integration of multiple networks) and use the output of the large network as a soft target to train the student network.
[0070] The technical solution of this application is described in detail below with reference to the accompanying drawings.
[0071] See also Figure 1 , Figure 1 This is a schematic diagram of an application environment of a virtual fitting method provided in an embodiment of the present application;
[0072] like Figure 1 As shown, the application environment 100 includes: a terminal 101 and a server 102, and the terminal 101 and the server 102 communicate through wired or wireless communication.
[0073] The terminal 101 may be, but is not limited to, a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc. The terminal 101 may be provided with a client, which may be a video client, a browser client, an online shopping client, an instant messaging client, etc. This application does not limit the type of the client.
[0074] The terminal 101 and the server 102 can be connected directly or indirectly via wired or wireless communication, which is not limited in this application. The terminal 101 can receive clothing images sent by the server 102 and display the clothing images on a visual interface. The terminal 101 can also set a corresponding try-on button at each clothing image to provide a try-on function. The user can browse the clothing images and trigger a try-on instruction for the clothing image by triggering the try-on button corresponding to any clothing image. The terminal can respond to the try-on instruction and obtain a human body image through an image acquisition device. The image acquisition device can be built into the terminal 101 or externally connected to the terminal 101, which is not limited in this application.
[0075] The terminal 101 can send the fitting instruction and the collected human body image to the server 102, and receive the target clothing image returned by the server 102, and then display the target clothing image on the visual interface so that the user can understand the effect of the clothes on the body.
[0076] It is understood that terminal 101 may generally refer to one of multiple terminals. The embodiments of this application only use terminal 101 as an example. Those skilled in the art will appreciate that the number of terminals may be greater or lesser. For example, there may be only one terminal, or there may be dozens, hundreds, or even more terminals. The embodiments of this application do not limit the number of terminals or device types.
[0077] Among them, server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), as well as big data and artificial intelligence platforms.
[0078] The server 102 and the terminal 101 may be connected directly or indirectly via wired or wireless communication, which is not limited in this application. The server 102 may maintain a clothing image database for storing multiple clothing images. The server 102 may receive a try-on instruction and a human body image sent by the terminal 101, and based on the try-on instruction, obtain the clothing image corresponding to the try-on instruction from the clothing image database, and generate a target clothing image based on the clothing image and the human body image, and then send the target clothing image to the terminal 101.
[0079] It is understandable that the number of the above servers 102 can be more or less, and the embodiment of the present application does not limit this. Of course, the server 102 can also include other functional servers to provide more comprehensive and diversified services.
[0080] Example 1
[0081] See also Figure 2 , Figure 2 This is a flowchart of a virtual fitting method provided in Example 1 of the present application;
[0082] The virtual fitting method is applied to a server. Specifically, the execution subject of the virtual fitting method is one or more processors of the server.
[0083] like Figure 2 As shown, the virtual fitting method includes:
[0084] Step S201: Acquire a first human body image and a first clothing image;
[0085] Specifically, the first human image is an image of a person wearing clothing, i.e., the first human image includes both person information and clothing information. It is understood that the first human image also includes a background portion. The person in the first human image can have various poses or postures, such as with hands on hips or with hands hanging naturally, and this embodiment of the application does not limit this.
[0086] See also Figure 3 , Figure 3 This is a schematic diagram of the principle of a virtual fitting method provided in Example 1 of the present application;
[0087] like Figure 3 As shown, the first human body image (I) and the first clothing image (I c ).
[0088] Step S202: obtaining human body image analysis information according to the first human body image;
[0089] like Figure 3 As shown, after obtaining the first human body image (I), processing is performed to obtain human body image analysis information (p * ), specifically, the human body image analysis information includes: a human body mask map, a human body dense point segmentation map and a human body key point map, wherein the human body mask map is obtained by segmenting the human body in the first human body image and removing the upper part; the human body dense point segmentation map is the human body dense posture key point (Densepose), and the human body dense point segmentation map is processed by the human semantic parser (Human Parsing); the human body key point map is obtained by acquiring the human body key points by the human posture estimator of the body tracking system (OpenPose). Optionally, other methods can also be used to acquire the human body key points, which is not limited in the embodiments of the present application.
[0090] Step S203: processing the human body image analysis information and the first clothing image through the appearance flow deformation module to generate first deformation information;
[0091] Specifically, the human body image parsing information and the first clothing image are input into the appearance flow deformation module, which processes the first deformation information. The appearance flow deformation module includes: a parser-based appearance flow warping module (PB-AFWM), such as Figure 3 As shown, the parser appearance flow deformation module (PB-AFWM) is used to parse the human body image information (p* ) and the first clothing image (I c ), and processes to obtain first deformation information (uf). The parser-based appearance flow warping module (PB-AFWM) includes an appearance flow warping module (AFWM), which is used to predict a dense correspondence between the clothing image and the person image so as to deform the clothing. The output of the appearance flow warping module is the appearance flow, i.e., the first deformation information (uf), which is a set of two-dimensional coordinate vectors, each vector indicating which pixels in the clothing image should be used to fill a specific pixel in the person image.
[0092] Please refer to Figure 4 , Figure 4 is a schematic diagram of an appearance flow deformation module provided in Example 1 of the present application;
[0093] like Figure 4 As shown in Figure 3, the appearance flow warping module (AFWM) consists of two pyramid feature extraction networks (PFEN) and an appearance flow estimation network (AFEN). PFEN extracts two pyramid deep features from two inputs. Then, at each pyramid level, AFEN learns to generate a coarse appearance flow that is then refined at the next pyramid level. A second-order smoothness constraint is also employed when learning the appearance flow to further preserve clothing features such as logos and stripes.
[0094] Among them, the pyramid feature extraction network (PFEN) includes a feature extraction network and a feature fusion structure, such as backbone+FPN. Backbone is a commonly used feature extraction network, including vgg network, resnet network, etc. FPN is a feature fusion structure, which can better extract features.
[0095] Among them, the appearance flow estimation network (AFEN) consists of N appearance flow generation networks (flow networks, FN), such as: FN-1, FN-2, FN-3, and estimates the appearance flow from N levels of pyramid features, such as: first extract the pyramid features (C N , P N ) Input the appearance flow generation network FN-1 at the Nth pyramid layer to estimate the initial appearance flow f1, then input f1 and the pyramid features at the N-1th layer into the appearance flow generation network FN-2 to obtain a better appearance flow f2, and so on, until the best appearance flow f is obtained N , and according to the appearance flow f NThe first clothing image is deformed to obtain first deformation information.
[0096] Specifically, each appearance flow generation network (FN) performs pixel-by-pixel feature matching to produce a rough flow estimate and optimizes it at each pyramid level. Taking FN-2 as an example, its input is two pyramid features (c2, p2) and the appearance flow f1 of the previous pyramid level. The operation of FN can be roughly divided into four stages:
[0097] In the first stage, the initial appearance flow f1 is upsampled to obtain the appearance flow f′1, and then the vectors in c2 are sampled to transform c2 into c′2, where the sampling position is specified by f′1;
[0098] In the second stage, the association map r2 is calculated based on c′2 and p2, where the jth point in r2 is a vector that represents the result of the vector matrix product of the jth point in c′2 and the local displacement region centered at the jth point in p2. In this case, the number of channels in r2 is equal to the number of points in the local displacement region;
[0099] In the third stage, once r2 is obtained, it is fed into the Conv network to predict the residual appearance flow of f′″2 and added to f′1 as the coarse appearance flow f″2;
[0100] In the fourth stage, c2 is deformed into c″2 according to the newly generated appearance flow f″2. Then, c″2 and p2 are connected and input into the Conv network to calculate the remaining appearance flow f′2. By adding the appearance flow f′2 to the appearance flow f″2, the final appearance flow f2 is obtained in the appearance flow generation network FN-2.
[0101] In the embodiment of the present application, the appearance flow estimation network (AFEN) gradually corrects the estimated appearance flow through the cascaded N appearance flow generation networks (FN) to capture the long-distance correspondence between the clothing image and the character image, which is beneficial for processing misalignment and deformation and enhancing the effect of clothing deformation.
[0102] Step S204: obtaining a second clothing image, and deforming the second clothing image according to the first deformation information to generate a third clothing image;
[0103] Specifically, the second clothing image is a target image to be virtually dressed for the user, and the first deformation information (u f ) for the second clothing image Perform warping to obtain the third clothing image The third clothing image is an image obtained by deforming the second clothing image.
[0104] Step S205: fusing the third clothing image and the human body image analysis information to generate a first fitting image;
[0105] like Figure 3 As shown, the third clothing image is fused and human body image analysis information (p * ), generate the first try-on image Specifically, fusing the third clothing image and the human body image analysis information to generate the first fitting image includes:
[0106] The third clothing image and the human body image analysis information are input into the first generation module to fuse the third clothing image and the human body image analysis information to generate a first fitting image.
[0107] The first generation module includes a parser-based generative module (PB-GM), such as Figure 3 As shown, the input of the parser generation module (PB-GM) is the third clothing image The first try-on image is obtained by fusing the third clothing image and the human body image analysis information through the parser generation module (PB-GM).
[0108] Step S206: deforming the first clothing image according to the first deformation information to generate a fourth clothing image;
[0109] Specifically, such as Figure 3 As shown, through the first deformation information (u f ) for the first clothing image (I c ) is deformed (warped) to generate a fourth clothes image (Sw), wherein the fourth clothes image (Sw) is an image obtained after the first clothes image is deformed.
[0110] Step S207: Fusing the fourth clothing image and the first fitting image to generate a target clothing image.
[0111] Specifically, fusing the fourth clothing image and the first try-on image to generate a target clothing image includes:
[0112] The fourth clothing image and the first try-on image are input into the second generation module to fuse the fourth clothing image and the first try-on image to generate a target clothing image.
[0113] The second generation module includes a parser-free generative module (PF-GM), such as Figure 3 As shown, the input of the parser-free generation module (PF-GM) is the first try-on image and the fourth clothing image (S W ), wherein the fourth clothing image (Sw) is obtained by the first deformation information (u f ) for the first clothing image (I c ) is deformed (warped) to obtain it.
[0114] In an embodiment of the present application, the method further includes:
[0115] The first generation module performs knowledge distillation on the second generation module, so that the second generation module learns the feature extraction method and feature fusion method of the first generation module.
[0116] Specifically, the first generation module includes a parser generation module (PB-GM), and the second generation module includes a parser-free generation module (PF-GM), such as Figure 3 As shown, the parser-free generation module (PF-GM) is distilled and learned by the parser-based generation module (PB-GM), including: the parser-based generation module (PB-GM) is a teacher network, and the parser-free generation module (PF-GM) is a student network. The loss or intermediate features of the teacher network (PB-GM) are used to constrain the loss and intermediate features of the student network (PF-GM), thereby learning the feature extraction method and feature fusion method of the teacher network (PB-GM). It can be understood that the first generation module and the second generation module are both generation networks, such as deep learning image segmentation networks, such as UNet networks, which are used to generate target images from different input images through loss functions or constraints.
[0117] Specifically, the first generation module includes a first appearance flow deformation module, the second generation module includes a second appearance flow deformation module, and knowledge distillation is performed on the second generation module through the first generation module, including:
[0118] The feature layer in the first appearance flow warping module is used to calculate the distance between the feature layer in the second appearance flow warping module and the feature layer in the second appearance flow warping module. The calculated distance is used as a loss function to train the second generation module so that the feature layer in the second appearance flow warping module is close to the feature layer in the first appearance flow warping module. The first appearance flow warping module and the second appearance flow warping module are both appearance flow warping modules (AFWMs) with the same structure.
[0119] For example, the feature layer (p1, p2, p3) in the PFN module in the appearance flow warping module (AFWM) is used. PF-GM and PB-GM share the same appearance flow warping module. The Euclidean distance between (p1, p2, p3) in PF-GM and PB-GM is then calculated and used as the loss function to train the PF-GM model. This allows the feature layer (p1, p2, p3) in PF-GM to gradually approach the feature layer (p1, p2, p3) in PB-GM, thus completing the distillation learning process.
[0120] The first generation module performs knowledge distillation on the second generation module, so that the second generation module can learn the feature extraction and feature fusion methods of the first generation module, which can improve the network expression ability of the parser-free generation module (PF-GM) and is conducive to improving the effect of clothing fusion.
[0121] In an embodiment of the present application, the method further includes:
[0122] Obtain a generative adversarial network, wherein the generative adversarial network includes a generative network and a discriminative network; obtain a first loss function corresponding to the generative network and a second loss function of the discriminative network; and train a second generation module based on the first loss function and the second loss function.
[0123] Specifically, the generative adversarial network (GAN) is a new framework for estimating generative models through an adversarial process. GANs are a type of deep generative model that uses adversarial training. Both the discriminator and the generator can use different network structures based on different generative tasks. In a GAN, the generator and the discriminator engage in a non-cooperative zero-sum game. The generator captures the underlying distribution of real data samples and generates new data samples; the discriminator is a binary classifier that distinguishes whether the input is real data or generated samples. Both the generator and the discriminator can use perceptrons or deep learning models. The optimization process of the GAN is a minimax game problem, with the goal of reaching a Nash equilibrium, that is, until the discriminator cannot distinguish whether fake samples generated by the generator are real or fake. Upon convergence, if the discriminator can no longer distinguish the source of a sample, it means that the generator can generate samples that conform to the real data distribution.
[0124] Specifically, by obtaining a first loss function corresponding to the generation network and a second loss function of the discrimination network; based on the first loss function and the second loss function, adversarial training is performed on the second generation module.
[0125] In an embodiment of the present application, when training the second generation module, an adversarial training method is adopted, which is beneficial to improving the generation ability of the second generation module, so that the second generation module has better retention ability and fusion effect, which is beneficial to generating realistic target clothing images.
[0126] In an embodiment of the present application, a virtual fitting method is provided, which includes: obtaining a first human body image and a first clothing image; obtaining human body image analysis information based on the first human body image; processing the human body image analysis information and the first clothing image through an appearance flow deformation module to generate first deformation information; obtaining a second clothing image, deforming the second clothing image through the first deformation information to generate a third clothing image; fusing the third clothing image and the human body image analysis information to generate a first fitting image; deforming the first clothing image through the first deformation information to generate a fourth clothing image; and fusing the fourth clothing image and the first fitting image to generate a target clothing image.
[0127] On the one hand, by using the first human image, human image parsing information is obtained, and based on the human image parsing information and the first clothing image, the first deformation information is processed by the appearance flow deformation module to generate first deformation information. The appearance flow can be used for distillation learning to retain clothing features and details;
[0128] On the other hand, by obtaining a second clothing image, deforming the second clothing image according to the first deformation information, a third clothing image is generated; fusing the third clothing image and the human body image analysis information, a first try-on image is generated; deforming the first clothing image according to the first deformation information, a fourth clothing image is generated; fusing the fourth clothing image and the first try-on image, a target clothing image is generated, and the first deformation information can be used to deform the first clothing image, and the human body image analysis information can be fused to obtain the first try-on image, and the fourth clothing image can be fused, which can improve the clothing fusion effect.
[0129] Please refer to Figure 5 , Figure 5 This is a structural diagram of a virtual fitting device provided in Example 1 of the present application;
[0130] The virtual fitting device is applied to a server. Specifically, the virtual fitting device is applied to one or more processors of the server.
[0131] like Figure 1 As shown, the virtual fitting device 50 includes:
[0132] The acquisition unit 501 is configured to acquire a first human body image and a first clothing image; and acquire human body image analysis information based on the first human body image;
[0133] The deformation unit 502 is configured to process the human body image parsing information and the first clothing image through the appearance flow deformation module to generate first deformation information; and to deform the first clothing image using the first deformation information to generate a fourth clothing image;
[0134] a generating unit 503 configured to obtain a second clothing image, deform the second clothing image according to the first deformation information, and generate a third clothing image;
[0135] The fusion unit 504 is configured to fuse the third clothing image and the human body image analysis information to generate a first fitting image; and to fuse the fourth clothing image and the first fitting image to generate a target clothing image.
[0136] In the embodiments of the present application, the virtual fitting device can also be constructed by hardware devices. For example, the virtual fitting device can be constructed by one or more chips, and the chips can work in coordination with each other to complete the virtual fitting methods described in the above embodiments. For another example, the virtual fitting device can also be constructed by various logic devices, such as a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a single-chip microcomputer, an ARM (Acorn RISC Machine) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of these components.
[0137] The virtual fitting device in the embodiment of the present application can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The non-mobile electronic device can be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), an ATM, or an kiosks, etc., which are not specifically limited in the embodiment of the present application.
[0138] The virtual fitting device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.
[0139] The virtual fitting device provided in the embodiment of the present application can achieve Figure 2 To avoid repetition, the various implementation processes will not be described here.
[0140] It should be noted that the virtual fitting device described above can execute the virtual fitting method provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects of the execution method. For technical details not fully described in the embodiments of the boot management device, please refer to the virtual fitting method provided in the embodiments of this application.
[0141] In an embodiment of the present application, a virtual fitting device is provided, which includes: an acquisition unit for acquiring a first human body image and a first clothing image; and, based on the first human body image, acquiring human body image analysis information; a deformation unit for processing the human body image analysis information and the first clothing image through an appearance flow deformation module to generate first deformation information; and, for deforming the first clothing image through the first deformation information to generate a fourth clothing image; a generation unit for acquiring a second clothing image, deforming the second clothing image through the first deformation information to generate a third clothing image; a fusion unit for fusing the third clothing image and the human body image analysis information to generate a first try-on image; and, fusing the fourth clothing image and the first try-on image to generate a target clothing image. On the one hand, by using the first human body image, the human body image parsing information is obtained, and based on the human body image parsing information and the first clothing image, the appearance flow deformation module is used for processing to generate the first deformation information, and the appearance flow can be used for distillation learning to retain the characteristics and details of the clothing; on the other hand, by obtaining the second clothing image, the second clothing image is deformed by the first deformation information to generate the third clothing image; the third clothing image and the human body image parsing information are fused to generate the first try-on image; the first clothing image is deformed by the first deformation information to generate the fourth clothing image; the fourth clothing image and the first try-on image are fused to generate the target clothing image, and the first deformation information can be used to deform the first clothing image, and the human body image parsing information can be fused to obtain the first try-on image, and the fourth clothing image can be fused to improve the clothing fusion effect.
[0142] Example 2
[0143] See also Figure 6 , Figure 6 This is a flowchart of a virtual fitting method provided in Example 2 of the present application;
[0144] The virtual fitting method is applied to a server. Specifically, the execution subject of the virtual fitting method is one or more processors of the server.
[0145] like Figure 6 As shown, the virtual fitting method includes:
[0146] Step S601: Acquire a first human body image and a first clothing image;
[0147] Specifically, the first human image is an image of a person wearing clothing, i.e., the first human image includes both person information and clothing information. It is understood that the first human image also includes a background portion. The person in the first human image can have various poses or postures, such as with hands on hips or with hands hanging naturally, and this embodiment of the application does not limit this.
[0148] Step S602: obtaining human body image analysis information according to the first human body image;
[0149] Please refer to Figure 3 ,like Figure 3 As shown, after obtaining the first human body image (I), processing is performed to obtain human body image analysis information (p * ), specifically, the human body image analysis information includes: a human body mask map, a human body dense point segmentation map and a human body key point map, wherein the human body mask map (densemask) is obtained by segmenting the human body in the first human body image and removing the upper part; the human body dense point segmentation map is the human body dense posture key point (Densepose), and the human body dense point segmentation map is processed by the human semantic parser (Human Parsing); the human body key point map is obtained by acquiring the human body key points by the human posture estimator of the body tracking system (OpenPose). Optionally, other methods can also be used to acquire the human body key points, and the embodiments of the present application are not limited to this.
[0150] Step S603: Processing the human body image analysis information and the first clothing image through the appearance flow deformation module to generate first deformation information;
[0151] Specifically, the human body image parsing information and the first clothing image are input into the appearance flow warping module, which processes the first deformation information. The appearance flow warping module includes a parser-based appearance flow warping module (PB-AFWM), such as Figure 3 As shown, the parser appearance flow deformation module (PB-AFWM) is used to parse the human body image information (p * ) and the first clothing image (I c ), and process to obtain the first deformation information (u fThe parser-based appearance flow warping module (PB-AFWM) includes an appearance flow warping module (AFWM) for predicting dense correspondences between clothing images and person images so as to deform clothing. The output of the appearance flow warping module is an appearance flow, i.e., first deformation information (uf), which is a set of two-dimensional coordinate vectors, each vector indicating which pixels in the clothing image should be used to fill a specific pixel in the person image.
[0152] Specifically, regarding the relevant processing process of the appearance flow deformation module (AFWM), reference may be made to the relevant description of the above embodiment 1, which will not be repeated here.
[0153] Step S604: obtaining a second clothing image, and deforming the second clothing image according to the first deformation information to generate a first deformed clothing image;
[0154] For details, please refer to Figure 7 , Figure 7 This is a schematic diagram of the principle of a virtual fitting method provided in Example 2 of the present application;
[0155] like Figure 7 As shown, the second clothing image is the target image to be virtually dressed for the user, and the first deformation information (u f ) for the second clothing image Perform warping to obtain the first deformed clothing image The first deformed clothing image is an image obtained by deforming the second clothing image.
[0156] Step S605: Acquire a first hand image according to the first human body image, wherein the first hand image includes hand information and arm information;
[0157] like Figure 7As shown, by processing the first human body image (I), a first hand image is obtained, wherein the first hand image includes hand information and arm information. Specifically, the first human body image is processed, and the first hand image is obtained by identifying the hand information and arm information by a neural network, and extracting the hand information and arm information from the first hand image, and segmenting the first hand image to obtain the first hand image, wherein the neural network includes: a convolutional neural network, wherein the convolutional neural network includes a convolutional layer, a pooling layer and a fully connected layer, and the convolutional neural network is composed of a cross-stack of a convolutional layer, a pooling layer and a fully connected layer. The function of the convolutional layer is to extract the features of a local area, and different convolution kernels are equivalent to different feature extractors. The pooling layer (Pooling Layer) is also called a subsampling layer (Subsampling Layer), and its function is to perform feature selection, reduce the number of features, and thus reduce the number of parameters. Preferably, the convolutional neural network in the embodiment of the present application is a deep convolution generative adversarial network.
[0158] Step S606: fusing the first deformed clothing image, the first hand image, and the first human body image to determine a first fused clothing image;
[0159] Specifically, fusing the first deformed clothing image, the first hand image, and the first human body image to determine a first fused clothing image includes:
[0160] The first deformed clothing image, the first hand image, and the human body dense point segmentation map are fused through the first generation module to determine a first fused clothing image.
[0161] The first generation module includes a parser-based generative module (PB-GM).
[0162] like Figure 7 As shown, the input of the parser-based generative module (PB-GM) is the first deformed clothing image The first hand image and the human body dense point segmentation map are output as the first fused clothing image.
[0163] In the embodiment of the present application, since the clothing information, hand information, and arm information are specifically fused, the original features of the clothing information, hand information, and arm information are fully preserved and the fusion effect is improved.
[0164] Step S607: acquiring a second human body image based on the first human body image, wherein the second human body image is obtained by removing clothing information and arm information from the first human body image;
[0165] Specifically, the clothing information and arm information in the first human body image are identified by a neural network, and the clothing information and arm information in the first human body image are removed to obtain the second human body image. The neural network includes a convolutional neural network.
[0166] In an embodiment of the present application, the method further includes:
[0167] Obtain a generative adversarial network, wherein the generative adversarial network includes a generative network and a discriminative network;
[0168] Based on the generative adversarial network, the first generation module is trained adversarially.
[0169] Generative adversarial networks (GANs) are a new framework for estimating generative models through an adversarial process. GANs are a class of deep generative models that use adversarial training. Both the discriminator and the generator can use different network structures based on the generative task. In a GAN, the generator and the discriminator engage in a non-cooperative zero-sum game. The generator captures the underlying distribution of real data samples and generates new data samples; the discriminator is a binary classifier that distinguishes whether the input is real data or generated samples. Both the generator and the discriminator can use perceptrons or deep learning models. The optimization process of the GAN is a minimax game, with the goal of reaching a Nash equilibrium, that is, until the discriminator cannot distinguish whether fake samples generated by the generator are real or fake. Upon convergence, if the discriminator can no longer distinguish the source of a sample, it means that the generator can generate samples that conform to the real data distribution.
[0170] Specifically, by obtaining a first loss function corresponding to the generation network and a second loss function of the discrimination network; based on the first loss function and the second loss function, adversarial training is performed on the first generation module.
[0171] In an embodiment of the present application, when training the first generation module, an adversarial training method is adopted, which is beneficial to improving the generation ability of the first generation module, so that the first generation module has better retention ability and fusion effect, which is beneficial to generating realistic target clothing images.
[0172] Step S608: Fusing the first deformed clothing image, the first fused clothing image, the first human body image, and the second human body image to obtain a target clothing image.
[0173] Specifically, fusing the first deformed clothing image, the first fused clothing image, the first human body image, and the second human body image to obtain a target clothing image includes:
[0174] The target clothing image is obtained by fusing the first deformed clothing image, the first fused clothing image, the human body dense point segmentation map, and the second human body image through a first generation module, wherein the first generation module includes a parser generation module (PB-GM).
[0175] like Figure 7 As shown, the input of the parser generation module (PB-GM) is the first deformed clothing image, the first fused clothing image, the human body dense point segmentation map and the second human body image, and the output is the target clothing image.
[0176] In an embodiment of the present application, by utilizing a first fused clothing image including hand information and arm information, in combination with a second human body image from which clothing information and arm information are removed, and in combination with a human body dense point segmentation map, the hand information and arm information can be better retained, thereby improving the fusion effect. Moreover, when changing from long sleeves to short sleeves, the fusion effect at the arms can be improved, thereby generating a target clothing image with a realistic fusion effect.
[0177] In an embodiment of the present application, a virtual fitting method is provided, including: obtaining a first human body image and a first clothing image; obtaining human body image analysis information based on the first human body image; processing the human body image analysis information and the first clothing image through an appearance flow deformation module to generate first deformation information; obtaining a second clothing image, deforming the second clothing image through the first deformation information to generate a first deformed clothing image; obtaining a first hand image based on the first human body image, wherein the first hand image includes hand information and arm information; fusing the first deformed clothing image, the first hand image and the first human body image to determine a first fused clothing image; obtaining a second human body image based on the first human body image, wherein the second human body image is obtained by removing clothing information and arm information from the first human body image; fusing the first deformed clothing image, the first fused clothing image, the first human body image and the second human body image to obtain a target dressing image.
[0178] A first hand image and a second human body image are obtained by using the first human body image, and a first fused clothing image is determined by using the first deformed clothing image, the first hand image, and the first human body image. The first human body image, the first deformed clothing image, and the second human body image are fused to obtain a target clothing image. Since the first deformed clothing image includes clothing information and the first hand image includes hand information and arm information, the fused target clothing image can fully retain the features of the clothing, hand, and arm, thereby improving the clothing fusion effect.
[0179] Please refer to Figure 8 , Figure 8 This is a structural diagram of a virtual fitting device provided in Example 2 of the present application;
[0180] The virtual fitting device is applied to a server. Specifically, the virtual fitting device is applied to one or more processors of the server.
[0181] like Figure 8 As shown, the virtual fitting device 80 includes:
[0182] The acquisition unit 801 is configured to acquire a first human body image and a first clothing image; and acquire human body image analysis information based on the first human body image;
[0183] The deformation unit 802 is configured to process the human body image analysis information and the first clothing image through the appearance flow deformation module to generate first deformation information; obtain a second clothing image, and deform the second clothing image using the first deformation information to generate a first deformed clothing image;
[0184] The fusion unit 803 is used to obtain a first hand image based on the first human body image, wherein the first hand image includes hand information and arm information; fuse the first deformed clothing image, the first hand image and the first human body image to determine a first fused clothing image; obtain a second human body image based on the first human body image, wherein the second human body image is obtained by removing the clothing information and arm information from the first human body image; and fuse the first deformed clothing image, the first fused clothing image, the first human body image and the second human body image to obtain a target clothing image.
[0185] In the embodiments of the present application, the virtual fitting device can also be constructed by hardware devices. For example, the virtual fitting device can be constructed by one or more chips, and the chips can work in coordination with each other to complete the virtual fitting methods described in the above embodiments. For another example, the virtual fitting device can also be constructed by various logic devices, such as a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a single-chip microcomputer, an ARM (Acorn RISC Machine) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of these components.
[0186] The virtual fitting device in the embodiment of the present application can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The non-mobile electronic device can be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), an ATM, or an kiosks, etc., which are not specifically limited in the embodiment of the present application.
[0187] The virtual fitting device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.
[0188] The virtual fitting device provided in the embodiment of the present application can achieve Figure 6 To avoid repetition, the various implementation processes will not be described here.
[0189] It should be noted that the virtual fitting device described above can execute the virtual fitting method provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects of the execution method. For technical details not fully described in the embodiments of the boot management device, please refer to the virtual fitting method provided in the embodiments of this application.
[0190] In an embodiment of the present application, a virtual fitting device is provided, which includes: an acquisition unit for acquiring a first human body image and a first clothing image; acquiring human body image analysis information based on the first human body image; a deformation unit for processing the human body image analysis information and the first clothing image through an appearance flow deformation module to generate first deformation information; acquiring a second clothing image, deforming the second clothing image through the first deformation information to generate a first deformed clothing image; a fusion unit for acquiring a first hand image based on the first human body image, wherein the first hand image includes hand information and arm information; fusing the first deformed clothing image, the first hand image and the first human body image to determine a first fused clothing image; acquiring a second human body image based on the first human body image, wherein the second human body image is obtained by removing clothing information and arm information from the first human body image; fusing the first deformed clothing image, the first fused clothing image, the first human body image and the second human body image to obtain a target clothing image. A first hand image and a second human body image are obtained by using the first human body image, and a first fused clothing image is determined by using the first deformed clothing image, the first hand image, and the first human body image. The first human body image, the first deformed clothing image, and the second human body image are fused to obtain a target clothing image. Since the first deformed clothing image includes clothing information and the first hand image includes hand information and arm information, the fused target clothing image can fully retain the features of the clothing, hand, and arm, thereby improving the clothing fusion effect.
[0191] Example 3
[0192] See also Figure 9 , Figure 9 This is a flowchart of a virtual fitting method provided in Example 3 of the present application;
[0193] The virtual fitting method is applied to a server. Specifically, the execution subject of the virtual fitting method is one or more processors of the server.
[0194] like Figure 9 As shown, the virtual fitting method includes:
[0195] Step S901: Acquire a first target human body image and a first target clothing image, wherein the first target human body image and the first target clothing image are both high-resolution images;
[0196] Specifically, the first target human image is an image of a person wearing clothing, i.e., the first target human image includes both person information and clothing information. It is understood that the first target human image also includes a background portion. The person in the first human image can have various poses or postures, such as with hands on hips or with hands hanging naturally, and this embodiment of the application is not limited thereto.
[0197] In the embodiment of the present application, the first target human image and the first target clothing image are both high-resolution images, including high-definition images. For example, the resolutions of the first target human image and the first target clothing image are both 512*384, 1024*768, 1280*720, or 1920*1080. It is understood that the resolutions of the first target human image and the first target clothing image in the embodiment of the present application can also be other resolutions, which are not limited here.
[0198] Step S902: scaling the first target human body image and the first target clothing image to generate a second target human body image and a second target clothing image;
[0199] Specifically, the first target human body image and the first target clothing image are scaled to obtain low-resolution images. That is, the high-resolution first target human body image is scaled to a low-resolution second target human body image, and the high-resolution first target clothing image is scaled to a low-resolution second target clothing image. For example, the second target human body image and the second target clothing image are both 256*192. It is understood that the resolution of the second target human body image is lower than that of the first target human body image, and the resolution of the second target clothing image is lower than that of the first target clothing image. The resolutions of the second target human body image and the second target clothing image in the embodiments of the present application can also be other resolutions, which are not limited here.
[0200] Step S903: generating a first target fused image according to the second target human body image and the second target clothing image;
[0201] Specifically, generating a first target fused image according to the second target human body image and the second target clothing image includes:
[0202] The second target human body image and the second target clothing image are input into the appearance flow deformation module and the generation module to generate the first target fusion image.
[0203] Specifically, the second target human body image and the second target clothing image are input into the appearance flow deformation module and the generation module to generate the first target fused image, including:
[0204] Inputting the second target human body image and the second target clothing image into the appearance flow deformation module to generate second deformation information;
[0205] deforming the second target clothing image using the second deformation information to generate a third target clothing image;
[0206] The second target human body image and the third target clothing image are input into the generation module to generate a first target fusion image.
[0207] See also Figure 10 , Figure 10 This is a schematic diagram of the principle of generating a first target fused image provided in Example 3 of the present application;
[0208] like Figure 10 As shown, the appearance flow warping module includes: a parser-free appearance flow warping module (PF-AFWM), and the generation module includes: a parser-free generative module (PF-GM), which resizes the first target human image and the first target clothing image to generate a second target human image and a second target clothing image, inputs the second target human image and the second target clothing image into the parser-free appearance flow warping module (PF-AFWM) to obtain second deformation information (Sf), warps the second target clothing image (Ic) according to the second deformation information (Sf) to generate a third target clothing image (Sw), and inputs the second target human image and the third target clothing image (Sw) into the parser-free generative module (PF-GM) to obtain a first target fused image (output1).
[0209] In an embodiment of the present application, the appearance flow warping module includes: a first appearance flow warping module and a second appearance flow warping module, wherein the first appearance flow warping module includes a parser-based appearance flow warping module (PB-AFWM), and the second appearance flow warping module includes: a parser-free appearance flow warping module (PF-AFWM). The second appearance flow warping module is used to generate second deformation information, and the second deformation information is a set of two-dimensional coordinate vectors, each vector indicating which pixels in the clothing image should be used to fill specific pixels in the character image.
[0210] The method further includes:
[0211] Training the second appearance flow deformation module includes:
[0212] Acquire a first human body image and a first clothing image;
[0213] Acquiring human body image analysis information according to the first human body image;
[0214] Processing the human body image analysis information and the first clothing image through a first appearance flow deformation module to generate first deformation information;
[0215] Acquire a second clothing image, and deform the second clothing image according to the first deformation information to generate a third clothing image;
[0216] Fusing the third clothing image and the human body image analysis information to generate a first fitting image;
[0217] A second appearance flow deformation module is trained based on the first clothing image and the first try-on image.
[0218] The method further includes:
[0219] During the training process of the second appearance manifold deformation module, knowledge distillation is performed on the second appearance manifold deformation module through the first appearance manifold deformation module.
[0220] For details, please refer to Figure 11 , Figure 11 Schematic diagram of the principle of training a second appearance flow deformation module provided in the third embodiment of the present application;
[0221] like Figure 11 As shown, the first human body image (I) and the first clothing image (I c ), according to the first human body image, obtain human body image analysis information (p * ), the human body image analysis information includes: human body mask map, human body dense point segmentation map and human body key point map, according to the human body image analysis information (p * ) and the first clothing image (I c ), processed by the first appearance flow deformation module to generate the first deformation information (u f ), wherein the first appearance flow warping module includes: a parser-based appearance flow warping module (PB-AFWM), which uses the first deformation information (u f ) for the second clothing image Perform warping to obtain the first deformed clothing image Then the first deformed clothing image is fused through the parser-based generative module (PB-GM) And human body image analysis information (p * ), get the first try-on image And based on the first try-on image And the first clothing image (I c), train a second appearance flow warping module, wherein the second appearance flow warping module includes: a parser-free appearance flow warping module (PF-AFWM).
[0222] Specifically, during the training process of the parser-free appearance flow deformation module (PF-AFWM), knowledge distillation (knowledge distillation), such as adjustable knowledge distillation (adjustable knowledge distillation), is performed through the parser-based appearance flow deformation module (PB-AFWM).
[0223] Step S904: generating a first target processed image according to the first target human body image, wherein the first target processed image does not include clothing information and arm information;
[0224] Specifically, the clothing information and arm information in the first target human body image are identified through a neural network, and the clothing information and arm information in the first target human body image are removed to obtain a first target processed image, wherein the neural network includes a convolutional neural network; or, the first target human body image is obtained, processed to obtain a human body mask image corresponding to the first target human body image, and the human body mask image is used to remove the clothing information and arm information in the first target human body image to obtain the first target processed image.
[0225] Step S905: inputting the first target fusion image, the first target clothing image and the first target processed image into the first clothing model to generate a target clothing image.
[0226] Specifically, the first clothing model is a pre-trained model, which is used to fuse the input first target fusion image, the first target clothing image and the first target processed image to obtain a target clothing image, wherein the target clothing image is the final virtual clothing image, which is used to be presented to the user's terminal device, such as a mobile terminal or a fixed terminal.
[0227] In an embodiment of the present application, the method further includes:
[0228] Pre-train the first clothing model, specifically including:
[0229] Acquire a first data set, wherein the first data set includes a first human body image and a first clothing image, wherein the first human body image and the first clothing image are both high-resolution images;
[0230] Scaling the first human body image and the first clothing image in the first data set to generate a second data set, wherein the second data set includes the scaled first human body image and the scaled first clothing image;
[0231] generating a third data set according to the scaled first human body image and the scaled first clothing image in the second data set, wherein the third data set includes the first fused image;
[0232] generating a first processed image according to the first human body image, wherein the first processed image does not include clothing information and arm information;
[0233] Scaling the first fused image to generate a second fused image, wherein the second fused image has the same resolution as the first human body image;
[0234] A first human body image, a first processed image, and a second fused image are input to train a first clothing model.
[0235] For details, please refer to Figure 12 , Figure 12 This is a schematic diagram of the principle of training a first clothing model provided in Example 3 of the present application;
[0236] First, a first data set (data1) is obtained, wherein the first data set includes a first human body image and a first clothing image, wherein the first human body image and the first clothing image are both high-resolution images, and the first data set also includes a human body mask image;
[0237] The images in the first dataset (data1) are scaled and processed into a second dataset (data2) with low resolution images. For example, the resolution of the images in the first dataset (data1) is 512*384, and the resolution of the images in the second dataset (data2) is 256*192.
[0238] A third data set is generated based on the scaled first human image and the scaled first clothing image in the second data set, wherein the third data set includes the first fused image. Specifically, the scaled first human image and the scaled first clothing image in the second data set are input into a pre-trained parser-free appearance flow warping module (PF-AFWM) and a pre-trained parser-free generative module (PF-GM) to obtain multiple first fused images (output1), and the multiple first fused images are combined to form the third data set. For the specific processing process, please refer to the relevant description of the above embodiment and the drawings in the specification. Figure 10, I will not go into details here.
[0239] Then, a first processed image is generated based on the first human body image and the human body mask image corresponding to the first human body image, wherein the first processed image does not include clothing information and arm information, that is, the first processed image is an image with the clothing information and arm information removed, and is recorded as (un-cloth);
[0240] The first fused image (output1) is then resized to generate a second fused image (output2), wherein the second fused image (output2) has the same resolution as the first human body image. That is, the first fused image (output1) is enlarged to the same resolution as the images in the first data set (data1);
[0241] Finally, the first human body image, the first processed image (un-cloth), and the second fused image (output2) are input to train the first clothing model (cloth-HD), wherein the first clothing model is used to output a high-resolution target clothing image. Specifically, the first clothing model (cloth-HD) can adopt a commonly used Unet structure generation model and be trained using a loss function. At the same time, the first human body image in the first data set (data1) is used as the learning object (label) during training. It can be understood that the loss function is a non-negative real number function used to quantify the difference between the predicted label and the true label predicted by the model.
[0242] For example: the loss function is L = α*L1 loss function + β*perceptual loss function + γ*GAN loss function, that is, L = α*L1 loss + β*Perceptual loss + γ*GAN loss, where the L1 loss function, namely the mean absolute error (MAE), is a loss function used for regression models. MAE is the sum of the absolute differences between the target variable and the predicted variable, so it measures the average error size in a set of predicted values, regardless of their direction; the perceptual loss function, namely the human eye perception loss, is calculated in the feature space. The perceptual loss function is based on the high-level features provided by the pre-trained network to train a forward propagation network for the image conversion task. During the training process, the perceptual error measures the similarity between images and can run in real time during testing; the GAN loss function, namely the loss function of the generative adversarial network, generally adopts the cross entropy loss function. There are two neural networks and two loss functions in its training process, namely the generating network and its corresponding loss function, and the discriminant network and its corresponding loss function. In the embodiment of the present application, the GAN loss function includes the Relativistic GAN loss function.
[0243] In the embodiment of the present application, α, β and γ are all adjustable parameters, and their parameters are set according to specific needs, for example: setting α=0.1, β=0.1, γ=2.
[0244] In an embodiment of the present application, a virtual fitting method is provided, including: obtaining a first target human body image and a first target clothing image, wherein the first target human body image and the first target clothing image are both high-resolution images; scaling the first target human body image and the first target clothing image to generate a second target human body image and a second target clothing image; generating a first target fused image based on the second target human body image and the second target clothing image; generating a first target processed image based on the first target human body image, wherein the first target processed image does not include clothing information and arm information; and inputting the first target fused image, the first target clothing image, and the first target processed image into a first clothing model to generate a target dressing image.
[0245] By acquiring a high-resolution first target human body image and a second target clothing image and scaling them, a second target human body image and a second target clothing image are obtained to generate a first target fused image, and a first target processed image excluding clothing information and arm information is generated based on the first target human body image, and then the first target fused image, the first target clothing image and the first target processed image are input into the first clothing model to generate a target clothing image. The embodiment of the present application can generate a high-resolution target clothing image and improve the fusion effect of clothing.
[0246] Please refer to Figure 13 , Figure 13 This is a structural diagram of a virtual fitting device provided in Example 3 of the present application;
[0247] The virtual fitting device is applied to a server. Specifically, the virtual fitting device is applied to one or more processors of the server.
[0248] like Figure 13 As shown, the virtual fitting device 130 includes:
[0249] An acquiring unit 1301 is configured to acquire a first target human body image and a first target clothing image, wherein both the first target human body image and the first target clothing image are high-resolution images;
[0250] A scaling unit 1302 is configured to scale the first target human body image and the first target clothing image to generate a second target human body image and a second target clothing image;
[0251] A generating unit 1303 is configured to generate a first target fused image based on the second target human body image and the second target clothing image; and generate a first target processed image based on the first target human body image, wherein the first target processed image does not include clothing information and arm information;
[0252] The fusion unit 1304 is configured to input the first target fused image, the first target clothing image, and the first target processed image into the first clothing model to generate a target clothing image.
[0253] In the embodiments of the present application, the virtual fitting device can also be constructed by hardware devices. For example, the virtual fitting device can be constructed by one or more chips, and the chips can work in coordination with each other to complete the virtual fitting methods described in the above embodiments. For another example, the virtual fitting device can also be constructed by various logic devices, such as a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a single-chip microcomputer, an ARM (Acorn RISC Machine) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of these components.
[0254] The virtual fitting device in the embodiment of the present application can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The non-mobile electronic device can be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), an ATM, or an kiosks, etc., which are not specifically limited in the embodiment of the present application.
[0255] The virtual fitting device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.
[0256] The virtual fitting device provided in the embodiment of the present application can achieve Figure 9 To avoid repetition, the various implementation processes will not be described here.
[0257] It should be noted that the virtual fitting device described above can execute the virtual fitting method provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects of the execution method. For technical details not fully described in the embodiments of the boot management device, please refer to the virtual fitting method provided in the embodiments of this application.
[0258] In an embodiment of the present application, a virtual fitting device is provided, comprising: an acquisition unit for acquiring a first target human image and a first target clothing image, wherein the first target human image and the first target clothing image are both high-resolution images; a scaling unit for scaling the first target human image and the first target clothing image to generate a second target human image and a second target clothing image; a generation unit for generating a first target fused image based on the second target human image and the second target clothing image; a first target processed image based on the first target human image, wherein the first target processed image does not include clothing information and arm information; and a fusion unit for inputting the first target fused image, the first target clothing image, and the first target processed image into a first clothing model to generate a target dressing image. By acquiring the high-resolution first target human image and the second target clothing image and scaling them to obtain the second target human image and the second target clothing image to generate the first target fused image, and generating the first target processed image based on the first target human image without clothing information and arm information, and then inputting the first target fused image, the first target clothing image, and the first target processed image into the first clothing model to generate the target dressing image, the embodiment of the present application can generate a high-resolution target dressing image and improve the fusion effect of the clothing.
[0259] This application embodiment provides a server, see Figure 14 , Figure 14 A schematic diagram of the hardware structure of a server provided in an embodiment of the present application, specifically, Figure 14 As shown, the server 140 includes at least one processor 1401 and a memory 1402 ( Figure 14 (a bus connection and a processor are used as an example).
[0260] The processor 1401 is used to provide computing and control capabilities to control the server 140 to perform corresponding tasks, for example, to control the server 140 to perform the virtual fitting method in any of the above method embodiments.
[0261] It is understandable that the processor 1401 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0262] The memory 1402, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer executable programs and modules, such as the program instructions / modules corresponding to the virtual fitting method in the embodiment of the present application. The processor 1401 can implement the virtual fitting method in any of the following method embodiments by running the non-transitory software programs, instructions and modules stored in the memory 1402. Specifically, the memory 1402 may include a high-speed random access memory and may also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage device.
[0263] In some embodiments, the memory 1402 may also include a memory remotely located relative to the processor, which may be connected to the processor via a network. Examples of the aforementioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0264] In some embodiments, the server 140 may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input and output. The server 140 may also include other components for implementing device functions, which will not be described in detail here.
[0265] The present application also provides a computer-readable storage medium, such as a memory including program code, which can be executed by a processor to perform the virtual fitting method in the above embodiment. For example, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CDROM), a magnetic tape, a floppy disk, or an optical data storage device.
[0266] The present application also provides a computer program product comprising one or more program codes stored in a computer-readable storage medium. A processor of a server reads the program code from the computer-readable storage medium and executes the program code to perform the steps of the virtual fitting method provided in the above embodiment.
[0267] Those skilled in the art will understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or by hardware related to program code, and the program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a disk or an optical disk, etc.
[0268] Through the description of the above embodiments, it can be clearly understood by those skilled in the art that each embodiment can be implemented by means of software plus a general hardware platform, or of course by hardware. It can be understood by those skilled in the art that all or part of the processes in the above embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.
[0269] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Based on the idea of the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the present application as mentioned above. For the sake of simplicity, they are not provided in detail. Although the present application has been described in detail with reference to the aforementioned embodiments, ordinary technicians in this field should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A virtual fitting method, characterized in that: include: Acquire a first target human body image and a first target clothing image, wherein the first target human body image and the first target clothing image are both high-resolution images; scaling the first target human body image and the first target clothing image to generate a second target human body image and a second target clothing image; generating a first target fused image according to the second target human body image and the second target clothing image; generating a first target processed image according to the first target human body image, wherein the first target processed image does not include clothing information and arm information; inputting the first target fused image, the first target clothing image, and the first target processed image into a first clothing model to generate a target clothing image; Generating a first target fused image according to the second target human body image and the second target clothing image includes: Inputting the second target human body image and the second target clothing image into the appearance flow deformation module and the generation module to generate a first target fused image; The step of inputting the second target human body image and the second target clothing image into the appearance flow deformation module and the generation module to generate the first target fused image includes: Inputting the second target human body image and the second target clothing image into the appearance flow deformation module to generate second deformation information; deforming the second target clothing image according to the second deformation information to generate a third target clothing image; The second target human body image and the third target clothing image are input into the generating module to generate the first target fused image.
2. The method according to claim 1, characterized in that The method further includes: pre-training a first clothing model, specifically including: Acquire a first data set, wherein the first data set includes a first human body image and a first clothing image, wherein the first human body image and the first clothing image are both high-resolution images; Scaling the first human body image and the first clothing image in the first data set to generate a second data set, wherein the second data set includes the scaled first human body image and the scaled first clothing image; generating a third data set according to the scaled first human body image and the scaled first clothing image in the second data set, wherein the third data set includes the first fused image; generating a first processed image according to the first human body image, wherein the first processed image does not include clothing information and arm information; Scaling the first fused image to generate a second fused image, wherein the second fused image has the same resolution as the first human body image; The first human body image, the first processed image, and the second fused image are input to train a first clothing model.
3. The method according to claim 1, characterized in that The appearance flow deformation module includes: a first appearance flow deformation module and a second appearance flow deformation module, wherein the second appearance flow deformation module is used to generate the second deformation information. The method further includes: Training the second appearance flow deformation module includes: Acquire a first human body image and a first clothing image; Acquiring human body image analysis information according to the first human body image; Processing the human body image analysis information and the first clothing image through a first appearance flow deformation module to generate first deformation information; acquiring a second clothing image, and deforming the second clothing image according to the first deformation information to generate a third clothing image; fusing the third clothing image and the human body image analysis information to generate a first fitting image; The second appearance flow deformation module is trained according to the first clothing image and the first try-on image.
4. The method according to claim 3, characterized in that The method further comprises: During the training process of the second appearance manifold deformation module, knowledge distillation is performed on the second appearance manifold deformation module through the first appearance manifold deformation module.
5. The method according to claim 3 or 4, characterized in that The first appearance flow deformation module includes a parser-based appearance flow deformation module, and the second appearance flow deformation module includes a parser-free appearance flow deformation module.
6. A virtual fitting device, characterized in that: include: an acquiring unit, configured to acquire a first target human body image and a first target clothing image, wherein the first target human body image and the first target clothing image are both high-resolution images; a scaling unit, configured to scale the first target human body image and the first target clothing image to generate a second target human body image and a second target clothing image; a generating unit, configured to generate a first target fused image based on the second target human body image and the second target clothing image; and generate a first target processed image based on the first target human body image, wherein the first target processed image does not include clothing information and arm information; a fusion unit, configured to input the first target fused image, the first target clothing image, and the first target processed image into a first clothing model to generate a target clothing image; The generating unit is specifically configured to: Inputting the second target human body image and the second target clothing image into the appearance flow deformation module and the generation module to generate a first target fused image includes: Inputting the second target human body image and the second target clothing image into the appearance flow deformation module to generate second deformation information; deforming the second target clothing image according to the second deformation information to generate a third target clothing image; The second target human body image and the third target clothing image are input into the generating module to generate the first target fused image.
7. A server, characterized in that: include: A memory and one or more processors, wherein the one or more processors are used to execute one or more computer programs stored in the memory, and when the one or more processors execute the one or more computer programs, the server implements the method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Video virtual try-on method and device based on mixed optical flow
CN111275518A
Image processing method and device, computer equipment and storage medium
CN111339918A