Training method of virtual fitting model and related device

CN116109892BActive Publication Date: 2026-07-21SHENZHEN SHULIAN TIANXIA INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN SHULIAN TIANXIA INTELLIGENT TECH CO LTD
Filing Date
2023-01-31
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing virtual fitting technologies have high computational and learning costs when processing high-resolution images, resulting in insufficient detail in clothing textures and a lack of realism in the model's fitting experience.

Method used

A virtual try-on model, including a clothing distortion network and a try-on generation network, is used to generate high-resolution try-on effect images through feature extraction, optical flow graph processing, and loss function training.

Benefits of technology

While reducing the difficulty for the model to learn from high-resolution images, it improves the texture details of clothing in virtual try-on images and the realism of the models trying on clothes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116109892B_ABST
    Figure CN116109892B_ABST
Patent Text Reader

Abstract

The embodiment of the application relates to the technical field of image processing, and discloses a training method of a virtual fitting model and related devices, the virtual fitting model comprising a clothes warping network and a fitting generation network, the training method trains the clothes warping network based on a clothes warping network loss function, the clothes warping network loss function comprises a clothes edge loss function and a clothes deformation perception loss function, so that the clothes warping network performs feature extraction and fusion on a first image to obtain a low-resolution optical flow map, and the training method trains the fitting generation network based on a fitting generation network loss function, the fitting generation network loss function comprises a variable loss function and a perception loss function, so that the fitting generation network performs image segmentation on a second image to obtain a high-resolution fitting effect picture, so that the application can reduce the difficulty of model learning a high-resolution image, and improve the clothes texture details of a virtual fitting image and the authenticity of model fitting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a training method and related apparatus for a virtual fitting model. Background Technology

[0002] Virtual try-on is a technology that allows users to see how clothes look without having to remove them. Shoppers simply upload their photos and select the clothes they want to try on; the technology then displays the desired effect.

[0003] Existing virtual try-on technologies typically employ generative adversarial networks (GANs) to generate realistic images. However, as image size increases, the required computing power and model learning costs also rise exponentially. Therefore, considering algorithmic costs and the real-time nature of the service, existing virtual try-on algorithms are based on low-resolution applications, and the detail of clothing textures and the realism of the models wearing the clothes need improvement. Summary of the Invention

[0004] The embodiments of this application aim to provide a training method and related apparatus for a virtual fitting model, so as to reduce the difficulty of the model learning high-resolution images while improving the texture details of clothing in virtual fitting images and the realism of the model trying on clothes.

[0005] The embodiments of this application provide the following technical solutions:

[0006] In a first aspect, embodiments of this application provide a method for training a virtual fitting model, the virtual fitting model including a clothing distortion network and a fitting generation network, the method comprising:

[0007] Obtain an image dataset, which includes clothing images and model images that match the clothing images;

[0008] Each clothing image and model image undergoes data preprocessing to obtain the first image;

[0009] Based on the clothing distortion network, feature extraction and fusion are performed on the first image to obtain a low-resolution optical flow map;

[0010] Sampling operations are performed on low-resolution optical flow maps to obtain high-resolution deformable clothing;

[0011] The high-resolution images of deformed clothing and the human body's preserved regions are fused to obtain a second image;

[0012] Construct a clothing distortion network loss function, train the clothing distortion network until the clothing distortion network loss function converges, and generate the trained clothing distortion network. The clothing distortion network loss function includes a clothing edge loss function and a clothing deformation perception loss function.

[0013] Based on the virtual try-on generation network, the second image is segmented to obtain a high-resolution virtual try-on effect image;

[0014] A loss function for the virtual try-on generation network is constructed, and the virtual try-on generation network is trained until the loss function converges to generate the trained virtual try-on generation network. The loss function of the virtual try-on generation network includes a variable loss function and a perceptual loss function.

[0015] Secondly, embodiments of this application provide a method for predicting virtual fitting images, including:

[0016] Obtain the clothing image and model image to be predicted;

[0017] The clothing image to be predicted and the model image are input into the virtual fitting model to obtain a high-resolution fitting effect image corresponding to the clothing image to be predicted and the model image. The virtual fitting model is trained based on the training method of the virtual fitting model in the first aspect.

[0018] Thirdly, embodiments of this application provide an electronic device, including:

[0019] At least one processor, and

[0020] A memory that is communicatively connected to at least one processor, wherein,

[0021] The memory stores instructions that can be executed by at least one processor, which enables the at least one processor to perform either a training method for a virtual fitting model in the first aspect or a prediction method for a virtual fitting image in the second aspect.

[0022] Fourthly, embodiments of this application provide a non-volatile computer-readable storage medium storing computer-executable instructions for causing an electronic device to execute a training method for a virtual fitting model according to the first aspect or a prediction method for a virtual fitting image according to the second aspect.

[0023] The beneficial effects of this application's embodiments: Unlike existing technologies, this application provides a training method for a virtual try-on model. This virtual try-on model includes a clothing distortion network and a try-on generation network. The training method includes: acquiring an image dataset, wherein the image dataset includes clothing images and model images matching the clothing images; performing data preprocessing on each clothing image and model image to obtain a first image; performing feature extraction and fusion on the first image based on the clothing distortion network to obtain a low-resolution optical flow map; performing sampling operations on the low-resolution optical flow map to obtain a high-resolution deformed clothing; and processing the high-resolution deformed clothing and human body preservation region images. The images are fused to obtain a second image. A clothing distortion network loss function is constructed and trained until it converges to generate a trained clothing distortion network. The clothing distortion network loss function includes a clothing edge loss function and a clothing deformation perception loss function. Based on the virtual try-on generation network, the second image is segmented to obtain a high-resolution virtual try-on effect image. A virtual try-on generation network loss function is constructed and trained until it converges to generate a trained virtual try-on generation network. The virtual try-on generation network loss function includes a variable loss function and a perception loss function.

[0024] On the one hand, by using a clothing distortion network, feature extraction and fusion are performed on the first image to obtain a low-resolution optical flow map. The low-resolution optical flow map is then sampled to obtain a high-resolution deformed clothing. This application can better extract and fuse spatial features of different sizes of clothing images and model images, thereby enabling low-resolution optical flow information to guide high-resolution clothing deformation.

[0025] On the other hand, by training a clothing distortion network based on a clothing distortion network loss function, which includes a clothing edge loss function and a clothing deformation perception loss function, and by training a virtual try-on generation network based on a virtual try-on generation network loss function, a high-resolution virtual try-on effect image is obtained. This virtual try-on generation network loss function includes a variable loss function and a perception loss function. This allows the present application to improve the clothing texture details and the realism of the model trying on clothes in virtual try-on images while reducing the difficulty of the model learning high-resolution images. Attached Figure Description

[0026] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0027] Figure 1This is a schematic diagram of an application environment provided in an embodiment of this application;

[0028] Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0029] Figure 3 This is a flowchart illustrating a training method for a virtual fitting model provided in an embodiment of this application;

[0030] Figure 4 yes Figure 3 A detailed flowchart of step S302 is shown below:

[0031] Figure 5 yes Figure 3 A detailed flowchart of step S303 is shown below:

[0032] Figure 6 yes Figure 3 A detailed flowchart of step S304 is shown below:

[0033] Figure 7 This is a schematic diagram illustrating how to obtain high-resolution deformable clothing according to an embodiment of this application:

[0034] Figure 8 yes Figure 3 A detailed flowchart of step S305 is shown below:

[0035] Figure 9 yes Figure 3 A detailed flowchart of step S306 is shown below:

[0036] Figure 10 yes Figure 3 A detailed flowchart of step S308 is shown below:

[0037] Figure 11 This is a schematic diagram illustrating how to obtain a high-resolution fitting effect image according to an embodiment of this application:

[0038] Figure 12 This is a flowchart illustrating a method for predicting virtual fitting images provided in an embodiment of this application. Detailed Implementation

[0039] The present application will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present application, but do not limit the present application in any way. It should be noted that those skilled in the art can make several modifications and improvements without departing from the concept of the present application. These all fall within the protection scope of the present application.

[0040] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0041] It should be noted that, unless there is a conflict, the various features in the embodiments of this application can be combined with each other, all of which are within the protection scope of this application. Furthermore, although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than the module division in the device or the order in the flowchart. In addition, the terms "first," "second," and "third" used herein do not limit the data or execution order, but only distinguish identical or similar items with essentially the same function and effect.

[0042] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The term "and / or" as used in this specification includes any and all combinations of one or more of the associated listed items.

[0043] Furthermore, the technical features involved in the various embodiments of this application described below can be combined with each other as long as they do not conflict with each other.

[0044] Before introducing the embodiments of this application, a brief introduction will be given to the virtual try-on method known to the inventors of this application, so as to facilitate the understanding of the embodiments of this application later.

[0045] Some virtual try-on methods employ generative adversarial networks to generate realistic images. However, as image size increases, the required computing power and model learning costs also increase exponentially. Existing virtual try-on algorithms are based on low-resolution applications. To ensure real-time service and reduce algorithm costs, high-resolution try-on is a major challenge. However, with the improvement of mobile phone image quality, for 256-pixel and 1024-pixel images, the texture details of clothing displayed by low-resolution virtual try-on algorithms are drastically different from those of real clothing, resulting in a significantly different user experience.

[0046] To address the aforementioned problems, this application provides a training method for a virtual try-on model. The virtual try-on model includes a clothing distortion network and a try-on generation network. The training method includes: acquiring an image dataset, wherein the image dataset includes clothing images and model images matching the clothing images; performing data preprocessing on each clothing image and model image to obtain a first image; performing feature extraction and fusion on the first image based on the clothing distortion network to obtain a low-resolution optical flow map; performing sampling operations on the low-resolution optical flow map to obtain a high-resolution deformed clothing; and fusing the high-resolution deformed clothing and the human body preservation region image to obtain... The second image is used for image segmentation. A clothing distortion network loss function is constructed and trained until it converges to generate the trained clothing distortion network. This loss function includes a clothing edge loss function and a clothing deformation perception loss function. Based on the virtual try-on generation network, the second image is segmented to obtain a high-resolution virtual try-on effect image. A virtual try-on generation network loss function is then constructed and trained until it converges to generate the trained virtual try-on generation network. This loss function includes a variable loss function and a perception loss function.

[0047] On the one hand, by extracting and fusing features from the first image based on the clothing distortion network to obtain a low-resolution optical flow map, and then sampling the low-resolution optical flow map to obtain a high-resolution deformed clothing, this application can better extract and fuse spatial features of different sizes of clothing images and model images, thereby realizing the guidance of low-resolution optical flow information for high-resolution clothing deformation.

[0048] On the other hand, by training a clothing distortion network based on a clothing distortion network loss function, which includes a clothing edge loss function and a clothing deformation perception loss function, and by training a virtual try-on generation network based on a virtual try-on generation network loss function, a high-resolution virtual try-on effect image is obtained. This virtual try-on generation network loss function includes a variable loss function and a perception loss function. This allows the present application to improve the clothing texture details and the realism of the model trying on clothes in virtual try-on images while reducing the difficulty of the model learning high-resolution images.

[0049] In the embodiments of this application, the training method of the virtual fitting model and the prediction method of the virtual fitting image can be executed by an electronic device with computing power. The following describes an exemplary application of the electronic device provided in the embodiments of this application for training the virtual fitting model or for predicting the virtual fitting image. It can be understood that the electronic device can both train the virtual fitting model and use the virtual fitting model to predict the virtual fitting image.

[0050] In this embodiment, the electronic device can be a server, such as a server deployed in the cloud. When the server is used to train a virtual try-on model, it constructs a virtual try-on model based on image datasets, clothing distortion networks, and try-on generation networks provided by other devices or those skilled in the art. The clothing distortion network loss function and the try-on generation network loss function are used for iterative training to determine the final model parameters. When the server is used to predict virtual try-on images, it calls the built-in virtual try-on model to process the clothing images and model images to be predicted provided by other devices or the user, obtaining high-resolution try-on effect images corresponding to the clothing images and model images to be predicted.

[0051] In this embodiment, the electronic device can also be various types of terminals such as laptops, desktop computers, or mobile devices. When the terminal is used to train the virtual try-on model, those skilled in the art input a prepared image dataset into the terminal, design a clothing distortion network and a try-on generation network on the terminal, and construct loss functions for the clothing distortion network and the try-on generation network, so that the terminal iteratively trains the clothing distortion network using the clothing distortion network loss function and the try-on generation network using the try-on generation network loss function to determine the final model parameters. When the terminal is used to predict virtual try-on images, it calls the built-in virtual try-on model, processes the user-input clothing image and model image to be predicted accordingly, and obtains a high-resolution try-on effect image corresponding to the clothing image and model image to be predicted.

[0052] Before providing a detailed description of this application, the nouns and terms used in the embodiments of this application are explained, and the nouns and terms used in the embodiments of this application shall be interpreted as follows:

[0053] (1) A neural network, also known as a neural network (NNs) or connection model, is an algorithmic mathematical model that mimics the behavioral characteristics of animal neural networks to perform distributed parallel information processing. Neural networks rely on the complexity of the system to process information by adjusting the interconnections between a large number of internal nodes. Specifically, a neural network can be composed of neural units, which can be understood as a neural network with an input layer, hidden layers, and an output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the intermediate layers are hidden layers. Neural networks with many hidden layers are called deep neural networks (DNNs). The work of each layer in a neural network can be described by the mathematical expression y = a(W·x + b). From a physical perspective, the work of each layer in a neural network can be understood as transforming the input space (the set of input vectors) to the output space (i.e., from the row space to the column space of a matrix) through five operations on the input space: 1. Dimensional increase / decrease; 2. Magnification / reduction; 3. Rotation; 4. Translation; 5. "Bending". Operations 1, 2, and 3 are performed by "W·x", operation 4 by "+b", and operation 5 by "a()". The term "space" is used here because the objects being classified are not individual things, but a class of things; space refers to the set of all individuals within that class. W is the weight matrix of each layer in the neural network, where each value represents the weight of a neuron in that layer. This matrix W determines the spatial transformation from the input space to the output space, meaning that the W of each layer in the neural network controls how the space is transformed. The purpose of training the neural network is to ultimately obtain the weight matrices of all layers in the trained neural network. Therefore, the training process of a neural network is essentially learning how to control spatial transformation, more specifically, learning the weight matrix.

[0054] It should be noted that in the embodiments of this application, the models used for machine learning tasks are essentially neural networks. Common components in neural networks include convolutional layers, activation function layers, and batch normalization layers. By assembling these common components in neural networks, a model is designed. When the model parameters (weight matrices of each layer) are determined such that the model error meets a preset condition or the number of model parameters is adjusted to reach a preset threshold, the model converges.

[0055] The convolutional layer is configured with multiple convolutional kernels, each with a corresponding stride, to perform convolution operations on the image. The purpose of convolution is to extract different features from the input image. The first convolutional layer may only extract some low-level features such as edges, lines, and corners, while deeper convolutional layers can iteratively extract more complex features from low-level features.

[0056] Activation function layers are used to allow each neuron in a neural network to accept the output value of the previous layer as its input value and pass the processing result to the next layer. Commonly used activation functions include, but are not limited to, the Rectified Linear Unit (ReLU) function, the Swish function, and the Parametric Rectified Linear Unit (PReLU) function.

[0057] Batch Normalization (BN) layers are used to standardize the features of a certain layer in a network. Their purpose is to solve the problem of numerical instability in deep neural networks. Specifically, as the number of network layers increases, parameter updates during training can easily cause drastic changes in the feature outputs near the output layer, which is not conducive to training an effective neural network.

[0058] (2) A loss function is a function that maps the values ​​of a random event or its related random variables to non-negative real numbers to represent the "risk" or "loss" of that random event. The loss function is a non-negative real number function used to quantify the difference between the predicted label and the true label. In applications, the loss function is often used as a learning criterion in relation to optimization problems, i.e., solving and evaluating the model by minimizing the loss function. For example, it is used for parametric estimation in statistics and machine learning. During the training of a neural network, because we want the output of the neural network to be as close as possible to the actual predicted value, we can compare the current network's predicted value with the actual target value, and then update the weight matrix of each layer of the neural network based on the difference between the two (however, there is usually an initialization process before the first update, i.e., pre-configuring the parameters for each layer in the neural network). For example, if the network's predicted value is too high, the weight matrix is ​​adjusted to make it predict lower, and this adjustment continues until the neural network can predict the actual target value. Therefore, it is necessary to predefine "how to compare the difference between the predicted value and the target value," which is the loss function or objective function. These are important equations used to measure the difference between the predicted value and the target value. Taking the loss function as an example, a higher output value (loss) of the loss function indicates a greater difference, so training the neural network becomes the process of minimizing this loss as much as possible.

[0059] The technical solution of this application will be described in detail below with reference to the accompanying drawings.

[0060] Please see Figure 1 , Figure 1This is a schematic diagram of an application environment provided in an embodiment of this application;

[0061] like Figure 1 As shown, the application environment 100 includes an electronic device 10 and a server 20. The electronic device 10 is connected to the server 20 via network communication, wherein the network includes wired networks and / or wireless networks. It is understood that the network includes wireless networks such as 2G, 3G, 4G, 5G, wireless LAN, and Bluetooth, and may also include wired networks such as serial cables and network cables.

[0062] In this embodiment, the electronic device 10 is communicatively connected to the server 20 and is used to acquire image datasets, construct clothing distortion networks and fitting room generation networks. For example, those skilled in the art can download multiple clothing images and model images matching the clothing images to the electronic device, and construct clothing distortion networks and fitting room generation networks. It is understood that the electronic device 10 can also be used to acquire clothing images and model images to be predicted. For example, a user inputs clothing images and model images to be predicted through an input interface; after input, the electronic device automatically acquires the clothing images and model images to be predicted. Alternatively, the electronic device 10 may have a camera to capture clothing images and model images, or the electronic device 10 may store a clothing image library and a model image library, from which the user can select the clothing images and model images to be predicted. The electronic device in this embodiment includes, but is not limited to, various terminals with computing capabilities such as laptops, desktop computers, or mobile devices. Preferably, the electronic device is a smartphone.

[0063] In this embodiment, the server 20 is communicatively connected to the electronic device 10 and is used to train a virtual fitting model. Alternatively, it can acquire the clothing image and model image to be predicted input by the user on the electronic device 10, call the built-in virtual fitting model to process the clothing image and model image to be predicted, obtain a high-resolution fitting effect image corresponding to the clothing image and model image, and then send the high-resolution fitting effect image to the electronic device 10. The number of servers 20 can also be multiple, and multiple servers can form a server cluster. For example, the server cluster includes: a first server, a second server, ..., an Nth server; or, the server cluster can be a cloud computing service center, which includes several servers. The servers in this embodiment include, but are not limited to: tower servers, rack servers, blade servers, and cloud servers. Preferably, the server is a cloud server (Elastic Compute Service, ECS).

[0064] It is understood that, in this embodiment of the application, the electronic device 10 is also used to display the high-resolution virtual try-on effect image on its own display interface after receiving the high-resolution virtual try-on effect image sent by the server, so as to inform the user, or to locally execute the training method of the virtual try-on model or the prediction method of the virtual try-on image provided in this embodiment of the application.

[0065] Example 1

[0066] Please see Figure 2 , Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0067] like Figure 2 As shown, the electronic device 200 includes one or more processors 201 and a memory 202. Wherein, Figure 2 Take a processor 201 as an example.

[0068] The processor 201 and the memory 202 can be connected via a bus or other means. Figure 2 Taking the example of a connection between China and Israel via a bus.

[0069] Processor 201 is configured to provide computational and control capabilities to control electronic device 200 to perform corresponding tasks, such as controlling electronic device 200 to perform a training method for a virtual fitting model in any of the following method embodiments, including: acquiring an image dataset, wherein the image dataset includes clothing images and model images matching the clothing images; performing data preprocessing on each clothing image and model image to obtain a first image; performing feature extraction and fusion on the first image based on a clothing distortion network to obtain a low-resolution optical flow map; performing a sampling operation on the low-resolution optical flow map to obtain a high-resolution deformed clothing; and processing the high-resolution deformed clothing and human body preservation region images. The images are fused to obtain a second image. A clothing distortion network loss function is constructed and trained until it converges to generate a trained clothing distortion network. The clothing distortion network loss function includes a clothing edge loss function and a clothing deformation perception loss function. Based on the virtual try-on generation network, the second image is segmented to obtain a high-resolution virtual try-on effect image. A virtual try-on generation network loss function is constructed and trained until it converges to generate a trained virtual try-on generation network. The virtual try-on generation network loss function includes a variable loss function and a perception loss function.

[0070] Alternatively, the control electronic device 200 executes the virtual fitting image prediction method in any of the following method embodiments, including: acquiring a clothing image to be predicted and a model image; inputting the clothing image to be predicted and the model image into a virtual fitting model to obtain a high-resolution fitting effect image corresponding to the clothing image to be predicted and the model image, wherein the virtual fitting model is trained based on the training method of the virtual fitting model in any of the following embodiments.

[0071] Processor 201 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The aforementioned PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0072] The memory 202, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the virtual fitting model training method or the virtual fitting image prediction method in the embodiments of this application. The processor 201, by running the non-transitory software programs, instructions, and modules stored in the memory 202, can implement the virtual fitting model training method or the virtual fitting image prediction method in any of the following method embodiments. Specifically, the memory 202 may include volatile memory (VM), such as random access memory (RAM); the memory 202 may also include non-volatile memory (NVM), such as read-only memory (ROM), flash memory, hard disk drive (HDD), solid-state drive (SSD), or other non-transitory solid-state storage devices; the memory 202 may also include combinations of the above types of memory.

[0073] Memory 202 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 202 may optionally include memory remotely located relative to processor 201, and these remote memories may be connected to processor 201 via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0074] One or more modules are stored in memory 202. When executed by one or more processors 201, they perform the training method for the virtual try-on model or the prediction method for the virtual try-on image in any of the following method embodiments. For example, they perform the following... Figure 3 The steps shown.

[0075] In the embodiments of this application, the electronic device 200 may also have wired or wireless network interfaces, input / output interfaces and other components to perform input and output. The electronic device 200 may also include other components for implementing device functions, which will not be described in detail here.

[0076] The training method of the virtual fitting model provided in this application embodiment is described below with reference to exemplary applications and implementations of the electronic device provided in the embodiments of this application.

[0077] Please see Figure 3 , Figure 3 This is a flowchart illustrating a training method for a virtual fitting model provided in an embodiment of this application;

[0078] The training method for the virtual try-on model is applied to electronic devices, such as terminals and servers. Specifically, the execution subject of the training method for the virtual try-on model is one or at least two processors in the electronic device. The following uses a server as an example to illustrate the training method for the virtual try-on model.

[0079] Specifically, the virtual try-on model includes a clothing distortion network and a try-on generation network.

[0080] like Figure 3 As shown, the training method for this virtual try-on model includes:

[0081] Step S301: Obtain the image dataset;

[0082] Specifically, the image dataset includes clothing images and model images that match the clothing images. The clothing images are paired one-to-one with the model images. The clothing images are required to be flat lay images with the front facing up, and the model images are required to be images taken in a frontal pose. The image size of both the clothing images and the model images must be no less than 1024*768.

[0083] Specifically, an image dataset is formed by downloading multiple clothing images and matching model images from the Internet using an electronic device. The number of clothing images and model images in the image dataset can be in the tens of thousands, for example, 20,000, or the number of clothing images and model images can be determined by those skilled in the art based on the actual situation.

[0084] Step S302: Perform data preprocessing on each clothing image and model image to obtain the first image;

[0085] Please see Figure 4 , Figure 4 yes Figure 3 A detailed flowchart of step S302 is shown below:

[0086] like Figure 4 As shown, step S302: Perform data preprocessing on each clothing image and model image to obtain a first image, including:

[0087] Step S3021: Extract the outer contour of each clothing image to obtain the edge image corresponding to each clothing image;

[0088] Specifically, the outer contour of each clothing image is extracted using an image segmentation network to obtain the edge image corresponding to each clothing image. The edge image is a black and white image, where the white area in the edge image represents the area to be retained for clothing, and the black area in the edge image represents the background area other than clothing.

[0089] Specifically, image segmentation networks include U 2 -NET network, U 2 The U-Net network is a two-layer nested U-shaped structure. It designs a new U-shaped residual block (RSU) structure at the bottom layer and a U-Net-like structure at the top layer. Each level is filled with RSU. The RSU replaces the ordinary single-stream convolution with U-Net and replaces the original features with local features composed of a weight layer. This design change enables the network to extract features directly from multiple scales of each residual block.

[0090] Specifically, U 2 The -Net network consists of three parts: a six-level encoder, a five-level decoder, and a saliency map fusion module. The saliency map fusion module is connected to the decoder and the last-level encoder. The six-level encoder includes: En_1, En_2, En_3, En_4, En_5, and En_6, and the five-level decoder includes: De_1, De_2, De_3, De_4, and De_5.

[0091] In encoders En_1, En_2, En_3, and En_4, RSU structures of RSU-7, RSU-6, RSU-5, and RSU-4 are used, respectively. The preceding numbers, such as 7, 6, 5, and 4, represent the height L of the RSU. L is typically configured based on the spatial resolution of the input feature maps. In En_5 and En_6, the feature maps have relatively low resolution, and further downsampling of these feature maps would lead to the loss of useful context. Therefore, in the RSU-5 and RSU-6 stages, RSU-4F is used, where F indicates that RSU is an extended version, replacing the merging and upsampling operations with extended convolutions. This means that all intermediate feature maps in RSU-4F have the same resolution as their input feature maps.

[0092] The decoding stage has a structure similar to the symmetric encoding stage in En_6. In De_5, the RSU-4F extended board is also used, similar to that used in the encoding stages of En_5 and En_6. Each decoder stage takes a concatenation of the upsampled feature map from the previous stage and the feature map from its symmetric encoder stage as input.

[0093] The saliency map fusion module is used to generate a saliency probability map. 2 The -NET network first generates six output saliency probability maps from En_6, De_5, De_4, De_3, De_2, and De_1 using 3x3 convolutions and a sigmoid function. Then, the logistic graph of the output saliency map (the convolution output, before the sigmoid function) is upsampled to the same size as the input image and fused through a cascade operation. Finally, it passes through a 1x1 convolutional layer and a sigmoid function to generate the final saliency probability map, which is the edge image corresponding to the clothing image.

[0094] Step S3022: Perform background subtraction on each clothing image to obtain a high-resolution clothing image corresponding to each clothing image;

[0095] Specifically, the image segmentation network is also used to perform background subtraction on each clothing image to obtain a high-resolution clothing image corresponding to each clothing image. The high-resolution clothing image is a white-background clothing image obtained by removing useless background information from the clothing image.

[0096] In this embodiment of the application, by adopting U 2 As an image segmentation network, the -NET network in this application can fuse features from receptive fields of different scales, capture more contextual information, and increase the depth of the entire architecture without significantly increasing computational costs.

[0097] Step S3023: Scale each high-resolution clothing image using an image scaling function to obtain a low-resolution clothing image corresponding to each high-resolution clothing image;

[0098] Specifically, the image resizing function includes the resize(src,dsize,dst,fx,fy,interpolation) function from the open-source computer vision library (opencv) based on the computer programming language (Python), which scales down a high-resolution clothing image proportionally to a low-resolution clothing image.

[0099] In the image scaling function, src represents the original image, dsize represents the required size of the output image, fx represents the scaling factor along the horizontal axis, fy represents the scaling factor along the vertical axis, and interpolation represents the interpolation method, which includes, but is not limited to, nearest neighbor interpolation, bilinear interpolation, etc.

[0100] Step S3024: Detect each model image using a human pose evaluation algorithm to obtain the human key point region map corresponding to each model image;

[0101] Specifically, human pose estimation algorithms include, but are not limited to, the human skeletal keypoint detection (OpenPose) algorithm. Preferably, the OpenPose algorithm is used in the embodiments of this application.

[0102] Specifically, the OpenPose algorithm is an algorithm that relies on convolutional neural networks and supervised learning to achieve human pose assessment. Its main advantage is that it is applicable to open-source real-time systems for multi-person 2D pose detection and can accurately and quickly identify human key points.

[0103] Specifically, the OpenPose algorithm takes the entire image of a person as input to the network, then predicts a confidence map for body part detection and a part affinity field (PAF) for part association. After that, it performs a set of binary matching on the associated body part candidates through a parsing step, and finally builds a human skeleton to connect human key points and assembles them into the complete pose of all people in the image, thus obtaining the human key point region map corresponding to each model image.

[0104] The human body key points are divided into 25 categories, each labeled with 0-24. The categories are: {0, "Nose"}, {1, "Neck"}, {2, "Right Upper Arm"}, {3, "Right Elbow"}, {4, "Right Wrist"}, {5, "Left Upper Arm"}, {6, "Left Elbow"}, {7, "Left Wrist"}, {8, "Mid Hip"}, {9, "Right Hip"}, {10, "Right Knee"}, {11, "Right Ankle ...12, "Right Upper Arm"}, {13, "Right Upper Arm"}, {14, "Right Upper Arm"}, {15, "Right Upper Arm"}, {16, "Right Upper Arm"}, {17, "Right Upper Arm"}, {18, "Right Upper Arm"}, {19, "Right Upper Arm"}, {12, "Left hip (Lhip)"}, {13, "Left knee (LKnee)"}, {14, "Left ankle (LAnkle)"}, {15, "Right eye (Reye)"}, {16, "Left eye (Leye)"}, {17, "Right ear (Rear)"}, {18, "Left ear (Lear)"}, {19, "Left big toe (LBigToe)"}, {20, "Left little toe (LSmallToe)"}, {21, "Left heel (LHeel)"}, {22, "Right big toe (RBigToe)"}, {23, "Right little toe (RSmallToe)"}, {24, "Right heel (Rheel)"}.

[0105] Step S3025: Extract the human body image of each model using a human body analysis algorithm to obtain the human body analysis map corresponding to each model image;

[0106] Specifically, human body parsing algorithms include, but are not limited to, the general human body parsing (Graphonomy) algorithm performed through graph transfer learning. Preferably, the Graphiconomy algorithm is used in the embodiments of this application.

[0107] Specifically, the Graphicality algorithm first learns and propagates a compact high-level semantic graph representation in one dataset through intra-graph inference. Then, through inter-graph transfer driven by an explicit hierarchical semantic label structure, it transfers and fuses semantic information across multiple datasets to obtain a human anatomy map corresponding to each model image. By extracting a general semantic graph representation into each specific task, the Graphicality algorithm can predict labels at all levels within a system without increasing complexity.

[0108] Step S3026: The edge image, low-resolution clothing image, human body key point region map, and human body analytical map are stitched together using a stitching function to obtain the first image.

[0109] Specifically, the concatenation function includes the `torch.cat(input, concatenation dimension)` function from the scientific computing package (pytorch) based on the Python programming language. The `torch.cat` function can perform concatenation operations on the input tensor sequence along a given dimension. The parameters in parentheses are: input (input) represents the tensor sequence to be concatenated, and concatenation dimension (dim) represents the selected expansion dimension value, which allows concatenation of tensor sequences along this dimension.

[0110] Furthermore, by using the edge image, low-resolution clothing image, human body key point region map, and human body analytical map as inputs to the stitching function and setting the stitching dimension to 1, the first output image can be obtained.

[0111] Step S303: Based on the clothing distortion network, perform feature extraction and fusion on the first image to obtain a low-resolution optical flow map;

[0112] Specifically, the clothing distortion network is a convolutional neural network based on Feature Pyramid Networks (FPN), taking the first image as an input image and a low-resolution optical flow map as the output. The low-resolution optical flow map is a graph similar to optical flow used for clothing deformation. Each point in the graph is a two-dimensional vector, recording which point in the original clothing should be sampled from for each point in the deformed clothing.

[0113] The FPN network structure includes a bottom-up path, a top-down path, and lateral connections. The bottom-up process is the normal forward propagation process of a neural network. The top-down process is to upsample the more abstract and semantically stronger high-level feature maps. The lateral connections are to fuse the upsampled results with the feature maps of the same size generated from the bottom-up process.

[0114] Specifically, the clothing distortion network includes an encoding module and a decoding module. The encoding module includes convolutional layers, activation function layers, and batch normalization layers, while the decoding module includes transposed convolutional layers, activation function layers, and batch normalization layers.

[0115] Please see Figure 5 , Figure 5 yes Figure 3 A detailed flowchart of step S303 is shown below:

[0116] like Figure 5 As shown, step S303: Based on the clothing distortion network, feature extraction and fusion are performed on the first image to obtain a low-resolution optical flow map, including:

[0117] Step S3031: Based on the encoding module, perform feature extraction and fusion on the first image to obtain three feature maps of different sizes;

[0118] Specifically, the mathematical expression formula corresponding to the encoding module of the clothing distortion network is as follows:

[0119]

[0120] in, Let i represent the i-th feature image of the (x+1)-th layer, and BN denotes batch normalization. This represents the ReLU activation function. This represents the i-th feature image of the x-th layer. This represents the standard convolutional kernel of the (x+1)th layer. This indicates the bias term.

[0121] Specifically, the coding modules of the clothing distortion network include: a first convolutional layer (with 32 standard convolutional kernels), a first activation function layer, a first batch normalization layer, a second convolutional layer (with 64 standard convolutional kernels), a second activation function layer, a second batch normalization layer, a third convolutional layer (with 128 standard convolutional kernels), a third activation function layer, a third batch normalization layer, a fourth convolutional layer (with 256 standard convolutional kernels), a fourth activation function layer, a fourth batch normalization layer, a fifth convolutional layer (with 512 standard convolutional kernels), a fifth activation function layer, and a fifth batch normalization layer.

[0122] In this module, each convolutional layer has a standard convolutional kernel size of 3*3 and a stride of 2. The activation function layers use ReLU activation functions. It can be understood that by extracting and fusing features from the first image through the encoding module, the resulting three feature maps of different sizes are 8*6*512, 16*12*256, and 32*24*128, respectively.

[0123] Step S3032: Based on the decoding module, decode the three feature maps of different sizes output by the encoding module to obtain a low-resolution optical flow map.

[0124] Specifically, the decoding module is a convolutional neural network that takes three feature maps of different sizes output by the encoding module as input and a low-resolution optical flow map of size 256*192*2 as output. It includes transposed convolutional layers, activation function layers, and batch normalization layers.

[0125] Specifically, the decoding module includes a first transposed convolutional layer (512 kernels), a first activation function layer, a first batch normalization layer, a second transposed convolutional layer (256 kernels), a second activation function layer, a second batch normalization layer, a third transposed convolutional layer (128 kernels), a third activation function layer, a third batch normalization layer, a fourth transposed convolutional layer (64 kernels), a fourth activation function layer, a fourth batch normalization layer, a fifth transposed convolutional layer (32 kernels), a fifth activation function layer, a fifth batch normalization layer, a sixth transposed convolutional layer (2 kernels), a sixth activation function layer, and a sixth batch normalization layer. Each convolutional layer has a 3x3 kernel size and a stride of 2. The activation function used in the activation function layers includes the PReLU activation function.

[0126] In this embodiment of the application, by extracting and fusing features from the first image based on the clothing distortion network to obtain a low-resolution optical flow map, this application can better extract and fuse spatial features of different sizes of clothing images and model images.

[0127] Step S304: Sampling operation is performed on the low-resolution optical flow map to obtain high-resolution deformable clothing;

[0128] Please see Figure 6 , Figure 6 yes Figure 3 A detailed flowchart of step S304 is shown below:

[0129] like Figure 6 As shown, step S304: Sampling the low-resolution optical flow map to obtain a high-resolution deformable garment, including:

[0130] Step S3041: Upsample the low-resolution optical flow map using an array sampling operation function to obtain a high-resolution optical flow map;

[0131] Specifically, the array sampling operation function includes the `F.interpolate` function, which is used to upsample a low-resolution optical flow map to obtain a high-resolution optical flow map. Here, `input` represents the input, `scale_factor` represents the upsampling factor, and `mode` represents the upsampling algorithm. The upsampling algorithm includes, but is not limited to, the nearest neighbor algorithm and the bilinear algorithm. In this embodiment, the input is a low-resolution optical flow map, the upsampling factor is 2, and the upsampling algorithm is the bilinear algorithm.

[0132] Understandably, the F.interpolate function performs upsampling / downsampling operations on the input tensor array through interpolation, that is, it scientifically and reasonably changes the size of the array in order to maintain the integrity of the data as much as possible.

[0133] Step S3042: Based on the two-dimensional feature coordinate information of the high-resolution optical flow map, the high-resolution clothing is distorted and deformed using a bilinear interpolation function to obtain high-resolution deformed clothing.

[0134] Specifically, the bilinear interpolation function includes the `F.grid_sample` function, which is used to distort and deform high-resolution clothing to obtain high-resolution deformed clothing. Here, `input` represents the input, `grid` represents the optical flow map, containing the grid size of the output feature map and the sampling points of each grid corresponding to the input feature map, `mode` represents the upsampling algorithm, including but not limited to nearest neighbor algorithms, bilinear algorithms, etc., and `padding_mode` represents the padding mode, including zero padding, boundary padding, and symmetric padding. In this embodiment, the input is high-resolution clothing, the optical flow map is a high-resolution optical flow map, the upsampling algorithm is a bilinear algorithm, and the padding mode is boundary padding.

[0135] Understandably, the F.grid_sample function uses the two-dimensional feature coordinate information of the high-resolution optical flow map to guide the high-resolution clothing to undergo distortion and deformation, resulting in high-resolution deformed clothing.

[0136] In this embodiment of the application, by sampling the low-resolution optical flow map, a high-resolution deformable garment is obtained. This application can realize the deformation of high-resolution garment guided by low-resolution optical flow information.

[0137] Please see Figure 7 , Figure 7 This is a schematic diagram illustrating how to obtain high-resolution deformable clothing according to an embodiment of this application:

[0138] like Figure 7 As shown, the edge image, low-resolution clothing image, human body key point region map, and human body analytical map are stitched together and used as input to the clothing distortion network. The low-resolution optical flow map is used as output to the clothing distortion network. After upsampling, the low-resolution optical flow map generates a high-resolution optical flow map. The high-resolution clothing is distorted and deformed according to the high-resolution optical flow map to obtain high-resolution deformed clothing.

[0139] In this embodiment of the application, the method further includes:

[0140] The background is removed from each model image based on the image segmentation network to obtain the white background model image corresponding to each model image.

[0141] The human body analysis image and the white background model image are multiplied by a matrix to obtain the human body preserved region image.

[0142] Specifically, the image segmentation network is also used to perform background subtraction on each model image to obtain a white-background model image corresponding to each model image. The image segmentation network in this step has the same structure as the image segmentation network used in step S3022, and the background subtraction on each model image based on the image segmentation network to obtain a white-background model image corresponding to each model image is similar to the specific implementation method of step S3022: performing background subtraction on each clothing image to obtain a high-resolution clothing image corresponding to each clothing image, which will not be described in detail here.

[0143] Furthermore, a matrix multiplication is performed between the anatomy analysis image and the white-background model image to obtain the human body preservation image. In this matrix, the black areas in the anatomy analysis image correspond to a value of zero, while the non-black areas correspond to a value of 1. This is because, since 0 multiplied by any number equals 0, and 1 multiplied by any number equals itself, the matrix multiplication between the anatomy analysis image and the white-background model image allows for the removal of the model's clothing and torso, resulting in the human body preservation image.

[0144] Step S305: Fuse the high-resolution deformed clothing and human body preservation area images to obtain a second image;

[0145] For details, please refer to Figure 8 , Figure 8 yes Figure 3 A detailed flowchart of step S305 is shown below:

[0146] like Figure 8 As shown, step S305: fusing the high-resolution deformed clothing and human body preservation region images to obtain a second image, including:

[0147] Step S3051: Add the high-resolution deformed clothing and human body preserved area images by matrix addition to obtain the second image.

[0148] Specifically, both the matrix of the high-resolution deformable clothing and the matrix of the human body preservation region image are RGB three-channel matrices, where R (red) represents the red channel, G (green) represents the green channel, and B (blue) represents the blue channel. The values ​​of each channel in the high-resolution deformable clothing matrix and the human body preservation region image matrix are added together to obtain the second image. In this second image, the black areas correspond to a value of 0 in the matrix. It can be understood that the black areas in both the high-resolution deformable clothing and the human body preservation region image have RGB values ​​of (0, 0, 0) in the matrix, and the sum remains (0, 0, 0).

[0149] Step S306: Construct the clothing distortion network loss function, train the clothing distortion network until the clothing distortion network loss function converges, so as to generate the trained clothing distortion network;

[0150] Specifically, the clothing distortion network loss function is used to train the clothing distortion network, adjust the parameters of the clothing distortion network until the clothing distortion network loss function converges, and the clothing distortion network at this point is considered the trained clothing distortion network.

[0151] Specifically, the loss function for the clothing distortion network includes the clothing edge loss function and the clothing deformation perception loss function, expressed by the following formula:

[0152]

[0153] Where Loss1 represents the loss function of the clothing distortion network, L edge Let λ represent the loss function for clothing edges, and λ1 represent the first hyperparameter, which can be set to 0.05. This represents the loss function for perceiving clothing deformation.

[0154] The clothing edge loss function includes:

[0155]

[0156] Among them, L edge This represents the loss function for clothing edges, where edge represents the true contour sample. Represents the predicted contour sample. This represents the absolute difference between the true contour sample and the predicted contour sample.

[0157] Understandably, real contour samples include high-resolution deformed clothing corresponding to clothing images and model images in the image dataset, while predicted contour samples include high-resolution deformed clothing corresponding to clothing images predicted by the clothing distortion network and model images.

[0158] The clothing deformation perception loss function includes:

[0159]

[0160] in, Let M represent the loss function for clothing deformation perception, and let N represent the number of layers in the VGG network. i C represents the number of feature elements in the i-th layer of the VGG network. warp_gt Representing a true high-resolution sample of deformable clothing, C warp_high VGG represents the predicted high-resolution deformable clothing sample. i (C warp_gt ) represents the feature image of a real high-resolution deformable clothing sample in the i-th layer of the VGG network. i (C warp_high Let represent the feature image of the predicted high-resolution deformable clothing sample in the i-th layer of the VGG network, ||VGG i (C warp_gt -VGG i (C warp_high )|| L1 This represents the absolute difference between the feature image of the real high-resolution deformable clothing sample in the i-th layer of the VGG network and the feature image of the predicted high-resolution deformable clothing sample in the i-th layer of the VGG network.

[0161] The VGG network includes the VGG19 network, a convolutional neural network with 19 hidden layers (16 convolutional layers and 3 fully connected layers). The VGG19 network is used to extract image features and can reflect the difference between predicted and real samples under the same VGG19 feature extraction, i.e., perceptual loss. Furthermore, it increases the network depth while maintaining the same receptive field, thus improving the performance of the neural network to some extent. It can be understood that real high-resolution deformed clothing samples include high-resolution deformed clothing corresponding to clothing images and model images in the image dataset, while predicted high-resolution deformed clothing samples include high-resolution deformed clothing corresponding to clothing images predicted by the clothing distortion network and model images.

[0162] For details, please refer to [link / reference]. Figure 9 , Figure 9 yes Figure 3 A detailed flowchart of step S306 is shown below:

[0163] In this embodiment, the Adam algorithm (Adaptive Moment Estimation Algorithm) is used to optimize the parameters of the clothing distortion network. For example, the upper limit of the number of iterations is set to 10,000, the initial learning rate is set to 0.0005, and the weight decay is set to 0.0005. After obtaining the adjusted clothing distortion network parameters output by the Adam algorithm, these parameters are used for the next training iteration until the clothing distortion network loss function converges. It can be understood that the convergence of the clothing distortion network loss function refers to the convergence of the sum of the clothing edge loss function and the clothing deformation perception loss function.

[0164] Understandably, the Adam algorithm (Adaptive Moment Estimation Algorithm) can be seen as a combination of the momentum method and the RMSprop algorithm. It not only uses momentum as a parameter to update the direction, but also can adaptively adjust the learning rate.

[0165] In this embodiment, while optimizing the parameters of the clothing distortion network using the Adam algorithm, the loss function value curve of the clothing distortion network is monitored in real time. If the loss function value curve of the clothing distortion network does not decrease significantly within a preset number of iterations, the loss function of the clothing distortion network converges, indicating that the training of the clothing distortion network is complete. The parameters of the clothing distortion network at this time are then output, thereby obtaining the trained clothing distortion network.

[0166] like Figure 9 As shown, step S306: Constructing a clothing distortion network loss function, training the clothing distortion network until the loss function converges to generate the trained clothing distortion network, including:

[0167] Step S3061: Save the loss function value of the clothing distortion network for each training session;

[0168] Specifically, the loss function value of the clothing distortion network includes the sum of the clothing edge loss function value and the clothing deformation perception loss function value.

[0169] Step S3062: When the number of iterations is greater than the first preset number of iterations, sort the currently trained clothing distortion network loss function value and the saved clothing distortion network loss function value in ascending order;

[0170] Specifically, the first preset number of iterations can be 100. When the number of iterations is greater than 100, the loss function value of the clothing distortion network obtained by training is sorted in ascending order with the lost function value of the clothing distortion network that has been saved, so as to obtain the ranking of the loss function value of the clothing distortion network obtained by training.

[0171] Step S3063: If the loss function value of the clothing distortion network obtained by training is greater than the first preset sequence value, then the clothing distortion network loss function converges.

[0172] Specifically, the first preset sequence value can be 100. If the ranking of the clothing distortion network loss function value obtained from the current training is greater than 100, it indicates that the clothing distortion network loss function value curve has not decreased significantly within 100 iterations, and the clothing distortion network loss function has converged.

[0173] Furthermore, if the loss function of the clothing distortion network converges, it signifies that the training of the clothing distortion network is complete. The parameters of the clothing distortion network at this point are then output, thus obtaining the trained clothing distortion network.

[0174] In the embodiments of this application, in the first stage of model training, the clothing distortion network is trained based on the clothing distortion network loss function, which includes the clothing edge loss function and the clothing deformation perception loss function. This application can better enable the clothing distortion network to learn the realism of clothing try-on, thereby improving the deformation effect of high-resolution deformable clothing.

[0175] Step S307: Based on the virtual try-on generation network, perform image segmentation on the second image to obtain a high-resolution virtual try-on effect image;

[0176] Specifically, the virtual try-on generation network is a convolutional neural network that takes a second image as input and a high-resolution virtual try-on effect image as output. The network includes an encoding module and a decoding module. The encoding module consists of convolutional layers, activation function layers, and batch normalization layers, while the decoding module consists of transposed convolutional layers, activation function layers, and batch normalization layers.

[0177] Specifically, the encoding modules of the virtual fitting generation network include: a first convolutional layer (32 standard convolutional kernels), a first activation function layer, a first batch normalization layer, a second convolutional layer (64 standard convolutional kernels), a second activation function layer, a second batch normalization layer, a third convolutional layer (128 standard convolutional kernels), a third activation function layer, a third batch normalization layer, a fourth convolutional layer (256 standard convolutional kernels), a fourth activation function layer, a fourth batch normalization layer, a fifth convolutional layer (512 standard convolutional kernels), a fifth activation function layer, a fifth batch normalization layer, a sixth convolutional layer (1024 standard convolutional kernels), a sixth activation function layer, and a sixth batch normalization layer.

[0178] In this network, each convolutional layer has a standard kernel size of 3*3 and a stride of 2. The activation function layers use ReLU activation functions. It can be understood that by segmenting the second image through the encoding module of the network, a feature map with a size of 16*16*1024 can be obtained.

[0179] Furthermore, the decoding module of the virtual try-on generation network is a convolutional neural network that takes a 16*16*1024 feature map output from the encoding module as input and a high-resolution virtual try-on image as output. It includes transposed convolutional layers, activation function layers, and batch normalization layers. The high-resolution virtual try-on image is a three-channel RGB image with a size of 1024*768.

[0180] Specifically, the decoding module of the virtual fitting generation network includes a first transposed convolutional layer (1024 kernels), a first activation function layer, a first batch normalization layer, a second transposed convolutional layer (512 kernels), a second activation function layer, a second batch normalization layer, a third transposed convolutional layer (256 kernels), a third activation function layer, a third batch normalization layer, a fourth transposed convolutional layer (128 kernels), a fourth activation function layer, a fourth batch normalization layer, a fifth transposed convolutional layer (64 kernels), a fifth activation function layer, a fifth batch normalization layer, a sixth transposed convolutional layer (3 kernels), a sixth activation function layer, and a sixth batch normalization layer. Each convolutional layer has a 3x3 kernel size and a stride of 2. The activation function used in the activation function layers includes the PReLU activation function.

[0181] Step S308: Construct the loss function for the virtual try-on generation network, train the virtual try-on generation network until the loss function converges, and generate the trained virtual try-on generation network.

[0182] Specifically, it is used to train the virtual try-on network, adjust the parameters of the virtual try-on network until the loss function of the virtual try-on network converges, and then take the virtual try-on network at this point as the trained virtual try-on network.

[0183] Specifically, the loss function of the virtual try-on generation network includes a variable loss function and a perceptual loss function, expressed by the following formula:

[0184]

[0185] Where Loss2 represents the loss function of the virtual try-on generation network, L con Let represent the variable loss function, and λ2 represent the second hyperparameter, which can be set to 0.1. This represents the perceptual loss function.

[0186] The variable loss function includes:

[0187] L con =||P gt -P1|| L1 +λ3||P gt -P1|| L2

[0188] Among them, L con P represents the variable loss function. gt Let P represent the real sample, P1 represent the predicted sample, and λ3 represent the third hyperparameter, which can be set to 0.005. gt -P1|| L1 ||P represents the absolute difference between the actual sample and the predicted sample. gt -P1|| L2 This represents the squared error between the actual sample and the predicted sample.

[0189] Understandably, real samples include white-background model images corresponding to model images in the image dataset, while predicted samples include clothing images predicted by the virtual fitting generation network and high-resolution virtual fitting effect images corresponding to model images.

[0190] The perceptual loss function includes:

[0191]

[0192] in, Let M represent the perceptual loss function, M represent the number of layers in the VGG network, Ni represent the number of feature elements in the i-th layer of the VGG network, and P represent the perceptual loss function. gt VGG represents the real sample, P1 represents the predicted sample, and VGG represents the predicted sample. i (P gt ) represents the feature image of a real sample in the i-th layer of the VGG network, VGG i (P1) represents the feature image of the predicted sample in the i-th layer of the VGG network, ||VGG i (P gt -VGG i (P1)|| L1 This represents the absolute difference between the feature image of the real sample in the i-th layer of the VGG network and the feature image of the predicted sample in the i-th layer of the VGG network.

[0193] The VGG network includes the VGG19 network, a convolutional neural network with 19 hidden layers (16 convolutional layers and 3 fully connected layers). The VGG19 network is used to extract image features and can reflect the difference between predicted and real samples under the same VGG19 feature extraction, i.e., perceptual loss. Furthermore, it increases the network depth while maintaining the same receptive field, thus improving the performance of the neural network to some extent. Understandably, real samples include white-background model images corresponding to model images in the image dataset, while predicted samples include clothing images predicted by the virtual try-on generation network and high-resolution virtual try-on images corresponding to the model images.

[0194] In this embodiment, during the second stage of model training, the virtual try-on generation network is trained based on the loss function of the virtual try-on generation network to obtain a high-resolution virtual try-on effect image. The loss function of the virtual try-on generation network includes a variable loss function and a perceptual loss function. This application can better enable the virtual try-on generation network to learn the realism of the high-resolution virtual try-on effect image and improve the quality of the generated high-resolution virtual try-on effect image.

[0195] Please see Figure 10 , Figure 10 yes Figure 3 A detailed flowchart of step S308 is shown below:

[0196] In this embodiment, the Adam algorithm (Adaptive Moment Estimation Algorithm) is used to optimize the parameters of the virtual try-on generation network. For example, the upper limit of the number of iterations is set to 100,000, the initial learning rate is set to 0.001, and the weight decay is set to 0.0008. After obtaining the adjusted virtual try-on generation network parameters output by the Adam algorithm, these parameters are used for the next training iteration until the loss function of the virtual try-on generation network converges. It can be understood that the convergence of the loss function of the virtual try-on generation network refers to the convergence of the sum of the variable loss function and the perceptual loss function.

[0197] In this embodiment, while optimizing the parameters of the virtual try-on generation network using the Adam algorithm, the loss function curve of the virtual try-on generation network is monitored in real time. If the loss function curve of the virtual try-on generation network does not decrease significantly within a preset number of iterations, the loss function of the virtual try-on generation network converges, indicating that the training of the virtual try-on generation network is complete. The parameters of the virtual try-on generation network at this time are then output, thereby obtaining the trained virtual try-on generation network.

[0198] like Figure 10 As shown, step S308: Constructing the loss function for the virtual try-on generation network, training the virtual try-on generation network until the loss function converges, to generate the trained virtual try-on generation network, including:

[0199] Step S3081: Save the loss function value of the fitting room generation network for each training session;

[0200] Specifically, the loss function value of the virtual fitting generation network includes the sum of the variable loss function value and the perceptual loss function value.

[0201] Step S3082: When the number of iterations is greater than the second preset number of iterations, sort the loss function value of the currently trained virtual try-on generation network and the saved virtual try-on generation network loss function values ​​in ascending order;

[0202] Specifically, the second preset number of iterations can be 1000. When the number of iterations is greater than 1000, the loss function value of the currently trained virtual try-on generation network is sorted in ascending order with the saved loss function value of the virtual try-on generation network to obtain the ranking of the loss function value of the currently trained virtual try-on generation network.

[0203] Step S3083: If the loss function value of the currently trained virtual try-on generation network is greater than the second preset sequence value, then the loss function of the virtual try-on generation network converges.

[0204] Specifically, the second preset sequence value can be 1000. If the ranking of the loss function value of the currently trained fitting room generation network is greater than the 1000th, it indicates that the curve of the loss function value of the fitting room generation network does not decrease significantly within 1000 iterations, and the loss function of the fitting room generation network converges.

[0205] Furthermore, if the virtual try-on generation loss function converges, it signifies that the virtual try-on generation network training is complete. The network parameters at this point are then output, resulting in the trained virtual try-on generation network. The completion of training for both the clothing distortion network and the virtual try-on generation network signifies the completion of the virtual try-on model training. At this point, the trained virtual try-on model can be used to predict the clothing image and the model image to obtain the corresponding high-resolution virtual try-on effect image.

[0206] Please see Figure 11 , Figure 11 This is a schematic diagram illustrating how to obtain a high-resolution fitting effect image according to an embodiment of this application:

[0207] like Figure 11 As shown, a second image is generated by fusing the high-resolution deformed clothing and the human body preservation area image. The second image is then processed by the fitting generation network to generate a high-resolution fitting effect image.

[0208] After training a virtual fitting model using the training method provided in this application, the virtual fitting model can be used to predict virtual fitting images. The virtual fitting image prediction method provided in this application can be implemented by various types of electronic devices with computing capabilities, such as smart terminals and servers.

[0209] In this embodiment, a training method for a virtual try-on model is provided. The virtual try-on model includes a clothing distortion network and a try-on generation network. The training method includes: acquiring an image dataset, wherein the image dataset includes clothing images and model images matching the clothing images; performing data preprocessing on each clothing image and model image to obtain a first image; performing feature extraction and fusion on the first image based on the clothing distortion network to obtain a low-resolution optical flow map; performing sampling operations on the low-resolution optical flow map to obtain a high-resolution deformed clothing; and fusing the high-resolution deformed clothing and the human body preservation region image to obtain a second... The second image is used to generate a clothing distortion network. A clothing distortion network loss function is constructed and trained until it converges, resulting in a trained clothing distortion network. This loss function includes a clothing edge loss function and a clothing deformation perception loss function. Based on the clothing try-on generation network, the second image is segmented to obtain a high-resolution try-on effect image. A clothing try-on generation network loss function is then constructed and trained until it converges, resulting in a trained try-on generation network. This loss function includes a variable loss function and a perception loss function.

[0210] On the one hand, by extracting and fusing features from the first image based on the clothing distortion network to obtain a low-resolution optical flow map, and then sampling the low-resolution optical flow map to obtain a high-resolution deformed clothing, this application can better extract and fuse spatial features of different sizes of clothing images and model images, thereby realizing the guidance of low-resolution optical flow information for high-resolution clothing deformation.

[0211] On the other hand, by training a clothing distortion network based on a clothing distortion network loss function, which includes a clothing edge loss function and a clothing deformation perception loss function, and by training a virtual try-on generation network based on a virtual try-on generation network loss function, a high-resolution virtual try-on effect image is obtained. This virtual try-on generation network loss function includes a variable loss function and a perception loss function. This allows the present application to improve the clothing texture details and the realism of the model trying on clothes in virtual try-on images while reducing the difficulty of the model learning high-resolution images.

[0212] Example 2

[0213] After training the virtual try-on model using the method provided in the above embodiments, a trained virtual try-on model is obtained, which can be used to predict virtual try-on images.

[0214] It is understood that the above embodiment one is the training stage of the virtual fitting model, and embodiment two of this application uses the virtual fitting model to predict the virtual fitting image in order to obtain a high-resolution fitting effect image corresponding to the clothing image to be predicted and the model image.

[0215] The following describes the virtual fitting image prediction method provided in this application embodiment, with reference to the exemplary application and implementation of the terminal provided in the embodiments of this application.

[0216] Please see Figure 12 , Figure 12 This is a flowchart illustrating a method for predicting virtual fitting images provided in an embodiment of this application;

[0217] like Figure 12 As shown, the prediction method for the virtual fitting image includes:

[0218] Step S1201: Obtain the clothing image and model image to be predicted;

[0219] Specifically, the electronic device acquires images of the clothing and the model to be predicted. For example, the user inputs the images of the clothing and the model through an input interface, and the electronic device automatically acquires them after input. Alternatively, the electronic device has a camera that captures images of the clothing and the model. Or, the electronic device stores a library of clothing images and a library of model images, from which the user can select the images of the clothing and the model to be predicted. Or, the electronic device receives images of the clothing and the model uploaded by the user via the network.

[0220] Step S1202: Input the clothing image to be predicted and the model image into the virtual fitting model to obtain a high-resolution fitting effect image corresponding to the clothing image to be predicted and the model image.

[0221] Specifically, the virtual try-on model is trained using the training method described in Example 1. The image of the clothing to be predicted and the image of the model are input into the trained virtual try-on model to obtain a high-resolution try-on effect image output by the model.

[0222] It is understood that the virtual fitting model is trained using the same training method as the virtual fitting model in the above embodiments, and has the same structure and function as the virtual fitting model in the above embodiments, which will not be described in detail here.

[0223] In this embodiment, by inputting the clothing image to be predicted and the model image into the virtual fitting model, a high-resolution fitting effect image corresponding to the clothing image to be predicted and the model image is obtained, which can improve the clothing texture details of the virtual fitting image and the realism of the model trying on the clothes.

[0224] Example 3

[0225] This application embodiment also provides a computer-readable storage medium storing computer-executable instructions for causing an electronic device to execute the training method of the virtual fitting model provided in the above embodiments, for example, such as... Figure 3 The training method for the virtual fitting model shown, or the prediction method for virtual fitting images provided in the above embodiments, for example, such as... Figure 12 The method for predicting virtual fitting images shown.

[0226] In the embodiments of this application, the storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a device including one or any combination of the above-mentioned memories.

[0227] In the embodiments of this application, the executable instructions may take the form of a program, software, software module, script or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as a standalone program or as a module, component, subroutine or other unit suitable for use in a computing environment.

[0228] As an example, executable instructions may, but do not necessarily, correspond to files in the file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple collaborative files (e.g., a file that stores one or more modules, subroutines, or code sections).

[0229] As an example, executable instructions can be deployed to execute on a single computing device (including devices such as smart terminals and servers), or on multiple computing devices located in one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network.

[0230] This application also provides a computer-readable storage medium storing a computer program, which includes program instructions. When executed by an electronic device, the program instructions cause the electronic device to perform the training method for the virtual fitting model or the prediction method for the virtual fitting image as described in the above embodiments.

[0231] This application also provides a computer program product comprising one or more lines of program code stored in a computer-readable storage medium. A processor of an electronic device reads the program code from the computer-readable storage medium and executes the program code to complete the method steps of the virtual fitting model training method or the virtual fitting image prediction method provided in the above embodiments.

[0232] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware, or by a program or program code related to hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0233] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0234] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and not to limit them; under the concept of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of this application as described above. For the sake of brevity, they are not provided in detail; although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A training method for a virtual fitting model, characterized in that, The virtual try-on model includes a clothing distortion network and a try-on generation network, and the method includes: Obtain an image dataset, wherein the image dataset includes clothing images and model images that match the clothing images; Each of the clothing images and the model images is preprocessed to obtain a first image; Based on the clothing distortion network, feature extraction and fusion are performed on the first image to obtain a low-resolution optical flow map; The low-resolution optical flow map is sampled to obtain high-resolution deformable clothing; The high-resolution deformed clothing and the human body preservation area image are fused to obtain a second image; A clothing distortion network loss function is constructed, and the clothing distortion network is trained until the clothing distortion network loss function converges to generate a trained clothing distortion network. The clothing distortion network loss function includes a clothing edge loss function and a clothing deformation perception loss function. Based on the virtual try-on generation network, the second image is segmented to obtain a high-resolution virtual try-on effect image; A loss function for the virtual try-on generation network is constructed, and the virtual try-on generation network is trained until the loss function converges to generate the trained virtual try-on generation network. The loss function of the virtual try-on generation network includes a variable loss function and a perceptual loss function.

2. The method according to claim 1, characterized in that, The step of preprocessing each of the clothing images and the model images to obtain a first image includes: The outer contour of each clothing image is extracted to obtain the edge image corresponding to each clothing image; Background subtraction is performed on each of the clothing images to obtain a high-resolution clothing image corresponding to each clothing image; The high-resolution clothing image is scaled using an image scaling function to obtain a low-resolution clothing image corresponding to each high-resolution clothing image. The human pose evaluation algorithm is used to detect each model image to obtain the human key point region map corresponding to each model image; The human body analysis algorithm is used to extract the human body analysis image corresponding to each model image. The first image is obtained by stitching together the edge image, the low-resolution clothing image, the human body key point region map, and the human body analytical map using a stitching function.

3. The method according to claim 2, characterized in that, The clothing distortion network includes an encoding module and a decoding module. The encoding module includes a convolutional layer, an activation function layer, and a batch normalization layer. The decoding module includes a transposed convolutional layer, an activation function layer, and a batch normalization layer. The step of extracting and fusing features from the first image based on the clothing distortion network to obtain a low-resolution optical flow map includes: Based on the encoding module, feature extraction and fusion are performed on the first image to obtain three feature maps of different sizes; The low-resolution optical flow map is obtained by decoding the three feature maps of different sizes output by the encoding module based on the decoding module.

4. The method according to claim 3, characterized in that, The sampling operation on the low-resolution optical flow map to obtain high-resolution deformable clothing includes: The low-resolution optical flow map is upsampled using an array sampling operation function to obtain a high-resolution optical flow map. Based on the two-dimensional feature coordinate information of the high-resolution optical flow map, the high-resolution clothing is distorted and deformed using a bilinear interpolation function to obtain the high-resolution deformed clothing.

5. The method according to claim 4, characterized in that, The method further includes: Background subtraction is performed on each model image based on an image segmentation network to obtain a white background model image corresponding to each model image. The human body analysis image and the white background model image are multiplied by a matrix to obtain the human body preserved region image; The process of fusing the high-resolution deformed clothing and the human body preservation area image to obtain a second image includes: The high-resolution deformed clothing and the image of the human body's preserved region are matrix-added to obtain a second image.

6. The method according to claim 1, characterized in that... The process of constructing a clothing distortion network loss function and training the clothing distortion network until the loss function converges to generate the trained clothing distortion network includes: Save the loss function value of the clothing distortion network during each training session; When the number of iterations exceeds the first preset number of iterations, sort the currently trained clothing distortion network loss function value and the saved clothing distortion network loss function value in ascending order. If the loss function value of the currently trained clothing distortion network is greater than the first preset sequence value, then the clothing distortion network loss function converges.

7. The method according to claim 1, characterized in that... The step of constructing a loss function for the virtual try-on generation network and training the virtual try-on generation network until the loss function converges to generate the trained virtual try-on generation network includes: Save the loss function value of the virtual fitting generation network for each training session; When the number of iterations exceeds the second preset number, the loss function value of the currently trained virtual try-on generation network is sorted in ascending order with the saved loss function value of the virtual try-on generation network. If the loss function value of the currently trained virtual try-on generation network is greater than the second preset sequence value, then the loss function of the virtual try-on generation network converges.

8. The method according to any one of claims 1-7, characterized in that, The clothing distortion network includes: in, Represents the (x+1)th level For each feature image, BN represents the batch normalization operation. This represents the ReLU activation function. Represents the xth layer. Each feature image, This represents the standard convolutional kernel of the (x+1)th layer. This indicates the bias term.

9. The method according to any one of claims 1-7, characterized in that, The loss function of the clothing distortion network includes: in, This represents the loss function of the clothing distortion network. This represents the loss function for clothing edges. Indicates the first hyperparameter. This represents the loss function for perceiving clothing deformation.

10. The method according to claim 9, characterized in that, The clothing edge loss function includes: in, This represents the loss function for clothing edges. Represents a real contour sample. Represents the predicted contour sample. This represents the absolute difference between the true contour sample and the predicted contour sample.

11. The method according to claim 9, characterized in that, The clothing deformation perception loss function includes: in, Let M represent the loss function for clothing deformation perception, and M represent the number of layers in the VGG network. This represents the number of feature elements in the i-th layer of the VGG network. This represents a real, high-resolution sample of deformable clothing. This represents a high-resolution sample of the predicted deformed clothing. This represents a feature image of a real high-resolution deformable clothing sample in the i-th layer of the VGG network. This represents the feature image of the predicted high-resolution deformable clothing sample in the i-th layer of the VGG network. This represents the absolute difference between the feature image of the real high-resolution deformable clothing sample in the i-th layer of the VGG network and the feature image of the predicted high-resolution deformable clothing sample in the i-th layer of the VGG network.

12. The method according to any one of claims 1-7, characterized in that, The loss function of the virtual fitting generation network includes: in, This represents the loss function of the virtual fitting generation network. Represents the variable loss function, This represents the second hyperparameter. This represents the perceptual loss function.

13. The method according to claim 12, characterized in that, The variable loss function includes: in, Represents the variable loss function, Represents a real sample. Indicates the predicted sample, Indicates the third hyperparameter. This represents the absolute difference between the actual sample and the predicted sample. This represents the squared error between the actual sample and the predicted sample.

14. The method according to claim 12, characterized in that, The perception loss function includes: in, Let M represent the perceptual loss function, and M represent the number of layers in the VGG network. This represents the number of feature elements in the i-th layer of the VGG network. Represents a real sample. Indicates the predicted sample, This represents the feature image of a real sample in the i-th layer of the VGG network. This represents the feature image of the predicted sample in the i-th layer of the VGG network. This represents the absolute difference between the feature image of the real sample in the i-th layer of the VGG network and the feature image of the predicted sample in the i-th layer of the VGG network.

15. A method for predicting virtual fitting images, characterized in that, The method includes: Obtain the clothing image and model image to be predicted; The clothing image to be predicted and the model image are input into the virtual fitting model to obtain a high-resolution fitting effect image corresponding to the clothing image to be predicted and the model image. The virtual fitting model is trained based on the method described in any one of claims 1-14.

16. An electronic device, characterized in that, include: At least one processor, and The memory communicatively connected to the at least one processor, wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method according to any one of claims 1-15.