Method for training fitting model, method for generating fitting image and related device
Through the method of training fitting models, the fitting network analyzes and integrates the dressing and user images in the existing technology, and solves the problem of inappropriate combination of dressing and user images, and realizes the style of dressing and the fitting effect of dressing and the real and naturalness of the dressing.
Patent Information
- Application Number
- CN202210286675.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-22
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-03-22
AI Technical Summary
In the online trial installation of the existing technology, the combination of the trial clothes and the user image is not appropriate enough, and it is easy to lose the style of the trial clothes, resulting in a poor shopping experience.
Using the method of training fitting models, through the fitting network, including the clothing encoding network and the analytical fusion network, the trained fitting model can make the dressing fitting appropriately combined with the user and maintain the style of the dressing. The method includes obtaining the training set, analyzing and encoding the real fitting image and clothing image using human analytical algorithm and clothing encoding network, inputting the analytical fusion network for fusion analysis, calculating the loss function and iteratively training the fitting network.
It achieves the appropriate combination of clothes to try on and users, maintains the style of clothes to try on and the fitting effect is real and natural, and improves the shopping experience.
Smart Images

Figure CN114821220B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of image processing technology, and in particular to a method for training a fitting model, a method for generating a fitting image, and related devices. Background Art
[0002] With the continuous advancement of modern technology, the scale of online shopping continues to increase. Users can buy clothes on online shopping platforms through their mobile phones. However, since the information about clothes for sale obtained by users is generally two-dimensional display pictures, users cannot know the effect of wearing these clothes on themselves, which may result in buying clothes that are not suitable for them, resulting in a poor shopping experience.
[0003] Online fitting generally involves taking a user image, selecting the target clothing provided by the system, and automatically replacing it. Specifically, most of them collect body data and information about the clothing they want to try on, and reshape the user image and target clothing through 3D modeling. However, the combination of the target clothing and the user image is often not appropriate enough, and the style of the clothes being tried on is easily lost. Summary of the invention
[0004] The main technical problem solved by the embodiments of the present application is to provide a method for training a fitting model, a method for generating a fitting image and related devices. The fitting model trained by this method can make the tried-on clothes fit the user closely, maintain the style of the tried-on clothes, and achieve a real and natural fitting effect.
[0005] In order to solve the above technical problems, in a first aspect, a method for training a fitting model is provided in an embodiment of the present application, wherein the fitting network includes a clothing encoding network and a parsing fusion network;
[0006] The method includes:
[0007] Obtain a training set, the training set includes multiple training data, the training data includes clothing images and real fitting images, the real fitting images include images of models wearing corresponding clothes in the clothing images;
[0008] The human body analysis algorithm is used to analyze the real fitting images to obtain a preliminary human body analysis map;
[0009] Use the clothing encoding network to downsample and encode the clothing image to obtain the clothing feature map;
[0010] The preliminary human body parsing map and clothing feature map are input into the parsing fusion network for fusion parsing to obtain the predicted human body parsing map;
[0011] The loss function is used to calculate the difference between the preliminary human body parsing map and the predicted human body parsing map corresponding to each real fitting image in the training set, wherein the loss function includes a constraint function for constraining the positional relationship between the clothes in the clothing image and the human body in the preliminary human body parsing map;
[0012] According to the differences, the fitting network is iteratively trained until the fitting network converges to obtain the fitting model.
[0013] In some embodiments, the aforementioned parsing fusion network includes an encoder, a fusion module and a decoder.
[0014] The preliminary human body analysis map and clothing feature map are input into the analysis fusion network for fusion to obtain the predicted human body analysis map, including:
[0015] Input the preliminary human body analysis map into the encoder for downsampling encoding to obtain a human body feature map;
[0016] The human body feature map and the clothing feature map are input into the fusion module for fusion, so as to obtain a fused feature map with clothing style information;
[0017] The fused feature map is input into the decoder for upsampling decoding to obtain the predicted human body parsing map.
[0018] In some embodiments, the above-mentioned inputting the human body feature map and the clothing feature map into the fusion module for fusion to obtain a fused feature map fused with clothing style information includes:
[0019] The pixels in the human feature map and the pixels in the clothing feature map are multiplied by corresponding positions to obtain a fused feature map.
[0020] In some embodiments, the aforementioned constraint function includes at least one preset area constraint function, and the preset area constraint function is used to constrain the distance between a preset clothing area and a preset human body area.
[0021] In some embodiments, the method further comprises:
[0022] The human body key point algorithm is used to detect the key points of the real fitting images and obtain the key points of the model's body;
[0023] Use clothing key point algorithm to detect key points of clothing images and obtain clothing key points;
[0024] The key points of the model's body and clothing are substituted into the constraint function for calculation to satisfy the distance relationship constrained by the constraint function.
[0025] In some embodiments, the aforementioned at least one preset area constraint function includes the following inequality:
[0026]
[0027]
[0028] (|f 2 -C 4 |+|f 5 -C 5 |)<ε 3
[0029]
[0030] Among them, ε 1 is the distance constraint value between the left arm and the left sleeve, f 2 and f 4 are the key point coordinates of the left arm, C 4 and C 6 are the key point coordinates of the left sleeve, ε 2 is the distance constraint value between the right arm and the right sleeve, f 5 and f 7 are the key point coordinates of the right arm, C 5 and C 8 are the key point coordinates of the right sleeve, ε 3 is the distance constraint value between the shoulder and the shoulder of the clothing, ε 4 is the distance constraint between the hip and the hem of the garment, f 8 and f 11 are the key point coordinates of the hip, C 12 and C 13 They are the key point coordinates of the hem of the clothes.
[0031] In some embodiments, before the human body analysis algorithm is used to analyze the real fitting image to obtain a preliminary human body analysis diagram, the method further includes:
[0032] The real fitting image is filled so that the resolution of the real fitting image meets the preset ratio.
[0033] In order to solve the above technical problems, in a second aspect, an embodiment of the present application provides a method for generating a fitting image, comprising:
[0034] Obtaining images of clothes to be tried and user images;
[0035] Use the human body analysis algorithm to analyze the user's image and obtain the user's preliminary human body analysis map;
[0036] Inputting the image of the clothes to be tried on and the preliminary human body analysis diagram of the user into the fitting model to obtain a new human body analysis diagram of the user, wherein the fitting model is trained using the method for training the fitting model in the first aspect above;
[0037] The regional categories of the user's new human body analysis image are hidden to obtain a fitting image.
[0038] In order to solve the above technical problems, in a third aspect, an embodiment of the present application provides a computer device, including:
[0039] at least one processor, and
[0040] a memory communicatively coupled to at least one processor, wherein:
[0041] The memory stores instructions that can be executed by at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method in the first aspect as described above.
[0042] To solve the above technical problems, in a fourth aspect, a computer-readable storage medium is provided in an embodiment of the present application, and the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer device to execute the method of the first aspect as above.
[0043] Beneficial effects of the embodiments of the present application: Different from the prior art, the method for training a fitting model provided by the embodiments of the present application, the fitting network as the network structure of the fitting model includes a clothing encoding network and a parsing fusion network, firstly, a training set is obtained, the training set includes a plurality of training data, each of which includes a clothing image and a real fitting image of a model wearing the corresponding clothing in the clothing image, that is, the real fitting image can be used as a real label. Then, for each training data, a human body parsing algorithm is used to parse the real fitting image to obtain a preliminary human body parsing graph. The clothing encoding network is used to downsample and encode the clothing image to obtain a clothing feature graph. The preliminary human body parsing graph and the clothing feature graph are input into the parsing fusion network for fusion parsing to obtain a predicted human body parsing graph, so that the predicted human body parsing graph is fused with the characteristics of the fitting clothes and the human body characteristics of the model, that is, it can maintain the style of the fitting clothes and retain the human body characteristics of the model without distortion. The difference between the preliminary human body parsing graph and the predicted human body parsing graph corresponding to each real fitting image in the training set is calculated using a loss function, and according to the difference, the fitting network is iteratively trained until the fitting network converges to obtain a fitting model. Among them, the loss function includes a constraint function for constraining the positional relationship between the clothes in the clothes image and the human body in the preliminary human body analysis diagram. Under the constraint of the constraint function, the clothes to be tried on can be accurately positioned and fit the model. In addition, the clothes can also maintain their own style characteristics and are not affected by the original clothes characteristics of the model. For example, if the model is wearing tight clothes, and the clothes to be tried on are loose clothes such as fleece or coat, then under the action of the constraint function, the loose clothes will still maintain the original loose style and will not be affected by the original tight style of the model. With the continuous iterative training of the fitting network, the pre-test clothes image of the fusion analysis will continue to approach the real fitting image (real label), and a fitting model that can maintain the style of the clothes to be tried on and fit the human body can be obtained. Therefore, the fitting image of the user generated by the fitting model can make the clothes to be tried on fit the user, maintain the style of the clothes to be tried on, and the fitting effect is real and natural. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] One or more embodiments are exemplarily described by pictures in the corresponding drawings, and these exemplified descriptions do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0045] Figure 1 A flowchart of a method for training a fitting model provided in some embodiments of the present application;
[0046] Figure 2 Human body analysis diagram in some embodiments of the present application;
[0047] Figure 3 for Figure 1 A schematic diagram of a sub-process of step S40 in the method shown;
[0048] Figure 4 A flowchart of a method for training a fitting model provided in some embodiments of the present application;
[0049] Figure 5 Schematic diagram of key points of the human body in some embodiments of the present application;
[0050] Figure 6 Schematic diagram of key points of clothes in some embodiments of the present application;
[0051] Figure 7 A schematic flow chart of a method for generating a fitting image in some embodiments of the present application;
[0052] Figure 8 This is a schematic diagram of the structure of a computer device in some embodiments of the present application. DETAILED DESCRIPTION
[0053] The present application is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present application, but do not limit the present application in any form. It should be noted that, for those of ordinary skill in the art, several variations and improvements can also be made without departing from the concept of the present application. These all belong to the protection scope of the present application.
[0054] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0055] It should be noted that, if there is no conflict, the various features in the embodiments of the present application can be combined with each other, all within the scope of protection of the present application. In addition, although the functional module division is performed in the device schematic diagram and the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in a sequence different from the module division in the device or the flow chart. In addition, the words "first", "second", "third", etc. used herein do not limit the data and execution order, but only distinguish the same items or similar items with basically the same functions and effects.
[0056] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used in this specification and in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The term "and / or" used in this specification includes any and all combinations of one or more of the related listed items.
[0057] In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as they do not conflict with each other.
[0058] To facilitate understanding of the method provided in the embodiments of the present application, the nouns involved in the embodiments of the present application are first introduced:
[0059] (1) Neural Network
[0060] A neural network can be composed of neural units. Specifically, it can be understood as a neural network with an input layer, a hidden layer, and an output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the layers in between are all hidden layers. Among them, a neural network with many hidden layers is called a deep neural network (DNN). The work of each layer in a neural network can be described by the mathematical expression y=a(W·x+b). From a physical level, the work of each layer in a neural network can be understood as completing the transformation from input space to output space (i.e., the row space to column space of a matrix) through five operations on the input space (a set of input vectors). These five operations include: 1. Dimensionality increase / reduction; 2. Zoom in / out; 3. Rotation; 4. Translation; 5. "Bending". Among them, operations 2 and 3 are completed by "W·x", operation 4 is completed by "+b", and operation 5 is implemented by "a()". The word "space" is used here because the classified object is not a single thing, but a class of things. Space refers to the collection of all individuals of this type of thing. Among them, W is the weight matrix of each layer of the neural network. Each value in the matrix represents the weight value of a neuron in this layer. The matrix W determines the spatial transformation from the input space to the output space mentioned above, that is, the W of each layer of the neural network controls how to transform the space. The purpose of training a neural network is to finally obtain the weight matrix of all layers of the trained neural network. Therefore, the training process of a neural network is essentially to learn how to control spatial transformation, and more specifically to learn the weight matrix.
[0061] It should be noted that in the embodiments of the present application, the models used in the machine learning tasks are essentially neural networks. Common components in neural networks include convolutional layers, pooling layers, normalization layers, and inverse convolutional layers. By assembling these common components in neural networks, a model is designed. When the model parameters (weight matrices of each layer) are determined so that the model error meets the preset conditions or the number of model parameters is adjusted to reach a preset threshold, the model converges.
[0062] The convolution layer is configured with multiple convolution kernels, each of which is set with a corresponding step size to perform convolution operations on the image. The purpose of the convolution operation is to extract different features of the input image. The first convolution layer may only extract some low-level features such as edges, lines, and corners. Deeper convolution layers can iteratively extract more complex features from low-level features.
[0063] The deconvolution layer is used to map a low-dimensional space to a high-dimensional space while maintaining the connection relationship / pattern between them (the connection relationship here refers to the connection relationship during convolution). The deconvolution layer is configured with multiple convolution kernels, each of which is set with a corresponding step size to perform deconvolution operations on the image. Generally, the framework library used to design neural networks (such as the PyTorch library) has a built-in upsumple() function, which can be called to achieve low-dimensional to high-dimensional spatial mapping.
[0064] Pooling layer is to imitate the human visual system to reduce the dimension of data or represent the image with higher-level features. Common operations of pooling layer include maximum pooling, mean pooling, random pooling, median pooling and combined pooling. Generally speaking, pooling layers are periodically inserted between convolutional layers of neural networks to achieve dimensionality reduction.
[0065] The normalization layer is used to normalize all neurons in the intermediate layer to prevent gradient explosion and gradient disappearance.
[0066] (2) Loss Function
[0067] In the process of training a neural network, because we want the output of the neural network to be as close as possible to the value we really want to predict, we can compare the current network's predicted value with the target value we really want, and then update the weight matrix of each layer of the neural network according to the difference between the two (of course, there is usually an initialization process before the first update, that is, pre-configuring parameters for each layer in the neural network). For example, if the network's predicted value is high, adjust the weight matrix to make it predict lower, and keep adjusting until the neural network can predict the target value we really want. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or objective function, which are important equations used to measure the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference, so the training of the neural network becomes a process of minimizing this loss as much as possible.
[0068] (3) Human body analysis
[0069] Human body parsing refers to segmenting a person captured in an image into multiple semantically consistent regions, such as body parts and clothing, or subdivided categories of body parts and subdivided categories of clothing, etc. That is, the input image is recognized at the pixel level, and each pixel in the image is labeled with the object category to which it belongs. For example, a neural network is used to distinguish various elements in a picture of a human body (including hair, face, limbs, clothing, and background, etc.).
[0070] Before introducing the embodiments of the present application, a virtual fitting method known to the inventor of the present application is briefly introduced to facilitate subsequent understanding of the embodiments of the present application.
[0071] Offline fitting generally uses interactive fitting mirror equipment. Shoppers stand in front of the interactive fitting mirror equipment and choose to try on clothes through its screen. The interactive fitting mirror equipment will collect user images and display the overall clothing according to the basic characteristics and clothing styles of the user. Users can switch clothes freely. Online fitting generally takes user images and selects target clothing provided by the system for automatic replacement. However, whether it is an interactive fitting mirror device or an online virtual fitting, most of them are reshaped by collecting human body data and reshaping the user image through 3D modeling, that is, building a 3D human body model based on a two-dimensional human body image, and setting the try-on clothes in the corresponding position of the 3D human body model. Often, the combination of clothes and user images is often not appropriate enough, lacking a natural feel, and it is easy to lose the style of the clothes being tried on (for example, loose clothes become tight after trying on), resulting in a poor user experience. At the same time, collecting 3D information of the human body and clothing is usually costly and cumbersome.
[0072] In view of the above problems, the present application provides a method for training a fitting model. The embodiments of the present application are described below in conjunction with the accompanying drawings. It is known to those skilled in the art that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0073] See also Figure 1 , Figure 1 The flowchart of the method for training a fitting model provided in the embodiment of the present application is as follows. The fitting network as the network structure of the fitting model includes a clothing encoding network and a parsing fusion network. The method S100 may specifically include the following steps:
[0074] S10: Obtain a training set.
[0075] The training set includes a plurality of training data, each of which includes clothing images and real fitting images, and the real fitting images include images of models wearing corresponding clothing in the aforementioned clothing images.
[0076] It is understandable that a training data set includes image pairs consisting of clothing images and real fitting images. In some embodiments, the number of training data sets is in the tens of thousands, for example, 20,000, which is conducive to training to obtain an accurate general model. Those skilled in the art can determine the number of training data sets according to actual conditions.
[0077] In the image pair consisting of the clothing image and the real fitting image, the clothing image includes the clothing that one wants to try on, for example, clothing image 1# includes a green short-sleeved shirt. In the real fitting image, the model wears the clothing in the corresponding clothing image, for example, in the real fitting image corresponding to clothing image 1#, the model wears the green short-sleeved shirt.
[0078] It is understandable that the training data including clothing images and real fitting images can be collected in advance by those skilled in the art. For example, clothing images and corresponding images of models wearing the clothes (i.e., real fitting images) can be crawled from some clothing sales websites.
[0079] In some embodiments, the real fitting images are full-body photos of models, and their image ratios are similar to the length-to-width ratio of the human body. However, the length-to-width ratio of the human body does not meet the network training requirements, that is, the resolution ratio of the real fitting images does not meet the network training requirements. Therefore, it is necessary to pre-process the real fitting images in the training set so that the resolution ratio of the real fitting images meets the network training requirements.
[0080] In some embodiments, before training, the method further includes:
[0081] The real fitting image is filled so that the resolution of the real fitting image meets the preset ratio.
[0082] Since the real fitting image is a full-body photo of a model, the image ratio is close to the length-to-width ratio of the human body. In order not to change the model's posture, the width of the real fitting image is padded. The difference between the image length and width is used to fill the width with pixels in proportion. For example, pixels with a value of 0 can be filled, so that the resolution of the real fitting image meets the preset ratio.
[0083] In some embodiments, the preset ratio is 1:1, that is, the length and width of the real fitting image after the filling process are consistent. For example, the resolution of the real fitting image can be 512*512. In some embodiments, before training, the real fitting image after the filling process can be normalized to facilitate the training speed of the fitting network. Since the normalization operation is a common data processing method well known to those skilled in the art, the normalization operation will not be explained in detail here.
[0084] S20: using a human body analysis algorithm to analyze the real fitting image and obtain a preliminary human body analysis diagram.
[0085] As mentioned in “Term Introduction (3)”, human body analysis is to separate the various parts of the human body, such as Figure 2 As shown, different parts, such as hair, face, top, pants, arms, hat, shoes, etc., are identified and segmented and represented by different colors, thus obtaining a human body analysis diagram.
[0086] In some embodiments, the human body parsing algorithm may be an existing Graphonomy algorithm. The Graphomay algorithm will segment the image into 20 categories, which can be distinguished by different colors to divide the parts. In some embodiments, the 20 categories may also be divided by numbers 0-19, for example, 0 represents background, 1 represents hat, 2 represents hair, 3 represents gloves, 4 represents sunglasses, 5 represents top, 6 represents dress, 7 represents coat, 8 represents socks, 9 represents pants, 10 represents trunk skin, 11 represents scarf, 12 represents skirt, 13 represents face, 14 represents left arm, 15 represents right arm, 16 represents left leg, 17 represents right leg, 18 represents left shoe, and 19 represents right shoe. From the human body parsing diagram, the category to which each part in the image belongs can be determined.
[0087] In order to ensure that the analytical categories of the key areas of the human body in the preliminary human body analytical diagram and the new human body analytical diagram subsequently generated based on the preliminary human body analytical diagram are consistent, that is, to reduce the analytical error, in some embodiments, the analytical categories are simplified. Specifically, according to the analytical needs, the categories of human body areas such as the head, arms, legs, and feet are set to 1, and the rest are set to 0. It can be understood that reducing the analytical error in the above manner is conducive to improving the convergence speed and accuracy of the model.
[0088] S30: down-sample and encode the clothing image using a clothing encoding network to obtain a clothing feature map.
[0089] In order to fully integrate the clothing image and the preliminary human body analysis map, first, the clothing image is downsampled and encoded using a clothing encoding network, and features are extracted to obtain a clothing feature map. In some embodiments, the clothing feature map can be a 4*4*256 feature map obtained after the clothing image is downsampled.
[0090] In some embodiments, the structure of the clothing encoding network is shown in Table 1 below:
[0091] Table 1 Structure of clothing encoding network
[0092] layer Convolution kernel size Step Length The size of the output feature map Clothes ImageX - - 512*512*3 Convolutional Layer 1*1 2 256*256*16 Convolutional Layer 3*3 2 128*128*32 Convolutional Layer 3*3 2 64*64*32 Convolutional Layer 3*3 2 32*32*64 Convolutional Layer 3*3 2 16*16*64 Convolutional Layer 3*3 2 8*8*128 Convolutional Layer 3*3 2 4*4*256
[0093] As can be seen from Table 1 above, the clothing encoding network includes multiple convolutional layers, most of which are configured with 3*3 convolution kernels and the step size is set to 2, thereby achieving downsampling and dimensionality reduction operations.
[0094] S40: Input the preliminary human body analysis map and the clothing feature map into the analysis fusion network for fusion analysis to obtain a predicted human body analysis map.
[0095] It is understandable that the analytical fusion network is a neural network, and its structure and mechanism have been described in detail in the aforementioned "(1) Neural Network". It is understandable that the analytical fusion network includes multiple convolutional layers, through which the preliminary human body analytical map is upsampled and / or downsampled. In the process of upsampling and downsampling, feature maps of different sizes are generated. The clothing feature map can be fused with at least one of the feature maps, and the fused feature map continues to be convolutionally processed to output an image fused with the feature information of the tried-on clothes.
[0096] The parsing fusion network also has the function of implementing human body parsing, which has been described in detail in the above "(3) Human Body Parsing". Similar to the above-mentioned Graphonomy algorithm, the parsing fusion network will segment the image fused with the feature information of the tried-on clothes according to the preset categories. In some embodiments, the preset categories can be set in this way: the categories of human body regions such as the head, arms, legs, and feet are set to 1, and the rest are set to 0.
[0097] Finally, the parsing fusion network outputs the predicted human body parsing map, so that the predicted human body parsing map not only includes the parsed categories, but also incorporates the feature information of the clothes being tried on.
[0098] In some embodiments, the parsing fusion network includes an encoder, a fusion module, and a decoder.
[0099] See also Figure 3 , the aforementioned step S40 specifically includes:
[0100] S41: Input the preliminary human body analysis map into the encoder for downsampling encoding to obtain a human body feature map.
[0101] In some embodiments, the encoder includes 7 downsampling convolutional layers, which downsample the preliminary human body analysis map layer by layer, extract features, and generate feature maps of gradually decreasing sizes. Finally, the size of the output human body feature map can be 4*4*256.
[0102] S42: Inputting the human body feature map and the clothing feature map into a fusion module for fusion, thereby obtaining a fused feature map fused with clothing style information.
[0103] It is understandable that the size of the human feature map and the clothing feature map is consistent, for example, both are 4*4*256. The human feature map and the clothing feature map are input into the fusion module for fusion. Since the human feature map includes the feature information of the model's body and the clothing feature map includes the feature information of the clothes being tried on, the obtained fusion feature map includes not only the feature information of the model's body but also the clothing style information. For example, the loose-fitting style of the clothes being tried on still maintains the loose style in the fusion feature map.
[0104] In some embodiments, the fusion mode adopted by the fusion module includes linear fusion or nonlinear fusion. It is understood that the linear fusion here refers to the corresponding pixels of the two images being operated once to obtain a fused image, and the nonlinear fusion here refers to the corresponding pixels of the two images being operated twice or more to obtain a fused image.
[0105] In some embodiments, the aforementioned step S42 specifically includes: multiplying the pixels in the human body feature map and the pixels in the clothing feature map by corresponding positions to obtain a fused feature map.
[0106] Here, the corresponding positions are multiplied, and the formula R can be used. ij =P ij ×C ij Explain, where P ij represents the pixel value in the i-th row and j-th column of the human feature map, C ij Represents the pixel value of the i-th row and j-th column in the clothing feature map, R ij is the pixel value of the i-th row and j-th column in the fused feature map.
[0107] S43: Input the fused feature map into the decoder for upsampling decoding to obtain a predicted human body parsing map.
[0108] In some embodiments, the decoder includes 7 upsampling convolution layers, which upsample the fused feature map layer by layer, decode the features, and generate feature maps of gradually increasing sizes. The feature map output by the last convolution layer is a predicted human body analysis map of size 512*512*3. Based on the fused feature map, which includes not only the feature information of the model's body but also the clothing style information, the predicted human body analysis map integrates the features of the clothes being tried on and the features of the model's body, that is, it can maintain the style of the clothes being tried on and retain the human body features of the model without distortion.
[0109] In this embodiment, the features of the preliminary human body analysis map are extracted by an encoder to generate a human body feature map, the human body feature map and the clothing feature map are fused by a fusion module to generate a fused feature map, and the fused feature map is decoded by a decoder to obtain a predicted human body analysis map, so that the predicted human body analysis map is fused with the features of the tried-on clothes and the body features of the model, that is, it can maintain the style of the tried-on clothes and retain the body features of the model without distortion.
[0110] S50: using a loss function to calculate the difference between the preliminary human body analysis map and the predicted human body analysis map corresponding to each real fitting image in the training set, wherein the loss function includes a constraint function for constraining the positional relationship between clothes in the clothing image and the human body in the preliminary human body analysis map.
[0111] S60: According to the aforementioned differences, the fitting network is iteratively trained until the fitting network converges to obtain a fitting model.
[0112] It can be understood that if the difference between the preliminary human body analysis map and the predicted human body analysis map corresponding to each real fitting image in the training set is smaller, the preliminary human body analysis map and the predicted human body analysis map are more similar, indicating that the predicted human body analysis map can accurately restore the preliminary human body analysis map. Therefore, according to the difference between the preliminary human body analysis map and the predicted human body analysis map corresponding to each real fitting image in the training set, the model parameters of the aforementioned fitting network can be adjusted, and the fitting network can be iteratively trained. Based on the fitting network including the clothing encoding network and the analytical fusion network, the model parameters include the model parameters of the clothing encoding network and the model parameters of the analytical fusion network. That is, the above differences are back-propagated so that the predicted human body analysis map output by the fitting network continuously approaches the preliminary human body analysis map until the fitting network converges and the fitting model is obtained.
[0113] It can be understood that the convergence of the fitting network here can refer to that under certain model parameters, the sum of the differences between the preliminary human body analysis map and the predicted human body analysis map corresponding to each real fitting image in the training set is less than a preset threshold or fluctuates within a certain range.
[0114] In some embodiments, the ADAM algorithm is used to optimize the model parameters. For example, the number of iterations is set to 100,000 times, the initialization learning rate is set to 0.001, the weight decay of the learning rate is set to 0.0005, and the learning rate decays to 1 / 10 of the original value every 1,000 iterations. The learning rate and the difference between the preliminary human body analysis diagrams in the training set and the predicted human body analysis diagrams can be input into the ADAM algorithm to obtain the adjusted model parameters output by the ADAM algorithm. The adjusted model parameters are used for the next training until the training is completed, and the model parameters of the fitting network after convergence are output to obtain the fitting model.
[0115] It should be noted that in the embodiment of the present application, the training set includes multiple training data, for example, 20,000 training data, which cover different models and clothes and can cover the characteristics of most clothes on the market. Therefore, the trained fitting model is a universal model that can be widely used in virtual fitting and generating fitting images.
[0116] In this embodiment, a loss function is used to calculate the difference between the preliminary human body analysis map and the predicted human body analysis map corresponding to each real fitting image in the training set. The loss function has been introduced in detail in the aforementioned "Term Introduction (2)" and will not be repeated here. It can be understood that based on different network structures and training methods, the structure of the loss function can be set according to actual conditions.
[0117] In some embodiments, the loss function includes a key area loss and a reconstruction loss, wherein the key area loss reflects the analytical difference between the predicted human body parsing map and the preliminary human body parsing map, and the reconstruction loss reflects the pixel difference between the predicted human body parsing map and the preliminary human body parsing map.
[0118] In some embodiments, the loss function is:
[0119] L loss =α*L loc +β*L rec ;
[0120] Among them, L loc =‖Y mask -X mask ‖ L1 ,
[0121] Among them, L loss is the loss function, L loc is the critical area loss, L rec is the reconstruction loss, X mask is the key area in the preliminary human body analysis diagram, Y mask To predict the key areas in the human body analysis graph, the key areas here are the parts and categories segmented and identified along the way. i, Y is the pixel value of the i-th row and j-th column in the preliminary human body analysis image. i, To predict the pixel value of the i-th row and j-th column in the human body analysis image.
[0122] In some embodiments, the loss function also includes a constraint function for constraining the positional relationship between the clothes in the clothes image and the human body in the preliminary human body analysis image. Under the constraint of the constraint function, the clothes to be tried on can be accurately positioned and fit the model. On the other hand, the clothes can also maintain their own style characteristics and are not affected by the original clothing characteristics of the model. For example, if the model is wearing tight clothes and the clothes to be tried on are loose clothes such as fleece or coat, since the clothes to be tried on will not be deformed according to the original clothes, the clothes to be tried on can be adjusted according to their own style and adapted to the human body under the action of the constraint function, and the loose clothes will still maintain their original loose style and will not be affected by the original tight style of the model.
[0123] With the continuous iterative training of the fitting network, the predicted human body analysis map of the fusion analysis will continue to approach the preliminary human body analysis map (equivalent to the real human body analysis map), and a fitting model that can maintain the style of the tried-on clothes and fit the human body can be obtained.
[0124] In some embodiments, the constraint function includes at least one preset area constraint function, and the preset area constraint function is used to constrain the distance between a preset clothing area and a preset human body area.
[0125] Among them, the preset clothing area may include a cuff area, a shoulder area, a top hem area or a trouser leg area. The preset human body area may include an arm area, a shoulder area, a hip area or a leg area, etc. It can be understood that when a human body is trying on clothes, the preset clothing area and the preset human body area have a corresponding relationship, for example, the cuff area corresponds to the arm area, the shoulder area corresponds to the shoulder area, the top hem area corresponds to the hip area, etc. By constraining the distance between these corresponding clothing areas and human body areas through the preset area constraint function, the combination of the tried-on clothes and the human body can be made appropriate and natural.
[0126] In some embodiments, see Figure 4 , the method S100 further includes:
[0127] S70: Using a human body key point algorithm to perform key point detection on the real fitting image to obtain the model's human body key points.
[0128] S80: Using a clothing key point algorithm to perform key point detection on the clothing image to obtain clothing key points.
[0129] S90: Substituting the key points of the model's body and the key points of the clothes into the constraint function for calculation to satisfy the distance relationship constrained by the constraint function.
[0130] The key point detection algorithm of the human body is used to detect the key points of the real fitting images, and the key point information of the human body (i.e., several key points on the human body) can be located. Figure 5As shown, the key points include the head, shoulders, arms, legs and body trunk. In some embodiments, the human key point detection algorithm can adopt a 2D key point detection algorithm, such as Convolutional Pose Machine (CPM) or Stacked Hourglass Network (Hourglass).
[0131] By using the clothing key point detection algorithm to detect key points of clothing images, the key points of clothing (i.e., several key points on the clothing) can be located. Figure 6 As shown, key points of areas such as cuffs, necklines, shoulders and hems are included. In some embodiments, a neural network for implementing the clothing key point detection algorithm can be trained based on sample images, wherein the sample images include clothing and are annotated with reference key point coordinates of the clothing, and the neural network includes multiple convolutional layers, pooling layers or fully connected layers, etc., and its structure can adopt an existing deep convolutional neural network, or can be designed by those skilled in the art.
[0132] It can be understood that if the distance between the key points of the deformed clothes and the corresponding parts of the key points of the human body is appropriate, for example, the distance between the key points of the shoulder area of the clothes and the key points of the shoulder area of the human body is within a certain range, then when the clothes and the model are fused, the clothes can fit the human body and look real and natural.
[0133] Therefore, the constraint function can constrain the distance relationship between the tried-on clothes and the human body based on the model's body key points and the clothes key points. Specifically, the model's body key points and the clothes key points are substituted into the constraint function for calculation to satisfy the distance relationship constrained by the constraint function.
[0134] In some embodiments, the aforementioned at least one preset area constraint function includes the following inequality:
[0135]
[0136]
[0137] (|f 2 -C 4 |+|f 5 -C 5 |)<ε 3
[0138]
[0139] Among them, ε 1 is the distance constraint value between the left arm and the left sleeve, f 2 and f 4 are the key point coordinates of the left arm, C 4 and C6 are the key point coordinates of the left sleeve, ε 2 is the distance constraint value between the right arm and the right sleeve, f 5 and f 7 are the key point coordinates of the right arm, C 5 and C 8 are the key point coordinates of the right sleeve, ε 3 is the distance constraint value between the shoulder and the shoulder of the clothing, ε 4 is the distance constraint between the hip and the hem of the garment, f 8 and f 11 are the key point coordinates of the hip, C 12 and C 13 They are the key point coordinates of the hem of the clothes.
[0140] In the above four inequalities, the distance relationship between the key points of the left arm and the key points of the left sleeve is less than ε 1 , the distance between the key points of the right arm and the key points of the right sleeve is less than ε 2 , the distance between the shoulder and the shoulder of the garment is less than ε 3 , the distance between the hips and the hem of the clothes is less than ε 4 Under the constraints of these four distance constraints, the left sleeve of the clothes fits the left arm of the human body, the right sleeve of the clothes fits the right arm of the human body, the shoulder of the clothes fits the shoulder of the human body, and the hem of the clothes fits the hips of the human body. Therefore, the clothes can be accurately positioned when tried on and fit the model closely.
[0141] In this embodiment, the distance relationship between the clothes in the clothes image and the human body in the preliminary human body analysis image is constrained by a constraint function, so that the clothes can maintain their own style characteristics and are not affected by the original clothing characteristics of the model. For example, if the model is wearing tight clothes and the clothes being tried on are loose clothes such as sweaters or coats, then under the action of the constraint function, the loose clothes will still maintain the original loose style and will not be affected by the original tight style of the model.
[0142] In summary, the method for training a fitting model provided by an embodiment of the present application, the fitting network as the network structure of the fitting model includes a clothing encoding network and a parsing fusion network, firstly, a training set is obtained, the training set includes a plurality of training data, each training data includes a clothing image and a real fitting image of a model wearing the corresponding clothing in the clothing image, that is, the real fitting image can be used as a real label. Then, for each training data, a human body parsing algorithm is used to parse the real fitting image to obtain a preliminary human body parsing graph. The clothing encoding network is used to downsample and encode the clothing image to obtain a clothing feature graph. The preliminary human body parsing graph and the clothing feature graph are input into the parsing fusion network for fusion parsing to obtain a predicted human body parsing graph, so that the predicted human body parsing graph is fused with the characteristics of the fitting clothes and the human body characteristics of the model, that is, it can maintain the style of the fitting clothes and retain the human body characteristics of the model without distortion. The difference between the preliminary human body parsing graph and the predicted human body parsing graph corresponding to each real fitting image in the training set is calculated using a loss function, and according to the difference, the fitting network is iteratively trained until the fitting network converges to obtain a fitting model. Among them, the loss function includes a constraint function for constraining the positional relationship between the clothes in the clothes image and the human body in the preliminary human body analysis diagram. Under the constraint of the constraint function, the clothes to be tried on can be accurately positioned and fit the model. In addition, the clothes can also maintain their own style characteristics and are not affected by the original clothes characteristics of the model. For example, if the model is wearing tight clothes, and the clothes to be tried on are loose clothes such as fleece or coat, then under the action of the constraint function, the loose clothes will still maintain the original loose style and will not be affected by the original tight style of the model. With the continuous iterative training of the fitting network, the pre-test clothes image of the fusion analysis will continue to approach the real fitting image (real label), and a fitting model that can maintain the style of the clothes to be tried on and fit the human body can be obtained. Therefore, the fitting image of the user generated by the fitting model can make the clothes to be tried on fit the user, maintain the style of the clothes to be tried on, and the fitting effect is real and natural.
[0143] After the fitting model is trained by the method for training a fitting model provided in this application, the fitting model can be used for virtual fitting to generate a fitting image. Figure 7 , Figure 7 A flow chart of a method for generating a fitting image provided in an embodiment of the present application is shown as follows: Figure 7 As shown, the method S200 includes the following steps:
[0144] S201: Acquire images of clothes to be tried on and images of the user.
[0145] The clothes-to-be-tried image includes clothes, and the user image includes the user's body.
[0146] S202: Analyze the user image using a human body analysis algorithm to obtain a preliminary human body analysis diagram of the user.
[0147] Here, the user image may be analyzed with reference to the aforementioned step S20 to obtain a preliminary human body analysis diagram of the user, and the specific process will not be described in detail here.
[0148] S203: Inputting the image of the clothes to be tried on and the preliminary human body analysis diagram of the user into the fitting model to obtain a new human body analysis diagram of the user, wherein the fitting model is trained using the method for training a fitting model in any of the above embodiments.
[0149] It can be understood that the fitting model is trained by the method of training the fitting model in the above embodiment, and has the same structure and function as the fitting model in the above embodiment, which will not be described one by one here.
[0150] In short, the image of the clothes to be tried and the preliminary body analysis map of the user are input into the fitting model, and the fitting model will downsample and encode the image of the clothes to be tried to obtain a feature map of the clothes to be tried. The fitting model will downsample and encode the preliminary body analysis map of the user to obtain a body feature map of the user. Then, the body feature map of the user and the feature map of the clothes to be tried are fused to obtain a fused feature map that incorporates the style information of the clothes to be tried. After upsampling and decoding the fused feature map, a new body analysis map of the user is obtained.
[0151] Based on the fact that the fitting model has the same structure and function as the fitting model in the above-mentioned embodiment, the clothes being tried on and the user are closely combined in the output new human body analysis image, the style of the clothes being tried on is maintained, and the fitting effect is real and natural.
[0152] S204: Hide the region category of the user's new human body analysis image to obtain a fitting image.
[0153] It is understandable that the user's new human body analysis image also reflects the category information of each area in the image. Therefore, the area category of the user's new human body analysis image is hidden to obtain a fitting image.
[0154] Therefore, the clothes being tried on in the fitting image are closely combined with the user, the style of the clothes being tried on can be maintained, and the fitting effect is real and natural.
[0155] See also Figure 8 , Figure 8 1 is a schematic diagram of a computer device provided in an embodiment of the present application, wherein the computer device 50 includes a processor 501 and a memory 502. The processor 501 is connected to the memory 502, for example, the processor 501 may be connected to the memory 502 via a bus.
[0156] The processor 501 is configured to support the computer device 50 to execute Figure 1-Figure 6 method or Figure 7 The processor 501 may be a central processing unit (CPU), a network processor (NP), a hardware chip or any combination thereof. The hardware chip may be an application specific integrated circuit (ASIC), a programmable logic device (PLD) or a combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.
[0157] The memory 502 is a non-transitory computer-readable storage medium that can be used to store non-transitory software programs, non-transitory computer executable programs and modules, such as the program instructions / modules corresponding to the method for training a fitting model in the embodiments of the present application, or the program instructions / modules corresponding to the method for generating a fitting image. The processor 501 can implement the training in any of the above method embodiments by running the non-transitory software programs, instructions and modules stored in the memory 502.
[0158] The memory 502 may include a volatile memory (VM), such as a random access memory (RAM); the memory 1002 may also include a non-volatile memory (NVM), such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD); the memory 502 may also include a combination of the above types of memory.
[0159] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed by a computer, the computer executes the method for training a fitting model or the method for generating a fitting image in the aforementioned embodiment.
[0160] It should be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0161] Through the description of the above implementation methods, ordinary technicians in this field can clearly understand that each implementation method can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Ordinary technicians in this field can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.
[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Under the concept of the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other changes in different aspects of the present application as described above, which are not provided in detail for the sake of simplicity. Although the present application has been described in detail with reference to the aforementioned embodiments, a person of ordinary skill in the art should understand that the technical solutions described in the aforementioned embodiments can still be modified, or some of the technical features can be replaced by equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for training a fitting model, It is characterized in that The fitting network includes clothing encoding network and parsing fusion network; The method comprises: Acquire a training set, wherein the training set includes a plurality of training data, wherein the training data includes clothing images and real fitting images, wherein the real fitting images include images of a model wearing corresponding clothing in the clothing images; Using a human body analysis algorithm to analyze the real fitting image to obtain a preliminary human body analysis diagram; Down-sampling and encoding the clothing image using a clothing encoding network to obtain a clothing feature map; Inputting the preliminary human body analysis graph and the clothing feature graph into the analysis fusion network for fusion analysis to obtain a predicted human body analysis graph; A loss function is used to calculate the difference between the preliminary human body analysis map and the predicted human body analysis map corresponding to each of the real fitting images in the training set, wherein the loss function includes a constraint function for constraining the positional relationship between the clothes in the clothing image and the human body in the preliminary human body analysis map; the constraint function includes at least one preset area constraint function, and the preset area constraint function is used to constrain the distance between a preset clothing area and a preset human body area; Iteratively training the fitting network according to the difference until the fitting network converges to obtain the fitting model; During the training process, a human body key point algorithm is used to detect key points of the real fitting image to obtain the model's human body key points; a clothing key point algorithm is used to detect key points of the clothing image to obtain clothing key points; the model's human body key points and the clothing key points are substituted into the constraint function for calculation to satisfy the distance relationship constrained by the constraint function; The at least one preset area constraint function includes the following inequality: (|f 2 -C 4 |+|f 5 -C 5 |)<ε 3 Among them, ε 1 is the distance constraint value between the left arm and the left sleeve, f 2 and f 4 are the key point coordinates of the left arm, C 4 and C 6 are the key point coordinates of the left sleeve, ε 2 is the distance constraint value between the right arm and the right sleeve, f 5 and f 7 are the key point coordinates of the right arm, C 5 and C 8 are the key point coordinates of the right sleeve, ε 3 is the distance constraint value between the shoulder and the shoulder of the clothing, ε 4 is the distance constraint between the hip and the hem of the garment, f 8 and f 11 are the key point coordinates of the hip, C 12 and C 13 are respectively the key point coordinates of the hem of the clothes.
2. The method according to claim 1, It is characterized in that The parsing fusion network includes an encoder, a fusion module and a decoder. The step of inputting the preliminary human body parsing graph and the clothing feature graph into the parsing fusion network for fusion to obtain a predicted human body parsing graph comprises: Inputting the preliminary human body analysis map into the encoder for downsampling encoding to obtain a human body feature map; Inputting the human body feature map and the clothing feature map into the fusion module for fusion, so as to obtain a fused feature map fused with clothing style information; The fused feature map is input into the decoder for upsampling decoding to obtain the predicted human body parsing map.
3. The method according to claim 2, It is characterized in that The step of inputting the human body feature map and the clothing feature map into the fusion module for fusion to obtain a fused feature map fused with clothing style information includes: The pixels in the human body feature map and the pixels in the clothing feature map are multiplied at corresponding positions to obtain the fused feature map.
4. The method according to claim 1, It is characterized in that Before the human body analysis algorithm is used to analyze the real fitting image to obtain a preliminary human body analysis diagram, the method further includes: The real fitting image is filled so that the resolution of the real fitting image meets a preset ratio.
5. A method for generating a fitting image, It is characterized in that include: Obtaining images of clothes to be tried and user images; Using a human body analysis algorithm to analyze the user image, and obtain a preliminary human body analysis diagram of the user; Inputting the image of the clothes to be tried on and the preliminary human body analysis diagram of the user into a fitting model to obtain a new human body analysis diagram of the user, wherein the fitting model is trained using the method for training a fitting model as claimed in any one of claims 1 to 4; The region category of the new human body analysis image of the user is hidden to obtain the fitting image.
6. A computer device, It is characterized in that include: at least one processor, and a memory communicatively coupled to the at least one processor, wherein: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.
7. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer device to execute the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
A virtual garment try-on system
CN109934613A