A virtual fitting method and system based on neural network search
Through virtual trial-oning methods and systems based on neural network search, the clothing deformation field is automatically searched and the trial-oning results are generated by human body shape and semantic segmentation diagrams are combined to solve the spatial misalignment problem caused by the difference in clothing deformation in the prior art, and the freedom and success rate of trial-oning are improved.
Patent Information
- Application Number
- CN202111323293.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-09
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-11-09
AI Technical Summary
Existing virtual trial-on technology can easily lead to spatial misalignment, low degree of freedom and success rate when dealing with the deformation differences of different clothing.
Using virtual trial-on method and system based on neural network search, the human body key point map and semantic segmentation map are obtained through semantic prediction network and Openpose network, the deformation network search space is constructed and the fusion network search space is integrated, and the clothing deformation field is automatically searched and deformation is performed, and the try-on results are generated by combining the human body shape and semantic segmentation map.
It improves the freedom and success rate of virtual try-ons, and can generate high-quality try-on images in complex scenarios, solving the problem of spatial dislocation caused by differences in different clothing deformations.
Smart Images

Figure CN114202637B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of virtual fitting, and in particular to a virtual fitting method and system based on neural network search. Background Art
[0002] Due to the rapid development and great practical value of generative models, image-based virtual fitting has great application prospects. The core challenge of virtual fitting is to find a suitable deformation method to solve the spatial misalignment between the source and target clothing regions. Existing technologies mainly solve this problem by simulating non-rigid clothing deformation through TPS-based methods or flow-based methods. The TPS-based method transfers the clothing onto the human body by estimating the parameters of the thin plate spline transformation. Although the results are generally good, the TPS-based method has limited degrees of freedom, and when the clothing deformation is large or the occlusion is severe, it often leads to fitting failure. The flow-based method operates on fine-grained element pixels in the image space by estimating a dense flow field, and consistently uses a fixed deformation structure when facing different clothing categories, resulting in low transformation efficiency and restricting possible deformations to the local deformation subspace, and the fitting effect is not ideal in complex fitting scenarios. Summary of the Invention
[0003] The present invention provides a virtual fitting method and system based on neural network search, which solves the spatial misalignment caused by the deformation differences of different clothing, and improves the degrees of freedom and success rate of fitting.
[0004] To solve the above technical problems, an embodiment of the present invention provides one, including:
[0005] Obtain a human body shape map and a target clothing picture, and obtain a human body key point map, a human body semantic segmentation map, and a partial semantic segmentation map through a semantic prediction network and an Openpose network;
[0006] Construct a deformation network search space, connect the mask corresponding to the original clothing and the mask corresponding to the target clothing picture, and search for a clothing deformation field corresponding to the clothing type in the deformation network search space;
[0007] Deform the target clothing picture to the target form through the clothing deformation field;
[0008] Construct a fusion network search space, and through the fusion network search space, combine the human body shape map, the human body key point map, the target clothing picture in the target form, and the human body semantic segmentation map to obtain a fitting result.
[0009] Further, the obtaining of the human body shape map and the target clothing picture, and obtaining the human body key point map, the human body semantic segmentation map, and the partial semantic segmentation map through a semantic prediction network and an Openpose network are specifically:
[0010] Input the human body shape diagram, predict the human body key point diagram through the Openpose network, and predict the human body semantic segmentation diagram through the first semantic prediction network;
[0011] Intercept the human body shape diagram according to the human body key point diagram and the human body semantic segmentation diagram to obtain an image containing the human face and head, and convert the human body semantic segmentation diagram into a binary mask;
[0012] Connect the human body key point diagram, the image containing the human face and head, and the binary mask corresponding to the human body semantic segmentation diagram, and obtain a partial semantic segmentation diagram through the second semantic prediction network.
[0013] Further, the masks corresponding to the original clothing and the masks corresponding to the target clothing pictures are connected, and a clothing deformation field corresponding to the clothing type is searched in the deformation network search space, specifically:
[0014] Connect the masks corresponding to the original clothing and the masks corresponding to the target clothing pictures, extract the feature layer through the deformation network search space and predict to obtain a rough deformation flow field, and search for a refined clothing deformation field corresponding to the clothing type through several deformations of the feature layer and several refinements of the rough deformation flow field.
[0015] Further, the target clothing picture is deformed to the target form through the clothing deformation field, specifically:
[0016] According to each pixel in the clothing deformation field, perform x-direction offset calculation and y-direction offset calculation on each pixel point of the target clothing picture to obtain the target clothing picture in the target form.
[0017] Further, through the fusion network search space, combine the human body shape diagram, the human body key point diagram, the target clothing picture in the target form, and the human body semantic segmentation diagram to obtain a fitting result, specifically:
[0018] In the fusion network search space, train all the fusion networks, search for the trained fusion network through the genetic algorithm, and input the human body shape diagram, the human body key point diagram, the target clothing picture in the target form, and the human body semantic segmentation diagram into the searched fusion network to obtain a fitting result.
[0019] Correspondingly, the embodiment of the present invention also provides a virtual fitting system based on neural network search, including an acquisition module, a search module, a deformation module, and a fitting module; wherein,
[0020] The obtaining module is used to obtain a human body shape map and a target clothing picture, and through a semantic prediction network and an Openpose network, obtain a human body key point map, a human body semantic segmentation map, and a partial semantic segmentation map;
[0021] The searching module is used to construct a deformation network search space, connect the mask corresponding to the original clothing and the mask corresponding to the target clothing picture, and search for a clothing deformation field corresponding to the clothing type in the deformation network search space;
[0022] The deforming module is used to deform the target clothing picture to the target form through the clothing deformation field;
[0023] The try-on module is used to construct a fusion network search space, and through the fusion network search space, combine the human body shape map, the human body key point map, the target clothing picture in the target form, and the human body semantic segmentation map to obtain a try-on result.
[0024] Further, the obtaining module obtains a human body shape map and a target clothing picture, and through a semantic prediction network and an Openpose network, obtains a human body key point map, a human body semantic segmentation map, and a partial semantic segmentation map, specifically:
[0025] The obtaining module obtains the human body shape map, predicts the human body key point map through the Openpose network, and predicts the human body semantic segmentation map through the first semantic prediction network;
[0026] According to the human body key point map and the human body semantic segmentation map, intercept the human body shape map to obtain an image including the human face and head, and convert the human body semantic segmentation map into a binary mask;
[0027] According to the connection of the human body key point map, the image including the human face and head, and the binary mask corresponding to the human body semantic segmentation map, obtain a partial semantic segmentation map through the second semantic prediction network.
[0028] Further, the searching module connects the mask corresponding to the original clothing and the mask corresponding to the target clothing picture, and searches for a clothing deformation field corresponding to the clothing type in the deformation network search space, specifically:
[0029] The searching module connects the mask corresponding to the original clothing and the mask corresponding to the target clothing picture, extracts a feature layer through the deformation network search space and predicts to obtain a rough clothing deformation field, and through several deformations of the feature layer and several refinements of the rough clothing deformation field, searches for a refined clothing deformation field corresponding to the clothing type.
[0030] Further, the deformation module deforms the target clothing image to the target form through the clothing deformation field, specifically as follows:
[0031] The deformation module calculates the x-direction offset and y-direction offset for each pixel point of the target clothing image according to each pixel in the clothing deformation field, and obtains the target clothing image in the target form.
[0032] Further, the fitting module searches the fusion network search space, combines the human body shape map, the human body key point map, the target clothing image in the target form, and the human body semantic segmentation map, and obtains the fitting result, specifically as follows:
[0033] The fitting module trains all the fusion networks in the fusion network search space, searches for the trained fusion network through the genetic algorithm, and inputs the human body shape map, the human body key point map, the target clothing image in the target form, and the human body semantic segmentation map into the searched fusion network to obtain the fitting result.
[0034] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0035] The embodiments of the present invention provide a virtual fitting method and system based on neural network search. The method includes: obtaining a human body shape map and a target clothing image, and obtaining a human body key point map, a human body semantic segmentation map, and a partial semantic segmentation map through a semantic prediction network and an Openpose network; constructing a deformation network search space, connecting the mask corresponding to the original clothing and the mask corresponding to the target clothing image, and searching for a clothing deformation field corresponding to the clothing type in the deformation network search space; deforming the target clothing image to the target form through the clothing deformation field; constructing a fusion network search space, and combining the human body shape map, the human body key point map, the target clothing image in the target form, and the human body semantic segmentation map through the fusion network search space to obtain the fitting result. The present invention automatically searches through neural networks and a double-layer hierarchical search space, automatically searches for the optimal clothing deformation field for different clothing types, provides a virtual fitting method and system with high success rate and high degree of freedom in various complex scenarios, can solve the spatial misalignment caused by the deformation differences of different clothing, and has high transformation efficiency and flexibility. Description of the Drawings
[0036] Figure 1 : A flowchart of an embodiment provided for the virtual fitting method based on neural network search of the present invention.
[0037] Figure 2 : A comparison diagram of the clothing deformation effects of an embodiment provided for the virtual fitting method based on neural network search of the present invention and other methods.
[0038] Figure 3 : A comparison chart of the fitting effects of an embodiment provided by the virtual fitting method based on neural network search of the present invention and other methods.
[0039] Figure 4 : A comparison chart of the quality evaluation of an embodiment provided by the virtual fitting method based on neural network search of the present invention and other methods.
[0040] Figure 5 : A schematic structural diagram provided for the virtual fitting system based on neural network search of the present invention. Detailed implementation manners
[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.
[0042] Embodiment 1:
[0043] Please refer to Figure 1 , Figure 1 A virtual fitting method based on neural network search provided for an embodiment of the present invention, including steps S1 to S4; wherein,
[0044] Step S1, obtain a human body shape map and a target clothing picture, and obtain a human body key point map, a human body semantic segmentation map, and a partial semantic segmentation map through a semantic prediction network and an Openpose network. Specifically:
[0045] Step S101, obtain the human body shape map, and predict the human body key point map through the Openpose network. The human body key point map is combined by 18-channel heat maps, and each key point of each part of the human body is regarded as a heat map of a certain channel. In this embodiment, the neighborhood of each feature point is an 11*11 square area centered on the feature point. As an example of this embodiment, for a given human body shape image of a target person, 18 key points (such as elbows, heads, etc.) are detected using a human body key point detector, and each key point is characterized by a 1-channel heat map. A square area of 11*11 is defined centered on the key point, and the pixel values within the area are set to 1, and the pixel values outside the area are set to 0; the heat maps of each channel are spliced at the channel layer to obtain an 18-channel heat map of the human body. This heat map represents all the information of the positions of all key points of the human body, and then the human body semantic segmentation map is predicted through the first semantic prediction network or the human body semantic prediction network. The human body semantic segmentation map separates different regional parts of the human body and represents them with different pixel values.
[0046] Step S102: According to the human key point map and the human semantic segmentation map, intercept the human shape map to obtain an image containing the human face and head, and convert the human semantic segmentation map into a binary mask. Specifically:
[0047] According to the obtained human semantic segmentation map and key point map, intercept the human shape map to obtain an RGB image H containing the human face and head, and at the same time convert the human semantic segmentation map into a binary mask.
[0048] Step S103: Combine the obtained 18-channel human key point map, binary mask, and RGB image H containing the human face and head, and input them into the second semantic prediction network or a partial semantic segmentation prediction network to obtain a partial semantic segmentation map.
[0049] Step S2: Construct a deformation network search space, combine the mask corresponding to the original clothing and the mask corresponding to the target clothing image, and search for a clothing deformation field corresponding to the clothing type in the deformation network search space. Specifically:
[0050] Step S201: Perform a mask operation on the obtained partial human semantic segmentation map, perform a binarization operation on the pixel values of the clothing part area in the segmentation map, set the pixel points where the clothing part appears to 255, and the remaining pixel points to 0, to obtain a target clothing mask. Similarly, perform a mask operation on the input clothing image to obtain an input clothing mask.
[0051] Step S202: Construct a deformation network search space according to different transformation modules and mask encoders. Provide 3 deformation branches in the network layer, each deformation branch contains a different number of deformation modules, and provide 4 operations in the operation layer. Each deformation module can select one of the operations.
[0052] Step S203: Combine the obtained target clothing mask and the original clothing mask together in the channel layer and send them into the mask encoder to extract the feature layer F. Through the deformation modules in the deformation network search space, use the obtained rough flow field to continuously refine the prediction of F, and at the same time guide the deformation of the feature layer F. In the network layer search space, we search for different deformation branches to adapt to the deformation of different types of clothing. In the operation layer search space, we provide a predefined set of convolutions to search for the optimal convolution dimension for different transformation modules, specifically applied to human clothing alignment.
[0053] In this embodiment, the extracted feature layer is fed into five deformation modules. Each deformation module includes at least one network layer search space and at least one operation layer search space. The network layer search space of a module contains three deformation branches, each of which contains 1 / 2 / 3 deformation sub-boxes. The operation layer search space contains four operations, namely 1x1 convolution, depthwise separable 1x1 convolution, 3x3 convolution, and depthwise separable 3x3 convolution. The design of the search space fully considers the situations of different types of clothing. Using neural network search technology (NAS), for each type of clothing, the most suitable deformation module is found.
[0054] In this embodiment, in each deformation module, the original feature and the target feature are combined and fed into the convolutional layer to predict a deformation flow field map F. Then, under the guidance of the flow field map, the original feature R s After deformation, we get And it is fed into the next deformation module to predict the next clothing deformation flow. The specific process is as follows:
[0055]
[0056]
[0057]
[0058] Among them, F is the feature layer, R s Is the original feature layer, R t Is the target feature layer, Represents the connection of the deep layer, C(·) represents the convolution operation, W(·, F) is the deformation part that predicts the deformation flow field F according to the predicted deformation flow field. The predicted deformation flow field F is used to optimize the deformation flow field predicted by the previous module The finally obtained optimized deformation flow field will be directly used to deform the target clothing C.
[0059] Step S3, deform the target clothing picture to the target form through the clothing deformation field;
[0060] In this embodiment, each pixel point of the clothing deformation field obtained in step S2 contains two values, namely the offset in the x direction relative to the original picture position and the offset in the y direction relative to the original picture position. For each pixel value of the target clothing picture, the offset calculation in the x direction and the offset calculation in the y direction are respectively performed to obtain the deformed pixel points, which form the target clothing picture in the target form.
[0061] Step S4, construct a fusion network search space. Through the fusion network search space, combine the human body shape map, the human body key point map, the target clothing picture in the target form, and the human body semantic segmentation map to obtain the try-on result.
[0062] In this embodiment, a fused network search space is constructed. Similarly, a subspace set of the network layer and a subspace set of the operation layer are defined. The subspace set of the network layer includes various different skip connection methods. Two alternative operation sets are provided in the operation layer, which are respectively applied to the upsampling process and the downsampling process. One subset consists of 3x3 convolution, 4x4 convolution, and 5x5 convolution, and the other subset consists of bilinear interpolation + 3x3 convolution and bilinear interpolation + 5x5 convolution; all fused networks (the fused network is a super network) are trained and searched using the genetic algorithm for all fused networks, and the obtained human semantic segmentation map M h , human key points P, and the deformed clothing Part of the human body image I′ is fed into the searched fused network, and the predicted fused mask Mf is used for the deformed clothing and the initial result I c are fused to obtain the final try-on result The specific process is as follows:
[0063]
[0064] Preferably, in the training stage, the try-on result is generated by calculating the L1 loss between the generated try-on result and the real image It and the perceptual loss
[0065]
[0066] where is the target clothing mask, and Mf is the predicted fused mask
[0067] In the network transformation automatic search training stage in this embodiment, the shape of the deformation flow field is constrained by calculating the L1 loss between the deformed clothing mask and the target clothing mask, specifically:
[0068]
[0069] where L mask is the L1 loss, is the target clothing mask, is the deformed clothing mask.
[0070] At the same time, a perceptual loss function is used to constrain the distance between the deformed clothing and the clothing in the real image, specifically:
[0071]
[0072] where φ k (C t ) represents the image C tThe feature map obtained after the k-th layer of the VGG19 network. Specifically, k represents the 5 feature maps of VGG19 in turn, and λ k represents different learning rates for each layer, is the perceptual loss function.
[0073] Finally, the TV loss function is used to prevent the appearance of overly deformed clothing, specifically:
[0074]
[0075] where F x and F y represent the differences between adjacent regions of the x-axis and y-axis of the deformation flow field respectively. The total loss function of the network transformation automatic search module is expressed as follows:
[0076]
[0077] where λ perc , λ TV are hyperparameters, which are set to 0.1 and 0.3 respectively.
[0078] In the embodiment of the present invention, the virtual try-on dataset used is the VITON dataset. This dataset contains 16,235 pairs of front-view human body images and upper body clothing images, and there are three clothing categories, namely short-sleeved, long-sleeved, and vest. In order to increase the diversity of clothing, we follow the crawler protocol and expand the original dataset with pants and skirts. The expanded dataset is divided into a training set and a test set, with 18,312 and 2,487 pairs of image pairs respectively. During the network search process, the test set is further segmented to obtain a validation set containing 950 image pairs, and this validation set is not used for the final evaluation.
[0079] Next, the virtual try-on effect of the present invention will be described in conjunction with the accompanying drawings:
[0080] The virtual try-on effect of the present invention will be qualitatively and quantitatively analyzed below. Among them, the method of the embodiment of the present invention is quantitatively compared with three existing virtual try-on methods: VITON, CP-VTON, and ACGPN. SSIM and FID are used to evaluate the deformation and the final try-on results. The results of the data evaluation indicators show that WAS-VTON proposed in the embodiment of the present invention is superior to other methods, proving that WAS-VTON in the embodiment of the present invention can synthesize more realistic try-on results. The embodiment of the present invention also conducts a user evaluation study. Specifically, on a certain platform, a human body image and an example clothing image are shown to the platform staff, and the staff is asked to select a more realistic and more detailed virtual try-on picture from the virtual try-on pictures generated by two different generation methods. 63.07% of the users think that the result of WAS-VTON is more realistic. In order to compare with the existing methods more fairly, the embodiment of the present invention further tests WAS-VTON on the original virtual try-on data set without expanding the data set. The results show that the method of the embodiment of the present invention still achieves the best performance in all indicators.
[0081] In addition, the embodiment of the present invention also conducts an ablation experiment to verify the effectiveness of the neural network automatic search module. The embodiment of the present invention compares the widely used DARTS search framework and the single-path method applicable to WAS-VTON. Both SSIM and FID indicators show that the single-path method for WAS-VTON in the embodiment of the present invention has better results. The embodiment of the present invention first designs three ablation baselines for the transformation module. Each baseline contains a whole block with a fixed number (1 / 2 / 3), and the operation of each whole block is a 3*3 convolution. Compared with these three baselines, the network searched by the embodiment of the present invention can obtain the highest SSIM and a competitive FID score. By adding dynamic skip connections and thus using the searched network model, the best SSIM (0.8430) and FID (13.83) can be obtained.
[0082] To illustrate the effectiveness of the automatic search method of the neural network deformation module designed by the present invention for clothing deformation, the present invention compares the effect diagrams obtained by deforming the same piece of clothing with different deformation methods. Figure 2 Comparison of the deformation results of WAS-VTON with three deformation methods: VITON, CP-VTON, and ACGPN Figure 3 Comparison diagram of the try-on effects of the present invention and other methods Figure 4This is a comparison chart for the quality assessment of the present invention and other methods. As can be seen from the figure, VITON usually produces unrealistic deformation results, including visual illusions, blurred textures, and unreasonable clothing cuts. Although CPVTON can alleviate these problems to a certain extent, the texture details are usually still blurred. In addition, when the clothing is blocked by certain parts of the body, both VITON and CP-VTON are more likely to fail in deformation. ACGPN performs well in maintaining the invariance of clothing textures and human body shapes. At the same time, guided by the human semantic segmentation map, it can generate relatively reasonable try-on results when part of the human body is blocked. However, in ACGPN, the edge area of the deformed clothing is still blurred, and there are obvious artifacts in the collar area. Moreover, the above three methods all perform poorly in lower body try-on. On the contrary, WAS-VTON can well preserve the texture of the clothing and other parts of the body, thus producing high-quality full-body try-on effects, especially in the collar and clothing contour areas. At the same time, in some cases where part of the body is blocked and complex clothing deformation scenarios, WAS-VTON can also generate good results.
[0083] In summary, a virtual try-on method and system based on neural network search according to the present invention uses the neural network automatic search technology to find the most suitable deformation network for a specific clothing type. After obtaining the deformation flow field, the clothing is deformed, and then the deformed clothing is fused with the original image to finally generate a try-on result image, realizing a virtual try-on method that does not require complex prior knowledge and can achieve good deformation for specific clothing types in cases of partial human body occlusion and partial complex scenarios, and generate try-on images with clear textures.
[0084] Correspondingly, referring to Figure 5 , an embodiment of the present invention further provides a virtual try-on system based on neural network search, including an acquisition module 101, a search module 102, a deformation module 103, and a try-on module 104; wherein,
[0085] The acquisition module 101 is used to acquire a human body shape map and a target clothing picture, and obtain a human key point map, a human semantic segmentation map, and a partial semantic segmentation map through a semantic prediction network and an Openpose network;
[0086] The search module 102 is used to construct a deformation network search space, connect the mask corresponding to the original clothing and the mask corresponding to the target clothing picture, and search for a clothing deformation field corresponding to the clothing type in the deformation network search space;
[0087] The deformation module 103 is used to deform the target clothing picture to the target form through the clothing deformation field;
[0088] The fitting module 104 is used to construct a fusion network search space. Through the fusion network search space, combining the human body shape map, the human body key point map, the target clothing picture in the target form, and the human body semantic segmentation map, a fitting result is obtained.
[0089] In this embodiment, the acquisition module 101 acquires the human body shape map and the target clothing picture, and obtains the human body key point map, the human body semantic segmentation map, and the partial semantic segmentation map through the semantic prediction network and the Openpose network. Specifically:
[0090] The acquisition module 101 acquires the human body shape map, predicts the human body key point map through the Openpose network, and predicts the human body semantic segmentation map through the first semantic prediction network;
[0091] According to the human body key point map and the human body semantic segmentation map, the human body shape map is intercepted to obtain an image including the human face and head, and the human body semantic segmentation map is converted into a binary mask;
[0092] According to the connection of the human body key point map, the image including the human face and head, and the binary mask corresponding to the human body semantic segmentation map, a partial semantic segmentation map is obtained through the second semantic prediction network.
[0093] In this embodiment, the search module 102 connects the mask corresponding to the original clothing and the mask corresponding to the target clothing picture, and searches for a clothing deformation field corresponding to the clothing type in the deformation network search space. Specifically:
[0094] The search module 102 connects the mask corresponding to the original clothing and the mask corresponding to the target clothing picture, extracts the feature layer through the deformation network search space and predicts to obtain a rough clothing deformation field, and searches for a refined clothing deformation field corresponding to the clothing type through several deformations of the feature layer and several refinements of the rough clothing deformation field.
[0095] In this embodiment, the deformation module 103 deforms the target clothing picture to the target form through the clothing deformation field. Specifically: the deformation module 103 calculates the x-direction offset and the y-direction offset for each pixel point of the target clothing picture according to each pixel in the clothing deformation field, and obtains the target clothing picture in the target form.
[0096] In this embodiment, the fitting module 104 obtains the fitting result through the fusion network search space, combining the human body shape map, the human body key point map, the target clothing picture in the target form, and the human body semantic segmentation map. Specifically:
[0097] The fitting module 104 trains all the fusion networks in the fusion network search space, searches for the trained fusion network through a genetic algorithm, and inputs the human body shape map, the human body key point map, the target clothing picture in the target form, and the human body semantic segmentation map into the searched fusion network to obtain a fitting result.
[0098] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0099] The embodiments of the present invention provide a virtual fitting method and system based on neural network search. The method includes: obtaining a human body shape map and a target clothing picture, and obtaining a human body key point map, a human body semantic segmentation map, and a partial semantic segmentation map through a semantic prediction network and an Openpose network; constructing a deformation network search space, connecting the mask corresponding to the original clothing and the mask corresponding to the target clothing picture, and searching for a clothing deformation field corresponding to the clothing type in the deformation network search space; deforming the target clothing picture to the target form through the clothing deformation field; constructing a fusion network search space, and combining the human body shape map, the human body key point map, the target clothing picture in the target form, and the human body semantic segmentation map through the fusion network search space to obtain a fitting result. The present invention automatically searches through a neural network and a double-layer hierarchical search space, automatically searches for the optimal clothing deformation field for different clothing types, provides a virtual fitting method and system with high success rate and high degree of freedom in various complex scenarios, can solve the spatial misalignment caused by the deformation differences of different clothing, and has high transformation efficiency and flexibility.
[0100] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. In particular, for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A virtual try-on method based on neural network search, characterized in that, Including: Obtain a human body shape map and a target clothing image, and through a semantic prediction network and an Openpose network, obtain a human body key point map, a human body semantic segmentation map, and a partial semantic segmentation map; Construct a deformation network search space, connect the mask corresponding to the original clothing and the mask corresponding to the target clothing image, and search for a clothing deformation field corresponding to the clothing type in the deformation network search space; Deform the target clothing image to the target form through the clothing deformation field; Construct a fusion network search space, and through the fusion network search space, combine the human body shape map, the human body key point map, the target clothing image in the target form, and the human body semantic segmentation map to obtain a try-on result; The connecting the mask corresponding to the original clothing and the mask corresponding to the target clothing image, and searching for a clothing deformation field corresponding to the clothing type in the deformation network search space is specifically: Connect the mask corresponding to the original clothing and the mask corresponding to the target clothing image, extract a feature layer through the deformation network search space and predict to obtain a rough deformation flow field, and through several deformations of the feature layer and several refinements of the rough deformation flow field, search for a refined clothing deformation field corresponding to the clothing type; The deforming the target clothing image to the target form through the clothing deformation field is specifically: According to each pixel in the clothing deformation field, perform an x-direction offset calculation and a y-direction offset calculation on each pixel point of the target clothing image to obtain the target clothing image in the target form; The obtaining the try-on result by combining the human body shape map, the human body key point map, the target clothing image in the target form, and the human body semantic segmentation map through the fusion network search space is specifically: In the fusion network search space, train all the fusion networks, search for the trained fusion network through a genetic algorithm, and input the human body shape map, the human body key point map, the target clothing image in the target form, and the human body semantic segmentation map into the searched fusion network to obtain the try-on result.
2. The virtual fitting method based on neural network search according to claim 1, characterized in that, The obtaining the human body shape map and the target clothing image, and obtaining the human body key point map, the human body semantic segmentation map, and the partial semantic segmentation map through the semantic prediction network and the Openpose network is specifically: Obtain the human body shape map, predict the human body key point map through the Openpose network, and predict the human body semantic segmentation map through the first semantic prediction network; According to the human body key point map and the human body semantic segmentation map, intercept the human body shape map to obtain an image including the human face and head, and convert the human body semantic segmentation map into a binary mask; According to the connection of the human body key point map, the image including the human face and head, and the binary mask corresponding to the human body semantic segmentation map, obtain a partial semantic segmentation map through the second semantic prediction network.
3. A virtual fitting system based on neural network search, characterized in that, Including an acquisition module, a search module, a deformation module, and a try-on module; wherein, The acquisition module is used to obtain a human body shape map and a target clothing image, and through a semantic prediction network and an Openpose network, obtain a human body key point map, a human body semantic segmentation map, and a partial semantic segmentation map; The search module is used to construct a deformation network search space, connect the mask corresponding to the original clothing and the mask corresponding to the target clothing picture, and search for a clothing deformation field corresponding to the clothing type in the deformation network search space; The deformation module is used to deform the target clothing picture to the target form through the clothing deformation field; The try-on module is used to construct a fusion network search space, and through the fusion network search space, combine the human body shape map, the human body key point map, the target clothing picture in the target form, and the human body semantic segmentation map to obtain a try-on result; The search module connects the mask corresponding to the original clothing and the mask corresponding to the target clothing picture, and searches for a clothing deformation field corresponding to the clothing type in the deformation network search space, specifically: The search module connects the mask corresponding to the original clothing and the mask corresponding to the target clothing picture, extracts a feature layer through the deformation network search space and predicts a rough clothing deformation field, and through several deformations of the feature layer and several refinements of the rough clothing deformation field, searches for a refined clothing deformation field corresponding to the clothing type; The deformation module deforms the target clothing picture to the target form through the clothing deformation field, specifically: The deformation module calculates the x-direction offset and y-direction offset for each pixel point of the target clothing picture according to each pixel in the clothing deformation field, and obtains the target clothing picture in the target form; The try-on module combines the human body shape map, the human body key point map, the target clothing picture in the target form, and the human body semantic segmentation map through the fusion network search space to obtain a try-on result, specifically: The try-on module trains all fusion networks in the fusion network search space, searches for a trained fusion network through a genetic algorithm, and inputs the human body shape map, the human body key point map, the target clothing picture in the target form, and the human body semantic segmentation map into the searched fusion network to obtain a try-on result.
4. The virtual fitting system based on neural network search according to claim 3, wherein, The acquisition module acquires the human body shape map and the target clothing picture, and obtains the human body key point map, the human body semantic segmentation map, and a partial semantic segmentation map through a semantic prediction network and an Openpose network, specifically: The acquisition module acquires the human body shape map, predicts the human body key point map through the Openpose network, and predicts the human body semantic segmentation map through the first semantic prediction network; According to the human body key point map and the human body semantic segmentation map, the human body shape map is intercepted to obtain an image including the human face and head, and the human body semantic segmentation map is converted into a binary mask; According to the connection of the human body key point map, the image including the human face and head, and the binary mask corresponding to the human body semantic segmentation map, a partial semantic segmentation map is obtained through the second semantic prediction network.
Citation Information
Patent Citations
Virtual try-on method and device based on artificial intelligence, server and storage medium
CN111784845A
Human body posture transformation method and system for virtual fitting of clothes
CN113297944A