Image Generation Method, Apparatus, Server, and Medium
By predicting, segmenting and transforming the model images, and combining the backbone and branch networks to supplement the details of the clothing, a virtual try-on image with complete details and deformation fit is generated, solving the problem of missing clothing details in the virtual try-on and improving the try-on effect.
Patent Information
- Application Number
- CN202210598983.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-30
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-05-30
AI Technical Summary
The detailed information (such as logo or pattern) of the clothing to be tried on in the virtual try on clothing generated in the prior art is seriously missing, resulting in unsatisfactory trial on test.
By inputting the model image into the key point prediction model, obtaining the human body key point map, performing clothing segmentation and instance segmentation, combining spatial transformation and correction network, using the backbone and branch network of the clothing try-on model, supplementing the clothing details information, and generating a complete detailed virtual try-on image.
The virtual trial-on image is achieved with little or no details of the clothing, and the deformation degree fits the model's human posture, improving the trial-on experience.
Smart Images

Figure CN114998481B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and more specifically, to an image generation method, apparatus, server, and medium. Background Art
[0002] Virtual Try-on means: Given a clothing image including the clothing to be tried on and a model image including a model, a target photo of the model wearing the clothing to be tried on is generated based on the clothing image and the model image, so as to achieve the purpose of the model virtually trying on the clothing to be tried on.
[0003] In the existing related technologies, the detailed information of the clothing to be tried on in the generated target photo is severely missing. For example, the detailed information such as the logo or pattern of the clothing to be tried on is severely missing, resulting in a poor try-on result. Therefore, how to generate a target photo with less or even no missing detailed information of the clothing to be tried on is a technical problem that needs to be solved by those skilled in the art. Summary of the Invention
[0004] In view of this, the present application provides an image generation method, apparatus, server, and medium to solve the problem that the detailed information such as the logo or pattern of the clothing to be tried on is severely missing, resulting in a poor try-on result.
[0005] To achieve the above object, the present application provides the following technical solutions:
[0006] According to the first aspect of the embodiments of the present disclosure, an image generation method is provided, including:
[0007] Inputting the model image into a pre-constructed key point prediction model to obtain a human key point map output by the key point prediction model, where the human key point map includes the key points of the model in the model image;
[0008] Performing deformation processing on the clothing to be tried on in the clothing image including the clothing to be tried on according to the model image and the human key point map to obtain a clothing segmentation map;
[0009] Performing instance segmentation processing on the model image according to the human key point map to obtain a candidate human instance segmentation map, where the candidate human instance segmentation map includes a try-on instance area for wearing the clothing to be tried on;
[0010] Matching the clothing to be tried on included in the clothing segmentation map to the try-on instance area in the candidate human instance segmentation map to generate a human instance segmentation map;
[0011] Inputting the clothing image and the clothing segmentation map into a pre-constructed spatial transformation network to obtain a first clothing image output by the spatial transformation network;
[0012] Input the first clothing image and the clothing segmentation map into a calibration network to obtain a second clothing image output by the calibration network;
[0013] Input the human body instance segmentation map, the clothing segmentation map, the second clothing image, and a retained image into the backbone network of a pre-constructed clothing try-on model, where the retained image is the image in the model image except for the try-on instance area;
[0014] Input the first clothing image into a branch network of the clothing try-on model, where the branch network is used to send the detail information of the clothing to be tried on obtained from the first clothing image to the backbone network;
[0015] Output a target image through the clothing try-on model, where the target image includes the model wearing the clothing to be tried on.
[0016] According to a second aspect of the embodiments of the present disclosure, there is provided an image generation device, including:
[0017] A first acquisition module, configured to input a model image into a pre-constructed key point prediction model to obtain a human body key point map output by the key point prediction model, where the human body key point map includes the key points of the model in the model image;
[0018] A first processing module, configured to perform a deformation process on the clothing to be tried on in a clothing image including the clothing to be tried on according to the model image and the human body key point map to obtain a clothing segmentation map;
[0019] A second processing module, configured to perform instance segmentation processing on the model image according to the human body key point map to obtain a candidate human body instance segmentation map, where the candidate human body instance segmentation map includes a try-on instance area for wearing the clothing to be tried on;
[0020] A generation module, configured to match the clothing to be tried on included in the clothing segmentation map to the try-on instance area in the candidate human body instance segmentation map to generate a human body instance segmentation map;
[0021] A second acquisition module, configured to input the clothing image and the clothing segmentation map into a pre-constructed spatial transformation network to obtain a first clothing image output by the spatial transformation network;
[0022] A third acquisition module, configured to input the first clothing image and the clothing segmentation map into a calibration network to obtain a second clothing image output by the calibration network;
[0023] A first input module, configured to input the human instance segmentation map, the clothing segmentation map, the second clothing image, and the remaining image into a backbone network of a pre-constructed clothing try-on model, where the remaining image is the image in the model image except for the try-on instance area;
[0024] A second input module, configured to input the first clothing image into a branch network of the clothing try-on model, where the branch network is configured to send the detail information of the clothing to be tried on obtained from the first clothing image to the backbone network;
[0025] A fourth acquisition module, configured to output a target image through the clothing try-on model, where the target image includes the model wearing the clothing to be tried on.
[0026] According to a third aspect of the embodiments of the present disclosure, there is provided a server, including:
[0027] A processor;
[0028] A memory for storing executable instructions of the processor;
[0029] Wherein, the processor is configured to execute the instructions to implement the image generation method as described in the first aspect.
[0030] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the storage medium are executed by a processor of a server, enabling the server to execute the image generation method as described in the first aspect.
[0031] As can be seen from the above technical solutions, in the image generation method provided by this application, a model image is input into a pre-constructed key point prediction model to obtain a human key point map output by the key point prediction model; the to-be-tried-on clothing in the clothing image including the to-be-tried-on clothing is deformed according to the model image and the human key point map to obtain a clothing segmentation map; the model image is subjected to instance segmentation processing according to the human key point map to obtain a candidate human instance segmentation map, and the candidate human instance segmentation map includes a try-on instance area for wearing the to-be-tried-on clothing; the to-be-tried-on clothing included in the clothing segmentation map is matched to the try-on instance area in the candidate human instance segmentation map to generate a human instance segmentation map; the clothing image and the clothing segmentation map are input into a pre-constructed spatial transformation network to obtain a first clothing image output by the spatial transformation network; the first clothing image and the clothing segmentation map are input into a correction network to obtain a second clothing image output by the correction network; the human instance segmentation map, the clothing segmentation map, the second clothing image, and a retained image are input into the backbone network of a pre-constructed clothing try-on model, and the retained image is the image in the model image except for the try-on instance area; the first clothing image is input into the branch network of the clothing try-on model, and the branch network is used to send the detailed information of the to-be-tried-on clothing obtained from the first clothing image to the backbone network; a target image is output through the clothing try-on model, and the target image includes the model wearing the to-be-tried-on clothing. Since the deformation degree of the second clothing image is better, that is, it fits the human body posture of the model in the model image more, the model in the target image obtained through the backbone network wears the to-be-tried-on clothing more fittingly. Since the first clothing image contains more detailed information of the to-be-tried-on clothing, the branch network can obtain a lot of detailed information of the to-be-tried-on clothing, and the branch network sends the obtained detailed information of the to-be-tried-on clothing to the backbone network, so that the backbone network can supplement the second clothing image based on the detailed information output by the branch network, thereby achieving the purpose that the deformation degree of the to-be-tried-on clothing in the target image fits the human body posture of the model, and the lack of detailed information of the to-be-tried-on clothing is less or even non-existent. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0033] Figure 1 It is a structural diagram of an implementation manner of the hardware architecture related to the embodiments of the present application;
[0034] Figure 2 It is a flowchart of the image generation method provided by the embodiment of the present application;
[0035] Figure 3 It is a schematic diagram of the human key point map output by the key point prediction model provided by the embodiment of the present application;
[0036] Figure 4 It is a schematic diagram of the clothing image, candidate clothing segmentation map, and clothing segmentation map provided by the embodiment of the present application;
[0037] Figure 5 It is a schematic diagram of the model image and the candidate human instance segmentation map provided by the embodiment of the present application;
[0038] Figure 6 It is a schematic diagram of the human instance segmentation map provided by the embodiment of the present application;
[0039] Figure 7 It is a schematic diagram of the first clothing image obtained by the spatial transformation network provided by the embodiment of the present application;
[0040] Figure 8 It is a schematic diagram of the second clothing image output by the calibration network provided by the embodiment of the present application;
[0041] Figure 9 It is a structural diagram of an implementation manner of the clothing try-on model provided by the embodiment of the present application;
[0042] Figure 10 It is a network structural diagram of another implementation manner of the clothing try-on model provided by the embodiment of the present application;
[0043] Figure 11 It is a network structural diagram of another implementation manner of the clothing try-on model provided by the embodiment of the present application;
[0044] Figure 12 It is a further illustration of a specific embodiment of an example target image;
[0045] Figure 13 It is a limb connection topology map provided by the embodiment of the present application;
[0046] Figure 14 It is a labeled clothing segmentation map provided by the embodiment of the present application;
[0047] Figure 15 It is a structural diagram of the image generation device provided by the embodiment of the present application;
[0048] Figure 16 It is a block diagram of a server shown according to an exemplary embodiment. Detailed implementation manners
[0049] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.
[0050] The embodiments of the present application provide an image generation method, device, server, medium, and product. Before introducing the technical solutions provided by the embodiments of the present application, the hardware architecture involved in the embodiments of the present application will be described first.
[0051] As Figure 1 shown, it is a structural diagram of an implementation manner of the hardware architecture involved in the embodiments of the present application. The hardware architecture includes: an electronic device 11 and a server 12.
[0052] Exemplarily, the electronic device 11 can be any electronic product that can perform human-computer interaction with the user through one or more ways such as a keyboard, touchpad, touch screen, remote control, voice interaction, or handwriting device. For example, a mobile phone, a laptop computer, a tablet computer, a handheld computer, a personal computer, a wearable device, a smart TV, a PAD, etc.
[0053] It should be noted that Figure 1 is only an example, and there can be multiple types of electronic devices, not limited to Figure 1 the laptop computer, smart phone, and personal computer in
[0054] Exemplarily, the server 12 can be a single server, a server cluster composed of multiple servers, or a cloud computing server center. The server 12 can include a processor, a memory, and a network interface, etc.
[0055] Exemplarily, the electronic device 11 can establish a connection and communicate with the server 12 through a wireless communication network.
[0056] Exemplarily, the user can determine a clothing image including the clothing to be tried on through the electronic device 11.
[0057] Exemplarily, the user can take a photo through the electronic device 11 to obtain a model image including a model.
[0058] Exemplarily, the user can upload the model image and the clothing image through the electronic device 11. Thus, the electronic device 11 can send the model image and the clothing image to the server 12.
[0059] Exemplarily, the server 12 can execute the image generation method provided by the embodiments of the present application.
[0060] The embodiments of the present application can be applied to shopping application scenarios or movie-watching application scenarios. For example, in a shopping application scenario, one can virtually try on the clothes they are going to buy; in a movie-watching application scenario, one can virtually try on the clothes of the characters in a TV drama or movie.
[0061] Those skilled in the art should understand that the above-mentioned electronic devices and servers are only examples. Other existing or future possible electronic devices or servers that can be applied to the present disclosure should also be included within the protection scope of the present disclosure and are hereby incorporated herein by reference.
[0062] The image generation method provided by the embodiments of the present application will be described below in conjunction with the above-mentioned hardware architecture.
[0063] As Figure 2 shown, it is a flowchart of the image generation method provided by the embodiments of the present application. This method can be applied to a server as Figure 1 shown. The method includes the following steps S21 to S29 in the implementation process.
[0064] Step S21: Input the model image into a pre-constructed key point prediction model to obtain a human key point map output by the key point prediction model. The human key point map includes the key points of the model in the model image.
[0065] The key point prediction model is pre-constructed. For example, a neural network can be used. A large number of sample images are used as the training sample images of the neural network, and the corresponding annotated images of the sample images are used as the training targets to train the key point prediction model.
[0066] Optionally, the neural network can be a fully connected neural network (such as an MLP network, where MLP represents Multi-layer Perceptron, which means a multi-layer perceptron), or other forms of neural networks (such as a convolutional neural network, a deep neural network, etc.).
[0067] Exemplarily, in the process of training the key point prediction model, the sample image can be an image containing a human body, and the corresponding annotated image of the sample image is an image with the key points of the human body already annotated.
[0068] As Figure 3 shown, it is a schematic diagram of the key point prediction model outputting a human key point map provided by the embodiments of the present application.
[0069] Exemplarily, the key points may include: head key point, left shoulder key point, right shoulder key point, neck key point, left elbow key point, right elbow key point, left wrist key point, right wrist key point, left palm key point, right palm key point, left hip key point, right hip key point, left knee key point, right knee key point, left ankle key point, right ankle key point, left toe key point, and right toe key point.
[0070] Exemplarily, the number of key points may also be less than the above 18 key points, or more than the above 18 key points.
[0071] Take the model image 11 as the input of the key point prediction model 10; after data processing by the key point prediction model 10, the predicted positions of the key points of the model in the model image 11 are obtained. Exemplarily, the human key point map output by the key point prediction model 10 may be as shown in the image 12.
[0072] Step S22: Perform deformation processing on the to-be-tried-on clothing in the clothing image including the to-be-tried-on clothing according to the model image and the human key point map to obtain a clothing segmentation map.
[0073] In an alternative implementation, human body features can be extracted from the model image and the human key point map; the correlation between the clothing features and the human body features in the clothing image can be calculated to obtain a correlation tensor representing the correlation between the human body features and the clothing features. A candidate clothing segmentation map can be obtained from the clothing image based on the instance segmentation method; then, based on the correlation tensor and the candidate clothing segmentation map, the deformed clothing segmentation map can be obtained.
[0074] As Figure 4 shown, it is a schematic diagram of the clothing image, the candidate clothing segmentation map, and the clothing segmentation map provided by the embodiment of the present application.
[0075] As Figure 4 shown, assuming the clothing image is the clothing image 41, after performing instance segmentation on the clothing image 41, a candidate clothing segmentation map 42 can be obtained; then, combined with the pose of the model, a deformed clothing segmentation map 43 can be obtained.
[0076] Step S23: Perform instance segmentation processing on the model image according to the human key point map to obtain a candidate human instance segmentation map, and the candidate human instance segmentation map includes a fitting instance area for wearing the to-be-tried-on clothing.
[0077] Instance segmentation refers to automatically framing different instances from an image using an object detection method, and then performing pixel-by-pixel marking in different instance areas using a semantic segmentation method.
[0078] Exemplarily, the fitting instance area of the model can be determined as the upper body and / or the lower body in combination with the clothing image. If the clothing to be tried on in the clothing image is a short-sleeved shirt, the fitting instance area of the model is the upper body; if the clothing to be tried on in the clothing image is a pair of trousers, a skirt or shorts, the fitting instance area of the model is the lower body; if the clothing to be tried on in the clothing image is a dress, the fitting instance area of the model includes the upper body and the lower body.
[0079] For instance segmentation of the model image, in combination with the human key point map, the hair instance area where the hair of the model is located, the face instance area where the face is located, the left arm instance area where the left arm is located, the right arm instance area where the right arm is located, the fitting instance area to be tried on, the left leg instance area where the left leg is located, the right leg instance area where the right leg is located, the background instance area, and the clothing instance area of the clothing worn by the model can be framed from the model image by using the object detection method, that is, the hair instance area, the face instance area, the left arm instance area, the right arm instance area, the clothing instance area, the left leg instance area, the right leg instance area, and the background instance area of the model are all different instance areas.
[0080] Then, pixel labeling is performed on different instance areas through the semantic segmentation method to obtain the class labels of different instance areas, that is, the class labels of different instance areas are different.
[0081] In an optional implementation manner, if the fitting instance area of the model is the upper body, indicating that it is necessary to clearly label the clothing instance area where the clothing worn on the upper body of the model is located, the left arm instance area, the right arm instance area, the clothing instance area where the clothing worn by the model is located, the hair instance area, the face instance area, and the lower body instance area are respectively used as different instance areas; the lower body instance area includes: the left leg instance area and the right leg instance area, that is, the left leg instance area and the right leg instance area are the same instance area; the upper body instance area includes: the left arm instance area, the right arm instance area, and the clothing instance area where the clothing worn by the model is located.
[0082] In an optional implementation manner, if the fitting instance area of the model is the lower body, indicating that it is necessary to clearly label the clothing instance area where the clothing worn on the lower body of the model is located, the upper body instance area, the hair instance area, the face instance area, the left leg instance area, and the right leg instance area are respectively used as different instance areas; the upper body instance area includes: the left arm instance area and the right arm instance area, that is, the left arm instance area and the right arm instance area are the same instance area, and the lower body instance area includes: the left leg instance area, the clothing instance area where the clothing worn by the model is located, and the right leg instance area.
[0083] In an alternative implementation, if the fitting instance area of the model is the upper body and the lower body, the hair instance area, the face instance area, the left arm instance area, the right arm instance area, the clothing instance area, the left leg instance area, and the right leg instance area are all regarded as different instance areas.
[0084] To enable those skilled in the art to better understand the candidate human body instance segmentation map in the embodiments of the present application, an illustration is provided below. As Figure 5 shown, it is a schematic diagram of the model image and the candidate human body instance segmentation map provided by the embodiments of the present application.
[0085] Assume that the model image is as shown in Model Image 11, then the human body instance segmentation map is as Figure 5 shown.
[0086] As Figure 5 shown in the candidate human body instance segmentation map, the hair instance area, the face instance area, the left arm instance area, the right arm instance area, the clothing instance area, the background instance area, the left leg instance area, and the right leg instance area of the model are filled with different patterns respectively. Figure 5 Among them, to distinguish each instance area, the hair instance area is filled with multiple squares; the face instance area is filled with vertical lines, the right arm instance area is filled with diagonal lines, the left arm instance area is filled with rhombuses, the clothing instance area is filled with black, the right leg instance area is filled with small dots, the left leg instance area is filled with horizontal lines, and the background instance area is filled with white.
[0087] In the real human body instance segmentation map, the colors of different instance areas are different. Since colored images cannot be shown in the drawings, Figure 5 therefore, the above method is used to distinguish different instance areas.
[0088] In an alternative implementation, a human parser (LIP) can be used to obtain the candidate human body instance segmentation map. Exemplarily, LIP can estimate a reasonable human body analysis of the target image based on the approximate shapes of the body, face, hair, clothing, and target pose, and can effectively guide the synthesis of the precise areas of human body parts.
[0089] Step S24: Match the to-be-tried-on clothing included in the clothing segmentation map to the fitting instance area in the candidate human body instance segmentation map to generate a human body instance segmentation map.
[0090] To enable those skilled in the art to better understand Step S24, an example where the to-be-tried-on clothing is a long-sleeved one is used for illustration below. Assume that the candidate human body instance segmentation map is as Figure 5 shown, then the obtained human body instance segmentation map is as Figure 6 shown. From Figure 6It can be seen that the clothes to be tried on included in the clothing segmentation map have been matched to the fitting instance area in the candidate human body instance segmentation map.
[0091] Step S25: Input the clothing image and the clothing segmentation map into a pre-constructed spatial transformer network to obtain a first clothing image output by the spatial transformer network.
[0092] As a special network module, the Spatial Transformer Network (STN) can be embedded into a certain layer of the network, enabling the network to support spatial transformations (such as affine transformation and projection transformation), and providing properties such as rotational invariance and translational invariance for the network.
[0093] Exemplarily, the first clothing image obtained through the spatial transformer network belongs to an RGB image with three RGB channels.
[0094] As Figure 7 shown, it is a schematic diagram of obtaining the first clothing image by the spatial transformer network provided in the embodiment of the present application.
[0095] As Figure 7 shown, the details of the clothes to be tried on included in the first clothing image 44 obtained through the spatial transformer network are relatively comprehensive. For example, the patterns (such as patterns and LOGOs) are very comprehensive. However, the deformed shape of the clothes in the first clothing image 44 is not very ideal, that is, it does not fit well with the human body posture of the model in the model image. Based on this, the first clothing image 44 and the clothing segmentation map 43 are input into a correction network to obtain a second clothing image output by the correction network.
[0096] Step S26: Input the first clothing image and the clothing segmentation map into a correction network to obtain a second clothing image output by the correction network.
[0097] As Figure 8 shown, it is a schematic diagram of the correction network outputting the second clothing image provided in the embodiment of the present application.
[0098] As Figure 8 shown, the correction network is mainly used to correct the deformed shape of the first clothing image 44 based on the clothing segmentation map, so that the degree of deformation of the clothes in the obtained second clothing image 45 fits the posture of the model better.
[0099] However, although the degree of deformation of the clothes in the second clothing image 45 obtained through the correction network fits the posture of the model better, the patterns (such as patterns and LOGOs) of the clothes in the second clothing image 45 are severely missing.
[0100] Step S27: Input the human instance segmentation map, the clothing segmentation map, the second clothing image, and the remaining image into the backbone network of the pre-constructed clothing try-on model. The remaining image is the image in the model image except for the try-on instance area.
[0101] The remaining image belongs to an image in the RGB channel.
[0102] Step S28: Input the first clothing image into the branch network of the clothing try-on model. The branch network is used to send the detailed information of the clothing to be tried on obtained from the first clothing image to the backbone network.
[0103] Step S29: Output a target image through the clothing try-on model. The target image includes the model wearing the clothing to be tried on.
[0104] Exemplarily, the target image belongs to an image in the RGB channel.
[0105] As Figure 9 shown, it is a structural diagram of an implementation manner of the clothing try-on model provided by the embodiment of the present application. The clothing try-on model includes: a backbone network 91 and a branch network 92.
[0106] Exemplarily, the backbone network 91 includes an encoding module and a decoding module. Figure 9 It is described by taking the encoding module including 4 encoding layers and the decoding module including 4 decoding layers as an example.
[0107] As Figure 9 shown, the encoding layers are connected by solid thin arrows, and the decoding layers are connected by solid thick arrows. The encoding layers in the encoding module gradually extract the features of the human instance segmentation map, the clothing segmentation map, and the second clothing image. The decoding layers of the decoding module gradually amplify the features and restore the image to the original size. The dotted thin arrows are the skip connection parts. By directly connecting the shallow features of the encoding layers to the subsequent decoding layers, the network retains more input information.
[0108] Since the deformation degree of the second clothing image is better, the target image obtained through the backbone network shows that the model wearing the clothing to be tried on is more fitting.
[0109] Exemplarily, the branch network 92 includes an encoding module and a decoding module. Figure 9 It is described by taking the encoding module including 4 encoding layers and the decoding module including 4 decoding layers as an example.
[0110] The encoding layer in the encoding module of the branch network 92 gradually extracts the features of the first clothing image. Since the first clothing image contains more and more complete detail information of the clothing to be tried on, the branch network can obtain a lot of detail information of the clothing to be tried on. For example, pattern information (such as LOGO and patterns). The branch network sends the obtained detail information of the clothing to be tried on to the decoding layer of the backbone network through the decoding layer, so that the decoding layer of the backbone network can supplement the second clothing image based on the detail information output by the decoding layer of the branch network, so that the clothing to be tried on worn by the model in the target image fits well, the detail information in the clothing to be tried on is more comprehensive, that is, the lost detail information of the clothing to be tried on is less, or even no detail information is lost.
[0111] Exemplarily, the decoding module of the backbone network includes a first number of layer decoding layers, the decoding module of the branch network includes the first number of layer decoding layers, the i-th layer decoding layer in the decoding module of the branch network is connected to the i-th layer decoding layer in the decoding module of the backbone network; the decoding layer of the encoding module of the branch network is not connected to the decoding layer of the decoding module of the branch network, and the encoding layer of the encoding module of the backbone network is connected to the decoding layer of the decoding module of the backbone network.
[0112] In the image generation method provided by the embodiments of the present application, a model image is input into a pre-constructed key point prediction model to obtain a human key point map output by the key point prediction model; the to-be-tried-on clothing in the clothing image including the to-be-tried-on clothing is deformed according to the model image and the human key point map to obtain a clothing segmentation map; the model image is subjected to instance segmentation processing according to the human key point map to obtain a candidate human instance segmentation map, and the candidate human instance segmentation map includes a try-on instance area for wearing the to-be-tried-on clothing; the to-be-tried-on clothing included in the clothing segmentation map is matched to the try-on instance area in the candidate human instance segmentation map to generate a human instance segmentation map; the clothing image and the clothing segmentation map are input into a pre-constructed spatial transformation network to obtain a first clothing image output by the spatial transformation network; the first clothing image and the clothing segmentation map are input into a correction network to obtain a second clothing image output by the correction network; the human instance segmentation map, the clothing segmentation map, the second clothing image, and a retained image are input into the backbone network of a pre-constructed clothing try-on model, and the retained image is the image in the model image except for the try-on instance area; the first clothing image is input into the branch network of the clothing try-on model, and the branch network is used to send the detailed information of the to-be-tried-on clothing obtained from the first clothing image to the backbone network; a target image is output through the clothing try-on model, and the target image includes the model wearing the to-be-tried-on clothing. Since the deformation degree of the second clothing image is better, that is, it fits the human pose of the model in the model image more, the model in the target image obtained through the backbone network wears the to-be-tried-on clothing more fittingly. Since the first clothing image contains more detailed information of the to-be-tried-on clothing, the branch network can obtain a lot of detailed information of the to-be-tried-on clothing, and the branch network sends the obtained detailed information of the to-be-tried-on clothing to the backbone network, so that the backbone network can supplement the second clothing image based on the detailed information output by the branch network, thereby achieving the purpose that the deformation degree of the to-be-tried-on clothing in the target image fits the human pose of the model, and the lack of detailed information of the to-be-tried-on clothing is less or even non-existent.
[0113] In an alternative implementation, the network structures of the backbone network and the branch network included in the clothing try-on model may be the same or different.
[0114] The network structure of the clothing try-on model will be described below with a specific example.
[0115] As Figure 10 shown, it is a network structure diagram of another implementation of the clothing try-on model provided by the embodiments of the present application.
[0116] From Figure 10It can be seen that the network structures of the backbone network 101 and the branch network 102 are the same.
[0117] Among them, the encoding modules of the backbone network 101 and the branch network 102 both include 4 encoding layers, and one downsampling layer and one convolutional layer form an encoding layer; the decoding modules of the backbone network 101 and the branch network 102 both include 4 decoding layers, and one upsampling layer and one convolutional layer form a decoding layer.
[0118] Since the decoding module in the branch network is connected to the encoding module, and the decoding module and the encoding module of the backbone network are connected, the network operation amount of the clothing try-on model is greatly increased. In order to reduce the network operation amount, the present application provides another network structure of the clothing try-on model.
[0119] As Figure 11 shown, it is the network structure diagram of another implementation manner of the clothing try-on model provided by the embodiment of the present application.
[0120] Figure 11 Among them, the encoding modules of the backbone network 111 and the branch network 112 both include 4 encoding layers, and one downsampling layer and one convolutional layer form an encoding layer; the decoding modules of the backbone network 111 and the branch network 112 both include 4 decoding layers, and one upsampling layer and one convolutional layer form a decoding layer.
[0121] The network structures of the backbone network 111 and the backbone network 101 are the same, and the network structures of the branch network 112 and the branch network 102 are different. Among them, there is no skip layer connection between the decoding layer and the encoding layer in the branch network 112, that is, the shallow features obtained by the encoding layer of the branch network 112 are not input to the decoding layer.
[0122] It can be understood that the first clothing image contains more detailed information of the clothing to be tried on. The encoding layer of the branch network 112 can obtain a lot of detailed information of the clothing, but the deformation degree of the first clothing image is not good, and the degree of conformity with the human body posture of the model is poor. Therefore, too fine position information is not required, otherwise it will interfere with the position of the detailed information of the clothing to be tried on in the target image output by the backbone network and the deformation degree of the clothing to be tried on. In order not to send too fine position information to the backbone network, the skip layer connection between the encoding module and the decoding module of the branch network can be removed, as shown in the branch network 112. This can also reduce the network operation amount of the clothing try-on model.
[0123] Figure 9 、 Figure 10 and Figure 11The description is made by taking the example that the encoding module includes 4 encoding layers and the decoding module includes 4 decoding layers. In practical applications, the number of encoding layers included in the encoding module and the number of decoding layers included in the decoding module are not limited; illustratively, the number of encoding layers included in the encoding module in the branch network can be the same as or different from the number of encoding layers included in the encoding module in the trunk network.
[0124] The above method supplements the detail information of the clothing to be tried on in the target image output by the clothing fitting model, that is, the detail information of the clothing to be tried on obtained by the decoding layer of the branch network supplements the details of the clothing to be tried on in the target image obtained by the trunk network, thereby improving the generation effect of the target image and enhancing the fitting experience.
[0125] It is understandable that if the pose of the model in the model image is relatively complex, for example, with the hands crossed on the chest, the clothes to be tried on in the target image may cover the hands, or the hands may be confused with the clothes to be tried on.
[0126] like Figure 12 FIG. 1 is a schematic diagram showing a specific embodiment of an exemplary target image, that is, if the pose of the model is relatively complex, the limbs of the model in the target image are not clear.
[0127] like Figure 12 As shown, if the model image is image 1101 and the clothing image is image 1102, the target image expected to be obtained is image 1103. However, since the model's hands are crossed in front of the chest, the target image obtained may be image 1104 or image 1105.
[0128] As can be seen from the areas framed by dotted lines in images 1104 and 1105 , the model's arms in image 1104 are blocked by the clothing to be tried on, that is, her hands are confused with the clothing on her chest, and the two arms of the model in image 1105 are confused into one.
[0129] Through continuous research, the inventor of the present application finally proposed a method with better effect, that is, combining the limb connection topology map to obtain the human body instance segmentation map, and combining the limb connection topology map to obtain the clothing segmentation map. Since the limb connection topology map contains the connection information of each key point, the limb connection topology map is used as the limb connection constraint, and the model can learn and effectively identify the left arm, right arm, left leg and right leg, and distinguish the left arm, right arm, left leg and right leg from the clothing.
[0130] The detailed description is given below.
[0131] It can be understood that there are multiple ways to implement step S24. The embodiments of the present application provide but are not limited to the following ways. The process of obtaining the human instance segmentation map by combining the limb connection topology map will be described below. This method includes the following steps A11 to step A13.
[0132] Step A11: Connect the key points included in the human key point map based on the pre-set connection relationships of each key point to obtain a limb connection topology map.
[0133] Still taking Figure 3 as an example, the limb connection topology map is as Figure 13 shown.
[0134] Step A12: Input the human key point map, the limb connection topology map, the candidate human instance segmentation map, and the clothing image into a pre-constructed instance segmentation model to obtain the human instance segmentation map output by the instance segmentation model. In the human instance segmentation map, the regions where different parts in the fitting part of the model belong to different instance regions.
[0135] The fitting part of the model refers to the body part located in the fitting instance region. If the fitting instance region is the upper body, the fitting part includes the left arm and the right arm, that is, the left arm and the right arm belong to different parts in the fitting part.
[0136] Among them, the instance segmentation model takes the human key point map of the sample model, the limb connection topology map of the sample model, the candidate human instance segmentation map of the sample model, and the clothing image corresponding to the sample model including the clothing to be tried on as inputs, and takes the annotated human instance segmentation map corresponding to the sample model as the training target to train a machine learning model. In the annotated human instance segmentation map, different parts in the fitting part of the sample model are annotated as different instance regions.
[0137] In the human instance segmentation map, different limbs in the fitting instance region of the model belong to different instance regions.
[0138] Since the limb connection topology map is combined, the instance segmentation model can effectively identify different parts of the fitting part. For example, the left arm instance region and the right arm instance region are identified as different instance regions. And since the limb connection topology map is combined, the instance segmentation model can effectively identify the orientation of the parts of the fitting part. For example, whether the left arm and the right arm cross. Thus, the limbs in the fitting instance region in the obtained human instance segmentation map will not be considered to belong to the same instance region. Since the instance segmentation model will not determine the left arm instance region and the right arm instance region as the same instance region, the situation shown in Image 1105 will not occur.
[0139] Since the obtained human instance segmentation map is more accurate, and the hair instance area, face instance area, left arm instance area, right arm instance area, left leg instance area, right leg instance area, and background instance area in the human instance segmentation map are all different instance areas, in the subsequent process of obtaining the target image based on the human instance segmentation map, the situation of Figure 12 the image 1105 shown will not occur.
[0140] In an alternative implementation, a more accurate clothing segmentation map can be obtained by combining the limb connection topology map. The process of obtaining the clothing segmentation map in step S22 is described below, and this process includes the following steps B11.
[0141] Step B11: Input the human key point map, the human instance segmentation map, the limb connection topology map, and the clothing image into a pre-constructed clothing segmentation model to obtain the clothing segmentation map output by the clothing segmentation model;
[0142] Among them, the clothing segmentation model takes the human key point map of the sample model, the human instance segmentation map of the sample model, the limb connection topology map of the sample model, and the corresponding clothing image of the sample model as inputs, and takes the annotated clothing segmentation map corresponding to the sample model as the training target, and trains a machine learning model to obtain it. The area of the clothing to be tried on in the annotated clothing segmentation map does not include the area blocked by the body parts of the sample model.
[0143] The degree of deformation of the target clothing in the clothing segmentation map matches the human body posture of the model.
[0144] The following explains "the area of the clothing to be tried on in the annotated clothing segmentation map does not include the area blocked by the body parts of the sample model". If the model image of the sample model is as shown in image 1101 and the clothing image of the sample model is as shown in image 1102, then the annotated clothing segmentation map of the sample model can be as Figure 14 the clothing segmentation map 1301 shown.
[0145] It can be seen from the clothing segmentation map 1301 that since the clothing segmentation map no longer includes the area where the hands are crossed, when matching the clothing to be tried on in the clothing segmentation map to the fitting area of the human body segmentation map, the situation where the hands are blocked by the clothing as shown in image 1104 will not occur.
[0146] The embodiments of the present application involve multiple technologies in the field of artificial intelligence.
[0147] Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems that can perceive the environment, acquire knowledge, and use knowledge to achieve the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also involves researching the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0148] AI technology is an interdisciplinary subject that covers a wide range of fields, including both hardware and software technologies. Its basic technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, mechatronics, etc. AI software technologies mainly include computer vision technology, speech processing technology, Natural Language Processing (NLP) technology, and machine learning / deep learning in several major directions.
[0149] Among them, NLP technology mainly studies various theories and methods that can enable effective communication between humans and computers in natural language. It is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. NLP technology usually includes machine translation. As the name implies, machine translation technology refers to the technology of researching an intelligent machine that can translate languages in a way similar to human intelligence. Among them, a machine translation system usually consists of an encoder and a decoder. In addition to machine translation, NLP technology also includes technologies such as robot question answering, text processing, semantic understanding, and knowledge graphs.
[0150] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as object recognition, tracking, and measurement in machine vision, and further performing image processing to make the images processed by the computer more suitable for human eyes to observe or for transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, attempting to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0151] The key technologies of Speech Processing Technology include Automatic Speech Recognition (ASR), Text To Speech (TTS), and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction, and among them, speech has become one of the most promising human-computer interaction methods in the future.
[0152] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.
[0153] The following models mentioned in the embodiments of this application: key point prediction model, spatial transformation network, calibration network, clothing try-on model, instance segmentation model, clothing segmentation model may be obtained by training machine learning models.
[0154] During the process of training machine learning models, at least one of the technologies in machine learning such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration is involved.
[0155] Exemplarily, the machine learning model can be any one of a neural network model, a logistic regression model, a linear regression model, a support vector machine (SVM), Adaboost, XGboost, and a Transformer-Encoder model.
[0156] Exemplarily, the neural network model can be any one of a model based on a recurrent neural network, a model based on a convolutional neural network, and a classification model based on a Transformer-encoder.
[0157] Exemplarily, the machine learning model can be a deep hybrid model of a model based on a recurrent neural network, a model based on a convolutional neural network, and a classification model based on a Transformer-encoder.
[0158] Exemplarily, the machine learning model can be any one of an attention-based deep model, a memory network-based deep model, and a short text classification model based on deep learning.
[0159] The short text classification model based on deep learning is a recurrent neural network (RNN) or a convolutional neural network (CNN), or a variant based on a recurrent neural network or a convolutional neural network.
[0160] Exemplarily, some simple domain adaptation modifications can be made on a pre-trained model to obtain a machine learning model.
[0161] Exemplarily, the "simple domain adaptation modifications" include, but are not limited to, re-using large-scale unsupervised domain corpus for secondary pre-training on a pre-trained model, and / or compressing the pre-trained model by means of model distillation.
[0162] Exemplarily, the network structures of the instance segmentation model, the clothing segmentation model, and the calibration network can be the same as Figure 10 the network structure of the backbone network shown.
[0163] In the above embodiments disclosed in the present application, the method is described in detail. For the method of the present application, it can be implemented by means of devices in various forms. Therefore, the present application also discloses a device, and specific embodiments are given below for detailed description.
[0164] As Figure 15 shown, it is the structural diagram of the image generation device provided by the embodiment of the present application. The device includes: a first acquisition module 151, a first processing module 152, a second processing module 153, a generation module 154, a second acquisition module 155, a third acquisition module 156, a first input module 157, a second input module 158, and a fourth acquisition module 159, where:
[0165] The first acquisition module 151 is configured to input the model image into a pre-constructed key point prediction model to obtain a human key point map output by the key point prediction model, where the human key point map includes the key points of the model in the model image;
[0166] The first processing module 152 is configured to perform deformation processing on the to-be-tried-on clothing in the clothing image including the to-be-tried-on clothing according to the model image and the human key point map to obtain a clothing segmentation map;
[0167] The second processing module 153 is configured to perform instance segmentation processing on the model image according to the human key point map to obtain a candidate human instance segmentation map, where the candidate human instance segmentation map includes a fitting instance area for wearing the to-be-tried-on clothing;
[0168] The generation module 154 is configured to match the to-be-tried-on clothing included in the clothing segmentation map to the fitting instance area in the candidate human instance segmentation map to generate a human instance segmentation map;
[0169] The second acquisition module 155 is configured to input the clothing image and the clothing segmentation map into a pre-constructed spatial transformation network to obtain a first clothing image output by the spatial transformation network;
[0170] The third acquisition module 156 is configured to input the first clothing image and the clothing segmentation map into a calibration network to obtain a second clothing image output by the calibration network;
[0171] The first input module 157 is configured to input the human instance segmentation map, the clothing segmentation map, the second clothing image, and a retained image into the backbone network of a pre-constructed clothing try-on model, where the retained image is the image in the model image except for the fitting instance area;
[0172] The second input module 158 is configured to input the first clothing image into a branch network of the clothing try-on model, and the branch network is configured to send the detailed information of the to-be-tried-on clothing obtained from the first clothing image to the backbone network;
[0173] The fourth acquisition module 159 is configured to output a target image through the clothing try-on model, where the target image includes the model wearing the to-be-tried-on clothing.
[0174] In an optional implementation manner, the generation module includes:
[0175] The first acquisition unit is configured to connect the key points included in the human key point map based on the connection relationship of each preset key point to obtain a limb connection topology map;
[0176] A second acquisition unit, configured to input the human key point map, the limb connection topology map, the candidate human instance segmentation map, and the clothing image into a pre-constructed instance segmentation model, so as to obtain a human instance segmentation map output by the instance segmentation model, where regions where different parts in the fitting part of the model in the human instance segmentation map belong to different instance regions;
[0177] Wherein, the instance segmentation model takes the human key point map of a sample model, the limb connection topology map of the sample model, the candidate human instance segmentation map of the sample model, and the clothing image corresponding to the sample model including the clothing to be tried on as inputs, and takes the annotated human instance segmentation map corresponding to the sample model as the training target, and is obtained by training a machine learning model, and different parts in the fitting part of the sample model in the annotated human instance segmentation map are marked as different instance regions;
[0178] A generation unit, configured to match the clothing to be tried on included in the clothing segmentation map to the fitting instance region in the human instance segmentation map, and generate a human instance segmentation map..
[0179] In an optional implementation manner, the first processing module includes:
[0180] A fourth acquisition unit, configured to input the human key point map, the human instance segmentation map, the limb connection topology map, and the clothing image into a pre-constructed clothing segmentation model, so as to obtain the clothing segmentation map output by the clothing segmentation model;
[0181] Wherein, the clothing segmentation model takes the human key point map of a sample model, the human instance segmentation map of the sample model, the limb connection topology map of the sample model, and the corresponding clothing image of the sample model as inputs, and takes the annotated clothing segmentation map corresponding to the sample model as the training target, and is obtained by training a machine learning model, and the region where the clothing to be tried on is located in the annotated clothing segmentation map does not include the region blocked by the body parts of the sample model.
[0182] In an optional implementation manner, the backbone network includes a decoding module and an encoding module, and the branch network includes a decoding module and an encoding module;
[0183] The decoding module of the backbone network includes a first number of decoding layers, the decoding module of the branch network includes the first number of decoding layers, and the i-th decoding layer in the decoding module of the branch network is connected to the i-th decoding layer in the decoding module of the backbone network;
[0184] The decoding layers of the encoding module of the branch network are not connected to the decoding layers of the decoding module of the branch network, and the encoding layers of the encoding module of the backbone network are connected to the decoding layers of the decoding module of the backbone network.
[0185] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.
[0186] Figure 16 is a block diagram of a server shown according to an exemplary embodiment, as Figure 16 shown, the server includes but is not limited to: a processor 1501, a memory 1502, a network interface 1503, an I / O controller 1504, and a communication bus 1505.
[0187] It should be noted that those skilled in the art can understand that Figure 16 the structure of the server shown in Figure 16 does not constitute a limitation on the server. The server may include more or fewer components than
[0188] shown, or combine certain components, or have different component arrangements. Figure 16 The following specifically introduces each component of the server 1500 in combination with
[0189] The processor 1501 is the control center of the server, connecting various parts of the entire server through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 1502, and calling data stored in the memory 1502, it executes various functions of the server and processes data, thereby monitoring the server as a whole. The processor 1501 may include one or more processing units; optionally, the processor 1501 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above modem processor may not be integrated into the processor 1501.
[0190] The processor 1501 may be a central processing unit (CPU), or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention, etc.;
[0191] Memory 1502 may include internal memory, such as high-speed random access memory (RAM) 15021 and read-only memory (ROM) 15022, and may also include a mass storage device 15023, such as at least one magnetic disk storage device, etc. Of course, the server may also include other hardware required for other services.
[0192] Among them, the above-mentioned memory 1502 is used to store executable instructions of the above-mentioned processor 1501. The above-mentioned processor 1501 is configured to execute the above-mentioned image generation method.
[0193] A wired or wireless network interface 1503 is configured to connect the server 1500 to a network.
[0194] The processor 1501, the memory 1502, the network interface 1503, and the I / O controller 1504 may be interconnected via a communication bus 1505, which may be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc.
[0195] In an exemplary embodiment, the server 1500 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components, and is used to execute the above-mentioned image generation method.
[0196] In an exemplary embodiment, a storage medium including instructions is also provided, such as the memory 1502 including instructions, and the above-mentioned instructions can be executed by the processor 1501 of the server 1500 to complete the above-mentioned method. Optionally, the storage medium may be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium may be ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0197] In an exemplary embodiment, a computer program product is further provided. The computer program product can be directly loaded into the internal memory of a computer, such as the aforementioned memory 1502, and contains software code. After being loaded and executed by the computer, the computer program can implement the method shown in any of the embodiments of the above image generation method.
[0198] It should be noted that the features described in the respective embodiments of this specification can be replaced or combined with each other. For device or system type embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.
[0199] It should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0200] The steps of the method or algorithm described in connection with the embodiments disclosed herein can be implemented directly in hardware, in a software module executed by a processor, or in a combination of both. The software module can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0201] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An image generation method, characterized in that, Including: Inputting a model image into a pre - constructed key - point prediction model to obtain a human key - point map output by the key - point prediction model, where the human key - point map includes the key points of the model in the model image; Performing deformation processing on the to - be - tried - on clothing in the clothing image including the to - be - tried - on clothing according to the model image and the human key - point map to obtain a clothing segmentation map; Performing instance segmentation processing on the model image according to the human key - point map to obtain a candidate human instance segmentation map, where the candidate human instance segmentation map includes a try - on instance area for wearing the to - be - tried - on clothing; Matching the to - be - tried - on clothing included in the clothing segmentation map to the try - on instance area in the candidate human instance segmentation map to generate a human instance segmentation map; Inputting the clothing image and the clothing segmentation map into a pre - constructed spatial transformation network to obtain a first clothing image output by the spatial transformation network; Inputting the first clothing image and the clothing segmentation map into a calibration network to obtain a second clothing image output by the calibration network; Inputting the human instance segmentation map, the clothing segmentation map, the second clothing image, and a remaining image into the backbone network of a pre - constructed clothing try - on model, where the remaining image is the image in the model image except for the try - on instance area; Inputting the first clothing image into the branch network of the clothing try - on model, and the branch network is used to send the detail information of the to - be - tried - on clothing obtained from the first clothing image to the backbone network; Outputting a target image through the clothing try - on model, where the target image includes the model wearing the to - be - tried - on clothing.
2. The image generation method according to claim 1, wherein Matching the to - be - tried - on clothing included in the clothing segmentation map to the try - on instance area in the candidate human instance segmentation map to generate a human instance segmentation map, including: Connecting the key points included in the human key - point map based on the pre - set connection relationship of each key point to obtain a limb connection topology map; Inputting the human key - point map, the limb connection topology map, the candidate human instance segmentation map, and the clothing image into a pre - constructed instance segmentation model to obtain a human instance segmentation map output by the instance segmentation model, where the areas of different parts in the try - on parts of the model in the human instance segmentation map belong to different instance areas; Among them, the instance segmentation model takes the human key - point map of a sample model, the limb connection topology map of the sample model, the candidate human instance segmentation map of the sample model, and the clothing image corresponding to the sample model including the to - be - tried - on clothing as inputs, and takes the annotated human instance segmentation map corresponding to the sample model as the training target to train a machine - learning model to obtain, and different parts in the try - on parts of the sample model in the annotated human instance segmentation map are annotated as different instance areas.
3. The image generation method according to claim 2, wherein Performing deformation processing on the to - be - tried - on clothing in the clothing image including the to - be - tried - on clothing according to the model image and the human key - point map to obtain a clothing segmentation map, including: Input the human key point map, the human instance segmentation map, the limb connection topology map, and the clothing image into a pre-constructed clothing segmentation model to obtain the clothing segmentation map output by the clothing segmentation model; Among them, the clothing segmentation model takes the human key point map of the sample model, the human instance segmentation map of the sample model, the limb connection topology map of the sample model, and the corresponding clothing image of the sample model as inputs, and takes the annotated clothing segmentation map corresponding to the sample model as the training target, and trains a machine learning model to obtain it. The area of the clothing to be tried on in the annotated clothing segmentation map does not include the area blocked by the body parts of the sample model.
4. The image generation method according to any one of claims 1 to 3, characterized in that, The backbone network includes a decoding module and an encoding module, and the branch network includes a decoding module and an encoding module; The decoding module of the backbone network includes a first number of decoding layers. The decoding module of the branch network includes the first number of decoding layers. The i-th decoding layer in the decoding module of the branch network is connected to the i-th decoding layer in the decoding module of the backbone network; The decoding layers of the encoding module of the branch network are not connected to the decoding layers of the decoding module of the branch network, and the encoding layers of the encoding module of the backbone network are connected to the decoding layers of the decoding module of the backbone network.
5. An image generation device, characterized in that, It includes: The first acquisition module is used to input the model image into a pre-constructed key point prediction model to obtain the human key point map output by the key point prediction model. The human key point map includes the key points of the model in the model image; The first processing module is used to perform deformation processing on the clothing to be tried on in the clothing image including the clothing to be tried on according to the model image and the human key point map to obtain a clothing segmentation map; The second processing module is used to perform instance segmentation processing on the model image according to the human key point map to obtain a candidate human instance segmentation map. The candidate human instance segmentation map includes a fitting instance area for wearing the clothing to be tried on; The generation module is used to match the clothing to be tried on included in the clothing segmentation map to the fitting instance area in the candidate human instance segmentation map to generate a human instance segmentation map; The second acquisition module is used to input the clothing image and the clothing segmentation map into a pre-constructed spatial transformation network to obtain the first clothing image output by the spatial transformation network; The third acquisition module is used to input the first clothing image and the clothing segmentation map into a correction network to obtain the second clothing image output by the correction network; The first input module is used to input the human instance segmentation map, the clothing segmentation map, the second clothing image, and the remaining image into the backbone network of a pre-constructed clothing try-on model. The remaining image is the image in the model image except for the fitting instance area; The second input module is used to input the first clothing image into the branch network of the clothing try-on model. The branch network is used to send the detailed information of the clothing to be tried on obtained from the first clothing image to the backbone network; A fourth acquisition module, configured to output a target image through the clothing try-on model, where the target image includes the model wearing the clothing to be tried on.
6. The image generation device according to claim 5, characterized in that, The generation module includes: A first acquisition unit, configured to connect the key points included in the human key point map based on the connection relationships of the preset key points, so as to obtain a limb connection topology map; A second acquisition unit, configured to input the human key point map, the limb connection topology map, the candidate human instance segmentation map, and the clothing image into a pre-constructed instance segmentation model, so as to obtain a human instance segmentation map output by the instance segmentation model, where different regions in the fitting parts of the model in the human instance segmentation map belong to different instance regions; Wherein, the instance segmentation model takes the human key point map of the sample model, the limb connection topology map of the sample model, the candidate human instance segmentation map of the sample model, and the clothing image corresponding to the sample model as inputs, and takes the labeled human instance segmentation map corresponding to the sample model as the training target, and trains a machine learning model to obtain it. In the labeled human instance segmentation map, different parts in the fitting parts of the sample model are labeled as different instance regions.
7. The image generation device according to claim 6, wherein The first processing module includes: A fourth acquisition unit, configured to input the human key point map, the human instance segmentation map, the limb connection topology map, and the clothing image into a pre-constructed clothing segmentation model, so as to obtain the clothing segmentation map output by the clothing segmentation model; Wherein, the clothing segmentation model takes the human key point map of the sample model, the human instance segmentation map of the sample model, the limb connection topology map of the sample model, and the corresponding clothing image of the sample model as inputs, and takes the labeled clothing segmentation map corresponding to the sample model as the training target, and trains a machine learning model to obtain it. In the labeled clothing segmentation map, the area where the clothing to be tried on is located does not include the area blocked by the body parts of the sample model.
8. The image generation device according to any one of claims 5 to 7, characterized in that, The backbone network includes a decoding module and an encoding module, and the branch network includes a decoding module and an encoding module; The decoding module of the backbone network includes a first number of decoding layers, and the decoding module of the branch network includes the first number of decoding layers. The i-th decoding layer in the decoding module of the branch network is connected to the i-th decoding layer in the decoding module of the backbone network; The decoding layers of the encoding module of the branch network are not connected to the decoding layers of the decoding module of the branch network, and the encoding layers of the encoding module of the backbone network are connected to the decoding layers of the decoding module of the backbone network.
9. A server, characterized in that, Includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the image generation method according to any one of claims 1 to 4.
10. A computer-readable storage medium, when the instructions in the storage medium are executed by the processor of the server, enabling the server to execute the image generation method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Virtual try-on method and system capable of reserving details of example clothes
CN111062777A
Figure virtual clothes changing method, terminal equipment and storage medium
CN113436058A