Model Training Method, Apparatus, Electronic Device, and Storage Medium
By using initial clothing images, human instance segmented images and key point images in the virtual trial installation network model for training, the problem of relying on human analysis results in the existing technology is solved, and a more flexible and high-precision virtual trial installation effect is achieved.
Patent Information
- Application Number
- CN202210605229.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-30
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-05-30
AI Technical Summary
The existing virtual trial installation technology relies on human analysis results during the training and inference stages, resulting in inflexible solutions, limited application scenarios, and the trial installation accuracy is affected by human analysis accuracy.
By obtaining the initial virtual trial installation network model, the first clothing image, the first human instance segmented image and the first human key point image are used for model training, and the target virtual trial installation network model is obtained. The model does not rely on the human analysis results during the training stage and retains the human analysis feature information.
A virtual trial installation network model that does not rely on human analysis results during the training stage is realized, which enhances the flexibility of the solution and the breadth of application scenarios, and at the same time improves the trial installation accuracy, avoiding the impact of human analysis accuracy.
Smart Images

Figure CN114972919B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a model training method, apparatus, electronic device, and storage medium. Background Art
[0002] Virtual Try-on means: given a target clothing image and a character (such as a model) image, generating an image of the character wearing the target clothing, so as to achieve the purpose of virtual try-on of the target clothing for the character. In the e-commerce field, a good virtual try-on technology can not only provide consumers with a novel interactive experience, but also guide and stimulate consumers to make purchase decisions faster. In the video field, the virtual try-on technology can bring a new viewing experience to viewers. For example, while watching a movie, viewers can also try on the clothing of the characters in the movie, bringing an immersive viewing experience. Or they can freely switch the clothing worn by the characters in the movie, increasing the interactivity and fun of viewing, and also virtually making viewers stay active on the video platform for a longer time.
[0003] In related technologies, the virtual try-on technology includes the parser-based scheme. Among them, the parser-based scheme is characterized in that in both the training stage and the inference stage, it is necessary to perform human parsing on the input image to obtain the instance segmentation result of the person and the human key points (and the connection relationship of the human limbs) as the input. Since both the training stage and the inference stage rely on the result of human parsing, the scheme is not flexible, the application scenario is limited, and the try-on accuracy is also affected by the result of human parsing. If the result of human parsing is not accurate enough, the accuracy of virtual try-on will be greatly reduced. Summary of the Invention
[0004] To solve the above technical problems that since both the training stage and the inference stage rely on the result of human parsing, the scheme is not flexible, the application scenario is limited, and the try-on accuracy is also affected by the result of human parsing. If the result of human parsing is not accurate enough, the accuracy of virtual try-on will be greatly reduced, the embodiments of the present invention provide a model training method, apparatus, electronic device, and storage medium. The specific technical solutions are as follows:
[0005] In the first aspect of the embodiments of the present invention, first, a model training method is provided, and the method includes:
[0006] Obtain an initial virtual try-on network model, where the initial virtual try-on network model is obtained by training a model based on a first clothing image, a first human instance segmentation image, and a first human key point image;
[0007] Obtain a second clothing image, a second human body instance segmentation image, and a second human body key point image, and input them into the initial virtual fitting network model to obtain a first fitting result output by the initial virtual fitting network model;
[0008] Obtain a third clothing image, and perform model training on the initial virtual fitting network model based on the first fitting result and the third clothing image to obtain a target virtual fitting network model.
[0009] In an optional embodiment, the initial virtual fitting network model includes an initial deformation network, an initial transformation network, and an initial fitting network;
[0010] The step of obtaining a second clothing image, a second human body instance segmentation image, and a second human body key point image, and inputting them into the initial virtual fitting network model to obtain a first fitting result output by the initial virtual fitting network model includes:
[0011] Obtain a second clothing image, a second human body instance segmentation image, and a second human body key point image, and input them into the initial deformation network to obtain a deformed second clothing segmentation map output by the initial deformation network;
[0012] Input the second clothing image and the deformed second clothing segmentation map into the initial transformation network to obtain a transformed second clothing image output by the initial transformation network;
[0013] Process the second human body instance segmentation image;
[0014] Input the processed second human body instance segmentation image, the transformed second clothing image, and the deformed second clothing segmentation map into the initial fitting network to obtain a first fitting result output by the initial fitting network.
[0015] In an optional embodiment, the target virtual fitting network model includes a target deformation network, a target transformation network, and a target fitting network;
[0016] The step of performing model training on the initial virtual fitting network model based on the first fitting result and the third clothing image to obtain a target virtual fitting network model includes:
[0017] Perform model training on the initial deformation network based on the first fitting result and the third clothing image to obtain a target deformation network;
[0018] Keep the parameters of the initial transformation network fixed, or perform model training on the initial transformation network based on the first fitting result, the third clothing image, and the target deformation network to obtain a target transformation network;
[0019] Based on the first fitting result, the third clothing image, the target deformation network, and the initial transformation network or the target transformation network, model training is performed on the initial fitting network to obtain a target fitting network.
[0020] In an alternative embodiment, the initial deformation network includes an initial clothing branch sub-network, an initial human body branch sub-network, an initial processing sub-network, and an initial feature correction sub-network, and the target deformation network includes a target clothing branch sub-network, a target human body branch sub-network, a target processing sub-network, and a target feature correction sub-network;
[0021] The model training of the initial deformation network based on the first fitting result and the third clothing image to obtain the target deformation network includes:
[0022] Keep the parameters of the initial clothing branch sub-network fixed, input the first fitting result and the third clothing image into the initial deformation network to perform model training on the initial deformation network until the cross-entropy loss function of the deformed third clothing segmentation map output by the initial deformation network and the first distillation loss function between the initial processing sub-network and the target processing sub-network both converge, and obtain the target deformation network;
[0023] Wherein, the first distillation loss function includes:
[0024]
[0025] Includes the l-th layer feature in the target processing sub-network, Includes the l-th layer feature in the initial processing sub-network, and the first distillation loss function represents calculating the L1 loss between all L layer features of the initial processing sub-network and the target processing sub-network.
[0026] In an alternative embodiment, the model training of the initial transformation network based on the first fitting result, the third clothing image, and the target deformation network to obtain the target transformation network includes:
[0027] Input the first fitting result and the third clothing image into the target deformation network to obtain the deformed third clothing segmentation map output by the target deformation network;
[0028] Input the third clothing image and the deformed third clothing segmentation map into the initial transformation network to perform model training on the initial transformation network until the L1 loss function and the perceptual loss function of the transformed clothing image output by the initial transformation network both converge, and obtain the target transformation network.
[0029] In an alternative embodiment, training the initial fitting network based on the first fitting result, the third clothing image, the target deformation network, and the initial transformation network or the target transformation network to obtain a target fitting network includes:
[0030] Inputting the first fitting result and the third clothing image into the target deformation network to obtain a deformed third clothing segmentation map output by the target deformation network;
[0031] Inputting the third clothing image and the deformed third clothing segmentation map into the initial transformation network or the target transformation network to obtain a transformed third clothing image output by the initial transformation network or the target transformation network;
[0032] Inputting the transformed third clothing image and the first fitting result into the initial fitting network to train the model of the initial fitting network until the L1 loss function, the perceptual loss function of the second fitting result output by the initial fitting network, and the second distillation loss function between the initial fitting network and the target fitting network all converge, thereby obtaining the target fitting network;
[0033] Wherein, the second distillation loss function includes:
[0034]
[0035] The includes the l-th layer feature of the target fitting network, and the includes the l-th layer feature of the initial fitting network. The second distillation loss function represents calculating the L1 loss between all L layer features of the initial fitting network and the target fitting network.
[0036] In an alternative embodiment, the method further includes:
[0037] Obtaining a target clothing image and a target fitting result, where the target fitting result includes an image of a character wearing a specific piece of clothing;
[0038] Inputting the target clothing image and the target fitting result into the target deformation network to obtain a deformed target clothing segmentation map output by the target deformation network;
[0039] Inputting the target clothing image and the deformed target clothing segmentation map into the target transformation network to obtain a transformed target clothing image output by the target transformation network;
[0040] Input the target fitting result and the transformed target clothing image into the target fitting network to obtain the third fitting result output by the target fitting network.
[0041] In a second aspect of the embodiments of the present invention, there is also provided a model training device, which includes:
[0042] A model acquisition module, configured to acquire an initial virtual fitting network model, where the initial virtual fitting network model is obtained by training a model based on a first clothing image, a first human instance segmentation image, and a first human key point image;
[0043] An image acquisition module, configured to acquire a second clothing image, a second human instance segmentation image, and a second human key point image, and input them into the initial virtual fitting network model to obtain a first fitting result output by the initial virtual fitting network model;
[0044] A model training module, configured to acquire a third clothing image, and train the initial virtual fitting network model based on the first fitting result and the third clothing image to obtain a target virtual fitting network model.
[0045] In a third aspect of the embodiments of the present invention, there is also provided an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus;
[0046] The memory is used to store a computer program;
[0047] The processor is configured to implement the model training method described in any one of the first aspects when executing the program stored in the memory.
[0048] In a fourth aspect of the embodiments of the present invention, there is also provided a storage medium, in which instructions are stored. When the instructions run on a computer, the computer is enabled to execute the model training method described in any one of the first aspects.
[0049] In a fifth aspect of the embodiments of the present invention, there is also provided a computer program product containing instructions. When the computer program product runs on a computer, the computer is enabled to execute the model training method described above.
[0050] The technical solution provided by the embodiment of the present invention is to obtain an initial virtual fitting network model, where the initial virtual fitting network model is obtained by training a model based on a first clothing image, a first human instance segmentation image, and a first human key point image. Then, a second clothing image, a second human instance segmentation image, and a second human key point image are obtained and input into the initial virtual fitting network model to obtain a first fitting result output by the initial virtual fitting network model. Next, a third clothing image is obtained, and the initial virtual fitting network model is trained based on the first fitting result and the third clothing image to obtain a target virtual fitting network model. By obtaining the first virtual fitting network model trained based on the first clothing image, the first human instance segmentation image, and the first human key point image, and training the first virtual fitting network model based on the first fitting result and the third clothing image to obtain the target virtual fitting network model, the target virtual fitting network model does not depend on the result of human parsing during the training stage and can still retain the feature information of human parsing, making the solution more flexible, not restricted by the application scenario, and the fitting accuracy is no longer affected by the result of human parsing. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention.
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0053] Figure 1 It is a schematic flowchart of an implementation process of a model training method shown in an embodiment of the present invention;
[0054] Figure 2 It is a schematic structural diagram of an initial virtual fitting network model shown in an embodiment of the present invention;
[0055] Figure 3 It is a schematic flowchart of an implementation process of an inference method of an initial virtual fitting network model shown in an embodiment of the present invention;
[0056] Figure 4 It is a schematic flowchart of an implementation process of a training method of an initial virtual fitting network model shown in an embodiment of the present invention;
[0057] Figure 5 It is a schematic flowchart of an implementation process of another training method of an initial virtual fitting network model shown in an embodiment of the present invention;
[0058] Figure 6 This is a schematic diagram of the implementation process of a target virtual fitting network model inference method shown in an embodiment of the present invention;
[0059] Figure 7 This is a schematic structural diagram of a model training device shown in an embodiment of the present invention;
[0060] Figure 8 This is a schematic structural diagram of an electronic device shown in an embodiment of the present invention. Detailed implementation manners
[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0062] In the embodiments of the present invention, for the initial virtual fitting network model, the initial virtual fitting network model here can be regarded as the Teacher model. The training and inference of the initial virtual fitting network model both rely on the human body parsing result, that is, the initial virtual fitting network model is obtained by training based on the first clothing image, the first human body instance segmentation image, and the first human body key point image.
[0063] After the training of the initial virtual fitting network model is completed, it means that the features in each network of the initial virtual fitting network model all contain the feature information of human body parsing. At this time, the initial virtual fitting network model can be used as the initialization virtual fitting network model, which means that each network in the initialization virtual fitting network model is initialized with the initial virtual fitting network model, so that the target virtual fitting network model can be trained, which can be regarded as the Student model. Therefore, only by constraining the features of the corresponding levels in the target virtual fitting network model to be as similar as possible to those of the initial virtual fitting network model, the training and inference of the target virtual fitting network model can no longer rely on the human body parsing result, and the feature information of human body parsing can still be retained.
[0064] Based on the above inventive concept, as Figure 1 shown, this is a schematic diagram of the implementation process of a model training method provided by an embodiment of the present invention. This method is applied to an electronic device and specifically may include the following steps:
[0065] S101. Obtain an initial virtual fitting network model, where the initial virtual fitting network model is obtained by training a model based on a first clothing image, a first human instance segmentation image, and a first human key point image.
[0066] In an embodiment of the present invention, an initial virtual fitting network model can be obtained. The initial virtual fitting network model can be regarded as a Teacher model, which is obtained by training a model based on a first clothing image, a first human instance segmentation image, and a first human key point image.
[0067] For example, for the Teacher model, it is obtained by performing supervised model training based on a first clothing image C1, a first human instance segmentation image M1, and a first human key point image Kf. Specifically, each network in the Teacher model can be trained separately.
[0068] S102. Obtain a second clothing image, a second human instance segmentation image, and a second human key point image, and input them into the initial virtual fitting network model to obtain a first fitting result output by the initial virtual fitting network model.
[0069] In an embodiment of the present invention, a second clothing image, a second human instance segmentation image, and a second human key point image can be obtained, and the second clothing image, the second human instance segmentation image, and the second human key point image are input into the above-mentioned initial virtual fitting network model.
[0070] Among them, for clothing, for example, it can be any clothing such as an upper garment, trousers, shoes, etc. The embodiment of the present invention does not limit this. Thus, by inputting the second clothing image, the second human instance segmentation image, and the second human key point image into the above-mentioned initial virtual fitting network model, a corresponding fitting result can be obtained, that is, a first fitting result output by the initial virtual fitting network model.
[0071] S103. Obtain a third clothing image, and train the initial virtual fitting network model based on the first fitting result and the third clothing image to obtain a target virtual fitting network model.
[0072] In an embodiment of the present invention, a third clothing image can also be obtained. Thus, the first fitting result and the third clothing image can be used as training samples to participate in subsequent model training. Among them, the first fitting result can be understood as an image of a character wearing the second clothing.
[0073] Thus, based on the training samples, i.e., the first fitting result and the third clothing image, the initial virtual fitting network model can be trained until the relevant loss function converges to obtain the target virtual fitting network model. Among them, the initial virtual fitting network model can be supervised trained based on the first fitting result and the third clothing image, and the embodiments of the present invention do not limit this.
[0074] Through the description of the technical solution provided by the embodiments of the present invention above, an initial virtual fitting network model is obtained. Among them, the initial virtual fitting network model is obtained by training based on the first clothing image, the first human instance segmentation image, and the first human key point image. The second clothing image, the second human instance segmentation image, and the second human key point image are obtained and input into the initial virtual fitting network model to obtain the first fitting result output by the initial virtual fitting network model, and the third clothing image is obtained. The initial virtual fitting network model is trained based on the first fitting result and the third clothing image to obtain the target virtual fitting network model.
[0075] By obtaining the initial virtual fitting network model trained based on the first clothing image, the first human instance segmentation image, and the first human key point image, and training the initial virtual fitting network model based on the first fitting result and the third clothing image to obtain the target virtual fitting network model. In this way, the target virtual fitting network model does not depend on the result of human parsing during the training phase, and can also retain the feature information of human parsing, making the solution more flexible, not restricted by the application scenario, and the fitting accuracy is no longer affected by the result of human parsing.
[0076] Among them, for the initial virtual fitting network model, it can include an initial deformation network, an initial transformation network, and an initial fitting network. Among them, the initial deformation network includes an initial clothing branch sub-network, an initial human body branch sub-network, an initial processing sub-network, and an initial feature correction sub-network. The structure of the initial virtual fitting network model can be as Figure 2 shown.
[0077] Among them, a simple description of the training process of the initial virtual fitting network model (i.e., the Teacher model) is as follows:
[0078] (1) Train the initial deformation network. The initial deformation network includes two branches, namely the initial clothing branch sub-network and the initial human body branch sub-network. Among them, input 1 includes the first clothing image, and input 2 includes the first human instance segmentation image and the first human key point image. The first clothing image, the first human instance segmentation image, and the first human key point image are input into the initial deformation network, and the initial deformation network outputs the deformed first clothing image. The cross-entropy loss constraint L cross is used until convergence.
[0079] (2) Train the initial transformation network. The input of the initial transformation network is the first clothing image and the output of the initial deformation network. That is, the first clothing image, the first human instance segmentation image, and the first human key point image are input into the initial deformation network. The initial deformation network outputs the deformed first clothing image. The first clothing image and the deformed first clothing image are input into the initial transformation network. The initial transformation network outputs the transformed first clothing image. Use L1 and perceptual loss L perceptual , until convergence.
[0080] (3) Train the initial fitting network. The input of the initial fitting network is the transformed first clothing image output by the initial transformation network, the deformed first clothing image output by the initial deformation network, and the human information to be retained. Among them, the human information to be retained is generally the processed first human instance segmentation image. Among them, the first clothing image, the first human instance segmentation image, and the first human key point image are input into the initial deformation network. The initial deformation network outputs the deformed first clothing image. The first clothing image and the deformed first clothing image are input into the initial transformation network. The initial transformation network outputs the transformed first clothing image. The deformed first clothing image, the transformed first clothing image, and the human information to be retained are input into the initial fitting network. The fitting result output by the initial fitting network uses L1 and perceptual loss L perceptual , until convergence.
[0081] After training each network in the initial virtual fitting network model in this way, the initial virtual fitting network model can be used for inference. Specifically, step S102 above can be specifically as Figure 3 shown Figure 3 is a schematic flowchart of the implementation process of an initial virtual fitting network model inference method provided by an embodiment of the present invention, which may specifically include the following steps:
[0082] S301, obtain the second clothing image, the second human instance segmentation image, and the second human key point image, and input them into the initial deformation network to obtain the deformed second clothing segmentation map output by the initial deformation network.
[0083] S302, input the second clothing image and the deformed second clothing segmentation map into the initial transformation network to obtain the transformed second clothing image output by the initial transformation network.
[0084] S303, process the second human instance segmentation image.
[0085] S304. Input the processed second human instance segmentation image, the transformed second clothing image, and the deformed second clothing segmentation map into the initial fitting network to obtain the first fitting result output by the initial fitting network.
[0086] In an embodiment of the present invention, a second clothing image, a second human instance segmentation image, and a second human key point image can be obtained, and thus the second clothing image, the second human instance segmentation image, and the second human key point image can be input into the initial deformation network, and the initial deformation network can output a deformed second clothing segmentation map.
[0087] Obtain the deformed second clothing segmentation map output by the initial deformation network, input the second clothing image and the deformed second clothing segmentation map into the initial transformation network, and the initial transformation network outputs a transformed second clothing image. In addition, the second human instance segmentation image can be processed, and the processed second human instance segmentation image will retain the human information that needs to be retained, such as the lower limbs, hair, face, etc.
[0088] For the processed second human instance segmentation image, the transformed second clothing image, and the deformed second clothing segmentation map, the processed second human instance segmentation image, the transformed second clothing image, and the deformed second clothing segmentation map can be input into the initial fitting network, and the initial fitting network outputs the first fitting result, so that the first fitting result output by the initial fitting network can be obtained.
[0089] In addition, for the target virtual fitting network model, it can be understood as a copy version of the initial virtual fitting network model, and thus the corresponding target virtual fitting network model includes a target deformation network, a target transformation network, and a target fitting network.
[0090] Before training the model of the initial virtual fitting network model, the network structure and parameters of the target deformation network are the same as those of the initial deformation network, the network structure and parameters of the target transformation network are the same as those of the initial transformation network, and the network structure and parameters of the target fitting network are the same as those of the initial fitting network.
[0091] Based on this, the above step S104, as Figure 4 shown, Figure 4 is a schematic flowchart of the implementation process of a method for training an initial virtual fitting network model provided by an embodiment of the present invention. The method may specifically include the following steps:
[0092] S401. Train the model of the initial deformation network based on the first fitting result and the third clothing image to obtain a target deformation network.
[0093] In an embodiment of the present invention, for the first fitting result and the third clothing image, the initial deformation network can be trained based on the first fitting result and the third clothing image to obtain a target deformation network. Among them, the initial deformation network is trained in a supervised manner.
[0094] Among them, for the initial deformation network, it includes an initial clothing branch sub-network, an initial human body branch sub-network, an initial processing sub-network, and an initial feature correction sub-network. For the target deformation network, it includes a target clothing branch sub-network, a target human body branch sub-network, a target processing sub-network, and a target feature correction sub-network.
[0095] Specifically, since the input and output of the initial clothing branch sub-network remain unchanged, the parameters of the initial clothing branch sub-network can be kept fixed. The first fitting result and the third clothing image are input into the initial deformation network to train the initial deformation network until the cross-entropy loss function of the deformed third clothing segmentation map output by the initial deformation network and the first distillation loss function between the initial processing sub-network and the target processing sub-network both converge, and a target deformation network is obtained.
[0096] Among them, the first distillation loss function includes:
[0097]
[0098] It includes the l-th layer feature in the target processing sub-network. It includes the l-th layer feature in the initial processing sub-network. The first distillation loss function represents calculating the L1 loss between all L layer features of the initial processing sub-network and the target processing sub-network.
[0099] It should be noted that for the target processing sub-network, the middle L layer features are used as intermediate layer feature supervision to ensure that these middle L layer features are as similar as possible to the corresponding middle layer features in the initial processing sub-network.
[0100] S402. Keep the parameters of the initial transformation network fixed, or train the initial transformation network based on the first fitting result, the third clothing image, and the target deformation network to obtain a target transformation network.
[0101] In an embodiment of the present invention, for the initial transformation network, since its input and output remain unchanged, the parameters of the initial transformation network can be kept fixed, that is, the initial transformation network is used as the target transformation network, or the initial transformation network can be trained based on the first fitting result, the third clothing image, and the target deformation network to obtain a target transformation network.
[0102] Specifically, input the first fitting result and the third clothing image into the target deformation network to obtain the deformed third clothing segmentation map output by the target deformation network;
[0103] Input the third clothing image and the deformed third clothing segmentation map into the initial transformation network to train the model of the initial transformation network until both the L1 loss function and the perceptual loss function of the transformed clothing image output by the initial transformation network converge, obtaining the target transformation network.
[0104] S403. Train the model of the initial fitting network based on the first fitting result, the third clothing image, the target deformation network, and the initial transformation network or the target transformation network to obtain the target fitting network.
[0105] In the embodiment of the present invention, the model of the initial fitting network can be trained based on the first fitting result, the third clothing image, the target deformation network, and the initial transformation network or the target transformation network to obtain the target fitting network.
[0106] Specifically, input the first fitting result and the third clothing image into the target deformation network to obtain the deformed third clothing segmentation map output by the target deformation network;
[0107] Input the third clothing image and the deformed third clothing segmentation map into the initial transformation network or the target transformation network to obtain the transformed third clothing image output by the initial transformation network or the target transformation network;
[0108] Input the transformed third clothing image and the first fitting result into the initial fitting network to train the model of the initial fitting network until the L1 loss function, the perceptual loss function of the second fitting result output by the initial fitting network, and the second distillation loss function between the initial fitting network and the target fitting network all converge, obtaining the target fitting network.
[0109] Wherein, the second distillation loss function includes:
[0110]
[0111] Include the features of the l-th layer of the target fitting network, Include the features of the l-th layer of the initial fitting network. The second distillation loss function represents calculating the L1 loss between all L-layer features of the initial fitting network and the target fitting network.
[0112] It should be noted that for the target fitting network, the middle L-layer features are used as the middle layer feature supervision to ensure that these middle L-layer features are as similar as possible to the corresponding middle layer features in the initial fitting network.
[0113] Among them, through the above description of the training process of the initial virtual fitting network model, each network in the initial virtual fitting network model is trained separately. Among them, some networks need to be retrained, while some networks can be directly reused. Specifically, as Figure 5 shown.
[0114] Therefore, after the initial virtual fitting network model is trained, the trained initial virtual fitting network model (i.e., the target virtual fitting network model) can be used for model inference, as Figure 6 shown, Figure 6 The following figure shows a schematic flowchart of the implementation process of a target virtual fitting network model inference method provided by an embodiment of the present invention. This method may specifically include the following steps:
[0115] S601, Obtain a target clothing image and a target fitting result, where the target fitting result includes an image of a character wearing specific clothing.
[0116] S602, Input the target clothing image and the target fitting result into the target deformation network to obtain a deformed target clothing segmentation map output by the target deformation network.
[0117] S603, Input the target clothing image and the deformed target clothing segmentation map into the target transformation network to obtain a transformed target clothing image output by the target transformation network.
[0118] S604, Input the target fitting result and the transformed target clothing image into the target fitting network to obtain a third fitting result output by the target fitting network.
[0119] In the embodiment of the present invention, a target clothing image and a target fitting result can be obtained, where the target fitting result includes an image of a character wearing specific clothing, so that the target clothing image and the target fitting result can be input into the target deformation network, and thus the target deformation network outputs a deformed target clothing segmentation map.
[0120] Obtain the deformed target clothing segmentation map output by the target deformation network, input the target clothing image and the deformed target clothing segmentation map into the target transformation network, and thus the target transformation network outputs a transformed target clothing image.
[0121] Obtain the transformed target clothing image output by the target transformation network, input the target fitting result and the transformed target clothing image into the target fitting network, and thus the target fitting network outputs a third fitting result, and obtain the third fitting result output by the target fitting network.
[0122] In this way, the target virtual fitting network model does not depend on the result of human parsing during the inference stage, and can also retain the feature information of human parsing, making the solution more flexible, not restricted by the application scenario, and the fitting accuracy is no longer affected by the result of human parsing.
[0123] Corresponding to the above method embodiments, the embodiments of the present invention also provide a model training device, as Figure 7 shown. The device may include: a model acquisition module 710, an image acquisition module 720, and a model training module 730.
[0124] The model acquisition module 710 is configured to acquire an initial virtual fitting network model, where the initial virtual fitting network model is obtained by training a model based on a first clothing image, a first human instance segmentation image, and a first human key point image;
[0125] The image acquisition module 720 is configured to acquire a second clothing image, a second human instance segmentation image, and a second human key point image, and input them into the initial virtual fitting network model to obtain a first fitting result output by the initial virtual fitting network model;
[0126] The model training module 730 is configured to acquire a third clothing image, and train the initial virtual fitting network model based on the first fitting result and the third clothing image to obtain a target virtual fitting network model.
[0127] The embodiments of the present invention also provide an electronic device, as Figure 8 shown, including a processor 81, a communication interface 82, a memory 83, and a communication bus 84. Among them, the processor 81, the communication interface 82, and the memory 83 communicate with each other through the communication bus 84.
[0128] The memory 83 is used to store a computer program;
[0129] When the processor 81 is configured to execute the program stored in the memory 83, the following steps are implemented:
[0130] Acquire an initial virtual fitting network model, where the initial virtual fitting network model is obtained by training a model based on a first clothing image, a first human instance segmentation image, and a first human key point image; acquire a second clothing image, a second human instance segmentation image, and a second human key point image, and input them into the initial virtual fitting network model to obtain a first fitting result output by the initial virtual fitting network model; acquire a third clothing image, and train the initial virtual fitting network model based on the first fitting result and the third clothing image to obtain a target virtual fitting network model.
[0131] The communication bus mentioned in the above-mentioned electronic device may be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0132] The communication interface is used for communication between the above-mentioned electronic device and other devices.
[0133] The memory may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0134] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0135] In another embodiment provided by the present invention, a storage medium is further provided. Instructions are stored in the storage medium, and when it runs on a computer, the computer is caused to execute the model training method described in any one of the above embodiments.
[0136] In another embodiment provided by the present invention, a computer program product containing instructions is further provided. When it runs on a computer, the computer is caused to execute the model training method described in any one of the above embodiments.
[0137] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a storage medium or transmitted from one storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0138] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0139] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.
[0140] The above is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A model training method, characterized in that, The method includes: Obtaining an initial virtual fitting network model, where the initial virtual fitting network model is obtained by training a model based on a first clothing image, a first human body instance segmentation image, and a first human body key point image; the initial virtual fitting network model includes an initial deformation network, an initial transformation network, and an initial fitting network; Obtaining a second clothing image, a second human body instance segmentation image, and a second human body key point image, and inputting them into the initial virtual fitting network model to obtain a first fitting result output by the initial virtual fitting network model, including: obtaining a second clothing image, a second human body instance segmentation image, and a second human body key point image, and inputting them into the initial deformation network to obtain a deformed second clothing segmentation map output by the initial deformation network; inputting the second clothing image and the deformed second clothing segmentation map into the initial transformation network to obtain a transformed second clothing image output by the initial transformation network; processing the second human body instance segmentation image; inputting the processed second human body instance segmentation image, the transformed second clothing image, and the deformed second clothing segmentation map into the initial fitting network to obtain a first fitting result output by the initial fitting network; Obtaining a third clothing image, and training the initial virtual fitting network model based on the first fitting result and the third clothing image to obtain a target virtual fitting network model.
2. The method according to claim 1, characterized in that, The target virtual fitting network model includes a target deformation network, a target transformation network, and a target fitting network; The training the initial virtual fitting network model based on the first fitting result and the third clothing image to obtain a target virtual fitting network model includes: Training the initial deformation network based on the first fitting result and the third clothing image to obtain a target deformation network; Keeping the parameters of the initial transformation network fixed, or training the initial transformation network based on the first fitting result, the third clothing image, and the target deformation network to obtain a target transformation network; Training the initial fitting network based on the first fitting result, the third clothing image, the target deformation network, and the initial transformation network or the target transformation network to obtain a target fitting network.
3. The method according to claim 2, characterized in that, The initial deformation network includes an initial clothing branch sub-network, an initial human body branch sub-network, an initial processing sub-network, and an initial feature correction sub-network, and the target deformation network includes a target clothing branch sub-network, a target human body branch sub-network, a target processing sub-network, and a target feature correction sub-network; The training the initial deformation network based on the first fitting result and the third clothing image to obtain a target deformation network includes: Keep the parameters of the initial clothing branch sub-network fixed, and input the first fitting result and the third clothing image into the initial deformation network to train the model of the initial deformation network until the cross-entropy loss function of the deformed third clothing segmentation map output by the initial deformation network and the first distillation loss function between the initial processing sub-network and the target processing sub-network both converge, and obtain the target deformation network; Among them, the first distillation loss function includes: f l s , F2 including the features of the l-th layer in the target processing sub-network, f l t , F2 including the features of the l-th layer in the initial processing sub-network, and the first distillation loss function represents calculating the L1 loss between all L-layer features of the initial processing sub-network and the target processing sub-network.
4. The method according to claim 2, characterized in that, Training the model of the initial transformation network based on the first fitting result, the third clothing image, and the target deformation network to obtain the target transformation network, including: Input the first fitting result and the third clothing image into the target deformation network to obtain the deformed third clothing segmentation map output by the target deformation network; Input the third clothing image and the deformed third clothing segmentation map into the initial transformation network to train the model of the initial transformation network until the L1 loss function and the perceptual loss function of the transformed clothing image output by the initial transformation network both converge, and obtain the target transformation network.
5. The method according to claim 2, characterized in that, Training the model of the initial fitting network based on the first fitting result, the third clothing image, the target deformation network, and the initial transformation network or the target transformation network to obtain the target fitting network, including: Input the first fitting result and the third clothing image into the target deformation network to obtain the deformed third clothing segmentation map output by the target deformation network; Input the third clothing image and the deformed third clothing segmentation map into the initial transformation network or the target transformation network to obtain the transformed third clothing image output by the initial transformation network or the target transformation network; Input the transformed third clothing image and the first fitting result into the initial fitting network to train the model of the initial fitting network until the L1 loss function, the perceptual loss function of the second fitting result output by the initial fitting network, and the second distillation loss function between the initial fitting network and the target fitting network both converge, and obtain the target fitting network; Among them, the second distillation loss function includes: The f l s , C includes the l-th layer features of the target fitting network, the f l t , C includes the l-th layer features of the initial fitting network, and the second distillation loss function represents calculating the L1 loss between all L layer features of the initial fitting network and the target fitting network.
6. The method according to claim 5, characterized in that, The method further includes: Obtain a target clothing image and a target fitting result, where the target fitting result includes an image of a character wearing a specific clothing; Input the target clothing image and the target fitting result into the target deformation network to obtain the deformed target clothing segmentation map output by the target deformation network; Input the target clothing image and the deformed target clothing segmentation map into the target transformation network to obtain the transformed target clothing image output by the target transformation network; Input the target fitting result and the transformed target clothing image into the target fitting network to obtain the third fitting result output by the target fitting network.
7. A model training device, characterized in that, The device includes: A model acquisition module, configured to acquire an initial virtual fitting network model, where the initial virtual fitting network model is obtained by training a model based on a first clothing image, a first human instance segmentation image, and a first human key point image; the initial virtual fitting network model includes an initial deformation network, an initial transformation network, and an initial fitting network; An image acquisition module, configured to acquire a second clothing image, a second human instance segmentation image, and a second human key point image, and input them into the initial virtual fitting network model to obtain a first fitting result output by the initial virtual fitting network model, including: acquiring a second clothing image, a second human instance segmentation image, and a second human key point image, and inputting them into the initial deformation network to obtain a deformed second clothing segmentation map output by the initial deformation network; inputting the second clothing image and the deformed second clothing segmentation map into the initial transformation network to obtain a transformed second clothing image output by the initial transformation network; processing the second human instance segmentation image; inputting the processed second human instance segmentation image, the transformed second clothing image, and the deformed second clothing segmentation map into the initial fitting network to obtain a first fitting result output by the initial fitting network; A model training module, configured to acquire a third clothing image, and train the initial virtual fitting network model based on the first fitting result and the third clothing image to obtain a target virtual fitting network model.
8. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, where the processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used to store a computer program; The processor is configured to implement the method steps described in any one of claims 1-6 when executing the program stored on the memory.
9. A storage medium, on which a computer program is stored, characterized in that, When the program is executed by the processor, it implements the method described in any one of claims 1-6.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and storage medium
CN113570685A
Virtual wearing method and device, equipment, storage medium and program product
CN114067088A