Model training method, device and electronic equipment

By using clothing images and clothing segmentation diagrams in model training and combining the loss function of local images to train the model, the problem of missing details of the clothing to be tried is solved, and a better virtual trial-on effect is achieved.

CN114881245BActive Publication Date: 2025-05-13BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210600631.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-30
Publication Date
2025-05-13
Estimated Expiration
2042-05-30

AI Technical Summary

Technical Problem

The image details of the clothing to be tried on in the target photos obtained by the pre-constructed model in the prior art are seriously missing, resulting in the inability to produce good trial on results.

Method used

By inputting clothing images and clothing segmentation diagrams into machine learning models, predict clothing images are obtained, and the model is trained through the loss functions of multiple local images, so that the model focuses on the global and detailed information of the clothing to be tried on.

Benefits of technology

The detailed information of the clothing to be tried on in the target photo output by the model is achieved, which improves the effect of virtual trial on.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114881245B_ABST
    Figure CN114881245B_ABST
Patent Text Reader

Abstract

The present application provides a model training method, device and electronic device. In the process of training a machine learning model, multiple first partial images are obtained from the predicted clothing image obtained by the machine learning model. For each first partial image, a second partial image associated with the first partial image is obtained from the true clothing image. The areas of different first partial images are not completely the same. For each first partial image, a first loss function is obtained based on the first partial image and the second partial image associated with it. The machine learning model is trained based on the first loss functions corresponding to the multiple first partial images. Since the multiple first partial images include the first partial image composed of the detailed information of the clothing to be tried on, it means that the machine learning model can focus on the local area with detailed information in the clothing to be tried on, so that the machine learning model can obtain an image of the clothing to be tried on with relatively complete detailed information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and more specifically, to model training methods, devices and electronic equipment. Background Art

[0002] Virtual Try-on means: given a preset clothing image containing the clothing to be tried on and a model image containing the model, the preset clothing image and the model image are used as input of a pre-built model, and a target photo of the model wearing the clothing image to be tried on is generated through the model, thereby achieving the purpose of the model virtually trying on the clothing to be tried on.

[0003] The images of the clothing to be tried on in the target photo obtained by the pre-built model in the prior art are seriously missing details, for example, the details such as the logo or pattern of the clothing to be tried on are missing, resulting in failure to produce good fitting results. Therefore, how to train the model so that the detailed information of the clothing to be tried on in the target photo output by the model is not missing is a technical problem that those skilled in the art need to solve. Summary of the invention

[0004] In view of this, the present application provides a model training method, device and electronic device to solve the technical problem that the details of the image of the clothing to be tried on in the target photo obtained by the pre-built model in the prior art are seriously missing.

[0005] To achieve the above objectives, this application provides the following technical solutions:

[0006] According to a first aspect of an embodiment of the present disclosure, a model training method is provided, including:

[0007] At least a clothing image and a clothing segmentation map are input into a machine learning model, and a predicted clothing image is obtained through the machine learning model; the predicted clothing image is at least obtained by the machine learning model by deforming the clothing to be tried on contained in the clothing image according to the clothing segmentation map, the clothing segmentation map is obtained by deforming the clothing to be tried on in the clothing image according to a model image and a human body key point map corresponding to the model image, and the model image includes a model for trying on the clothing to be tried on;

[0008] Obtaining a plurality of first partial images in the predicted clothing image, wherein different first partial images are located in different locations of the predicted clothing image, and areas of different first partial images are not completely the same, and the plurality of first partial images include a first partial image composed of detailed information of the clothing to be tried on, and the plurality of first partial images can be combined into the predicted clothing image;

[0009] For each of the first partial images, a second partial image associated with the first partial image is obtained from a preset true value clothing image, wherein a position region of the second partial image in the true value clothing image is the same as a position region of the first partial image in the predicted clothing image, and the true value clothing image is a training target image output by the machine learning model;

[0010] For each of the first partial images, based on the first partial image and the second partial image associated with the first partial image, obtain a first loss function corresponding to the first partial image;

[0011] The machine learning model is trained using first loss functions corresponding to each of the first local images.

[0012] According to a second aspect of an embodiment of the present disclosure, a model training device is provided, including:

[0013] A first acquisition module is used to input at least a clothing image and a clothing segmentation map into a machine learning model, and obtain a predicted clothing image through the machine learning model; the predicted clothing image is at least obtained by the machine learning model deforming the clothing to be tried on contained in the clothing image according to the clothing segmentation map, the clothing segmentation map is obtained by deforming the clothing to be tried on in the clothing image according to a model image and a human body key point map corresponding to the model image, and the model image includes a model for trying on the clothing to be tried on;

[0014] A second acquisition module is used to obtain a plurality of first partial images in the predicted clothing image, wherein different first partial images are located in different locations of the predicted clothing image, and the areas of different first partial images are not completely the same, and the plurality of first partial images include a first partial image composed of detailed information of the clothing to be tried on, and the plurality of first partial images can be combined into the predicted clothing image;

[0015] A third acquisition module is used to obtain, for each of the first partial images, a second partial image associated with the first partial image from a preset true value clothing image, wherein a position region of the second partial image in the true value clothing image is the same as a position region of the first partial image in the predicted clothing image, and the true value clothing image is a training target image output by the machine learning model;

[0016] a fourth acquisition module, configured to obtain, for each of the first partial images, a first loss function corresponding to the first partial image based on the first partial image and the second partial image associated with the first partial image;

[0017] A training module is used to train the machine learning model through first loss functions corresponding to multiple first local images respectively.

[0018] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, characterized in that it includes:

[0019] processor;

[0020] a memory for storing instructions executable by the processor;

[0021] Wherein, the processor is configured to execute the instructions to implement the model training method as described in the first aspect.

[0022] Through the above technical solution, it can be known that in the model training method provided by the present application, at least the clothing image and the clothing segmentation map are input into the machine learning model, and the predicted clothing image is obtained through the machine learning model; multiple first partial images in the predicted clothing image are obtained; for each of the first partial images, a second partial image associated with the first partial image is obtained from a preset true value clothing image; for each of the first partial images, based on the first partial image and the second partial image associated with it, a first loss function corresponding to the first partial image is obtained; and the machine learning model is trained by the first loss function corresponding to multiple first partial images. Since different first partial images are located in different position areas of the predicted clothing image, the areas of different first partial images are not exactly the same, which means that there are first partial images with larger areas and first partial images with smaller areas. The first partial images with larger areas can constrain the machine learning model to learn the information of the global area of ​​the clothing to be tried, and the first partial images with smaller areas may only contain the local area of ​​the clothing to be tried. For example, the first partial image composed of the detailed information of the clothing to be tried only includes the detailed information of the clothing to be tried, and the detailed information of the clothing to be tried in the first partial image accounts for a large proportion, which can constrain the machine learning model to learn the detailed information of the clothing to be tried. By adopting the above-mentioned training method, in the process of training the machine learning model, the machine learning model can be made to focus on the global area of ​​the clothing to be tried on and the area where the detail information is located, so that the machine learning model can obtain images of the clothing to be tried on with detailed information and relatively comprehensive detail information, and finally obtain a target image containing the clothing to be tried on with detailed information and relatively comprehensive detail information. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0024] Figure 1 A schematic diagram of a network structure of a pre-built model involved in an embodiment of the present application;

[0025] Figure 2 A schematic diagram of a key point map of a human body output by a key point prediction model provided in an embodiment of the present application;

[0026] Figure 3 A schematic diagram of a target image provided in an embodiment of the present application;

[0027] Figure 4 A limb connection topology diagram provided for an embodiment of the present application;

[0028] Figure 5 A schematic diagram of a model image and a human body instance segmentation diagram provided in an embodiment of the present application;

[0029] Figure 6 A schematic diagram of a clothing image, a candidate clothing segmentation map, and a clothing segmentation map provided in an embodiment of the present application;

[0030] Figure 7 A clothing segmentation diagram provided in an embodiment of the present application;

[0031] Figure 8 A schematic diagram of a first clothing image obtained by a spatial transformation network provided in an embodiment of the present application;

[0032] Fig. 9 A schematic diagram of a correction network outputting a second clothing image provided by an embodiment of the present application;

[0033] Fig.10 A network structure diagram of a clothing fitting model provided in an embodiment of the present application;

[0034] Fig.11 A structural diagram of an implementation method of the hardware architecture involved in the embodiments of the present application;

[0035] Fig.12 A flowchart of a model training method provided in an embodiment of the present application;

[0036] Fig.13 A schematic diagram of a process for obtaining a predicted clothing image by setting a cropping frame provided in an embodiment of the present application;

[0037] Fig.14A schematic diagram of a process for obtaining a first partial image provided in an embodiment of the present application;

[0038] Fig.15 A structural diagram of a model training device provided in an embodiment of the present application;

[0039] Fig.16 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0040] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0041] The embodiments of the present application provide a model training method, device and electronic device. Before introducing the embodiments of the present application, the network structure and hardware architecture of the model involved in the embodiments of the present application are first described.

[0042] like Figure 1 , which is a schematic diagram of the network structure of the pre-built model involved in an embodiment of the present application.

[0043] The network structure of the pre-built model includes: a key point prediction model 11, an instance segmentation processing module 12, a clothing segmentation model 13, a clothing deformation model 14 and a clothing fitting model 15.

[0044] Among them, the key point prediction model is pre-constructed. For example, a neural network can be used, a large number of sample images are used as training sample images of the neural network, and the labeled images corresponding to the sample images are used as training targets to train and obtain the key point prediction model.

[0045] Optionally, the neural network can be a fully connected neural network (such as an MLP network, where MLP stands for Multi-layer Perceptron, which means a multi-layer perceptron), or other forms of neural networks (such as a convolutional neural network, a deep neural network, etc.).

[0046] Exemplarily, in the process of training to obtain the key point prediction model, the sample image may be an image containing a human body, and the annotated image corresponding to the sample image is an image in which each key point of the human body is annotated.

[0047] like Figure 2 As shown, it is a schematic diagram of the key point map of the human body output by the key point prediction model provided in an embodiment of the present application.

[0048] Exemplarily, the key points may include: head key point, left shoulder key point, right shoulder key point, neck key point, left elbow key point, right elbow key point, left wrist key point, right wrist key point, left palm key point, right palm key point, left hip key point, right hip key point, left knee key point, right knee key point, left ankle key point, right ankle key point, left toe key point, and right toe key point.

[0049] Exemplarily, the number of key points may be less than the above 18 key points, or more than the above 18 key points.

[0050] The model image 21 is used as the input of the key point prediction model 11; the key point prediction model 11 obtains the predicted positions of the key points of the model in the model image 21 after data processing. Exemplarily, the human body key point map output by the key point prediction model 11 may be shown in image 22.

[0051] In an optional implementation, there are multiple implementations of the instance segmentation processing module 12 , and the embodiments of the present application provide but are not limited to the following three.

[0052] The first method is to use an instance segmentation method to implement the function of the instance segmentation processing module 12. The instance segmentation method refers to automatically using an object detection method to frame different instances from an image, and then using a semantic segmentation method to perform pixel-by-pixel labeling in different instance regions.

[0053] Exemplarily, the fitting example area of ​​the model may be determined as the upper body and / or the lower body in combination with the preset clothing image. If the clothing to be tried on in the preset clothing image is a short-sleeved shirt, the fitting example area of ​​the model is the upper body; if the clothing to be tried on in the preset clothing image is pants, a skirt, or shorts, the fitting example area of ​​the model is the lower body; if the clothing to be tried on in the preset clothing image is a dress, the fitting example area of ​​the model includes the upper body and the lower body.

[0054] Instance segmentation is performed on the model image. Combined with the human body key point map, the target detection method can be used to frame out the hair instance region where the model's hair is located, the face instance region where the face is located, the left arm instance region where the left arm is located, the right arm instance region where the right arm is located, the clothing instance region where the model's clothes are located, the left leg instance region where the left leg is located, the right leg instance region where the right leg is located, and the background instance region from the model image. That is, the model's hair instance region, face instance region, left arm instance region, right arm instance region, clothing instance region where the model's clothes are located, left leg instance region, right leg instance region, and background instance region are all different instance regions.

[0055] Then, the pixels of different instance regions are marked by the semantic segmentation method to obtain the category labels of different instance regions, that is, the category labels of different instance regions are different.

[0056] In an optional implementation, if the fitting instance area of ​​the model is the upper body, it means that the clothing instance area where the clothing worn by the model's upper body is located needs to be clearly marked, and then the left arm instance area, the right arm instance area, the clothing instance area where the clothing worn by the model is located, the hair instance area, the face instance area and the lower body instance area are respectively regarded as different instance areas; the lower body area includes: the left leg instance area and the right leg instance area, that is, the left leg instance area and the right leg instance area are the same instance area; the upper body instance area includes: the left arm instance area, the right arm instance area, and the clothing instance area where the clothing worn by the model is located.

[0057] In an optional implementation, if the fitting instance area of ​​the model is the lower body, it means that the clothing instance area where the clothing worn by the model's lower body is located needs to be clearly marked, and then the upper body instance area, hair instance area, face instance area, left leg instance area, clothing instance area where the clothing worn by the model is located, and right leg instance area are respectively regarded as different instance areas; the upper body instance area includes: left arm instance area, right arm instance area, that is, the left arm instance area and the right arm instance area are the same instance area, and the lower body instance area includes: left leg instance area, clothing instance area where the clothing worn by the model is located, and right leg instance area.

[0058] In an optional implementation, if the model's fitting instance areas are the upper body and the lower body, the hair instance area, the face instance area, the left arm instance area, the right arm instance area, the clothing instance area (including the clothing instance area where the clothing worn by the model's upper body is located and the clothing instance area where the clothing worn by the model's lower body is located), the left leg instance area, and the right leg instance area are all treated as different instance areas.

[0059] The second method is to use the model to implement the function of the instance segmentation processing module 12.

[0060] It is understandable that if the pose of the model in the model image is relatively complex, for example, with the hands crossed on the chest, the clothes to be tried on in the target image may cover the hands, or the hands may be confused with the clothes to be tried on.

[0061] like Figure 3 As shown in Figure 1, it is a schematic diagram of the target image. Figure 3 It can be seen that if the model's posture is more complicated, the model's limbs (ie, body parts) in the target image are not clear.

[0062] like Figure 3As shown, if the model image is image 31, the preset clothing image is image 32, and the target image expected to be obtained is image 33. However, since the model's hands are crossed in front of the chest, the target image obtained may be image 34 or image 35.

[0063] As can be seen from the areas framed by dotted lines in images 34 and 35, in image 34, the model's arms are blocked by the clothing to be tried on, that is, her hands are confused with the clothing on her chest, and in image 35, the two arms of the model are confused into one.

[0064] Through continuous research, the inventor of the present application finally proposed a method with better effect, that is, combining the limb connection topology map to obtain the human body instance segmentation map, and combining the limb connection topology map to obtain the clothing segmentation map. Since the limb connection topology map contains the connection information of each key point, that is, the limb connection topology map is used as the limb connection constraint, the model can learn and effectively identify the left arm, right arm, left leg and right leg, and distinguish the left arm, right arm, left leg and right leg from the clothing. The limbs mentioned in the embodiments of the present application refer to at least one of the left arm, right arm, left leg and right leg.

[0065] The detailed description is given below.

[0066] The following describes a process of obtaining a human body instance segmentation map by combining a limb connection topology map. The method includes the following steps A11 to A13.

[0067] Step A11: connecting the key points included in the human body key point map based on the pre-set connection relationship of each key point to obtain a limb connection topology map.

[0068] Still Figure 2 For example, the limb connection topology diagram is as follows Figure 4 shown.

[0069] Step A12: performing instance segmentation processing on the model image according to the human body key point map to obtain a candidate human body instance segmentation map, wherein the candidate human body instance segmentation map includes a fitting instance area for wearing the garment to be tried on.

[0070] Exemplarily, the regions where different limbs of the fitting instance region of the model in the candidate human body instance image are located may belong to the same instance region. If the fitting instance region of the model is the upper body, the different limbs of the model in the fitting instance region include the left arm and the right arm; if the fitting instance region of the model is the lower body, the different limbs of the model in the fitting instance region include the left leg and the right leg; if the fitting instance region of the model is the lower body and the upper body, the different limbs of the model in the fitting instance region include the left leg, the right leg, the left arm and the right arm.

[0071] It is understandable that if the posture of the model in the model image is relatively complex, such as crossing her hands in front of her chest, the limbs of the fitting instance area in the candidate human instance segmentation map (for example, the left arm area and the right arm area) may belong to the same instance area.

[0072] Step A13: Input the human body key point map, the limb connection topology map, the candidate human body instance segmentation map and the preset clothing image into a pre-built instance segmentation model to obtain a human body instance segmentation map output by the instance segmentation model, wherein different areas of the model's fitting area in the human body instance segmentation map belong to different instance areas.

[0073] The fitting part of the model refers to the body part located in the fitting instance area. If the fitting instance area is the upper body, the fitting part includes the left arm and the right arm, that is, the left arm and the right arm belong to different parts of the fitting part.

[0074] Among them, the instance segmentation model is obtained by training a machine learning model by taking the human body key point map of the sample model, the limb connection topology map of the sample model, the candidate human body instance segmentation map of the sample model and the preset clothing image corresponding to the sample model containing the clothing to be tried on as input, and taking the annotated human body instance segmentation map corresponding to the sample model as the training target, wherein different parts of the sample model's fitting parts in the annotated human body instance segmentation map are annotated as different instance areas.

[0075] Exemplarily, the instance segmentation model has a function of matching the to-be-tried-on garment contained in the garment segmentation map to the fitting instance region in the human body instance segmentation map.

[0076] Different body parts, namely, limbs, of the model's fitting instance region in the human body instance segmentation graph belong to different instance regions.

[0077] Due to the combination of the limb connection topology map, the instance segmentation model can effectively identify different body parts in the fitting instance area, for example, identifying the left arm instance area and the right arm instance area as different instance areas. And due to the combination of the limb connection topology map, the instance segmentation model can effectively identify the direction of the body parts in the fitting instance area, for example, whether the left arm and the right arm are crossed. The different body parts of the instance area to be fitted in the human body instance segmentation map thus obtained will not be considered to belong to the same instance area. Since the instance segmentation model will not determine the left arm instance area and the right arm instance area as the same instance area, the situation shown in image 35 will not occur.

[0078] Since the obtained human instance segmentation map is more accurate, and the hair instance region, face instance region, left arm instance region, right arm instance region, left leg instance region, right leg instance region, and background instance region in the human instance segmentation map are all different instance regions, there will be no problem in the subsequent process of obtaining the target image based on the human instance segmentation map. Figure 3 The situation of image 35 is shown.

[0079] In summary, without combining the limb connection topology map, the instance segmentation processing module 12 outputs the above-mentioned candidate human body instance segmentation map; when combining the limb connection topology map, the instance segmentation processing module 12 outputs the above-mentioned human body instance segmentation map.

[0080] The third method is to use the human body parser (LIP) to implement the function of the instance segmentation processing module 12.

[0081] A human instance segmentation map can be obtained using a human parser (LIP). Exemplarily, LIP estimates a reasonable human analysis of a target image based on the approximate shapes of the body, face, hair, clothing, and human posture, which can effectively guide the synthesis of precise regions of human body parts.

[0082] In order for those skilled in the art to better understand the human body instance segmentation diagram provided in the embodiments of the present application, an example is given below to illustrate.

[0083] like Figure 5 , which is a schematic diagram of a model image and a human body instance segmentation map provided in an embodiment of the present application.

[0084] Assuming that the model image is shown as model image 21, if the clothing to be tried on is shown as preset clothing image 51, the human body instance segmentation image 52 is as follows: Figure 5 shown.

[0085] like Figure 5 In the human body instance segmentation diagram shown, the hair instance region, face instance region, left arm instance region, right arm instance region, clothing instance region, background instance region, left leg instance region and right leg instance region of the model are filled with different patterns respectively. Figure 5 In order to distinguish each instance area, the hair instance area is filled with multiple squares; the face instance area is filled with vertical lines, the right arm instance area is filled with oblique lines, the left arm instance area is filled with rhombuses, the clothing instance area is filled with black, the right leg instance area is filled with small dots, the left leg instance area is filled with horizontal lines, and the background instance area is filled with white.

[0086] In the real human instance segmentation map, different instance regions have different colors. Since the attached figure cannot display colored images, Figure 5 The above method is used to distinguish different instance areas.

[0087] In an optional implementation, there are multiple implementations of the clothing segmentation model 13, and the embodiments of the present application provide but are not limited to the following two.

[0088] The first implementation method of the clothing segmentation model 13 includes the following steps B11 to B14.

[0089] Step B11: extracting human features based on the model image, the human instance segmentation map and the human key point map.

[0090] Step B12: Calculate the correlation between clothing features and human body features in the preset clothing image to obtain a correlation tensor representing the human body features and clothing features.

[0091] Step B13: Obtain a candidate clothing segmentation map through a preset clothing image based on an instance segmentation method.

[0092] Step B14: Based on the correlation tensor and the candidate clothing segmentation map, a deformed clothing segmentation map is obtained.

[0093] In order for those skilled in the art to better understand the relationship between the preset clothing image, the candidate clothing segmentation map, and the clothing segmentation map, an example is given below to illustrate.

[0094] like Figure 6 , which is a schematic diagram of a preset clothing image, a candidate clothing segmentation map, and a clothing segmentation map provided in an embodiment of the present application.

[0095] like Figure 6 As shown, assuming that the preset clothing image is the preset clothing image 61, after performing instance segmentation on the preset clothing image 61, a candidate clothing segmentation map 62 can be obtained; and then combined with the model's posture, a deformed clothing segmentation map 63 can be obtained.

[0096] The second implementation method of the clothing segmentation model 13 includes the following steps B21.

[0097] Step B21: inputting the human body key point map, the human body instance segmentation map, the limb connection topology map, and the preset clothing image into a pre-constructed clothing segmentation model to obtain a clothing segmentation map output by the clothing segmentation model.

[0098] Exemplarily, the limb connection topology map may be combined to obtain a more accurate clothing segmentation map; exemplary, the limb connection topology map may not be combined in the process of obtaining the clothing segmentation map.

[0099] Among them, the clothing segmentation model is obtained by training a machine learning model by taking the human body key point map of the sample model, the human body instance segmentation map of the sample model, the limb connection topology map of the sample model, and the corresponding preset clothing image of the sample model as input, and taking the annotated clothing segmentation map corresponding to the sample model as the training target. The area where the clothing to be tried on is located in the annotated clothing segmentation map does not include the area blocked by the body parts of the sample model.

[0100] The degree of deformation of the target garment in the garment segmentation image matches the posture of the model.

[0101] The following is an explanation of “the region where the garment to be tried on is located in the labeled garment segmentation map does not include the region blocked by the body of the sample model”. If the model image of the sample model is shown in image 31, and the preset garment image of the sample model is shown in image 32, then the labeled garment segmentation map of the sample model can be as follows: Figure 7 The clothing segmentation diagram 71 is shown.

[0102] It can be seen from the clothing segmentation map 71 that since the clothing segmentation map no longer includes the area where the hands are crossed, the situation where the hands shown in image 34 are blocked by the clothing to be tried on in the clothing segmentation map will not occur.

[0103] In an optional implementation, there are multiple implementations of the clothing deformation model 14, and the embodiments of the present application provide but are not limited to the following two.

[0104] The first implementation of the clothing deformation model 14 includes the following steps C11 to C12.

[0105] The clothing deformation model 14 includes a spatial variation network and a correction network.

[0106] Step C11: inputting the preset clothing image and the clothing segmentation map into a pre-constructed spatial transformation network to obtain a first clothing image output by the spatial transformation network.

[0107] The Spatial Transformer Network (STN) can be embedded into a certain layer of the network as a special network module, so that the network can support spatial transformation (affine transformation, projection transformation), etc., and provide the network with properties such as rotation invariance and translation invariance.

[0108] Exemplarily, the first clothing image obtained by the spatial transformation network is an RGB image with three RGB channels.

[0109] like Figure 8 , which is a schematic diagram of obtaining a first clothing image using the spatial transformation network provided in an embodiment of the present application.

[0110] like Figure 8 As shown, the first clothing image obtained by the spatial transformation network contains more comprehensive details of the clothing to be tried on, such as a very comprehensive pattern (e.g., pattern and LOGO) without any missing, but the deformed shape of the clothing in the first clothing image is not very ideal, that is, it is not very consistent with the human body posture of the model in the model image. Based on this, the first clothing image and the clothing segmentation map are input into the correction network to obtain the second clothing image output by the correction network.

[0111] Step C12: inputting the first clothing image and the clothing segmentation map into a correction network to obtain a second clothing image output by the correction network.

[0112] like Fig. 9 , which is a schematic diagram of the correction network provided in an embodiment of the present application outputting a second clothing image.

[0113] like Fig. 9 As shown, the correction network is mainly used to correct the deformed shape of the first clothing image based on the clothing segmentation map, so that the deformation degree of the clothing in the obtained second clothing image is more consistent with the posture of the model.

[0114] However, although the deformation degree of the clothing in the second clothing image obtained by the correction network is more consistent with the posture of the model, the pattern (for example, pattern and LOGO) of the clothing in the second clothing image is seriously missing.

[0115] In summary, in the implementation of the first clothing deformation model 14, the generated image output by the clothing deformation model is the second clothing image.

[0116] The second implementation method of the clothing deformation model 14 includes the following steps C21.

[0117] Step C21: Using the clothing deformation model to perform shape context thin plate spline interpolation (TPS) deformation on the preset clothing image according to the clothing segmentation map, to obtain a deformed image of the clothing to be tried on (belonging to the RGB channel).

[0118] Shape Context is a contour shape descriptor. In the clothing deformation model, shape context descriptors of clothing segmentation map and preset clothing image are obtained respectively, and N pairs of matching point pairs are calculated.

[0119] Thin plate spline interpolation will calculate the TPS parameters based on these N pairs of matching point pairs. TPS is a common method for 2D shape deformation. For N pairs of matching point sets in two images, a deformation is calculated to simulate the 2D deformation so that after one of the images is deformed, the N pairs of matching points coincide. Finally, based on the calculated TPS parameters, the preset clothing image is transformed in the same way to obtain the deformed image of the clothing to be tried on (belonging to the RGB channel).

[0120] In an optional implementation, there are multiple implementations of the clothing fitting model 15, and the embodiments of the present application provide but are not limited to the following two.

[0121] The first implementation method of the clothing fitting model 15 includes the following steps D11 to D13.

[0122] The clothing fitting model 15 includes a backbone network and a branch network.

[0123] Step D11: input the human body instance segmentation map, the clothing segmentation map, the second clothing image and the retained image into the backbone network of the pre-constructed clothing fitting model, and the retained image is the image of the model image excluding the fitting instance area.

[0124] Keep the image belonging to the RGB channels.

[0125] Step D12: inputting the first clothing image into the branch network of the clothing fitting model, wherein the branch network is used to send the detailed information of the clothing to be fitted obtained from the first clothing image to the main network.

[0126] Step D13: outputting a target image through the clothing fitting model, wherein the target image includes the model wearing the clothing to be fitted.

[0127] Exemplarily, the target image is an image of RGB channels.

[0128] For example, Fig.10 As shown, the backbone network 1001 includes an encoding module and a decoding module. Fig.10 In the description, an example is given in which the encoding module includes 4 encoding layers and the decoding module includes 4 decoding layers.

[0129] like Fig.10As shown, the encoding layer is connected by a solid thin arrow, and the decoding layer is connected by a solid wide arrow. The encoding layer in the encoding module gradually extracts the features of the human body instance segmentation map, the clothing segmentation map, and the second clothing image. The decoding layer of the decoding module gradually enlarges the features and restores the image to the original size. The dotted thin arrow is the skip layer splicing part, which allows the network to retain more input information by directly connecting the shallow features of the encoding layer to the subsequent decoding layer.

[0130] Since the deformation degree of the second clothing image is better, the model in the target image obtained by the backbone network wears the clothing to be tried on more closely.

[0131] Exemplarily, the branch network 1002 includes an encoding module and a decoding module. Fig.10 In the description, an example is given in which the encoding module includes 4 encoding layers and the decoding module includes 4 decoding layers.

[0132] The encoding layer in the encoding module in the branch network 1002 gradually extracts the features of the first clothing image. Since the first clothing image contains more and more complete detail information of the clothing to be tried, the branch network can obtain a lot of detail information of the clothing to be tried, such as pattern information (e.g., LOGO and pattern). The branch network sends the obtained detail information of the clothing to be tried to the decoding layer of the trunk network through the decoding layer, so that the decoding layer of the trunk network can supplement the second clothing image based on the detail information output by the decoding layer of the branch network, so that the clothing to be tried on the model in the target image is fitted, and the detail information in the clothing to be tried is more comprehensive, that is, the detail information of the clothing to be tried is less, or even no detail information is lost.

[0133] In the second implementation of the clothing fitting model 15, the clothing fitting model 15 may only include a backbone network.

[0134] The embodiments of the present application involve multiple technologies in the field of artificial intelligence.

[0135] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0136] AI technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. Its basic technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. AI software technology mainly includes computer vision technology, speech processing technology, natural language processing (Nature Language Processing, NLP) technology, and machine learning / deep learning.

[0137] Among them, machine learning (ML) is a multi-disciplinary interdisciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.

[0138] The following models mentioned in the embodiments of the present application: key point prediction model, clothing deformation model, spatial transformation network, correction network, clothing fitting model, instance segmentation model, clothing segmentation model may be obtained by training a machine learning model.

[0139] The process of training a machine learning model involves at least one of the following techniques: artificial neural network, belief network, reinforcement learning, transfer learning, inductive learning, and method-based learning.

[0140] Exemplarily, the machine learning model can be any one of a neural network model, a logistic regression model, a linear regression model, a support vector machine (SVM), an Adaboost, an XGboost, and a Transformer-Encoder model.

[0141] Exemplarily, the neural network model can be any one of a recurrent neural network-based model, a convolutional neural network-based model, and a Transformer-encoder-based classification model.

[0142] Exemplarily, the machine learning model may be a deep hybrid model of a recurrent neural network-based model, a convolutional neural network-based model, and a Transformer-encoder-based classification model.

[0143] Exemplarily, the machine learning model can be any one of an attention-based deep model, a memory network-based deep model, and a short text classification model based on deep learning.

[0144] The short text classification model based on deep learning is a recurrent neural network (RNN) or a convolutional neural network (CNN) or a variant based on the recurrent neural network or the convolutional neural network.

[0145] For example, some simple domain adaptation modifications can be made on the pre-trained model to obtain a machine learning model.

[0146] Exemplarily, “simple domain adaptation” includes but is not limited to re-pre-training an already pre-trained model using large-scale unsupervised domain corpus, and / or compressing an already pre-trained model by model distillation.

[0147] The training method provided in the embodiment of the present application can be applied to the training of the above-mentioned model.

[0148] It can be understood that the purpose of the embodiment of the present application is to make the detail information of the clothing to be tried on in the target image output by the pre-constructed model less or even not missing. In the pre-constructed model, there are two models related to the detail information of the clothing to be tried on, namely: clothing deformation model and clothing try-on model. Therefore, in the process of training to obtain the pre-constructed model, the model training method provided in the embodiment of the present application is used to train the clothing deformation model and / or clothing try-on model.

[0149] In an optional implementation, the model training method provided by the embodiment of the present application can only train the clothing deformation model, or only train the clothing try-on model; in an optional implementation, the model training method provided by the embodiment of the present application can train the clothing deformation model and the clothing try-on model.

[0150] In the existing related technology, in the process of training a clothing deformation model, a generated image 1 output by the clothing deformation model is compared with a true image 1 (the true image 1 is a target training image output by the clothing deformation model corresponding to the generated image 1) pixel by pixel to obtain a first loss function 1, and the clothing deformation model is trained by the first loss function 1; in the process of training a clothing fitting model, a generated image 2 output by the clothing fitting model is compared with a true image 2 (the true image 2 is a target training image output by the clothing fitting model corresponding to the generated image 2) pixel by pixel to obtain a first loss function 2, and the clothing deformation model is trained by the first loss function 2.

[0151] It is understandable that the true image and the generated image include not only the clothes to be tried on, but may also include the body parts of the model and the background. When the true image and the generated image are compared pixel by pixel, the clothes to be tried on are not taken as the comparison focus. As a result, the detailed information of the clothes to be tried on in the generated image 1 output by the clothing deformation model is seriously missing, and the detailed information of the clothes to be tried on in the generated image 2 output by the clothing try-on model is seriously missing.

[0152] Based on this, an embodiment of the present application provides a model training method. In the process of training a clothing deformation model or a clothing fitting model, the clothing to be tried on is used as the training focus, so that the detail information of the clothing to be tried on in the generated image 1 output by the clothing deformation model is less missing or even not missing, and the detail information of the clothing to be tried on in the generated image 2 output by the clothing fitting model is less missing or even not missing.

[0153] The hardware architecture involved in the embodiments of the present application is described below.

[0154] like Fig.11 1 is a structural diagram of an implementation method of the hardware architecture involved in an embodiment of the present application, and the hardware architecture includes: a terminal device 1101 and an electronic device 1102.

[0155] Exemplarily, the terminal device 1101 can be any electronic product that can interact with a user through one or more methods such as a keyboard, touchpad, touch screen, remote control, voice interaction or handwriting device, such as a mobile phone, a laptop computer, a tablet computer, a PDA, a personal computer, a wearable device, a smart TV, a PAD, etc.

[0156] It should be noted that Fig.11 This is just an example. There can be many types of terminal devices, not limited to Fig.11 Laptops, smartphones, personal computers in.

[0157] Exemplarily, the electronic device 1102 may be a server, or a server cluster consisting of multiple servers, or a cloud computing server center. The electronic device 1102 may include a processor, a memory, a network interface, and the like.

[0158] Exemplarily, electronic device 1102 can be any electronic product that can interact with a user through one or more methods such as a keyboard, touchpad, touch screen, remote control, voice interaction or handwriting device, such as a mobile phone, a laptop computer, a tablet computer, a PDA, a personal computer, a wearable device, a smart TV, a PAD, etc.

[0159] Exemplarily, the user may determine, through the terminal device 1101 , a preset clothing image containing clothing to be tried on.

[0160] Exemplarily, the user may take a photo through the terminal device 1101 to obtain a model image including the model.

[0161] Exemplarily, the user may upload a model image via the terminal device 1101 .

[0162] Thereby, the terminal device 1101 can send the model image and the preset clothing image to the electronic device 1102 .

[0163] Exemplarily, the user may upload a preset true value image through the terminal device 1101. Each preset clothing image corresponds to a true value image, so that the terminal device 1101 sends the true value image to the electronic device 1102.

[0164] The ground truth image is the training target image for the machine learning model to be trained.

[0165] Exemplarily, the machine learning model is a clothing deformation model or a clothing fitting model.

[0166] Exemplarily, the electronic device 1102 can execute the model training method provided in the embodiments of the present application.

[0167] Those skilled in the art should understand that the above-mentioned terminal devices and electronic devices are only examples, and other existing or future terminal devices or electronic devices, if applicable to the present disclosure, should also be included in the protection scope of the present disclosure and are included here by reference.

[0168] The model training method provided in the embodiment of the present application is described below in combination with the network structure and hardware architecture of the above-mentioned pre-built model.

[0169] like Fig.12 As shown, it is a flowchart of the model training method provided in an embodiment of the present application, and the method includes the following steps S1201 to S1205.

[0170] Step S1201: at least input a clothing image and a clothing segmentation map into a machine learning model, and obtain a predicted clothing image through the machine learning model; the predicted clothing image is at least obtained by the machine learning model by deforming the clothing to be tried on contained in the clothing image according to the clothing segmentation map.

[0171] The clothing segmentation map is obtained by deforming the clothing to be tried on in the clothing image according to the model image and the human body key point map corresponding to the model image. The description of the clothing segmentation map can refer to the above description, which will not be repeated here.

[0172] The model image includes a model for trying on the garment to be tried on.

[0173] In an embodiment of the present application, the machine learning model includes: a clothing deformation model and / or a clothing fitting model.

[0174] The clothing deformation model and clothing fitting model can be trained separately.

[0175] If the machine learning model is a clothing deformation model, it is necessary to input the clothing image and the clothing segmentation map into the clothing deformation model, and output the second clothing image through the clothing deformation model. At this time, the clothing image mentioned in step S1201 includes the above-mentioned preset clothing image, see steps C11 to C12 for details.

[0176] If the machine learning model is a clothing fitting model, the human body instance segmentation map, the clothing segmentation map, the second clothing image, the retained image and the first clothing image need to be input into the clothing fitting model to obtain the target image output by the clothing fitting model. At this time, the clothing image mentioned in step S1201 includes the first clothing image and the second clothing image, see steps D11 to D13 for details.

[0177] In the process of training the clothing deformation model, the predicted clothing image corresponding to the clothing deformation model is the deformed second clothing image, for example, image 65, and the true clothing image corresponding to the predicted clothing image is a preset deformed image of the clothing to be tried on with detailed information.

[0178] In the process of training the clothing fitting model, the predicted clothing image corresponding to the clothing fitting model is the target image, and the true clothing image corresponding to the predicted clothing image is a preset image of the model wearing the deformed clothing to be tried on with detailed information.

[0179] Step S1202: obtaining a plurality of first partial images in the predicted clothing image, wherein different first partial images are located in different positions of the predicted clothing image, and areas of different first partial images are not completely the same, and the plurality of first partial images include a first partial image composed of detailed information of the clothing to be tried on, and the plurality of first partial images can be merged into the predicted clothing image.

[0180] In an optional implementation, the predicted clothing image is a generated image output by a machine learning model.

[0181] In an optional implementation, the predicted clothing image is obtained based on the generated image output by the machine learning model. Exemplarily, the predicted clothing image is an image of the clothing to be tried on in the generated image. That is, information irrelevant to the clothing to be tried on in the generated image is eliminated, and the image of the clothing to be tried on in the generated image is screened out to obtain the predicted clothing image.

[0182] In an optional implementation, different first partial images are located in different regions of the predicted clothing image, and the areas of different first partial images are not completely the same, which means that some first partial images have a large area and some first partial images have a small area.

[0183] For example, if the areas of the first partial images are area 1, area 2, and area 3 in descending order, then the number of first partial images with area 1 is less than the number of first partial images with area 2 and less than the number of first partial images with area 3.

[0184] It can be understood that the larger the area of ​​the first partial image is, the less content in the predicted clothing image that is not included in the first partial image is, and the more first partial images with larger areas are, the higher the overlap among multiple first partial images is, which is meaningless for obtaining the loss function.

[0185] It can be understood that the smaller the area of ​​the first partial image, the more likely it is that the first partial image contains only a partial area of ​​the clothing to be tried on. "Multiple first partial images include a first partial image composed of detailed information of the clothing to be tried on" indicates that multiple first partial images include at least one first partial image that only contains detailed information, for example, a logo. This means that in the first partial image, the logo of the clothing to be tried on accounts for a larger proportion, which makes it easier to constrain the machine learning model to pay attention to the detailed information of the clothing to be tried on, thereby making it easier for the machine learning model to recover the detailed information of the clothing to be tried on.

[0186] If the area of ​​the first partial image is larger, the content of the predicted clothing image that is not included in the first partial image is less. In the first partial image, there is a global area of ​​the clothing to be tried on, which can constrain the machine learning model to focus on the global area of ​​the clothing to be tried on, thereby making it easy for the machine learning model to restore the global area of ​​the clothing to be tried on.

[0187] The areas of the above-mentioned different first local images may be different, which means that in the process of training the machine learning model, the machine learning model can focus on the global area, medium local area, and small local area of ​​the clothing to be tried on, so that the machine learning model can obtain images of the clothing to be tried on with detailed information and relatively complete detailed information.

[0188] “Multiple first partial images can be merged into the predicted clothing image” indicates that the machine learning model can focus on the overall predicted clothing image.

[0189] Step S1203: For each of the first partial images, a second partial image associated with the first partial image is obtained from a preset true value clothing image, the position area of ​​the second partial image in the true value clothing image is the same as the position area of ​​the first partial image in the predicted clothing image, and the true value clothing image is a training target image output by the machine learning model.

[0190] It can be understood that since the true clothing image is a pre-set known image, and in the process of obtaining the first loss function, it is necessary to compare the second local image in the true clothing image with the first local image in the predicted clothing image, so the first local image and the second local image are one-to-one associated.

[0191] The first partial image and the second partial image associated therewith are explained below: the position area of ​​the first partial image in the true value clothing image is the same as the position area of ​​the second partial image in the predicted clothing image. For example, the coordinates of the position area corresponding to the first partial image are {(upper left vertex coordinate 1, upper right vertex coordinate 2, lower right vertex coordinate 3, lower left vertex coordinate 4)}, then the coordinates of the position area corresponding to the second partial image associated with the first partial image are {(upper left vertex coordinate 1, upper right vertex coordinate 2, lower right vertex coordinate 3, lower left vertex coordinate 4)}.

[0192] Step S1204: for each of the first partial images, based on the first partial image and the second partial image associated with the first partial image, obtain a first loss function corresponding to the first partial image.

[0193] In an optional implementation, based on the first partial image and the second partial image associated therewith, a method for obtaining a first loss function corresponding to the first partial image includes: based on the formula The first loss function is calculated, wherein p1 refers to the local generated image, p1(i) refers to the pixel value of the i-th pixel in the first local image, p2 refers to the local true value image, p2(i) refers to the pixel value of the i-th pixel in the second local image, and N is the total number of pixels contained in the first local image or the second local image.

[0194] It can be understood that the first partial image and the second partial image associated therewith contain the same total number of pixels.

[0195] Step S1205: training the machine learning model using first loss functions corresponding to the plurality of first local images respectively.

[0196] In an optional implementation, the machine learning model is trained using first loss functions corresponding to multiple first partial images respectively.

[0197] In an optional implementation, step S1205 may include: determining the mean of the first loss functions corresponding to the multiple first local images respectively as the second loss function; and training the machine learning model through the second loss function.

[0198] Assume that the number of first local images is 4, and they are: first local image 1 (corresponding weight 1), first local image 2 (corresponding weight 2), first local image 3 (corresponding weight 3), and first local image 4 (corresponding weight 4); then the second loss function = weight 1*first loss function of first local image 1 + weight 2*first loss function of first local image 2 + weight 3*first loss function of first local image 3 + weight 4*first loss function of first local image 4.

[0199] For example, the smaller the area of ​​the first partial image is, the greater the weight corresponding to the first partial image is, so that the machine learning model can pay more attention to the details of the clothing to be tried on.

[0200] In an optional implementation, before training the machine learning model, a target number of training times is set based on experience. If the machine learning model is trained at the target number of training times, the training is terminated.

[0201] In an optional implementation, before training the machine learning, a training termination condition can be set. For example, the training termination condition is the generated image output by the machine learning, so that the first loss function corresponding to the multiple first local images is less than or equal to the minimum value. If the output of the machine learning satisfies the training termination condition, the training is terminated.

[0202] In an optional implementation, before training the machine learning, a training termination condition can be set. For example, the training termination condition is that the generated image of the machine learning output makes the second loss function less than or equal to the minimum value. If the output of the machine learning satisfies the training termination condition, the training is terminated.

[0203] In the model training method provided in the embodiment of the present application, at least a clothing image and a clothing segmentation map are input into a machine learning model, and a predicted clothing image is obtained through the machine learning model; a plurality of first partial images in the predicted clothing image are obtained; for each of the first partial images, a second partial image associated with the first partial image is obtained from a preset true value clothing image; for each of the first partial images, a first loss function corresponding to the first partial image is obtained based on the first partial image and the second partial image associated with the first partial image; the machine learning model is trained by the first loss functions corresponding to the plurality of first partial images. Since different first partial images are located in different position areas of the predicted clothing image, the areas of different first partial images are not exactly the same, which means that there are first partial images with larger areas and first partial images with smaller areas. The first partial images with larger areas can constrain the machine learning model to learn the information of the global area of ​​the clothing to be tried on, and the first partial images with smaller areas may only contain the local area of ​​the clothing to be tried on. For example, the first partial image composed of the detailed information of the clothing to be tried on only includes the detailed information of the clothing to be tried on, and the detailed information of the clothing to be tried on in the first partial image accounts for a large proportion, which can constrain the machine learning model to learn the detailed information of the clothing to be tried on. By adopting the above-mentioned training method, in the process of training the machine learning model, the machine learning model can be made to focus on the global area of ​​the clothing to be tried on and the area where the detail information is located, so that the machine learning model can obtain images of the clothing to be tried on with detailed information and relatively comprehensive detail information, and finally obtain a target image containing the clothing to be tried on with detailed information and relatively comprehensive detail information.

[0204] In an optional implementation, there are multiple implementations of step S1201, and the embodiments of the present application provide but are not limited to the following two.

[0205] The first implementation method of step S1201 includes: determining the generated image output by the machine learning model as the predicted clothing image.

[0206] The second implementation of step S1201 includes the following steps E1 to E3.

[0207] Step E1: Get the generated image output by the machine learning model.

[0208] If the machine learning model is a clothing fitting model, the generated image may be the target image.

[0209] Step E2: Determine that the garment to be tried on is located in a first position area in the generated image.

[0210] It can be seen from image 65 or the target image that the generated image includes other contents besides the garment to be tried on, such as background, body parts, hair, and face of the model. In order to make the machine learning model learn the garment to be tried on in a targeted manner, the first position area where the garment to be tried on is located can be identified from the generated image.

[0211] In an optional implementation, the first position area where the garment to be tried on is located can be identified by a garment recognition model. In an optional implementation, the first position area where the garment to be tried on is located can be artificially drawn in the generated image. In an optional implementation, it can be known based on prior knowledge that the garment to be tried on is almost located at the center of the generated image, based on which, the specific implementation of step E12 includes the following steps E21 to E23.

[0212] Step E21: Determine the center position of the generated image.

[0213] Step E22: Taking the center position as the center of a set cropping frame, the set cropping frame is placed in the generated image.

[0214] Step E23: Determine that the position area where the set cropping frame is located in the generated image is the first position area.

[0215] Step E3: Determine that the image located in the first position area in the generated image is the predicted clothing image.

[0216] In an optional implementation, a preset true value clothing image is obtained from a preset true value image, that is, an image located in the first position area of ​​the preset true value image is determined to be the true value clothing image, and the true value image corresponds to the predicted clothing image.

[0217] The true value image is the training target image of the machine learning model, which is obtained in advance by humans.

[0218] like Fig.13 , which is a schematic diagram of a process for obtaining a predicted clothing image by setting a cropping frame provided in an embodiment of the present application.

[0219] Fig.13 In the example, the machine learning model is taken as a clothing fitting model, and it is assumed that the generated image is the target image 1301 and the preset true value image is 1302. It can be seen that the detail information of the clothing to be tried on in the generated image 1301 is missing, for example, the pattern is missing. The detail information of the clothing to be tried on in the true value image is more comprehensive, for example, the pattern is more comprehensive.

[0220] Assume that the first location area is Fig.13In the area enclosed by the dotted line, the predicted clothing image is image 1303 , and the true clothing image is image 1304 .

[0221] The above method can eliminate information irrelevant to the clothing to be tried on in the generated image and the true image, so that the machine learning model focuses on learning the information in the clothing to be tried on.

[0222] In an optional implementation, there are multiple implementations of step S1202. The embodiment of the present application provides but is not limited to the following method, which involves the following steps F1 to F3 during implementation.

[0223] Step F1: Obtain preset cropping frames with different areas.

[0224] Exemplarily, cropping frames of different areas may be obtained randomly; exemplary, cropping frames of different areas may be pre-set. Different areas of the cropping frames refer to different products of the length and width of the cropping frames.

[0225] Step F2: for each of the cropping frames, place the cropping frame in the predicted clothing image to obtain a second position area where the cropping frame is located in the predicted clothing image, and obtain the second position areas corresponding to different cropping frames.

[0226] Assume that the cropping frames of different areas are: the first cropping frame 1 with an area of ​​S1, the second cropping frame 2 with an area of ​​S2, the third cropping frame 3 with an area of ​​S3, ..., the cropping frame 4 with an area of ​​S q The qth cropping box q of .

[0227] Assume that the number of second position regions obtained by cropping frame 1 is M1, ..., and the number of second position regions obtained by cropping frame q is M q .

[0228] Exemplarily, the larger the area of ​​the cropping frame is, the smaller the number of the second position regions corresponding to the cropping frame is, that is, if S1>S2>…>S q , exemplary, M1 <M2<…<M q .

[0229] It can be understood that if the area of ​​the cropping box is larger and the number of first partial images obtained through the cropping box is larger, the overlap among multiple first partial images with the same area will be higher, which is meaningless for obtaining the loss function. Therefore, if the area of ​​the cropping box is larger, the number of first partial images obtained through the cropping box will be smaller.

[0230] It can be understood that if the area of ​​the cropping box is smaller, the first partial image with this area is more likely to contain only a local area of ​​the clothing to be tried on. For example, the first partial image may only include detailed information of the clothing to be tried on, such as a logo. This means that in the first partial image, the detailed information of the clothing to be tried on accounts for a larger proportion, and it is easier to constrain the machine learning model to pay attention to the detailed information, thereby making it easier for the machine learning model to recover the detailed information of the clothing to be tried on. Therefore, if the area of ​​the cropping box is smaller, the number of first partial images obtained through the cropping box is larger.

[0231] The different areas of the above-mentioned different cropping boxes mean that in the process of training the machine learning model, the machine learning model can focus on the global area, medium local area, and small local area of ​​the clothing to be tried on, so that the machine learning model can obtain images of the clothing to be tried on with detailed information and relatively comprehensive detail information.

[0232] For example, if S1>S2>…>S q , for example, M1, M2, ..., M q No size relationship.

[0233] Step F3: for each second position area, determine the image located in the second position area in the predicted clothing image as the first partial image.

[0234] In order for those skilled in the art to better understand the method provided in the embodiments of the present application, the following examples are provided for illustration.

[0235] like Fig.14 , which is a schematic diagram of the process of obtaining a first partial image provided in an embodiment of the present application.

[0236] Fig.14 The predicted clothing image is taken as image 1401 as an example for explanation. Fig.14 The square with a medium-thick solid line is the cropping frame, and the square with a thin solid line is the predicted clothing image 1401 .

[0237] For the first cropping frame 1 with an area of ​​S1, the cropping frame 1 is placed at different positions of the predicted clothing image to obtain M1 second position areas (but the areas of these second position areas are the same); based on the M1 second position areas, M1 first partial images can be obtained.

[0238] For the second cropping frame 2 with an area of ​​S2, the cropping frame 2 is placed at different positions of the predicted clothing image, and M2 second position areas can be obtained (but the areas of these second position areas are the same); based on the M2 second position areas, M2 first partial images can be obtained.

[0239] By analogy, for aq The qth cropping box q is placed at different positions of the predicted clothing image, and M q second location areas (but the areas of these second location areas are the same); based on M q The second location area can get M q The first partial image.

[0240] Correspondingly, in an optional implementation, obtaining a second local image associated with the first local image from a preset true-value clothing image includes: determining the second position area corresponding to the first local image as a position area in the true-value clothing image where the second local image associated with the first local image is located; and determining the image located in the second position area in the true-value clothing image as the second local image.

[0241] In an optional implementation, there are multiple implementations of step 1205. The present application embodiment provides but is not limited to the following: Based on the formula The second loss function L is calculated patchwise ;

[0242] Among them, q represents the total number of cropping boxes, M k is the total number of local generated images or local true value images cropped by the kth cropping frame; L kj 1(p1, p2) is the first loss function L1(p1, p2) obtained based on the j-th local generated image corresponding to the k-th cropping box and the j-th local true value image.

[0243] The method is described in detail in the embodiments disclosed in the above-mentioned application. The method of the application can be implemented by various forms of devices. Therefore, the application also discloses a device, and a specific embodiment is given below for detailed description.

[0244] like Fig.15 As shown, it is a structural diagram of a model training device provided in an embodiment of the present application, and the device includes: a first acquisition module 1501, a second acquisition module 1502, a third acquisition module 1503, a fourth acquisition module 1504 and a training module 1505, wherein:

[0245] The first acquisition module 1501 is used to input at least a clothing image and a clothing segmentation map into a machine learning model, and obtain a predicted clothing image through the machine learning model; the predicted clothing image is at least obtained by the machine learning model deforming the clothing to be tried on contained in the clothing image according to the clothing segmentation map, the clothing segmentation map is obtained by deforming the clothing to be tried on in the clothing image according to a model image and a human body key point map corresponding to the model image, and the model image includes a model for trying on the clothing to be tried on;

[0246] A second acquisition module 1502 is used to obtain a plurality of first partial images in the predicted clothing image, wherein different first partial images are located in different locations of the predicted clothing image, and the areas of different first partial images are not completely the same, and the plurality of first partial images include a first partial image composed of detailed information of the clothing to be tried on, and the plurality of first partial images can be combined into the predicted clothing image;

[0247] A third acquisition module 1503 is used to obtain, for each of the first partial images, a second partial image associated with the first partial image from a preset true value clothing image, wherein the second partial image is located in the same position area in the true value clothing image as the first partial image is located in the predicted clothing image, and the true value clothing image is a training target image output by the machine learning model;

[0248] A fourth acquisition module 1504 is configured to obtain, for each of the first partial images, a first loss function corresponding to the first partial image based on the first partial image and the second partial image associated with the first partial image;

[0249] The training module 1505 is used to train the machine learning model using the first loss functions corresponding to the plurality of first local images respectively.

[0250] In an optional implementation, the first acquisition module includes:

[0251] A first acquisition unit, used to acquire a generated image output by the machine learning model;

[0252] A first determining unit, configured to determine that the garment to be tried on is located in a first position area in the generated image;

[0253] The second determining unit is configured to determine that the image located in the first position area in the generated image is the predicted clothing image.

[0254] In an optional implementation, the method further includes:

[0255] The first determination module is used to determine that the image located in the first position area of ​​the preset true value image is the true value clothing image.

[0256] In an optional implementation, the first determining unit includes:

[0257] A first determining subunit, used to determine the center position of the generated image;

[0258] A second setting subunit is used to set the center position as the center of the set cropping frame and place the set cropping frame in the generated image;

[0259] The second determining subunit is used to determine that the position area where the set cropping frame is located in the generated image is the first position area.

[0260] In an optional implementation, the second acquisition module includes:

[0261] A second acquisition unit, used to acquire a preset cropping frame with different areas;

[0262] a third determining unit, for each of the cropping frames, placing the cropping frame in the predicted clothing image to obtain a second position area where the cropping frame is located in the predicted clothing image, and to obtain the second position areas corresponding to different cropping frames;

[0263] The fourth determining unit is configured to determine, for each second position area, an image located in the second position area in the predicted clothing image as the first partial image.

[0264] In an optional implementation, for each of the locally generated images, the training module includes:

[0265] a fifth determining unit, configured to determine the mean of the first loss functions respectively corresponding to the plurality of first partial images as the second loss function;

[0266] A training unit is used to train the machine learning model by using the second loss function.

[0267] In an optional implementation, for each of the first partial images, the fourth acquisition module includes:

[0268] The first calculation unit is used to calculate the The first loss function is calculated, wherein p1 refers to the local generated image, p1(i) refers to the pixel value of the i-th pixel in the first local image, p2 refers to the local true value image, p2(i) refers to the pixel value of the i-th pixel in the second local image, and N is the total number of pixels contained in the first local image or the second local image.

[0269] In an optional implementation manner, the fifth determining unit is specifically configured to:

[0270] The second calculation unit is used to calculate the The second loss function L is obtained patchwise ;

[0271] Among them, q represents the total number of cropping boxes, M k is the total number of local generated images or local true value images cropped by the kth cropping frame; L kj 1(p1, p2) is the first loss function L1(p1, p2) obtained based on the j-th local generated image corresponding to the k-th cropping box and the j-th local true value image.

[0272] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0273] Fig.16 is a block diagram of an electronic device according to an exemplary embodiment. Fig.16 As shown, the electronic device includes but is not limited to: a processor 1601 , a memory 1602 , a network interface 1603 , an I / O controller 1604 , and a communication bus 1605 .

[0274] It should be noted that those skilled in the art can understand that Fig.16 The structure of the electronic device shown in the figure does not constitute a limitation on the electronic device, and the electronic device may include Fig.16 More or fewer components may be shown, or certain components may be combined, or the components may be arranged differently.

[0275] Combine the following Fig.16 The components of the electronic device 1600 are described in detail:

[0276] The processor 1601 is the control center of the electronic device. It uses various interfaces and lines to connect various parts of the entire electronic device. By running or executing software programs and / or modules stored in the memory 1602 and calling data stored in the memory 1602, it performs various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole. The processor 1601 may include one or more processing units; optionally, the processor 1601 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 1601.

[0277] The processor 1601 may be a central processing unit (CPU), or an application specific integrated circuit ASIC (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention;

[0278] The memory 1602 may include a memory, such as a high-speed random access memory (RAM) 16021 and a read-only memory (ROM) 16022, and may also include a large-capacity storage device 16023, such as at least one disk storage, etc. Of course, the electronic device may also include hardware required for other services.

[0279] The memory 1602 is used to store instructions executable by the processor 1601. The processor 1601 is configured to execute any step in the above model training method embodiment.

[0280] A wired or wireless network interface 1603 is configured to connect the electronic device 1600 to a network.

[0281] The processor 1601, the memory 1602, the network interface 1603 and the I / O controller 1604 may be interconnected via a communication bus 1605, which may be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc.

[0282] In an exemplary embodiment, the electronic device 1600 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to perform the above-mentioned model training method.

[0283] In an exemplary embodiment, a storage medium including instructions is also provided, such as a memory 1602 including instructions, and the instructions can be executed by a processor 1601 of an electronic device 1600 to complete the above method. Optionally, the storage medium can be a non-transitory computer-readable storage medium, for example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0284] In an exemplary embodiment, a computer program product is also provided, which can be directly loaded into the internal memory of a computer, such as the above-mentioned memory 1602, and contains software code. After being loaded and executed by a computer, the computer program can implement the method shown in any embodiment of the above-mentioned model training method.

[0285] It should be noted that the features described in the various embodiments in this specification can be replaced or combined with each other. For the device or system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0286] It should also be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0287] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0288] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A model training method, characterized in that: include: At least inputting the clothing image and the clothing segmentation map into a machine learning model, and obtaining a predicted clothing image through the machine learning model; The predicted clothing image is at least obtained by deforming the clothing to be tried contained in the clothing image according to the clothing segmentation map by the machine learning model, and the clothing segmentation map is obtained by deforming the clothing to be tried in the clothing image according to the model image and the human body key point map corresponding to the model image, and the model image includes a model for trying on the clothing to be tried; Obtaining a plurality of first partial images in the predicted clothing image, wherein different first partial images are located in different locations of the predicted clothing image, and areas of different first partial images are not completely the same, and the plurality of first partial images include a first partial image composed of detailed information of the clothing to be tried on, and the plurality of first partial images can be combined into the predicted clothing image; For each of the first partial images, a second partial image associated with the first partial image is obtained from a preset true value clothing image, wherein a position region of the second partial image in the true value clothing image is the same as a position region of the first partial image in the predicted clothing image, and the true value clothing image is a training target image output by the machine learning model; For each of the first partial images, based on the first partial image and the second partial image associated with the first partial image, obtain a first loss function corresponding to the first partial image; The machine learning model is trained using first loss functions corresponding to each of the first local images.

2. The model training method according to claim 1, characterized in that: Obtaining a predicted clothing image through the machine learning model includes: Get the generated image output by the machine learning model; Determining that the garment to be tried on is located in a first position area in the generated image; An image located in the first position area in the generated image is determined as the predicted clothing image.

3. The model training method according to claim 2, characterized in that: Also includes: An image located in the first position area of ​​a preset true value image is determined as the true value clothing image.

4. The model training method according to claim 2, characterized in that: Determining that the garment to be tried on is located in a first position area in the generated image includes: Determining a center position of the generated image; Taking the center position as the center of a set cropping frame, placing the set cropping frame in the generated image; The position area where the set cropping frame is located in the generated image is determined as the first position area.

5. The model training method according to any one of claims 1 to 4, characterized in that: Obtaining a plurality of first partial images in the predicted clothing image comprises: Get the preset cropping frames with different areas; For each of the cropping frames, the cropping frame is placed in the predicted clothing image to obtain a second position area where the cropping frame is located in the predicted clothing image, so as to obtain the second position areas corresponding to different cropping frames respectively; For each second position area, an image located in the second position area in the predicted clothing image is determined as the first partial image.

6. The model training method according to claim 5, characterized in that: Training the machine learning model using first loss functions respectively corresponding to the plurality of first partial images includes: Determine the mean of the first loss functions respectively corresponding to the multiple first partial images as the second loss function; The machine learning model is trained using the second loss function.

7. The model training method according to claim 6, characterized in that: Obtaining a first loss function based on the first partial image and the second partial image associated therewith, comprising: Based on the formula The first loss function is calculated, wherein p1 refers to the local generated image, p1(i) refers to the pixel value of the i-th pixel in the first local image, p2 refers to the local true value image, p2(i) refers to the pixel value of the i-th pixel in the second local image, and N is the total number of pixels contained in the first local image or the second local image.

8. The model training method according to claim 7, characterized in that: Determining the mean of the first loss functions respectively corresponding to the plurality of first partial images as the second loss function includes: Based on the formula The second loss function L is calculated patchwise ; Among them, q represents the total number of cropping boxes, M k is the total number of local generated images or local true value images cropped by the kth cropping frame; L kj 1(p1, p2) is the first loss function L1(p1, p2) obtained based on the j-th local generated image corresponding to the k-th cropping box and the j-th local true value image.

9. A model training device, characterized in that: include: A first acquisition module, configured to input at least a clothing image and a clothing segmentation map into a machine learning model, and obtain a predicted clothing image through the machine learning model; The predicted clothing image is at least obtained by deforming the clothing to be tried contained in the clothing image according to the clothing segmentation map by the machine learning model, and the clothing segmentation map is obtained by deforming the clothing to be tried in the clothing image according to the model image and the human body key point map corresponding to the model image, and the model image includes a model for trying on the clothing to be tried; A second acquisition module is used to obtain a plurality of first partial images in the predicted clothing image, wherein different first partial images are located in different locations of the predicted clothing image, and the areas of different first partial images are not completely the same, and the plurality of first partial images include a first partial image composed of detailed information of the clothing to be tried on, and the plurality of first partial images can be combined into the predicted clothing image; A third acquisition module is used to obtain, for each of the first partial images, a second partial image associated with the first partial image from a preset true value clothing image, wherein a position region of the second partial image in the true value clothing image is the same as a position region of the first partial image in the predicted clothing image, and the true value clothing image is a training target image output by the machine learning model; a fourth acquisition module, configured to obtain, for each of the first partial images, a first loss function corresponding to the first partial image based on the first partial image and the second partial image associated with the first partial image; A training module is used to train the machine learning model through first loss functions corresponding to multiple first local images respectively.

10. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the model training method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Virtual fitting method, system and device and storage medium

    CN114119905A

  • Image processing method, image processing system and electronic equipment

    CN114266695A