A virtual fitting method, device, equipment and storage medium

By disassembling and deforming the target clothing and combining it with posture information to generate a realistic virtual try-on effect, the problem of clothing not fitting the posture of the person in the existing technology is solved, and a more realistic virtual try-on effect is achieved.

CN113283953BActive Publication Date: 2025-12-30BEIJING AIBI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202010104448.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-02-20
Publication Date
2025-12-30
Estimated Expiration
2040-02-20

AI Technical Summary

Technical Problem

In existing virtual try-on methods, the clothes in the generated try-on images are difficult to match with the person's posture, resulting in an unrealistic try-on effect and failing to provide consumers with reliable reference opinions.

Method used

By splitting the target clothing to be tried on into multiple clothing sub-images, and then deforming each clothing sub-image according to the pose information of the target person, the target clothing image is generated by combining the mask image, and finally a realistic try-on effect image is generated.

Benefits of technology

The virtual try-on effect has been improved, making the try-on images more realistic, better matching the posture of the target person, and providing reliable reference opinions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113283953B_ABST
    Figure CN113283953B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a virtual fitting method, device and equipment and a storage medium. The method comprises the following steps: obtaining a basic human body picture and a basic clothes picture, the basic human body picture comprising a target person, and the basic clothes picture comprising a target clothes to be fitted on the target person; extracting reference information of the target person in the basic human body picture, and generating a target human body picture according to the reference information, wherein the reference information at least comprises posture information of the target person; performing splitting processing on the target clothes to obtain a plurality of clothes sub-pictures, different clothes sub-pictures comprising different parts of the target clothes; performing morphing processing on the plurality of clothes sub-pictures according to the posture information to obtain a plurality of clothes sub-target pictures and their respective mask images; generating a target clothes picture according to the plurality of clothes sub-target pictures and their respective mask images; and generating a fitting effect picture according to the target human body picture and the target clothes picture. The method can obtain a more realistic fitting effect picture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to a virtual fitting method, apparatus, device, and storage medium. Background Technology

[0002] With the rapid development of internet technology, online shopping has become one of the main ways for consumers to shop. However, when buying clothes online, consumers often find it difficult to accurately predict how the clothes will look on them. To solve this problem, virtual try-on technology has emerged. This technology can organically combine images of clothing with images of people to generate images of people trying on clothes.

[0003] Current virtual try-on methods generally involve directly deforming the entire garment in a picture of the person to be tried on, and then combining the deformed garment with the picture of the person to generate the final try-on effect. However, the inventors of this application have found that the try-on effect generated in this way does not accurately match the posture of the person, resulting in an unrealistic try-on effect and failing to provide consumers with reliable reference opinions.

[0004] In conclusion, improving the effect of virtual try-on to make the final try-on images more realistic has become an urgent problem to be solved. Summary of the Invention

[0005] This application provides a virtual fitting method, apparatus, device, and storage medium, which can effectively improve the effect of virtual fitting and make the final generated fitting effect image more realistic.

[0006] In view of the above, the first aspect of this application provides a virtual try-on method, the method comprising:

[0007] Obtain basic human body images and basic clothing images; the basic human body images include the target person, and the basic clothing images include the target clothing to be tried on by the target person;

[0008] Extract reference information related to the target person from the base human image, and generate a target human image based on the reference information; the reference information includes at least the pose information of the target person.

[0009] The target clothing in the base clothing image is split into N clothing sub-images; where N is an integer greater than 1, and different clothing sub-images include different parts of the target clothing.

[0010] The N clothing sub-images are deformed based on the pose information to obtain N clothing sub-target images and their corresponding mask images; a target clothing image is generated based on the N clothing sub-target images and their corresponding mask images.

[0011] A try-on effect image is generated based on the target human body image and the target clothing image.

[0012] Optionally, the step of extracting the head information and posture information of the target person from the base human image, and generating the target human image based on the head information and posture information, includes:

[0013] The head and body information of the target person are extracted from the base human image using a human body analysis model.

[0014] The pose information is extracted from the base human image using the Openpose model;

[0015] The target human image is synthesized using a human body synthesis model based on the head information, body shape information, and posture information.

[0016] Optionally, the step of splitting the target clothing in the base clothing image to obtain N clothing sub-images includes:

[0017] The target clothing is split horizontally according to a preset ratio to obtain the N clothing sub-images.

[0018] Optionally, the step of splitting the target clothing in the base clothing image to obtain N clothing sub-images includes:

[0019] The target garment is divided into N sub-images based on its composition structure.

[0020] Optionally, the step of deforming the N clothing sub-images according to the pose information to obtain N clothing sub-target images and the mask images corresponding to each of the N clothing sub-target images includes:

[0021] The first image processing model determines the N clothing sub-target images and their corresponding mask images based on the pose information and the N clothing sub-images.

[0022] Optionally, generating a target clothing image based on the N clothing sub-target images and the mask images corresponding to each of the N clothing sub-target images includes:

[0023] For each clothing sub-target image, a clothing sub-try-on image is determined based on the clothing sub-target image and its corresponding mask image; the clothing sub-try-on image is used to characterize the try-on effect of the target clothing part corresponding to the clothing sub-target image on the target person.

[0024] By stitching together the corresponding clothing try-on images of the N clothing sub-target images, the target clothing image is obtained.

[0025] Optionally, generating a try-on effect image based on the target human body image and the target clothing image includes:

[0026] A target mask image is determined using a second image processing model based on the target human body image and the target clothing image; the target mask image is used to represent the occlusion relationship between the target person and the target clothing.

[0027] The try-on effect image is generated based on the target human body image, the target clothing image, and the target mask image.

[0028] A second aspect of this application provides a virtual fitting device, the device comprising:

[0029] The basic image acquisition module is used to acquire basic human body images and basic clothing images; the basic human body images include the target person, and the basic clothing images include the target clothing to be tried on by the target person;

[0030] A human body image processing module is used to extract reference information related to the target person from the basic human body image, and generate a target human body image based on the reference information; the reference information includes at least the pose information of the target person.

[0031] The clothing image segmentation module is used to segment the target clothing in the base clothing image to obtain N clothing sub-images; where N is an integer greater than 1, and different clothing sub-images include different parts of the target clothing;

[0032] The clothing image processing module is used to perform deformation processing on the N clothing sub-images according to the pose information to obtain N clothing sub-target images and the mask images corresponding to each of the N clothing sub-target images; and to generate a target clothing image based on the N clothing sub-target images and the mask images corresponding to each of the N clothing sub-target images.

[0033] The fitting effect determination module is used to generate a fitting effect image based on the target human body image and the target clothing image.

[0034] Optionally, the clothing image splitting module is specifically used for:

[0035] The target clothing is split horizontally according to a preset ratio to obtain the N clothing sub-images;

[0036] Alternatively, the target clothing can be broken down according to its composition to obtain the N clothing sub-images.

[0037] A third aspect of this application provides an apparatus comprising: a processor and a memory;

[0038] The memory is used to store computer programs;

[0039] The processor is configured to invoke the computer program to execute the virtual try-on method described in the first aspect.

[0040] A fourth aspect of this application provides a computer-readable storage medium for storing a computer program for performing the virtual try-on method described in the first aspect.

[0041] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:

[0042] This application provides a virtual try-on method. This method innovatively splits the target garment to be tried on into multiple sub-images, each containing different parts of the garment. Then, based on the pose information of the target person, each sub-image is deformed. Using a mask image determined during the deformation process, the multiple deformed sub-images are stitched together to obtain a complete target garment image that closely matches the target person's pose. Based on this target garment image, a try-on effect image is generated to simulate the target person trying on the garment, ensuring a more realistic try-on effect. Attached Figure Description

[0043] Figure 1 A flowchart illustrating the virtual try-on method provided in this application embodiment;

[0044] Figure 2 This is a schematic diagram of the structure of the virtual fitting device provided in the embodiments of this application;

[0045] Figure 3 This is a schematic diagram of the server structure provided in an embodiment of this application;

[0046] Figure 4 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation

[0047] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0048] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0049] Existing virtual try-on methods generally involve directly deforming the entire garment being tried on, and then combining the deformed garment with an image of the person to generate a try-on effect image. However, the inventors of this application have found that the clothing try-on effect simulated by the above methods is not ideal. The clothing tried on in the try-on effect image usually does not fit the person's posture naturally, and such try-on effect images are difficult to provide consumers with reliable reference opinions.

[0050] To address the problems existing in the prior art, this application provides a virtual try-on method. This method takes a different approach by splitting the target clothing to be tried on into smaller parts and deforming each sub-image of the clothing based on the posture information of the target person. This allows the target clothing to fully fit the posture of the target person, simulating a more realistic clothing try-on effect.

[0051] Specifically, in the virtual try-on method provided in this application embodiment, a basic human body image and a basic clothing image are first obtained. The basic human body image includes the target person, and the basic clothing image includes the target clothing to be tried on by the target person. Then, reference information related to the target person is extracted from the basic human body image, and a target human body image is generated based on this reference information. Here, the reference information includes at least the target person's posture information. The target clothing in the basic clothing image is split into N (N is an integer greater than 1) clothing sub-images, each containing different parts of the target clothing. Furthermore, the N clothing sub-images are deformed according to the target person's posture information to obtain a corresponding clothing sub-target image and a corresponding mask image for each clothing sub-target image. Using these N clothing sub-target images and their corresponding mask images, a complete target clothing image that fits the target person's posture is generated. Finally, a try-on effect image representing the target person trying on the target clothing is generated based on the target human body image and the target clothing image.

[0052] The above method innovatively splits the target clothing to be tried on into multiple sub-images, each containing different parts of the target clothing. Then, based on the pose information of the target person, each sub-image is deformed. Using the mask image determined during the deformation process, the multiple sub-images are stitched together to obtain a complete image of the target clothing that closely matches the pose of the target person. Based on this image, a try-on effect image is generated to simulate the target person trying on the target clothing, ensuring that the try-on effect shown in the image is more realistic.

[0053] It should be noted that the virtual try-on method provided in this application embodiment can be applied to various devices with image processing capabilities, such as terminal devices and servers. Terminal devices can include smartphones, computers, tablets, etc. Servers can be application servers or web servers; in specific deployments, servers can be standalone servers or cluster servers.

[0054] The virtual try-on method provided in this application will be described in detail below through embodiments.

[0055] See Figure 1 , Figure 1 This is a flowchart illustrating the virtual try-on method provided in an embodiment of this application. For ease of description, the following embodiment uses a server as the execution entity as an example. Figure 1 As shown, the virtual try-on method includes the following steps:

[0056] Step 101: Obtain a basic human body image and a basic clothing image; the basic human body image includes the target person, and the basic clothing image includes the target clothing to be tried on by the target person.

[0057] When a user needs to simulate the effect of a target person trying on target clothes, the user can upload a basic human image including the target person to the server. At the same time, the server can obtain a basic clothing image including the target clothes to be tried on via the network.

[0058] Taking a user purchasing clothes through a shopping app on a mobile device as an example, when a user wants to see how a target person would look wearing the clothes they are currently browsing, the user can click the try-on control on the clothing display screen and select a basic human image containing the target person from the images stored locally on the mobile device. After the user confirms, the mobile device can transmit this basic human image to the server via the network. Simultaneously, the server can retrieve the corresponding basic clothing image, which is usually a pre-uploaded image from the merchant specifically for simulating the try-on effect, based on the webpage address currently being viewed by the user.

[0059] It should be understood that, in order to ensure a good fitting effect, the above-mentioned basic human body image is usually an image that only includes the target person or an image in which the target person occupies the main part, and the image needs to completely retain the body part of the target person used to try on the target clothes. For example, assuming that the target clothes to be tried on are tops, the basic human body image needs to include the complete upper body of the target person.

[0060] It should be understood that in practical applications, other methods can be used to obtain the aforementioned basic human body images and basic clothing images, depending on the actual situation. For example, when the virtual try-on method provided in this application is executed by a terminal device, the terminal device can directly obtain the basic human body image selected by the user or capture the user's image in real time as the basic human body image, and obtain the basic clothing image by communicating with the server. This application does not limit the implementation method of obtaining the basic human body images and basic clothing images.

[0061] Step 102: Extract reference information related to the target person from the basic human body image, and generate a target human body image based on the reference information; the reference information includes at least the pose information of the target person.

[0062] After the server obtains the basic human image, it can extract reference information related to the target person from the basic human image, and then generate the target human image based on the reference information; the reference information here includes at least the pose information of the target person.

[0063] It should be understood that in practical applications, in order to ensure that the final generated image of the target person wearing the target clothing is more realistic and can provide users with more reliable reference information, the reference information here may also include the target person's head information and / or body shape information.

[0064] It should be noted that the target human body image mentioned above is different from the basic human body image. The person shown is usually simulated based on the head information, body shape information and posture information of the target person in the basic human body image. When generating the target human body image, the clothing information of the target person in the basic human body image is usually not referenced, so that the target clothing can be directly simulated and tried on later based on the target human body image.

[0065] It should be noted that the above posture information is actually determined based on the key point information of the target person in the base human body image. This key point information is the position information of the joints on the target person that can determine the posture of the target person. By extracting multiple key point information of the target person from the base human body image, the posture information of the target person can be determined based on these multiple key point information.

[0066] In practice, the server can extract head and body information from a basic human image using a human body parsing model, extract pose information from the basic human image using an Openpose model, and then synthesize a target human image using a human body synthesis model based on the head, body, and pose information.

[0067] That is, the server can input a basic human image P into a human body analysis model to obtain the head and body information of the target person extracted by the human body analysis model from the basic human image P; input the basic human image P into an Openpose model to obtain the pose information of the target person extracted by the Openpose model from the basic human image P; then, input the extracted head information, body information and pose information into a human body synthesis model, and use the human body synthesis model to organically combine the head information and pose information to generate the target human image P'.

[0068] It should be noted that the aforementioned human body analysis model, Openpose model, and human body synthesis model are all relatively mature neural network models in the existing technology. This application directly calls these mature models to generate target human body image P' based on the basic human body image P.

[0069] Step 103: Segment the target clothing in the base clothing image to obtain N clothing sub-images; where N is an integer greater than 1, and different clothing sub-images include different parts of the target clothing.

[0070] After the server obtains the basic clothing image, it can split the target clothing in the basic clothing image to obtain N clothing sub-images, each of which includes different parts of the target clothing.

[0071] In one possible implementation, the server can split the target clothing image horizontally according to a preset ratio to obtain the aforementioned N clothing sub-images. For example, assuming the preset ratio is 1:2:1, the server can divide the target clothing in the base clothing image into three parts horizontally in a 1:2:1 ratio, namely clothing sub-image a, clothing sub-image b, and clothing sub-image c.

[0072] It should be understood that in practical applications, the aforementioned preset ratio can be set according to actual needs, and this application does not specifically limit the preset ratio. Furthermore, in addition to splitting the target clothing horizontally, the server can also split the target clothing in other directions. For example, it can split the target clothing vertically according to a preset ratio; or, for example, it can determine the tilt direction of the target person's posture based on the target person's posture information, and then split the target clothing along that tilt direction according to a preset ratio, etc. This application also does not limit the direction of splitting the target clothing.

[0073] In another possible implementation, the server can break down the target garment according to its structural composition, resulting in N sub-images. For example, assuming the target garment is a long-sleeved top, the server can break it down into sub-images including the left sleeve, the main body of the top, and the right sleeve. Similarly, if the target garment is trousers, the server can break it down into sub-images including the left and right legs. And if the target garment is a sleeved dress, the server can break it down into sub-images including the left sleeve, the upper body, the lower body, and the right sleeve.

[0074] It should be understood that in practical applications, the server may also use other methods to split the target clothing in the base clothing image, and this application does not specifically limit the splitting method of the target clothing.

[0075] It should be noted that in practical applications, step 102 can be executed first and then step 103, or step 103 can be executed first and then step 102, or steps 102 and 103 can be executed simultaneously. This application does not impose any restrictions on the specific execution order of steps 102 and 103.

[0076] Step 104: Perform deformation processing on the N clothing sub-images according to the pose information to obtain N clothing sub-target images and the mask images corresponding to each of the N clothing sub-target images.

[0077] After extracting the pose information of the target person in the basic human image in step 102, and splitting the target clothing in the basic clothing image into N clothing sub-images in step 103, the server can further deform the N clothing sub-images according to the pose information of the target person, thereby obtaining the clothing sub-target images corresponding to each of the N clothing sub-images. Each clothing sub-target image represents the way the target clothing portion of its corresponding clothing sub-image conforms to the pose of the target person. Furthermore, during the deformation processing of the clothing sub-images, the server can also determine the corresponding mask image for each clothing sub-target image. The mask image corresponding to each clothing sub-target image represents the occlusion of the target clothing portion in that clothing sub-target image.

[0078] In practice, the server can use the first image processing model to determine N clothing sub-target images and the corresponding mask images for each of the N clothing sub-target images based on the pose information of the target person and N clothing sub-images.

[0079] That is, the server can input N clothing sub-images and the pose information of the target person into the first image processing model. The first image processing model processes these N clothing sub-images and the pose information of the target person accordingly, and outputs N deformed clothing sub-target images and their corresponding mask images. For example, assuming the clothing sub-images input to the first image processing model are clothing sub-image a, clothing sub-image b, and clothing sub-image c, the first image processing model will correspondingly output clothing sub-target image a' corresponding to clothing sub-image a, clothing sub-target image b' corresponding to clothing sub-image b, and clothing sub-target image c' corresponding to clothing sub-image c, as well as the mask image a' corresponding to clothing sub-target image a'. m The mask image b corresponding to the clothing sub-target image b' m and the mask image c corresponding to the clothing sub-target image c' m .

[0080] It should be understood that the first image processing model mentioned above is a pre-trained neural network model, which can perform deformation processing on the clothing sub-images containing various parts of the target clothing according to the input pose information, and determine the occlusion status of the target clothing parts in each clothing sub-target image obtained after deformation processing.

[0081] In practical applications, the first image processing model can be trained as follows: A training sample set containing a large number of training samples is obtained. Each training sample includes multiple clothing sub-images obtained by splitting the clothing to be tried on, as well as a human body image. Before training the first image processing model, the pose information of the person in the human body image from the training samples can be extracted using the OpenPose model. During training, this pose information and the multiple clothing sub-images from the training samples are input into a pre-built neural network model. The neural network model performs deformation processing on these multiple clothing sub-images according to the pose information, obtaining N clothing sub-target images and their corresponding mask images. Then, based on the mask images corresponding to these N clothing sub-target images, the N clothing sub-target images are stitched together to obtain a complete deformed image of the clothing to be tried on. The degree of deformation of this image is compared with the degree of deformation of the clothing in the human body image from the training samples. The model parameters of the pre-built neural network model are adjusted based on the difference in the degree of deformation between the two. Thus, the pre-built neural network model is iteratively trained using each training sample in the training sample set until the training termination condition is met. Step 105: Generate the target clothing image based on the N clothing sub-target images and the mask images corresponding to each of the N clothing sub-target images.

[0082] After obtaining N clothing sub-target images and their corresponding mask images through deformation processing, the N clothing sub-target images can be stitched together based on their respective mask images to obtain a complete target clothing image that fits the pose of the target person.

[0083] In practice, for each clothing sub-target image, the server can determine the corresponding clothing sub-try-on image based on the clothing sub-target image and its corresponding mask image. The clothing sub-try-on image can represent the try-on effect of the target clothing part corresponding to the clothing sub-target image on the target person. Then, the clothing sub-try-on images corresponding to these N clothing sub-target images are stitched together to obtain the target clothing image.

[0084] Let N clothing sub-target images be clothing sub-target image a', clothing sub-target image b', and clothing sub-target image c', and let the mask image corresponding to each of the N clothing sub-target images be mask image a'. m , mask image b m and mask image c m For example, the server can determine the target clothing image W' using equation (1):

[0085] W' = a m *a'+b m *b'+c m *c' (1)

[0086] Step 106: Generate a try-on effect image based on the target human body image and the target clothing image.

[0087] After obtaining the target human body image in step 102 and the target clothing image in step 105, the server can organically combine the target human body image and the target clothing image to generate a try-on effect image, which shows the effect of the target person in the basic human body image trying on the target clothing.

[0088] In practice, the server can determine a target mask image based on the target human body image and the target clothing image using a second image processing model. This target mask image can represent the occlusion relationship between the target person and the target clothing. Then, the server generates a try-on effect image based on the target human body image, the target clothing image, and the target mask image.

[0089] That is, the server inputs the target human image P' and the target clothing image W' into the second image processing model. The second image processing model analyzes and processes the target human image P' and the target clothing image W' accordingly, and outputs a target mask image β. This target mask image β can represent the occlusion relationship between the person in the target human image and the target clothing. Furthermore, the server can generate a try-on effect image I based on the target human image P', the target clothing image W', and the target mask image β using equation (2).

[0090] I=β*P'+(1-β)*W' (2)

[0091] In practical applications, the second image processing model can be trained as follows: A training sample set containing a large number of training samples is obtained. Each training sample includes an image of the clothing to be tried on (after deformation processing) and an image of the human body. Before training the second image processing model, the human body image can be processed using step 102 above to generate a target human body image based on the head, body shape, and posture information of the person in the image. During training, the target human body image and the clothing image to be tried on from the training samples are input into a pre-built neural network model. This neural network model processes the target human body image and the clothing image accordingly and outputs a corresponding mask image. Then, using this mask image, an image of the person wearing the clothing is generated based on the target human body image and the clothing image from the training samples. This image is compared with the human body images in the training samples, and the model parameters of the pre-built neural network model are adjusted based on the difference in the fit between the clothing and the person in the two images. In this way, the pre-built neural network model is repeatedly trained using each training sample in the training sample set until the training termination condition is met.

[0092] The aforementioned virtual try-on method innovatively breaks down the target garment to be tried on, obtaining multiple sub-images of the garment, each containing different parts of the target garment. Then, based on the target person's posture information, each sub-image of the garment is deformed. Based on the mask image determined during the deformation process, the multiple sub-images of the garment obtained after deformation are stitched together to obtain a complete image of the target garment that can better fit the target person's posture. Based on this image of the target garment, a try-on effect image is generated to simulate the target person trying on the target garment, ensuring that the try-on effect image shows a more realistic try-on effect.

[0093] In addition to the virtual try-on method described above, this application also provides a virtual try-on device so that the above-mentioned virtual try-on method can be applied and implemented in practice.

[0094] See Figure 2 , Figure 2 This is a schematic diagram of the structure of a virtual fitting device 200 provided in an embodiment of this application, as shown below. Figure 2 As shown, the virtual fitting device includes:

[0095] The basic image acquisition module 201 is used to acquire basic human body images and basic clothing images; the basic human body images include the target person, and the basic clothing images include the target clothing to be tried on by the target person;

[0096] The human body image processing module 202 is used to extract reference information related to the target person from the basic human body image and generate a target human body image based on the reference information; the reference information includes at least the pose information of the target person.

[0097] The clothing image splitting module 203 is used to split the target clothing in the base clothing image to obtain N clothing sub-images; where N is an integer greater than 1, and different clothing sub-images include different parts of the target clothing;

[0098] The clothing image processing module 204 is used to perform deformation processing on the N clothing sub-images according to the pose information to obtain N clothing sub-target images and the mask images corresponding to each of the N clothing sub-target images; and to generate a target clothing image based on the N clothing sub-target images and the mask images corresponding to each of the N clothing sub-target images.

[0099] The fitting effect determination module 205 is used to generate a fitting effect image based on the target human body image and the target clothing image.

[0100] Optionally, the human image processing module 202 is specifically used for:

[0101] The head and body information of the target person are extracted from the base human image using a human body analysis model.

[0102] The pose information is extracted from the base human image using the Openpose model;

[0103] The target human image is synthesized using a human body synthesis model based on the head information, body shape information, and posture information.

[0104] Optionally, the clothing image splitting module 203 is specifically used for:

[0105] The target clothing is split horizontally according to a preset ratio to obtain the N clothing sub-images.

[0106] Optionally, the clothing image splitting module 203 is specifically used for:

[0107] The target garment is divided into N sub-images based on its composition structure.

[0108] Optionally, the clothing image processing module 204 is specifically used for:

[0109] The first image processing model determines the N clothing sub-target images and their corresponding mask images based on the pose information and the N clothing sub-images.

[0110] Optionally, the clothing image processing module 204 is specifically used for:

[0111] For each clothing sub-target image, a clothing sub-try-on image is determined based on the clothing sub-target image and its corresponding mask image; the clothing sub-try-on image is used to characterize the try-on effect of the target clothing part corresponding to the clothing sub-target image on the target person.

[0112] By stitching together the corresponding clothing try-on images of the N clothing sub-target images, the target clothing image is obtained.

[0113] Optionally, the fitting effect determination module 205 is specifically used for:

[0114] A target mask image is determined using a second image processing model based on the target human body image and the target clothing image; the target mask image is used to represent the occlusion relationship between the target person and the target clothing.

[0115] The try-on effect image is generated based on the target human body image, the target clothing image, and the target mask image.

[0116] The aforementioned virtual try-on device innovatively breaks down the target garment to be tried on, obtaining multiple sub-images of the garment, each containing different parts of the target garment. Then, based on the target person's posture information, each sub-image of the garment is deformed. Based on the mask image determined during the deformation process, the multiple sub-images of the garment obtained after deformation are stitched together to obtain a complete image of the target garment that fits the target person's posture well. Based on this image of the target garment, a try-on effect image is generated to simulate the target person trying on the target garment, ensuring that the try-on effect image shows a more realistic try-on effect.

[0117] This application also provides a device for virtual try-on, which may specifically be a server or a terminal device. The server and terminal devices provided in this application will be described below from the perspective of hardware physicalization.

[0118] See Figure 3 , Figure 3This is a schematic diagram of the structure of a server 300 provided in an embodiment of this application. The server 300 can vary significantly due to different configurations or performance, and may include one or more central processing units (CPUs) 322 (e.g., one or more processors) and memory 332, and one or more storage media 330 (e.g., one or more mass storage devices) for storing application programs 342 or data 344. The memory 332 and storage media 330 can be temporary or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the server. Furthermore, the CPU 322 may be configured to communicate with the storage media 330 and execute the series of instruction operations stored in the storage media 330 on the server 300.

[0119] Server 300 may also include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input / output interfaces 358, and / or one or more operating systems 341, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0120] The steps performed by the server in the above embodiments can be based on this Figure 3 The server structure shown.

[0121] CPU 322 is used to perform the following steps:

[0122] Obtain basic human body images and basic clothing images; the basic human body images include the target person, and the basic clothing images include the target clothing to be tried on by the target person;

[0123] Extract reference information related to the target person from the base human image, and generate a target human image based on the reference information; the reference information includes at least the pose information of the target person.

[0124] The target clothing in the base clothing image is split into N clothing sub-images; where N is an integer greater than 1, and different clothing sub-images include different parts of the target clothing.

[0125] The N clothing sub-images are deformed based on the pose information to obtain N clothing sub-target images and their corresponding mask images; a target clothing image is generated based on the N clothing sub-target images and their corresponding mask images.

[0126] A try-on effect image is generated based on the target human body image and the target clothing image.

[0127] Optionally, CPU 322 can also be used to execute any of the steps of the virtual try-on method provided in the embodiments of this application.

[0128] See Figure 4 , Figure 4 This is a schematic diagram of a terminal device provided in an embodiment of this application. For ease of explanation, only the parts related to the embodiments of this application are shown; for specific technical details not disclosed, please refer to the method section of the embodiments of this application. The terminal can be any terminal device including computers, tablet computers, personal digital assistants (PDAs), etc. Taking a computer as an example:

[0129] Figure 4 This is a block diagram illustrating a portion of the structure of a computer associated with the terminal provided in an embodiment of this application. (Reference) Figure 4 The computer includes: a radio frequency (RF) circuit 410, a memory 420, an input unit 430, a display unit 440, a sensor 450, an audio circuit 460, a wireless fidelity (WiFi) module 470, a processor 480, and a power supply 490, among other components. Those skilled in the art will understand that... Figure 4 The computer architecture shown does not constitute a limitation on the computer and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0130] The memory 420 can be used to store software programs and modules. The processor 480 executes various computer functions and data processing by running the software programs and modules stored in the memory 420. The memory 420 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer (such as audio data, telephone directory, etc.). In addition, the memory 420 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0131] The processor 480 is the control center of the computer, connecting various parts of the computer through various interfaces and lines. It performs various computer functions and processes data by running or executing software programs and / or modules stored in the memory 420, and by calling data stored in the memory 420, thereby providing overall monitoring of the computer. Optionally, the processor 480 may include one or more processing units; preferably, the processor 480 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 480.

[0132] In this embodiment of the application, the processor 480 included in the terminal also has the following functions:

[0133] Obtain basic human body images and basic clothing images; the basic human body images include the target person, and the basic clothing images include the target clothing to be tried on by the target person;

[0134] Extract reference information related to the target person from the base human image, and generate a target human image based on the reference information; the reference information includes at least the pose information of the target person.

[0135] The target clothing in the base clothing image is split into N clothing sub-images; where N is an integer greater than 1, and different clothing sub-images include different parts of the target clothing.

[0136] The N clothing sub-images are deformed based on the pose information to obtain N clothing sub-target images and their corresponding mask images; a target clothing image is generated based on the N clothing sub-target images and their corresponding mask images.

[0137] A try-on effect image is generated based on the target human body image and the target clothing image.

[0138] Optionally, the processor 480 is further configured to execute steps of any implementation of the virtual try-on method provided in the embodiments of this application.

[0139] This application also provides a computer-readable storage medium for storing program code that executes any one of the implementation methods of the virtual try-on method described in the foregoing embodiments.

[0140] This application also provides a computer program product including instructions that, when run on a computer, cause the computer to execute any one of the implementation methods of the virtual try-on method described in the foregoing embodiments.

[0141] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0142] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0143] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0144] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0145] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing computer programs.

[0146] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0147] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A virtual fitting method, characterized by, The method comprises: obtaining a basic human body picture and a basic clothes picture; the basic human body picture comprises a target person, and the basic clothes picture comprises target clothes to be tried on the target person; extracting reference information related to the target person in the basic human body picture, and generating a target human body picture according to the reference information; the reference information at least comprises posture information of the target person; performing splitting processing on the target clothes in the basic clothes picture to obtain N clothes sub-pictures; N is an integer greater than 1, and different clothes sub-pictures comprise different parts of the target clothes; performing deformation processing on the N clothes sub-pictures according to the posture information to obtain N clothes sub-target pictures and mask images corresponding to the N clothes sub-target pictures respectively; and generating a target clothes picture according to the N clothes sub-target pictures and the mask images corresponding to the N clothes sub-target pictures respectively; generating a trying-on effect picture according to the target human body picture and the target clothes picture.

2. The method of claim 1, wherein, The extracting reference information related to the target person in the basic human body picture and the generating a target human body picture according to the reference information comprises: extracting head information and body shape information of the target person from the basic human body picture through a human body analysis model; extracting the posture information from the basic human body picture through an Openpose model; synthesizing the target human body picture according to the head information, the body shape information and the posture information through a human body synthesis model.

3. The method of claim 1, wherein, The splitting processing on the target clothes in the basic clothes picture to obtain N clothes sub-pictures comprises: splitting the target clothes in a horizontal direction according to a preset proportion to obtain the N clothes sub-pictures; or splitting the target clothes according to a composition structure of the target clothes to obtain the N clothes sub-pictures.

4. The method of claim 1, wherein, The deformation processing on the N clothes sub-pictures according to the posture information to obtain N clothes sub-target pictures and mask images corresponding to the N clothes sub-target pictures respectively comprises: determining the N clothes sub-target pictures and the mask images corresponding to the N clothes sub-target pictures respectively according to the posture information and the N clothes sub-pictures through a first image processing model.

5. The method of claim 1, wherein, The generating a target clothes picture according to the N clothes sub-target pictures and the mask images corresponding to the N clothes sub-target pictures respectively comprises: for each clothes sub-target picture, determining a clothes sub-trying-on picture corresponding to the clothes sub-target picture according to the clothes sub-target picture and the mask image corresponding to the clothes sub-target picture; the clothes sub-trying-on picture is used to represent a trying-on effect of a target clothes part corresponding to the clothes sub-target picture on the target person; splicing the clothes sub-trying-on pictures corresponding to the N clothes sub-target pictures respectively to obtain the target clothes picture.

6. The method of claim 1, wherein, The generating a trying-on effect picture according to the target human body picture and the target clothes picture comprises: determine a target mask image according to the target human body picture and the target clothes picture through a second image processing model; the target mask image is used to represent an occlusion relationship between the target character and the target clothes; generate the try-on effect picture according to the target human body picture, the target clothes picture and the target mask image.

7. A virtual fitting device, characterized in that, The device comprises: a basic picture acquisition module, configured to acquire a basic human body picture and a basic clothes picture; the basic human body picture comprises a target character, and the basic clothes picture comprises target clothes to be tried on by the target character; a human body picture processing module, configured to extract reference information related to the target character in the basic human body picture, and generate a target human body picture according to the reference information; the reference information at least comprises posture information of the target character; a clothes picture splitting module, configured to split the target clothes in the basic clothes picture to obtain N clothes sub-pictures; N is an integer greater than 1, and different clothes sub-pictures comprise different parts of the target clothes; a clothes picture processing module, configured to deform the N clothes sub-pictures according to the posture information to obtain N clothes sub-target pictures and mask images corresponding to the N clothes sub-target pictures respectively; and generate a target clothes picture according to the N clothes sub-target pictures and the mask images corresponding to the N clothes sub-target pictures respectively; a try-on effect determination module, configured to generate a try-on effect picture according to the target human body picture and the target clothes picture.

8. The apparatus of claim 7, wherein, The clothes picture splitting module is specifically configured to: split the target clothes in the horizontal direction according to a preset ratio to obtain the N clothes sub-pictures; or split the target clothes according to a component structure of the target clothes to obtain the N clothes sub-pictures.

9. An apparatus, comprising: The device comprises a processor and a memory; The memory is configured to store a computer program; The processor is configured to call the computer program to execute the virtual fitting method in any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium is configured to store a computer program, and the computer program is configured to execute the virtual fitting method in any one of claims 1 to 6.