Method for generating virtual fitting image and electronic equipment
By using the matching relationship between the target characters and the key points of the clothing in the virtual try-on technology to guide the model to generate virtual try-on images, the problem of spatial information destruction in complex postures and action scenes in the existing technology is solved, and the authenticity and display effect of the generation effect are improved.
Patent Information
- Application Number
- CN202411731042.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-05-06
AI Technical Summary
When existing virtual try-on techniques deal with models trying on pictures/videos containing complex poses, actions or scene changes, they can easily destroy the spatial information in the original model pictures/videos and affect the generation effect.
By determining multiple human posture key points marked by the target character in the source image, multiple clothing key points marked by the clothing image of the second clothing, and the matching relationship between the clothing key points and the human posture key points, these information are input into the pre-trained model to guide the model to generate a virtual try-on image.
It improves the generation effect of virtual try-on images, ensures that the key points of the clothing match more appropriate positions and areas on the target person, and enhances the authenticity and display effect of the image.
Smart Images

Figure CN119941923A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of virtual try-on, and in particular to a method and electronic device for generating virtual try-on images. Background Art
[0002] Virtual fitting is a technical means to view the effect of trying on clothing through virtual technology. In one scenario, the specific virtual fitting effect is that, given a model picture / video 1, the image content is a fitting effect picture / video of a model character S trying on a certain clothing A; in addition, given pictures of clothing B, C, D, etc., when the user selects the picture of clothing B, a virtual fitting picture / video 2 of the aforementioned model character S trying on clothing B can be generated through the model. In the above process, in the fitting effect pictures / videos before and after, except for the clothing part changing from A to B, other parts including the posture and movements of the model character remain unchanged.
[0003] In order to achieve the above purpose, in traditional implementation methods, most of them rely on masks (or "masks" or "masks", etc.) to determine the fitting area in model image / video 1, and then generate the aforementioned model image / video 2 based on the image of clothing B and the virtual fitting image / video 1 with the mask. This method works better when processing simple model images / videos, but it is easy to destroy the spatial information in the original model image / video when processing model fitting images / videos with complex postures, movements or scene changes, thus affecting the final generation effect. For example, assuming that in the aforementioned model image / video 1, Figure 1 As shown in (A), the model's pose is to put her hands on her chest, blocking part of the clothing image. At this time, when generating the mask, the hand image may also be covered by the mask. At this time, when generating the model image / video 2, the model's hand image needs to be completely generated by the model, but the generated effect may be relatively poor. For example, Figure 1 As shown in (B), the image of the model’s hand is obviously not realistic enough, which affects the display effect. Summary of the invention
[0004] The present application provides a method and electronic device for generating a virtual try-on image, which can enhance the generation effect of the virtual try-on image.
[0005] This application provides the following solutions:
[0006] A method for generating a virtual try-on image, comprising:
[0007] Determine a source image, wherein the image content of the source image includes an image of a target person wearing a first garment;
[0008] Determining a clothing image of a second clothing to be replaced in the source image for display;
[0009] Determining a plurality of clothing key points annotated based on the clothing image of the second clothing, a plurality of human body posture key points annotated based on the target person in the source image, and a matching relationship between the clothing key points and the human body posture key points;
[0010] The source image, the clothing image of the second clothing and the matching relationship information are input into a pre-trained model, so that in the process of the model generating a virtual try-on image, the matching relationship information is used to provide guidance information for the model; the virtual try-on image is used to show the try-on effect when the target person tries on the second clothing.
[0011] The model is specifically used to: extract human body posture features at multiple points from the source image, and extract clothing features at multiple points from the clothing image of the second clothing; according to the matching relationship information, integrate the clothing features at the key points of the clothing into the human body features at the matching key points of the human body posture according to preset weights, so as to perform a fusion calculation of the human body features and the clothing features, and generate the target fitting image; wherein the preset weight is greater than the feature correlation weight calculated by the multi-layer neural network for the corresponding position.
[0012] Among them, if the source image is a dynamic image, the human body posture key points are determined according to one frame image of the source video, and the positions of the human body posture key points in other frames of the source video are determined according to the human body posture key points marked in the frame image and the tracking algorithm, so as to generate multiple frames of fitting images and form a target fitting video according to the corresponding human body posture key points in each frame and the matching relationship with the clothing key points.
[0013] The source image includes a fitting image of a model trying on the first garment;
[0014] The generated virtual fitting image is used to show the fitting effect when the model character tries on the second garment.
[0015] The source image includes an image uploaded by a user, including an image of a target person wearing a first garment;
[0016] The generated virtual fitting image is used to show the fitting effect when the target person specified by the user tries on the second clothing.
[0017] A method for providing a virtual try-on service, comprising:
[0018] Provide a virtual try-on service option on the landing page;
[0019] Receiving a user's request through the service option, and determining a source image and a clothing image of a second clothing selected by the user to be replaced in the source image for display, wherein the image content of the source image includes an image of a target person wearing a first clothing;
[0020] A virtual try-on image is displayed, and the virtual try-on image is generated in the following manner: a plurality of clothing key points annotated based on the clothing image of the second clothing, a plurality of human posture key points annotated by the target person in the source image, and a matching relationship between clothing key points and human posture key points are determined, and the source image, the clothing image of the second clothing, and the matching relationship information are input into a pre-trained model, so that in the process of the model generating the virtual try-on image, the matching relationship information is used to provide guidance information for the model; the virtual try-on image is used to display the try-on effect when the target person tries on the second clothing.
[0021] The source image includes: an image provided by the system of a model wearing a first clothing state, or an image uploaded by the user of a target person wearing a first clothing state.
[0022] A model training method, comprising:
[0023] Acquire a plurality of training data pairs, wherein the training data pairs include a first image and a second image, wherein the image content of the first image includes an image of a target person wearing a first garment, and the image content of the second image includes an image of a target person wearing a second garment, wherein the target person included in the first image and the second image is the same person and has the same posture and / or action;
[0024] Acquire a clothing image of a second clothing included in the second image;
[0025] Determining a plurality of clothing key points included in the clothing image of the second clothing, a plurality of human body posture key points of the target person, and a matching relationship between the clothing key points and the human body posture key points;
[0026] The first image and the clothing image of the second clothing item are input into a target model, and the target model generates a virtual try-on image for displaying the target person trying on the second clothing item, and the training process of the model is supervised using the corresponding second image in the training data pair; wherein, in the process of the target model generating the virtual try-on image, the matching relationship information is used to provide guidance information for the target model.
[0027] The determining of the plurality of clothing key points included in the clothing image of the second clothing, the plurality of human body posture key points of the target person, and the matching relationship between the clothing key points and the human body posture key points includes:
[0028] A point matching algorithm is used to extract multiple clothing key points from the clothing image of the second clothing, multiple human body posture key points of the target person are extracted from the second image, and a matching relationship between the clothing key points and the human body posture key points is determined.
[0029] Wherein, the obtaining of multiple training data pairs includes:
[0030] Acquire a second image and a clothing image of the first clothing;
[0031] The first image is generated by inputting the second image and the clothing image of the first clothing into a model pre-trained in a mask-based manner, so that the first image and the second image form the training data pair.
[0032] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of any of the methods described above.
[0033] An electronic device, comprising:
[0034] one or more processors; and
[0035] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of any of the methods described above.
[0036] A computer program product comprises a computer program / computer executable instructions, wherein the computer program / computer executable instructions implement the steps of any of the aforementioned methods when executed by a processor in an electronic device.
[0037] According to the specific embodiments provided in this application, this application discloses the following technical effects:
[0038] Through the embodiment of the present application, in the process of generating a virtual try-on image in an unmasked manner, multiple human posture key points annotated based on the target person in the source image, multiple clothing key points annotated based on the clothing image of the second clothing, and the matching relationship between clothing key points and human posture key points can be determined. In this way, the source image, the clothing image of the second clothing, and the matching relationship information can be input into the pre-trained model, so that in the process of the model generating a virtual try-on image, the matching relationship information is used to provide guidance information for the model to guide the model to generate a virtual try-on image for showing the try-on effect when the target person tries on the second clothing. Through this point-enhanced spatial attention scheme, the matching relationship information between clothing key points and human posture key points can be used as the guidance information of the model to guide the model to match the second clothing to a more suitable position and area on the target person, thereby improving the generation effect of the virtual try-on image.
[0039] In addition, when the source image and the generated virtual try-on image are dynamic images such as videos or animated images, a point-enhanced temporal attention mechanism can also be provided, that is, the key points of the human body posture can be annotated based on one frame of images, and the positions of the key points of the human body posture can be determined by tracking algorithms in other frames, so as to enhance the coherence of the generated video.
[0040] Of course, any product implementing the present application does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0042] Figure 1 It is a defect diagram when the virtual try-on image is generated based on the mask method;
[0043] Figure 2 It is a defect diagram when the virtual try-on image is generated based on the ordinary non-mask method;
[0044] Figure 3 It is a schematic diagram of the system architecture provided by the embodiment of the present application;
[0045] Figure 4 is a flow chart of the first method provided in an embodiment of the present application;
[0046] Figure 5is a schematic diagram of key point matching provided in an embodiment of the present application;
[0047] Figure 6 is a flow chart of the second method provided in an embodiment of the present application;
[0048] Figure 7 is a flowchart of the third method provided in an embodiment of the present application;
[0049] Figure 8 It is a schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0050] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.
[0051] First of all, it should be noted that, in view of the problem that the spatial information may be destroyed in the method of determining the try-on area based on the mask mentioned in the background technology part and generating the virtual try-on picture / video, a solution can be: obtain multiple paired training data to train the model, each training data pair can include two try-on pictures / videos, there is no mask in these two try-on pictures / videos, and it is assumed that the clothing A is tried on in the try-on picture / video 1, and the clothing B is tried on in the try-on picture / video 2. The other parts of the two try-on pictures / models are the same except for the different clothing, including the same model characters and the same postures, actions, etc. In this way, the try-on picture / video 1 and clothing B can be used as the input of the model, so that the model generates a try-on picture / video in which the model character in the try-on picture / video 1 tries on clothing B, and the training process of the model is supervised by the try-on picture / video 2. In this way, since there is no mask occlusion in the training data, the influence of mask occlusion on model training and prediction effect can be avoided.
[0052] Of course, in actual applications, it may be difficult to obtain two pictures / videos that are exactly the same in terms of models, poses, movements, etc. except for the clothing by directly crawling from the Internet. Therefore, in this case, it is also possible to construct training data pairs that meet the above conditions.
[0053] Specifically, the training process of the model can be divided into two stages. The first stage is to first train the model in a masked manner; the second stage is to generate multiple paired training data using the model trained in the first stage, wherein when generating each pair of training data, firstly, a source image 1 can be prepared, and the source image 1 can be a real model picture or video, animated picture, etc. obtained by real shooting, etc., for presenting the effect of the person S trying on the clothing A; in addition, a picture of clothing B can be prepared. Then, by inputting the above-mentioned source image 1 and the picture of clothing B into the model trained in the above-mentioned mask-based manner, a virtual try-on image 2 (which can be in the format of a picture, video, animated picture, etc.) in which the person S tries on the clothing B can be generated. In this way, the source image 1 and the virtual try-on image 2 can form a pair of training data for training the model. At this time, the above-mentioned virtual try-on image 2 and the picture of clothing A can be used as the input data of the model, and the real source image 1 can be used to supervise the training process of the model. Of course, there can be many similar training data pairs, which can be constructed and supervised in the same way to complete the model training in the second stage. After that, the model trained in the second stage can be used to generate virtual try-on images / videos. That is, in the prediction stage, a source image and a picture of clothing (the clothing is different from the clothing in the source image) can be input into the model, and the model can generate new try-on images without adding masks or other processing.
[0054] Through the above-mentioned method, since it is no longer necessary to add a mask in the process of generating a virtual try-on image, the spatial information in the source image will not be destroyed (for example, the limbs of the model are partially blocked by the mask, etc.), and the complex human body movements and postures in the source image can be effectively retained, which is beneficial to improving the generation effect. However, in the process of realizing the present application, the inventor of the present application discovered that the mask actually has its value, that is, the try-on area in the picture can be marked by the mask. For example, the try-on area in a source image may be the upper body. The area can be marked by a mask. When generating a new virtual try-on image, the model can be guided to replace clothing in the try-on area. However, in the above-mentioned process of generating a virtual try-on image without adding a mask, since there is no mask, the specific try-on area cannot be marked. Therefore, the corresponding guidance information cannot be provided to the model, so that the try-on area may be inaccurate. For example, assuming that the source image is such as Figure 2 As shown in (A), the model is wearing a skirt, and the top part needs to be replaced with a T-shirt as shown in the lower left corner 21. The expected generated effect is as follows Figure 2However, since the top and skirt in the source image are a complete set, and are relatively consistent in color and pattern, when replacing without adding a mask, the skirt of the person may also be replaced, so the generated result may be as follows: Figure 2 (C) shown.
[0055] The embodiment of the present application provides a further optimization scheme based on the above-mentioned maskless virtual try-on image generation scheme. Specifically, in this scheme, the model can be guided to replace the required clothing to the desired try-on area by matching the key points of the clothing with the key points of the model's body, thereby improving the generation effect of the virtual try-on image.
[0056] Specifically, in the training stage, since there are target fitting images and clothing pictures in the training data pair, and the human features (including postures, movements, etc.) of the model in the source image and the target fitting image are the same, the existing point matching algorithm can be used to extract the clothing key points on the clothing from the clothing picture, extract the human body posture key points on the model body from the target fitting image, and establish a matching relationship between the clothing key points and the human body posture key points. For example, a clothing key point on the clothing matches a human body posture key point on the model body, which means that in the process of generating a virtual fitting image, the clothing key point needs to be matched to the position of the human body posture key point. In this way, in the process of feature extraction and fusion of the model, after the clothing features and human body features are extracted respectively, the matching relationship between the above key points can be used to forcibly add the clothing features at a certain clothing key point to the human body features at the human body posture key point with a matching key with it, so as to provide the model with guidance information about the fitting area to improve the training effect of the model.
[0057] In the process of prediction by the model obtained through the above training, since there are only source images and pictures of target clothing that need to be replaced in the source images, it can be implemented by manually marking key points in clothing pictures and source images. For example, in specific implementation, a specific operation interface can be provided so that users (for example, back-end technicians in the commodity information service system, etc.) can complete the key point marking process by performing click operations in clothing pictures and source images. After completing the marking of key points, the clothing pictures and source images with marked information can be input into the trained model. In the process of fusing clothing features and human body features, the matching relationship information between the marked key points can be used to provide the model with guidance information about the try-on area.
[0058] From the perspective of system architecture, Figure 3As shown, the embodiment of the present application can provide a virtual try-on service in the application related to the commodity information service system, etc. First, the model training can be completed on the server side, and then in the prediction stage, by inputting the pictures of the specific clothing to be tried on and the source image into the trained model, a virtual try-on image (picture or video, animated image, etc.) for displaying the target person trying on the above-mentioned clothing to be tried on is generated without generating a mask. Among them, the above-mentioned generation process of the virtual try-on image can be completed in advance by an offline method, and the specific generation result is saved on the server side. After that, an entrance to a specific virtual try-on service can also be provided in the client interface of the front end. After the user enters the service page through the entrance, multiple optional clothing (these services have completed the generation of virtual try-on images in advance) pictures and other information can be provided. The user can click on the picture of the specific clothing to view the virtual try-on image corresponding to the clothing. Alternatively, an entrance to upload source images can also be provided for users. Users can upload their own photos or videos, etc., and select the specific clothing pictures that need to be tried on, and then virtual try-on pictures / videos of the people in the photos or videos uploaded by the users trying on the clothing they selected can be generated for the users. In the process of model training and prediction, key point matching can be used to provide the model with guidance information about the matching position to improve the generation effect. In the scene where the source image is a video, since there are multiple frames in the same video, the target person in the video may be making certain actions. At this time, when marking or extracting key points, key point marking or extraction can be performed based on one frame, and the key point positions can be calculated using tracking algorithms for other frames to ensure the consistency of the key point matching relationship, thereby improving the fluency of the generated video.
[0059] The specific implementation scheme provided in the embodiments of the present application is described in detail below.
[0060] Embodiment 1
[0061] First, the first embodiment of the present application provides a method for generating a virtual try-on image. Figure 4 , the method may include:
[0062] S401: Determine a source image, wherein the image content of the source image includes an image of a target person wearing a first garment.
[0063] Among them, regarding the source image, there may be different ways of determining it in different scenarios. For example, in one way, the system may provide pictures / videos of clothing being tried on by models as source images. Such pictures / videos of trying on clothing may be specially taken for the virtual try-on service, or may be obtained from the public network, etc. Of course, in the method of obtaining the source image from the public network, it may also be used with the authorization of the creator or publisher. In this way, it is equivalent to providing users with the try-on effects of different clothing on models to facilitate users to make purchasing decisions.
[0064] In another way, the user may have the following requirements: before choosing a certain garment, the user may need to check how the specific garment looks on him or herself, or if the user needs to buy garments for others, the user may also need to check how the specific garment looks on others, etc. Therefore, the user may upload the source image, which may also include a target person who also wears a certain garment. The target person may be the current user himself or another person selected by the user.
[0065] It should be noted that the first clothing in the source image may be clothing provided by the system and corresponding to a specific product, or may not correspond to a specific product. For example, when the source image is obtained from a public network or when the source image is uploaded by a user, the clothing worn by the specific target person may be the target person's own clothing, which is not limited here.
[0066] In addition, the source image obtained in this step can be a picture, or a video, a dynamic image, or the like. In the case of a video, a dynamic image, or the like, the target person can make certain movements while wearing the first garment to show the effect of the first garment on the upper body in motion. Correspondingly, the generated virtual fitting image can also be a corresponding video or dynamic image, in which the movements made by the target person are consistent with those made in the source image.
[0067] S402: Determine a clothing image of a second clothing to be replaced in the source image for display.
[0068] The second clothing can also be determined in a variety of ways. In the scenario where the system generates a virtual try-on image offline in advance and then provides the user with the clothing to choose for virtual try-on through the front-end page, the specific second clothing can be selected from the product library by the system according to a certain strategy, or, if the front-end page is a store page of a certain merchant, the merchant can also select the specific clothing that needs to provide virtual try-on services, etc. In short, the second clothing can be determined in a variety of ways, and there can be multiple second clothing, and each corresponding virtual try-on image is generated.
[0069] After the second clothing is determined, a clothing image of the specific second clothing can also be determined. Such clothing image can be provided by the merchant, or, if the specific second clothing corresponds to a product that has been published in the product information service system, a suitable image can also be selected from the product details information, etc. Specifically, in order to obtain a better generation effect, the clothing image of the second clothing obtained in this step can be a white background image containing a front image of the second clothing.
[0070] S403: Determine a plurality of clothing key points annotated based on the clothing image of the second clothing, a plurality of human body posture key points annotated based on the target person in the source image, and a matching relationship between clothing key points and human body posture key points.
[0071] After determining the source image and the clothing image of the second clothing, as mentioned above, the embodiment of the present application needs to generate a virtual try-on image without adding a mask. However, if clothing feature extraction, human posture feature extraction, and clothing feature and human posture feature fusion calculation are directly performed, it may cause the try-on area recognition error, so that the generated virtual try-on image cannot show a good try-on effect. Therefore, in the embodiment of the present application, a point matching method can be used to provide guidance information for the model to solve the above problem.
[0072] Among them, since the only known information is the source image and the second clothing, and the clothing presented in the source image is different from the second clothing, it is impossible to extract the clothing key points and the human body posture key points through the point matching algorithm. In view of this situation, in the testing phase (that is, the phase of generating virtual try-on images using the trained model), the clothing key points and the human body posture key points can be determined by manual annotation.
[0073] For example, in a specific implementation, an interface for performing annotation operations can be provided to the user, in which the clothing image of the second clothing and the source image can be displayed respectively, wherein if the source image is a video or a moving picture, one of the image frames in the video or moving picture can be displayed. In addition, a tool for performing annotation operations can also be provided in the interface, and after the user clicks to select the tool, the cursor can change to an annotation state, and then the user can perform a click operation at a certain position in the clothing image of the second clothing, and the system will determine the position as a clothing key point (for the convenience of description, it can be recorded as clothing key point A1), and then the user can perform a click operation at a certain position in the source image, and the system will determine the position as a human posture key point (which can be recorded as human posture key point B1), and automatically establish a matching relationship between the human posture key point B1 and the previous clothing key point A1. Afterwards, the user can re-perform the click operation at another position on the clothing image of the second clothing, and the system will determine the position as another clothing key point (recorded as clothing key point A2). Afterwards, the user can also perform the click operation at another position in the source image, and the system will determine the position as another human body posture key point (recorded as human body posture key point B2), and automatically establish a matching relationship between the human body posture key point B2 and the previous clothing key point A2. By analogy, multiple clothing key points and human body posture key points with matching relationships can be obtained.
[0074] Of course, in the specific implementation, there may be other ways of marking key points. For example, multiple clothing key points may be marked on the clothing image of the second clothing, and multiple human posture key points may be marked in the source image. The user may then specify which human posture key point has a matching relationship with each clothing key point, and so on.
[0075] Specifically, when marking key points, the user can select the key point positions to be marked by observing the clothing image and the specific posture of the target person in the source image, etc. The goal is to more accurately "wear" the second clothing on the appropriate area of the target person according to the matching relationship of the key point positions. The number of key points can be determined according to the actual situation. For example, for more complex situations, more key points can be marked, and for simple situations, the number of key points can be reduced.
[0076] For example, suppose a clothing image of a second clothing item is Figure 5 As shown in (A), the source image is Figure 5 (B) shows the dots in the figure, where the dots are the locations of the marked key points. Figure 5 (A) is the key point of the clothing. Figure 5 (B) is the key point of human posture. The virtual fitting image to be generated can be as follows: Figure 5 (C) shown.
[0077] S404: Input the source image, the clothing image of the second clothing and the matching relationship information into a pre-trained model, so that when the model generates a virtual try-on image, the matching relationship information is used to provide guidance information for the model; the virtual try-on image is used to show the try-on effect when the target person tries on the second clothing.
[0078] After determining the key points of clothing and human posture that have a matching relationship, the source image, the clothing image of the second clothing and the matching relationship information can be input into a pre-trained model, and the model can generate a virtual try-on image. The generated virtual try-on image can be used to show the trying-on effect when the target person tries on the second clothing.
[0079] In the process of generating the above-mentioned virtual try-on image, if the traditional method is used, human posture features at multiple points can be first extracted from the source image, and clothing features at multiple points can be extracted from the clothing image of the second clothing, and then the human posture features can be fused and calculated with the clothing features. Specifically, the clothing features and human posture features can be input into a multi-layer neural network, and the clothing features at each point can be correlated with the human posture features at each point to obtain the correlation weights between the clothing features and the human posture features at each point, and then all clothing features can be integrated into the human posture features according to the correlation weights, and then the virtual try-on image can be generated according to the fusion results.
[0080] However, in the embodiment of the present application, after extracting the human features at multiple points from the source image and extracting the clothing features at multiple points from the clothing image of the second clothing, the correlation weight between the clothing features at the clothing key points and the human posture features at the matching human posture key points can be reset to a higher value according to the matching relationship information between the aforementioned clothing key points and the human posture key points. The higher weight value mentioned here is relative to the feature correlation weight calculated by the multi-layer neural network in the aforementioned traditional scheme. That is to say, assuming that the correlation weight between the clothing features at point A and the human posture features at point B is calculated by the aforementioned multi-layer neural network as w1, but since point A and point B are key points with a matching relationship, the correlation weight between the clothing features at these two positions and the human posture features can be forcibly set to a weight value w2 greater than w1, and then the clothing features at the position are integrated into the corresponding human posture features according to the correlation weight w2, so as to perform a fusion calculation of the human posture features and the clothing features, and generate the target fitting image.
[0081] This method can be called a point-enhanced spatial attention scheme, which allows the matching relationship information between clothing key points and human posture key points to be used as guidance information for the model to guide the model to match the second clothing to a more suitable position and area on the target person.
[0082] In addition, if the specific source image is a video image, since a video frame usually includes multiple video frames, the generated virtual try-on image also needs to include multiple different video frames. In a specific implementation, in one way, each video frame can be regarded as a static image, and the human body posture key points are marked for each video frame. However, since the position and posture of the target person in the image may be different between different video frames, if the human body posture key points are marked for each video frame separately, it may also lead to inconsistent key point matching relationships between adjacent frames, thereby affecting the coherence of the generated video.
[0083] Based on the above situation, in the preferred implementation mode of the embodiment of the present application, the human posture key points can be annotated according to one frame of the source video, and according to the human posture key points annotated in the frame image and the tracking algorithm, the human posture key points in other frames of the source video can be determined, so as to generate multiple frames of fitting images and form a target fitting video according to the human posture key points corresponding to each frame and the matching relationship with the clothing key points. In this way, the coherence of the generated video can be enhanced through the point-enhanced temporal attention mechanism.
[0084] After the virtual try-on image for the specific second clothing item is generated, it can be saved to the server. When the user chooses to view the virtual try-on image of a certain clothing item through the front-end interface, the virtual try-on image can be provided from the server to the client for display in the front-end interface. If the user uploads the source image and selects the second clothing item, the server can also generate a virtual try-on image of the second clothing item on the target person specified by the user and return it to the client for display.
[0085] In summary, through the embodiment of the present application, in the process of generating a virtual try-on image in an unmasked manner, multiple human posture key points annotated based on the target person in the source image, multiple clothing key points annotated based on the clothing image of the second clothing, and the matching relationship between clothing key points and human posture key points can be determined. In this way, the source image, the clothing image of the second clothing, and the matching relationship information can be input into the pre-trained model, so that in the process of the model generating a virtual try-on image, the matching relationship information is used to provide guidance information for the model to guide the model to generate a virtual try-on image for showing the try-on effect when the target person tries on the second clothing. Through this point-enhanced spatial attention scheme, the matching relationship information between clothing key points and human posture key points can be used as the guidance information of the model to guide the model to match the second clothing to a more suitable position and area on the target person, thereby improving the generation effect of the virtual try-on image.
[0086] In addition, when the source image and the generated virtual try-on image are dynamic images such as videos or animated images, a point-enhanced temporal attention mechanism can also be provided, that is, the key points of the human body posture can be annotated based on one frame of images, and the positions of the key points of the human body posture can be determined by tracking algorithms in other frames, so as to enhance the coherence of the generated video.
[0087] Embodiment 2
[0088] The second embodiment corresponds to the first embodiment and provides a method for providing a virtual try-on service from the perspective of the client. Figure 6 , the method may include:
[0089] S601: providing a virtual try-on service option in the target page;
[0090] S602: receiving a user's request through the service option, and determining a source image and a clothing image of a second clothing selected by the user to be replaced in the source image for display, wherein the image content of the source image includes an image of a target person wearing a first clothing;
[0091] S603: Display a virtual try-on image, wherein the virtual try-on image is generated in the following manner: determine a plurality of clothing key points annotated based on the clothing image of the second clothing, a plurality of human posture key points annotated by the target person in the source image, and a matching relationship between clothing key points and human posture key points, input the source image, the clothing image of the second clothing, and the matching relationship information into a pre-trained model, so that in the process of the model generating the virtual try-on image, the matching relationship information provides guidance information for the model; the virtual try-on image is used to display the try-on effect when the target person tries on the second clothing.
[0092] The source image may include: an image provided by the system of a model wearing the first clothing state, or an image uploaded by the user of a target person wearing the first clothing state.
[0093] Embodiment 3
[0094] This fact 3 provides a model training method, see Figure 7 , the method may include:
[0095] S701: Acquire a plurality of training data pairs, wherein the training data pairs include a first image and a second image, wherein the image content of the first image includes an image of a target person wearing a first garment, and the image content of the second image includes an image of a target person wearing a second garment, wherein the target person included in the first image and the second image is the same person and has the same posture and / or action;
[0096] S702: Acquire a clothing image of a second clothing included in the second image;
[0097] S703: Determine a plurality of clothing key points included in the clothing image of the second clothing, a plurality of human body posture key points of the target person, and a matching relationship between the clothing key points and the human body posture key points;
[0098] S704: Input the first image and the clothing image of the second clothing into the target model, and generate a virtual try-on image for displaying the target person trying on the second clothing by the target model, and supervise the training process of the model using the corresponding second image in the training data; wherein, in the process of the target model generating the virtual try-on image, the matching relationship information is used to provide guidance information to the target model.
[0099] In an optional implementation, a point matching algorithm can be used to extract multiple clothing key points from the clothing image of the second clothing, extract multiple human body posture key points of the target person from the second image, and determine the matching relationship between the clothing key points and the human body posture key points. How the specific point matching algorithm completes the extraction and matching of key points does not belong to the protection focus of the embodiments of this application, so it is not described in detail here.
[0100] In addition, when obtaining multiple training data pairs, the second image and the clothing image of the first clothing can be obtained first; then, the first image can be generated by inputting the second image and the clothing image of the first clothing into a model that has been pre-trained in a mask-based manner, so as to construct the first image and the second image into the training data pair.
[0101] For the undescribed parts in Embodiment 2 and Embodiment 3, please refer to the records in the aforementioned Embodiment 1 and other parts of this specification, and will not be repeated here.
[0102] It should be noted that the embodiments of the present application may involve the use of user data. In actual applications, user-specific personal data can be used in the scheme described herein within the scope permitted by applicable laws and regulations, subject to the requirements of applicable laws and regulations of the country where the user is located (for example, with the user's explicit consent, effective notification to the user, etc.).
[0103] Corresponding to the first embodiment, the embodiment of the present application further provides a device for generating a virtual try-on image, which may include:
[0104] A source image determining unit, configured to determine a source image, wherein the image content of the source image includes an image of a target person wearing a first garment;
[0105] A clothing image determining unit, used to determine a clothing image of a second clothing to be replaced in the source image for display;
[0106] a point matching relationship determination unit, configured to determine a plurality of clothing key points annotated based on the clothing image of the second clothing, a plurality of human body posture key points annotated based on the target person in the source image, and a matching relationship between the clothing key points and the human body posture key points;
[0107] A generation unit is used to input the source image, the clothing image of the second clothing and the matching relationship information into a pre-trained model, so that when the model generates a virtual try-on image, the matching relationship information is used to provide guidance information for the model; the virtual try-on image is used to show the try-on effect when the target person tries on the second clothing.
[0108] The model is specifically used to: extract human body posture features at multiple points from the source image, and extract clothing features at multiple points from the clothing image of the second clothing; according to the matching relationship information, integrate the clothing features at the key points of the clothing into the human body features at the matching key points of the human body posture according to preset weights, so as to perform a fusion calculation of the human body features and the clothing features, and generate the target fitting image; wherein the preset weight is greater than the feature correlation weight calculated by the multi-layer neural network for the corresponding position.
[0109] In addition, if the source image is a dynamic image, the human body posture key points are determined according to one frame of the source video, and the positions of the human body posture key points in other frames of the source video are determined according to the human body posture key points marked in the frame and the tracking algorithm, so as to generate multiple frames of fitting images and form a target fitting video based on the human body posture key points corresponding to each frame and the matching relationship with the clothing key points.
[0110] Specifically, the source image includes a fitting image of a model trying on the first clothing item; at this time, the generated virtual fitting image is used to show the fitting effect when the model tries on the second clothing item.
[0111] Alternatively, the source image includes an image uploaded by a user, including an image of a target person wearing a first garment; in this case, the generated virtual fitting image is used to show the fitting effect of the target person designated by the user when trying on the second garment.
[0112] Corresponding to the second embodiment, the embodiment of the present application further provides a device for providing a virtual try-on service, which may include:
[0113] A service option providing unit, used for providing a virtual try-on service option in a target page;
[0114] a request receiving unit, configured to receive a user's request through the service option, and determine a source image and a clothing image of a second clothing selected by the user to be replaced in the source image for display, wherein the image content of the source image includes an image of a target person wearing a first clothing;
[0115] A display unit is used to display a virtual try-on image, wherein the virtual try-on image is generated in the following manner: determining a plurality of clothing key points annotated based on the clothing image of the second clothing, a plurality of human posture key points annotated by the target person in the source image, and a matching relationship between clothing key points and human posture key points, inputting the source image, the clothing image of the second clothing, and the matching relationship information into a pre-trained model, so that in the process of the model generating the virtual try-on image, the matching relationship information provides guidance information for the model; the virtual try-on image is used to display the try-on effect when the target person tries on the second clothing.
[0116] The source image includes: an image provided by the system of a model wearing a first clothing state, or an image uploaded by the user of a target person wearing a first clothing state.
[0117] Corresponding to the third embodiment, the embodiment of the present application further provides a model training device, which may include:
[0118] a training data pair acquisition unit, configured to acquire a plurality of training data pairs, wherein the training data pairs include a first image and a second image, wherein the image content of the first image includes an image of a target person wearing a first garment, and the image content of the second image includes an image of a target person wearing a second garment, wherein the target person included in the first image and the second image is the same person and has the same posture and / or action;
[0119] a clothing image acquiring unit, configured to acquire a clothing image of a second clothing included in the second image;
[0120] a point matching information determining unit, configured to determine a plurality of clothing key points included in the clothing image of the second clothing, a plurality of human body posture key points of the target person, and a matching relationship between the clothing key points and the human body posture key points;
[0121] A training unit is used to input the first image and the clothing image of the second clothing into a target model, and the target model generates a virtual try-on image for showing the target person trying on the second clothing, and supervises the training process of the model using the corresponding second image in the training data; wherein, in the process of the target model generating the virtual try-on image, the matching relationship information is used to provide guidance information for the target model.
[0122] Specifically, the point matching information determination unit may be used to:
[0123] A point matching algorithm is used to extract multiple clothing key points from the clothing image of the second clothing, multiple human body posture key points of the target person are extracted from the second image, and a matching relationship between the clothing key points and the human body posture key points is determined.
[0124] The training data pair acquisition unit can be specifically used for:
[0125] Acquire a second image and a clothing image of the first clothing;
[0126] The first image is generated by inputting the second image and the clothing image of the first clothing into a model pre-trained in a mask-based manner, so that the first image and the second image form the training data pair.
[0127] In addition, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, the steps of any one of the methods in the aforementioned method embodiments are implemented.
[0128] And an electronic device, comprising:
[0129] one or more processors; and
[0130] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of the aforementioned method embodiments.
[0131] A computer program product includes a computer program / computer executable instructions, which, when executed by a processor in an electronic device, implement the steps of the method described in the aforementioned method embodiment.
[0132] in, Figure 8 The architecture of the electronic device is shown exemplarily, and may specifically include a processor 810, a video display adapter 811, a disk drive 812, an input / output interface 813, a network interface 814, and a memory 820. The processor 810, the video display adapter 811, the disk drive 812, the input / output interface 813, the network interface 814, and the memory 820 may be communicatively connected via a communication bus 830.
[0133] Among them, the processor 810 can be implemented by a general-purpose CPU (Central Processing Unit, processor), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solution provided in this application.
[0134] The memory 820 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 820 can store an operating system 821 for controlling the operation of the electronic device 800, and a basic input and output system (BIOS) for controlling the low-level operation of the electronic device 800. In addition, a web browser 823, a data storage management system 824, and a virtual try-on image generation processing system 825, etc. can also be stored. The above-mentioned virtual try-on image generation processing system 825 can be an application program that specifically implements the aforementioned steps in the embodiment of the present application. In short, when the technical solution provided by the present application is implemented by software or firmware, the relevant program code is stored in the memory 820 and is called and executed by the processor 810.
[0135] The input / output interface 813 is used to connect the input / output module to realize information input and output. The input / output module can be configured as a component in the device (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0136] The network interface 814 is used to connect to a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).
[0137] The bus 830 comprises a pathway for transmitting information between the various components of the device (eg, the processor 810, the video display adapter 811, the disk drive 812, the input / output interface 813, the network interface 814, and the memory 820).
[0138] It should be noted that, although the above device only shows a processor 810, a video display adapter 811, a disk drive 812, an input / output interface 813, a network interface 814, a memory 820, a bus 830, etc., in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include components necessary for implementing the solution of the present application, and does not necessarily include all the components shown in the figure.
[0139] It can be known from the description of the above implementation methods that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be essentially or partly contributed to the prior art in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application or certain parts of the embodiments.
[0140] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can refer to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without creative work.
[0141] The above is a detailed introduction to the method and electronic device for generating virtual try-on images provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. At the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present application.
Claims
1. A method for generating a virtual try-on image, characterized in that: include: Determine a source image, wherein the image content of the source image includes an image of a target person wearing a first garment; Determining a clothing image of a second clothing to be replaced in the source image for display; Determining a plurality of clothing key points annotated based on the clothing image of the second clothing, a plurality of human body posture key points annotated based on the target person in the source image, and a matching relationship between the clothing key points and the human body posture key points; The source image, the clothing image of the second clothing and the matching relationship information are input into a pre-trained model, so that in the process of the model generating a virtual try-on image, the matching relationship information is used to provide guidance information for the model; the virtual try-on image is used to show the try-on effect when the target person tries on the second clothing.
2. The method according to claim 1, characterized in that The model is specifically used to: extract human body posture features at multiple points from the source image, and extract clothing features at multiple points from the clothing image of the second clothing; According to the matching relationship information, the clothing features at the clothing key points are integrated into the human body features at the matching human body posture key points according to preset weights, so as to perform a fusion calculation of the human body features and the clothing features, and generate the target fitting image; wherein the preset weights are greater than the feature correlation weights calculated by the multi-layer neural network for the corresponding positions.
3. The method according to claim 1, characterized in that If the source image is a dynamic image, the human body posture key points are determined according to one frame of the source video, and the positions of the human body posture key points in other frames of the source video are determined according to the human body posture key points marked in the frame and the tracking algorithm, so as to generate multiple frames of fitting images and form a target fitting video based on the human body posture key points corresponding to each frame and the matching relationship with the clothing key points.
4. The method according to any one of claims 1 to 3, characterized in that: The source image includes a fitting image of a model trying on the first garment; The generated virtual fitting image is used to show the fitting effect when the model character tries on the second garment.
5. The method according to any one of claims 1 to 3, characterized in that: The source image includes an image uploaded by a user, including an image of a target person wearing a first garment; The generated virtual fitting image is used to show the fitting effect when the target person specified by the user tries on the second clothing.
6. A method for providing a virtual try-on service, characterized in that: include: Provide a virtual try-on service option on the landing page; Receiving a user's request through the service option, and determining a source image and a clothing image of a second clothing selected by the user to be replaced in the source image for display, wherein the image content of the source image includes an image of a target person wearing a first clothing; A virtual try-on image is displayed, wherein the virtual try-on image is generated in the following manner: a plurality of clothing key points annotated based on the clothing image of the second clothing item are determined, a plurality of human body posture key points annotated based on the target person in the source image, and a matching relationship between clothing key points and human body posture key points are input into a pre-trained model, so that in the process of the model generating the virtual try-on image, the matching relationship information is used to provide guidance information to the model; the virtual try-on image is used to display the try-on effect when the target person tries on the second clothing item.
7. The method according to claim 6, characterized in that The source image includes: an image provided by the system of a model character wearing a first clothing state, or an image uploaded by the user of a target character wearing a first clothing state.
8. A model training method, characterized in that: include: Acquire a plurality of training data pairs, wherein the training data pairs include a first image and a second image, wherein the image content of the first image includes an image of a target person wearing a first garment, and the image content of the second image includes an image of a target person wearing a second garment, wherein the target person included in the first image and the second image is the same person and has the same posture and / or action; Acquire a clothing image of a second clothing included in the second image; Determining a plurality of clothing key points included in the clothing image of the second clothing, a plurality of human body posture key points of the target person, and a matching relationship between the clothing key points and the human body posture key points; The first image and the clothing image of the second clothing item are input into a target model, and the target model generates a virtual try-on image for displaying the target person trying on the second clothing item, and the training process of the model is supervised using the corresponding second image in the training data pair; wherein, in the process of the target model generating the virtual try-on image, the matching relationship information is used to provide guidance information for the target model.
9. The method according to claim 8, characterized in that The determining of a plurality of clothing key points included in the clothing image of the second clothing, a plurality of human body posture key points of the target person, and a matching relationship between the clothing key points and the human body posture key points includes: A point matching algorithm is used to extract multiple clothing key points from the clothing image of the second clothing, multiple human body posture key points of the target person are extracted from the second image, and a matching relationship between the clothing key points and the human body posture key points is determined.
10. The method according to claim 8, characterized in that The obtaining of multiple training data pairs comprises: Acquire a second image and a clothing image of the first clothing; The first image is generated by inputting the second image and the clothing image of the first clothing into a model pre-trained in a mask-based manner, so that the first image and the second image form the training data pair.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 10 are implemented.
12. An electronic device, characterized in that: include: one or more processors; as well as A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of claims 1 to 10.
13. A computer program product comprising a computer program / computer executable instructions, characterized in that: When the computer program / computer executable instructions are executed by a processor in an electronic device, the steps of the method according to any one of claims 1 to 10 are implemented.