Model image generation method and device, equipment, medium and product
By acquiring images of people in clothing and using model templates with similar poses and deformation feature parameters to generate clothing masks, the problems of distortion and inaccurate adaptation in virtual clothing swapping technology are solved, achieving efficient generation of model clothing images and reducing the cost of product release on e-commerce platforms.
Patent Information
- Application Number
- CN202211215394.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-30
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2042-09-30
AI Technical Summary
In existing virtual clothing-changing technologies, the limited amount of clothing data and redundant input image information lead to model overfitting, resulting in unnatural distortions and deformations in the generated virtual clothing-changing results. Furthermore, the clothing cannot accurately fit the model's body, resulting in poor synthesis effects.
By acquiring images of people in clothing, using model template matching with similar poses, determining deformation feature parameters, generating clothing masks, and compositing the clothing transformation images into the model's body image, the effect of the clothing is predicted using the key point information of the model's body image, and generating images of the model in clothing.
The generated model clothing images are closer to the posture of real people in clothing images, avoiding distortion and deformation, improving the compositing effect, simplifying business processes, reducing product release costs, and improving product listing efficiency.
Smart Images

Figure CN115512085B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to virtual clothes changing technology, and in particular to a model image generation method and device, equipment, medium and product. BACKGROUND
[0002] In the process of listing clothing commodities in an e-commerce platform, an important link is to make and provide model display information. The traditional making method needs to find real models to shoot commodity pictures, which is very costly, and the picture shooting cycle is also very long. Virtual clothes changing technology emerges as the times require, which can achieve "cost reduction and efficiency improvement", and virtual model commodity pictures can also be used for commodity measurement, reducing the investment cost in the early stage of commodity listing.
[0003] In the traditional virtual clothes changing technology, a deep learning model is usually used to process various information. The basic principle is to first obtain the deformation parameters of the clothing image to be published relative to the model body image, and then synthesize the clothing image into the model body image according to the deformation parameters to generate the final effect picture.
[0004] In practice, due to the small amount of clothes changing data, there is a lot of redundant information in the image information input into each model. These redundant information plays an interference role, which easily leads to model overfitting in the model training process, and the convergence process is dominated by a certain supervision signal, causing the virtual clothes changing result obtained by the model to be very prone to unnatural distortion and deformation, and the actual synthesis success rate is extremely low.
[0005] In addition, according to the deformation parameters, the original clothing image is directly synthesized into the model body image. Since the deformation parameters are based on and also refer to the information of the human body image in the original image, the deformation parameters cannot effectively focus on the actual effect after being migrated to the model body image, thus causing the clothes in the final effect picture to not accurately fit the model human body.
[0006] Therefore, it is necessary to comprehensively improve the existing virtual clothes changing technology in order to obtain good application effect. SUMMARY
[0007] The present application aims to solve the above problems and provide a model image generation method and its corresponding device, equipment, non-volatile readable storage medium, and computer program product.
[0008] According to one aspect of the present application, a model image generation method is provided, comprising the following steps:
[0009] Obtaining a person dressed image;
[0010] Recalling a model template with a similar pose of the human body in the person dressed image from a model gallery, the model template comprising a body image of a model;
[0011] determine a deformation feature parameter of an original clothing image in the person dressing image relative to the body map, transform the original clothing image according to the deformation feature parameter to obtain a clothing transformed image;
[0012] predict a clothing mask corresponding to the body map according to the clothing transformed image and the body map;
[0013] synthesize the clothing transformed image into the body map according to the clothing mask to obtain a model dressing image.
[0014] Optionally, the person dressing image is obtained, including:
[0015] reading a preview video stream generated by a camera, performing human body recognition on each image frame in the preview video stream, and determining an image frame containing a human body image as a person image;
[0016] determining a key point feature map of each person image by using a human key point detection model, and calculating semantic similarity between the key point feature map and a key point feature map of a body map of a model template in a model gallery;
[0017] when the highest semantic similarity obtained by the person image exceeds a preset threshold, determining the person image as the person dressing image.
[0018] Optionally, a model template having a similar pose to a human body in the person dressing image is recalled from the model gallery, and the model template includes a body map of a model, including:
[0019] detecting and determining a key point feature map of the person dressing image by using a human key point detection model;
[0020] calculating semantic similarity between the key point feature map of the person dressing image and a key point feature map of a body map of each model template in a model gallery, the key point feature map of the body map of each model template being determined in advance by using the human key point detection model;
[0021] screening at least one model template having high semantic similarity as a model template having a similar pose to the human body in the person dressing image.
[0022] Optionally, the deformation feature parameter of the original clothing image in the person dressing image relative to the body map is determined, including:
[0023] constructing the original clothing image and the key point feature map in the person dressing image as user input information, and inputting a first feature extraction model to extract first deep semantic information;
[0024] The original image, the body image, the hairstyle image, and the key point feature image of the model included in the model template are configured as template input information, and a second feature extraction model is input to extract second deep semantic information;
[0025] The morphing feature parameter is predicted according to the first deep semantic information and the second deep semantic information, and the morphing feature parameter is used to pay attention to the morphing amplitude information of the original clothing image in the person dressing image relative to the body image.
[0026] Optionally, the clothing mask corresponding to the body image is predicted according to the clothing transformation image and the body image, and the clothing mask comprises:
[0027] The key point feature image corresponding to the body image of the model template is configured as a clothing transformation image information together with the clothing transformation image;
[0028] The transformation image information is input into a preset image segmentation model for image segmentation, and a clothing mask corresponding to the clothing transformation image adjusted to the body image is predicted.
[0029] Optionally, the clothing transformation image is synthesized into the body image according to the clothing mask to obtain a model dressing image, and the clothing mask comprises:
[0030] The clothing transformation image is corrected according to the clothing mask, and the clothing transformation image is synthesized into the body image of the model in the model template to obtain a preliminary dressing image;
[0031] The preliminary dressing image is corrected by using the hairstyle image and the torso image in the model template to obtain a body state correction image;
[0032] The body state correction image is synthesized as a foreground with a background image in the model template to obtain a model dressing image.
[0033] Optionally, before the original clothing image and the key point feature image of the person dressing image are configured as user input information, the method further comprises:
[0034] The person dressing image is subjected to image segmentation to obtain the original clothing image therein;
[0035] An image inpainting model is used to inpaint the occluded part of the original clothing image in the person dressing image.
[0036] According to another aspect of the present application, a model image generation device is provided, comprising:
[0037] An image acquisition module is configured to acquire a person dressing image;
[0038] a template recall module configured to recall a model template with a similar pose to the human body in the person dressing image from a model library, the model template including a body map of a model;
[0039] a deformation obtaining module configured to determine a deformation feature parameter of an original clothing image in the person dressing image relative to the body map, and transform the original clothing image according to the deformation feature parameter to obtain a clothing transformed image;
[0040] a mask generation module configured to predict a clothing mask corresponding to the body map according to the clothing transformed image and the body map;
[0041] an image synthesis module configured to synthesize the clothing transformed image into the body map according to the clothing mask to obtain a model dressing image.
[0042] According to another aspect of the present application, there is provided a model image generation device, comprising a central processing unit and a memory, the central processing unit being configured to invoke a computer program stored in the memory to execute the steps of the model image generation method described in the present application.
[0043] According to another aspect of the present application, there is provided a non-volatile readable storage medium storing a computer program implemented according to the model image generation method in the form of computer readable instructions, the computer program being invoked by a computer to execute the steps included in the method when running.
[0044] According to another aspect of the present application, there is provided a computer program product comprising computer program / instructions, the computer program / instructions being executed by a processor to implement the steps of the method described in any one of the embodiments of the present application.
[0045] Compared with the prior art, the present application has many technical advantages, including but not limited to:
[0046] Firstly, the present application takes the person dressing image as the source of the original clothing image, and uses the pose of the human body in the person dressing image to match the model template with a similar pose first, and then performs virtual clothing transformation based on the selected model template, which can make the generated model dressing image closer to the pose of the person dressing image and maximize the degree of avoiding distortion of the model dressing image.
[0047] Secondly, after obtaining the morphing feature parameters of the original clothing image in the character dressing image corresponding to the body map of the model template, the application transforms the original clothing image according to the morphing feature parameters to obtain a clothing transformation image, and then uses the clothing transformation image and the body map of the model template as inputs. The human key point information implied by the body map can be used to predict the clothing effect applied to the body map, so as to generate a corresponding clothing mask. Then, the clothing transformation image is adjusted and synthesized into the body map according to the clothing mask, to obtain a model dressing image. Because the morphing feature parameters and the key point information of the body map of the model provide sufficient information, the corresponding clothing mask can effectively cover the corresponding area of the body map. When the clothing transformation image and the body map are synthesized by using the clothing mask, the generated clothing image can effectively mask the body parts of the body map, avoiding the edge alignment disorder effect that occurs when the clothing transformation image is directly superimposed on the body map. Thus, the actual effect of the character dressing image can be effectively transferred to the model template to generate a high-quality model dressing image.
[0048] In addition, the application only needs to provide a character dressing image by the user, and can automatically generate a model dressing image in one station, so as to realize the migration of the original clothing image in the character dressing image to the model dressing image, simplify the business process, and can be deployed in an e-commerce platform to generate model dressing related display image information for clothing goods, thereby saving the cost of publishing goods and improving the efficiency of processing goods on the shelf. BRIEF DESCRIPTION OF DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.
[0050] Figure 1 The network architecture schematic diagram of the application environment of the technical solutions of the application;
[0051] Figure 2 The principle schematic diagram of the exemplary model architecture of the application;
[0052] Figure 3 The flow schematic diagram of one embodiment of the model image generation method of the application;
[0053] Figure 4 The flow schematic diagram of determining the character dressing image in the embodiment of the application;
[0054] Figure 5 The flow schematic diagram of the preferred model template in the embodiment of the application;
[0055] Figure 6 A flowchart of a process for solving deformation feature parameters in an embodiment of the present application;
[0056] Figure 7 A flowchart of a process for obtaining a garment mask in an embodiment of the present application;
[0057] Figure 8 A flowchart of a process for synthesizing a model dressed image according to a garment mask in an embodiment of the present application;
[0058] Figure 9 A principle block diagram of a model image generation apparatus of the present application;
[0059] Figure 10 A structure diagram of a model image generation apparatus of the present application. DETAILED DESCRIPTION
[0060] The models referred to or possibly referred to in the present application, including traditional machine learning models or deep learning models, can be deployed on a remote server and remotely called by a client, or directly called by a client with sufficient device capability, unless otherwise specified. In some embodiments, when running on a client, the corresponding intelligence can be obtained through transfer learning to reduce the requirement for client hardware running resources and avoid excessive occupation of client hardware running resources.
[0061] Referring to Figure 1 , the network architecture adopted in the exemplary application scenario of the present application includes a terminal device 80, a service server 81 and an application server 82. The application server 82 can be used to deploy a virtual clothes changing service and provide a corresponding functional plug-in to an online store implemented on the service server 81 to call the virtual clothes changing service. When a user accesses a corresponding page of the online store of the service server 81 from the terminal device 80, the functional plug-in can be used to upload a person dressed image, and then the virtual clothes changing service opened by the application server 82 is called by the service server 81 to generate a model dressed image based on the person dressed image.
[0062] The virtual clothes changing service of the present application can be implemented by executing the model image generation method of the present application. Specifically, the model image generation method of the present application can be implemented as a computer program product and installed in a corresponding device such as the application server. After running, the method can be executed to open the virtual clothes changing service.
[0063] Of course, in another example network architecture of the present application, the virtual clothes changing service of the present application can be deployed in a terminal device for implementation, as long as the terminal device can have the software and hardware resources to implement the virtual clothes changing service.
[0064] Please refer to Figure 2 , Figure 2 An example model architecture is shown for implementing the model image generation method of the present application, which includes a first feature extraction model, a second feature extraction model, a deformation solving model, a human key point detection model, and an image segmentation model. According to the model architecture, one route of input information using a person dressed image and another route of input information using a model template are used, and through the model architecture, a garment mask can be obtained, which is used to guide the generation of the model dressed image required by the present application.
[0065] The first feature extraction model and the second feature extraction model can be constructed using a convolution-based neural network model, such as a convolutional neural network (CNN), a deep convolutional network structure (VGG), a ResNet (Residual Network) series residual network model, etc., which is used to implement semantic feature extraction on the input image to obtain corresponding deep semantic information.
[0066] The deformation solving model can be mathematically modeled based on traditional machine learning principles, or it can be implemented using a neural network model based on deep learning. Its purpose is to generate deformation feature parameters that focus on the deformation amplitude information of the original garment image in the person dressed image relative to the body map, using the two-way deep semantic information obtained by the first feature extraction model and the second feature extraction model through a linear regression algorithm.
[0067] The human key point detection model can be used to detect the human key point information in the original image and / or body map of the person dressed image and the model in the model template, to obtain the corresponding key point feature map, which is used to construct part of the input information of the first feature extraction model, the second feature extraction model, the image segmentation model, etc.
[0068] The image segmentation model can use a U-net, U 2a basic model implementation such as net, and input information is formed by combining the key point feature maps corresponding to the garment transformation image and the model template obtained according to the morphing feature parameters, and the input information is sequentially down-sampled and up-sampled by the image segmentation model, and finally a garment mask of the garment image corresponding to the model body adjusted to the model template is obtained, so as to synthesize the garment transformation image into the body map of the model according to the garment mask, to obtain a corresponding model dressed image.
[0069] The above exemplary model architecture can be obtained by joint training of corresponding training samples to a convergent state, and the training samples can be various image information obtained from a person dressed image and a model original image in a model template, which also includes a key point feature map obtained by using the human body key point detection model, and a body map of the model, etc., which can be flexibly set.
[0070] In one embodiment, in order to construct the input of the first feature extraction model, the person dressed image in the training sample can be processed into first path input information including a plurality of channels of key point feature maps and a plurality of channels of original images, and then input into the first feature extraction model, and the plurality of channel information of the model original image in the model dressed image and the hairstyle map, body map, key point feature map, etc. therein are constructed into second path input information and then input into the second feature extraction model. Among them, the body map in the model template can also include or be the mask information of the body part in the model original image.
[0071] Correspondingly, for the output obtained by the model architecture processing the training sample, the overall loss of the model can be calculated by using the corresponding annotation information, and in one embodiment, the overall loss of the model can be composed of three parts, wherein the first part is the L1 norm loss value between the garment transformation image obtained by processing the garment area image in the person dressed image by the morphing solving model and the actual garment area image in the model dressed image, the second part is the model perception loss value of the original image or the garment area image of each other in the person dressed image and the model dressed image in the training sample, and the third part is the L1 norm loss value between the garment mask predicted by the image segmentation model relative to the body map of the model and the garment mask actually annotated for the model dressed image. In each iteration training, the three loss values are weighted and summarized as the overall loss of the model, whether the model reaches the preset convergence condition is judged according to the overall loss of the model, when the model architecture does not converge, the model architecture is updated according to the overall loss of the model, and then the next training sample is used for iteration training, until the model converges. Otherwise, when the model architecture reaches the convergence state, the training of the model architecture is terminated, so that the model architecture can be put into the inference stage.
[0072] It should be noted that the above model architecture is an exemplary architecture, and in actual implementation, the various components thereof can be flexibly decoupled or replaced with components having equivalent functions, as long as the relevant components can directly or equivalently achieve the same functions required by the application. Similarly, for the various input information of the application, the preprocessing stage before inputting the entire model architecture can also be flexibly adjusted, as long as the model architecture can obtain the same inference ability as the above examples through training, for example, in constructing the input information of the second feature extraction model, in some embodiments, the hairstyle map and torso map of the model can not be provided, and in other embodiments, although the hairstyle map and torso map of the model are provided in the second input information, the hairstyle map and torso map can be mask information obtained by image segmentation of the original image of the model by means of an image segmentation model. Similarly, those skilled in the art can make appropriate modifications based on the above principle examples without departing from the scope of the inventive spirit of the present application.
[0073] Please refer to Figure 3 According to the model image generation method provided by the present application, in one embodiment, the method comprises the following steps:
[0074] Step S1100, obtaining a person dressing image;
[0075] The person dressing image can be an image obtained by photographing a person wearing a clothing commodity, and the image content usually includes the overall information of the person image and the clothing worn by the person.
[0076] The person dressing image can be uploaded by a user or obtained from a preview video stream generated by a camera of a terminal device of the user.
[0077] In one embodiment, a user can take a person dressing image by himself / herself, submit the person dressing image to a business server through a page for generating a commodity image provided by an online store, drive the business server to call an interface of an application server, and execute the method of the present application to generate a model dressing image in which the clothing image in the person dressing image is migrated to the model.
[0078] In another embodiment, after a computer program product implementing the present application is run, the camera of the terminal device is started, the preview video stream generated by the camera is intelligently detected to determine a person dressing image with better quality, and then the method of the present application is executed in the background to generate a corresponding model dressing image.
[0079] The character dressing image can be pre-processed to generate its related image information, for example, a key point feature map in it can be detected by means of a human key point detection model pre-trained to a convergent state. The key point feature map can be multi-channel feature data. For another example, an original clothing image and / or a mask thereof in the character dressing image can be obtained by means of an image segmentation model pre-trained to a convergent state. Further, the original clothing image can be further subjected to image inpainting by means of a conventional image inpainting model to inpaint the clothing image content of the part of the original clothing image occluded by the character torso.
[0080] Step S1200, recalling a model template with a similar pose to the human body in the character dressing image from a model gallery, the model template including a body map of a model;
[0081] The present application is provided with a model gallery containing a large number of model templates, each model template containing an original image of a model taken as a prototype. Generally, the original image is preferably a figure without wearing clothes, for example, a figure of the model wearing only a bra and underpants to expose the body contour information of the model to the maximum within a reasonable range.
[0082] On the basis of the original image in the model template, other related image information can also be obtained by pre-processing, for example, a body map, a hairstyle map, a torso map, etc. in the original image can be determined by means of an image segmentation model pre-trained to a convergent state. The body map, the hairstyle map, the torso map, etc. can also be represented as corresponding mask information. For another example, a key point feature map corresponding to the model original image or the body map can be pre-detected by means of a human key point detection model pre-trained to a convergent state. All these image information can be stored in the model template for calling as needed in each specific aspect of the technical solution of the present application.
[0083] It is not difficult to understand that when the pose of the character in the character dressing image is closer to the pose of the model in the original image in the model template, the better the image migration effect will be when the original clothing image in the character dressing image is migrated to the model body map. Accordingly, the present application utilizes the semantic similarity between the key point feature map obtained from the character dressing image and the key point feature map of the model template in the model gallery to optimally recall one or more model templates for the character dressing image. The recalled model template is generally the model template with the highest similarity to the character dressing image, indicating that the pose of the model in the model template is closer to the pose of the character in the character dressing image. The subsequent virtual dressing process based on these model templates can provide a plurality of model dressing images in batches. It should be noted that, for the convenience of illustration and understanding, a single model template will be taken as an example for illustration in the present application.
[0084] Step S1300, determining a deformation feature parameter of an original garment image in the person dressing image relative to the body map, transforming the original garment image according to the deformation feature parameter to obtain a garment transformed image;
[0085] In order to implement the image dressing processing, it is necessary to determine a deformation feature parameter of an original garment image in the person dressing image relative to the body map in the model template, and to transform the original garment image according to the deformation feature parameter to obtain a garment transformed image.
[0086] To this end, the original garment image or its mask information in the person dressing image can be obtained by image segmentation, and then the corresponding key point feature map of the human body in the person dressing image is merged to construct one-way input information. On the basis of obtaining the body map or its mask information of the model template, the corresponding key point feature map of the body map is merged to form another way of input information. The information from the two sources is constructed as input information respectively, and a pre-modeled deformation solving model is used for linear regression prediction, so as to solve the deformation feature parameter that the original garment image needs to generate relative to the body map. Since the prediction ability of the deformation solving model can be obtained by training the model architecture of the example of the present application, the deformation feature parameter can be effectively determined by the deformation solving model.
[0087] It is not difficult to understand that the role of the deformation feature parameter is to focus on the deformation amplitude information that the original garment image in the person dressing image should generate relative to the body map. Therefore, image deformation operation can be implemented on the basis of the original garment image according to the deformation feature parameter, and through the image deformation operation, the corresponding garment transformed image is obtained. Theoretically, the garment transformed image is suitable for the body of the model in the model template, and the image of the original garment image in the person dressing image is obtained after being worn on the body of the model.
[0088] Step S1400, predicting a garment mask corresponding to the body map according to the garment transformed image and the body map;
[0089] In actual use, the edge alignment of the garment transformation image obtained by transforming the original garment image according to the deformation feature parameter is often inaccurate. If the garment transformation image is directly combined with the body map or the original image in the model template, the quality of the obtained image is poor, which is caused by the fact that, in order to effectively determine the deformation feature parameter, the deformation solving model must rely on the key point information of the person and the model, and thus the key point feature map corresponding to the person and the model must be referenced in the two-way input information, so as to enable the deformation solving model to predict the deformation feature parameter corresponding to the completely covered body region of the model, and thus to determine the garment transformation image completely covering the body region. However, due to the existence of the key point information of the person and the model, the deformation solving model cannot determine the garment transformation image completely covering the body region of the model, for example, the garment transformation image corresponding to the shoulder, elbow and leg is smaller or larger than the corresponding region of the body map of the model, which leads to poor final synthesis effect.
[0090] It is not difficult to understand that when the garment transformation image is locally larger than the corresponding region of the body map, it is not easy to see in the finally synthesized model dressed image, but if it is smaller than the corresponding region of the body map, the skin will be exposed in the finally synthesized model dressed image, which is obviously not expected.
[0091] In order to solve this problem, the garment transformation image, which can also be represented as mask information of the garment transformation image, is combined with the body map or the mask information thereof in the model template to construct input information, which is provided to the image segmentation model in the model architecture of the present application to implement image segmentation and obtain the garment mask corresponding to the adjustment of the garment transformation image to the body map.
[0092] In the input information of the image segmentation model, one of the bases is the model body map, which can also be replaced by the model original image. Since the key point information of the model is implied in the body map or the original image, the image segmentation model can effectively mine these key point information to moderately adjust the garment transformation image, so as to obtain the garment mask corresponding to the adjustment of the garment transformation image to the body region of the model. The garment mask can more accurately indicate how the garment transformation image in the person dressed image adjusts the contour and position distribution to more accurately match the body map of the model.
[0093] According to the above principle, in another embodiment, the body map in the image segmentation model can be replaced by the key point feature map obtained by detecting the original image or the body map of the template by using the human key point detection model. Since the body map is intercepted from the original image, it is also essentially dependent on the key point feature map corresponding to the body map.
[0094] Step S1500, synthesizing the clothes transform image into the body map according to the clothes mask, to obtain a model dressed image.
[0095] The clothes mask generally indicates the clothes area and non-clothes area in the body map of the model template in the form of 1, 0, according to which, the clothes transform image can be synthesized into the body map according to the area information indicated by the clothes mask, so as to obtain a corresponding model dressed image. It is not difficult to understand that the model dressed image obtained in this case has a more neat dressing effect, and has a high probability of not appearing inaccurate alignment of clothes and body.
[0096] In 1000 real samples actually measured by the applicant, combined with subjective evaluation, only a small part appears local alignment inaccurate, and the synthesis qualified rate is as high as 86.35%, which has extremely high practical significance for an artificial intelligence model based on deep learning.
[0097] According to the above embodiments, the present application has many technical advantages, including but not limited to:
[0098] Firstly, the present application takes the person dressed image as the source of the original clothes image, matches the model template with similar posture by using the posture of the human body in the person dressed image, and implements virtual clothes changing transformation on the basis of the selected model template, so that the generated model dressed image is closer to the posture of the person dressed image, and the distortion of the model dressed image is maximized.
[0099] Secondly, after obtaining the deformation feature parameters of the original clothes image in the person dressed image corresponding to the body map of the model template, the present application uses the clothes transform image and the body map of the model template as input, uses the key point information of the human body implied by the body map to predict the clothes effect applied to the body map, generates a corresponding clothes mask, and adjusts and synthesizes the clothes transform image into the body map according to the clothes mask, to obtain a model dressed image. Because the deformation feature parameters and the key point information of the body map of the model provide sufficient information, the corresponding clothes mask can effectively cover the corresponding area of the body map, and when the clothes mask is used to synthesize the clothes transform image and the body map, the generated clothes image can effectively mask the body part of the body map, avoiding the edge alignment disorder effect of directly superimposing the clothes transform image on the body map, so as to effectively transfer the actual effect of the person dressed image to the model template to generate a high-quality model dressed image.
[0100] In addition, the present application only needs to provide a person dressing image by a user, and can automatically generate a model dressing image in one station, realizes migration of an original clothing image in the person dressing image to the model dressing image, simplifies a business process, and can be deployed in an e-commerce platform to generate a model dressing related display image information for a clothing commodity, thereby saving commodity release cost and improving commodity listing processing efficiency.
[0101] On the basis of any embodiment of the present application, please refer to Figure 4 , obtaining a person dressing image, comprising:
[0102] Step S1110, reading a preview video stream generated by a camera device, performing human body recognition on each image frame, and determining an image frame containing a human body image as a person image;
[0103] The present embodiment can intelligently obtain a person dressing image by means of a camera device in a terminal device. For this purpose, the user usually manually operates or starts through a background instruction to make the camera device process a working state and start to record a preview video stream, so as to facilitate reading each image frame in the preview video stream from the image space thereof, and intelligently recognizing on the basis of these image frames to determine an effective person dressing image.
[0104] For the preview video stream generated by the camera device, each image frame can be sequentially read, and then a preset human body detection model is used for detection to determine whether a human body image exists therein. When a human body image exists in an image frame, the image frame can be initially determined as a person image. The human body detection model can be obtained by using a mature known model or self-construction and training.
[0105] Step S1120, determining a key point feature map in each person image by using a human key point detection model, and calculating semantic similarity of the key point feature map and a key point feature map of a body map of a model template in a model library;
[0106] In one embodiment, a certain detection time length can be preset, for example, 1 second. Within the detection time length, a plurality of person images can be obtained. Therefore, these person images can be selected to determine one or more person dressing images therefrom.
[0107] For each of the person image, a human key point detection model described in the present application can be used to determine the key point feature map in real time, and the key point feature map is essentially a feature matrix. For the convenience of calculation, the feature matrix of each person image can be spliced into a corresponding feature vector row by row. Similarly, the original image or the body image in the model template in the model gallery of the present application can also use the same human key point detection model to determine its feature matrix and even its corresponding feature vector.
[0108] For the feature vector of each person image, the data distance between the feature vector and the feature vector of each model template in the model gallery can be calculated, and the data distance is converted into semantic similarity. For this purpose, each person image can obtain the semantic similarity of each model template, and the highest semantic similarity can be used to determine whether the corresponding person image can be regarded as the person dressing image of the present application.
[0109] The data distance between the two feature vectors can be calculated by using any one of the following algorithms: cosine similarity, Euclidean distance, vector dot product, Pearson correlation coefficient, and Jaccard coefficient.
[0110] The above method of determining the semantic similarity between two key point feature maps can also be applied to other scenarios of the present application.
[0111] Step S1130, when the highest semantic similarity obtained by the person image exceeds the preset threshold, the person image is determined as the person dressing image.
[0112] When determining whether a person image can be used as the person dressing image of the present application, the highest semantic similarity of the person image corresponding to each model template in the model gallery can be compared with a preset threshold. When the highest semantic similarity is higher than the preset threshold, the person image can be determined as the person dressing image of the present application.
[0113] According to the above method, there can be multiple person dressing images in a detection time range. In this case, these person dressing images can be regarded as pre-selected images, and further optimization can be performed on the basis of this. For this purpose, the person dressing image with the highest semantic similarity can be selected as the final person dressing image.
[0114] In other alternative embodiments, a mature image sharpness analysis model can be used to analyze the image sharpness of multiple pre-selected images, and the pre-selected image with the highest image sharpness can be determined as the final person dressing image.
[0115] In further embodiments, the final image of the character dress to be used can be determined by taking into account both the clarity and the semantic similarity, specifically, the clarity and the highest semantic similarity of the preselected image can be matched with different weights respectively, wherein the weight of the clarity is higher than the weight of the highest semantic similarity, or vice versa, and then the total score is calculated by adding, and the preselected image with the highest total score is selected as the final image of the character dress to be used.
[0116] According to the above embodiments, there are various ways to intelligently obtain high-quality images of characters in dress from the camera, avoiding the influence of manual operation on obtaining high-quality model dress images, such as image defocus and shaking, and more easily matching a model with a similar pose to the character in the image of the character dress from the model template, to ensure that high-quality model dress images are obtained based on high-quality images of characters in dress.
[0117] Based on any of the embodiments of the present application, please refer to Figure 5 , recalling a model template with a similar pose to the human body in the image of the character dress from the model gallery, the model template including a body map of the model, comprising:
[0118] Step S1210, detecting and determining the key point feature map of the image of the character dress by using the human key point detection model;
[0119] As mentioned above, if the pose of the model in the model template is similar to the pose of the human body in the image of the character dress, the model architecture of the present application can be expected to more accurately transfer the image of the dress in the image of the character dress to the body map of the model template, to obtain an image effect of accurate alignment of the dress and the body. Therefore, by analogy, the human key point detection model can be used to detect the key point information of the image of the character dress, to determine the corresponding key point feature map.
[0120] Step S1220, calculating the semantic similarity between the key point feature map of the image of the character dress and the key point feature map of the body map of each model template in the model gallery, the key point feature map of the body map of each model template being pre-detected and determined by using the human key point detection model on the body map;
[0121] The way to determine the semantic similarity between the key point feature map of the image of the character dress and the key point feature map of the body map (or its original image) of each model template in the model gallery is the same as the implementation manner of the previous embodiment, which can be determined by calculating the data distance of the feature vectors of the two key point feature maps, and will not be described here. It should be emphasized that the key point feature map corresponding to the original image or the body map of the model template in the model template has been pre-stored, and the key point feature map has been detected and determined by means of the human key point detection model in advance.
[0122] Step S1230, screening out at least one model template with high semantic similarity as a model template with a similar pose to the human body in the character dressing image.
[0123] After the processing of the above steps, for a character dressing image, each model template in the model gallery can obtain a corresponding semantic similarity. The higher the semantic similarity, the more similar the pose of the human body in the model template to the pose of the human body in the character dressing image. Accordingly, each model template can be sorted according to the semantic similarity from large to small, and then one or more model templates in the front of the sorting can be selected according to a preset default number or a user-specified number, for implementing virtual dressing, and one or more model dressing images are obtained accordingly.
[0124] In one embodiment, a preset threshold can also be set. For one or more model templates in the front of the sorting, only the model template with a semantic similarity higher than the preset threshold is used to generate a model dressing image, and the model template with a semantic similarity lower than the preset threshold is not used. In this way, a minimum threshold can be set for the selection of the model template, to ensure the effective generation of the model dressing image.
[0125] According to the above embodiments, it can be seen that the present application does not need to rely on manual determination of the model template, but can automatically match one or more model templates according to the key point information contained in the character dressing image. The pose of the human body of the model in these matched model templates is closer to the pose of the human body in the character dressing image, and thus it is more helpful to generate a high-quality model dressing image.
[0126] On the basis of any embodiment of the present application, please refer to Figure 6 to determine the deformation feature parameters of the original clothing image in the character dressing image relative to the body map, including:
[0127] Step S1310, obtaining the original clothing image and key point feature map in the character dressing image to construct user input information, and inputting the first feature extraction model to extract the first deep semantic information;
[0128] Please also understand this embodiment in combination with the model architecture shown in Figure 2 According to the model architecture of Figure 2 , the input information of the first feature extraction model needs to be constructed by using the original clothing image and the key point feature map in the character dressing image finally used.
[0129] The original clothing image, as described above, can be preprocessed by using a preset image segmentation model on the character dressing image to segment the clothing image region therein and obtain the corresponding original clothing image.
[0130] The key point feature map can also be detected by a preset human key point detection model to obtain the key point feature map.
[0131] In one embodiment, according to the specific function realized by the corresponding model, the original clothing image can be processed into, for example, 3 channels, and the key point feature map can be processed into 15 channels, thereby forming 18 channels of user input information, which is provided to the first feature extraction model for image feature extraction.
[0132] The first feature extraction model is realized by a convolutional neural network as the bottom layer, and the image feature information is extracted by performing convolution operation on the user input information to obtain the first deep semantic information.
[0133] In step S1320, the original image, the body image, the hairstyle image, and the key point feature map of the model included in the model template are constructed as template input information, which is input into the second feature extraction model to extract the second deep semantic information.
[0134] Similarly, for the second feature extraction model, the model template to be input needs to be constructed as template input information, and the second feature extraction model extracts image features based on the template input information to obtain the corresponding second deep semantic information.
[0135] When constructing the template input information, in one embodiment, the original image in the model template can be processed into 3 channels of information, and then combined with the 1 channel of body image, 1 channel of hairstyle image obtained by preprocessing the original image, and combined with the 13 channels of key point feature map determined by detecting the original image or the body image by using the human key point detection model, to construct 18 channels of template input information. In other embodiments, the above examples can be appropriately modified, for example, the body image can be replaced by the mask information corresponding to the body part in the original image, and the hairstyle image is the same. In other embodiments, the template's torso image or its mask information obtained by preprocessing the original image can also be added when constructing the template input information. And so on, the template input information can be flexibly constructed, and the key elements are the image resources representing the model's body image, such as the original image, the body image or its mask information, and the key point feature map representing the model's body posture. According to this principle, those skilled in the art can modify the template input information according to the actual situation, as long as the input consistency is maintained during the model architecture training and reasoning stage of the present application.
[0136] Step S1330, predicting the deformation feature parameter according to the first deep semantic information and the second deep semantic information, the deformation feature parameter being used to pay attention to the deformation amplitude information of the original clothing image in the person dressed image relative to the body map.
[0137] As described above, the deformation solving model of the present application can solve the deformation feature parameter required for image transformation of the original clothing image in the person dressed image when the clothing image in the person dressed image is correspondingly migrated to the model body image based on the two-way input information correspondingly constructed based on the person dressed image and the model template, which has previously obtained such reasoning ability by mathematical modeling and training. Therefore, after the first feature extraction model and the second feature extraction model respectively input the first deep semantic information and the second deep semantic information obtained by themselves into the deformation solving model, the deformation solving model can obtain the corresponding deformation feature parameter by linear regression, pay attention to the deformation amplitude information of the original clothing image in the person dressed image relative to the body map through the deformation feature parameter, so that the corresponding clothing transformation image of the original clothing image after being assumed to migrate to the body map of the model template can be generated according to the deformation feature parameter.
[0138] As can be seen from the above embodiments, after the person dressed image and the model template are respectively constructed to input information and image feature extraction is performed to obtain their respective deep semantic information, the deformation feature parameter corresponding to the clothing image migration in the virtual clothing changing process can be efficiently obtained by linear regression, the clothing transformation image can be quickly determined according to the deformation feature parameter, the intelligent degree is high, the centralized service is conveniently provided, the deformation feature parameter and the corresponding clothing transformation image can be quickly obtained by calling the corresponding service interface, so that the model dressed image can be quickly obtained.
[0139] On the basis of any embodiment of the present application, please refer to Figure 7 , predicting the clothing mask corresponding to the body map according to the clothing transformation image and the body map, comprising:
[0140] Step S1410, constructing the key point feature map corresponding to the body map of the model template and the clothing transformation image as a changed clothing image information;
[0141] In the present embodiment, in order to improve Figure 2 the ability of the image segmentation model in the model architecture shown in the figure to predict the clothing image contour corresponding to the clothing image migrated to the body map of the model, the key point feature map corresponding to the body map or the original image of the model template is obtained in advance, and the clothing transformation image obtained according to the deformation feature parameter is combined with the key point feature Figure 1The constructed image transformation information is input into the image segmentation model.
[0142] In one embodiment, the clothing transformation image in the image transformation information can also be represented as mask information, which is represented in the form of binary values, and can reduce the operation amount of the image segmentation model. The key point feature map can provide deformation information for the clothing transformation image, so as to guide the image segmentation model to make posture correction and corresponding adjustment on the clothing transformation image or the mask information thereof.
[0143] In step S1420, the image transformation information is input into a preset image segmentation model for image segmentation, and a clothing mask corresponding to the adjustment of the clothing transformation image to the body map is predicted.
[0144] After the image transformation information is input into the image segmentation model, the image segmentation model performs down-sampling and up-sampling on the image transformation information at different scales to obtain feature maps corresponding to each scale, and then a clothing mask represented in binary form is obtained by synthesizing these feature maps. Since the image segmentation model has learned the ability to adaptively adjust the clothing transformation image according to the key point information provided by the key point feature map corresponding to the posture of the model template through prior training, the clothing mask obtained by the image segmentation model can be effectively aligned to each pixel in the body map of the model to provide more accurate contour and position distribution description, effectively and accurately describe the correct image position of the original clothing image in the model template after the clothing image in the character dressing image is migrated to the body map of the model template, and ensure that the clothing image migration effect with accurate alignment can be obtained.
[0145] According to the above embodiments, under the optimization of the image segmentation model in the model architecture of the present application, instead of directly synthesizing the model dressing image according to the deformation feature parameters, a clothing mask that more accurately represents the position and contour of the original clothing image in the body map of the model template is obtained by means of the image segmentation model. When the model dressing image is synthesized according to the clothing mask, an excellent edge alignment effect can be obtained.
[0146] On the basis of any embodiment of the present application, please refer to Figure 8 The clothing transformation image is synthesized into the body map according to the clothing mask to obtain a model dressing image, including:
[0147] In step S1510, the clothing transformation image is corrected according to the clothing mask, and the clothing transformation image is synthesized into the body map of the model in the model template to obtain a preliminary dressing image.
[0148] As mentioned above, the garment mask obtained by the image segmentation model can accurately specify the correct position and contour of the migrated garment image in the body map of the model template, and according to the binary information provided by the garment mask, the pixels in the garment transform image are extracted and correspondingly synthesized into the body map of the model template as the foreground of the body map, thereby obtaining a preliminary dressed image with the garment image attached on the basis of the body map.
[0149] Step S1520, the hairstyle map and the torso map in the model template are used to correct the preliminary dressed image to obtain a body state corrected image.
[0150] Since the image segmentation model does not consider the occlusion of hairstyle, torso and other body parts on the garment when generating the garment mask, it is easy to cause the phenomenon of improper image of the garment covering the body or the body covering the garment. Therefore, the hairstyle map and the torso map obtained by preprocessing the model original image in the model template can be further used to correct the preliminary dressed image, and the preliminary dressed image is adjusted according to the original relationship of the hairstyle and the torso in the model original image to obtain a body state corrected image.
[0151] In one embodiment, the body state corrected image can be directly used as the model dressed image without considering the existence of background of the model dressed image.
[0152] Step S1530, the body state corrected image is synthesized with the background map in the model template as the foreground to obtain a model dressed image.
[0153] Further, in the embodiment in which the model template provides a background map, the background map of the model template can be called as the background, and then the body state corrected image is synthesized as the foreground to obtain a final model dressed image.
[0154] According to the above embodiments, the garment transform image is migrated into the body map of the model template according to the garment mask, which can make the position and contour relationship of the garment and the body map more accurate, and further combined with necessary image correction processing, a high-quality model dressed image can be obtained, making the model dressed image more practical.
[0155] In another embodiment further extended on the basis of any embodiment of the present application, a mature face swapping model can also be used to swap faces in the model dressing image, replace the face image in the model dressing image with the face image uploaded by the user, enrich the technical implementation means for the user to obtain the model dressing image, and make the corresponding business logic more perfect. As an example, the face swapping model can be implemented using faceshifter, whose original title is: towards high fidelity and occlusion aware face swapping. Of course, other face replacement models with equivalent functions can also be used to implement it, and those skilled in the art can handle it flexibly.
[0156] Before the original clothing image and key point feature map of the character dressing image are constructed as user input information on the basis of any embodiment of the present application, the following steps are included:
[0157] Step S2100, performing image segmentation on the character dressing image to obtain the original clothing image therein;
[0158] As described above, a mature image segmentation model can be used to segment the character dressing image to obtain the original clothing image therein, so as to pre-process the character dressing image for constructing user input information for the first feature extraction model.
[0159] Step S2200, using an image inpainting model to inpaint the occluded part of the original clothing image in the character dressing image.
[0160] The image segmentation model usually does not have the ability to correctly identify the occlusion relationship of the clothing when performing image segmentation. Therefore, a corresponding trained image inpainting model can be used to perform image inpainting processing on the original clothing image to inpaint the occluded part therein, so that the clothing style in the original clothing image is more complete.
[0161] The image inpainting model used in the present application can use a resolution-robust large mask inpainting model based on Fourier convolution for inpainting. The model was published in 2021, and the original title is: Resolution-robust Large Mask Inpainting with Fourier Convolutions. Of course, it is also feasible to replace it with other models with equivalent functions.
[0162] According to the above embodiments, it can be seen that the original clothing image is repaired on the basis of the original clothing image, the image information of the original clothing image is perfected, the adverse effects caused by the incomplete original clothing image due to human body actions and shooting angles are avoided, and effective model dressing images can be ensured to be obtained.
[0163] Please refer to Figure 9 According to an aspect of the present application, a model image generation device is provided, which comprises an image acquisition module 1100, a template recall module 1200, a deformation calculation module 1300, a mask generation module 1400, and an image synthesis module 1500. The image acquisition module 1100 is configured to acquire a person dressing image. The template recall module 1200 is configured to recall a model template with a similar pose to a human body in the person dressing image from a model image library, and the model template comprises a body map of a model. The deformation calculation module 1300 is configured to determine a deformation feature parameter of an original clothing image in the person dressing image relative to the body map, and transform the original clothing image according to the deformation feature parameter to obtain a clothing transformed image. The mask generation module 1400 is configured to predict a clothing mask corresponding to the body map according to the clothing transformed image and the body map. The image synthesis module 1500 is configured to synthesize the clothing transformed image into the body map according to the clothing mask to obtain a model dressing image.
[0164] On the basis of any embodiment of the present application, the image acquisition module 1100 comprises a preview extraction unit configured to read a preview video stream generated by a camera, perform human body recognition on each image frame, and determine an image frame containing a human body image as a person image. A similarity calculation unit is configured to determine a key point feature map in each person image by using a human body key point detection model, and calculate semantic similarity of the key point feature map and a key point feature map of a body map of a model template in a model image library. An automatic capture unit is configured to determine the person image as the person dressing image when the highest semantic similarity obtained by the person image exceeds a preset threshold.
[0165] On the basis of any embodiment of the present application, the template recall module 1200 comprises: a key point detection unit configured to detect a key point feature map of the character dressing image by using a human key point detection model; a semantic matching unit configured to calculate semantic similarity between the key point feature map of the character dressing image and a key point feature map of a body map of each model template in the model template library, the key point feature map of the body map of each model template being detected and determined in advance by using the human key point detection model on the body map; and a template optimization unit configured to select at least one model template with high semantic similarity as a model template with a similar pose to the human body in the character dressing image.
[0166] On the basis of any embodiment of the present application, the deformation obtaining module 1300 comprises: a first extraction unit configured to obtain an original clothing image in the character dressing image and a key point feature map as user input information, and input a first feature extraction model to extract first deep semantic information; a second extraction unit configured to obtain an original image, a body map, a hairstyle map and a key point feature map of a model included in the model template as template input information, and input a second feature extraction model to extract second deep semantic information; and a parameter solving unit configured to predict the deformation feature parameter according to the first deep semantic information and the second deep semantic information, the deformation feature parameter being used to focus on deformation amplitude information of the original clothing image in the character dressing image relative to the body map.
[0167] On the basis of any embodiment of the present application, the mask generation module 1400 comprises: a dressing construction unit configured to construct a dressing image information by using the key point feature map corresponding to the body map of the model template and the clothing transformation image; and a mask prediction unit configured to input the transformation image information into a preset image segmentation model to perform image segmentation and predict a clothing mask corresponding to the clothing transformation image adjusted to the body map.
[0168] On the basis of any embodiment of the present application, the image synthesis module 1500 comprises: a preliminary synthesis unit configured to correct the clothing transformation image according to the clothing mask, synthesize the clothing transformation image into the body map of the model in the model template, and obtain a preliminary dressing image; a detail correction unit configured to correct the preliminary dressing image by using the hairstyle map and the torso map in the model template, and obtain a body state correction image; and a panoramic synthesis unit configured to synthesize the body state correction image as a foreground with a background map in the model template to obtain a model dressing image.
[0169] On the basis of any embodiment of the present application, before the first extraction unit, comprising: an original segmentation unit configured to perform image segmentation on the person dressed image to obtain an original clothing image therein; and an original inpainting unit configured to repair an occluded part of the original clothing image in the person dressed image by using an image inpainting model.
[0170] Another embodiment of the present application also provides a model image generation device. As shown in the figure, a schematic diagram of the internal structure of the model image generation device. The model image generation device includes a processor, a computer readable storage medium, a memory and a network interface connected by a system bus. Among them, the computer readable non-volatile readable storage medium of the model image generation device, the operating system, the database and the computer readable instructions are stored, the database can store information sequence, and the computer readable instructions are executed by the processor, so that the processor realizes a model image generation method. Figure 10
[0171] The processor of the model image generation device is used to provide computing and control ability to support the operation of the whole model image generation device. The memory of the model image generation device can store computer readable instructions, which can make the processor execute the model image generation method of the present application when executed by the processor. The network interface of the model image generation device is used to connect and communicate with the terminal.
[0172] Those skilled in the art can understand that Figure 10 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the model image generation device to which the scheme of the present application is applied. The specific model image generation device can include more or less components than those shown in the figure, or combine certain components, or have different component arrangement.
[0173] The processor in the embodiment is used to execute the specific functions of each module in Figure 9 The memory stores the program code and various data required for executing the above-mentioned modules or sub-modules. The network interface is used to realize the data transmission between the user terminal or the server. The non-volatile readable storage medium in the embodiment of the present application stores the program code and data required for executing all modules in the model image generation device of the present application. The server can call the program code and data of the server to execute the functions of all modules.
[0174] The present application also provides a non-volatile readable storage medium storing computer readable instructions, which are executed by one or more processors to make one or more processors execute the steps of the model image generation method of any embodiment of the present application.
[0175] The application also provides a computer program product comprising computer programs / instructions which, when executed by one or more processors, implement the steps of the method according to any of the embodiments of the application.
[0176] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments of the application can be completed by a computer program instructing relevant hardware, and the computer program can be stored in a non-volatile readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments of the method. The storage medium can be a computer readable storage medium such as a magnetic disc, an optical disc, a read-only memory (ROM), or a random access memory (RAM).
[0177] In summary, the application can accurately migrate the original clothing image in the character dressing image to the model body map of the model template to synthesize an effective model dressing image, facilitate the e-commerce platform to quickly generate the model dressing image corresponding to the clothing commodity when publishing the clothing commodity, and thus improve the commodity publishing efficiency.
Claims
1. A model image generation method characterized by comprising: The method comprises the following steps: obtaining a person dressing image, wherein the person dressing image is an image obtained after an arbitrary person wears a clothing commodity; recalling a model template with a similar pose to a human body in the person dressing image from a model library, wherein the model template comprises a body map and a hairstyle map of a model; determining a deformation feature parameter of an original clothing image in the person dressing image relative to the body map, and transforming the original clothing image according to the deformation feature parameter to obtain a clothing transformation image, comprising: obtaining the original clothing image and key point feature map in the person dressing image to construct user input information, and inputting a first feature extraction model to extract first deep semantic information; obtaining the original image, body map, hairstyle map and key point feature map of the model included in the model template to construct template input information, and inputting a second feature extraction model to extract second deep semantic information; predicting the deformation feature parameter according to the first deep semantic information and the second deep semantic information, wherein the deformation feature parameter is used to pay attention to the deformation amplitude information of the original clothing image in the person dressing image relative to the body map; predicting a clothing mask corresponding to the body map according to the clothing transformation image and the body map; synthesizing the clothing transformation image into the body map according to the clothing mask to obtain a model dressing image, comprising: correcting the clothing transformation image according to the clothing mask, and synthesizing the clothing transformation image into the body map of the model in the model template to obtain a preliminary dressing image; correcting the preliminary dressing image by using the hairstyle map and the torso map in the model template to obtain a body state correction image; synthesizing the body state correction image as a foreground with a background image in the model template to obtain a model dressing image.
2. The model image generation method according to claim 1, characterized by, recalling a model template with a similar pose to a human body in the person dressing image from a model library, wherein the model template comprises a body map of a model, comprising: detecting and determining a key point feature map of the person dressing image by using a human key point detection model; calculating semantic similarity between the key point feature map of the person dressing image and the key point feature map of the body map of each model template in the model library, wherein the key point feature map of the body map of each model template is detected and determined by using the human key point detection model in advance; screening at least one model template with higher semantic similarity as a model template with a similar pose to the human body in the person dressing image.
3. The model image generation method according to claim 1, characterized by, predicting a clothing mask corresponding to the body map according to the clothing transformation image and the body map, comprising: constructing the key point feature map corresponding to the body map of the model template and the clothing transformation image as a dressing image information; inputting the dressing image information into a preset image segmentation model for image segmentation to predict a clothing mask corresponding to the clothing transformation image adjusted to the body map.
4. The model image generation method according to claim 1, characterized by, Before the user input information is constructed by obtaining the original clothing image and the key point feature map of the person dressing image, comprising: performing image segmentation on the person dressing image to obtain the original clothing image; An image inpainting model is used to inpaint the occluded part of the original clothing image in the person-dressed image.
5. The model image generation method according to any one of claims 1 to 4, characterized by, The person-dressed image is obtained, including: A preview video stream generated by a camera is read, human body recognition is performed on each image frame, and an image frame containing a human body image is determined as a person image; A human body key point detection model is used to determine a key point feature map in each person image, and a semantic similarity is calculated between the key point feature map and a key point feature map of a body map of a model template in a model gallery; When the highest semantic similarity obtained by the person image exceeds a preset threshold, the person image is determined as the person-dressed image.
6. A model image generating apparatus characterized by comprising: It includes: An image acquisition module is configured to obtain a person-dressed image, the person-dressed image being an image obtained after a person wears a clothing commodity; A template recall module is configured to recall a model template with a similar posture to a human body in the person-dressed image from a model gallery, the model template including a body map and a hairstyle map of a model; A deformation obtaining module is configured to determine a deformation feature parameter of an original clothing image in the person-dressed image relative to the body map, and to transform the original clothing image according to the deformation feature parameter to obtain a clothing transformed image, including: obtaining the original clothing image and the key point feature map in the person-dressed image as user input information, and inputting a first feature extraction model to extract first deep semantic information; obtaining the original image, the body map, the hairstyle map, and the key point feature map of the model included in the model template as template input information, and inputting a second feature extraction model to extract second deep semantic information; predicting the deformation feature parameter according to the first deep semantic information and the second deep semantic information, the deformation feature parameter being used to focus on deformation amplitude information of the original clothing image in the person-dressed image relative to the body map; A mask generation module is configured to predict a clothing mask corresponding to the body map according to the clothing transformed image and the body map; An image synthesis module is configured to synthesize the clothing transformed image into the body map according to the clothing mask to obtain a model-dressed image, including: correcting the clothing transformed image according to the clothing mask, and synthesizing the clothing transformed image into the body map of the model in the model template to obtain a preliminary dressed image; correcting the preliminary dressed image using the hairstyle map and the torso map in the model template to obtain a body state corrected image; and synthesizing the body state corrected image as a foreground with a background image in the model template to obtain the model-dressed image.
7. The model image generation apparatus according to claim 6, characterized by The device further includes: An original segmentation unit is configured to perform image segmentation on the person-dressed image to obtain an original clothing image therein; An original inpainting unit is configured to use an image inpainting model to inpaint an occluded part of the original clothing image in the person-dressed image.
8. The model image generation apparatus according to claim 6 or 7, characterized by The image acquisition module includes: A preview extraction unit is configured to read a preview video stream generated by a camera, perform human body recognition on each image frame, and determine an image frame containing a human body image as a person image; The similarity calculation unit is configured to determine a key point feature map of each of the person images by using a human key point detection model, and calculate semantic similarity between the key point feature map and a key point feature map of a body map of a model template in a model gallery; The automatic capture unit is configured to determine the person image as the person dressing image when a highest semantic similarity obtained by the person image exceeds a preset threshold.
9. A model image generation device comprising a central processing unit and a memory, characterized by The central processing unit is configured to invoke a computer program stored in the memory to execute steps of the method according to any one of claims 1 to 5.
10. A non-volatile readable storage medium, characterized by The computer program is stored in the form of computer readable instructions, and when the computer program is invoked and run by a computer, steps included in the corresponding method are executed.
Citation Information
Patent Citations
3D garment virtual garment system
CN109919727A
KR20220066564A