A training-free virtual fitting image generation method and system based on a main posture and dense human body surface mapping
Patent Information
- Application Number
- CN202610766881.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-29
- Publication Date
- 2026-09-08
AI Technical Summary
1,服装结构保真度不足:扩散模型可能改变目标服装的纹理、图案、袖长、领口或裙摆结构;2,人体姿态与服装对齐不稳定:目标人物和服装参考图像之间存在姿态差异时,服装区域容易发生错位;3,稀疏姿态约束不足:仅利用人体关键点难以准确描述躯干、四肢、裙摆等局部区域的密集对应关系;4,人体语义区域混淆:服装区域、皮肤区域、头发区域和背景区域之间容易互相污染;5,训练成本高:针对虚拟试衣任务重新训练或微调扩散模型需要大量数据和计算资源,不利于快速部署;6,不同服装类型适应性差:上衣、裤装、裙装、连衣裙等服装具有不同的形变规律,统一处理容易导致局部畸变
1,本发明可直接利用预训练扩散式图像修复模型进行虚拟试衣生成,降低训练成本和部署成本。
Smart Images

Figure CN122714596A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of virtual fitting technology, and in particular relates to a training-free virtual fitting image generation method and system based on master pose and dense human body surface mapping. Background Technology
[0002] Virtual try-on technology aims to naturally transfer target clothing into an image of a target person, allowing the person to retain their original identity features, posture, body contours, background information, and appearance outside the clothing area, while presenting the visual effect of wearing the target clothing. This technology can be widely applied in e-commerce clothing displays, online try-on, intelligent shopping guides, digital human generation, advertising material generation, film and animation production, and personalized clothing recommendations.
[0003] Existing virtual try-on methods can be broadly categorized into geometric deformation-based methods, generative adversarial network (GAN)-based methods, and diffusion model-based methods. Geometric deformation-based methods typically deform clothing using keypoints, thin-plate spline transformations, or homography transformations, then fit the deformed clothing onto the human body. However, these methods are prone to issues such as clothing distortion, unnatural edges, and discontinuities in the human body structure under conditions of large pose changes, complex clothing structures, occlusion, and non-rigid deformation. GAN-based methods can improve image realism to some extent, but they usually rely on large amounts of paired or unpaired training data and lack generalization ability in scenarios involving cross-category clothing, complex human poses, multiple figures, or non-realistic figures.
[0004] In recent years, pre-trained diffusion-based image generation models have demonstrated strong generative capabilities in image inpainting, image editing, and conditional generation tasks. However, when directly using general diffusion-based image inpainting models for virtual try-on, the following problems still exist: 1. Insufficient fabric structure fidelity: The diffusion model may alter the texture, pattern, sleeve length, neckline, or hem structure of the target garment. 2. Unstable alignment between human pose and clothing: When there are pose differences between the target person and the clothing reference image, the clothing area is prone to misalignment. 3. Insufficient sparse pose constraints: It is difficult to accurately describe the dense correspondence of local areas such as the torso, limbs, and hem using only human keypoints. 4. Confusion of human semantic regions: Clothing areas, skin areas, hair areas, and background areas are prone to mutual contamination. 5. High training cost: Retraining or fine-tuning the diffusion model for virtual try-on tasks requires a large amount of data and computing resources, which is not conducive to rapid deployment. 6. Poor adaptability to different clothing types: Tops, trousers, skirts, dresses, and other garments have different deformation patterns, and uniform processing can easily lead to local distortions.
[0005] Therefore, there is an urgent need for a virtual fitting image generation method that does not require retraining of the diffusion model and can integrate master pose guidance, dense human body surface mapping, semantic partitioning deformation and multi-condition image restoration, so as to improve the clothing fidelity, human body pose consistency, edge naturalness and cross-scene generalization ability of the virtual fitting results. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a training-free virtual fitting image generation method and system based on master pose and dense human body surface mapping. By constructing the master pose representation of the target clothing under the reference human body pose, and combining human body key points, human body parsing information, dense human body surface mapping information and clothing category information, semantic partitioning deformation and dense alignment guidance are performed on the clothing area. Furthermore, a pre-trained diffusion-type image restoration model is used to generate an image of the target person wearing the target clothing.
[0007] To achieve the above technical solution, the technical solution adopted by the present invention is as follows: A training-free virtual fitting image generation method based on master pose and dense human body surface mapping includes the following steps: Acquire images of the target person and the target clothing; The target person image and target clothing image are preprocessed to generate multimodal conditional information, which includes human body-independent mask, clothing mask, human body pose key points, human body analytical map, human body dense surface information, clothing category information and clothing text description information. A main pose reference image of the garment is constructed based on the target garment image. The main pose reference image of the garment is used to represent the wearing appearance of the target garment under the reference human body pose. Based on the key points of human posture, human body analysis diagram and clothing category information, semantic partitioning deformation is performed on the clothing area in the main clothing posture reference image to obtain deformed clothing image and deformed clothing mask. Based on the dense human body surface information corresponding to the target person image and the dense human body surface information corresponding to the clothing main pose reference image, a dense mapping relationship between the reference human body surface and the target human body surface is constructed, and a dense human body surface guidance map is generated based on the dense mapping relationship. Input the target person image, target clothing image, human body-independent mask, clothing mask, deformed clothing image, deformed clothing mask, human body analytical map, dense human body surface guide map, and clothing text description information into a pre-trained diffusion-based image inpainting model to generate a fitting image of the target person wearing the target clothing. Among them, the pre-trained diffusion-based image inpainting model does not need to be retrained for the virtual fitting task when generating fitting images.
[0008] Furthermore, the multimodal conditional information also includes skin color information; the skin color information is calculated based on the skin semantic region in the target person image and is used to constrain the skin color consistency in the clothing replacement area during the diffusion-based image restoration process.
[0009] Preferably, the construction of the clothing master pose reference image includes at least one of the following methods: Use the image of the model wearing the target garment as the reference image for the main pose of the garment; Generate a dummy image that shows the effect of wearing the target clothing from the target clothing image; Retrieve reference images of clothing that match the target garment from a clothing image library; A reference image of the main pose of the clothing is synthesized based on the target clothing image and the preset human body pose.
[0010] Preferably, semantic partitioning deformation of the clothing area includes: Based on the clothing category information, determine whether the target clothing is a top, bottoms, skirt, or dress; Based on the human body analysis diagram, the clothing region in the main pose reference image of clothing is divided into at least one semantic sub-region; Calculate the human body proportions based on the key points of the target human body posture and the key points of the reference human body posture. Geometrically deform the semantic sub-region according to the human body proportions; The deformed semantic sub-regions are then combined to obtain the deformed clothing image.
[0011] Preferably, when the target garment is a top, the semantic sub-region includes at least two of the torso region, the left upper arm region, the left lower arm region, the right upper arm region, and the right lower arm region; The geometric deformation includes constructing a local transformation region based on at least one of the shoulder key point, elbow key point, wrist key point, neck key point or waist key point, and performing at least one of the affine transformation, perspective transformation, homography transformation or non-rigid transformation on the local transformation region. When the target garment is a bottom garment, the semantic sub-region includes at least two of the waist region, thigh region, and calf region; The geometric deformation includes constructing a local transformation region based on at least one of the waist key points, knee key points, and ankle key points, and deforming the local transformation region according to the waist-width ratio, leg-length ratio, or leg-direction difference between the target human body and the reference human body. When the target garment is a dress, the upper and lower body areas of the dress are deformed respectively, and the deformed upper and lower body areas are subjected to occlusion judgment, overlapping area cropping or fusion processing to obtain the deformed dress image.
[0012] Preferably, the dense human body surface information includes human body component identifiers and local coordinates of the human body surface; The construction of a dense mapping relationship between the reference human body surface and the target human body surface includes: Extract the reference human body surface point set and the target human body surface point set based on the human body part identifiers; For each human body part, establish a matching relationship between reference human body surface points and target human body surface points in the local coordinate space of the human body surface; A sampling grid is generated based on the matching relationship; the matching relationship is any one of nearest neighbor matching relationship, weighted nearest neighbor matching relationship, interpolation matching relationship, or matching relationship predicted based on a learning model; The main pose reference image of the clothing, the reference clothing mask, or the reference prompt image are resampled using the sampling grid to obtain a dense human body surface guide image.
[0013] Preferably, the deformable clothing mask is fused with the dense human body surface guidance map to obtain a joint posture guidance map; the joint posture guidance map is used to simultaneously constrain the geometric position of the clothing and the local correspondence of the human body surface.
[0014] Furthermore, the clothing text description information includes at least one of the following attribute information: clothing category, color, material, pattern, cut, collar type, sleeve type, length, or style.
[0015] Preferably, the pre-trained diffusion-based image inpainting model is a diffusion generation model that supports image mask conditional input; The pre-trained diffusion-based image inpainting model determines the area to be redrawn in the target person image based on the human body-independent mask, and generates the image content of the area to be redrawn based on the target clothing image, the deformed clothing image, the dense human body surface guide map, and the clothing text description information.
[0016] Preferably, during the sampling process of the pre-trained diffusion-based image restoration model, the latent variables are decomposed into pose structure information and texture detail information, and the sampling noise or sampling direction is adjusted according to the pose structure information corresponding to the main pose reference image of the clothing.
[0017] Preferably, after generating the fitting image, the method further includes: Based on a human-independent mask or a clothing-changing area mask, the virtual fitting image and the target person image are fused at the boundary to obtain the final virtual fitting image; the boundary fusion includes at least one of feathering fusion, Poisson fusion, color matching, brightness matching, or edge smoothing.
[0018] Preferably, a training-free virtual fitting image generation system based on master pose and dense human body surface mapping includes: The image acquisition module is used to acquire images of the target person and the target clothing. The condition generation module is used to preprocess the target person image and the target clothing image to generate multimodal condition information; The main pose reference construction module is used to construct a main pose reference image for the garment based on the target garment image. The semantic partitioning clothing deformation module is used to perform semantic partitioning deformation on the clothing area in the main clothing pose reference image based on human body pose key points, human body analysis diagram and clothing category information, to obtain deformed clothing image and deformed clothing mask. The dense human body surface mapping module is used to construct a dense mapping relationship between the reference human body surface and the target human body surface based on the dense human body surface information, and generate a dense human body surface guidance map. The diffusion-based image inpainting generation module is used to input the target person image, target clothing image, human body-independent mask, clothing mask, deformed clothing image, deformed clothing mask, human body analytical map, dense human body surface guide map, and clothing text description information into the pre-trained diffusion-based image inpainting model to generate a fitting image. The fusion output module is used to fuse the fitting image and the target person image to output the final virtual fitting image.
[0019] The beneficial effects of this invention are as follows: 1. This invention can directly utilize a pre-trained diffusion-based image restoration model for virtual fitting generation, reducing training and deployment costs.
[0020] 2. By using the main pose reference image of the clothing, the text description information of the clothing, and the deformed clothing image to jointly constrain the generation process, the problem of clothing texture, pattern, and structure being arbitrarily changed by the diffusion model can be reduced.
[0021] 3. By establishing the correspondence between the source clothing posture and the target human posture through key points of human posture and dense human surface mapping information, the generated clothing is more in line with the posture of the target person.
[0022] 4. By employing a semantic partitioning deformation strategy, different local transformation methods can be used for tops, bottoms, skirts, and dresses, which can reduce distortion in local areas.
[0023] 5. By using human-independent masks, human body analysis maps, skin color constraints, and post-blending processing, problems such as skin leakage, edge breakage, and background contamination can be reduced.
[0024] 6. This invention does not rely on specific paired training data and can be applied to various scenarios such as ordinary people, multi-person images, partial human body images, anime characters, or digital human images. Attached Figure Description
[0025] Figure 1 The flowchart illustrates a training-free virtual fitting image generation method based on master pose and dense human body surface mapping, as provided in an embodiment of the present invention.
[0026] Figure 2 This is a schematic diagram of the virtual fitting image generation system provided in an embodiment of the present invention. Detailed Implementation
[0027] Example 1: like Figure 1 As shown, a training-free virtual fitting image generation method based on master pose and dense human body surface mapping includes the following steps: Acquire images of the target person and the target clothing; The target person image and target clothing image are preprocessed to generate multimodal conditional information, which includes human body-independent mask, clothing mask, human body pose key points, human body analytical map, human body dense surface information, clothing category information and clothing text description information. A main pose reference image of the garment is constructed based on the target garment image. The main pose reference image of the garment is used to represent the wearing appearance of the target garment under the reference human body pose. Based on the key points of human posture, human body analysis diagram and clothing category information, semantic partitioning deformation is performed on the clothing area in the main clothing posture reference image to obtain deformed clothing image and deformed clothing mask. Based on the dense human body surface information corresponding to the target person image and the dense human body surface information corresponding to the clothing main pose reference image, a dense mapping relationship between the reference human body surface and the target human body surface is constructed, and a dense human body surface guidance map is generated based on the dense mapping relationship. Input the target person image, target clothing image, human body-independent mask, clothing mask, deformed clothing image, deformed clothing mask, human body analytical map, dense human body surface guide map, and clothing text description information into a pre-trained diffusion-based image inpainting model to generate a fitting image of the target person wearing the target clothing. Among them, the pre-trained diffusion-based image inpainting model does not need to be retrained for the virtual fitting task when generating fitting images.
[0028] Furthermore, the multimodal conditional information also includes skin color information; the skin color information is calculated based on the skin semantic region in the target person image and is used to constrain the skin color consistency in the clothing replacement area during the diffusion-based image restoration process.
[0029] Preferably, the construction of the clothing master pose reference image includes at least one of the following methods: Use the image of the model wearing the target garment as the reference image for the main pose of the garment; Generate a dummy image that shows the effect of wearing the target clothing from the target clothing image; Retrieve reference images of clothing that match the target garment from a clothing image library; A reference image of the main pose of the clothing is synthesized based on the target clothing image and the preset human body pose.
[0030] Preferably, semantic partitioning deformation of the clothing area includes: Based on the clothing category information, determine whether the target clothing is a top, bottoms, skirt, or dress; Based on the human body analysis diagram, the clothing region in the main pose reference image of clothing is divided into at least one semantic sub-region; Calculate the human body proportions based on the key points of the target human body posture and the key points of the reference human body posture. Geometrically deform the semantic sub-region according to the human body proportions; The deformed semantic sub-regions are then combined to obtain the deformed clothing image.
[0031] Preferably, when the target garment is a top, the semantic sub-region includes at least two of the torso region, the left upper arm region, the left lower arm region, the right upper arm region, and the right lower arm region; The geometric deformation includes constructing a local transformation region based on at least one of the shoulder key point, elbow key point, wrist key point, neck key point or waist key point, and performing at least one of the affine transformation, perspective transformation, homography transformation or non-rigid transformation on the local transformation region. When the target garment is a bottom garment, the semantic sub-region includes at least two of the waist region, thigh region, and calf region; The geometric deformation includes constructing a local transformation region based on at least one of the waist key points, knee key points, and ankle key points, and deforming the local transformation region according to the waist-width ratio, leg-length ratio, or leg-direction difference between the target human body and the reference human body. When the target garment is a dress, the upper and lower body areas of the dress are deformed respectively, and the deformed upper and lower body areas are subjected to occlusion judgment, overlapping area cropping or fusion processing to obtain the deformed dress image.
[0032] Preferably, the dense human body surface information includes human body component identifiers and local coordinates of the human body surface; The construction of a dense mapping relationship between the reference human body surface and the target human body surface includes: Extract the reference human body surface point set and the target human body surface point set based on the human body part identifiers; For each human body part, establish a matching relationship between reference human body surface points and target human body surface points in the local coordinate space of the human body surface; A sampling grid is generated based on the matching relationship; the matching relationship is any one of nearest neighbor matching relationship, weighted nearest neighbor matching relationship, interpolation matching relationship, or matching relationship predicted based on a learning model; The main pose reference image of the clothing, the reference clothing mask, or the reference prompt image are resampled using the sampling grid to obtain a dense human body surface guide image.
[0033] Preferably, the deformable clothing mask is fused with the dense human body surface guidance map to obtain a joint posture guidance map; the joint posture guidance map is used to simultaneously constrain the geometric position of the clothing and the local correspondence of the human body surface.
[0034] Furthermore, the clothing text description information includes at least one of the following attribute information: clothing category, color, material, pattern, cut, collar type, sleeve type, length, or style.
[0035] Preferably, the pre-trained diffusion-based image inpainting model is a diffusion generation model that supports image mask conditional input; The pre-trained diffusion-based image inpainting model determines the area to be redrawn in the target person image based on the human body-independent mask, and generates the image content of the area to be redrawn based on the target clothing image, the deformed clothing image, the dense human body surface guide map, and the clothing text description information.
[0036] Preferably, during the sampling process of the pre-trained diffusion-based image restoration model, the latent variables are decomposed into pose structure information and texture detail information, and the sampling noise or sampling direction is adjusted according to the pose structure information corresponding to the main pose reference image of the clothing.
[0037] Preferably, after generating the fitting image, the method further includes: Based on a human-independent mask or a clothing-changing area mask, the virtual fitting image and the target person image are fused at the boundary to obtain the final virtual fitting image; the boundary fusion includes at least one of feathering fusion, Poisson fusion, color matching, brightness matching, or edge smoothing.
[0038] Preferably, a training-free virtual fitting image generation system based on master pose and dense human body surface mapping includes: The image acquisition module is used to acquire images of the target person and the target clothing. The condition generation module is used to preprocess the target person image and the target clothing image to generate multimodal condition information; The main pose reference construction module is used to construct a main pose reference image for the garment based on the target garment image. The semantic partitioning clothing deformation module is used to perform semantic partitioning deformation on the clothing area in the main clothing pose reference image based on human body pose key points, human body analysis diagram and clothing category information, to obtain deformed clothing image and deformed clothing mask. The dense human body surface mapping module is used to construct a dense mapping relationship between the reference human body surface and the target human body surface based on the dense human body surface information, and generate a dense human body surface guidance map. The diffusion-based image inpainting generation module is used to input the target person image, target clothing image, human body-independent mask, clothing mask, deformed clothing image, deformed clothing mask, human body analytical map, dense human body surface guide map, and clothing text description information into the pre-trained diffusion-based image inpainting model to generate a fitting image. The fusion output module is used to fuse the fitting image and the target person image to output the final virtual fitting image.
[0039] Example 2: This invention provides a training-free virtual fitting image generation method based on master pose and dense human body surface mapping, including steps S101 to S108.
[0040] S101: Obtain the input image: Obtain target person image and target clothing images .
[0041] Among them, the target person image Includes a character to be dressed up and an image of the target clothing. This includes the clothing to be transferred to the target person. The target clothing image can be a flat lay image of the clothing, an image of a model wearing the clothing, a product display image, or an image of a cropped clothing area.
[0042] S102: Generate multimodal condition information: The target person image and target clothing image are preprocessed to generate multimodal conditional information, which includes, but is not limited to: Human body-independent mask : Used to represent the clothing area in the target character image that needs to be replaced or redrawn; clothing mask : Used to represent the clothing area in the target clothing image; target human pose key points Used to describe the skeletal structure of the target person; references key points of human posture. : Used to describe the human skeleton structure in the main pose reference image; target human body analytical image Used to distinguish semantic regions such as head, torso, arms, legs, clothing, skin, and background; refer to human anatomy diagrams. Used to distinguish semantic regions of the human body in the master pose reference image; dense surface information of the target human body. Used to describe the identification and local coordinates of dense components on the surface of a target human body; references dense surface information of the human body. Used to describe the identification and local coordinates of densely packed parts on a reference human body surface; textual description information for clothing. Used to describe the target garment's attributes such as color, material, category, pattern, sleeve length, and collar type; skin tone information. Used to ensure color consistency of skin or exposed areas of the body in the changing area; clothing category information. Used to indicate that the target garment belongs to the category of top, bottom, skirt, dress, or other clothing.
[0043] The aforementioned conditional information can be generated through image segmentation models, human body parsing models, human body pose estimation models, human body dense surface estimation models, image description models, or rule-based algorithms. This invention does not limit specific models or tools.
[0044] S103: Constructing the main pose reference image for the clothing: Construct a reference image of the main pose of the clothing based on the target clothing image. .
[0045] The garment master pose reference image is used to represent the appearance of the target garment when worn in a reference human pose. This reference image can be obtained in any of the following ways: The process involves using existing product images of models wearing the target garment as reference images; generating or compositing flat-lay garment images into human wearing images; cropping or reconstructing wearing effect images from garment product images; generating pseudo-wearing images of the target garment in the reference human pose using an image generation model; and retrieving reference wearing images that match the target garment from a historical fitting image database.
[0046] By constructing a reference image of the main posture of the garment, the subsequent garment deformation process can be based on the image of "the garment is already in the state of being worn by the human body", rather than directly performing complex non-rigid deformation on the flat garment image, thereby improving the fidelity of the garment texture and structure.
[0047] S104: Clothing Deformation Based on Semantic Partitioning Based on clothing category information Key points of target human posture Reference human posture key points Target human body analysis diagram and reference human anatomy diagram The clothing region in the main pose reference image of the clothing is semantically partitioned and deformed to obtain a deformed clothing image. and deformable clothing mask .
[0048] 1. Deformation of the upper garment: When the clothing category is top, the clothing area should be divided into at least the torso area, left upper arm area, left lower arm area, right upper arm area, and right lower arm area.
[0049] Based on the reference human body posture key points, calculate the reference shoulder width, torso length, arm length and other body proportions. Based on the target human body posture key points, calculate the target shoulder width, target torso length and target arm length, and adjust the reference clothing skeleton according to the proportional relationship between the two.
[0050] For the torso region, the neck, left and right shoulder points, left and right waist points, or hip points can be selected as control points. The geometric transformation matrix from the reference posture to the target posture is calculated to deform the torso region.
[0051] For the sleeve area, a local quadrilateral area can be constructed based on the shoulder point, elbow point, and wrist point, and local perspective transformation or affine transformation can be performed on each sleeve area.
[0052] The deformed torso and sleeve areas are cut and merged according to the target human body analysis diagram to obtain the deformed upper garment.
[0053] 2. Deformation of the lower garment: When the clothing category is bottoms, the clothing area should be divided into at least the waist area, left thigh area, left calf area, right thigh area, and right calf area.
[0054] Based on the waist width, thigh length, and calf length ratios of the target and reference human bodies, the proportions of the reference garment skeleton are corrected.
[0055] For trousers, local geometric transformations are performed on the left and right leg areas to ensure that the trouser leg area is consistent with the direction and length of the target human leg.
[0056] For skirts, overall or partial deformation areas can be constructed based on key points at the waist, hips, and hem to maintain a relatively continuous skirt structure and avoid incorrectly splitting the skirt into a leg structure.
[0057] 3. Dress deformation: When the clothing category is a dress, the dress is split into an upper body clothing area and a lower body clothing area. The upper body area is processed according to the upper garment deformation strategy; the lower body area is processed according to the skirt or bottom garment deformation strategy; the overlapping areas of the two are occluded and merged; and the complete dress deformation result is obtained.
[0058] By using the above-mentioned classification and deformation methods, the present invention can adopt different transformation strategies for different types of clothing to reduce local misalignment and unnatural stretching of clothing.
[0059] S105: Constructing dense human body surface mapping relationships: Based on the dense surface information of the target human body and reference to dense human body surface information Construct a dense mapping relationship from the reference human body surface to the target human body surface.
[0060] In one embodiment, human body dense surface information includes human body part identification. and local two-dimensional surface coordinates For each human body part, extract the set of pixels belonging to that part from the reference human body surface and the set of pixels belonging to that part from the target human body surface.
[0061] For the For individual body parts, let the set of reference points on the human body surface be: The target set of points on the human body surface is: For each point on the target human body surface. Searching for points on the reference human body surface The nearest reference point in space: Construct a sampling grid based on the correspondence. The sampling grid is then used to resample the reference clothing area, reference clothing mask, or reference human body surface cue map to obtain a dense human body surface guidance map. .
[0062] The dense human body surface guidance map is used to supplement the local surface correspondences that cannot be expressed by sparse human body key points, thereby improving the clothing alignment accuracy in scenarios with large pose changes and partial occlusion.
[0063] S106: Input for constructing a diffusion-based image inpainting model: Input the following conditions into the pre-trained diffusion image inpainting model: target person image Target clothing image Human body-independent mask Clothing mask Deformation clothing images Deformation clothing mask Target human body analysis diagram Dense human body surface guidance map Skin color information Clothing text description information .
[0064] Among them, the human-independent mask is used to indicate the areas of the model that need to be redrawn; the target clothing image and clothing text description information are used to constrain the appearance of the target clothing; the deformed clothing image is used to provide cues for the clothing structure after geometric alignment; the human body analytical map is used to constrain the human body semantic region; the dense human body surface guide map is used to provide local human body surface alignment constraints; and the skin color information is used to maintain the consistency of skin region color.
[0065] S107: Training-free diffusion-based image inpainting generation: A pre-trained diffusion-based image inpainting model is used to generate an initial fitting image by processing the clothing area of the target person's image. .
[0066] The pre-trained diffusion-based image inpainting model can be a diffusion-based image inpainting model, which does not require retraining or fine-tuning for the virtual try-on task when executing the method of this invention. The model can perform conditional image generation based on text conditions, image conditions, mask conditions, pose conditions, or other auxiliary conditions.
[0067] In an optional embodiment, during the diffusion sampling process, the pose structure information and texture detail information in the latent variables can be further decomposed, and the sampling noise or sampling direction can be adjusted using the master pose reference information, so that the generated result better preserves the pose of the target person and the texture of the target clothing. This step is an optional enhancement step and does not limit the basic implementation of the present invention.
[0068] S108: Image Fusion and Output For the initial fitting image and the original target person image The final virtual fitting image is obtained by performing a fusion process. .
[0069] The blending process can include: region blending based on human-independent masks; edge feathering blending; Poisson blending; color consistency correction; brightness and contrast matching; and boundary smoothing between generated and non-generated regions.
[0070] By using fusion processing, issues such as abrupt changes in the edges of the clothing change area, color inconsistencies, and background contamination can be reduced, resulting in a more natural final image.
[0071] Example 3: like Figure 2 As shown, the present invention also provides a training-free virtual fitting image generation system based on master pose and dense human body surface mapping, comprising: The input layer is used to receive images of the target person, images of the target clothing, user parameters, and clothing category information, forming an input data stream.
[0072] The processing layer includes a condition generation module, a main pose reference construction module, a semantic partitioning clothing deformation module, a dense human body surface mapping module, and a multi-condition fusion module. The condition generation module calls the pose estimation, human body parsing, dense surface estimation and image description model of the model resource layer to generate human body-independent mask, clothing mask, human pose key points, human body parsing map, human body dense surface information, skin color information and clothing text description information. The main pose reference construction module constructs a main pose reference image for the clothing based on the target clothing image. The semantic partitioning clothing deformation module performs semantic partitioning deformation on the main posture reference image of clothing based on clothing category information, key points of human posture, and human body analysis diagram, generating deformed clothing images and deformed clothing masks. The dense human body surface mapping module establishes a dense mapping relationship and generates a dense human body surface guidance map based on the dense surface information of the target human body and the dense surface information of the reference human body. The multi-condition fusion module fuses deformable clothing images, deformable clothing masks, dense human body surface guide maps, and other multimodal conditions to generate the multi-condition generation flow required for the diffusion model.
[0073] The output layer is generated, including a diffusion-based image inpainting generation module, an image fusion module, and a result output module; The diffusion-based image restoration generation module calls the pre-trained diffusion-based image restoration model in the model resource layer to generate an initial fitting image based on the target person image, the target clothing image, and a multi-condition generation flow. The image fusion module performs edge blending processing on the initial fitting image and the original target person image; The results output module outputs the final virtual try-on image.
[0074] The model resource layer provides algorithmic support for each module of the system, including pose estimation model, human body analysis model, dense surface estimation model, image description model, and pre-trained diffusion image inpainting model. Each processing module calls the above model resources through the dotted lines.
[0075] The present invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the training-free virtual fitting image generation method based on master pose and dense human body surface mapping described in any of the above embodiments.
[0076] The electronic device may be a server, cloud computing device, personal computer, workstation, mobile terminal, edge computing device or embedded computing device.
[0077] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the training-free virtual fitting image generation method based on master pose and dense human body surface mapping as described in any of the above embodiments.
[0078] The computer-readable storage medium may include a read-only memory, a random access memory, a disk, an optical disk, a solid-state drive, flash memory, or other media capable of storing program code.
Claims
1. A training-free virtual fitting image generation method based on master pose and dense human body surface mapping, characterized in that, Includes the following steps: Acquire images of the target person and the target clothing; The target person image and target clothing image are preprocessed to generate multimodal conditional information, which includes human body-independent mask, clothing mask, human body pose key points, human body analytical map, human body dense surface information, clothing category information and clothing text description information. A main pose reference image of the garment is constructed based on the target garment image. The main pose reference image of the garment is used to represent the wearing appearance of the target garment under the reference human body pose. Based on the key points of human posture, human body analysis diagram and clothing category information, semantic partitioning deformation is performed on the clothing area in the main clothing posture reference image to obtain deformed clothing image and deformed clothing mask. Based on the dense human body surface information corresponding to the target person image and the dense human body surface information corresponding to the clothing main pose reference image, a dense mapping relationship between the reference human body surface and the target human body surface is constructed, and a dense human body surface guidance map is generated based on the dense mapping relationship. Input the target person image, target clothing image, human body-independent mask, clothing mask, deformed clothing image, deformed clothing mask, human body analytical map, dense human body surface guide map, and clothing text description information into a pre-trained diffusion-based image inpainting model to generate a fitting image of the target person wearing the target clothing. Among them, the pre-trained diffusion-based image inpainting model does not need to be retrained for the virtual fitting task when generating fitting images.
2. The training-free virtual fitting image generation method based on master pose and dense human body surface mapping according to claim 1, characterized in that, The construction of the clothing master pose reference image includes at least one of the following methods: Use the image of the model wearing the target garment as the reference image for the main pose of the garment; Generate a dummy image that shows the effect of wearing the target clothing from the target clothing image; Retrieve reference images of clothing that match the target garment from a clothing image library; A reference image of the main pose of the clothing is synthesized based on the target clothing image and the preset human body pose.
3. The training-free virtual fitting image generation method based on master pose and dense human body surface mapping according to claim 1, characterized in that, The semantic partitioning and deformation of the clothing area includes: Based on the clothing category information, determine whether the target clothing is a top, bottoms, skirt, or dress; Based on the human body analysis diagram, the clothing region in the main pose reference image of clothing is divided into at least one semantic sub-region; Calculate the human body proportions based on the key points of the target human body posture and the key points of the reference human body posture. Geometrically deform the semantic sub-region according to the human body proportions; The deformed semantic sub-regions are then combined to obtain the deformed clothing image.
4. The training-free virtual fitting image generation method based on master pose and dense human body surface mapping according to claim 3, characterized in that, When the target garment is a top, the semantic sub-region includes at least two of the following regions: torso region, left upper arm region, left lower arm region, right upper arm region, and right lower arm region. The geometric deformation includes constructing a local transformation region based on at least one of the shoulder key point, elbow key point, wrist key point, neck key point or waist key point, and performing at least one of the affine transformation, perspective transformation, homography transformation or non-rigid transformation on the local transformation region. When the target garment is a bottom garment, the semantic sub-region includes at least two of the waist region, thigh region, and calf region; The geometric deformation includes constructing a local transformation region based on at least one of the waist key points, knee key points, and ankle key points, and deforming the local transformation region according to the waist-width ratio, leg-length ratio, or leg-direction difference between the target human body and the reference human body. When the target garment is a dress, the upper and lower body areas of the dress are deformed respectively, and the deformed upper and lower body areas are subjected to occlusion judgment, overlapping area cropping or fusion processing to obtain the deformed dress image.
5. The training-free virtual fitting image generation method based on master pose and dense human body surface mapping according to claim 1, characterized in that, The dense surface information of the human body includes human body component identifiers and local coordinates of the human body surface; The construction of a dense mapping relationship between the reference human body surface and the target human body surface includes: Extract the reference human body surface point set and the target human body surface point set based on the human body part identifiers; For each human body part, establish a matching relationship between reference human body surface points and target human body surface points in the local coordinate space of the human body surface; A sampling grid is generated based on the matching relationship; the matching relationship is any one of nearest neighbor matching relationship, weighted nearest neighbor matching relationship, interpolation matching relationship, or matching relationship predicted based on a learning model; The main pose reference image of the clothing, the reference clothing mask, or the reference prompt image are resampled using the sampling grid to obtain a dense human body surface guide image.
6. The training-free virtual fitting image generation method based on master pose and dense human body surface mapping according to claim 1, characterized in that, The deformable clothing mask is fused with the dense human body surface guidance map to obtain a joint posture guidance map; the joint posture guidance map is used to simultaneously constrain the geometric position of the clothing and the local correspondence of the human body surface.
7. The training-free virtual fitting image generation method based on master pose and dense human body surface mapping according to claim 1, characterized in that, The pre-trained diffusion-based image inpainting model is a diffusion generation model that supports image mask conditional input; The pre-trained diffusion-based image inpainting model determines the area to be redrawn in the target person image based on the human body-independent mask, and generates the image content of the area to be redrawn based on the target clothing image, the deformed clothing image, the dense human body surface guide map, and the clothing text description information.
8. The training-free virtual fitting image generation method based on master pose and dense human body surface mapping according to claim 1, characterized in that, During the sampling process of the pre-trained diffusion-based image restoration model, latent variables are decomposed into pose structure information and texture detail information, and the sampling noise or sampling direction is adjusted according to the pose structure information corresponding to the main pose reference image of the clothing.
9. The training-free virtual fitting image generation method based on master pose and dense human body surface mapping according to claim 1, characterized in that, After generating the fitting image, the process also includes: Based on a human-independent mask or a clothing-changing area mask, the virtual fitting image and the target person image are fused at the boundary to obtain the final virtual fitting image; the boundary fusion includes at least one of feathering fusion, Poisson fusion, color matching, brightness matching, or edge smoothing.
10. A training-free virtual fitting image generation system based on master pose and dense human body surface mapping, characterized in that, include: The image acquisition module is used to acquire images of the target person and the target clothing. The condition generation module is used to preprocess the target person image and the target clothing image to generate multimodal condition information; The main pose reference construction module is used to construct a main pose reference image for the garment based on the target garment image. The semantic partitioning clothing deformation module is used to perform semantic partitioning deformation on the clothing area in the main clothing pose reference image based on human body pose key points, human body analysis diagram and clothing category information, to obtain deformed clothing image and deformed clothing mask. The dense human body surface mapping module is used to construct a dense mapping relationship between the reference human body surface and the target human body surface based on the dense human body surface information, and generate a dense human body surface guidance map. The diffusion-based image inpainting generation module is used to input the target person image, target clothing image, human body-independent mask, clothing mask, deformed clothing image, deformed clothing mask, human body analytical map, dense human body surface guide map, and clothing text description information into the pre-trained diffusion-based image inpainting model to generate a fitting image. The fusion output module is used to fuse the fitting image and the target person image to output the final virtual fitting image.