A method, system, device and medium for generating AI model images

By decoupling clothing and pose in an AI model image generation method, and utilizing multimodal data and pose skeleton graphs, a LoRA diffusion model with added pose condition injection channels is used to generate natural and rich model poses and realistic clothing details. This solves the problem of stiff poses in existing technologies and improves the aesthetics and realism of the images.

CN121685752BActive Publication Date: 2026-05-26HANGZHOU LINRUN INTELLIGENT TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU LINRUN INTELLIGENT TECHNOLOGY CO LTD
Filing Date
2026-02-09
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In existing AI model image generation technologies, the binding of clothing to mannequin poses results in stiff and unnatural poses in the generated models, affecting the aesthetics of the images.

Method used

By collecting multimodal data, selecting a reference image that matches the desired pose and extracting the pose skeleton image, adding a LoRA diffusion model with an independent pose condition injection channel, decoupling clothing from pose, and combining clothing deformation network to generate clothing images that match the target pose.

Benefits of technology

It achieves a greater naturalness and richness in model poses, a high degree of fidelity in clothing style and material details, and a significant improvement in the aesthetics and realism of the generated images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121685752B_ABST
    Figure CN121685752B_ABST
Patent Text Reader

Abstract

This application relates to an AI model image generation method, system, device, and medium, belonging to the technical field of image processing. It involves collecting multimodal data of the target garment, including mannequin images, flat static images, and material close-up images; selecting a reference image from a target pose library whose pose matches the desired pose, and extracting the corresponding pose skeleton image as the target pose skeleton image; inputting the flat static image and the material close-up image into a visual language model to generate a text description; and inputting the target pose skeleton image, mannequin image, and text description into a pre-trained LoRA diffusion model to generate a preliminary dressed image. The LoRA diffusion model's input layer includes an independent pose condition injection channel, through which the target pose skeleton image is input into the LoRA diffusion model. This application has the beneficial technical effect of avoiding stiff and unnatural AI model poses, thereby improving the aesthetics of the generated images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to an AI model image generation method, system, device and medium. Background Technology

[0002] In recent years, with the development of deep learning and generative artificial intelligence (AI), generating high-quality images using diffusion models has become a popular technology. In the apparel e-commerce sector, using AI to generate virtual model images can significantly reduce shooting costs and shorten the new product launch cycle.

[0003] One mainstream technical approach is to use Low-Rank Adaptation (LoRA) fine-tuning. This method involves injecting a small number of trainable parameters into a pre-trained diffusion model to customize the learning for a specific concept (such as a set of clothing). In practice, developers take a set (e.g., 10-20) photos of each outfit worn by a mannequin and use these photos to train a dedicated LoRA model for that outfit.

[0004] Because mannequin poses are very fixed and rigid, they can usually only perform simple rotations and cannot strike rich, natural, and graceful poses like real models. Furthermore, due to the limited variety of poses in the training data, the LoRA model inevitably forcibly associates (couples) the clothing with the rigid mannequin poses while learning clothing features. This results in the generated AI model poses being severely similar to the mannequin poses used in training, even when given descriptive words describing graceful poses, appearing stiff and unnatural, seriously affecting the aesthetics of the generated images. Summary of the Invention

[0005] In order to avoid stiff and unnatural poses of AI models and improve the aesthetics of generated images, this application provides an AI model image generation method, system, device and medium.

[0006] Firstly, this application provides an AI model image generation method, which adopts the following technical solution:

[0007] An AI-powered model image generation method includes:

[0008] Collect multimodal data of the target garment, including mannequin images, static flat lay images, and close-up images of the material.

[0009] Select a reference image whose pose matches the desired pose from the target pose library, and extract the pose skeleton image corresponding to the reference image as the target pose skeleton image;

[0010] The tiling static image and the material close-up image are input into the visual language model to generate a text description;

[0011] The target pose skeleton diagram, mannequin image, and text description are input into the pre-trained LoRA diffusion model to generate a preliminary dressed image. An independent pose condition injection channel is added to the input layer of the LoRA diffusion model, and the target pose skeleton diagram is input into the LoRA diffusion model through the pose condition injection channel.

[0012] By adopting the above technical solution, a reference image with the desired posture is obtained and the target posture skeleton image is extracted. An independent posture condition injection channel is added to the input layer of the LoRA diffusion model to input the target posture skeleton image into the model, thereby decoupling the posture and clothing. This allows the generated preliminary clothing image to present the desired posture, avoiding the stiffness and unnaturalness of the AI ​​model's posture and improving the aesthetics of the generated image.

[0013] Optionally, the training method for the LoRA diffusion model includes:

[0014] Collect multi-angle real-life photos of mannequins wearing the relevant clothing;

[0015] Each captured mannequin image is processed to extract a skeleton diagram containing pose information;

[0016] For each mannequin image, a text description is matched and paired with the corresponding skeleton image to form training samples;

[0017] A pre-trained LoRA model is used as the base model. An independent attitude condition injection channel is added to the input layer to form a LoRA diffusion model. The parameters of the base model and the attitude encoder weights are frozen.

[0018] The training samples are input into the LoRA diffusion model to train the parameters of the LoRA adaptation layer and the projection layer output by the attitude encoder.

[0019] Calculate the difference between the noise predicted by the LoRA diffusion model and the actual noise, and update the LoRA parameters based on this difference through backpropagation.

[0020] By employing the aforementioned technical solution, the LoRA diffusion model is explicitly informed that the pose information in an image is defined by the input "pose skeleton map." This allows the attention layer of the diffusion model to attribute the learning of pose features to skeleton map conditions, while enabling the LoRA network to focus more on learning visual features beyond pose, such as core details like clothing style, color, texture, and folds. Because pose is stripped away as a known prior, the LoRA learning task becomes purer and more efficient, potentially improving convergence speed.

[0021] Optionally, the steps for extracting the skeleton map containing pose information include:

[0022] For each captured mannequin image, perform 2D human pose estimation to obtain the coordinates of key human points, and connect them to generate the corresponding initial pose skeleton diagram.

[0023] The initial posture skeleton diagram is input into the posture repair network, which optimizes and completes the key points based on human kinematic priors to generate the repaired skeleton diagram.

[0024] By adopting the above technical solution, the prior knowledge of human kinematics includes knowledge such as the reasonable range of human joint movement and interrelationships. Using this knowledge, unreasonable key point positions that may exist in the initial posture skeleton diagram can be corrected, so that the generated repaired skeleton diagram is more in line with the actual movement and posture of the human body, and the accuracy of posture information can be improved.

[0025] Optionally, the steps for building the target pose library include:

[0026] Collect diverse real-life pose reference images from the internet, fashion image libraries, and professional photography materials;

[0027] A consistent pose estimation algorithm, used in the same phase as the training phase, is used to extract skeleton maps corresponding to real-person pose reference images. These skeleton maps are then labeled and categorized according to action type to establish a searchable target pose database.

[0028] By adopting the above technical solutions, a rich and diverse target pose library, which is classified and organized, provides more different types of pose data for model training.

[0029] Optionally, the steps following the generation of the initial dressed image include:

[0030] The mannequin image is semantically segmented using a clothing analysis model to generate a pixel-aligned clothing region segmentation mask.

[0031] Based on the clothing region segmentation mask and the text description, the physical attribute semantic tags of the clothing are labeled to form a comprehensive clothing representation that includes appearance, structure and physical characteristics.

[0032] The integrated clothing representation, the target pose skeleton diagram, and the preliminary clothing image are input into the pre-trained Clothing Deformation Network (GDN) to generate a deformation field.

[0033] Based on the deformation field, pixel-level deformation distortion is performed on the clothing area in the preliminary clothing image to output the final clothing display image;

[0034] The physical attribute semantic tags are embedded as prior knowledge into the Garment Deformation Network (GDN) to guide the GDN in generating differentiated deformation patterns based on different material responses.

[0035] The physical attribute semantic tags are automatically labeled through visual understanding of the large model, including material semantic tags and pattern semantic tags;

[0036] Among them, the material semantic tag is used to describe the optical and mechanical properties of the fabric; the pattern semantic tag is used to describe the overall silhouette and cutting features of the garment.

[0037] By adopting the above technical solutions, different fabrics and patterns will exhibit different deformation behaviors when subjected to external forces. This method can simulate more realistic clothing deformation effects, making the final clothing display image more lifelike. It allows the clothing to deform naturally according to the target posture, better conforming to the human body, improving the realism and aesthetics of the clothing display. The generated final clothing display image can more realistically reflect the appearance and physical characteristics of the clothing in different postures, providing users with a more intuitive and accurate clothing display.

[0038] Optionally, the training method for the garment deformation network (GDN) includes:

[0039] Collect images of the test mannequin wearing the relevant clothing;

[0040] The test mannequin image is segmented into regions, and a clothing segmentation mask is defined and the physical properties of the clothing are labeled to form a comprehensive test clothing representation.

[0041] Extract the skeleton image of the test mannequin in the current training pose;

[0042] The skeleton diagrams under the standard posture, the skeleton diagrams under the current training posture, and the test integrated clothing representation are input into the GDN to predict the deformation field from the standard posture to the current posture.

[0043] The deformation field is applied to the image of the test mannequin, and a reconstructed image is obtained through a differentiable image deformation operation. The deformation operation is performed only on the clothing area defined by the clothing segmentation mask.

[0044] Calculate the pixel-level difference loss between the reconstructed image and the original real image, and update the GDN parameters based on the difference loss through backpropagation.

[0045] By adopting the above technical solution, the accurately trained GDN can generate a high-quality deformation field. In the process of generating AI model images, applying such a deformation field to the clothing area can make the clothing fit the model's posture more naturally, avoiding mismatches and unnaturalness between clothing and posture, thereby improving the quality and aesthetics of the final generated AI model images and providing users with a better visual experience.

[0046] Optionally, the AI ​​model image generation method includes:

[0047] Extract ambient lighting information from the original reference image to construct a lighting map;

[0048] The target pose skeleton map and the clothing region segmentation mask are input into a pre-trained shadow generation model to predict the shadow distribution of the clothing in the target pose.

[0049] The lighting map and shadow distribution are fused into the final clothing display image to generate a lighting rendering image.

[0050] By employing the aforementioned technical solution and adding lighting maps, the generated model images can present lighting effects that more closely resemble real-world scenes, enhancing the realism and immersiveness of the images. Furthermore, the appropriate presentation of light and shadow helps to better showcase the material, color, and style of the clothing.

[0051] Secondly, this application provides an AI model image generation system, which adopts the following technical solution:

[0052] An AI model image generation system includes:

[0053] The data acquisition module is used to collect multimodal data of the target garment, including mannequin images, flat static images, and close-up images of the material.

[0054] The skeleton extraction module is used to extract the corresponding posture skeleton map as the target posture skeleton map after selecting a reference map whose posture matches the expectation from the target posture library.

[0055] The multimodal semantic parsing module is used to input the tiled static image and the material close-up image into the visual language model to generate a text description;

[0056] The model processing module is used to input the target pose skeleton diagram, mannequin diagram and text description into the pre-trained LoRA diffusion model to generate a preliminary dressed image; the input layer of the LoRA diffusion model is equipped with an independent pose condition injection channel, and the target pose skeleton diagram is input into the LoRA diffusion model through the pose condition injection channel.

[0057] Thirdly, this application provides a computer device that adopts the following technical solution:

[0058] A computer device includes a memory, a processor, and a computer program stored in the memory, the processor executing the computer program to implement the AI ​​model image generation method as described in the first aspect.

[0059] Fourthly, this application provides a computer-readable storage medium, which adopts the following technical solution:

[0060] A computer-readable storage medium storing a computer program capable of being loaded by a processor and executing the AI ​​model image generation method as described in the first aspect.

[0061] In summary, this application includes at least one of the following beneficial technical effects:

[0062] 1. It achieves effective decoupling of clothing and pose, fundamentally solving the problem of binding clothing features with mannequin pose; allowing the LoRA model to focus on the clothing itself, while pose becomes an independent control variable that can be freely replaced.

[0063] 2. Greatly enhances the richness and aesthetics of poses: Users can provide reference images of any pose to control the poses of the generated model. Whether it is an elegant standing posture, a dynamic walking posture or a complex dance movement, it can be easily achieved, completely getting rid of the stiffness of the mannequin.

[0064] 3. Ensures high fidelity in clothing reproduction: Because the LoRA model has a more focused learning objective, the style, pattern, and material details of the clothing are fully learned and preserved.

[0065] 4. By introducing material and pattern semantic tags as priors, the Garment Deformation Network (GDN) is guided to generate differentiated deformations that conform to material characteristics, significantly improving the realism and artistic expression of the generated results. Attached Figure Description

[0066] Figure 1 This is a first flowchart of an embodiment of the method of this application;

[0067] Figure 2 This is a second flowchart of an embodiment of the method of this application;

[0068] Figure 3 This is a third flowchart of an embodiment of the method of this application;

[0069] Figure 4 This is the fourth flowchart of an embodiment of the method of this application;

[0070] Figure 5 This is the fifth flowchart of an embodiment of the method of this application. Detailed Implementation

[0071] To make the purpose, technical solution, and advantages of this application clearer, the following description is provided in conjunction with the appendix. Figure 1-5 The present application will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the application.

[0072] The first embodiment of this application discloses an AI model image generation method. (Refer to...) Figure 1 The AI ​​model image generation method may include steps S110-S140:

[0073] S110 collects multimodal data of the target garment, including mannequin images, flat static images, and close-up images of the material.

[0074] S120: Select a reference image whose posture matches the desired posture from the target posture library, and extract the corresponding posture skeleton image as the target posture skeleton image.

[0075] S130: Input the tiled static image and material close-up image into the visual language model to generate a text description;

[0076] S140: Input the target pose skeleton map, mannequin image, and text description into the pre-trained LoRA diffusion model to generate a preliminary dressed image; The input layer of the LoRA diffusion model adds an independent pose condition injection channel, and the target pose skeleton map is input into the LoRA diffusion model through the pose condition injection channel.

[0077] Specifically, in the S110 stage of acquiring multimodal data of the target garment, it is first necessary to systematically acquire mannequin images, flat lay static images, and material close-up images of the target garment. Mannequin images are taken in a professional photography studio to ensure uniform lighting and a clean background. A 360-degree rotating shooting stand can be used to obtain mannequin images from multiple angles to fully display the three-dimensional structure of the garment. Flat lay static images require the garment to be laid flat and photographed from directly above using a high-definition camera to ensure accurate proportions of each part of the garment and to prevent wrinkles from interfering with key design details. Material close-up images use a macro lens to focus on the texture, weave, luster, and detailed decorations of the garment fabric (such as buttons, zippers, embroidery, etc.). During shooting, the lighting angle can be adjusted appropriately to highlight the differences in material texture. All images must have a uniform resolution of at least 2048×2732 pixels to ensure the accuracy of subsequent processing.

[0078] In S120, the steps for constructing the target pose library include S121-S122:

[0079] S121 collects diverse real-life pose reference images from the Internet, fashion image libraries, and professional photography materials;

[0080] S122, using the same pose estimation algorithm as in the training phase, extract the skeleton map corresponding to the real person pose reference image, and label and classify it according to the action category to establish a searchable target pose database.

[0081] Specifically, firstly, diverse real-life pose reference images were collected from fashion blogger accounts (such as Pinterest and Instagram), fashion trend platforms like WGSN, and professional stock image websites like Shutterstock using web crawling technology. This ensured coverage of different action categories, including standing, sitting, walking, and dynamic displays, while also taking into account the pose diversity of models of different body types and ages. High-definition, unobstructed, and high-quality images were selected while avoiding copyright risks. Subsequently, the collected real-life images were processed using the OpenPose pose estimation algorithm, which was completely consistent with the subsequent training phase. The two-dimensional coordinates of 18 key nodes (including head, neck, and limb joints) were extracted. The Bresenham algorithm was used to connect the key points to generate a standardized skeleton diagram. The skeleton lines were marked with solid red lines, and the key points were marked with solid green circles. Then, the images were manually labeled according to action categories (such as "casual standing pose," "runway walk," and "business sitting pose"). Cluster analysis was performed based on the key point coordinate features of the skeleton diagram to optimize the classification system. Finally, a MySQL database was used to store image paths, skeleton diagram data, and action labels, constructing a target pose database that supports keyword retrieval and pose similarity matching.

[0082] When generating text descriptions, the S130 inputs a tiled static image and a close-up image of the material into the CLIP visual language model. First, the images are preprocessed (including normalization and resizing to 224×224 pixels). The model's visual encoder extracts image features, which are then matched cross-modally with the lexical features in the text encoder to generate a preliminary description. Subsequently, the description is manually optimized by adding key information such as clothing style (e.g., A-line skirt, suit jacket), collar type (e.g., lapel, round neck), sleeve type (e.g., bell sleeve, three-quarter sleeve), color (e.g., "smog blue", "burgundy"), and pattern (e.g., "stripes", "polka dots", "checkered") to form a structured text description that includes appearance, structure, and material. For example, "Dark blue denim jacket, lapel design, double pockets on the front, metal buttons, long sleeves, 100% cotton fabric with medium thickness and slight elasticity."

[0083] When generating the initial dressed image using S140, the LoRA diffusion model input needs to integrate the target pose skeleton map, mannequin image, and text description. An independent pose condition injection channel is added to the input layer of the diffusion model. A VAE model is used to compress the skeleton map features to the same dimension as the image features, and then fuse them with the image features through a spatial attention mechanism. The target pose skeleton map is preprocessed (converted to a grayscale image of the same size as the input image, with the skeleton line pixel value set to 255 and the background set to 0) and then injected into the model through this channel. The text description is converted into a 768-dimensional feature vector by the CLIP text encoder and then fused with the pose features and image features through cross-attention at different levels of the LoRA diffusion model. During model inference, the sampling steps are set to 50, the DDIM sampler is used, and the CFG value is set to 7.5, generating an initial dressed image with a resolution of 1024×1365. This image already possesses the basic fusion effect of the target pose and clothing.

[0084] Pose estimation algorithms should be of a fixed version (e.g., OpenPose v1.7.0) and run on the same hardware environment (e.g., NVIDIA Tesla V100 GPU) to avoid errors in skeleton map extraction caused by algorithm differences. Classification and labeling can be combined with a crowdsourcing platform for manual verification to improve label accuracy.

[0085] Reference Figure 2 The training methods for the LoRA diffusion model include S210-S260:

[0086] S210, collect multi-angle real-shot images of mannequins for the relevant clothing;

[0087] S220 processes each captured mannequin image to extract a skeleton diagram containing posture information;

[0088] S230: Match text descriptions for each mannequin image and pair them with the corresponding skeleton diagram to form training samples;

[0089] S240 uses a pre-trained LoRA model as the base model, adds an independent attitude condition injection channel to the input layer to form a LoRA diffusion model, and freezes the base model parameters and attitude encoder weights.

[0090] S250, input the training samples into the LoRA diffusion model to train the parameters of the LoRA adaptation layer and the projection layer output by the attitude encoder;

[0091] S260, calculate the difference between the noise predicted by the LoRA diffusion model and the actual noise, and update the LoRA parameters based on the difference through backpropagation.

[0092] In S220, the steps for extracting the skeleton map containing pose information include S221-S222:

[0093] S221, Perform 2D human pose estimation on each acquired mannequin image to obtain the coordinates of key human points and connect them to generate the corresponding initial pose skeleton diagram.

[0094] S222, the initial pose skeleton map is input into the pose repair network. The pose repair network optimizes and completes the key points based on human kinematics priors to generate the repaired skeleton map.

[0095] Specifically, the training phase of the LoRA diffusion model requires strict control over data quality and training parameters. In S210, when collecting multi-angle real-world images of the mannequin, several to dozens of images of each garment are selected, ensuring consistency in lighting and background from multiple angles. In S220, the same pose estimation algorithm as S122 is used to extract key points and generate an initial pose skeleton map. A pose restoration network based on the diffusion model is constructed, with the input being a grayscale image of the skeleton map (size adjusted to 256x256). The network design includes an encoder-decoder structure, where the encoder uses ResNet-18 to extract features, and the decoder combines human kinematic priors (such as joint angle constraints and limb length ratios) with a custom loss function (e.g., adding a kinematic regularization term to penalize unreasonable angle deviations) to optimize key point positions. The model is implemented in PyTorch, with pre-trained weights from the MPII Human Pose dataset. The output is a restored skeleton map (e.g., filling in occluded points or correcting distorted poses), which is saved as a high-resolution image to ensure skeleton continuity.

[0096] For each mannequin image, a detailed text description is matched and associated with a repaired skeleton diagram. The text description covers information such as clothing category (e.g., dress, blazer), style features (e.g., A-line skirt, lapel design), color (e.g., bright red, navy blue), material (e.g., silk, denim), and details (e.g., metal buttons, embroidery patterns). A structured description template (e.g., "A [color][material][category] item with [style features] and [details]") is used to ensure a high degree of matching between the text and image content. Subsequently, the mannequin image, corresponding skeleton diagram, and text description are packaged into training sample pairs and divided into training and validation sets in an 8:2 ratio.

[0097] The LoRA model architecture is modified by adding an independent pose condition injection channel to the input layer. This channel uses a convolutional neural network (CNN) to encode the pose skeleton map into a conditional vector with the same dimension as the image features, and then fuses it with the low-level visual features of the diffusion model through a spatial attention mechanism. Simultaneously, a pre-trained LoRA parameter module is loaded, the basic model network parameters are frozen, and only the parameters related to the LoRA adaptation layer and the newly added pose condition injection channel are retained as trainable.

[0098] During training, training set samples are input into the model in batches. The skeleton map is fed in through an independent pose condition injection channel, the mannequin image serves as the primary visual input, and the text description is converted into text embedding vectors by the CLIP text encoder. These three inputs together guide the model to learn the generative mapping from the mannequin image to the clothing image. Training employs the AdamW optimizer with an initial learning rate of 2e-4, dynamically adjusted using a cosine annealing strategy. The batch size is set to 8-16 based on the GPU memory. The total number of training epochs is determined based on the sample size (typically 100-300 epochs). Every 10 epochs, the pose consistency and clothing detail retention of the generated images are evaluated on the validation set. The mean squared error (MSE) between the model's predicted noise and the actual Gaussian noise is used as the primary loss function. Perceptual loss is also introduced to improve the visual quality of the generated images. The parameters of the LoRA adaptation layer and the projection layer of the pose condition injection channel are updated based on the loss value through backpropagation until the images generated by the model on the validation set achieve a balance between pose accuracy and clothing reproduction.

[0099] Reference Figure 3 The steps following the generation of the initial clothing image include S310-S340:

[0100] S310 uses a clothing analysis model to perform semantic segmentation on the mannequin image and generates a pixel-aligned clothing region segmentation mask.

[0101] S320, based on the clothing region segmentation mask and text description, marks the physical attribute semantic tags of the clothing to form a comprehensive clothing representation that includes appearance, structure and physical characteristics;

[0102] S330 inputs the integrated clothing representation, preliminary dressing image and target pose skeleton map into the pre-trained clothing deformation network GDN to generate the deformation field;

[0103] S340, perform pixel-level deformation distortion on the clothing area in the preliminary clothing image according to the deformation field, and output the final clothing display image.

[0104] Specifically, when performing semantic segmentation on mannequin images using the clothing parsing model, the S310 employs a clothing region segmentation model based on an improved Mask R-CNN. This model, pre-trained on the COCO dataset and fine-tuned using the FashionAI dataset, focuses on optimizing the boundary discrimination ability between clothing and human body regions. It supports pixel-level segmentation of at least 20 semantic categories of clothing, including the body, sleeves, collar, pockets, and buttons. After inputting the mannequin image into the model, multi-scale features are extracted through a Feature Pyramid Network (FPN), and candidate boxes are generated by combining it with a Region Proposal Network (RPN). After RoIAlign operation, a pixel-level clothing region segmentation mask is output. In the mask, the pixel value of the clothing region is set to 1, while the background and mannequin regions are set to 0, ensuring that the mask is aligned with the original image at the sub-pixel level. After segmentation, morphological closing operations are used to eliminate holes in the mask, and edge detection algorithms are used to extract the clothing outline, further optimizing the mask accuracy and providing accurate region localization for subsequent annotation of physical attribute semantic labels.

[0105] The physical property semantic labels in S320 are automatically annotated using a large visual understanding model, including material semantic labels and pattern semantic labels. Material semantic labels describe the optical and mechanical properties of the fabric, while pattern semantic labels describe the overall silhouette and tailoring features of the garment. These physical property semantic labels are embedded as prior knowledge into the Garment Deformation Network (GDN), guiding the GDN to generate differentiated deformation patterns based on different material responses.

[0106] Specifically, when labeling the semantic tags of clothing physical attributes, the BLIP-2 visual understanding big data model is used for automatic labeling based on the segmentation mask generated by S310 and the text description of S130.

[0107] First, close-up images of the materials and segmentation masks are input into the model. The model learns by comparison to identify the fabric's optical properties (e.g., gloss "matte / glossy", transparency "opaque / semi-transparent") and mechanical properties (e.g., drape "soft / crisp", elasticity "high elasticity / non-elasticity", wrinkle resistance "wrinkle-prone / wrinkle-resistant"), generating material semantic labels. Simultaneously, by combining the contour features of flat lay images and mannequin images, the model analyzes the overall garment silhouette (e.g., "X-shape", "H-shape", "O-shape"), cutting methods (e.g., "cinched waist", "loose fit", "asymmetrical cut"), and detailed structures (e.g., "princess seam" "yoke design"), generating pattern semantic labels. These labels are converted into vectors and embedded as prior knowledge into the input layer of the Garment Deformation Network (GDN). An attention mechanism guides the network to generate differentiated deformation patterns for different materials (e.g., the drape deformation of silk, the stiff deformation of denim).

[0108] The processing flow of the Garment Deformation Network (GDN) in S330 is to take the comprehensive garment representation (including segmentation mask, physical attribute semantic label vector), the preliminary dressing image, and the target pose skeleton map as input.

[0109] GDN employs a U-Net-based deformation prediction architecture. The input layer simultaneously receives visual features from a preliminary dressed image and a target pose skeleton map, along with semantic features from a comprehensive clothing representation. The encoder extracts multi-scale image and semantic features. In the decoder, deconvolution and skip connections generate a 2D displacement field with the same resolution as the image. The vector value of each pixel in the displacement field represents the offset of that location in the horizontal and vertical directions. GDN utilizes embedded physical attribute semantic labels to generate differentiated deformations for clothing regions of different materials and styles: for example, for a silk dress with excellent drape, a larger, natural drape deformation is generated at the hem; for a structured suit jacket, a slight, ergonomic fit deformation is generated at the shoulders and waist while preserving the fabric's crispness.

[0110] After generating the deformation field, the S340 uses a bilinear interpolation algorithm to perform pixel-level deformation distortion on the clothing areas marked by the segmentation mask in the initial clothing image, while the non-clothing areas remain unchanged. This allows the clothing to produce deformation effects that conform to physical laws according to the movement trends of the human body in the target posture (such as the stretching of sleeves when the arms are raised and the wrinkling of trouser legs when the legs are bent). Finally, a 2048×2732 pixel final clothing display image is output, with clear clothing details, natural posture, and realistic fabric characteristics.

[0111] Reference Figure 4 The training methods for the Garment Deformation Network (GDN) include S410-S460:

[0112] S410, acquires images of the test mannequin for the relevant clothing;

[0113] S420 performs region segmentation on the test mannequin image, defines the clothing segmentation mask and labels the physical attributes of the clothing to form a comprehensive test clothing representation;

[0114] S430, extract the skeleton map of the test mannequin image under the current training posture;

[0115] S440 inputs the skeleton map in the standard pose, the skeleton map in the current training pose, and the test integrated clothing representation into the clothing deformation network GDN to predict the deformation field from the standard pose to the current pose.

[0116] S450, apply the deformation field to the test mannequin image, and obtain the reconstructed image through a differentiable image deformation operation, performing the deformation operation only on the clothing area defined by the clothing segmentation mask;

[0117] S460 calculates the pixel-level difference loss between the reconstructed image and the original real image, and updates the GDN parameters based on this difference loss through backpropagation.

[0118] Specifically, the training of the Garment Deformation Network (GDN) (S310-S360) requires the construction of a dedicated training dataset. S410 collects mannequin images of 500 garments, each containing standard poses (e.g., standing at attention) and 5 non-standard poses (e.g., side bend, arm raise). S420 generates garment segmentation masks through a combination of manual annotation and automatic segmentation, with physical attribute labels provided by professional fashion designers. S430 extracts skeleton images under different poses as the basis for pose changes. During training, the standard pose skeleton image, the current training pose skeleton image, and the test garment representation are input into the GDN. After predicting the deformation field, deformation is applied to the standard pose mannequin image. The pixel-level L1 loss between the reconstructed image and the real non-standard pose mannequin image is calculated, while a deformation field smoothing loss is added (to avoid drastic local deformation). The SGD optimizer is used with a learning rate of 1e-3, momentum of 0.9, and training for 200 epochs, enabling the network to generate realistic deformation fields based on pose changes and physical attributes.

[0119] Reference Figure 5 Furthermore, the AI ​​model image generation methods include S510-S530:

[0120] S510: Extract ambient lighting information from the original reference image and construct a lighting map;

[0121] S520: Input the target pose skeleton map and clothing region segmentation mask into the pre-trained shadow generation model to predict the shadow distribution of the clothing in the target pose.

[0122] The S530 integrates the light map and shadow distribution into the final garment display image to generate a light-rendered image.

[0123] Specifically, in the lighting and shadow optimization stages S510-S530, S510 extracts ambient lighting information from the original reference image (such as the real-person reference image corresponding to the target pose selected in S120), and uses a deep learning-based lighting estimation model (such as Neural Illumination Estimation) to analyze the brightness distribution, highlight position, and shadow direction in the image, constructing a lighting map containing the light source direction, intensity, and color temperature. The lighting map is stored in the form of an environment map, supporting global lighting adjustments to the generated image. S520 inputs the target pose skeleton map and clothing region segmentation mask into a pre-trained shadow generation model. This model predicts the shadow distribution of clothing under the target pose based on the relative positional relationship between the human body pose and the clothing region, combined with a physical lighting model (such as the Phong lighting model), including the attached shadows of the clothing in contact with the human body (such as underarms and waist), the projected shadows formed by the clothing's own folds, and the projection relationship with the ground / background. The intensity and softness of the shadows are dynamically adjusted according to the light source distance and surface roughness. The S530 integrates the ambient lighting parameters of the lighting map and the predicted shadow distribution into the final clothing display image using image fusion technology (such as alpha blending). By adjusting the light intensity coefficient (0.3-0.7) and shadow transparency (0.4-0.6), the generated image has a realistic sense of light and shadow hierarchy, improving the overall visual realism.

[0124] Based on the above method embodiments, the second embodiment of this application discloses an AI model image generation system. The AI ​​model image generation system of this application embodiment can implement any of the above-described AI model image generation methods, and the specific working process of each module in the AI ​​model image generation system can be referred to the corresponding process in the above method embodiments.

[0125] For ease of understanding, an example is as follows: An AI model image generation system includes:

[0126] The data acquisition module is used to collect multimodal data of the target garment, including mannequin images, flat lay static images, and material close-up images.

[0127] The skeleton extraction module is used to extract the corresponding posture skeleton map as the target posture skeleton map after selecting a reference map whose posture matches the expectation from the target posture library.

[0128] The multimodal semantic parsing module is used to input tiled static images and material close-up images into the visual language model to generate text descriptions;

[0129] The model processing module is used to input the target pose skeleton map, mannequin image and text description into the pre-trained LoRA diffusion model to generate a preliminary dressed image; the input layer of the LoRA diffusion model adds an independent pose condition injection channel, and the target pose skeleton map is input into the LoRA diffusion model through the pose condition injection channel.

[0130] The third embodiment of this application provides a computer device, which may include a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement an AI model image generation method.

[0131] The memory can communicate with the processor via a communication bus, which can be an address bus, a data bus, a control bus, etc.

[0132] Additionally, the memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device.

[0133] Furthermore, the processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0134] The fourth embodiment of this application provides a computer-readable storage medium storing a computer program that can be loaded by a processor and executed as an AI model image generation method.

[0135] The computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device; the program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0136] It should be noted that the computer device and storage medium in the embodiments of this application are respectively electronic devices and storage media that apply the above-described AI model image generation method. Therefore, all embodiments of the above-described AI model image generation method are applicable to the computer device and storage medium, and can achieve the same or similar beneficial effects. For the computer device / storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple; relevant details can be found in the descriptions of the method embodiments.

[0137] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, disclosure, and appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce a good effect.

[0138] The above are all preferred embodiments of this application and are not intended to limit the scope of protection of this application. Any feature disclosed in this specification (including the abstract and drawings) may be replaced by other equivalent or similar features unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is only one example of a series of equivalent or similar features.

Claims

1. A method for generating AI model images, characterized in that, include: Collect multimodal data of the target garment, including mannequin images, static flat lay images, and close-up images of the material. Select a reference image whose pose matches the desired pose from the target pose library, and extract the pose skeleton image corresponding to the reference image as the target pose skeleton image; The tiling static image and the material close-up image are input into the visual language model to generate a text description; The target pose skeleton diagram, mannequin image, and text description are input into the pre-trained LoRA diffusion model to generate a preliminary dressed image. An independent pose condition injection channel is added to the input layer of the LoRA diffusion model, and the target pose skeleton diagram is input into the LoRA diffusion model through the pose condition injection channel. The steps following the generation of the initial clothing image include: The mannequin image is semantically segmented using a clothing analysis model to generate a pixel-aligned clothing region segmentation mask. Based on the clothing region segmentation mask and the text description, the physical attribute semantic tags of the clothing are labeled to form a comprehensive clothing representation that includes appearance, structure and physical characteristics. The integrated clothing representation, the target pose skeleton diagram, and the preliminary clothing image are input into the pre-trained Clothing Deformation Network (GDN) to generate a deformation field. Based on the deformation field, pixel-level deformation distortion is performed on the clothing area in the preliminary clothing image to output the final clothing display image; The physical attribute semantic tags are embedded as prior knowledge into the Garment Deformation Network (GDN) to guide the GDN in generating differentiated deformation patterns based on different material responses. The physical attribute semantic tags include material semantic tags and pattern semantic tags. The material semantic tags are used to describe the optical and mechanical properties of the fabric and are generated by the visual understanding big model recognizing close-up material images and segmentation masks. The pattern semantic tags are used to describe the overall silhouette and cutting features of the garment and are generated by the visual understanding big model recognizing flat lay images and mannequin images. The training method for the garment deformation network (GDN) includes: Collect images of the test mannequin wearing the relevant clothing; The test mannequin image is segmented into regions, and a clothing segmentation mask is defined and the physical properties of the clothing are labeled to form a comprehensive test clothing representation. Extract the skeleton image of the test mannequin in the current training pose; The skeleton diagrams under the standard posture, the skeleton diagrams under the current training posture, and the test integrated clothing representation are input into the GDN to predict the deformation field from the standard posture to the current posture. The deformation field is applied to the test mannequin image, and a reconstructed image is obtained through a differentiable image deformation operation. The deformation operation is performed only on the clothing area defined by the clothing segmentation mask. Calculate the pixel-level difference loss between the reconstructed image and the original real image, and update the GDN parameters based on the difference loss through backpropagation.

2. The AI ​​model image generation method according to claim 1, characterized in that, The training method for the LoRA diffusion model includes: Collect multi-angle real-life photos of mannequins wearing the relevant clothing; Each captured mannequin image is processed to extract a skeleton diagram containing pose information; For each mannequin image, a text description is matched and paired with the corresponding skeleton image to form training samples; A pre-trained LoRA model is used as the base model. An independent attitude condition injection channel is added to the input layer to form a LoRA diffusion model. The parameters of the base model and the attitude encoder weights are frozen. The training samples are input into the LoRA diffusion model to train the parameters of the LoRA adaptation layer and the projection layer output by the attitude encoder. Calculate the difference between the noise predicted by the LoRA diffusion model and the actual noise, and update the LoRA parameters based on the difference through backpropagation.

3. The AI ​​model image generation method according to claim 2, characterized in that, The steps for extracting a skeleton map containing pose information include: For each captured mannequin image, perform 2D human pose estimation to obtain the coordinates of key human points, and connect them to generate the corresponding initial pose skeleton diagram. The initial posture skeleton diagram is input into the posture repair network, which optimizes and completes the key points based on human kinematic priors to generate the repaired skeleton diagram.

4. The AI ​​model image generation method according to claim 1, characterized in that, The steps to build a target pose library include: Collect diverse real-life pose reference images from the internet, fashion image libraries, and professional photography materials; A consistent pose estimation algorithm, used in the same phase as the training phase, is used to extract skeleton maps corresponding to real-person pose reference images. These skeleton maps are then labeled and categorized according to action type to establish a searchable target pose database.

5. The AI ​​model image generation method according to claim 1, characterized in that, The AI ​​model image generation method also includes: Extract ambient lighting information from the original reference image to construct a lighting map; The target pose skeleton map and the clothing region segmentation mask are input into a pre-trained shadow generation model to predict the shadow distribution of the clothing in the target pose. The lighting map and shadow distribution are fused into the final clothing display image to generate a lighting rendering image.

6. An AI model image generation system, characterized in that, Performing the AI ​​model image generation method as described in any one of claims 1 to 5 includes: The data acquisition module is used to collect multimodal data of the target garment, including mannequin images, flat static images, and close-up images of the material. The skeleton extraction module is used to extract the corresponding posture skeleton map as the target posture skeleton map after selecting a reference map whose posture matches the expectation from the target posture library. The multimodal semantic parsing module is used to input the tiled static image and the material close-up image into the visual language model to generate a text description; The model processing module is used to input the target pose skeleton diagram, mannequin diagram and text description into the pre-trained LoRA diffusion model to generate a preliminary dressed image; the input layer of the LoRA diffusion model is equipped with an independent pose condition injection channel, and the target pose skeleton diagram is input into the LoRA diffusion model through the pose condition injection channel.

7. A computer device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the AI ​​model image generation method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer program is stored that can be loaded by a processor and executed as described in any one of claims 1 to 5 to generate AI model images.