Costume design system and method based on multi-modal AIGC and storage medium thereof

The clothing design AIGC system based on multimodal deep learning solves the problems of low efficiency and lack of innovation in traditional clothing design, realizes intelligent and automated clothing design, improves material collection efficiency and design quality assessment, and supports user feedback optimization.

CN120765778APending Publication Date: 2025-10-10重庆对外经贸学院
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510837568.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

The traditional clothing design process is inefficient and lacks innovation. Designers need to spend a lot of time searching for materials. In addition, the existing three-dimensional clothing design system has a high threshold for use and the cost of designer revisions is high.

Method used

The clothing design AIGC system based on multimodal deep learning is adopted, including heterogeneous multi-source input processing module, clothing feature extraction module, clothing semantic understanding module, multimodal fusion module, few-sample style transfer module, clothing design generation module, design evaluation module and user feedback module. Intelligent and automated clothing design is achieved through LoRA low-rank adaptation technology and multi-level conditional diffusion model.

Benefits of technology

It improves the efficiency and accuracy of design material collection, reduces data requirements, achieves precise control from clothing outline to details, optimizes design generation through user feedback, and supports multi-dimensional design quality evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765778A_ABST
    Figure CN120765778A_ABST
Patent Text Reader

Abstract

The invention relates to a costume design system and method based on multi-modal AIGC and a storage medium thereof. Comprising a multi-source input processing module used for receiving and processing multiple modal inputs including costume design text description, a costume style drawing, a fabric sample image and user preference data; the clothing feature extraction module is used for extracting clothing structure features and visual style features from the multi-source input; the semantic understanding module is used for carrying out semantic understanding and style analysis on the clothing features; the multi-modal fusion module is used for mapping the clothing feature representations of different modals to a unified semantic space and generating a fusion feature vector; and the style migration module is used for realizing design generation of a specific style based on a small number of garment samples through a LoRA low-rank adaptation technology, and is used for generating a garment design drawing through a multi-stage conditional diffusion model based on the fusion feature vector.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, in particular to a clothing design system and method based on multi-modal AIGC and a storage medium thereof. BACKGROUND

[0002] Traditional clothing design usually involves multiple steps such as style design, fabric selection, plate making, and cutting, which takes a long time period, and the cost of modifying the design after the designer finalizes the design is high, and the space for modification is small. With the development of three-dimensional modeling technology in the field of clothing design, designers can now quickly design and modify design drafts through CAD and other three-dimensional modeling technologies; at the same time, based on three-dimensional flexible simulation technology, many shopping websites have developed three-dimensional virtual fitting functions, and consumers can try on clothes online through 3D virtual fitting, saving the time and space cost of offline fitting.

[0003] However, the existing three-dimensional clothing design system model is complex and has high professional degree, with a high use threshold. At the same time, the design materials in the system are few, and designers need to spend a lot of time searching for materials and finding inspiration, which limits the application of three-dimensional modeling technology in the field of clothing design.

[0004] With the development of AIGC (Artificial Intelligence Generated Content) technology in recent years, especially in the image creation processing capability, it is possible to develop a clothing rapid design system based on AIGC technology. SUMMARY

[0005] To solve the above problems, the present application provides a clothing design system and method based on AIGC and image recognition technology. To solve the problems of low efficiency, lack of innovation and inflexible processing of design elements in the existing clothing design process, and to realize intelligent, automated and innovative clothing design.

[0006] To achieve the above purpose, the technical solution provided by the present application is:

[0007] A clothing design AIGC system based on multimodal deep learning includes: a heterogeneous multi-source input processing module for receiving and processing multi-modal inputs including clothing design text descriptions, clothing style diagrams, fabric sample images and user preference data; a clothing feature extraction module for extracting clothing structure features and visual style features from the multi-source inputs; a clothing semantic understanding module for performing semantic understanding and style analysis on the clothing features; a multimodal fusion module for mapping clothing feature representations of different modalities to a unified semantic space and generating a fused feature vector; a few-sample style transfer module for achieving design generation of a specific style based on a small number of clothing samples through LoRA low-rank adaptation technology; a clothing design generation module for generating clothing design drawings based on the fused feature vectors through a multi-level conditional diffusion model; a design evaluation module for evaluating the quality of the generated design from multiple dimensions of aesthetic design, manufacturability and market adaptability; and a user feedback module for obtaining user feedback on the generated results and guiding system optimization.

[0008] Furthermore, the heterogeneous multi-source input processing module includes: a text processing unit for processing clothing design text descriptions and extracting design intentions and style requirements; an image processing unit for processing clothing style drawings and reference images and extracting visual features; a fabric property analysis unit for processing fabric sample images and extracting fabric texture, glossiness, drape and other properties; and a user preference analysis unit for processing user historical selections and explicit preferences and generating user preference representations.

[0009] Furthermore, the user preference analysis unit includes a variational autoencoder structure for processing user historical choices and explicit preferences and learning the potential distribution representation of user preferences. The variational autoencoder includes an encoder network, a latent variable sampling mechanism, a decoder network and a regularization mechanism, and realizes probabilistic modeling of user preferences by minimizing reconstruction error and KL divergence.

[0010] Furthermore, the user preference analysis unit also includes a Transformer-based user behavior sequence modeling unit, which is used to capture the temporal dependencies of user historical choices, analyze user interest change trends through a multi-head self-attention mechanism, and distinguish long-term stable preferences from short-term intentions.

[0011] Furthermore, the clothing feature extraction module includes: a structural feature extraction unit, which is used to identify the silhouette, cutting lines, pattern structure and other features of the clothing; the structural feature extraction unit uses an image segmentation algorithm combined with Mask R-CNN to generate a clothing contour line map, separates the clothing area and the non-clothing area through key point detection, and uses non-uniform rational B-spline curve fitting to generate a parameterized and editable clothing contour line map; and a visual feature extraction unit, which is used to identify the color, texture, decoration and other features of the clothing; the visual feature extraction unit uses an improved YOLOv7 target detection model to extract color histogram clustering, pattern texture coding and decorative component positioning data.

[0012] Furthermore, the clothing semantic understanding module includes: a style recognition unit for identifying the design style of clothing, such as classical, modern, simple, luxurious, and sporty; a design element analysis unit for analyzing the semantic attributes of each design element in clothing; and a semantic mapping unit for achieving semantic alignment of text and images based on the CLIP model, mapping design features to a shared semantic space.

[0013] Furthermore, the multimodal fusion module adopts a multi-head attention mechanism and a cross-modal Transformer architecture to achieve the fusion of different modal information, through the formula:

[0014]

[0015] Where: E i 、E j represents the eigenvectors from different modes; W q 、W k represents the learnable weight matrix, d represents the dimension of the feature vector, A i,j represents the attention weight of the i-th eigenvector to the j-th eigenvector, and softmax represents the normalization function;

[0016] Calculate the attention distribution between different modal information and generate a fusion representation vector:

[0017]

[0018] Among them, E f Represents the unified feature vector after fusion; m∈M={1,...n}, M represents the modality set (such as M={text, image, audio}, m is a single modality in the modality set, E m The original feature vector of the mth modality (such as word embedding of text, CNN features of image); α m The dynamic weight coefficient of the mth mode, P m (E m ) represents the original feature E mProjection function mapping to fusion space;

[0019] where α m is the weight coefficient dynamically assigned to each modality, calculated using the following formula:

[0020]

[0021] where g(E m ) represents the feature transformation function of vector E m , and ω m is a learnable parameter vector. Further, the few-shot style transfer module, based on the LoRA low-rank adaptation technology, realizes the learning and transfer of specific clothing styles by inserting a low-rank adaptation layer in the pre-trained diffusion model; the parameter update formula of the LoRA adaptation layer is:

[0022] W = W0 + ΔW = W0 + BA

[0023] where W0 is the pre-trained weight matrix, ΔW is the low-rank update, B ∈ R (d×r) , A ∈ R (r×k) , and r << min(d, k). By controlling the rank parameter r, the system can balance the learning ability and the risk of overfitting.

[0024] In clothing design, the rank parameter r of LoRA is dynamically adjusted according to the complexity of the clothing style:

[0025] r = log2(N s *N v )

[0026] where N s is the complexity of the clothing structure elements, and N v is the complexity of the clothing visual style.

[0027] The style transfer module is combined with the multi-level conditional diffusion model through the following steps: first, collect a small number of clothing samples (usually 5-10 images) representing a specific style; extract the style feature representation in the samples; fine-tune the specific layers of the pre-trained diffusion model, especially the cross-attention layer, through the LoRA technology; integrate the fine-tuned model parameters into the multi-level conditional diffusion model of the clothing design generation module; during the generation process, the fine-tuned model can apply the learned specific style features to new design generation.

[0028] This combination allows the system to quickly learn and apply the style of a specific designer, the design language of a specific brand, or the popular trends of a specific period, while maintaining the basic generation ability of the diffusion model, greatly enhancing the adaptability and creativity of the system.

[0029] Furthermore, the clothing design generation module adopts a multi-level conditional diffusion model to achieve hierarchical design generation from outline to details.

[0030] Furthermore, the clothing design generation module operates as follows:

[0031] The contour generation unit generates the basic contour of the garment based on the initial random noise and structural feature conditions through the first diffusion model, and determines the overall structural framework of the garment;

[0032] The detail generation unit uses the second diffusion model to input the initial random noise, structural feature conditions, and style feature conditions, and combines them with the basic outline generated in the first step to refine the style features and texture details of the clothing and generate an intermediate result with a specific style and texture;

[0033] The rendering optimization unit uses the third diffusion model to render the material texture of the clothing based on the initial random noise, structural feature conditions, style feature conditions, material feature conditions, and combined with the intermediate results with style texture generated in the second step, and finally generates a finished clothing design with complete outline, style, texture and material effects.

[0034] Furthermore, the clothing design generation module adopts a multi-level conditional diffusion model to achieve hierarchical design generation from outline to details.

[0035] Furthermore, the design evaluation module includes:

[0036] Aesthetics scoring unit, which evaluates the design's aesthetics and innovation;

[0037] The manufacturability assessment unit evaluates the structural rationality, manufacturing complexity and material compatibility of the design; and the market adaptability assessment unit evaluates the matching degree of the design with the target market and consumer groups.

[0038] Furthermore, the user feedback module adopts the DDPG algorithm to achieve design parameter optimization based on user feedback; the DDPG algorithm includes: an actor network, which is used to generate design parameter adjustment actions based on the current design state and user preference representation; a critic network, which is used to estimate the action value function and evaluate the expected benefits of design adjustments; an experience replay buffer, which is used to store user interaction experience and improve sample utilization efficiency; and a target network update mechanism, which stabilizes the learning process through a soft update strategy.

[0039] Furthermore, the DDPG algorithm also includes: a state representation mechanism that combines the current design parameters and user preference representation into a state vector; continuous action space modeling that supports fine-grained adjustment of design parameters; deterministic policy gradient calculation that directly maximizes expected returns; and a noise exploration strategy that generates time-correlated noise through the Ornstein-Uhlenbeck process.

[0040] An AIGC method for clothing design based on multimodal deep learning, including:

[0041] Receive and process multimodal inputs such as clothing design text descriptions, clothing style images, fabric sample images, and user preference data;

[0042] Extract clothing structure features and visual style features from input data;

[0043] Perform semantic understanding and style analysis on the extracted clothing features;

[0044] Map clothing feature representations of different modalities into a unified semantic space and generate a fused feature vector;

[0045] Learn specific styles based on a small number of clothing samples through LoRA low-rank adaptation technology;

[0046] Based on the fusion feature vector, the clothing design image is generated through the multi-level conditional diffusion model;

[0047] Evaluate the quality of generative design from multiple dimensions including aesthetic design, manufacturability, and market adaptability;

[0048] Obtain user feedback on the generated results, use the DDPG algorithm to process user feedback, and optimize the design generation parameters through off-policy learning.

[0049] As an optional implementation, the above method further includes:

[0050] The contour line graph of clothing structural features is used as the conditional input of ControlNet;

[0051] Encode clothing style features into model prompt words through Textual Inversion;

[0052] The embedding vector of Textual Inversion is updated through gradient backpropagation, driving the generated results to evolve towards the target style.

[0053] As another optional implementation, the above method further includes:

[0054] Based on the U-Net architecture in the diffusion model, the intermediate states during the conditional injection process are tracked;

[0055] Visualize the attention map to analyze the model's attention distribution on each design element; guide the fine-tuning of design elements based on the attention map.

[0056] Beneficial effects

[0057] The clothing design system and method of the present invention, by combining AIGC and image recognition technology, can automatically extract rich clothing design elements from input media containing videos of people wearing clothes, greatly improving the efficiency and accuracy of design material collection. The present invention can simultaneously process multiple heterogeneous data such as text descriptions, clothing images, fabric samples, and user preferences, providing more comprehensive design information support. The low-rank adaptation method based on LoRA technology only requires a small number of clothing samples to learn specific styles, significantly reducing data requirements. Through a multi-level conditional diffusion model and ControlNet technology, precise control from clothing outline to details is achieved. A built-in clothing design evaluation system evaluates design quality from multiple dimensions such as aesthetics, manufacturability, and market adaptability. The system supports user feedback and iterative optimization, and continuously improves system performance through a hybrid reinforcement learning strategy. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solution of the present invention, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0059] The methods, systems, and / or programs in the accompanying drawings will be further described according to exemplary embodiments. These exemplary embodiments will be described in detail with reference to the drawings. These exemplary embodiments are non-limiting exemplary embodiments, wherein example numerals represent similar structures in the various views of the drawings.

[0060] Figure 1 This is a system structure framework diagram provided by an embodiment of the present application;

[0061] Figure 2 This is a system flow chart provided in an embodiment of the present application. DETAILED DESCRIPTION

[0062] In order to better understand the above technical solution, the technical solution of the present application is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present application and the specific features in the embodiments are detailed descriptions of the technical solution of the present application, rather than limitations on the technical solution of the present application. In the absence of conflict, the embodiments of the present application and the technical features in the embodiments can be combined with each other.

[0063] The following is combined with Figure 1 The embodiments of the present application are described.

[0064] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0065] like Figure 1 As shown, the clothing design AIGC system based on multimodal deep learning provided by the present invention includes a multi-source input processing module 10, a feature extraction module 20, a semantic understanding module 30, a multimodal fusion module 40, a style transfer module 50, a clothing design generation module 60, a design evaluation module 70 and a user feedback module 80.

[0066] The multi-source input processing module 10 is used to receive and process multi-modal inputs, including clothing design text descriptions, clothing style diagrams, fabric sample images, and user preference data. The feature extraction module 20 is used to extract clothing structural features and visual style features from the multi-source inputs. The semantic understanding module 30 is used to perform semantic understanding and style analysis on the extracted clothing features. The multimodal fusion module 40 is used to map clothing feature representations from different modalities into a unified semantic space and generate a fused feature vector. The style transfer module 50 is used to generate designs of a specific style based on a small number of clothing samples using the LoRA low-rank adaptation technique. The clothing design generation module 60 is used to generate clothing design drawings based on the fused feature vectors using a multi-level conditional diffusion model. The design evaluation module 70 is used to evaluate the quality of the generated design from multiple dimensions: aesthetic design, manufacturability, and market adaptability. The user feedback module 80 is used to obtain user feedback on the generated results and guide system optimization.

[0067] The specific implementation of each module is described in detail below.

[0068] Multi-source input processing module

[0069] The heterogeneous multi-source input processing module 10 includes a text processing unit 11 , an image processing unit 12 , a fabric property analysis unit 13 and a user preference analysis unit 14 .

[0070] The text processing unit 11 is used to process the text description of clothing design and extract the design intention and style requirements. In a specific implementation, the text processing unit 11 uses a pre-trained language model, such as BERT or T5, to convert the clothing design text into a semantic vector representation. The text content may include a description of the design style (such as "simple and modern", "retro and elegant", etc.), the target audience (such as "25-35 years old professional women"), occasions (such as "business formal", "leisure vacation", etc.) and design element requirements (such as "puff sleeves", "high waist design", etc.). The text processing unit 11 not only extracts keywords, but also understands the semantic relationship between words to form a structured design requirement representation.

[0071] The image processing unit 12 is used to process clothing style images and reference images to extract visual features. Specifically, the image processing unit 12 processes the input images using a multi-layer convolutional neural network to extract multi-scale visual features. For clothing style images, the analysis focuses on the garment's silhouette, cut lines, and structural relationships; for reference images, the analysis focuses on extracting style elements, texture features, and the distribution of visual elements. The image processing unit 12 also supports multi-image input, processing clothing images from multiple angles and views simultaneously to form a more comprehensive visual feature representation.

[0072] The fabric property analysis unit 13 is used to process fabric sample images and extract fabric properties such as texture, gloss, and drape. In practice, the fabric property analysis unit 13 uses a neural network specifically trained for material analysis, combined with texture feature extraction and reflected light analysis, to identify the physical properties of the fabric. By analyzing sample images under multi-angle lighting conditions, the system can assess the fabric's light transmittance, reflectivity, and texture changes. The fabric property analysis unit 13 also maintains a fabric knowledge base, matching extracted features with known fabric properties to infer fabric attributes such as drape, elasticity, and thickness.

[0073] The user preference analysis unit 14 processes historical user selections and explicit preferences to generate a user preference representation. In practice, this unit builds a user preference model through a combination of collaborative filtering and content feature analysis. The system records historical user interaction data, including selections, ratings, and feedback, and extracts user preference characteristics for color, style, fabric, and other aspects. For new users or those with unclear preferences, the system proactively guides users through simple preference selections to quickly build an initial preference model. The user preference representation serves as a special conditional input, influencing the subsequent design generation process.

[0074] Clothing feature extraction module

[0075] The clothing feature extraction module 20 includes a structural feature extraction unit 21 and a visual feature extraction unit 22 .

[0076] The structural feature extraction unit 21 is used to identify features such as the silhouette, cutting lines, and pattern structure of clothing. In a specific implementation, the structural feature extraction unit 21 uses an image segmentation algorithm combined with Mask R-CNN to generate a clothing contour line map. First, the human body posture is estimated for the selected image, and the clothing area and the non-clothing area are separated by key point detection. Then, a parametric and editable clothing contour line map is generated by non-uniform rational B-spline (NURBS) curve fitting. The structural feature extraction process includes: edge detection, region segmentation, contour extraction, and parametric modeling. The system subdivides the clothing structure into areas such as the collar, sleeves, body parts, and hem, extracts specific structural parameters for each area, and constructs a skeleton representation of the clothing.

[0077] The visual feature extraction unit 22 is used to identify clothing features such as color, texture, and decoration. In practice, the visual feature extraction unit 22 uses an improved YOLOv7 object detection model to extract color histogram clustering, pattern texture encoding, and decorative component location data. Color analysis uses an adaptive color quantization algorithm to extract the clothing's primary color, color scheme, and color distribution. Texture analysis uses a multi-scale texture filter bank to extract texture features, including dimensions such as directionality, regularity, density, and contrast. Decorative component detection focuses on identifying the location, shape, and style of decorative elements such as buttons, pleats, embroidery, and pockets. These visual features collectively constitute the visual style representation of the clothing.

[0078] Semantic Understanding Module

[0079] The clothing semantic understanding module 30 includes a style recognition unit 31 , a design element analysis unit 32 and a semantic mapping unit 33 .

[0080] The style recognition unit 31 is used to identify clothing design styles, such as classic, modern, minimalist, luxurious, and sporty. In practice, this unit is based on a pre-trained style classification model trained on a large-scale clothing style dataset and capable of identifying dozens of mainstream clothing style categories. This model utilizes a feature pyramid architecture, simultaneously considering global style features and local design details to improve the accuracy and robustness of style recognition. In addition to the classification results, the system also outputs a style probability distribution, indicating the design's similarity to various styles, providing a reference for subsequent style fusion and transfer.

[0081] The design element analysis unit 32 is used to analyze the semantic attributes of each design element in a garment. In practice, the unit decomposes the garment into multiple design elements, such as collar type, sleeve type, waistline, skirt length / trouser type, and performs fine-grained semantic analysis on each element. The system maintains a design element semantic knowledge base, recording the professional terminology, style affiliation, and usage context of each design element. Through image-text matching learning, the system can associate visual design elements with textual descriptions, supporting visual-semantic correspondence between professional terms such as "V-neck," "puff sleeve," and "fishtail skirt."

[0082] The semantic mapping unit 33 achieves semantic alignment between text and images based on the CLIP model, mapping design features into a shared semantic space. In specific implementations, the semantic mapping unit 33 is based on the pre-trained CLIP model and fine-tuned for the clothing design domain using domain adaptation technology. This unit implements a bidirectional mapping between text descriptions and clothing images, enabling the system to understand complex queries such as "This dress is bohemian style" or "Find a minimalist top similar to this image." By constructing a shared semantic space, the system enables information from different modalities (text descriptions, image features, user preferences, etc.) to be compared and integrated within the same semantic framework.

[0083] Multimodal fusion module

[0084] The multimodal fusion module 40 adopts a dynamic multi-head attention mechanism to calculate the attention distribution between different modal information and generate a fusion representation vector.

[0085] In the specific implementation, the multimodal fusion module 40 first converts the input of different modalities into feature vectors through modality-specific encoders. The multimodal fusion module adopts a multi-head attention mechanism and a cross-modal Transformer architecture to achieve the fusion of different modal information, through the formula:

[0086]

[0087] Where: E i 、E j represents the eigenvectors from different modes; W q 、W k represents the learnable weight matrix, d represents the dimension of the feature vector, A i,j represents the attention weight of the i-th eigenvector to the j-th eigenvector, and softmax represents the normalization function;

[0088] Calculate the attention distribution between different modal information and generate a fusion representation vector:

[0089]

[0090] Among them, E f Represents the unified feature vector after fusion; m∈M={1,...n}, M represents the modality set (such as M={text, image, audio}, m is a single modality in the modality set, E m The original feature vector of the mth modality (such as word embedding of text, CNN features of image); α m The dynamic weight coefficient of the mth mode, P m (E m ) represents the original feature E m Projection function mapped to the fusion space;

[0091] Among them, α m The weight coefficient dynamically assigned to each mode is calculated using the following formula:

[0092]

[0093] Among them, g(E m ) represents the vector E m The characteristic conversion function, ω m is the learnable parameter vector.

[0094] For example, for the text feature E text , image feature E image , fabric feature E material and user preference characteristics E user , the system is based on the formula Compute the cross-modal attention distribution between them.

[0095] Based on attention distribution, Calculate and generate the fusion representation vector: where m∈M, E m are the eigenvectors of different modes.

[0096] Weight coefficient α m Dynamic calculation through context-aware mechanism: where w _m is the learning parameter, g(E m ) represents the vector E m The feature conversion function of .

[0097] This dynamic weighting mechanism enables the system to adaptively adjust the importance of each modality based on the specific design task and input content. For example, when the user provides a detailed text description but no reference image, the system increases the weight of the text features; when the user uploads a clear reference image, the system increases the weight of the image features.

[0098] Style Transfer Module

[0099] The style transfer module 50 is based on the LoRA low-rank adaptation technology, which enables the learning and transfer of specific clothing styles by inserting a low-rank adaptation layer into the pre-trained diffusion model.

[0100] In specific implementations, the style transfer module 50 uses LoRA (Low-Rank Adaptation) technology to quickly learn specific clothing styles without retraining the entire model. The system inserts a low-rank adaptation layer into the cross-attention layer of the pre-trained StableDiffusion model and updates the model weights in the following way:

[0101] W = W0 + ΔW = W0 + BA;

[0102] where W0 is the pre-trained weight matrix, ΔW is the low-rank update, B is the basis vector matrix, B ∈ R (d×r) , A is the coefficient matrix, A ∈ R (r×k) , and r << min(d, k). By controlling the rank parameter r, the system can balance learning ability and overfitting risk.

[0103] In fashion design, the rank parameter r of LoRA is dynamically adjusted according to the complexity of the clothing style:

[0104] r = log2(N s × N v );

[0105] where N s is the complexity of the clothing structure elements, and N v is the complexity of the clothing visual style.

[0106] The style transfer process includes the following steps: collect a small number of clothing samples representing a specific style, which can be 5-10 images according to actual conditions; extract the style feature representation in the samples; fine-tune the specific layers of the pre-trained model through LoRA technology; verify the style transfer effect and adjust the parameters; apply the learned style knowledge to new design generation.

[0107] This method enables the system to quickly learn and apply the style of a specific designer, the design language of a specific brand, or the popular trends of a specific period, greatly enhancing the adaptability and creativity of the system.

[0108] Clothing design generation module

[0109] The clothing design generation module 60 includes a contour generation unit 61, a detail generation unit 62, and a rendering optimization unit 63.

[0110] The contour generation unit 61 generates the basic contour of the clothing based on the clothing structure features. In specific implementation, the contour generation unit 61 uses the conditional diffusion model, inputs the structure feature representation as a condition, and generates the basic contour of the clothing. This stage mainly focuses on the contour, proportion, and main cutting line of the clothing, and does not involve texture and details. The system uses ControlNet technology to input the extracted contour line graph as a spatial constraint condition, ensuring that the generated clothing conforms to the expected structure features.

[0111] Detail generation unit 62 supplements garment details based on visual style features. Specifically, this unit adds texture, embellishments, and visual elements to the garment's outline. This unit employs an attention-based detail generation mechanism, adding appropriate design details based on key elements in the visual style features. For example, elements such as pleats, ruffles, embroidery, or patterns can be added based on the style requirements. The detail generation process respects the spatial constraints of the outline stage, ensuring that detail enhancement does not disrupt the underlying structure.

[0112] The rendering optimization unit 63 performs texture rendering and effect optimization based on fabric properties. In practice, the rendering optimization unit 63 considers the physical properties of the fabric, such as drape, reflectivity, and texture details, to perform the final rendering of the generated garment image. This unit utilizes a specialized material rendering network to simulate the visual effects of different materials under different lighting conditions based on fabric parameters. Furthermore, the rendering optimization unit 63 performs overall color balancing, contrast adjustment, and detail enhancement to improve the visual quality and realism of the generated image.

[0113] The clothing design generation module 60 adopts a multi-level conditional diffusion model, and its operation process is as follows:

[0114] Through the first diffusion model, based on the initial random noise and structural characteristic conditions E struct (such as the basic version and outline of the clothing, etc.), generate the basic outline of the clothing, and determine the overall structural framework of the clothing, such as the general style of a dress, jacket or pants.

[0115] With the help of the second diffusion model, the initial random noise x is input T , structural characteristic condition E struct , style feature condition E style (such as retro, simple, street style elements), and combined with the basic outline generated in the first step x shape , refine the style features and texture details of the clothing and generate an intermediate result x with a specific style and texture texture , such as adding patterns such as plaid and polka dots to the basic outline, or adjusting the design style of the neckline and cuffs.

[0116] Using the third diffusion model, based on the initial random noise x T , structural characteristic condition E struct , style feature condition E style , Material characteristic condition E material (such as silk, denim, wool and other material properties), and combined with the intermediate result x with style texture generated in the second step texture , rendering the material texture of the clothing, and finally generating a finished clothing design with complete outline, style, texture and material effect final, such as simulating the sheen of silk or the rough texture of denim, giving designs more realism and detail.

[0117] In practical applications, the clothing design generation module 60 also supports the following functions: the contour line graph of the structural elements is used as the conditional input of ControlNet; the style elements are encoded as model prompt words through Textual Inversion; and the embedding vector of the semantic label is updated through gradient backpropagation to drive the generation results to continuously evolve towards the target style.

[0118] Design Evaluation Module

[0119] The design evaluation module 70 includes an aesthetic scoring unit 71 , a manufacturability evaluation unit 72 , and a market suitability evaluation unit 73 .

[0120] Aesthetic scoring unit 71 evaluates the aesthetics and innovation of the design. Specifically, this unit evaluates the aesthetic quality of the generated design based on a scoring model trained using a large-scale clothing aesthetic evaluation dataset. Evaluation criteria include proportional coordination, color harmony, innovation, and overall aesthetic quality. The system combines objective computational metrics (such as compositional balance and color contrast) with learned subjective aesthetic preferences to generate a comprehensive aesthetic score.

[0121] The manufacturability assessment unit 72 evaluates the design's structural rationality, manufacturing complexity, and material compatibility. In practice, the manufacturability assessment unit 72 primarily examines the actual manufacturability of the generated design. The system analyzes the geometric complexity of the garment design, including the number of curves, corners, and irregular shapes; analyzes the compatibility of the design form with the physical properties of the selected fabric; and assesses the complexity of the process technology required to implement the design. The manufacturability score is calculated using the following formula:

[0122] S m =w1*S structure +w2*S fabric +w3*S technique

[0123] Among them, w1, w2, w3 are weight coefficients, S structure is the structural rationality score, S fabric is the fabric compatibility score, S technique is the process complexity score.

[0124] The market suitability assessment unit 73 evaluates the design's compatibility with the target market and consumer group. In practice, this unit assesses the design's commercial potential based on market trend data and a target demographic preference model. The system analyzes the design's compatibility with current trends, its compatibility with the target price range, and its appeal to the target consumer group. This market suitability assessment combines quantitative analysis (such as similarity calculation) with qualitative judgment to provide a multi-dimensional market assessment.

[0125] The design evaluation module 70 combines the evaluation results of the three units to form a comprehensive design evaluation report. The evaluation results not only serve as the basis for design screening, but also guide the optimization of the design generation process through a gradient feedback mechanism.

[0126] User feedback module

[0127] The user feedback module 80 uses a hybrid reinforcement learning strategy to optimize system generation parameters based on user feedback.

[0128] In a specific implementation process, the user feedback module 80 can support two feedback modes: one is policy gradient update based on explicit user feedback, and the other is reward inference based on implicit user behavior.

[0129] For explicit feedback, the system collects user ratings, choices, and suggested modifications for the generated designs, and directly calculates the feedback loss. The feedback loss is calculated by associating the probability of the model policy taking an action in a specific state with the user feedback reward.

[0130] For implicit behaviors, the system analyzes behavioral data such as user browsing time, click order, and gaze retention, using inverse reinforcement learning to infer user preferences. The reward inference logic is as follows: based on user behavior data D, the inferred reward function R, the entropy of the reward function H(R), and the tuning parameter λ, the inferred reward function Rθ is determined by maximizing the product of the logarithm of the probability of the user behavior data under the reward function, the entropy of the reward function, and the tuning parameter λ.

[0131] The clothing design AIGC system based on multimodal deep learning provided by the present invention has the following significant advantages and technical effects:

[0132] The system simultaneously processes multimodal information, including text descriptions, images, and fabric physical properties. Through an innovative attention fusion mechanism, it effectively integrates this heterogeneous data, providing comprehensive information support for apparel design. Utilizing LoRA low-rank adaptation technology, the system can understand and apply specific design styles with only a small number of samples, significantly reducing training data requirements. Using a multi-level conditional diffusion model and ControlNet technology, it achieves precise control from garment outline to detail. Evaluation results provide reliable guidance for design screening and optimization.

[0133] By deeply integrating AIGC technology with professional knowledge of clothing design, this invention innovatively solves the problems of creative generation, professional evaluation and customized design in clothing design, providing strong technical support for the digital transformation of the clothing design industry.

[0134] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A clothing design system based on multimodal AIGC technology, characterized in that: include: Multi-source input processing module, which receives and processes multiple modal inputs including clothing design text descriptions, clothing style diagrams, fabric sample images, and user preference data; A feature extraction module extracts clothing structural features and visual style features from the multiple modal inputs; A semantic understanding module, which performs semantic understanding and style analysis on the clothing features; Multimodal fusion module, used to map clothing feature representations of different modalities into a unified semantic space and generate a fused feature vector; The style transfer module uses LoRA low-rank adaptation technology to achieve design transfer of a specific style based on a small number of clothing samples; Clothing design generation module, which is used to generate clothing design images based on the fused feature vectors through a multi-level conditional diffusion model; Design evaluation module, which evaluates the quality of generative design from multiple dimensions including aesthetic design, manufacturability, and market adaptability; The user feedback module obtains user feedback information on the generated results, and feeds the evaluation quality and feedback information back to the clothing design generation module to obtain an optimized clothing design drawing.

2. The system according to claim 1, wherein: The clothing feature extraction module includes: the multi-source input processing module includes: a text processing unit for processing text information such as clothing design descriptions, style requirements and functional requirements; an image processing unit for processing clothing style drawings, reference design drawings and fabric images; and a user data processing unit for obtaining user feedback on the generated results and feeding back system optimization.

3. The system according to claim 1, wherein: The clothing feature extraction module includes: a structural feature extraction unit for extracting structural features of clothing such as cutting lines, stitching structures, and component shapes, and generating a clothing outline map through an image segmentation algorithm; a visual feature extraction unit for extracting visual features of clothing such as color, texture, and decoration; and a fabric property extraction unit for extracting physical properties and visual effect features of the fabric.

4. The system according to claim 1, wherein: The clothing semantic understanding module includes: a text encoder for encoding clothing description text into a feature vector; an image encoder for encoding clothing images into a feature vector; and a semantic alignment unit for aligning text features with image features in a shared embedding space.

5. The system according to claim 1, wherein: The multimodal fusion module adopts a multi-head attention mechanism and a cross-modal Transformer architecture to achieve the fusion of information from different modalities. The formula is: Where: E i 、E j represents the eigenvectors from different modes; W q 、W k represents the learnable weight matrix, d represents the dimension of the feature vector, A i,j represents the attention weight of the i-th eigenvector to the j-th eigenvector, and softmax represents the normalization function; Calculate the attention distribution between different modal information and generate a fusion representation vector: Among them, E f Represents the unified feature vector after fusion; m∈M={1,...n}, M represents the modality set (such as M={text, image, audio}, m is a single modality in the modality set, E m The original feature vector of the mth modality (such as word embedding of text, CNN features of image); α m The dynamic weight coefficient of the mth mode, P m (E m ) represents the original feature E m Projection function mapped to the fusion space; Among them, α m The weight coefficient dynamically assigned to each mode is calculated using the following formula: Among them, g(E m ) represents the vector E m The characteristic conversion function, ω m is the learnable parameter vector.

6. The system according to claim 1, wherein: The clothing design generation module is implemented using a multi-stage conditional diffusion model, including: a contour structure generation stage, which guides the diffusion process based on structural conditions to generate the basic contour of the clothing; a style and detail enhancement stage, which generates clothing designs with texture and details based on style conditions and contour results; and a physical effect rendering stage, which simulates the real wearing effect of the clothing based on the physical properties of the fabric. Among them, each stage adopts a conditional injection U-Net architecture, and injects conditional information into the diffusion process through the CrossAttention layer.

7. The system according to claim 6, characterized in that The clothing design generation module also includes a ControlNet unit, which controls the generation process through the following steps: binarizing the clothing structure contour line map; using the processed contour line map as the input of ControlNet; fusing it with the feature maps of each layer of U-Net through a zero convolution layer; and adjusting the noise prediction results in real time during the diffusion sampling process.

8. The system according to claim 1, wherein: The design evaluation module includes: an aesthetic evaluation unit, which evaluates the visual appeal of the design based on a pre-trained aesthetic scoring model; a manufacturability evaluation unit, which evaluates the structural rationality and production feasibility of the design; a market adaptability evaluation unit, which evaluates the design's compatibility with current trends and target user groups; and a feedback optimization unit, which adjusts generation parameters through gradient backpropagation based on the evaluation results to optimize the generation results.

9. The system according to claim 8, characterized in that The user feedback module uses the DDPG algorithm to process user feedback and optimizes the design generation parameters through off-policy learning.

10. A clothing design AIGC method based on multimodal deep learning, characterized in that: The following steps are involved: Receive and process multimodal inputs such as clothing design text descriptions, clothing style images, fabric sample images, and user preference data; Extract clothing structure features and visual style features from input data; Perform semantic understanding and style analysis on the extracted clothing features; Map clothing feature representations of different modalities into a unified semantic space and generate a fused feature vector; Learn specific styles based on a small number of clothing samples through LoRA low-rank adaptation technology; Based on the fusion feature vector, the clothing design image is generated through the multi-level conditional diffusion model; Evaluate the quality of generative design from multiple dimensions including aesthetic design, manufacturability, and market adaptability; Obtain user feedback on the generated results, use the DDPG algorithm to process user feedback, and optimize the design generation parameters through off-policy learning.

11. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the method according to claim 10 when executed by a processor.

Citation Information

Cited By

  • Method and device for generating costume design scheme embedded with cultural decoration symbols and medium

    CN121188852A

  • Ancient furniture decoration pattern generation method based on multi-modal large model

    CN121414902A

  • Gift personalized customization method and system based on multi-mode AIGC

    CN121561998A

  • Garment image generation method based on basic element retrieval and replacement

    CN122049119A