A clothing design incremental generation method and system based on DiT architecture

Through the incremental generation method of clothing design of DiT architecture, combined with global and local optimization, the existing AI tools have solved the problem of insufficient design continuity, accuracy and interactivity, and achieved an efficient clothing design process.

CN119578226BActive Publication Date: 2025-08-26CHENGDU RAINCHAIN TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411629569.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-15
Publication Date
2025-08-26
Estimated Expiration
2044-11-15

AI Technical Summary

Technical Problem

Existing generative AI tools have insufficient design continuity, accuracy problems, local tunability limitations and insufficient feedback and iterative capabilities in clothing design, resulting in inefficient design efficiency and poor quality.

Method used

The incremental generation method of clothing design based on DiT architecture is adopted. Through phased incremental generation, conditional control and multi-path diffusion, combined with global and local optimization, the Transformer model is used for text word meaning analysis and encoding, and image generation and local adjustment are carried out in combination with diffusion model and DiT model, supporting designers' immediate feedback and optimization.

Benefits of technology

The gradual refinement and adjustment of the design is achieved, the continuity, accuracy and local adjustability of the design are improved, the interaction ability between designers and systems is enhanced, and the design efficiency and quality are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119578226B_ABST
    Figure CN119578226B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for incremental generation of clothing design based on the DiT architecture, which relates to the technical field of fashion design. The method utilizes the powerful generation capability of the DiT architecture, and can, on the basis of the initial design, gradually adjust local details through multi-stage incremental generation and refinement, thereby achieving continuity of design and local fine-tuning. By combining the fusion of conditional control technology and multi-path diffusion model, designers can flexibly modify specific areas in the design process and dynamically optimize the overall design scheme based on feedback. This method not only significantly improves the accuracy and controllability of clothing design, but also effectively solves the problem of insufficient flexibility in traditional design processes, and provides designers with a more intelligent and efficient design tool. The present invention can achieve gradual refinement and adjustment of the design through the combination of phased incremental generation, conditional control and multi-path diffusion, thereby significantly improving the shortcomings of traditional design tools.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of fashion design technology, and in particular to a method and system for incrementally generating clothing designs based on a DiT architecture. Background Art

[0002] In the fields of intelligent design and computer-aided design, particularly in fashion design, existing generative artificial intelligence (Generative AI) tools such as StyleGAN and Variational Autoencoders (VAE) have begun to be applied experimentally. These tools generate novel design artwork by learning from a large number of design samples, significantly promoting design innovation. However, despite their outstanding performance in generating novel designs, these technologies still have some significant drawbacks, as analyzed below:

[0003] Lack of design continuity: Traditional generative AI tools typically use a full-scale generation approach, requiring each design update to start from scratch. This approach results in a lack of continuity between design versions, often causing designers to lose creative coherence when adjusting existing designs, resulting in low creative efficiency. Faced with multiple revisions, designers are unable to retain completed design elements, making it difficult to form a complete system design.

[0004] Accuracy Issues: In traditional design processes, designers often rely on manual editing to adjust local details, while AI-based generative tools have limited control over these details. This is because generative models often struggle to understand the designer's specific intent during local generation, resulting in significant discrepancies between the generated results and the designer's expectations. This lack of accuracy stems primarily from the model's inability to understand complex design concepts, limiting its effective application in professional design.

[0005] Limited local adjustability: Existing generative AI tools often lack flexible local adjustment mechanisms. Most tools only support global parameter adjustments, preventing detailed modifications to specific parts. This limitation prevents designers from receiving immediate feedback and optimizing specific design details, thus affecting the quality and creative expression of the final design. Traditional tools are typically based on statically generated frameworks, which cannot support complex interactive design processes.

[0006] Inadequate feedback and iteration capabilities: Existing design tools often lack real-time interaction with designers and are unable to respond promptly to feedback. This lack of flexibility in the design process forces designers to spend significant time on subsequent optimizations after making modifications. This delayed feedback reduces overall design efficiency. Structurally, these tools often utilize a single generation path and lack mechanisms for dynamic adjustments based on feedback. Summary of the Invention

[0007] The purpose of the present invention is to provide a method and system for incremental generation of clothing design based on the DiT architecture, which can achieve gradual refinement and adjustment of the design through the combination of phased incremental generation, conditional control and multi-path diffusion, thereby significantly improving the shortcomings of traditional design tools.

[0008] To achieve the above object, the present invention provides the following solutions:

[0009] A method for incremental generation of clothing designs based on the DiT architecture, including:

[0010] Selecting target elements in a database to determine initial design requirements; the database includes a design style model library, a user preference library, and a material library;

[0011] Based on the initial design requirements, the dialog box module of the design prompt word builder is used to automatically generate AI prompt words for the design drawing; the AI ​​prompt words reflect the original intention, background, design style, user preferences and clothing materials of the design;

[0012] Based on the AI ​​prompt word, a global image generator is used to perform text word meaning analysis and encoding, and output semantic encoding information with global context; the global image generator has a built-in Transformer model; the Transformer model processes the input text through word segmentation, word embedding, and encoding based on word meaning context;

[0013] According to the designer's needs, global design image generation or local optimization is determined to determine the final output design image; when it is decided to generate a global design image, the first processing flow is executed and the diffusion model is used for optimization and adjustment. When it is decided to perform local optimization, the second processing flow is executed and the DiT model is used for optimization and adjustment.

[0014] Optionally, the style model library stores classifications, characteristics, examples and related attribute information of different styles; the user preference library stores historical design data and preference information corresponding to user types, and supports personalized recommendations and dynamic updates; the material library stores the characteristics, applicable scenarios and supplier information of various materials involved in clothing design.

[0015] Optionally, the generating a corresponding clothing design drawing according to the semantic coding information using a diffusion model specifically includes:

[0016] The semantically encoded information is used as the conditional input of the diffusion model. The model then uses the initial Gaussian noise image as a starting point and combines the design requirement vector with the initial noise through conditional mapping, aligning the potential features in the noise with the design requirement. This guides image generation in a specific direction, completing conditional injection.

[0017] After the injection conditions, conditional control and image constraints are performed on each stage of image generation, and denoising is performed in each diffusion step to obtain the corresponding clothing design image.

[0018] Optionally, the first processing flow specifically includes:

[0019] Generate a corresponding clothing design drawing according to the semantic coding information using a diffusion model;

[0020] The clothing design drawing is provided to the designer in high-definition image through GEI's image display module, and a feedback interface is provided. If the designer's feedback is not satisfactory, the clothing design drawing is adjusted and optimized to determine the final output design image; the adjustment and optimization process includes editing the modification area, generating the modification area mask and local iterative design.

[0021] Optionally, the process of editing and modifying the area includes:

[0022] Designers use the "Free Selection Tool" in the interactive image editing module to draw irregular polygonal or rectangular areas on the image to select the parts that need to be modified. It supports drawing single or multiple discontinuous selections to facilitate local optimization. During operation, designers can use the mouse or stylus to draw on any part of the image, and the selected area can be a pattern, material texture, color block or cropping detail. The system will automatically identify and generate the boundary lines of the selection to ensure accuracy, and allow designers to adjust the shape and size of the selection by dragging the edges for precise control. At the same time, it supports undo and redraw functions to quickly correct unsatisfactory selections.

[0023] The process of generating the modified area Mask includes:

[0024] After the designer selects and adjusts the area to be modified, the system automatically generates a mask with the selected area in white and the overall image in black. This mask is used to precisely mark the modified area and assist in subsequent optimization. A mask is a mask layer that covers the image, ensuring that the modification operation only affects the selected area without affecting other parts. When the designer completes the selection of a local area, the system automatically generates a mask with the same shape as the selection to isolate and mark the area. The generated mask is separated from the background image, and the designer can operate on it independently to ensure that the modification is limited to the selected area.

[0025] The generated mask image and the original design drawing will be saved together in the continuous design timing library for subsequent multiple continuous and precise design work;

[0026] The local iterative design process includes:

[0027] After the mask image and the original design drawing are saved in the continuous design timing library, the designer will repeat the previous steps of "generating a global design image or performing local optimization based on the designer's needs to determine the final output design image."

[0028] Optionally, the second processing flow specifically includes:

[0029] Based on the semantic coding information, the DiT model is used to generate an image for the designed local area and seamlessly merge it with the original image; the processing process of the DiT model includes: adding noise, image merging, local generation condition injection, local generation condition control, local image feature sampling, local image feature recovery and output image.

[0030] Optionally, the process of adding noise includes:

[0031] When generating a locally optimized image, the system first adds Gaussian noise to both the original image and the masked area of ​​the local optimization. This step provides an initial state for the subsequent image generator, which gradually constructs an image that meets the designer's requirements from the noise. To ensure that the generated image retains global information while focusing on local optimization, the system independently adds noise to the original image and the local area covered by the mask. The intensity of the noise injection is controlled by preset parameters to ensure that the clarity of the unmodified area is not disturbed, while providing sufficient freedom for adjustment and optimization in the generation of the local image.

[0032] The image merging process includes:

[0033] After adding noise to the original image and the local mask area respectively, the system merges the encoding information of the two to generate a complete input image encoding;

[0034] The process of local generation condition injection includes:

[0035] Based on the design requirement semantics encoded by Transformer, the system first extracts the designer's local optimization requirements through the design requirement encoding, and combines the requirements with the local mask region encoding to generate a conditional local generation vector;

[0036] The process of controlling the local generation conditions includes:

[0037] Through a multi-level condition control mechanism, the system continuously re-injects and adjusts the design conditions at each step of diffusion generation to ensure that the generation of local areas is consistent with expectations;

[0038] The process of local image feature sampling includes:

[0039] Through sampling technology, the system extracts detailed features such as color, texture, light and shadow in local areas and compares them with the global information of the original image to ensure that local modifications not only meet the designer's needs but also maintain the same visual style as the unmodified parts.

[0040] The process of restoring local image features includes:

[0041] By adopting the inverse steps of the diffusion model and performing boundary smoothing transition processing, the system avoids visual discontinuity or conflict between local modifications and surrounding areas, and gradually restores clear image details from the initial noise.

[0042] The present invention also provides a clothing design incremental generation system based on the DiT architecture, which is applied to the clothing design incremental generation method as described above, comprising: a prompt word constructor, a GEI image displayer, a GEI image editor, a global image generator, and a DiT local image generator;

[0043] Among them, the prompt word constructor is respectively connected to the design style model library, the user preference library, the material library and the global image generator; the global image generator is also respectively connected to the GEI image displayer and the DiT local image generator; the DiT local image generator is also respectively connected to the GEI image displayer and the GEI image editor.

[0044] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0045] The present invention discloses a method and system for incremental generation of clothing designs based on a DiT architecture, the method comprising: selecting a target element in a database to determine initial design requirements; based on the initial design requirements, automatically generating AI prompt words for the design drawing using a dialog box module of a design prompt word builder; based on the AI ​​prompt words, performing text semantic analysis and encoding using a global image generator, and outputting semantic encoding information with a global context; the global image generator having a built-in Transformer model; the Transformer model processing input text through word segmentation, word embedding, and encoding based on word semantic context; determining whether to perform global design image generation or local optimization based on designer requirements, and determining the final output design image; when deciding to perform global design image generation, executing a first processing flow and performing optimization adjustment using a diffusion model; when deciding to perform local optimization, executing a second processing flow and performing optimization adjustment using a DiT model. The present invention can achieve gradual refinement and adjustment of the design through a combination of phased incremental generation, conditional control, and multi-path diffusion, thereby significantly improving the shortcomings of traditional design tools. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0047] Figure 1 It is a logic diagram for the overall implementation of the present invention;

[0048] Figure 2 Schematic diagram of the system implementation mechanism structure in this embodiment. DETAILED DESCRIPTION

[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0050] The purpose of the present invention is to provide a method and system for incremental generation of clothing design based on the DiT architecture, which can achieve gradual refinement and adjustment of the design through the combination of phased incremental generation, conditional control and multi-path diffusion, thereby significantly improving the shortcomings of traditional design tools.

[0051] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0052] like Figure 1-Figure 2 As shown, the present invention provides a method for incremental generation of clothing designs based on the DiT architecture. The specific processing steps of each step are as follows:

[0053] Step 1. Choose a design style

[0054] Before AI generates complex design drawings, fashion designers can first select the clothing style, user preferences, and materials to be designed from the design material model library to facilitate the rapid generation of design requirements. Designers interact with the system through the design prompt word constructor. First, they select a style from the design style classification provided by the system, such as modern style, retro style, street style, or high-end fashion style. They can also quickly search for a specific style using keywords. The system will automatically recommend style options that match the designer's preferences based on their historical design data, and recommend relevant design elements such as color, version, and cut based on the user preference model. Next, designers can select clothing materials from the material library, such as cotton, silk, wool, and leather, and filter and select based on material properties such as breathability and elasticity. After selecting the design style, user preferences, and materials, the system combines this information into structured design prompt words to guide the subsequent AI design generation process.

[0055] (1) Design style model library. This stores the classification, characteristics, examples, and related attribute information of different styles to help designers quickly find the right style for their designs. The style_model data structure is as follows:

[0056] style_id: The unique identifier of the style, used to uniquely identify each design style.

[0057] style_name: The name of the style, such as "modern style", "retro style", etc.

[0058] style_category: The category to which the style belongs, such as "street fashion", "haute couture", "minimalism", etc.

[0059] description: A brief description of the style, providing textual information to help designers understand the style characteristics.

[0060] examples: Contains sample images and brief descriptions of the style for designers to refer to.

[0061] Attributes: Key characteristics of a design style, including color, shape, pattern, and target audience (e.g., men, women, children, etc.).

[0062] related_styles: Style IDs related to this style, used to provide recommendations for similar styles.

[0063] (2) User preference library. This stores historical design data and preference information corresponding to user types, supporting personalized recommendations and dynamic updates. The preferred_model data structure is as follows:

[0064] user_type: A unique identifier for each user type, used to associate preferences with user type information.

[0065] preferred_styles: A list of styles that the user prefers, including the style ID and the designer's preference for the style.

[0066] preferred_colors: The user's preferred colors and their corresponding preference levels, expressed using the color's hexadecimal code.

[0067] preferred_materials: The materials preferred by the user and their degree of preference.

[0068] preferred_shapes: The user's preferred clothing cut, style, or shape, and their preference.

[0069] historical_designs: The user's past purchase history, including the style and materials used in the design.

[0070] recent_interactions: The user's recent interactions with the system, including styles and materials viewed, liked, or used.

[0071] (3) Material library. This stores the characteristics, applicable scenarios, and supplier information of various materials involved in clothing design, helping designers select materials that meet their needs. The material_model data structure is as follows:

[0072] material_id: Unique identifier of the material.

[0073] material_name: The name of the material, such as "cotton", "silk", "wool", etc.

[0074] material_category: Material classification, such as "natural materials", "synthetic materials", etc.

[0075] description: A brief description of the material, outlining its characteristics and scope of application.

[0076] Properties: The physical properties of materials, including breathability, elasticity, durability, texture, gloss and other information, help designers choose suitable materials.

[0077] usage_scenarios: The scenarios or types of clothing the material is suitable for, such as "summer wear", "formal wear", "sports wear", etc.

[0078] environmental_impact: The environmental properties of the material, including sustainability and recyclability.

[0079] color_options: The color options available for this material, using hexadecimal color encoding.

[0080] cost_per_meter: The cost of material per meter, usually a decimal number for easy cost calculation.

[0081] supplier_info: Relevant information of the material supplier, including supplier name and contact information, to facilitate subsequent material procurement.

[0082] Step 2. Generate design requirements

[0083] Designers enter their initial design requirements in natural language in the chatbox module of the design prompt builder. Combined with the design style selected in step 1, AI prompts are automatically generated for the design drawing. The prompts will reflect detailed design requirements, including the original intention, background, design style, user preferences, and clothing materials. The structured field structure of the design requirement prompt, design_prompt, is as follows:

[0084] (1) style

[0085] style_type (style type): Specific clothing design style type, such as "minimalism", "retro style", "street fashion", etc.

[0086] style_elements: A list of specific elements involved in the style, such as "geometric pattern", "symmetrical design", "futuristic sense", etc.

[0087] (2)garment_type (clothing type)

[0088] garment_name (garment name): the type of garment to be designed, such as "dress", "jacket", "shirt", etc.

[0089] occasion: The occasion the garment is suitable for, such as "dinner party", "daily leisure", "formal occasion", etc.

[0090] (3) material_texture (material and texture)

[0091] material_type: The type of material used for the garment, such as "cotton", "silk", "wool", "leather", etc.

[0092] material_properties: A list of material properties, such as "breathability", "elasticity", "glossiness", "lightness", etc.

[0093] (4) color_tone (color and tone)

[0094] primary_color: The main color of the clothing, such as "blue", "red", etc.

[0095] secondary_color: The secondary auxiliary color, which may be a matching color or an accent color.

[0096] color_mood: The overall emotion or atmosphere brought by the color tone, such as "soft", "bright", "cool", "warm", etc.

[0097] (5) pattern_decoration (pattern and decoration)

[0098] pattern_type: The type of pattern used on the garment, such as "stripes", "checkered", "floral", etc.

[0099] decoration_elements: Decorative details on clothing, such as "embroidery", "tassels", "sequins", "buttons", etc.

[0100] (6) cut_silhouette (cutting and contouring)

[0101] fit_type: The style or fit of the garment, such as "slim," "loose," "A-line," etc.

[0102] silhouette_features: The overall silhouette features of the garment, such as "high waist", "drape", "layering", etc.

[0103] structural_details: Structural details in clothing design, such as "V-neck", "slit", "bat sleeves", etc.

[0104] (7) functional_requirements

[0105] Seasonality: The season to which the clothing is suitable, such as "spring", "summer", "autumn and winter", etc.

[0106] functional_properties: functional requirements of clothing, such as "warmth", "waterproof", "breathable", "lightweight", etc.

[0107] target_audience: The target audience of the clothing design, such as "male", "female", "children", "neutral", etc.

[0108] (8)overall_style_mood (overall style and mood)

[0109] mood_inclination: The overall emotion or atmosphere conveyed by the clothing, such as "romantic", "elegant", "cool", "casual", etc.

[0110] theme_concept (theme concept): the theme concept of clothing design, such as "natural style", "technological future", "retro fashion", etc.

[0111] (9)special_requirements (special requirements)

[0112] Sustainability: Environmental or sustainable requirements for clothing, such as "environmentally friendly materials", "renewable resources", "non-toxic dyes", etc.

[0113] production_technique: The specific production technique used to make the garment, such as "handmade," "laser cutting," or "3D printing."

[0114] (10) brand_culture (Brand and culture)

[0115] brand_positioning: The positioning of the brand of the clothing, such as "high-end", "luxury", "affordable fashion", etc.

[0116] cultural_background: The cultural background of the design reference, such as "Eastern elements", "European and American classics", "African tribal style", etc.

[0117] The content of each field is described in natural language, but the overall design requirement prompt word design_prompt can be structured. The design requirement prompt word will be transmitted to the global image generator through the API in JSON format to control the generation of design renderings.

[0118] Input: design style model style_m (JSON), user preference model preferred_m (JSON), clothing material model material_m (JSON).

[0119] Output: design requirement prompt word design_prompt (JSON).

[0120] Step 3. Transformer encoding

[0121] After the design requirements submitted by the designer (structured prompt words organized in JSON format) are received by the global image generator, the system first needs to perform semantic analysis and encoding on these texts. To ensure that the design requirements can be correctly interpreted and used to generate images that meet the expectations, the Transformer model processes the input text through word segmentation, word embedding, and encoding based on word semantic context. This process will use the global shared vector space design_global_vector, that is, the vector space pre-trained by a deep learning model (such as BERT or CLIP) to ensure that each word or phrase has a unified semantic representation in the image generation task. The specific sub-steps include:

[0122] (1) Word segmentation. The input design requirements are structured data in JSON format, which contains the designer's requirements (such as design style, material, color, tailoring, decoration, etc.). Word segmentation is to split the design requirement text into basic words or subword units (tokens), which will serve as input for subsequent word embedding and encoding. The Transformer model used in the present invention adopts subword tokenization (Subword Tokenization), such as BPE (Byte-Pair Encoding) or WordPiece algorithm, to decompose words into smaller units, especially for uncommon words.

[0123] Input: Design requirement prompt word design_prompt (JSON).

[0124] Output: design requirement tokens design_tokens (string array).

[0125] (2) Word embedding. The tokens after word segmentation will be mapped to the shared vector space of clothing design images, which is established in advance through pre-trained models (such as BERT, GPT, CLIP, etc.). It is a multimodal vector space trained through a large-scale pre-trained corpus (such as image-text pairs), so that images and text can be matched in the same semantic space. Each word, phrase or concept can be represented by a vector, and similar semantics will be mapped to adjacent positions.

[0126] This paper uses the CLIP framework (Contrastive Language-Image Pre-training) based on the Transformer model for word embedding. It aligns text and image generation and can map design requirement text into a vector space related to image generation.

[0127] Input: design requirement tokens design_tokens (string array).

[0128] Output: design requirement word embedding design_embeddings (32-bit floating point matrix).

[0129] (3) Context encoding. In order to better understand the exact semantics of each design word in the design requirement description, the present invention uses a multi-head self-attention mechanism to calculate the relationship between each design word and other words in the sequence. The context of each design word will be used to adjust its embedding vector so that it contains relevant information about other words in the sentence. For example, when "geometric patterns" and "modern minimalism" are put together, the overall design intention may emphasize the style of "simple and technological" rather than other possible interpretations. The system integrates the word embedding vectors that have undergone context encoding to generate a global 512-dimensional feature vector to represent the semantics of the entire design requirement.

[0130] For example, each part of the design requirements (style, material, color, etc.) will be converted into feature vectors with global context, which will be used in subsequent image generation tasks.

[0131] Input: design requirement word embedding design_embeddings (32-bit floating point matrix); clothing image design shared vector space design_global_vector (32-bit floating point matrix).

[0132] Output: self-attention output attention_output (32-bit floating point matrix); global representation design_representation (512-bit vector).

[0133] Step 4. Diffustion image generation

[0134] At this point, a logical determination is needed to determine whether a global design image should be generated. If so, proceed to this step; otherwise, jump to step 9. In this step, we will use the DiffusionModel to generate the clothing design images required by the designer. The basic principle of this model is to gradually generate high-quality images that meet the design requirements from noise through a reverse diffusion process. Specifically, we will combine the semantic encoding information (such as style, material, color, etc.) from the third step with the shared vector space of the design image, design_global_vector, and use a conditional control mechanism to guide the diffusion model to gradually generate the image.

[0135] The key sub-steps involved in this step include: conditional injection, conditional control, reverse diffusion, and output image. The following is a detailed description.

[0136] (1) Conditional injection. The generation process of the diffusion model requires the input of initial noise, but in order to make the generated image meet the specific needs of the designer, conditional injection is required. First, the semantic vector design_representation (containing elements such as design style, clothing type, color, material, etc.) generated by Transformer encoding in step 3 is used as the conditional input of the diffusion model; then, the model uses the initial Gaussian noise image as the starting point and combines the vector of design requirements with the initial noise through conditional mapping, so that the potential features in the noise are consistent with the design requirements, thereby guiding the image generation to develop in a specific direction. For example: If the designer's requirements are "street style, black jacket, with geometric patterns", then these conditions will be injected into the initial noise of the model to ensure that the final generated image meets these characteristics.

[0137] Input: global semantic vector design_representation, the 512-dimensional feature vector from step 3, which contains the semantic vector space expression of all the designer's design requirements.

[0138] Output: conditional vector condition_vector (512-dimensional feature vector) is used to control each step of the generation process; initial noise image initial_noise (32-bit floating point matrix), Gaussian noise image of the input diffusion model.

[0139] (2) Condition control. In order to ensure that the diffusion process generates images that meet the design requirements, condition control is required at each stage. This is achieved specifically through multi-level condition control. In the process of gradually generating images, the diffusion model will reintroduce design conditions at each step (T step) to ensure that the generated images meet the requirements, such as "street style" or "black jacket". At the same time, the model uses the self-attention mechanism to capture the dependencies between design requirements and coordinate the associations between elements such as clothing color and material. In addition, the model will also impose image constraints to ensure that the generated images meet functional requirements such as material texture and color matching while satisfying visual aesthetics. For example: When the diffusion model generates an image in the middle stage (that is, the middle stage of the T step), the system will again check whether the generated part of the image has a geometric pattern and ensure that the material texture meets the "leather" texture selected by the designer.

[0140] Input: condition vector condition_vector (512-dimensional feature vector), initial noise image initial_noise.

[0141] Output: Diffusion step diffusion_t (integer), progress information of each step in the diffusion process (from T to 0).

[0142] (3) Reverse diffusion. The core mechanism of the diffusion model is reverse diffusion, which is to generate a clear image by gradually removing noise. First, the model starts with a completely random Gaussian noise image, and in each diffusion step (from T to 0), it gradually removes noise according to conditional control and design requirements, making the image clearer and clearer. As the denoising progresses, the model continues to integrate design elements: the early stage generates basic outlines, such as the overall shape and pattern of the clothes; the mid-stage refines visual details such as color and material texture; the later stage adds details such as pattern, decoration, and tailoring. Ultimately, reverse diffusion generates a high-quality image that contains all design requirements, including style, material, color, pattern, and decoration.

[0143] Input: Diffusion step diffusion_t.

[0144] Output: generated image generated_image, the final generated clothing design image.

[0145] (4) Output image. After reverse diffusion, the system will output the final design image, which will be saved and presented to the designer for further adjustment or confirmation. The image has the following characteristics: First, it meets the design requirements entered by the designer in step 3, including elements such as style, color, and material; second, the refinement of the diffusion model makes the image have a resolution of 300DPI or higher, which is suitable for subsequent design decisions or presentations; finally, the generated image file will be transferred to the GEI image displayer through the API. The designer can further adjust the generated results or request the system to generate new design drawings to expand the design inspiration.

[0146] Input: generated image generated_image.

[0147] Output: preliminary design rendering file initial_design_file, which can be in standard image format such as PNG, JPEG, etc., with a resolution of 300DPI or higher.

[0148] Step 5. Display the design

[0149] The garment design file (initial_design_file) generated in step 4 is provided to designers via GEI's Image Display Module (GV). The system also provides a feedback interface where designers can provide immediate feedback on the generated image (e.g., "like" or "dislike," with additional comments).

[0150] If the designer is satisfied with the design at this point, the design process ends. If not, especially with the design details, then step 6 is entered to further adjust and optimize the design. Of course, if the designer is not satisfied with the overall design, step 6 is not required, and step 1 is repeated to generate a new clothing design.

[0151] Input: preliminary design rendering file initial_design_file or final_design_file.

[0152] Output: interface interaction.

[0153] Step 6. Edit the modification area

[0154] Designers use the "Free Selection Tool" in the Image Editing Interaction Module (GEI) to draw irregular polygons or rectangles on the image, accurately selecting the parts that need to be modified. It supports the drawing of single or multiple discontinuous selections to facilitate local optimization. During operation, designers can use a mouse or stylus to draw on any part of the image. The selected area can be a pattern, material texture, color block, or visual elements such as cropping details. The system will automatically identify and generate the boundary lines of the selection to ensure accuracy, and allow designers to adjust the shape and size of the selection by dragging the edges for precise control. At the same time, it supports undo and redraw functions, which are convenient and quick to correct unsatisfactory selections.

[0155] Input: preliminary design rendering file initial_design_file.

[0156] Output: interface interaction.

[0157] Step 7. Generate the modified area mask

[0158] After the designer selects and adjusts the area that needs to be modified, the system will automatically generate a mask with the selected area in white and the overall image in black with a white background. This is used to accurately mark the modified area and assist in subsequent optimization processing. Mask is a mask layer covering the image to ensure that the modification operation only acts on the selected area without affecting other parts. When the designer completes the selection of the local area, the system will automatically generate a mask with the same shape as the selection to isolate and mark the area. The generated mask is separated from the background image, and the designer can operate it separately to ensure that the modification is limited to the selected area. The entire mask generation process supports high-precision adjustment to ensure the accuracy of complex details (such as irregular edges or gradient transitions).

[0159] The generated mask image and the original design drawing will be saved together in the continuous design timing library for subsequent multiple continuous and precise design tasks.

[0160] Input: preliminary design rendering file initial_design_file.

[0161] Output: mask image file mask_design_file with black background and white area.

[0162] Step 8. Local iterative design

[0163] After the mask image and the original design are saved to the continuous design timing library, the designer will repeat the above design steps 1, 2, and 3. A new design requirement prompt word design_prompt and a global vector design_representation are generated for the selected area of ​​the design drawing for local optimization.

[0164] Step 9. DiT local image generation

[0165] After the designer's requirements are encoded via the Transformer (step 3), the system will determine whether to perform local optimization (step 9) or global design image generation (step 4) based on the requirements. If the requirement is local optimization, the system will proceed to step 9, using the Diffusion Transformer (DiT) model to generate an image for the local area of ​​the design and seamlessly integrate it with the original image. Specific sub-steps include: adding noise, image merging, injecting local generation conditions, controlling local generation conditions, sampling local image features, restoring local image features, and outputting the image.

[0166] (1) Adding noise

[0167] In the process of generating a locally optimized image, the system first adds Gaussian noise to the original image and the local optimized area (the same as the mask generation process). This step provides an initial state for the subsequent image generator, enabling it to gradually construct an image that meets the designer's requirements from the noise. In order to ensure that the generated image retains global information while focusing on local optimization, the system independently adds noise to the original image and the local area covered by the mask. The intensity of the noise injection is controlled by preset parameters to ensure that the clarity of the unmodified area is not disturbed, while providing sufficient freedom for adjustment and optimization of the local image generation.

[0168] Input: initial design image initial_design_file, mask image of local optimization area mask_design_file.

[0169] Output: the original image with noise noise_initial_image, the mask area image with noise noise_mask_image.

[0170] (2) Image merging

[0171] After adding noise to the original image and the local mask region, the system merges their encoding information to generate a complete input image encoding. The merged encoding retains the overall feature information of the global image while also incorporating the specific features of the locally optimized region. By combining the global image encoding with the local mask region encoding, the system creates a comprehensive image representation (merged_image_encoding). This encoding ensures that in the subsequent generation process, the overall image consistency is maintained while the local region can be refined and optimized.

[0172] Input: Noisy original image noised_initial_image, noisy Mask area image noised_mask_image.

[0173] Output: merged_image_encoding, a comprehensive image representation containing global image information and local optimized area information - a 32-bit floating point matrix.

[0174] (3) Local generation condition injection

[0175] During the local image generation process, the system injects the designer's specific modification requirements into the local generation conditions based on the design requirement semantics encoded by the Transformer to ensure that the details of the generated image meet expectations. This process involves fine-grained control of elements such as the style, color, material, and cropping of the local area. The system first extracts the designer's local optimization requirements (such as color adjustment, material change, or decorative addition) through the design requirement encoding, and combines these requirements with the local Mask area encoding to generate a conditional local generation vector. This conditional vector will guide the image generator to gradually implement the modifications specified by the designer in the subsequent generation steps to ensure that the final image can reflect the key elements in the requirements and is consistent with the overall design style.

[0176] Input: Local design semantic vector design_representation, the 512-dimensional feature vector from step 3, contains the semantic vector space representation of the designer's design requirements for the current local area. This includes the designer's specific modification requirements, such as style, color, material, and cut.

[0177] Output: Conditional local generation vector (local_condition_vector): A 512-bit vector containing the local generation condition vector of the design requirements, which combines the designer's requirements and local area features to guide the local image generation process.

[0178] (4) Local generation condition control

[0179] At each stage of image generation, the system continuously monitors and adjusts local generation conditions to ensure that the generated results align with the designer's requirements. Through a multi-level conditional control mechanism, the system precisely manages local details at every step of the generation process. Specifically, the system continuously reinjects and adjusts design conditions at each step of the diffusion generation process (from initial noise to the reverse diffusion process of gradually generating a clear image) to ensure that the generation of local areas is consistent with expectations.

[0180] For example, if a designer requests the addition of geometric patterns or changes to material textures in a local area, the system will gradually introduce these features through various generation stages to ensure they are accurately reflected in the final image. This gradual control mechanism not only ensures the accuracy of local optimization but also maintains the consistency of the overall design, allowing local modifications to blend naturally with unmodified areas.

[0181] Input: Condition vector local_condition_vector (512-dimensional feature vector).

[0182] Output: Diffusion step diffusion_t (integer), progress information of each step in the diffusion process (from T to 0).

[0183] (5) Local image feature sampling

[0184] During the local image generation process, the system samples the feature encoding of the original image and the feature encoding of the local mask area to ensure that the generated local image is consistent with the overall style and forms a natural transition with the unmodified area. Through sampling technology, the system extracts detailed features such as color, texture, light and shadow in the local area and compares them with the global information of the original image, ensuring that the local modification not only meets the designer's needs but also maintains a consistent visual style with the unmodified part.

[0185] This sampling process is crucial, ensuring that the local region retains the designer's intended modification during generation while blending with the overall style of the image, preventing disconnection between the local region and the global image. The resulting local feature sampling vectors guide the natural transition and consistency of the local image generation stage.

[0186] Input: Original image feature encoding (initial_design_features): global feature information extracted from the original design image, including visual elements such as color, texture, light and shadow; Local Mask area feature encoding (mask_design_features): feature encoding of the local optimized area, including the color, material, light and shadow information of the area.

[0187] Output: local feature sampling vector local_feature_vector: combines the sampling results of local area features and global image style to guide style consistency and natural transition during local generation.

[0188] (6) Local image feature recovery

[0189] During local image feature recovery, the system gradually removes noise, restoring the local area to a high-quality image that meets the designer's requirements. This process uses the reverse steps of the diffusion model to gradually recover clear image details from the initial noise. In this process, the system not only restores the local image's color, texture, and lighting characteristics, but also ensures that the local area seamlessly blends with the original image boundary. By smoothing the boundary transition, the system avoids visual disconnection or conflict between the local modification and the surrounding area. This process still requires the support of training data from the design image shared vector space design_global_vector to ensure the authenticity of the image feature recovery.

[0190] Ultimately, the generated denoised local image is naturally connected to the global image, and the system outputs a fused image to ensure that the locally optimized area is consistent with the overall design style and meets the designer's expectations.

[0191] Input: Local feature sampling vector local_feature_vector: A feature vector that combines local and global information generated from local image feature sampling.

[0192] Original noise image noised_initial_image, Mask local noise image noised_mask_image.

[0193] Local generation condition vector local_condition_vector: the designer's modification requirements and local generation conditions, used to guide the noise removal process.

[0194] Output: denoised_local_image, a high-quality local image generated after gradual denoising that meets design requirements and seamlessly blends with the original image. Final_blended_image, the final output complete image, smoothly transitions between the modified local area and the global image, achieving visual harmony and consistency.

[0195] (7) Output image

[0196] Once the locally optimized image is generated, the system outputs the final design blended image, final_blended_image, which includes the designer's modifications to the local area. The system combines the locally optimized image with the original image to generate the final image file. The output image meets the designer's local optimization requirements, with high resolution and complete detail, making it suitable for subsequent presentation, feedback, or further modification.

[0197] The final garment design will be exported to the GEI Image Viewer (GV) via the API for visual presentation. Designers can repeat step 5 to continuously and accurately create a design that satisfies them.

[0198] Input: fused image final_blended_image.

[0199] Output: Final design rendering file (final_design_file). This file can be in standard image formats such as PNG and JPEG, with a resolution of 300 DPI or higher.

[0200] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0201] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the core concept of the present invention. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.

Claims

1. A clothing design incremental generation method based on DiT architecture, characterized by: include: Selecting target elements in a database to determine initial design requirements; the database includes a design style model library, a user preference library, and a material library; Based on the initial design requirements, the dialog box module of the design prompt word builder is used to automatically generate AI prompt words for the design drawing; the AI ​​prompt words reflect the original intention, background, design style, user preferences and clothing materials of the design; Based on the AI ​​prompt word, a global image generator is used to perform text word meaning analysis and encoding, and output semantic encoding information with global context; the global image generator has a built-in Transformer model; the Transformer model processes the input text through word segmentation, word embedding, and encoding based on word meaning context; Determine whether to generate a global design image or perform local optimization based on the designer's needs, and determine the final output design image; when it is decided to generate a global design image, execute the first processing flow and use the diffusion model for optimization and adjustment; when it is decided to perform local optimization, execute the second processing flow and use the DiT model for optimization and adjustment; The second processing flow specifically includes: Based on the semantic coding information, the DiT model is used to generate an image for the designed local area and seamlessly merge it with the original image; the processing process of the DiT model includes: adding noise, image merging, local generation condition injection, local generation condition control, local image feature sampling, local image feature recovery and output image; The process of adding noise includes: When generating a locally optimized image, the system first adds Gaussian noise to both the original image and the mask of the locally optimized area. The intensity of the noise injection is controlled by preset parameters. This step provides an initial state for the subsequent image generator, which gradually constructs an image that meets the designer's requirements from the noise. This ensures that the clarity of the unmodified area is not disturbed, while providing sufficient freedom for adjustment and optimization of the local image generation. The image merging process includes: After adding noise to the original image and the local mask area respectively, the system merges the encoding information of the two to generate a complete input image encoding; The process of local generation condition injection includes: Based on the design requirement semantics encoded by Transformer, the system first extracts the designer's local optimization requirements through the design requirement encoding, and combines the requirements with the local mask region encoding to generate a conditional local generation vector; The process of controlling the local generation conditions includes: Through a multi-level condition control mechanism, the system continuously re-injects and adjusts the design conditions at each step of diffusion generation to ensure that the generation of local areas is consistent with expectations; The process of local image feature sampling includes: Through sampling technology, the system extracts the color, texture, and light and shadow of the local area and compares it with the global information of the original image, ensuring that the local modification not only meets the designer's needs but also maintains the same visual style as the unmodified part; The process of restoring local image features includes: By adopting the inverse steps of the diffusion model and performing boundary smoothing transition processing, the system avoids visual discontinuity or conflict between local modifications and surrounding areas, and gradually restores clear image details from the initial noise.

2. The incremental generation method of clothing design based on DiT architecture according to claim 1 is characterized in that: The style model library stores classifications, characteristics, examples and related attribute information of different styles; The user preference library stores historical design data and preference information corresponding to user types, and supports personalized recommendations and dynamic updates; The material library stores the characteristics, applicable scenarios and supplier information of various materials involved in clothing design.

3. The incremental generation method of clothing design based on DiT architecture according to claim 1 is characterized in that: The first processing flow specifically includes: Generate a corresponding clothing design drawing according to the semantic coding information using a diffusion model; The clothing design drawing is provided to the designer in high-definition image through GEI's image display module, and a feedback interface is provided. If the designer's feedback is not satisfactory, the clothing design drawing is adjusted and optimized to determine the final output design image; the adjustment and optimization process includes editing the modification area, generating the modification area mask and local iterative design.

4. The incremental generation method of clothing design based on DiT architecture according to claim 3 is characterized in that: The step of generating a corresponding clothing design drawing according to the semantic coding information using a diffusion model specifically includes: The semantically encoded information is used as the conditional input of the diffusion model. The model then uses the initial Gaussian noise image as a starting point and combines the design requirement vector with the initial noise through conditional mapping, aligning the potential features in the noise with the design requirement. This guides image generation in a specific direction, completing conditional injection. After the injection conditions, conditional control and image constraints are performed on each stage of image generation, and denoising is performed in each diffusion step to obtain the corresponding clothing design image.

5. The incremental generation method of clothing design based on DiT architecture according to claim 3 is characterized in that: The process of editing and modifying a region includes: Designers use the "Free Selection Tool" in the interactive image editing module to draw irregular polygonal or rectangular areas on the image to select the parts that need to be modified. It supports drawing single or multiple discontinuous selections, facilitating local optimization. During operation, designers can use a mouse or stylus to draw on any part of the image, selecting areas such as patterns, material textures, color blocks, or cropping details. The system automatically identifies and generates selection boundaries to ensure accuracy, and allows designers to adjust the shape and size of the selection by dragging the edges for precise control. It also supports undo and redraw functions, making it easy to quickly correct unsatisfactory selections. The process of generating the modified area Mask includes: After the designer selects and adjusts the area to be modified, the system automatically generates a mask with the selected area in white and the overall image in black. This mask is used to precisely mark the modified area and assist in subsequent optimization. A mask is a mask layer that covers the image, ensuring that the modification operation only affects the selected area without affecting other parts. When the designer completes the selection of a local area, the system automatically generates a mask with the same shape as the selection to isolate and mark the area. The generated mask is separated from the background image, and the designer can operate on it independently to ensure that the modification is limited to the selected area. The generated mask image and the original design drawing will be saved together in the continuous design timing library for subsequent multiple continuous and precise design work; The local iterative design process includes: After the mask image and the original design drawing are saved in the continuous design timing library, the designer will repeat the steps before "determining global design image generation or local optimization based on designer needs to determine the final output design image." 6. A clothing design incremental generation system based on the DiT architecture, applied to the clothing design incremental generation method according to any one of claims 1 to 5, characterized in that: include: Prompt word builder, GEI image displayer, GEI image editor, global image generator and DiT local image generator; Among them, the prompt word constructor is respectively connected to the design style model library, the user preference library, the material library and the global image generator; the global image generator is also respectively connected to the GEI image displayer and the DiT local image generator; the DiT local image generator is also respectively connected to the GEI image displayer and the GEI image editor.

Citation Information

Patent Citations

  • Garment secondary design system and method

    CN117576246A

  • Multi-modal data driven generation type fashion compatible costume design method and system

    CN117951763A

  • Virtual fitting method of deep learning 2D picture

    CN118505835A