A text-driven intelligent method for generating clothing patterns
By building a multi-module collaborative clothing model intelligent generation model and learning-based three-dimensional virtual fitting technology, the problems of low efficiency of clothing model generation and difficulty in dynamic optimization are solved, and the efficient, automated and personalized design of clothing model is achieved.
Patent Information
- Application Number
- CN202510669798.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-05-23
AI Technical Summary
The existing intelligent generation technology of clothing model is lacking effective design rules and constraints, the generation efficiency is low and the results are difficult to control, and it cannot meet the needs of the modern clothing market for rapid response to diversification and personalization. It is difficult to automate clothing model optimization in dynamic scenarios.
Using text-driven intelligent generation method for clothing templates, we use multi-module collaborative intelligent generation model for clothing templates, combined with learning-based three-dimensional virtual fitting technology, we introduce structured design rules and clothing template parameter configuration files to realize automatic generation and dynamic optimization of clothing template data.
It significantly improves the efficiency and quality of clothing model generation, enhances the stability and diversity of clothing model generation, and can realize the automated optimization of clothing model in dynamic scenarios to meet the design needs of personalized and functional clothing.
Smart Images

Figure CN120197247B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent clothing pattern making, and relates to a method for intelligently generating clothing patterns driven by text. Background Art
[0002] As the core link in clothing production, the quality of clothing pattern design directly affects the fit, comfort and aesthetics of the finished clothing. In the traditional clothing manufacturing field, pattern making mainly relies on manual drawing. This method not only has low efficiency, but also highly depends on the experience accumulation of pattern makers, with high learning costs and long time consumption, making it difficult to meet the rapid response to the diverse and personalized needs of the modern clothing market. Although the application of computer-aided design (CAD) software has improved the pattern-making efficiency to a certain extent, the mode of adjusting based on fixed templates still has obvious limitations in terms of flexibility and creativity realization.
[0003] In recent years, the booming development of artificial intelligence technology has opened up a new path for the intelligent generation of clothing patterns. The text-driven method with natural language processing as the core makes it possible to generate complex clothing patterns through intuitive descriptive text. SewingGPT proposed in the literature (DressCode: Autoregressively Sewing and Generating Garments from Text Guidance, ACM Trans. Graph., 2024.) first explored the technology of generating clothing patterns guided by descriptive text and verified the feasibility of this technical route. However, the generated clothing styles are relatively simple, and due to the lack of effective rule constraints, the effectiveness and quality of the generated results need to be improved.
[0004] Similarly, the patent application with the publication number CN119337442A provides a method for generating clothing pattern data. It extracts and decodes features through pattern geometric data, but does not introduce systematic pattern design rule constraints. Due to the complex diversity and lack of standardized expression of clothing pattern geometric data, the generated results are difficult to control and cannot meet the basic pattern design specifications. In addition, during the geometric feature decoding process, the data dimension is high, the training difficulty is large, and problems such as generation failure are likely to occur, seriously reducing the stability and practical value of the generation system.
[0005] In summary, due to the lack of effective design rule constraints, the current intelligent clothing pattern generation technology has problems such as low generation efficiency and difficult result control, and it is urgent to explore more effective solutions to promote the continuous development of clothing pattern design technology. Summary of the Invention
[0006] The purpose of the present invention is to solve the above problems existing in the prior art and provide a method for intelligently generating clothing patterns driven by text.
[0007] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0008] A text-driven intelligent clothing pattern generation method inputs the user-defined clothing pattern feature description text D1* into the trained clothing pattern intelligent generation model G (a deep neural network model), and outputs the generated clothing pattern data X2*.
[0009] G consists of a text embedding feature extractor G1, a clothing pattern parallel graph neural network encoder G2, a clothing pattern parallel autoregressive decoder G3, and a parameterized pattern generation program G4. G1 and G2 are both connected to G3, and G3 is connected to G4. G2 is only enabled during the model training process, G4 is only enabled during the model application process, and G1 and G3 are enabled during both the training process and the application process;
[0010] The training of G uses the clothing pattern feature description text D1 in the clothing pattern dataset and the clothing pattern parameter configuration file D2 in the clothing pattern dataset. The process is as follows:
[0011] (a) Input D1 into G1, and it extracts the text embedding feature D3;
[0012] (b) Map D2 through G2 to the clothing pattern latent space to generate the clothing pattern latent space encoding D4;
[0013] (c) Input D3 and D4 together into G3 for decoding to obtain a new clothing pattern parameter configuration file D5;
[0014] (d) Calculate the parameter configuration file loss value The formula is as follows:
[0015]
[0016] In the formula, x i and y i respectively represent the values of the i-th parameter in D5 and D2, and N P represents the total number of parameters in D5 or D2 (the total number of parameters in D5 is the same as the total number of parameters in D2);
[0017] (e) Determine whether the number of iterations in steps (a) to (d) is greater than or equal to 10, and whether it is less than 0.0003. If both are yes, end. Otherwise, update the parameters of G2 and G3 and return to step (a).
[0018] A text-driven intelligent method for generating clothing patterns realizes the automatic generation of high-quality clothing pattern data from the text description of clothing pattern features customized by users by constructing an intelligent clothing pattern generation model G that collaborates multiple modules. The intelligent clothing pattern generation model G consists of a text embedding feature extractor G1, a parallel graph neural network encoder G2 for clothing patterns, a parallel autoregressive decoder G3 for clothing patterns, and a parametric pattern generation program G4, which can convert the text description of clothing pattern features input by users into a clothing pattern parameter configuration file and generate clothing pattern data that meets the requirements.
[0019] The present invention performs feature modeling through a clothing pattern parameter configuration file and incorporates structured design rules using the parametric pattern generation program G4 to effectively constrain the clothing pattern generation process. Compared with directly processing complex geometric data, the clothing pattern parameter configuration file provides a more compact and clearly expressed feature representation method, significantly improving the generation efficiency of clothing pattern data and the controllability of the model. At the same time, by using the parallel graph neural network encoder G2 for clothing patterns and the parallel autoregressive decoder G3 for clothing patterns, the clothing pattern components are partitioned and encoded / decoded in parallel, which not only optimizes the computational efficiency of the intelligent clothing pattern generation model but also enhances the stability and diversity of clothing pattern generation.
[0020] As a preferred technical solution:
[0021] In the above-mentioned text-driven intelligent method for generating clothing patterns, in step (b), G2 includes two parts: quantization encoding and feature encoding. Among them, quantization encoding normalizes the default value of D2 to a standardized value according to the value range, and the encoding dimension is the same as the number of pattern parameters without data compression. Feature encoding maps the quantization encoding to the latent space in parallel according to multiple components to form D4.
[0022] In the above-mentioned text-driven intelligent method for generating clothing patterns, the specific process of step (c) is as follows: First, align D3 and D4, and then decode the aligned encoding into a quantization encoding of clothing pattern parameters in parallel according to multiple components to generate D5.
[0023] In the above-mentioned text-driven intelligent method for generating clothing patterns, X2* is also dynamically optimized, and the process is as follows:
[0024] a) Input X2* and the text description of the action features X1* customized by the user into the trained learning-based 3D virtual fitting model T (a deep neural network model), and use the parameter settings of non-elastic fabrics to output the generated 3D dynamic virtual fitting grid data V1*.
[0025] b) Calculate the structural line offset loss value The formula is as follows:
[0026]
[0027] In the formula, ω1 represents the weight value of the key motion frames in V1*, and N k represents the total number of key motion frames in V1*, represents the difference of the structure lines under the i-th key motion frame in V1*, ω2 represents the weight value of the process action frames in V1*, and N p represents the total number of process action frames in V1*, represents the difference of the structure lines under the j-th process action frame in V1*;
[0028] c) Judge whether the number of iterations in steps a) to b) is greater than or equal to 10 steps, whether it converges to 10% of the initial value, whether it is less than or equal to 5 mm, and whether it is less than or equal to 5 mm. If all are yes, then take the X2* of the last iteration in steps a) to b) as the garment sample data V2* applicable to non-stretch fabrics and proceed to the next step. Otherwise, update X2* and return to step a);
[0029] d) Input V2* into the trained T, and use the real stretch fabric parameters to set, and generate the three-dimensional static virtual fitting grid data V3* from its output;
[0030] e) Calculate the fabric deformation loss value The formula is as follows:
[0031]
[0032] In the formula, N v represents the number of sampled vertices of V3* or V2* (the number of sampled vertices of V3* is the same as that of V2*), p i represents the spatial coordinate of the i-th sampled vertex of V3*, q i represents the spatial coordinate of the i-th sampled vertex of V2*;
[0033] f) Judge whether the number of iterations in steps d) to e) is greater than or equal to 10 steps, and whether it is less than or equal to 5 mm. If all are yes, then take the V2* of the last iteration in steps d) to e) as the garment sample data V4* applicable to real stretch fabrics. Otherwise, update V2* and return to step d);
[0034] e) Fine-tune V4* according to personalized needs and save it.
[0035] In the production of the clothing industry, the optimization and adjustment of clothing patterns is a crucial step. The initial version of the pattern usually needs to be adjusted multiple times before it can be put into production. Traditional optimization methods mainly rely on manual physical fitting and manual modification. This method not only consumes a large amount of time and labor costs, but it is also difficult to accurately evaluate the wearing effect of clothing in dynamic movement scenarios. Especially in the design and development of functional clothing and sports clothing, the dynamic wearing performance directly affects the functionality and comfort of the product, making the limitations of traditional methods more prominent.
[0036] Although existing related research has tried to innovate, for example, the literature (Development of upper cycling clothes using 3D-to-2D flattening technology and evaluation of dynamic wear comfort from the aspect of clothing pressure, International Journal of Clothing Science and Technology, 2016) guides pattern optimization by measuring the clothing pressure in specific sports postures. However, this method can only reflect the clothing state in static postures and cannot comprehensively evaluate the adaptability changes of clothing during dynamic movement. In addition, such research still highly relies on manual experience judgment and is difficult to achieve an automated production process. At the same time, the research generally regards fabric properties as fixed values, ignoring the significant impact of fabric elasticity, a key variable, on clothing pattern design, while fabric elasticity will significantly affect the actual wearing shape of clothing.
[0037] Existing clothing simulation technologies are mostly based on physical simulation fitting programs, and their operation processes are not differentiable, so they cannot automatically adjust the pattern backward according to the fitting results. The patent application with the publication number CN119337442A provides a method for generating clothing panel data. This method has not established a quantitative evaluation system for pattern fitting and cannot optimize fitting in dynamic postures. At the same time, the model does not consider the influence of fabric elasticity on the finished clothing effect, which may lead to pattern wearing deformation, decreased fitting, and damaged aesthetics under different fabrics. Related adjustments still rely on manual experience and manual operations and are difficult to automate, greatly limiting the popularization and application of this method in actual production.
[0038] How to quantitatively optimize clothing patterns in dynamic scenarios to improve the dynamic adaptability of clothing has become one of the key challenges in intelligent clothing pattern generation technology.
[0039] With the rapid development of artificial intelligence technology, the learning-based virtual fitting method has become an important development trend in the field of three-dimensional virtual fitting. The learning-based virtual fitting method uses machine learning algorithms to deeply analyze a large amount of clothing and human interaction data, learn and predict the deformation effect of clothing on the human body surface, and then achieve an efficient virtual fitting experience. The learning-based virtual fitting method has significant computational efficiency advantages. By using a feedforward neural network, it can achieve end-to-end generation of three-dimensional clothing deformation effects, successfully avoiding the complex cloth iterative solution process in traditional physical simulations. In addition, due to the differentiability of its calculation process, it can automatically perform reverse optimization based on the fitting results and adjust the clothing pattern, which opens up a new path for intelligent clothing pattern generation.
[0040] The present invention combines the learning-based three-dimensional virtual fitting technology to establish a quantitative evaluation system for the fit of clothing patterns with the offset of clothing-human structure lines as the index. By introducing dynamic action features, it can detect changes in structure lines under different poses and iteratively optimize them to achieve automatic optimization and adjustment of clothing patterns in the dynamic wearing state. In addition, a loss function is introduced to precisely control the generation quality of clothing patterns, ensuring that the clothing pattern data output by the intelligent clothing pattern generation model is highly consistent with the real data. The optimization system also considers the influence of fabric elasticity on the effect of clothing patterns. By simulating the deformation degree of real elastic fabrics, it further optimizes the clothing pattern structure to ensure fit and consistency under different elastic fabrics, effectively improving the automation level and practical performance of clothing pattern design and expanding the actual application scenarios of clothing pattern generation technology in the clothing industry.
[0041] A text-driven intelligent clothing pattern generation method as described above, T is composed of a dynamic human body generation module T1, a pattern static mapping module T2, a three-dimensional clothing dynamic deformation mapping module T3, and a three-dimensional clothing explicit decoder T4. T1 and T2 are both connected to T3, and T3 is connected to T4;
[0042] In the application process, when the non-elastic fabric parameter setting is adopted, T1, T2, T3, and T4 are all enabled; when the real elastic fabric parameter setting is adopted, T2 is not enabled, and T1, T3, and T4 are all enabled;
[0043] During the training process, T1, T2, T3, and T4 are all enabled;
[0044] Training T uses the action feature description text X1 in the virtual fitting dataset, the three-dimensional dynamic virtual fitting grid data X6 in the virtual fitting dataset, and the clothing pattern data X2 in the clothing pattern dataset. The process is as follows:
[0045] (ⅰ) Input X1 into T1, and it outputs a dynamic three-dimensional virtual human model X3;
[0046] (ii) Convert X2 into three-dimensional clothing static implicit data X4 through T2;
[0047] (iii) Input X3 and X4 into T3 together, and its output is three-dimensional clothing dynamic implicit data X5;
[0048] (iv) Input X5 into T4, and its output is three-dimensional dynamic virtual fitting grid data V1;
[0049] (v) Calculate the chamfer loss value The formula is as follows:
[0050]
[0051] In the formula, A represents the set of sampled vertices of V1, B represents the set of sampled vertices of X6, N a represents the number of sampled vertices in A, N b represents the number of sampled vertices in B, a ∈ A and b ∈ B respectively represent a single vertex of A and B;
[0052] (vi) Determine whether the number of iterations in steps (i) to (v) is greater than or equal to 10 steps, and whether it is less than or equal to 5 mm. If both are yes, end. Otherwise, after updating the parameters of T2 and T3, return to step (i).
[0053] The learning-based three-dimensional virtual fitting model T of the present invention combines dynamic human body generation, sample static mapping and three-dimensional clothing dynamic deformation mapping technologies, uses the action feature description text and the generated clothing sample data provided by the user to generate dynamic or static virtual fitting effects; and through optimization mechanisms such as structural line offset loss and fabric deformation loss to precisely control the generation quality of clothing sample data, ensure that the clothing sample data output by the clothing sample intelligent generation model is highly consistent with the real data, further improve the adaptability of clothing sample generation to different fabric characteristics, and ensure its accurate and reliable performance in real scenarios.
[0054] For a text-driven intelligent clothing sample generation method as described above, the construction steps of the clothing sample data set and the virtual fitting data set are as follows:
[0055] S1.1. Select a three-dimensional virtual human model A1, measure its virtual human model size data A2, set multiple clothing sample parameter configuration files D2 based on A2, input D2 into the parametric sample generation program G4, and generate various styles of clothing sample data X2;
[0056] S1.2. Manually construct multiple action feature description texts X1, use the text-action generation model A3 to generate three-dimensional virtual human action data A4, and import A4 into A1 in step S1.1, thereby constructing multiple dynamic three-dimensional virtual human models X3;
[0057] S1.3. Combine X3 in step S1.2, and use the physical simulation-based fitting program A5 to virtually sew X2 generated in step S1.1 to generate high-precision three-dimensional virtual dynamic fitting grid data X6;
[0058] S1.4. Map X6 in step S1.3 to multi-view images A6, and combine with X2 in step S1.1 to fine-tune the multi-modal large model A7 so that it can recognize the features at the garment and pattern levels and automatically generate the corresponding garment pattern feature description text D1 in batches;
[0059] S1.5. Integrate A2, D2, X2 in step S1.1 and D1 in step S1.4 to construct a garment pattern dataset; integrate X1, A4, X3 in step S1.2 and X6 in step S1.3 to construct a virtual fitting dataset.
[0060] In recent years, text-driven digital clothing generation technology has made certain progress. The literature (AIpparel: a large multimodal generative model for digital garments, arxiv, 2024) realizes the generation of complex clothing styles through fine coding of the garment pattern structure; the literature (Design2GarmentCode: turning design concepts to tangible garments through program synthesis, arxiv, 2024) uses a multi-modal model to generate garment pattern configuration files, improving the rule constraints and success rate of pattern generation. However, existing text-driven garment pattern generation methods still have significant defects: one is that the generation speed of complex-style garment patterns is slow, making it difficult to meet the requirements of rapid design iteration; the other is that existing models are mostly trained based on garment-level description texts, lacking the ability to understand and finely control the features at the pattern level, and unable to accurately adjust and optimize the pattern. For example, it is difficult to execute specific editing instructions such as "lengthen the front center line of the front piece pattern by 2 cm".
[0061] A method for generating garment pattern data provided in a patent application with the publication number CN119337442A trains a model based on descriptive text at the garment level and fails to deeply understand and model the characteristics at the pattern level, resulting in a serious lack of fine control ability in the garment pattern design and adjustment process. For specific instructions in actual garment pattern design and editing, such as "extend the front center line of the front piece pattern by 2 cm", the method provided by this patent application is difficult to accurately parse and execute, greatly limiting its practicality and operation accuracy, and it is difficult to be effectively applied to actual production design scenarios.
[0062] To support the efficient training and application of G and T, the present invention designs a multi-modal data construction process including a garment pattern data set and a virtual fitting data set to ensure data diversity and high precision; when constructing the garment pattern data set, in addition to introducing descriptions at the garment level (such as "loose fit of T-shirt"), refined text information at the pattern level (such as "V-neck depth 5 cm") is also added with the help of a multi-modal large model. The introduction of pattern-level features enables G to deeply understand and master the fine-grained design elements at the garment pattern level, thereby achieving precise control of garment pattern design and significantly improving the flexibility and accuracy of garment pattern generation and adjustment.
[0063] Beneficial effects:
[0064] First of all, by combining multi-modal large model technology and a parallel autoregressive generation model, the present invention can quickly generate high-quality garment pattern data according to the descriptive text of garment pattern features input by users. Compared with traditional pattern-making methods, the method of the present invention greatly improves the generation efficiency and design flexibility of garment patterns, reduces the dependence on professional experience, and meets the needs of the modern garment market for diversified and personalized designs. At the same time, the data set of the present invention contains descriptive text at the garment pattern level, further supporting the precise adjustment of local features of garment patterns, effectively solving the deficiencies in detail control of the existing technology, and providing higher precision and controllability for garment pattern design.
[0065] Secondly, the present invention innovatively integrates text-driven motion generation technology and learning-based virtual fitting technology, can accurately analyze the characteristics of structural line offset and fabric deformation in a dynamic scenario, and automatically adjust the garment pattern data through multi-level reverse iterative optimization, significantly improving the fit and comfort of the garment, especially suitable for the design of functional clothing and sportswear, filling the gaps in dynamic optimization and fabric property adaptability of the existing technology.
[0066] Finally, the combination of G and T in the present invention realizes the full-process automation from text input to dynamic optimization of garment patterns, greatly reducing manual intervention and effectively improving the pattern-making efficiency, precision and quality.
[0067] Generally speaking, the present invention can not only promote the rapid iteration and personalized customization of clothing design, but also provide innovative solutions for the intelligentization and industrialization of the clothing industry, demonstrating outstanding technical advantages and broad application prospects, and can be widely applied to fields such as clothing design, virtual fitting, and customized production. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 is a flowchart of a text-driven intelligent clothing pattern generation method according to an embodiment of the present invention;
[0069] Figure 2 is a flowchart of constructing a clothing pattern dataset and a virtual fitting dataset according to an embodiment of the present invention;
[0070] Figure 3 is a flowchart of training G according to an embodiment of the present invention;
[0071] Figure 4 is a flowchart of generating clothing pattern data according to an embodiment of the present invention;
[0072] Figure 5 is a flowchart of training T according to an embodiment of the present invention;
[0073] Figure 6 is a flowchart of dynamically optimizing X2* according to an embodiment of the present invention;
[0074] Figure 7 is a result diagram of clothing patterns generated by an embodiment of the present invention and a comparative example;
[0075] Figure 8 is a comparison diagram of three-dimensional static virtual fitting grid data of clothing patterns generated by an embodiment of the present invention and a comparative example in an A-type standing posture (feet shoulder-width apart, hands hanging naturally);
[0076] Figure 9 is a comparison diagram of three-dimensional static virtual fitting grid data of clothing patterns generated by an embodiment of the present invention and a comparative example in a running posture. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0077] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0078] The following are the test methods for the relevant performance indicators of the embodiments and the comparative examples:
[0079] (1) Virtual stitching experiment: Use the CLO3D clothing simulation software and the SMPL human model to perform virtual stitching on the clothing pattern.
[0080] (2) Three-dimensional virtual fitting experiment for running movements: Use the CLO3D clothing simulation software.
[0081] (3) Success rate of clothing pattern generation: First, divide the test set from the clothing pattern dataset at a ratio of 20%; then, provide the input parts of each sample in the test set to the clothing pattern generation model; finally, count the percentage of the number of times the clothing pattern generation model can correctly generate the corresponding target instance data in the total number of samples in the test set.
[0082] (4) Clothing pattern generation speed: It is defined as the time taken to generate the clothing pattern data, that is, timing starts from when the clothing pattern feature description text is completed and input into the clothing pattern generation model, and ends when the clothing pattern generation model completes the output of the clothing pattern data.
[0083] (5) Description text consistency: First, use the CLO3D clothing simulation software to virtually stitch the generated clothing pattern into a 3D clothing and intercept its three views; then, input the images of these three views and the description text corresponding to the pattern into the Qwen2.5-VL multimodal large model to extract their respective feature embeddings; finally, calculate the description text consistency by calculating the cosine similarity between the image feature embedding and the text feature embedding.
[0084] (6) Human body dynamic fit: By virtually fitting the clothing pattern generated in the CLO3D clothing simulation software and calculating the offset loss value between it and the human body structure line, take 10 human body structure lines to calculate the offset loss value respectively, and then take the average value to obtain the average offset loss value. This average offset loss value reflects the degree of fit between the clothing and the human body structure line in the simulation state, which is the human body dynamic fit.
[0085] (7) Industrial applicability: Meet software compatibility and fabric considerations;
[0086] Software compatibility means that the file format of the generated clothing pattern needs to conform to the clothing industry standard and support seamless docking with existing mainstream clothing industry software (such as clothing pattern CAD software, etc.);
[0087] Fabric considerations mean that the clothing pattern design adapts to the characteristics of different fabrics, or the information contained in the clothing pattern reflects the characteristics of different fabrics, meeting the diverse needs of actual production.
[0088] (8) Initial alignment accuracy:
[0089] Measure the positional deviation between the virtual garment after virtual stitching and the corresponding structural lines of the human body model, and calculate the percentage of the number of structural lines with a deviation less than 5 mm in the total number of structural lines, and that is obtained.
[0090] (9) Triangular face self-intersection rate:
[0091] First, count the number of triangular patches that intersect with each other (i.e., one patch passes through another patch) in the virtual garment after virtual stitching, and then calculate the percentage of it in the total number of triangular patches, and that is the triangular face self-intersection rate.
[0092] (10) Vertex normal vector consistency:
[0093] First, calculate the angle difference between the normal vector (the vector representing the surface orientation) of each vertex in the virtual garment after virtual stitching and its adjacent vertices, and then calculate the percentage of the number of vertices with an angle difference less than 10° in the total number of vertices, and that is the vertex normal vector consistency.
[0094] (11) Fabric penetration rate:
[0095] First, count the number of vertices that penetrate into the interior of the virtual human body model in the virtual garment after virtual stitching, and then calculate the percentage of the number of these vertices in the total number of vertices of the virtual garment, and that is the fabric penetration rate.
[0096] (12) Decoding accuracy rate:
[0097] After training G is completed, use 20% of the matching (D1, D2) pairs reserved from the garment pattern dataset as the test set. Input D1 in the test set into G, and calculate the percentage of the number of times the corresponding D5 is correctly generated in the total number of samples in the test set, and that is the decoding accuracy rate.
[0098] Embodiment
[0099] A text-driven intelligent generation method for garment patterns, as Figure 1 shown, the specific steps are as follows:
[0100] (1) Construct a garment pattern dataset and a virtual fitting dataset, as Figure 2 shown;
[0101] S1.1. Select the SMPL human body model as the three-dimensional virtual human body model A1, mark 10 structural lines (including chest circumference line, waist circumference line, hip circumference line, shoulder line, neck circumference line, arm circumference line, front center line, back center line, nipple side length line, front shoulder oblique length line), measure height, chest circumference, waist circumference, hip circumference, back length, front center length, nipple distance, nipple side length, shoulder width, chest width, back width, front shoulder oblique length, waist length, arm length, shoulder inclination, leg length and inseam height, and use these measurements as the virtual human body model size data A2. Based on A2 Through geometric mapping calculations, multiple garment pattern parameter configuration files D2 for five types of clothing (T-shirts, vests, pants, skirts, and dresses) were determined. Fifty garment patterns were generated for each type of clothing, for a total of 250 garment patterns. 1,250 pattern data were specified. D2 included combinable components (such as necklines, sleeves, and trouser legs). D2 was input into a parametric pattern generation program G4 to generate multiple styles of garment pattern data X2. The data was saved in JSON format and included pattern shapes, connection relationships, three-dimensional spatial positions, and fabric stretch settings.
[0102] S1.2. Manually input action feature description text, covering five main sports types ("running", "yoga", "cycling", "walking", and "jumping"), to construct multiple action feature description texts X1. Use the text-action generation model A3 from the literature (LGTM: Local-to-Global Text-Driven Human Motion Diffusion Model, SIGGRAPH, 2024) to generate 3D virtual human motion data A4. The "running" action includes typical postures such as starting, stepping, and leg retraction. The "running" action has a total of 1600 frames. Import A4 into A1 in step S1.1 to construct multiple dynamic 3D virtual human models X3.
[0103] S1.3. Combine X3 from step S1.2 and use the physics-based fitting program A5. Use fabrics with different tensile stiffnesses (randomly selected within the range of 1-10,000) to virtually stitch X2 generated in step S1.1 to generate high-precision three-dimensional virtual dynamic fitting mesh data X6. This data is saved in GLB format with an initial alignment accuracy of 92%, a triangle self-intersection rate of 0.8%, a vertex normal consistency of 96%, and a fabric penetration rate of 4%. This meets the requirements of initial alignment accuracy ≥ 90%, triangle self-intersection rate ≤ 1%, vertex normal consistency ≥ 95%, and fabric penetration rate ≤ 5%.
[0104] S1.4. Map X6 in step S1.3 to a multi-view image A6 with a sampling interval of 90° and a resolution of 1024×1024 pixels. Combine X2 in step S1.1 and use the LLaMA-Factory tool to fine-tune the Qwen 2.5-VL large model to obtain a fine-tuned multi-modal large model A7, enabling it to recognize features at the clothing level and the pattern level. Clothing-level features such as "loose fit of T-shirt", and pattern-level features such as "V-neck depth 5 cm", and automatically batch-generate corresponding clothing pattern feature description text D1. The cosine similarity of the cross-modal feature vectors is 0.75, meeting the requirement of ≥0.7.
[0105] S1.5. Integrate A2, D2, X2 in step S1.1 and D1 in step S1.4 to construct a clothing pattern dataset; integrate X1, A4, X3 in step S1.2 and X6 in step S1.3 to construct a virtual fitting dataset. Dedup with a cosine similarity of 0.7 as the threshold. After deduping, the clothing pattern dataset contains 10250 independent samples, meeting the requirement of ≥5000, and the virtual fitting dataset contains 1000 dynamic fitting samples, meeting the requirement of ≥800, with uniform sample distribution.
[0106] (2) Construct and train a clothing pattern intelligent generation model G;
[0107] G consists of a text embedding feature extractor G1, a clothing pattern parallel graph neural network encoder G2, a clothing pattern parallel autoregressive decoder G3, and a parameterized pattern generation program G4. G1 and G2 are both connected to G3, and G3 is connected to G4. G2 is only enabled during the training of G, and G4 is only enabled during the application of G. G1 and G3 are enabled both during the training of G and during the application of G.
[0108] G1 uses the text embedding module of the Qwen 2.5-VL model;
[0109] G2 is a network composed of five parallel graph neural networks, which encode the patterns of the upper body, sleeves, pants, skirt, and accessories respectively. All patterns of different clothing parts are simultaneously input into their corresponding graph neural networks. Inside each graph neural network, the clothing patterns are regarded as nodes, and the connection relationships and three-dimensional spatial position information between the patterns are used to construct the graph structure of the graph neural network.
[0110] G3 uses the decoder of the Transformer model;
[0111] G4 uses the parameterized pattern generation program in the literature (GarmentCodeData: A Dataset of 3D Made-to-Measure GarmentsWith Sewing Patterns, ECCV, 2024);
[0112] The training of G uses the clothing pattern feature description text D1 in the clothing pattern dataset, the clothing pattern parameter configuration file D2 in the clothing pattern dataset, and the Adam optimizer. The learning rate is 0.001, the batch size is 32, and it is trained for 50 epochs. After the training of G is completed, the decoding accuracy rate of D5 is 95.8%, meeting the requirement that the decoding accuracy rate ≥ 95%;
[0113] As Figure 3 shown, the process of training G is as follows:
[0114] (a) Input D1 into G1, and it extracts the 512×77-dimensional text embedding feature D3;
[0115] (b) Map D2 through G2 to the clothing pattern latent space to generate the clothing pattern latent space encoding D4; G2 consists of two parts: quantization encoding and feature encoding. Among them, the quantization encoding normalizes the default value of D2 to the standardized value [0,1] according to the value range, generating a 702×1-dimensional standardized value. The encoding dimension is the same as the number of pattern parameters, and no data compression is performed; the feature encoding maps the quantization encoding to the latent space in parallel according to multiple components to form 128×3-dimensional D4;
[0116] (c) Input D3 and D4 into G3 together for decoding to obtain the new clothing pattern parameter configuration file D5. The specific process is as follows: First, align D3 and D4, and then decode the aligned encoding to the clothing pattern parameter configuration quantization encoding in parallel according to multiple components to generate D5;
[0117] (d) Calculate the parameter configuration file loss value The formula is as follows:
[0118]
[0119] In the formula, x i and y i respectively represent the values of the i-th parameter in D5 and D2, and N P represents the total number of parameters in D5 or D2;
[0120] (e) Determine whether the number of iterations in steps (a) to (d) is greater than or equal to 10, and whether it is less than 0.0003. If both are yes, end; otherwise, update the neural network parameters of G2 and G3, and return to step (a);
[0121] (3) Generate clothing pattern data, as Figure 4 shown;
[0122] The user-defined clothing pattern feature description text D1*"The upper garment is a loose women's short-sleeved T-shirt with a V-neck and the side line is expanded by 1.5 cm; the lower garment is loose shorts for running" is input into the trained G, and the clothing pattern data X2* generated by G takes 18 s;
[0123] (4) Construct and train a learning-based 3D virtual fitting model T;
[0124] T consists of a dynamic human body generation module T1, a pattern static mapping module T2, a 3D clothing dynamic deformation mapping module T3, and a 3D clothing explicit decoder T4. T1 and T2 are connected to T3 simultaneously, and T3 is connected to T4;
[0125] T1 consists of an SMPL human body model and a text-action generation model. The action parameters generated by the text-action generation model are input into the SMPL human body model to obtain a dynamic 3D virtual human body; the SMPL human body model is from the literature (SMPL: a skinned multi-person linear model, SIGGRAPH Asia, 2015.), and the text-action generation model is from the literature (LGTM: Local-to-Global Text-Driven Human Motion Diffusion Model, SIGGRAPH, 2024.). The frame rate of the dynamic 3D virtual human body generated by T1 is 30 frames / s;
[0126] T2 uses a fully connected neural network;
[0127] T3 uses a residual neural network;
[0128] T4 uses MeshUDF in the literature (MeshUDF: Fast and Differentiable Meshing of Unsigned Distance Field Networks, ECCV, 2022.), and T4 generates 3D virtual fitting mesh data;
[0129] To train T, use the action feature description text X1 in the virtual fitting dataset, the 3D dynamic virtual fitting mesh data X6 in the virtual fitting dataset, the clothing pattern data X2 in the clothing pattern dataset, and the Adam optimizer. The learning rate is 0.001, the batch size is 8, and train for 50 epochs. During the training of T, T1, T2, T3, and T4 are all enabled;
[0130] In the trained T, the mapping speed from the clothing pattern to the 3D virtual fitting mesh is 7.5 ms / frame, and the chamfer distance between the fitting result using A5 and the fitting result using the trained T is 3.3 mm;
[0131] As shown in Figure 5 the following is the process of training T:
[0132] (i) Input X1 into T1, and its output is the dynamic three-dimensional virtual human model X3;
[0133] (ii) Convert X2 into three-dimensional clothing static implicit data X4 through T2. X4 is represented by an unsigned distance field;
[0134] (iii) Input X3 and X4 into T3 together, and its output is the three-dimensional clothing dynamic implicit data X5. X5 is represented by an unsigned distance field;
[0135] (iv) Input X5 into T4, and its output is the three-dimensional dynamic virtual fitting grid data V1;
[0136] (v) Calculate the chamfer loss value The formula is as follows:
[0137]
[0138] In the formula, A represents the set of sampled vertices of V1, B represents the set of sampled vertices of X6, N a represents the number of sampled vertices in A, N b represents the number of sampled vertices in B, a ∈ A and b ∈ B respectively represent a single vertex in A and B;
[0139] (vi) Judge whether the number of iterations in steps (i) to (v) is greater than or equal to 10 steps, and whether it is less than or equal to 5 mm. If both are yes, end. Otherwise, after updating the neural network parameters of T2 and T3, return to step (i);
[0140] (5) Perform dynamic optimization on X2*, as Figure 6 shown, the specific process is as follows:
[0141] a) Input X2* and the user-defined action feature description text X1* "running" into the trained learning-based three-dimensional virtual fitting model T. Set the non-elastic fabric parameter with a stretching stiffness of 10000. T1, T2, T3, and T4 are all enabled, and the three-dimensional dynamic virtual fitting grid data V1* generated by T4 is output;
[0142] b) Select key action frames (starting and stepping), sample once every 10 frames, and calculate the structural line offset loss value The formula is as follows:
[0143]
[0144] Where, ω1 represents the weight value of the key motion frames in V1*, ω1 = 0.05, N k represents the total number of key motion frames in V1*, represents the difference of the structure lines under the i-th key motion frame in V1*, ω2 represents the weight value of the process action frames in V1*, ω2 = 0.001, N p represents the total number of process action frames in V1*, represents the difference of the structure lines under the j-th process action frame in V1*;
[0145] c) Determine whether the number of iterations in steps a) to b) is greater than or equal to 10, whether it converges to 10% of the initial value, whether it is less than or equal to 5 mm, and whether it is less than or equal to 5 mm. If all are yes, then take the X2* of the last iteration in steps a) to b) as the clothing sample data V2* applicable to non-elastic fabrics and proceed to the next step. Otherwise, update X2* and return to step a);
[0146] d) Input V2* into T, set the parameters of the real elastic fabric with a tensile stiffness of 100, do not enable T2, and enable T1, T3, and T4. With an A-type standing posture, generate the three-dimensional static virtual fitting grid data V3* output by T4;
[0147] e) Calculate the fabric deformation loss value The formula is as follows:
[0148]
[0149] Where, N v represents the number of sampled vertices of V3* or V2*, with a sampling rate of 0.5 and 1548 sampling points. The sampling points meet the requirement of ≥1000 and are evenly distributed in the key areas of the upper and lower garments; p i represents the spatial coordinate of the i-th sampled vertex of V3*, q i represents the spatial coordinate of the i-th sampled vertex of V2*;
[0150] f) Whether the number of iterations in steps d) to e) is greater than or equal to 10, and whether it is less than or equal to 5 mm. If all are yes, then take the V2* of the last iteration in steps d) to e) as the clothing sample data V4* applicable to real elastic fabrics. Otherwise, update V2* and return to step d);
[0151] e) Fine-tune V4* according to personalized needs. Adjust the V-neck depth from 5 cm to 6 cm and save it to meet the industrial production standards.
[0152] The clothing sample results finally generated by the embodiment are as follows Figure 7 shown in part (a) of Figure 7 . The clothing sample is a running sportswear sample;
[0153] The Figure 7 clothing sample shown in part (a) of Figure 7 was subjected to a virtual sewing experiment, and three-dimensional static virtual fitting grid data was generated in the A-type standing posture, as shown in Figure 8 part (a) of Figure 8 ;
[0154] The Figure 7 clothing sample in part (a) of Figure 7 was applied to a three-dimensional virtual fitting experiment of a running motion, and three-dimensional static virtual fitting grid data in the running posture was generated, as shown in Figure 9 part (a) of Figure 9 .
[0155] Comparative example
[0156] A method for generating a clothing sample specifically uses the SewingGPT model proposed in the literature (DressCode: Autoregressively Sewing and Generating Garments from Text Guidance, ACM Trans. Graph., 2024) to generate clothing samples. The SewingGPT model supports inputting short English words to generate clothing samples. However, the SewingGPT model has the following technical limitations: it does not support Chinese input; it cannot parse or respond to descriptions of specific wearing states or action scenarios of clothing (such as the clothing form in the running posture); its input text format has strict requirements and lacks flexibility;
[0157] In order to compare the SewingGPT model with the intelligent clothing sample generation model G provided by the embodiment of the present invention, the clothing sample feature description text input into the SewingGPT model is the specific content of the clothing sample that is simplified and translated into English based on the Chinese input content in the embodiment of the present invention: "[\"shirt, short sleeves, V-neck, loose\", \"pants, short, loose fit\"]"; Subsequently, the clothing sample generated by the SewingGPT model was imported into the CLO3D clothing simulation software, and the same running motion simulation was performed on the clothing sample in the same way as the intelligent clothing sample generation model G in the embodiment of the present invention.
[0158] The clothing sample results finally generated by the comparative example are as follows Figure 7 shown in part (b) of Figure 7 . The clothing sample is a running sportswear sample; The Figure 7 clothing sample shown in part (b) of Figure 7 was subjected to a virtual sewing experiment, and three-dimensional static virtual fitting grid data was generated in the A-type standing posture, asFigure 8 as shown in part (b) of
[0159] Apply Figure 7 the clothing pattern of part (b) to a three-dimensional virtual fitting experiment of a running motion, and generate three-dimensional static virtual fitting grid data in a running posture, as Figure 9 shown in part (b) of
[0160] From Figure 7 the clothing pattern results shown in Figure 8 and the comparison results of the three-dimensional static virtual fitting grid data shown in
[0161] From Figure 9 the comparison diagram of the three-dimensional static virtual fitting grid data of the running posture shown in
[0162] In addition, the key indicators of the example and the comparative example were compared and evaluated, and the results are shown in Table 1.
[0163] The comparison results in Table 1 show that the clothing pattern generation method provided by the present invention is superior to SewingGPT in terms of clothing pattern generation success rate, clothing pattern generation speed, description text consistency, human body dynamic fit, and industrial applicability.
[0164] Table 1 Comparison of indicators for clothing pattern generation
[0165]
Claims
1. A text-driven clothing pattern intelligent generation method, characterized in that: Input the user-defined clothing sample feature description text D1* into the trained clothing sample intelligent generation model G, which outputs the generated clothing sample data X2*; G consists of a text embedding feature extractor G1, a clothing sample parallel graph neural network encoder G2, a clothing sample parallel autoregressive decoder G3, and a parameterized sample generator G4. G1 and G2 are connected to G3 at the same time, and G3 is connected to G4. G2 is only enabled during model training, G4 is only enabled during model application, and G1 and G3 are enabled during both training and application. Training G uses the clothing sample feature description text D1 in the clothing sample dataset and the clothing sample parameter configuration file D2 in the clothing sample dataset. The process is as follows: (a) Input D1 into G1, which extracts the text embedding feature D3; (b) Map D2 to the clothing sample latent space through G2 to generate the clothing sample latent space encoding D4; (c) Input D3 and D4 together into G3 for decoding to obtain a new clothing sample parameter configuration file D5; (d) Calculate the parameter profile loss value The formula is as follows: Where x i with y i Represents the value of the i-th parameter in D5 and D2 respectively, N P Indicates the total number of parameters in D5 or D2; (e) Determine whether the number of iterations of steps (a) to (d) is greater than or equal to 10, and Is it less than 0.0003? If yes, then end; otherwise, update the parameters of G2 and G3 and return to step (a).
2. A text-driven clothing pattern intelligent generation method according to claim 1, characterized in that: In step (b), G2 consists of two parts: quantization coding and feature coding. Quantization coding normalizes the default value of D2 to a standardized value according to the value range. The coding dimension is consistent with the number of template parameters, and no data compression is performed. Feature coding maps the quantization code to the latent space in parallel according to multiple components to form D4.
3. A text-driven clothing pattern intelligent generation method according to claim 1, characterized in that: The specific process of step (c) is as follows: first align D3 and D4, then decode the aligned codes into clothing pattern parameter configuration quantization codes in parallel according to multiple components to generate D5.
4. A text-driven clothing pattern intelligent generation method according to claim 1, characterized in that: X2* is also dynamically optimized, and the process is as follows: a) X2* and the user-defined motion feature description text X1* are input into the trained learning-based 3D virtual fitting model T, and the 3D dynamic virtual fitting mesh data V1* is generated by the model using non-stretch fabric parameter settings; b) Calculate the structural line offset loss value The formula is as follows: Where ω1 represents the weight value of the key motion frame in V1*, N k represents the total number of key motion frames in V1*, represents the difference of the structure line under the i-th key motion frame in V1*, ω2 represents the weight value of the process action frame in V1*, N p Indicates the total number of process action frames in V1*, represents the difference in the structural line under the j-th process action frame in V1*; c) Determine whether the number of iterations of steps a) to b) is greater than or equal to 10 steps, Whether it converges to 10% of the initial value, Is it less than or equal to 5mm, and Is it less than or equal to 5mm? If yes, use X2* from the last iteration of steps a) to b) as the garment sample data V2* suitable for non-stretch fabrics and proceed to the next step. Otherwise, update X2* and return to step a); d) Input V2* into the trained T, use the real stretch fabric parameter settings, and generate the three-dimensional static virtual fitting mesh data V3* from its output; e) Calculate the fabric deformation loss value The formula is as follows: Where N v Indicates the number of sampling vertices of V3* or V2*, p i represents the spatial coordinates of the i-th sampling vertex of V3*, q i Represents the spatial coordinates of the i-th sampling vertex of V2*; f) Determine whether the number of iterations of steps d) to e) is greater than or equal to 10 steps, and Is it less than or equal to 5mm? If yes, use V2* of the last iteration of steps d) to e) as the garment sample data V4* suitable for real stretch fabrics. Otherwise, update V2* and return to step d); e) Fine-tune V4* according to your individual needs and save it.
5. A text-driven clothing pattern intelligent generation method according to claim 4, characterized in that: T consists of a dynamic human body generation module T1, a template static mapping module T2, a 3D clothing dynamic deformation mapping module T3, and a 3D clothing explicit decoder T4. T1 and T2 are connected to T3 at the same time, and T3 is connected to T4. During application, when the non-stretch fabric parameter setting is used, T1, T2, T3 and T4 are all enabled; When using the real stretch fabric parameter setting, T2 is not enabled, and T1, T3 and T4 are all enabled; During training, T1, T2, T3, and T4 are all enabled; Training T uses the action feature description text X1 in the virtual fitting dataset, the three-dimensional dynamic virtual fitting mesh data X6 in the virtual fitting dataset, and the clothing sample data X2 in the clothing sample dataset. The process is as follows: (i) Input X1 into T1, which outputs a dynamic three-dimensional virtual human body model X3; (ii) convert X2 into three-dimensional clothing static implicit data X4 through T2; (iii) X3 and X4 are inputted into T3, which then outputs the three-dimensional clothing dynamic implicit data X5; (iv) X5 is input into T4, which outputs three-dimensional dynamic virtual fitting mesh data V1; (v) Calculation of chamfer loss The formula is as follows: Where A represents the set of sampling vertices of V1, B represents the set of sampling vertices of X6, and N a Indicates the number of sampled vertices in A, N b represents the number of sampled vertices in B, a∈A and b∈B represent single vertices of A and B respectively; (vi) Determine whether the number of iterations of steps (i) to (v) is greater than or equal to 10 steps, and Is it less than or equal to 5mm? If yes, then end; otherwise, update the parameters of T2 and T3 and return to step (i).
6. A text-driven clothing pattern intelligent generation method according to claim 5, characterized in that: The steps for constructing the clothing sample dataset and virtual fitting dataset are as follows: S1.
1. Select a 3D virtual human body model A1, measure its dimensional data A2, set multiple styles D2 based on A2, input D2 into a parametric pattern generation program G4, and generate multiple styles X2; S1.2, manually construct multiple X1, use the text-action generation model A3 to generate three-dimensional virtual human motion data A4, import A4 into A1 in step S1.1, and thus construct multiple dynamic three-dimensional virtual human models X3; S1.3, combining X3 from step S1.2, using the physical simulation-based fitting program A5, virtually stitching X2 generated from step S1.1 to generate X6; S1.4, map X6 from step S1.3 to a multi-view image A6, and combine it with X2 from step S1.1 to fine-tune the multimodal large model A7 so that it can recognize features at the clothing and pattern levels and automatically generate the corresponding D1 in batches; S1.5, integrate A2, D2, X2 in step S1.1 and D1 in step S1.4 to construct a clothing sample dataset; Integrate X1, A4, X3 in step S1.2 and X6 in step S1.3 to construct the virtual fitting dataset.
Citation Information
Patent Citations
Garment sheet data generation method and device
CN119337442A
Method, device and equipment for generating virtual clothing through multi-modal fusion and storage medium
CN114723843A
Virtual fitting method based on text-driven image generation
CN116205786A