Text-driven garment template intelligent generation method
Through multi-module collaborative clothing model intelligent generation model and learning-based three-dimensional virtual fitting technology, the problems of low generation efficiency and difficult to control results in the existing technology are solved, and high-quality, personalized and diversified clothing model automatic generation and dynamic optimization are achieved.
Patent Information
- Application Number
- CN202510669798.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-23
AI Technical Summary
The existing intelligent generation technology of clothing model is lacking effective design rules and constraints, resulting in low generation efficiency and difficult to control results, which cannot meet the rapid response of the modern clothing market to diversified and personalized needs.
A multi-module collaborative clothing template intelligent generation model is adopted, and through text embedding feature extractor, clothing template parallel graph neural network encoder, parallel autoregressive decoder and parameterized template generation program, combined with learning-based three-dimensional virtual fitting technology, a quantitative evaluation system and optimization mechanism is established to realize the automated generation and dynamic optimization of user-defined clothing template feature description text to high-quality clothing template data.
It significantly improves the efficiency and quality of clothing model generation, enhances the controllability and stability of the model, and can achieve accurate structural line offset and fabric deformation optimization in dynamic scenarios, meeting personalized and diversified design needs.
Smart Images

Figure CN120197247A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent garment pattern making, and relates to a method for intelligently generating garment patterns driven by text. Background Art
[0002] As the core link in garment production, the quality of garment pattern design directly affects the fitness, comfort and aesthetics of the finished garment. In the traditional garment manufacturing field, pattern making mainly relies on manual drawing. This method is not only inefficient, but also highly dependent on the experience accumulation of pattern makers, with high learning costs and long time consumption, making it difficult to meet the rapid response of the modern garment market to diverse and personalized demands. Although the application of computer-aided design (CAD) software has improved the pattern making efficiency to a certain extent, the mode of adjusting based on fixed templates still has obvious limitations in terms of flexibility and creativity realization.
[0003] In recent years, the booming development of artificial intelligence technology has opened up a new path for the intelligent generation of garment patterns. The text-driven method with natural language processing as the core makes it possible to generate complex garment patterns through intuitive descriptive text. The SewingGPT proposed in the literature (DressCode: Autoregressively Sewing and Generating Garments from TextGuidance, ACM Trans. Graph., 2024.) was the first to explore the technology of generating garment patterns guided by descriptive text, verifying the feasibility of this technical route. However, the generated garment styles are relatively simple, and due to the lack of effective rule constraints, the effectiveness and quality of the generated results need to be improved.
[0004] Similarly, the patent application with the publication number CN119337442A provides a method for generating garment panel data. It extracts and decodes features through the geometric data of the pattern, but does not introduce systematic pattern design rule constraints. Due to the complex diversity of the geometric data of garment patterns and the lack of a standardized expression method, the generated results are difficult to control and cannot meet the basic pattern design specifications. In addition, during the geometric feature decoding process, the data dimension is high, the training difficulty is large, and problems such as generation failure are likely to occur, seriously reducing the stability and practical value of the generation system.
[0005] In summary, due to the lack of effective design rule constraints, the current intelligent garment pattern generation technology has problems such as low generation efficiency and difficult result control, and it is urgent to explore more effective solutions to promote the continuous development of garment pattern design technology. Summary of the Invention
[0006] The purpose of the present invention is to solve the above problems existing in the prior art and provide a method for intelligently generating garment patterns driven by text.
[0007] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0008] A text-driven intelligent clothing pattern generation method inputs the user-defined clothing pattern feature description text D1* into the trained clothing pattern intelligent generation model G (a deep neural network model), and outputs the generated clothing pattern data X2*;
[0009] G is composed of a text embedding feature extractor G1, a clothing pattern parallel graph neural network encoder G2, a clothing pattern parallel autoregressive decoder G3, and a parameterized pattern generation program G4. G1 and G2 are both connected to G3, and G3 is connected to G4. G2 is only enabled during the model training process, G4 is only enabled during the model application process, and G1 and G3 are both enabled during the training process and the application process;
[0010] The training of G uses the clothing pattern feature description text D1 in the clothing pattern dataset and the clothing pattern parameter configuration file D2 in the clothing pattern dataset. The process is as follows:
[0011] (a) Input D1 into G1, and it extracts the text embedding feature D3;
[0012] (b) Map D2 through G2 to the clothing pattern latent space to generate the clothing pattern latent space encoding D4;
[0013] (c) Input D3 and D4 into G3 together for decoding to obtain a new clothing pattern parameter configuration file D5;
[0014] (d) Calculate the parameter configuration file loss value , and the formula is as follows:
[0015] ;
[0016] In the formula, and respectively represent the values of the th parameter in D5 and D2, represents the total number of parameters in D5 or D2 (the total number of parameters in D5 is the same as that in D2);
[0017] (e) Determine whether the number of iterations in steps (a) to (d) is greater than or equal to 10, and is less than 0.0003. If both are yes, end. Otherwise, update the parameters of G2 and G3, and return to step (a).
[0018] A text-driven intelligent clothing pattern generation method of the present invention realizes the automatic generation of high-quality clothing pattern data from the text description of clothing pattern features defined by users by constructing an intelligent clothing pattern generation model G with multi-module collaboration; the intelligent clothing pattern generation model G is composed of a text embedding feature extractor G1, a clothing pattern parallel graph neural network encoder G2, a clothing pattern parallel autoregressive decoder G3, and a parametric pattern generation program G4, and can convert the text description of clothing pattern features input by users into a clothing pattern parameter configuration file and generate clothing pattern data that meets the requirements;
[0019] The present invention performs feature modeling through a clothing pattern parameter configuration file and incorporates structured design rules using the parametric pattern generation program G4 to effectively constrain the clothing pattern generation process; compared with directly processing complex geometric data, the clothing pattern parameter configuration file provides a more compact and clearly expressed feature representation method, significantly improving the generation efficiency and model controllability of clothing pattern data; at the same time, by adopting the clothing pattern parallel graph neural network encoder G2 and the clothing pattern parallel autoregressive decoder G3, the clothing pattern components are partitioned and parallel encoded and decoded, which not only optimizes the computational efficiency of the intelligent clothing pattern generation model but also enhances the stability and diversity of clothing pattern generation.
[0020] As a preferred technical solution:
[0021] In the text-driven intelligent clothing pattern generation method as described above, in step (b), G2 includes two parts: quantization encoding and feature encoding. Among them, quantization encoding normalizes the default value of D2 into a standardized value according to the value range, and the encoding dimension is the same as the number of pattern parameters without data compression; feature encoding maps the quantization encoding to the latent space in parallel according to multiple components to form D4.
[0022] In the text-driven intelligent clothing pattern generation method as described above, the specific process of step (c) is as follows: first align D3 and D4, and then decode the aligned encoding into the quantization encoding of the clothing pattern parameter configuration in parallel according to multiple components to generate D5.
[0023] In the text-driven intelligent clothing pattern generation method as described above, X2* is also dynamically optimized, and the process is as follows:
[0024] a) Input X2* and the text description of the action features X1* defined by the user into the trained learning-based 3D virtual fitting model T (a deep neural network model) at the same time, and set the non-elastic fabric parameters, and output the generated 3D dynamic virtual fitting grid data V1* from it;
[0025] b) Calculate the structural line offset loss value , and the formula is as follows:
[0026] ;
[0027] Wherein, represents the weight value of the key motion frames in V1*, represents the total number of key motion frames in V1*, represents the difference in the structure line under the th key motion frame in V1*, represents the weight value of the process action frames in V1*, represents the total number of process action frames in V1*, represents the difference in the structure line under the th process action frame in V1*;
[0028] c) Determine whether the number of iterations in steps a) to b) is greater than or equal to 10 steps, whether it converges to 10% of the initial value, whether it is less than or equal to 5 mm, and whether it is less than or equal to 5 mm. If all are yes, then take the X2* of the last iteration of steps a) to b) as the garment sample data V2* applicable to non-elastic fabrics and proceed to the next step. Otherwise, update X2* and return to step a);
[0029] d) Input V2* into the trained T, and use the real elastic fabric parameters to set, and output the generated three-dimensional static virtual fitting grid data V3*;
[0030] e) Calculate the fabric deformation loss value , and the formula is as follows:
[0031] ;
[0032] Wherein, represents the number of sampled vertices of V3* or V2* (the number of sampled vertices of V3* is the same as that of V2*), represents the spatial coordinate of the th sampled vertex of V3*, represents the spatial coordinate of the th sampled vertex of V2*;
[0033] f) Determine whether the number of iterations in steps d) to e) is greater than or equal to 10 steps, and whether it is less than or equal to 5 mm. If all are yes, then take the V2* of the last iteration of steps d) to e) as the garment sample data V4* applicable to real elastic fabrics. Otherwise, update V2* and return to step d);
[0034] e) Fine-tune V4* according to personalized needs and save.
[0035] In the production of the garment industry, the optimization and adjustment of garment patterns is a key link. The initial patterns usually need to be adjusted multiple times before they can be put into production. Traditional optimization methods mainly rely on manual physical fitting and manual modification. This method not only consumes a large amount of time and labor costs, but also makes it difficult to accurately evaluate the wearing effect of clothing in dynamic movement scenarios. Especially in the design and development of functional clothing and sports clothing, the dynamic wearing performance directly affects the functionality and comfort of the product, making the limitations of traditional methods more prominent.
[0036] Although existing related research has tried to innovate, for example, the literature (Development of upper cycling clothes using 3D-to-2D flattening technology and evaluation of dynamic wear comfort from the aspect of clothing pressure, International Journal of Clothing Science and Technology, 2016) guides pattern optimization by measuring the clothing pressure in specific movement postures. However, this method can only reflect the clothing state in static postures and cannot comprehensively evaluate the adaptability changes of clothing during dynamic movement. In addition, such research still highly relies on manual experience judgment and is difficult to achieve an automated production process. At the same time, the research generally regards fabric properties as fixed values and ignores the significant impact of fabric elasticity, a key variable, on garment pattern design, while fabric elasticity will significantly affect the actual wearing form of clothing.
[0037] Existing clothing simulation technologies are mostly based on physical simulation fitting programs, whose operation process is non-differentiable and cannot automatically reverse-adjust the pattern according to the fitting results. The patent application with the publication number CN119337442A provides a method for generating garment panel data. This method has not yet established a quantitative evaluation system for pattern fitness and cannot perform fitness optimization in dynamic postures. At the same time, the model does not consider the influence of fabric elasticity on the finished garment effect, which may lead to pattern wearing deformation, decreased fitness, and damaged aesthetics under different fabrics. The relevant adjustments still rely on manual experience and manual operations and are difficult to automate, greatly limiting the popularization and application of this method in actual production.
[0038] How to quantitatively optimize garment patterns in dynamic scenarios to improve the dynamic adaptability of clothing has become one of the key challenges in intelligent garment pattern generation technology.
[0039] With the rapid development of artificial intelligence technology, the learning-based virtual fitting method has become an important development trend in the field of 3D virtual fitting. The learning-based virtual fitting method uses machine learning algorithms to deeply analyze a large amount of clothing and human interaction data, learn and predict the deformation effect of clothing on the human body surface, and then achieve an efficient virtual fitting experience. The learning-based virtual fitting method has significant computational efficiency advantages. By using a feedforward neural network, it can achieve end-to-end generation of 3D clothing deformation effects, successfully avoiding the complex cloth iterative solution process in traditional physical simulations. In addition, due to the differentiability of its computational process, it can automatically perform reverse optimization based on the fitting results and adjust the clothing pattern, which opens up a new path for intelligent clothing pattern generation.
[0040] The present invention combines the learning-based 3D virtual fitting technology to establish a quantitative evaluation system for the fitness of clothing patterns with the offset of clothing-human structure lines as the index. By introducing dynamic action features, it can detect the changes in structure lines under different postures and iteratively optimize, realizing the automatic optimization and adjustment of clothing patterns in the dynamic wearing state. In addition, a loss function is introduced to precisely control the generation quality of clothing patterns, ensuring that the clothing pattern data output by the intelligent clothing pattern generation model is highly consistent with the real data. The optimization system also considers the influence of fabric elasticity on the clothing pattern effect. By simulating the deformation degree of real elastic fabrics, it further optimizes the clothing pattern structure to ensure fitness and consistency under different elastic fabrics, effectively improving the automation level and practical performance of clothing pattern design, and expanding the actual application scenarios of clothing pattern generation technology in the clothing industry.
[0041] A text-driven intelligent clothing pattern generation method as described above, T is composed of a dynamic human body generation module T1, a pattern static mapping module T2, a 3D clothing dynamic deformation mapping module T3, and a 3D clothing explicit decoder T4. T1 and T2 are both connected to T3, and T3 is connected to T4;
[0042] In the application process, when the non-elastic fabric parameter setting is adopted, T1, T2, T3, and T4 are all enabled; when the real elastic fabric parameter setting is adopted, T2 is not enabled, and T1, T3, and T4 are all enabled;
[0043] During the training process, T1, T2, T3, and T4 are all enabled;
[0044] Training T uses the action feature description text X1 in the virtual fitting dataset, the 3D dynamic virtual fitting grid data X6 in the virtual fitting dataset, and the clothing pattern data X2 in the clothing pattern dataset. The process is as follows:
[0045] (i) Input X1 into T1, and it outputs a dynamic 3D virtual human body X3;
[0046] (ii) Convert X2 into three-dimensional clothing static implicit data X4 through T2;
[0047] (iii) Input X3 and X4 into T3 together, and output three-dimensional clothing dynamic implicit data X5 from it;
[0048] (iv) Input X5 into T4, and output three-dimensional dynamic virtual fitting grid data V1 from it;
[0049] (v) Calculate the chamfer loss value , and the formula is as follows:
[0050] ;
[0051] In the formula, A represents the set of sampled vertices of V1, B represents the set of sampled vertices of X6, represents the number of sampled vertices in A, represents the number of sampled vertices in B, a ∈ A and b ∈ B respectively represent a single vertex of A and B;
[0052] (vi) Determine whether the number of iterations in steps (i) to (v) is greater than or equal to 10 steps, and whether it is less than or equal to 5 mm. If both are yes, end. Otherwise, after updating the parameters of T2 and T3, return to step (i).
[0053] The learning-based three-dimensional virtual fitting model T of the present invention combines dynamic human body generation, sample static mapping, and three-dimensional clothing dynamic deformation mapping technologies, uses the action feature description text provided by the user and the generated clothing sample data to generate dynamic or static virtual fitting effects; and through optimization mechanisms such as seam line offset loss and fabric deformation loss to precisely control the generation quality of clothing sample data, ensure that the clothing sample data output by the clothing sample intelligent generation model is highly consistent with the real data, further improve the adaptability of clothing sample generation to different fabric characteristics, and ensure its accurate and reliable performance in real scenarios.
[0054] For a text-driven intelligent clothing sample generation method as described above, the construction steps of the clothing sample data set and the virtual fitting data set are as follows:
[0055] S1.1. Select a three-dimensional virtual human body model A1, measure its virtual human body size data A2, set multiple clothing sample parameter configuration files D2 based on A2, and input D2 into the parametric sample generation program G4 to generate various styles of clothing sample data X2;
[0056] S1.2. Manually construct multiple action feature description texts X1, use the text-action generation model A3 to generate three-dimensional virtual human body action data A4, and import A4 into A1 in step S1.1 to construct multiple dynamic three-dimensional virtual human body models X3;
[0057] S1.3. Combine X3 in step S1.2, and use the virtual fitting program A5 based on physical simulation to virtually stitch X2 generated in step S1.1 to generate high-precision three-dimensional virtual dynamic fitting grid data X6;
[0058] S1.4. Map X6 in step S1.3 to multi-view images A6, and combine with X2 in step S1.1 to fine-tune a specific multi-modal large model A7 so that it can recognize the features at the clothing and pattern levels and automatically batch-generate the corresponding clothing pattern feature description text D1;
[0059] S1.5. Integrate A2, D2, X2 in step S1.1 and D1 in step S1.4 to construct a clothing pattern data set; integrate X1, A4, X3 in step S1.2 and X6 in step S1.3 to construct a virtual fitting data set.
[0060] In recent years, text-driven digital clothing generation technology has made certain progress. The literature (AIpparel: a large multimodal generative model for digital garments, arxiv, 2024) realizes the generation of complex clothing styles by finely encoding the clothing pattern structure; the literature (Design2GarmentCode: turning design concepts to tangible garments through program synthesis, arxiv, 2024) uses a multi-modal model to generate clothing pattern configuration files, improving the rule constraints and success rate of pattern generation. However, the existing text-driven clothing pattern generation methods still have significant defects: First, the generation speed of clothing patterns with complex styles is slow, making it difficult to meet the requirements of rapid design iteration; second, the existing models are mostly trained based on clothing-level description texts, lacking the understanding and fine control ability of pattern-level features, and unable to accurately adjust and optimize the patterns. For example, it is difficult to execute specific editing instructions such as "lengthen the front center line of the front pattern piece by 2 cm".
[0061] A clothing pattern data generation method provided by a patent application with the publication number CN119337442A trains a model based on clothing-level description texts, fails to deeply understand and model pattern-level features, resulting in a serious lack of fine control ability in the clothing pattern design and adjustment links. For specific instructions in actual clothing pattern design and editing, such as "extend the front center line of the front pattern piece by 2 cm", the method provided by this patent application is difficult to accurately parse and execute, greatly limiting its practicability and operation accuracy, and it is difficult to be effectively applied to actual production design scenarios.
[0062] To support the efficient training and application of G and T, the present invention designs a multi-modal data construction process that includes a clothing pattern dataset and a virtual fitting dataset to ensure data diversity and high precision. When constructing the clothing pattern dataset, in addition to introducing clothing-level descriptions (such as "loose fit T-shirt"), refined text information at the pattern level (such as "V-neck depth 5 cm") is added with the help of a multi-modal large model. The introduction of pattern-level features enables G to deeply understand and master the fine-grained design elements at the clothing pattern level, thereby achieving precise control over clothing pattern design and significantly improving the flexibility and accuracy of clothing pattern generation and adjustment.
[0063] Beneficial effects:
[0064] First of all, by combining multi-modal large model technology and parallel autoregressive generation models, the present invention can quickly generate high-quality clothing pattern data based on the text description of clothing pattern features input by users. Compared with traditional pattern-making methods, the method of the present invention greatly improves the generation efficiency and design flexibility of clothing patterns, reduces the dependence on professional experience, and meets the needs of the modern clothing market for diversified and personalized designs. At the same time, the dataset of the present invention contains descriptive texts at the clothing pattern level, which further supports the precise adjustment of local features of clothing patterns, effectively solves the deficiencies in detail control of the existing technology, and provides higher precision and controllability for clothing pattern design.
[0065] Secondly, the present invention innovatively integrates text-driven motion generation technology and learning-based virtual fitting technology, which can accurately analyze the offset of structural lines and fabric deformation characteristics in dynamic scenarios, and automatically adjust clothing pattern data through multi-level reverse iterative optimization, significantly improving the fit and comfort of clothing, especially suitable for the design of functional clothing and sportswear, filling the gaps in dynamic optimization and fabric property adaptability of the existing technology.
[0066] Finally, the combination of G and T of the present invention realizes the full-process automation from text input to dynamic optimization of clothing patterns, greatly reducing manual intervention and effectively improving the pattern-making efficiency, precision and quality.
[0067] Generally speaking, the present invention can not only promote the rapid iteration and personalized customization of clothing design, but also provide an innovative solution for the intelligentization and industrialization of the clothing industry, demonstrating outstanding technical advantages and broad application prospects, and can be widely applied in the fields of clothing design, virtual fitting and customized production. Brief Description of the Drawings
[0068] Figure 1 is a flowchart of a text-driven intelligent clothing pattern generation method according to an embodiment of the present invention;
[0069] Figure 2It is a flowchart for constructing a clothing sample dataset and a virtual fitting dataset according to an embodiment of the present invention;
[0070] Figure 3 It is a flowchart for training G according to an embodiment of the present invention;
[0071] Figure 4 It is a flowchart for generating clothing sample data according to an embodiment of the present invention;
[0072] Figure 5 It is a flowchart for training T according to an embodiment of the present invention;
[0073] Figure 6 It is a flowchart for dynamically optimizing X2* according to an embodiment of the present invention;
[0074] Figure 7 It is a clothing sample result diagram generated according to an embodiment of the present invention and a comparative example;
[0075] Figure 8 is a comparison diagram of three-dimensional static virtual fitting grid data of clothing samples generated according to an embodiment of the present invention and a comparative example in an A-type standing posture (feet shoulder-width apart, hands hanging naturally);
[0076] Figure 9 It is a comparison diagram of three-dimensional static virtual fitting grid data of clothing samples generated according to an embodiment of the present invention and a comparative example in a running posture. Detailed Embodiments
[0077] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0078] The following are the test methods for relevant performance indicators of the embodiments and comparative examples:
[0079] (1) Virtual sewing experiment: Use the CLO3D clothing simulation software and the SMPL human model to perform virtual sewing on the clothing sample.
[0080] (2) Three-dimensional virtual fitting experiment for running actions: Use the CLO3D clothing simulation software.
[0081] (3) Success rate of clothing sample generation: First, divide the test set from the clothing sample dataset at a ratio of 20%; then, provide the input parts of each sample in the test set to the clothing sample generation model; finally, count the percentage of the number of times the clothing sample generation model can correctly generate the corresponding target instance data in the total number of test set samples.
[0082] (4)Garment pattern generation speed: It is defined as the time taken to generate garment pattern data, that is, starting from the completion of inputting the garment pattern feature description text into the garment pattern generation model and ending when the garment pattern generation model finishes outputting the garment pattern data.
[0083] (5)Description text consistency: First, use the CLO3D garment simulation software to virtually stitch the generated garment pattern into a 3D garment and capture its three views; then, input the images of these three views and the description text corresponding to the pattern into the Qwen 2.5-VL multimodal large model to extract their respective feature embeddings; finally, calculate the description text consistency by computing the cosine similarity between the image feature embedding and the text feature embedding.
[0084] (6)Human body dynamic fit: By virtually trying on the generated garment pattern in the CLO3D garment simulation software and calculating the offset loss value between it and the human body structure lines, calculate the offset loss values for 10 human body structure lines respectively, and then take the average value to obtain the average offset loss value. This average offset loss value reflects the degree of fit between the garment and the human body structure lines in the simulated state, which is the human body dynamic fit.
[0085] (7)Industrial applicability: Meet software compatibility and fabric considerations;
[0086] Software compatibility means that the file format of the generated garment pattern needs to conform to the clothing industry standard and support seamless docking with existing mainstream clothing industry software (such as garment pattern CAD software, etc.);
[0087] Fabric considerations mean that the garment pattern design adapts to the characteristics of different fabrics, or the information contained in the garment pattern reflects the characteristics of different fabrics, meeting the diverse needs of actual production.
[0088] (8)Initial alignment accuracy:
[0089] Measure the position deviation between the virtual garment after virtual stitching and the corresponding structure lines of the human body model, and calculate the percentage of the number of structure lines with a deviation less than 5mm in the total number of structure lines, which is obtained.
[0090] (9)Triangular face self-intersection rate:
[0091] First, count the number of triangular faces that intersect with each other (that is, one face passes through another face) in the virtual garment after virtual stitching, and then calculate the percentage of it in the total number of triangular faces, which is the triangular face self-intersection rate.
[0092] (10)Vertex normal vector consistency:
[0093] First, calculate the angular difference between the normal vectors (vectors representing the surface orientation) of each vertex in the virtual garment after virtual stitching and its adjacent vertices, and then calculate the percentage of the number of vertices with an angular difference less than 10° in the total number of vertices, which is the vertex normal vector consistency.
[0094] (11) Fabric penetration rate:
[0095] First, count the number of vertices in the virtual garment that penetrate into the interior of the virtual human body model after virtual stitching, and then calculate the percentage of this number of vertices in the total number of vertices of the virtual garment, which is the fabric penetration rate.
[0096] (12) Decoding accuracy rate:
[0097] After training G is completed, use 20% of the matching (D1, D2) pairs reserved from the garment pattern dataset as the test set. Input D1 in the test set into G, and calculate the percentage of the number of times the corresponding D5 is correctly generated in the total number of samples in the test set, which is the decoding accuracy rate.
[0098] Embodiment
[0099] A text-driven intelligent garment pattern generation method, as Figure 1 shown, the specific steps are as follows:
[0100] (1) Construct a garment pattern dataset and a virtual fitting dataset, as Figure 2 shown;
[0101] S1.1. Select the SMPL human body model as the three-dimensional virtual human body model A1, mark 10 structural lines (including chest line, waist line, hip line, shoulder line, neck base line, arm root line, front center line, back center line, side long line of nipple, oblique long line of front shoulder point), measure height, chest circumference, waist circumference, hip circumference, back length, front center length, nipple spacing, side length of nipple, shoulder width, chest width, back width, oblique length of front shoulder point, waist length, arm length, shoulder inclination, leg length and crotch height, use these measured values as the virtual human body model size data A2, based on A2, calculate and determine the parameter configuration files D2 of multiple garment patterns of 5 types of garments (T-shirts, vests, trousers, skirts, dresses) through geometric mapping. Generate 50 garment patterns for each type of garment, with a total of 250 garment patterns, specify 1250 pattern data. D2 includes combinable components (such as necklines, sleeves, trouser legs). Input D2 into the parametric pattern generation program G4 to generate garment pattern data X2 of various styles, save it in JSON format, including pattern shape, connection relationship, three-dimensional spatial position and fabric elasticity setting value;
[0102] S1.2, manually input action feature description texts, covering 5 main types of sports ("running", "yoga", "cycling", "walking", "jumping"), construct multiple action feature description texts X1, and use the text-action generation model A3 in the literature (LGTM: Local-to-GlobalText-Driven Human Motion Diffusion Model, SIGGRAPH, 2024) to generate 3D virtual human action data A4, where the "running" action includes typical postures such as starting, stepping, and leg retraction. The "running" action has a total of 1600 frames. Import A4 into A1 in step S1.1 to construct multiple dynamic 3D virtual human models X3;
[0103] S1.3. Combine X3 in step S1.2, adopt the fitting program A5 based on physical simulation, use fabrics with different tensile stiffness (random values in the range of 1-10000), virtually sew X2 generated in step S1.1, and generate high-precision three-dimensional virtual dynamic fitting mesh data X6, which is saved in GLB format, with an initial alignment accuracy of 92%, a triangle self-intersection rate of 0.8%, a vertex normal vector consistency of 96%, and a fabric penetration rate of 4%, meeting the requirements of initial alignment accuracy ≥ 90%, triangle self-intersection rate ≤ 1%, vertex normal vector consistency ≥ 95%, and fabric penetration rate ≤ 5%;
[0104] S1.4, map X6 in step S1.3 to a multi-view image A6, with a sampling interval of 90° and a resolution of 1024×1024 pixels, and combine it with X2 in step S1.1, use the LLaMA-Factory tool to fine-tune the Qwen 2.5-VL large model to obtain a fine-tuned multimodal large model A7, so that it can recognize clothing-level and pattern-level features, clothing-level features such as "T-shirt loose version", pattern-level features such as "V-neck depth 5cm", and automatically batch generate corresponding clothing pattern feature description text D1, and the cross-modal feature vector cosine similarity is 0.75, meeting the requirement of ≥0.7;
[0105] S1.5, integrate A2, D2, X2 of step S1.1 and D1 of step S1.4 to construct a clothing sample data set; integrate X1, A4, X3 of step S1.2 and X6 of step S1.3 to construct a virtual fitting data set, and remove duplicates with a cosine similarity of 0.7 as the threshold. After deduplication, the clothing sample data set contains 10250 independent samples, meeting the requirement of ≥5000, and the virtual fitting data set contains 1000 dynamic fitting samples, meeting the requirement of ≥800, and the samples are evenly distributed;
[0106] (2) Build and train a clothing sample intelligent generation model G;
[0107] G consists of a text embedding feature extractor G1, a garment pattern parallel graph neural network encoder G2, a garment pattern parallel autoregressive decoder G3, and a parametric pattern generation program G4. G1 and G2 are both connected to G3, and G3 is connected to G4. G2 is only enabled during the training of G, and G4 is only enabled during the application of G. G1 and G3 are enabled during both the training and application of G.
[0108] G1 uses the text embedding module of the Qwen 2.5-VL model.
[0109] G2 is a network composed of five parallel graph neural networks, which encode the patterns of the upper garment, sleeves, pants, skirt, and accessories respectively. All the patterns of different clothing parts are simultaneously input into their corresponding graph neural networks. Inside each graph neural network, the garment patterns are regarded as nodes, and the connection relationships and three-dimensional spatial position information between the patterns are used to construct the graph structure of the graph neural network.
[0110] G3 uses the decoder of the Transformer model.
[0111] G4 uses the parametric pattern generation program in the literature (GarmentCodeData: A Dataset of 3D Made-to-Measure Garments With Sewing Patterns, ECCV, 2024).
[0112] The training of G uses the garment pattern feature description text D1 in the garment pattern dataset, the garment pattern parameter configuration file D2 in the garment pattern dataset, and the Adam optimizer. The learning rate is 0.001, the batch size is 32, and it is trained for 50 epochs. After the training of G is completed, the decoding accuracy rate of D5 is 95.8%, meeting the requirement that the decoding accuracy rate ≥ 95%.
[0113] As Figure 3 shown, the training process of G is as follows:
[0114] (a) Input D1 into G1, and it extracts a 512×77-dimensional text embedding feature D3.
[0115] (b) Map D2 through G2 to the garment pattern latent space to generate a garment pattern latent space encoding D4. G2 consists of two parts: quantization encoding and feature encoding. Among them, quantization encoding normalizes the default values of D2 to standardized values [0,1] within the value range, generating 702×1-dimensional standardized values. The encoding dimension is the same as the number of pattern parameters, and no data compression is performed. Feature encoding maps the quantization encoding to the latent space in parallel according to multiple components to form a 128×3-dimensional D4.
[0116] (c) Input D3 and D4 jointly into G3 for decoding to obtain a new clothing sample parameter configuration file D5. The specific process is as follows: First, align D3 and D4, and then parallelly decode the aligned codes into clothing sample parameter configuration quantization codes for multiple components to generate D5;
[0117] (d) Calculate the parameter configuration file loss value , and the formula is as follows:
[0118] ;
[0119] In the formula, and respectively represent the values of the th parameter in D5 and D2, and represents the total number of parameters in D5 or D2;
[0120] (e) Determine whether the number of iterations in steps (a) to (d) is greater than or equal to 10 and is less than 0.0003. If both are true, end the process. Otherwise, update the neural network parameters of G2 and G3 and return to step (a);
[0121] (3) Generate clothing sample data, as Figure 4 shown;
[0122] Input the user-defined clothing sample feature description text D1* "The upper garment is a loose women's short-sleeved T-shirt with a V-neck and the side line expands outwards by 1.5 cm; the lower garment is loose shorts for running" into the trained G, and G outputs the generated clothing sample data X2*, which takes 18 s;
[0123] (4) Construct and train a learning-based 3D virtual fitting model T;
[0124] T consists of a dynamic human body generation module T1, a sample static mapping module T2, a 3D clothing dynamic deformation mapping module T3, and a 3D clothing explicit decoder T4. T1 and T2 are both connected to T3, and T3 is connected to T4;
[0125] T1 consists of an SMPL human model and a text-action generation model. The action parameters generated by the text-action generation model are input into the SMPL human model to obtain a dynamic 3D virtual human. The SMPL human model is sourced from the literature (SMPL: A Skinned Multi-Person Linear Model, SIGGRAPH Asia, 2015.), and the text-action generation model is sourced from the literature (LGTM: Local-to-Global Text-Driven Human Motion Diffusion Model, SIGGRAPH, 2024.). The frame rate of the dynamic 3D virtual human generated by T1 is 30 frames / s;
[0126] T2 uses a fully connected neural network;
[0127] T3 uses a residual neural network;
[0128] T4 uses MeshUDF from the literature (MeshUDF: Fast and Differentiable Meshing of Unsigned Distance Field Networks, ECCV, 2022.), and T4 generates 3D virtual try-on mesh data;
[0129] Training T uses the action feature description text X1 in the virtual fitting dataset, the 3D dynamic virtual try-on mesh data X6 in the virtual fitting dataset, the clothing sample data X2 in the clothing sample dataset, and the Adam optimizer, with a learning rate of 0.001, a batch size of 8, and 50 epochs of training. During the training of T, T1, T2, T3, and T4 are all enabled;
[0130] In the trained T, the mapping speed from the clothing sample to the 3D virtual try-on mesh is 7.5 ms / frame, and the chamfer distance between the fitting result using A5 and the fitting result using the trained T is 3.3 mm;
[0131] As Figure 5 shown, the training process of T is as follows:
[0132] (i) Input X1 into T1, and its output is the dynamic 3D virtual human X3;
[0133] (ii) Transform X2 into 3D clothing static implicit data X4 through T2, and X4 is represented by an unsigned distance field;
[0134] (iii) Input X3 and X4 together into T3, and its output is the 3D clothing dynamic implicit data X5, and X5 is represented by an unsigned distance field;
[0135] (iv) Input X5 into T4, and output the three-dimensional dynamic virtual fitting grid data V1 from it;
[0136] (v) Calculate the chamfer loss value , and the formula is as follows:
[0137] ;
[0138] In the formula, A represents the set of sampled vertices of V1, B represents the set of sampled vertices of X6, represents the number of sampled vertices in A, represents the number of sampled vertices in B, a ∈ A and b ∈ B respectively represent a single vertex in A and B;
[0139] (vi) Determine whether the number of iterations in steps (i) to (v) is greater than or equal to 10 steps, and whether it is less than or equal to 5 mm. If both are yes, end. Otherwise, after updating the neural network parameters of T2 and T3, return to step (i);
[0140] (5) Dynamically optimize X2*, as Figure 6 shown, the specific process is as follows:
[0141] a) Input X2* and the user-defined action feature description text X1* "running" into the trained learning-based three-dimensional virtual fitting model T. Set the non-elastic fabric parameter with a stretching stiffness of 10,000, and enable T1, T2, T3, and T4. Output the three-dimensional dynamic virtual fitting grid data V1* generated by T4;
[0142] b) Select the key action frames (starting and stepping), sample once every 10 frames, and calculate the structural line offset loss value , and the formula is as follows:
[0143] ;
[0144] In the formula, represents the weight value of the key motion frames in V1*, 0.05, represents the total number of key motion frames in V1*, represents the difference in the structural line under the th key motion frame in V1*, represents the weight value of the process action frames in V1*, 0.001, represents the total number of process action frames in V1*, represents the difference in the structural line under the th process action frame in V1*;
[0145] c) Determine whether the number of iterations in steps a) to b) is greater than or equal to 10, whether it converges to 10% of the initial value, whether it is less than or equal to 5 mm, and whether it is less than or equal to 5 mm. If all are yes, then take the X2* of the last iteration in steps a) to b) as the garment sample data V2* applicable to non-elastic fabrics and proceed to the next step. Otherwise, update X2* and return to step a);
[0146] d) Input V2* into T, set the parameters of the real elastic fabric with a tensile stiffness of 100, do not enable T2, and enable T1, T3, and T4. In the A-type standing posture, generate the three-dimensional static virtual fitting grid data V3* output by T4;
[0147] e) Calculate the fabric deformation loss value , and the formula is as follows:
[0148] ;
[0149] In the formula, represents the number of sampled vertices of V3* or V2*. The sampling rate is 0.5, and there are 1548 sampling points. The sampling points meet the requirement of ≥1000 and are evenly distributed in the key areas of the upper and lower garments; represents the spatial coordinate of the th sampled vertex of V3*, represents the spatial coordinate of the th sampled vertex of V2*;
[0150] f) Whether the number of iterations in steps d) to e) is greater than or equal to 10, and whether it is less than or equal to 5 mm. If all are yes, then take the V2* of the last iteration in steps d) to e) as the garment sample data V4* applicable to real elastic fabrics. Otherwise, update V2* and return to step d);
[0151] e) Fine-tune V4* according to personalized needs. Adjust the V-neck depth from 5 cm to 6 cm and save it to meet the industrial production standard.
[0152] The final garment sample result generated by the embodiment is as shown in Figure 7 part (a). The garment sample is a running sportswear sample;
[0153] Perform a virtual sewing experiment on the garment sample shown in Figure 7 part (a). Generate three-dimensional static virtual fitting grid data in the A-type standing posture, as shown in Figure 8 part (a);
[0154] Perform Figure 7The garment pattern in part (a) was applied to a 3D virtual fitting experiment for running movements, generating 3D static virtual fitting grid data in a running posture, as shown in Figure 9 shown in part (a).
[0155] Comparative example
[0156] A method for generating a garment pattern specifically uses the SewingGPT model proposed in the literature (DressCode: Autoregressively Sewing and Generating Garments from Text Guidance, ACM Trans. Graph., 2024) to generate a garment pattern. The SewingGPT model supports inputting short English words to generate a garment pattern. However, the SewingGPT model has the following technical limitations: it does not support Chinese input; it cannot parse or respond to descriptions of specific wearing states or action scenarios of garments (such as the shape of garments in a running posture); its input text format has strict requirements and lacks flexibility;
[0157] To compare the SewingGPT model with the intelligent garment pattern generation model G provided in the embodiments of the present invention, the garment pattern feature description text input into the SewingGPT model is the specific content of the garment pattern that is simplified and translated into English with reference to the Chinese input content in the embodiments of the present invention: "[\"shirt, short sleeves, V-neck, loose\", \"pants, short, loose fit\"]"; Subsequently, the garment pattern generated by the SewingGPT model was imported into the CLO3D garment simulation software, and the same running movement simulation was performed on the garment pattern in the same way as the intelligent garment pattern generation model G in the embodiments of the present invention.
[0158] The final garment pattern result generated by the comparative example is as shown in Figure 7 part (b), and the garment pattern is a running sportswear pattern; Figure 7 The virtual sewing experiment was performed on the garment pattern shown in part (b), generating 3D static virtual fitting grid data in an A-type standing posture, as shown in Figure 8 part (b).
[0159] The Figure 7 garment pattern in part (b) was applied to a 3D virtual fitting experiment for running movements, generating 3D static virtual fitting grid data in a running posture, as shown in Figure 9 part (b).
[0160] From the Figure 7 shown garment pattern results andFigure 8 From the comparison results of the three-dimensional static virtual fitting grid data shown, it can be seen that the garment pattern generated by the embodiment can accurately match the input garment and motion feature descriptions. The overall shape is a loose fit, meeting the requirements of running sports. The pattern shapes of the collar and sleeves also conform to the existing garment industry standards. After virtual sewing, the fit with the human body is good, and the garment has the necessary amount of looseness without any penetration phenomenon. However, the garment pattern generated by the SewingGPT model of the comparative example is short in the upper body and severely expands outward at the lower end, and the lower garment is too long, which does not conform to the input garment feature descriptions. Moreover, the pattern shapes of the collar and sleeves do not conform to the existing garment industry standards.
[0161] From Figure 9 From the comparison diagram of the three-dimensional static virtual fitting grid data of the running posture shown, it can be seen that the garment pattern generated by the embodiment is overall loose and shows natural wrinkling and stretching effects in the running motion simulation, providing sufficient activity space for the movement. However, the garment pattern generated by the SewingGPT model of the comparative example has an unbalanced structure, is too tight in the crotch, and is restricted in movement during dynamic fitting. At the same time, the generated garment pattern data of the comparative example lacks annotation information, and it is necessary to rely on the position of the chest line provided in the SewingGPT training dataset to manually draw the chest line for the generated garment pattern. Moreover, from the analysis of the structural line offset diagram, the matching degree of the structural lines of the garment pattern generated by the embodiment with the human body structural lines during dynamic movement is significantly better than that of the comparative example. Therefore, the garment pattern of the embodiment shows better human body dynamic fitness.
[0162] In addition, the key indicators of the embodiment and the comparative example were compared and evaluated, and the results are shown in Table 1.
[0163] The comparison results in Table 1 show that the garment pattern generation method provided by the present invention is superior to SewingGPT in terms of the success rate of garment pattern generation, the speed of garment pattern generation, the consistency of the description text, the human body dynamic fitness, and the industrial applicability.
[0164] Table 1 Comparison of indicators for garment pattern generation
[0165]
Claims
1. A text-driven intelligent method for generating clothing patterns, characterized in that Input the user-defined clothing pattern feature description text D1* into the trained clothing pattern intelligent generation model G, and output the generated clothing pattern data X2* by it; G consists of a text embedding feature extractor G1, a clothing pattern parallel graph neural network encoder G2, a clothing pattern parallel autoregressive decoder G3, and a parametric pattern generation program G4. G1 and G2 are both connected to G3, and G3 is connected to G4. G2 is only enabled during the model training process, G4 is only enabled during the model application process, and G1 and G3 are both enabled during the training process and the application process; To train G, use the clothing pattern feature description text D1 in the clothing pattern dataset and the clothing pattern parameter configuration file D2 in the clothing pattern dataset. The process is as follows: (a) Input D1 into G1, and extract the text embedding feature D3 by it; (b) Map D2 to the clothing pattern latent space through G2 to generate the clothing pattern latent space encoding D4; (c) Input D3 and D4 together into G3 for decoding to obtain a new clothing pattern parameter configuration file D5; (d) Calculate the parameter profile loss value , and the formula is as follows: ; wherein, and respectively represent the values of the th parameter in D5 and D2, and represents the total number of parameters in D5 or D2; (e) Determine whether the number of iterations in steps (a) to (d) is greater than or equal to 10 and less than 0.0003. If both conditions are met, end the process; otherwise, update the parameters of G2 and G3 and return to step (a).
2. The intelligent generation method of a text-driven clothing pattern according to claim 1, characterized in that In step (b), G2 includes two parts: quantization encoding and feature encoding. Among them, quantization encoding normalizes the default value of D2 to a standardized value according to the value range, and the encoding dimension is the same as the number of pattern parameters without data compression; feature encoding maps the quantization encoding to the latent space in parallel according to multiple components to form D4.
3. A method for intelligently generating a clothing pattern driven by text according to claim 1, characterized in that, The specific process of step (c) is: first align D3 and D4, and then decode the aligned encoding into the clothing pattern parameter configuration quantization encoding in parallel according to multiple components to generate D5.
4. A method for intelligent generation of a text-driven clothing pattern according to claim 1, characterized in that, Also perform dynamic optimization on X2*, and the process is as follows: a) Input X2* and the user-defined action feature description text X1* together into the trained learning-based 3D virtual fitting model T, and use the non-elastic fabric parameter setting, and output the generated 3D dynamic virtual fitting grid data V1* by it; b) Calculate the offset loss value of the structural line , and the formula is as follows: ; Wherein, represents the weight value of the key motion frames in V1*; represents the total number of key motion frames in V1*; represents the difference of the structure line under the th key motion frame in V1*; represents the weight value of the process action frames in V1*; represents the total number of process action frames in V1*; represents the difference of the structure line under the th process action frame in V1*. c) Determine whether the number of iterations in steps a) - b) is greater than or equal to 10 steps, whether it converges to 10% of the initial value, whether it is less than or equal to 5 mm, and whether it is less than or equal to 5 mm. If all are yes, then take the X2* of the last iteration in steps a) - b) as the garment pattern data V2* applicable to non-elastic fabrics and proceed to the next step. Otherwise, update X2* and return to step a); d) Input V2* into the trained T, and use the real elastic fabric parameter setting, and output the generated 3D static virtual fitting grid data V3* by it; e) Calculate the fabric deformation loss value , and the formula is as follows: ; In the formula, represents the number of sampling vertices of V3* or V2*, represents the th spatial coordinate of the sampling vertex of V3*, represents the th spatial coordinate of the sampling vertex of V2*; f) Determine whether the number of iterations in steps d) to e) is greater than or equal to 10 steps, and whether it is less than or equal to 5 mm. If both are yes, then use the V2* of the last iteration in steps d) to e) as the garment pattern data V4* applicable to the real elastic fabric. Otherwise, update V2* and return to step d); e) Fine-tune V4* according to personalized needs and save it.
5. A text-driven intelligent clothing pattern generation method according to claim 4, characterized in that T consists of a dynamic human body generation module T1, a pattern static mapping module T2, a 3D clothing dynamic deformation mapping module T3, and a 3D clothing explicit decoder T4. T1 and T2 are both connected to T3, and T3 is connected to T4; During the application process, when using the non-elastic fabric parameter setting, T1, T2, T3, and T4 are all enabled; When using the real elastic fabric parameter setting, T2 is not enabled, and T1, T3, and T4 are all enabled; During the training process, T1, T2, T3, and T4 are all enabled; To train T, use the action feature description text X1 in the virtual fitting dataset, the 3D dynamic virtual fitting grid data X6 in the virtual fitting dataset, and the clothing pattern data X2 in the clothing pattern dataset. The process is as follows: (ⅰ) Input X1 into T1, and output the dynamic 3D virtual human body X3 by it; (ⅱ) Convert X2 into 3D clothing static implicit data X4 through T2; (iii) Input X3 and X4 into T3 together, and let it output the three-dimensional clothing dynamic implicit data X5; (iv) Input X5 into T4, and let it output the three-dimensional dynamic virtual fitting grid data V1; (v) Calculate the chamfer loss value , and the formula is as follows: ; Wherein, A represents the set of sampling vertices of V1, and B represents the set of sampling vertices of X6, represents the number of sampling vertices in A, represents the number of sampling vertices in B, and a ∈ A and b ∈ B respectively represent individual vertices of A and B; (vi) Determine whether the number of iterations in steps (i) to (v) is greater than or equal to 10 steps, and whether it is less than or equal to 5 mm. If both are yes, end; otherwise, update the parameters of T2 and T3 and return to step (i).
6. The intelligent generation method of a text-driven clothing pattern according to claim 5, wherein, The construction steps of the clothing pattern dataset and the virtual fitting dataset are as follows: S1.
1. Select a three-dimensional virtual human model A1, measure its virtual human model size data A2, set multiple D2 based on A2, and input D2 into the parametric pattern generation program G4 to generate various styles of X2; S1.
2. Manually construct multiple X1, use the text-action generation model A3 to generate three-dimensional virtual human motion data A4, and import A4 into A1 in step S1.1 to construct multiple dynamic three-dimensional virtual human models X3; S1.
3. Combine X3 in step S1.2, and use the fitting program A5 based on physical simulation to virtually stitch X2 generated in step S1.1 to generate X6; S1.
4. Map X6 in step S1.3 to multi-view images A6, and combine with X2 in step S1.1 to fine-tune a specific multi-modal large model A7 so that it can recognize the features at the clothing and pattern levels and automatically generate the corresponding D1 in batches; S1.
5. Integrate A2, D2, X2 in step S1.1 and D1 in step S1.4 to construct the clothing pattern dataset; integrate X1, A4, X3 in step S1.2 and X6 in step S1.3 to construct the virtual fitting dataset.
Citation Information
Patent Citations
Garment template generation method based on cross-domain matching
CN111858997A
Method, device and equipment for generating virtual clothing through multi-modal fusion and storage medium
CN114723843A
Virtual fitting method based on text-driven image generation
CN116205786A
Virtual fitting model training method, virtual fitting method and electronic equipment
CN116416416A
Garment style fusion method and system based on diffusion model
CN117315417A
Cited By
Three-dimensional costume design drawing generation method and device and electronic equipment
CN120726232A
Three-dimensional garment design drawing generation method, device and electronic equipment
CN120726232B
Motion generation method and device based on multi-modal signal, equipment and storage medium
CN120931775A