Model training method, clothing pattern image generation method, device and readable medium

By designing and training a clothing layout prediction model based on the Transformer submodule, the problem of generating industrial clothing layouts in the prior art is solved, and user-friendly input and generation speed are improved.

CN119850782BActive Publication Date: 2025-06-06ZHEJIANG SHENFU ARTIFICIAL INTELLIGENCE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510329430.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-06
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

The prior art is difficult to generate industrial clothing styles suitable for industrial production, and cannot provide user-friendly input, and the generation speed is also slow.

Method used

By designing and training a clothing layout prediction model based on a Transformer submodule, the DiT architecture can accept multiple input types and generate clothing layout images suitable for industrial production.

Benefits of technology

It realizes the generation of industrial clothing styles suitable for industrial production, provides user-friendly input (supports text and image input), and significantly improves generation speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119850782B_ABST
    Figure CN119850782B_ABST
Patent Text Reader

Abstract

The present application discloses a model training method, a clothing pattern image generation method, a device and a readable medium, which obtain training data including text data and / or image data and encode the training data to obtain a training conditional coding vector; determine the original token sequence and the token position coding according to the initial data of the clothing pattern and the preset clothing pattern coding rule; obtain a noisy token sequence according to each original token sequence and a preset time step; input each training conditional coding vector, the token position coding, the preset time step and each noisy token sequence into an initial clothing pattern prediction model to obtain a denoised token sequence output by the initial clothing pattern prediction model for each training conditional coding vector; the denoised token sequence can be used to generate a clothing pattern image corresponding to the training data; and adjust the parameters of the initial clothing pattern prediction model until a converged and fitted clothing pattern prediction model is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular, to a model training method, a clothing pattern image generation method, a device and a readable medium. Background Art

[0002] At present, the generation of garment sewing patterns mainly relies on data-driven, that is, using a variety of deep learning networks for training and prediction on large-scale garment pattern datasets provided by academia.

[0003] In 2021 and 2024, the academic community released clothing pattern databases of 20,000 and 100,000 respectively, which included pattern definition files, clothing 3D models, renderings, etc. On this basis, a variety of technical solutions for generating 2D clothing patterns have been proposed in this field, such as prediction methods based on standard pose images, point cloud direct prediction technology, prediction methods based on hierarchical deep network architecture, sewing pattern generation methods based on potential diffusion models, proxy technology based on domain-specific languages, prediction technology based on fine-tuning multimodal large models, and so on.

[0004] However, these technical solutions generate academic patterns instead of industrial patterns that can be formally used in industrial production, and they cannot provide user-friendly input (only text input is supported but not clothing hand-drawn line drawing input, or although text input is supported, the text granularity is too low), and the speed of generating clothing patterns is limited. Therefore, there is an urgent need for a clothing pattern generation method that can generate industrial clothing patterns, provide user-friendly input, and generate faster. Summary of the invention

[0005] The present application aims to solve one of the technical problems in the related art to a certain extent. To this end, the present application provides a model training method, a clothing pattern image generation method, a device and a readable medium.

[0006] As a first aspect of the present application, a method for training a clothing pattern prediction model is provided, wherein the method comprises:

[0007] Acquire training data, and encode the training data to obtain a training conditional encoding vector; wherein the type of the training data includes text data and / or image data;

[0008] According to the initial data of the clothing pattern and the preset clothing pattern coding rules, determine the original token sequence and token position coding;

[0009] According to each of the original token sequences and the preset time step, a noisy token sequence is constructed;

[0010] Input each of the training condition coding vectors, the token position coding, the preset time step and each of the noisy token sequences into an initial clothing pattern prediction model to obtain a denoised token sequence output by the initial clothing pattern prediction model for each of the training condition coding vectors; wherein the clothing pattern prediction model is a diffusion model based on a Transformer submodule, and the denoised token sequence can be used to generate a clothing pattern image corresponding to the training data;

[0011] According to each of the noisy token sequences, each of the training conditional coding vectors and each of the denoised token sequences, the parameters of the initial clothing pattern prediction model are adjusted until a converged and fitted clothing pattern prediction model is obtained.

[0012] Optionally, the determining of the original token sequence and the token position coding according to the initial data of the clothing pattern and the preset clothing pattern coding rule comprises:

[0013] According to a preset clothing pattern coding rule, the initial data of the clothing pattern is encoded to obtain the coding data of the clothing pattern; wherein the preset clothing pattern coding rule includes: each clothing pattern includes a first preset number of patterns, each of the patterns includes a second preset number of edges, each of the edges includes a starting point position parameter, a Bezier curve control point position parameter, an edge arc parameter, a stitching point parameter and a stitching mark, and the edge arc parameter includes an edge arc radius, a primary and secondary arc parameter and a sweep direction;

[0014] Each of the encoded data is converted into an original token sequence, and a token position code is generated.

[0015] Optionally, converting each of the encoded data into an original token sequence includes:

[0016] Each edge data in each encoded data is converted into a token one by one, and an original token sequence corresponding to each encoded data is obtained; wherein the token position code includes a plate-level position code and an edge-level position code, the plate-level position code is used to distinguish the tokens corresponding to different plates, and the edge-level position code is used to distinguish the order of different tokens corresponding to the same plate.

[0017] Optionally, the clothing pattern prediction model includes a plurality of DiT blocks and a multi-layer perceptron, the DiT block includes a decoupled cross attention module, the decoupled cross attention module includes a text key value projection matrix, an image key value projection matrix and a query projection matrix,

[0018] The text key value projection matrix is ​​used to generate key values ​​according to the conditional coding vector corresponding to the text data, the image key value projection matrix is ​​used to generate key values ​​according to the conditional coding vector corresponding to the image data, and the query projection matrix is ​​used to generate query probes according to the noisy token sequence.

[0019] Optionally, acquiring training data includes:

[0020] Acquire clothing design parameter files and / or three-dimensional clothing model images;

[0021] Filtering the clothing design parameter file, and inputting the filtered clothing design parameter file into a preset large language model to obtain text training data output by the preset large language model;

[0022] Rendering the three-dimensional clothing model image, and inputting the rendering result into a preset stable diffusion model to obtain a stylized sketch output by the preset stable diffusion model;

[0023] The stylized sketch is binarized to obtain image training data; wherein the training data includes the text training data and / or the image training data.

[0024] As a second aspect of the present application, a clothing pattern generation method is provided, wherein the clothing pattern generation method comprises:

[0025] Acquire user input data, and encode the user input data to obtain an input conditional encoding vector; wherein the type of the user input data includes text data and / or image data;

[0026] Inputting the input conditional coding vector into the clothing pattern prediction model trained by the clothing pattern prediction model training method provided in the first aspect of the present application, and obtaining a denoised token sequence output by the clothing pattern prediction model for the input conditional coding vector;

[0027] A clothing pattern image corresponding to the user input data is generated according to the denoising token sequence.

[0028] Optionally, the method further comprises:

[0029] Performing data cleaning on each of the user input data and the corresponding clothing pattern images to obtain a feedback data set;

[0030] According to the feedback data set, the clothing pattern prediction model is adjusted until the clothing pattern image generated by the clothing pattern prediction model after the adjustment meets the preset conditions.

[0031] Optionally, the method further comprises:

[0032] Generate an initial industrial sample image according to a preset seam addition rule and the generated garment pattern image;

[0033] According to preset design parameter rules, the initial industrial sample drawing is updated to obtain a final industrial sample drawing.

[0034] As a third aspect of the present application, an electronic device is provided, wherein the electronic device includes:

[0035] one or more processors;

[0036] A memory having one or more computer programs stored thereon, wherein when the one or more computer programs are executed by the one or more processors, the one or more processors implement any of the following:

[0037] The training method of the clothing pattern prediction model provided in the first aspect of the present application;

[0038] The second aspect of the present application provides a method for generating a clothing pattern image.

[0039] As a fourth aspect of the present application, a computer-readable medium is provided, on which a computer program is stored, wherein when the computer program is executed by a processor, any of the following is implemented:

[0040] The training method of the clothing pattern prediction model provided in the first aspect of the present application;

[0041] The second aspect of the present application provides a method for generating a clothing pattern image.

[0042] Through the training method of the clothing pattern prediction model provided by the embodiment of the present application, training data including text data and / or image data is obtained, and the training data is encoded to obtain a training conditional encoding vector; training data is obtained, and the training data is encoded to obtain a training conditional encoding vector; according to the initial data of the clothing pattern and the preset clothing pattern encoding rules, the original token sequence and the token position encoding are determined; according to each of the original token sequences and the preset time step, a noisy token sequence is constructed; each of the training conditional encoding vectors, the token position encoding, the preset time step and Each of the noisy token sequences is input into the initial clothing pattern prediction model to obtain the denoised token sequence output by the initial clothing pattern prediction model for each of the training conditional coding vectors; wherein the clothing pattern prediction model is a diffusion model based on the Transformer submodule, and the denoised token sequence can be used to generate a clothing pattern image corresponding to the training data; according to each of the noisy token sequences, each of the training conditional coding vectors and each of the denoised token sequences, the initial clothing pattern prediction model is adjusted until a converged and fitted clothing pattern prediction model is obtained. The clothing pattern prediction model can be further used to quickly and accurately generate clothing pattern images suitable for industrial production based on the user's text input and / or image input, so as to achieve the generation of industrial clothing patterns that can be suitable for industrial production, provide user-friendly input and increase the speed of generating industrial clothing patterns. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The present application is further described below in conjunction with the accompanying drawings:

[0044] Figure 1 It is a flowchart of an implementation method of a training method of a clothing pattern prediction model provided in an embodiment of the present application;

[0045] Figure 2 It is a flowchart of another implementation of the training method of the clothing pattern prediction model provided in the embodiment of the present application;

[0046] Figure 3 It is a schematic diagram of an implementation method of the coding data of the clothing pattern provided in the embodiment of the present application;

[0047] Figure 4 It is a flowchart of another implementation of the method for training the clothing pattern prediction model provided in the embodiment of the present application;

[0048] Figure 5 It is a flowchart of an implementation method of the clothing pattern image generation method provided in the embodiment of the present application;

[0049] Figure 6is a flowchart of another implementation of the method for generating a clothing pattern image provided in an embodiment of the present application;

[0050] Figure 7 It is a data network schematic diagram provided in an embodiment of the present application;

[0051] Figure 8 It is a flowchart of another implementation of the method for generating a clothing pattern image provided in an embodiment of the present application;

[0052] Fig. 9 is a schematic diagram of an implementation method of using the training DiT architecture provided in the embodiment of the present application to generate clothing pattern images;

[0053] Fig.10 It is a module diagram of an implementation of an electronic device provided in an embodiment of the present application;

[0054] Fig.11 It is a schematic diagram of a computer-readable medium provided in an embodiment of the present application.

[0055] Description of Reference Numerals

[0056] 101: Processor 102: Memory

[0057] 103: I / O interface 104: bus DETAILED DESCRIPTION

[0058] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements with the same or similar functions. The embodiments based on the embodiments are intended to be used to explain the present application and cannot be understood as limiting the present application.

[0059] References to "one embodiment" or "an example" or "an example" in this specification mean that a particular feature, structure, or characteristic described in conjunction with the embodiment itself may be included in at least one embodiment disclosed in the present application. The appearance of the phrase "in one embodiment" in various places in the specification does not necessarily refer to the same embodiment.

[0060] At present, the generation of garment sewing patterns mainly relies on data-driven, that is, using a variety of deep learning networks for training and prediction on large-scale garment pattern datasets provided by academia.

[0061] In 2021 and 2024, the academic community released clothing pattern databases of 20,000 and 100,000 respectively, which included pattern definition files, clothing 3D models, renderings, etc. On this basis, the field has proposed a variety of technical solutions for generating 2D clothing patterns, such as the following technical solutions:

[0062] 1. Prediction method based on standard posture pictures: Use a special image recognition program (convolutional neural network) to predict the size parameters of each pattern required for making clothes by analyzing photos of standard standing postures of human bodies;

[0063] 2. Point cloud direct prediction technology: Using a deep learning network, by learning clothing store cloud data, it can directly generate the flat pattern outline of the clothes and automatically mark the stitching positions between different pieces of clothing;

[0064] 3. Prediction method based on hierarchical deep network architecture: Using advanced image understanding system, the input image is first analyzed and encoded, and then through a multi-level decoding process, the complete structure of the clothing pattern including the shape and connection relationship of each piece of clothing is gradually restored;

[0065] 4. Sewing pattern generation method based on latent diffusion model: First, an intelligent compression program is trained to compress complex clothing pattern information into concise digital codes, and then an intelligent generation program is trained to gradually remove noise from these codes to generate complete clothing patterns;

[0066] 5. Agent technology based on domain-specific language: First, the text and image inputs are understood through a pre-trained multimodal large model, and then the trained network is used to predict clothing pattern parameters;

[0067] 6. Prediction technology based on fine-tuning multimodal large models: Through supervised fine-tuning of large models that already have text and visual capabilities, it is enabled to have the ability to generate numerical values, and then form clothing patterns through customized rules.

[0068] However, these technical solutions still have some serious drawbacks:

[0069] 1. The clothing patterns that can be generated are still only academic patterns, and there is still a certain distance from actual industrial production in terms of pattern complexity and category diversity;

[0070] 2. Unable to provide user-friendly input. Some solutions only support text input but not image input (such as hand-drawn clothing line drawings). Although some solutions support text input, the text granularity of the paired annotation of clothing pattern and text is not fine enough, and the complete design features of the clothing pattern are missing, or the daily description terms of the clothing pattern are missing;

[0071] 3. The speed of generating clothing patterns is limited because the encoding compression of the clothing patterns is not enough or the generation pipeline is not end-to-end (the understanding speed of the large model used). As a result, the time to generate a pattern is often too long, exceeding the time limit that users are willing to wait.

[0072] In this regard, the applicant of this application proposes to re-characterize the clothing pattern based on the initial data of the clothing pattern by using pre-set clothing pattern encoding rules, and to make full use of the DiT (Diffusion Transformer) architecture that can accept multiple input types to control the characteristics of image generation, to design and train a lightweight model architecture, which can break through the limitations of pattern complexity and category diversity, break through the bottleneck of imperfect user input modal support, and remove the limitations on the generation speed of clothing patterns.

[0073] As a first aspect of the embodiment of the present application, a method for training a clothing pattern prediction model is provided, such as Figure 1 As shown, the method may include the following steps:

[0074] Step S110, acquiring training data, and encoding the training data to obtain a training conditional encoding vector; wherein the type of the training data includes text data and / or image data;

[0075] Step S120, determining the original token sequence and token position code according to the initial data of the clothing pattern and the preset clothing pattern coding rule;

[0076] Step S130, constructing a noisy token sequence according to each of the original token sequences and the preset time step;

[0077] Step S140, inputting each of the training condition coding vectors, the token position coding, the preset time step and each of the noisy token sequences into an initial clothing pattern prediction model, and obtaining a denoised token sequence output by the initial clothing pattern prediction model for each of the training condition coding vectors; wherein the clothing pattern prediction model is a diffusion model based on a Transformer submodule, and the denoised token sequence can be used to generate a clothing pattern image corresponding to the training data;

[0078] Step S150, adjusting the parameters of the initial clothing pattern prediction model according to each of the noisy token sequences, each of the training conditional coding vectors, and each of the denoised token sequences, until a converged and fitted clothing pattern prediction model is obtained.

[0079] The embodiment of the present application does not impose any special limitation on the execution order of the above-mentioned step S110 and steps S120 to S130, and they can be executed synchronously or asynchronously.

[0080] Among them, it can be understood that the above steps S110-S150 are equivalent to forming a DiT architecture, and the clothing pattern prediction model is a diffusion model based on the Transformer sub-module, which is only the core model in the entire DiT architecture. In particular, the core model needs to be iteratively trained, that is, the above steps S130, S140, and S150 need to be executed multiple times until the clothing pattern prediction model meets the conditions for terminating training (usually the model converges and fits).

[0081] The embodiment of the present application does not specifically limit how to execute the above step S110. For example, the training data may be obtained first, and then the training data may be encoded using a text and image encoder.

[0082] The initial data of clothing patterns refers to academic data that has been published in this field, such as pattern definition files, clothing three-dimensional models, renderings, and the like.

[0083] The embodiment of the present application does not specifically limit the preset clothing pattern coding rules. For example, the clothing pattern coding rules can be preset in the following ways: how many pieces each clothing pattern includes, how many edges each piece includes, what parameters each edge specifically includes, etc. Through these ways, a large amount of different clothing pattern data can be uniformly encoded, which is beneficial to subsequent data processing and industrial production.

[0084] Among them, constructing a noisy token sequence according to each original token sequence and a preset time step means shuffling the tokens in each original token sequence and gradually adding noise to it according to the preset time step, which belongs to the forward diffusion process in the DiT architecture, while the core model (i.e., the clothing pattern prediction model) predicts how much noise should be removed, which belongs to the reverse denoising process in the DiT architecture.

[0085] Through the training method of the clothing pattern prediction model provided by the embodiment of the present application, training data including text data and / or image data is obtained, and the training data is encoded to obtain a training conditional encoding vector; training data is obtained, and the training data is encoded to obtain a training conditional encoding vector; according to the initial data of the clothing pattern and the preset clothing pattern encoding rules, the original token sequence and the token position encoding are determined; according to each of the original token sequences and the preset time step, a noisy token sequence is constructed; each of the training conditional encoding vectors, the token position encoding, the preset time step and Each of the noisy token sequences is input into the initial clothing pattern prediction model to obtain the denoised token sequence output by the initial clothing pattern prediction model for each of the training conditional coding vectors; wherein the clothing pattern prediction model is a diffusion model based on the Transformer submodule, and the denoised token sequence can be used to generate a clothing pattern image corresponding to the training data; according to each of the noisy token sequences, each of the training conditional coding vectors and each of the denoised token sequences, the initial clothing pattern prediction model is adjusted until a converged and fitted clothing pattern prediction model is obtained. The clothing pattern prediction model can be further used to quickly and accurately generate clothing pattern images suitable for industrial production based on the user's text input and / or image input, so as to achieve the generation of industrial clothing patterns that can be suitable for industrial production, provide user-friendly input and increase the speed of generating industrial clothing patterns.

[0086] The applicant of the present application further proposes that the initial data of the clothing pattern can be encoded according to the preset clothing pattern encoding rules, and then each edge in the encoded data can be converted into a basic data unit - token that can be recognized by the core model. Accordingly, in some embodiments, the original token sequence and token position encoding (i.e., what step S120 involves) are determined according to the initial data of the clothing pattern and the preset clothing pattern encoding rules, such as Figure 2 As shown, the following steps may be included:

[0087] Step S210, encoding the initial data of the clothing pattern according to a preset clothing pattern encoding rule to obtain the encoding data of the clothing pattern; wherein the preset clothing pattern encoding rule includes: each clothing pattern includes a first preset number of panels, each panel includes a second preset number of edges, each edge includes a starting point position parameter, a Bezier curve control point position parameter, an edge arc parameter, a stitching point parameter and a stitching mark, and the edge arc parameter includes an edge arc radius, a primary and secondary arc parameter and a sweep direction;

[0088] Step S220, converting each of the encoded data into an original token sequence respectively, and generating a token position code.

[0089] Among them, the embodiment of the present application encodes the clothing pattern as follows: each clothing pattern is composed of multiple panels, each panel is a closed shape, and is composed of multiple edges. In order to simplify the representation, a Bezier curve is used to record the geometric shape of each edge. Specifically, the geometric shape of each edge can be represented by the following parameters: starting point position parameter (the position of the starting point of the edge in three-dimensional space), Bezier curve control point position parameter (three-dimensional coordinates of Bezier curve control point), edge arc parameter, stitching point parameter (calculated as the average value of the stitching edge endpoints) and stitching flag (a binary flag used to indicate whether the edge needs to be stitched). If the edge is an arc, three parameters are used to represent the radius of the arc, the main arc or the secondary arc, and the sweeping direction). In order to represent all clothing patterns in a standardized manner, the total number of panels included in each clothing pattern and the total number of edges included in each panel are also specified.

[0090] like Figure 3 As shown, it is a schematic diagram of an implementation method of the coding data of the clothing pattern provided in an embodiment of the present application. The coding data of each clothing pattern includes multiple panel data, each of which has a panel position code 1, 2, 3, ..., i, ..., M, and each panel data includes multiple edge data. For example, the i-th panel data includes a panel position code (the "i" sequence shown in the figure), an edge position code (the "1, 2, 3, 4, 5, ..., N" sequence shown in the figure) and edge data (the " , , , , , ..., 0" sequence). For example, the edge data Including the starting position parameters (shown in the figure as " , , ”), Bezier curve control point position parameters (shown in the figure “ , , , , , ”), edge arc parameters (shown in the figure “ , , ”), stitching point parameters (shown in the figure “ , , ”) and stitching marks (shown in the figure as “ ”).

[0091] It should be emphasized that Figure 3 This is only an exemplary description, and the embodiment of the present application does not specifically limit the specific data structure of the encoded data, as long as it complies with the preset clothing pattern encoding rules.

[0092] The preset clothing pattern coding rules stipulate the total number of patterns included in each clothing pattern and the total number of edges included in each pattern. Accordingly, the coded data includes a first preset number (denoted as M) of pattern data, and the pattern data includes a second preset number (denoted as N) of edge data. The applicant of the present application also proposes that for coded data with less than M pattern data or less than N edge data, zero vectors can be used for padding, and for coded data with more than M pattern data or more than N edge data, some redundant data can be removed, so as to ensure that all coded data have the same data length, that is, a total of MxN edge data.

[0093] In some embodiments, converting each of the encoded data into an original token sequence may include: converting each of the edge data in the encoded data into a token one by one, and obtaining an original token sequence corresponding to each of the encoded data one by one; wherein the token position encoding includes panel-order embedding and edge-order embedding, the panel-order position encoding is used to distinguish the tokens corresponding to different panels, and the edge-order position encoding is used to distinguish the order of different tokens corresponding to the same panel.

[0094] The encoding data of each clothing pattern includes a plurality of edge data, and each edge is converted into a token, so that the encoding data of each clothing pattern is converted into an original token sequence.

[0095] Among them, the embodiment of the present application does not specifically limit the specific data structure of the token position encoding, as long as the function of "distinguishing the tokens corresponding to different plates" and the function of "distinguishing the order of different tokens corresponding to the same plate" can be realized.

[0096] The applicant of this application further proposed that in order to more accurately control the generation of images, decoupled cross-attention layers can be used to process text features and image features through two independent "key" and "value" projection matrices, and process the edge features of the pattern through a shared query projection matrix, helping the DiT architecture to better understand and fuse information from different input modalities (text, images), thereby facilitating the generation of higher quality clothing pattern images.

[0097] Accordingly, in some embodiments, the clothing pattern prediction model includes multiple DiT blocks and a multi-layer perceptron, the DiT block includes a decoupled cross-attention module, the decoupled cross-attention module includes a text key value projection matrix, an image key value projection matrix and a query projection matrix, the text key value projection matrix is ​​used to generate a key value according to the conditional coding vector corresponding to the text data, the image key value projection matrix is ​​used to generate a key value according to the conditional coding vector corresponding to the image data, and the query projection matrix is ​​used to generate a query probe according to the noisy token sequence.

[0098] The decoupled cross-attention module establishes a connection between semantic and visual / structural information by learning the correspondence between the conditional encoding vector and the encoded data of the clothing pattern (corresponding token), thereby better understanding the meaning of the data.

[0099] The applicant of the present application further proposes that in order to help the DiT architecture better understand the user's design requirements, multi-level text descriptions and stylized design sketches can be generated as training data, so that the trained DiT architecture has better performance and more accurate processing results. Figure 4 As shown, the obtaining of training data may include the following steps:

[0100] Step S310, obtaining a clothing design parameter file and / or a three-dimensional clothing model image;

[0101] Step S320, filtering the clothing design parameter file, and inputting the filtered clothing design parameter file into a preset large language model to obtain text training data output by the preset large language model;

[0102] Step S330, rendering the three-dimensional clothing model image, and inputting the rendering result into a preset stable diffusion model to obtain a stylized sketch output by the preset stable diffusion model;

[0103] Step S340, binarizing the stylized sketch to obtain image training data; wherein the training data includes the text training data and / or the image training data.

[0104] Among them, the embodiment of the present application divides the text description of the clothing pattern into two levels. The first level is a brief and informative description, which mainly introduces the types and basic design features of clothing, and the second level is a more detailed description, involving the specific design of each part of the clothing. In order to generate these multi-level text descriptions, a data annotation process can be designed to automatically generate these design descriptions using the contextual understanding ability of the large language model. Specifically, first remove the content that is not related to the design from the clothing design parameter file (that is, filter the clothing design parameter file), and then use the large language model combined with the filtered version of the clothing design parameter file to automatically generate a detailed description of the clothing pattern (as close to the description habits in daily life as possible).

[0105] In addition, in order to support image modality input, the 3D clothing model image is first rendered, and then a fine-tuned stable diffusion model is used in combination with edge control to generate a stylized sketch, and finally binarization is performed to make the image training data as close as possible to the hand-drawn style of professional clothing designers.

[0106] As a second aspect of the embodiment of the present application, a method for generating a clothing pattern image is provided, such as Figure 5 As shown, the clothing pattern generation method may include the following steps:

[0107] Step S410, obtaining user input data, and encoding the user input data to obtain an input conditional encoding vector; wherein the type of the user input data includes text data and / or image data;

[0108] Step S420, inputting the input conditional coding vector into a clothing pattern prediction model trained according to a clothing pattern prediction model training method, and obtaining a denoised token sequence output by the clothing pattern prediction model for the input conditional coding vector;

[0109] Step S430: generating a clothing pattern image corresponding to the user input data according to the denoising token sequence.

[0110] Among them, it can be understood that the clothing pattern prediction model after training is based on the learned denoising mode to iteratively denoise the input (that is, the input conditional encoding vector). After the training is completed, the clothing pattern prediction model has learned the mapping logic of "text semantics → pattern features" and / or "image vision → pattern structure". At this time, input text (describing requirements) and / or images (visual references), and the clothing pattern prediction model will automatically generate a denoising token sequence corresponding to the input conditional encoding vector based on the associated knowledge learned during training.

[0111] Among them, the embodiment of the present application does not make any specific limitation on how to execute step S430, which belongs to the conventional operation of the DiT architecture. For example, it can be implemented through a series of operations such as denormalization, decoding, semantic parsing, image drawing, and image rendering, so it will not be repeated here.

[0112] In addition, the structure and other contents of the clothing pattern prediction model have been described in detail above, so they will not be repeated here.

[0113] The method for generating a clothing pattern image provided by the embodiment of the present application first obtains training data including text data and / or image data, encodes the training data, and obtains a training conditional encoding vector; obtains training data, encodes the training data, and obtains a training conditional encoding vector; determines an original token sequence and a token position encoding according to initial data of the clothing pattern and a preset clothing pattern encoding rule; constructs a noisy token sequence according to each of the original token sequences and a preset time step; and converts each of the training conditional encoding vectors, the token position encoding, the preset time step, and Each of the noisy token sequences is input into the initial clothing pattern prediction model to obtain the denoised token sequence output by the initial clothing pattern prediction model for each of the training conditional coding vectors; wherein the clothing pattern prediction model is a diffusion model based on the Transformer submodule, and the denoised token sequence can be used to generate a clothing pattern image corresponding to the training data; according to each of the noisy token sequences, each of the training conditional coding vectors and each of the denoised token sequences, the initial clothing pattern prediction model is adjusted until a converged and fitted clothing pattern prediction model is obtained. Finally, the user input data is obtained, and the user input data is encoded, the input conditional coding vector is obtained and input into the clothing pattern prediction model, and the denoised token sequence corresponding to the input conditional coding vector is obtained, and the clothing pattern image corresponding to the user input data is generated accordingly, so as to realize the generation of industrial clothing patterns suitable for industrial production, provide user-friendly input and improve the speed of generating industrial clothing patterns.

[0114] The applicant of the present application further proposes that after the trained clothing pattern prediction model is put into use, the clothing pattern prediction model can be further trained using the user input data fed back by the user and its corresponding clothing pattern image, so as to further improve the performance of the clothing pattern prediction model. Figure 6 As shown, the method may further include the following steps:

[0115] Step S510, performing data cleaning on each of the user input data and its corresponding clothing pattern image to obtain a feedback data set;

[0116] Step S520: adjusting the parameters of the clothing pattern prediction model according to the feedback data set until the clothing pattern image generated by the clothing pattern prediction model after the adjustment meets the preset conditions.

[0117] The embodiments of the present application do not specifically limit how to perform data cleaning. For example, it may include anomaly detection, standardization processing, etc.

[0118] Among them, the embodiments of the present application are not limited to adjusting the parameters of the clothing pattern prediction model according to the feedback data set, but can also appropriately adjust the parameters of other parts of the entire DiT architecture except the clothing pattern prediction model, such as the parameters of the encoder, the parameters of the decoder, etc.

[0119] The embodiment of the present application does not specifically limit the preset conditions involved in step S520. For example, the similarity between the generated clothing pattern image and the actual clothing pattern image may reach a certain threshold.

[0120] like Figure 7 As shown, it is a schematic diagram of a data network provided in an embodiment of the present application. It can be seen that the data storage layer includes an original pattern database (which can be used to store the initial data of the clothing pattern), a multimodal storage center (which can be used to store the constructed text training data and image training data), and a user feedback data pool (which can be used to store the user input data of user feedback and its corresponding clothing pattern images). The data processing layer includes a data completion engine (which can be used to construct text training data and image training data), and a data cleaning pipeline (which can be used to clean the user input data of user feedback and its corresponding clothing pattern images). The model layer includes a multimodal generation model (that is, the complete DiT architecture provided in an embodiment of the present application), and a continuous learning framework (which can be used to learn the user input data of user feedback and its corresponding clothing pattern images, and then adjust the parameters of the multimodal generation model). The data flow in the entire data network is divided into an initial data flow and a user feedback loop data flow. Figure 7In the figure, the solid arrows represent the initial data flow, corresponding to the initial training process of the DiT architecture, and the dotted arrows represent the user feedback loop data flow, corresponding to the re-training process of the DiT architecture.

[0121] The applicant of the present application further proposes that after the garment pattern image corresponding to the user input data is generated, industrial seams can be further added and design parameter checks can be performed to make the final industrial pattern image more accurate. Figure 8 As shown, the method may further include the following steps:

[0122] Step S610, generating an initial industrial sample image according to a preset seam addition rule and the generated garment pattern image;

[0123] Step S620, updating the initial industrial sample drawing according to preset design parameter rules to obtain a final industrial sample drawing.

[0124] Among them, the embodiments of the present application do not make specific limitations on the preset seam allowance adding rules and the preset design parameter rules. For example, a 2 cm seam allowance can be added to each edge, and the ranges of various design parameters such as sleeve width and neckline can be preset.

[0125] like Fig. 9As shown, it is a schematic diagram of an implementation method of using the training DiT architecture provided in the embodiment of the present application to generate clothing pattern images. First, three types of inputs are constructed: a brief description of the text (such as a mid-length sleeve slim dress, knee-length), a detailed description of the text (such as a slim dress, knee-length. The collar is a half-high collar, the sleeves are slightly tight three-quarter sleeves. The skirt is relatively tight...), and a hand-drawn sketch. Then, one of the brief description and the detailed description of the text, as well as the hand-drawn sketch, is used as a condition, and the condition is encoded into a training conditional encoding vector using a contrastive language-image multimodal model (Contrastive Language-Image Pre-Processing, CLIP); according to the preset clothing pattern encoding rules, the initial data of the clothing pattern, and the preset time step, a noisy token sequence, a pattern-level position encoding, and an edge-level position encoding are constructed, and used together with the time step encoding as the model input. The model input at this time is equivalent to an incomplete pattern. Furthermore, the model input and the training conditional encoding vector are used together to train the clothing pattern prediction model. The clothing pattern prediction model is a diffusion model based on the Transformer submodule, including multiple DiT blocks and multilayer perceptrons (MLP). Each DiT block includes a normalization layer, a self-attention layer, a decoupled cross attention layer, and a linear layer. The denoised token sequence output by the clothing pattern prediction model can be used to generate a clothing pattern image corresponding to the conditional input.

[0126] As a third aspect of the embodiments of the present application, an electronic device is provided, wherein: Fig.10 As shown, the electronic device includes:

[0127] One or more processors 101;

[0128] The memory 102 stores one or more computer programs. When the one or more computer programs are executed by the one or more processors 101, the one or more processors 101 implement any of the following:

[0129] The training method of the clothing pattern prediction model provided in the first aspect of the embodiment of the present application;

[0130] A second aspect of an embodiment of the present application provides a method for generating a clothing pattern image.

[0131] The electronic device may further include one or more I / O interfaces 103 connected between the processor 101 and the memory 102 and configured to implement information interaction between the processor 101 and the memory 102 .

[0132] Among them, the processor 101 is a device with data processing capabilities, including but not limited to a central processing unit (CPU), etc.; the memory 102 is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory (FLASH); the I / O interface (read-write interface) is connected between the processor and the memory, and can realize information exchange between the processor and the memory, including but not limited to a data bus (Bus), etc.

[0133] In some embodiments, the processor 101 , the memory 102 , and the I / O interface 103 are connected to each other via a bus 104 , and further connected to other components of the computing device.

[0134] As a fourth aspect of the embodiment of the present application, Fig.11 As shown, a computer readable medium is provided, on which a computer program is stored, wherein when the computer program is executed by a processor, any of the following is implemented:

[0135] The training method of the clothing pattern prediction model provided in the first aspect of the embodiment of the present application;

[0136] A second aspect of an embodiment of the present application provides a method for generating a clothing pattern image.

[0137] It will be understood by those skilled in the art that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. Accordingly, the computer program can be stored in a non-volatile computer-readable storage medium, and the computer program can implement the method of any of the above-mentioned embodiments when executed. Among them, any reference to memory, storage, database or other media used in the embodiments provided in the embodiments of the present application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0138] The above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Those skilled in the art should understand that the present application includes but is not limited to the contents described in the drawings and the above specific embodiments. Any modification that does not deviate from the functional and structural principles of the present application will be included in the scope of the claims.

Claims

1. A training method for a clothing pattern prediction model, characterized in that: The method comprises: Acquire training data, and encode the training data to obtain a training conditional encoding vector; wherein the type of the training data includes text data and / or image data; According to the initial data of the clothing pattern and the preset clothing pattern coding rules, determine the original token sequence and token position coding; According to each of the original token sequences and the preset time step, a noisy token sequence is constructed; Input each of the training condition coding vectors, the token position coding, the preset time step and each of the noisy token sequences into an initial clothing pattern prediction model to obtain a denoised token sequence output by the initial clothing pattern prediction model for each of the training condition coding vectors; wherein the clothing pattern prediction model is a diffusion model based on a Transformer submodule, and the denoised token sequence can be used to generate a clothing pattern image corresponding to the training data; According to each of the noisy token sequences, each of the training conditional coding vectors, and each of the denoised token sequences, adjusting the parameters of the initial clothing pattern prediction model until a converged and fitted clothing pattern prediction model is obtained; The clothing pattern prediction model includes a plurality of DiT blocks and a multi-layer perceptron, wherein the DiT block includes a decoupled cross attention module, and the decoupled cross attention module includes a text key value projection matrix, an image key value projection matrix and a query projection matrix. The text key value projection matrix is ​​used to generate key values ​​according to the conditional coding vector corresponding to the text data, the image key value projection matrix is ​​used to generate key values ​​according to the conditional coding vector corresponding to the image data, and the query projection matrix is ​​used to generate query probes according to the noisy token sequence.

2. The method according to claim 1, characterized in that The method of determining the original token sequence and the token position encoding according to the initial data of the clothing pattern and the preset clothing pattern encoding rules includes: According to a preset clothing pattern coding rule, the initial data of the clothing pattern is encoded to obtain the coding data of the clothing pattern; wherein the preset clothing pattern coding rule includes: each clothing pattern includes a first preset number of patterns, each of the patterns includes a second preset number of edges, each of the edges includes a starting point position parameter, a Bezier curve control point position parameter, an edge arc parameter, a stitching point parameter and a stitching mark, and the edge arc parameter includes an edge arc radius, a primary and secondary arc parameter and a sweep direction; Each of the encoded data is converted into an original token sequence, and a token position code is generated.

3. The method according to claim 2, characterized in that The converting each of the encoded data into an original token sequence comprises: The data of each edge in each encoded data is converted into a token one by one, and an original token sequence corresponding to each encoded data is obtained; wherein the token position code includes a plate-level position code and an edge-level position code, the plate-level position code is used to distinguish the tokens corresponding to different plates, and the edge-level position code is used to distinguish the order of different tokens corresponding to the same plate.

4. The method according to any one of claims 1 to 3, characterized in that The obtaining of training data comprises: Acquire clothing design parameter files and / or three-dimensional clothing model images; Filtering the clothing design parameter file, and inputting the filtered clothing design parameter file into a preset large language model to obtain text training data output by the preset large language model; Rendering the three-dimensional clothing model image, and inputting the rendering result into a preset stable diffusion model to obtain a stylized sketch output by the preset stable diffusion model; The stylized sketch is binarized to obtain image training data; wherein the training data includes the text training data and / or the image training data.

5. A method for generating a clothing pattern image, characterized in that: The clothing pattern generation method comprises: Acquire user input data, and encode the user input data to obtain an input conditional encoding vector; wherein the type of the user input data includes text data and / or image data; Inputting the input conditional coding vector into a clothing pattern prediction model trained by the clothing pattern prediction model training method according to any one of claims 1 to 4, and obtaining a denoised token sequence output by the clothing pattern prediction model for the input conditional coding vector; A clothing pattern image corresponding to the user input data is generated according to the denoising token sequence.

6. The method according to claim 5, characterized in that The method further comprises: Performing data cleaning on each of the user input data and the corresponding clothing pattern images to obtain a feedback data set; According to the feedback data set, the clothing pattern prediction model is adjusted until the clothing pattern image generated by the clothing pattern prediction model after the adjustment meets the preset conditions.

7. The method according to claim 5, characterized in that The method further comprises: Generate an initial industrial sample image according to a preset seam addition rule and the generated garment pattern image; According to preset design parameter rules, the initial industrial sample drawing is updated to obtain a final industrial sample drawing.

8. An electronic device, characterized in that: The electronic device comprises: one or more processors; A memory having one or more computer programs stored thereon, wherein when the one or more computer programs are executed by the one or more processors, the one or more processors implement any of the following: The training method of the clothing pattern prediction model according to any one of claims 1 to 4; A method for generating a clothing pattern image according to any one of claims 5-7.

9. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, any of the following is achieved: The training method of the clothing pattern prediction model according to any one of claims 1 to 4; A method for generating a clothing pattern image according to any one of claims 5-7.

Citation Information

Patent Citations

  • Image style migration method and device, electronic equipment and storage medium

    CN118570050A

  • DiT architecture-based costume design increment generation method and system

    CN119578226A