An interactive multi-parameter fabric texture generation method and system

CN121353287BActive Publication Date: 2026-09-22DONGHUA UNIV
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202511917423.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-09-22
Estimated Expiration
2045-12-18

AI Technical Summary

Benefits of technology

本发明提出了一种可交互的多参数织物纹理生成方法与系统,通过将稳定扩散模型与低秩自适应(LoRA)微调技术结合,在潜在空间中实现了多维度物理参数的精确映射,克服了传统纹理合成方法生成结果机械化、GAN类方法生成多样性不足且难以编辑的缺陷,让生成模型可解释,易控制。进一步地,本发明通过引入DSL Agent模块,在保证生成质量的同时显著降低了操作门槛,能够将非专业用户输入的模糊自然语言描述自动转化为可被模型识别的标准化面料设计指令,从而扩展了系统在大众化设计、教育及产业应用中的适用范围。通过构建包含丰富组织结构、密度、细度与色彩信息的专业数据集,并采用数据增强与多层次质量评估体系,本发明在图像保真度、色彩还原性与结构准确性上均显著优于现有方法,尤其能够精确呈现平纹、斜纹等组织结构以及飞数等专业参数,为虚拟服装设计提供了一种基于参数驱动的低门槛、低成本且可编辑、高质量的织物纹理生成方案,能够广泛应用于虚拟服装设计、数字时尚、家纺产品开发及智能制造等领域。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121353287B_ABST
    Figure CN121353287B_ABST
Patent Text Reader

Abstract

The present application relates to an interactive multi-parameter fabric texture generation method and system, the method comprising: collecting high-definition images of real fabric samples and measuring their physical parameters, and constructing an image-parameter paired dataset; fine-tuning a pre-trained stable diffusion model using low-rank adaptive technology to enable it to learn the precise mapping relationship from physical parameter text description to fabric image; receiving user input of target parameter text description, and generating the corresponding ultra-high-definition fabric image using the fine-tuned model; and finally, performing multi-dimensional quality assessment on the generated image. The present application solves the problems of existing fabric image generation techniques, such as dependence on physical samples, lack of microstructure parameter control capability, single editing method, and inability to reversibly adjust local texture. The present application achieves independent and fine control of fabric micro-physical parameters through natural language interaction, and the generated fabric image is significantly superior to existing methods in terms of visual authenticity and structural accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an interdisciplinary field of generative artificial intelligence and textile engineering, and particularly to an interactive multi-parameter fabric texture generation method and system. Background Technology

[0002] In virtual clothing design, digital fashion, and metaverse scenarios, the demand for highly realistic fabric texture images is increasing. Existing related technologies mainly suffer from the following problems: 1. Images generated by traditional texture synthesis methods lack variation and struggle to represent realistic yarn-level structures and complex material details; 2. While physical scanning-based methods can acquire high-precision fabric images, they heavily rely on physical samples, resulting in high costs and a lack of editing flexibility; 3. While generative adversarial networks (GANs) have made progress in generating macroscopic patterns for clothing, they fall short in terms of fidelity, diversity, and parameter controllability of microscopic fabric textures.

[0003] In recent years, diffusion models have gradually become the mainstream for multimodal generation due to their stability in high-fidelity image generation. However, in the field of fabric texture generation, existing methods are mostly focused on macroscopic patterns or style transfer, and have not yet achieved independent and joint control of multidimensional physical parameters such as warp and weft density, yarn fineness, and fly count, resulting in a lack of realism and editability in the microstructure of the generated results.

[0004] Furthermore, in existing research related to fabric image generation, some methods attempt to synthesize fabric defect or texture images using generative models. For example, Chinese invention patent application number 20211156706.0, entitled "A Method for Generating Defective Fabrics Based on a DPGAN Model," aims to expand detection samples or simulate actual defect scenarios. However, the generated results often exhibit repetitive textures or patterned artifacts, failing to reflect the detailed differences in warp and weft interlacing and yarn thickness variations of real fabrics. In some dyeing and finishing or printing simulation scenarios, such as Chinese invention patent application number 202210265143.X, entitled "A Rapid Generation System and Method for Patterns of Yunjin Brocade Satin Fabric," some systems predict the diffusion trajectory of dye on the fabric surface based on diffusion principles or physical models, or generate local color overlay effects. However, these models are more inclined towards process simulation, and the generation scope is usually limited to the color level, lacking the ability to model the overall texture structure. Furthermore, in the field of virtual try-on or clothing image generation, technical solutions combining human images and text descriptions for clothing replacement have emerged. In the area of ​​decorative fabric or tapestry image simulation, some solutions generate images resembling embroidery or weaving textures through style transfer or sample splicing, such as the Chinese invention patent application number 202510794137.7, entitled "An Intelligent Method and System for Generating Silk Thread Texture Images for Decorative Tapestries." However, these often rely on fixed pattern templates and lack the ability to freely adjust the texture structure based on parameters. It is evident that although existing related technologies have made some progress in fabric image generation, dyeing and finishing simulation, and texture style transfer, a unified fabric image generation mechanism that flexibly drives multi-dimensional physical parameters such as fabric structure, density, yarn fineness, and material color has not yet been formed. Summary of the Invention

[0005] To address the shortcomings of existing fabric image generation technologies, such as reliance on physical samples, lack of microstructural parameter control, limited editing methods, and inability to reversibly adjust local textures, this invention proposes an interactive, multi-parameter fabric texture generation method and system. This method can simultaneously respond to combined inputs of multiple physical attributes, including warp and weft density, yarn fineness, weave structure, material, and color, achieving synchronous modeling of fabric images at both the macroscopic style and microscopic texture levels. In one embodiment, this invention further introduces a fine-grained color control module, a structural consistency discrimination module, and a physical parameter feedback module to enhance the adaptability of the generated image in identifying subtle color differences, restoring warp and weft directions, and associating mechanical properties.

[0006] The technical solution of this invention is as follows: In view of the above problems, this invention proposes a fabric image generation scheme that balances image fidelity and parameter controllability. This scheme uses a stable diffusion model as the core of generation. By constructing a training dataset containing real fabric microscopic images and systematic physical parameter labels, and introducing a low-rank adaptive fine-tuning mechanism, the model can accurately understand and respond to natural language descriptions composed of multiple dimensions such as weave structure, warp and weft density, yarn fineness, material, and color. This achieves fine and controllable generation of fabric textures and reduces the uncertainty of neural network generation. Unlike synthesis methods that rely solely on texture matching or style transfer, this invention does not limit the generation process to overall image style transfer or fixed template filling. Instead, it progressively models the microstructure based on a conditional diffusion process within the latent space, ensuring that the generated image is consistent with the real fabric in both macroscopic visual effects and yarn-level details. Simultaneously, this invention establishes an objective quality evaluation system to quantitatively verify the generation results from aspects such as pixel fidelity, structural similarity, color accuracy, and weave structure recognition, ensuring that the generated images have reliable consistency and usability.

[0007] A multi-parameter driven method for generating fabric texture images includes the following steps: S1. Acquire high-resolution images of real fabric samples and measure multiple physical parameters corresponding to each sample. The physical parameters should include five categories of physical parameters: fabric structure, warp and weft density, yarn fineness, material and color. Perform data augmentation on the acquired high-resolution images to form a fabric image-physical parameter pair dataset for model training. The construction of the fabric image-physical parameter dataset in step S1 specifically includes: S1.1. Acquire original high-definition digital images of various fabric samples based on high-precision scanning equipment, and simultaneously record the physical parameters of each sample. The physical parameters are acquired by manual measurement or instrument measurement and stored in a structured manner. S1.2. Perform one or more operations on the acquired raw high-resolution digital images, including rotation, flipping, brightness adjustment, color transformation, and noise addition, to expand the dataset size and improve the model's generalization ability; the data augmentation operations can be formally represented as:

[0008] in, Represents rotation and flip operators. These are the brightness scaling and offset coefficients, respectively. It is Gaussian noise; S1.3. Write a corresponding normalized label description for each image. This normalized label description systematically includes the corresponding physical parameter information.

[0009] The innovation of this invention lies in the explicit coupling of the fabric's microstructural features with the visual generation process through a multi-parameter driven mechanism. Traditional research often relies on text prompts or image priors for fabric appearance simulation, while this invention is based on a multi-dimensional mapping model of fabric physical parameters and image generation. Parameters such as warp and weft density, yarn fineness, and fly count are input into the diffusion model in a structured manner, thereby achieving a reversible and precisely controllable fabric texture generation mechanism.

[0010] S2. A fabric image generation model is constructed based on the pre-trained stable diffusion model (refer to the paper Rombach et al., “High-Resolution Image Synthesis with Latent Diffusion Models”, CVPR 2022, and publicly available at https: / / github.com / CompVis / stable-diffusion). The stable diffusion model has a dual network structure of Base and Refiner, and the fabric image generation model is fine-tuned using low-rank adaptive (LoRA) technology. The image-physical parameter pairs (i.e., image-text pairs) in the dataset constructed in step S1 are used as training samples, where the text is a normalized description including the physical parameters. Through training, the fine-tuned model learns the accurate mapping relationship from the text description of the fabric physical parameters to the corresponding high-definition fabric image. The fine-tuning of the fabric image generation model employs a low-rank adaptive technique, specifically including: In the U-Net architecture of the pre-trained stable diffusion model, a pair of trainable low-rank matrices with rank r = 4-8 are introduced at the linear projection weight matrix of each cross-attention layer. , , making the original weights It is replaced during forward computation. ,in Let be the rank of the low-rank matrix. This is the scaling factor; the original weights are frozen during fine-tuning. Only for , The parameters are updated through backpropagation; this allows the model to adapt to fabric texture generation tasks while maintaining the original image generation capability. The training process typically involves 150-300 rounds, with weights dynamically adjusted based on the performance on the validation set.

[0011] The U-Net architecture includes symmetrical downsampling and upsampling paths. The downsampling path consists of four downsampling stages, each consisting of two 3×3 convolutional layers, a downsampling operation with a stride of 2, and a progressively increasing number of channels from 64, 128, 256 to 512. The upsampling path consists of four upsampling stages, each consisting of a 2×2 deconvolution operation and two 3×3 convolutional layers, with the number of channels progressively decreasing to a final 64. The intermediate layer is a dual-channel cross-attention block that connects the downsampling and upsampling paths, enabling cross-scale feature fusion.

[0012] During the training phase, the text containing fabric physical parameters in the image-text pair is converted into a semantic vector by a text encoder (CLIP text encoder), and randomly initialized latent noise is input together with the semantic vector into the U-Net architecture. At each time step, the U-Net architecture performs conditional denoising on the noise tensor through a multi-scale feature fusion mechanism, outputting a denoised latent representation. This latent representation is then decoded into a pixel-level image by an image decoder. The denoising process can be represented as follows:

[0013] in, It is a time step. For noise samples, For noise predictor, For the diffused noise figure, It is the noise adjustment parameter (variance). For conditional vectors, Using standard Gaussian variables, U-Net gradually obtains denoised latent variables in the latent space; The loss function is calculated based on the pixel-level image (i.e., the generated image obtained from decoding) and the corresponding target image. It includes two parts: pixel-level reconstruction loss and perceptual loss. This loss function is used... Only for the low-rank matrix , Perform gradient descent updates; iterate the above training phase process until the convergence condition is met.

[0014] After training is completed, the components are fused proportionally during the inference phase. The original weights are dynamically controlled to fine-tune the effect, and the low-rank adaptive technique injection can be applied to the Base subnetwork and Refiner subnetwork of the fabric image generation model, or to the cross-attention layer at multiple scales, to achieve hierarchical refinement and multi-scale control.

[0015] S3. Input a new text description containing the physical parameters of the target fabric into the fabric image generation model trained in step S2; the fabric image generation model encodes the text in the latent space, fuses the semantic vector of the text encoding with random noise, and performs multi-step iterative denoising through the U-Net architecture to generate denoised latent variables; finally, the latent variables are decoded into a pixel-level ultra-high-definition fabric image with a resolution of 1024×1024 pixels through an image decoder (VAE decoder).

[0016] S4. The quality of the fabric image generated in step S3 is evaluated using objective evaluation indicators, which include at least the following: Peak signal-to-noise ratio (PSNR) is used to evaluate the pixel-level fidelity between a generated image and a real image; its calculation formula is:

[0017] in, It is the maximum possible value of a pixel. It is the mean squared error, which is calculated based on the height of the image. m and width n , Is it a real image in position? pixel values, Is the generated image at location Pixel values; The Structural Similarity Index (SSIM) is used to simulate human visual perception to assess the structural similarity between generated images and real images.

[0018] in, These represent image patches at corresponding positions in the real image and the generated image, respectively. and The average brightness of the real image and the generated image are respectively. and The corresponding luminance variance, Let the covariance of the two be , To prevent constants with zero denominators, SSIM takes values ​​from 0 to 1, with values ​​closer to 1 indicating higher structural similarity.

[0019] Color similarity evaluation based on K-means clustering measures the accuracy of color reproduction by calculating the Euclidean distance between the generated image and the dominant color set of the real image; the calculation formula is as follows:

[0020] in, and Let represent the color vectors of the k-th primary color in the real image and the generated image, respectively. It is the number of clusters; For the first real image k A primary color, For generating the first image k The smaller the distance, the more consistent the dominant color of the generated image is with the real image.

[0021] Fabric structure accuracy is assessed by automated algorithms or manual judgment to evaluate the correctness of the fabric structure (such as plain weave, twill weave, satin weave) and key parameters (such as fly count) in the generated images.

[0022] A multi-parameter-driven fabric texture image generation system, used to execute the multi-parameter-driven fabric texture image generation method as described above, includes: The dataset building module is used to perform the acquisition of fabric sample images, the determination of physical parameters, data augmentation, and text description generation to build the training dataset. The model fine-tuning training module is used to load the pre-trained SD stable diffusion model and fine-tune the model using LoRA technology based on the training dataset. The image generation module is used to receive the text description of the physical parameters of the fabric input by the user, and generate the corresponding fabric image using the trained model. Through parameter-driven approach, the black-box fabric generation is transformed into an intuitive and controllable generation mode. The structural consistency discrimination module is used to determine and constrain the warp and weft features and weave structure type of the fabric during the image generation process.

[0023] The physical parameter feedback module is used to infer the mechanical properties of the fabric based on the generated fabric image and form a closed-loop control.

[0024] The quality assessment module is used to perform multi-dimensional quality assessments on the generated fabric images; The DSL intelligent agent module is used to parse the natural language fabric description input by the user into a standardized and structured professional design language, and to perform parameter rationality and consistency checks during the parsing process to ensure the quality of the generated product and lower the barrier to entry for users.

[0025] The system details are as follows: The dataset construction module is used to acquire original images of fabric samples using high-precision scanning equipment, and simultaneously measure and label physical parameters such as the weave structure, warp and weft density, and yarn fineness of each sample. It expands the sample size through various image enhancement techniques and generates standardized label descriptions containing various physical attributes for each image. The model fine-tuning training module, based on the pre-trained SD stable diffusion model, introduces LoRA technology and uses the training dataset to optimize the parameters of the cross-attention layer in the model, so that it can quickly adapt to the specific task requirements of fabric image generation without adjusting most of the weights of the original model. The image generation module receives the target fabric parameter description in natural language input, and performs conditional denoising operation in the latent space through the trained generation model to output an ultra-high-definition fabric image that meets the parameter requirements. The parameter-driven generation method transforms the originally complex and uncertain fabric generation task into an appearance-controllable generation task, better matching the user's needs. The quality assessment module quantifies the generated images from multiple dimensions, including pixel fidelity, structural similarity, color reproduction, and fabric structure accuracy. The assessment metrics include peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), color distance based on K-means clustering, and fabric structure judgment accuracy, to comprehensively verify the quality of the generated results.

[0026] The system also includes a fine-grained color control module, which introduces numerical color parameters in addition to natural language descriptions to improve the model's ability to identify subtle color differences. This fine-grained color control module can use color spaces such as RGB, Lab, or HSV, setting the primary or secondary color of the target fabric as a continuous color vector, and concatenating or weightedly fusing it with the semantic vector output by the text encoder in the latent space. During model training, the system further constrains the primary color set of the generated image and the real sample based on a color distance loss function to reduce the deviation between similar colors such as dark blue and blue-black, or off-white and warm white, thereby improving the stability and consistency of color reproduction.

[0027] The structural consistency discrimination module includes a directional feature extraction unit for identifying the warp and weft directions, and constructs a warp-weft differentiation model based on self-supervised contrastive learning to determine whether the arrangement direction of the warp and weft in the generated image conforms to the input parameters. For plain weave, twill weave, satin weave and other weave structures, the generated image can be screened a second time by pattern template matching or convolution kernel response features. When the recognition result does not match the target parameters, a structural constraint signal is applied to the generation model to improve the restoration accuracy at the weave structure level.

[0028] The physical parameter feedback module constructs an estimation model from the image to the elastic coefficient or friction coefficient, and performs physical rationality verification on the generated texture results in the absence of actual fabric samples. When the estimation result deviates from the preset parameters by more than a threshold, mechanical constraints are applied to the latent variables of the generated model, so that the generated image not only conforms to the input description in visual appearance, but also remains consistent with the real fabric in terms of macroscopic deformation trend and surface roughness. This closed-loop feedback mechanism provides a foundation for the in-depth expansion of application scenarios such as virtual try-on, digital weaving, and physical simulation rendering.

[0029] Furthermore, to improve the interaction efficiency between ordinary users and the professional design system, this invention introduces a DSL Agent (Domain Specific Language) module into the system. The DSL Agent module, acting as the system's language intermediary layer, automatically converts user-input natural language descriptions or coarse semantic commands into professional fabric descriptions (fine descriptions) recognizable by fashion designers or fabric engineers. Based on a knowledge base and large language model technology in the apparel field, the DSL Agent internally constructs a dedicated semantic graph and terminology mapping rules for the fabric design domain. It can automatically complete missing parameters, standardize terminology expressions, and correct non-professional expressions. For example, when an ordinary user inputs "I want a light blue, glossy summer fabric," the DSL Agent will convert it into a structured, executable fabric design language: "Weft-oriented silk twill, fly count 2 up 2 down, warp diameter 0.18mm, weft diameter 0.20mm, warp and weft densities 180 threads / inch and 160 threads / inch respectively, main color Pantone 290C."

[0030] Furthermore, the DSL Agent not only enhances the system's usability but also introduces a quality constraint mechanism during semantic parsing. Through contextual consistency checks and parameter rationality verification, it ensures that the generative model receives complete and professional input conditions, thereby significantly reducing generation errors caused by fuzzy input from non-professional users. This quality constraint mechanism, while guaranteeing generation quality, lowers the interaction threshold, enabling an automatic transition from fuzzy natural language descriptions to precise professional design parameters. It provides scalable natural language interface support for apparel companies, textile design schools, and digital twin manufacturing scenarios, giving this invention universal applicability and intelligent features for mass design.

[0031] Furthermore, the introduction of the DSL Agent enables the generation system of this invention to achieve a balance between interactivity and generation stability: Firstly, the standardization of the semantic layer ensures the consistency of input parameters such as fabric structure, density, and fly count; Secondly, the interpretable parameter mapping mechanism enhances the model's ability to analyze and respond to complex design instructions; Third, by reducing the need for manual parameter adjustments, the design efficiency and the reproducibility of the generated results are effectively improved.

[0032] Therefore, the DSL Agent module is not only a language conversion tool, but also an intelligent control core that connects user intent with professional generation models, significantly improving the system's usability and intelligent interaction level while ensuring the quality of generated images.

[0033] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the multi-parameter driven fabric texture image generation method as described above.

[0034] A computer-readable storage medium having stored thereon computer program instructions, which, when executed by a processor, implement the multi-parameter driven fabric texture image generation method as described above.

[0035] In the field of fabric image generation, existing technologies either rely on physically based modeling methods to obtain high-fidelity texture images, a cumbersome process that is highly dependent on real-world samples; or they utilize deep learning models such as Generative Adversarial Networks (GANs) to achieve controllable generation to some extent. However, these methods generally suffer from problems such as blurred details in the generated images and imprecise control over microscopic parameters such as warp and weft density and flyback number, resulting in fabric images that lack realism and fail to meet the high level of detail required for virtual clothing design. Especially when multiple parameters such as fabric structure, yarn fineness, and color need to be independently controlled by adjusting text descriptions, existing methods often struggle to achieve precise editing and offer poor interactivity.

[0036] For example, if the input text description does not explicitly specify the fly count of the twill fabric, the generated model may produce textures with structural errors, failing to accurately distinguish between "two-up-one-down" twill and "two-up-two-down" twill. Similarly, when the description of yarn fineness is too general, the model may fail to reflect the difference in thickness between warp and weft yarns, resulting in the generated fabric image lacking the necessary mechanical appearance characteristics. Furthermore, inaccurate color control may cause the intended "indigo denim" to appear with an unnatural purple hue. These deficiencies in pixel fidelity, structural similarity, and color accuracy ultimately lead to 3D clothing models generated based on such textures lacking the feel of realistic fabrics, affecting the overall expressiveness of digital clothing.

[0037] Furthermore, while general diffusion models perform well in image generation, their pre-trained models are not designed for specialized fabric image generation tasks. Directly applying such models to generate fabric textures often fails to understand the specialized, structured physical parameter descriptions, leading to risks of texture distortion, missing details, or inconsistencies with the description. Attempts to adapt the base model to fabric image generation tasks through fine-tuning face significant computational resource requirements, lengthy training times, and a high risk of overfitting.

[0038] To address these challenges, this invention provides a method for generating ultra-fine, editable fabric images based on a stable diffusion model. For example... Figure 1 As shown, the method includes: acquiring high-resolution images of real fabric samples and measuring their physical parameters to construct an image-parameter pairing dataset; fine-tuning a pre-trained SD stable diffusion model using low-rank adaptive (LoRA) technology to enable it to learn the precise mapping relationship from textual descriptions of physical parameters to fabric images; receiving user-inputted textual descriptions of target parameters and generating corresponding ultra-high-resolution fabric images using the fine-tuned model; and finally, performing multi-dimensional quality evaluation on the generated images. In this invention, by constructing a high-quality professional dataset and injecting domain knowledge into the model, and by employing LoRA technology for efficient and low-cost model fine-tuning, the model can accurately parse and respond to complex physical parameter descriptions; finally, a rigorous objective evaluation system ensures that the generated images meet requirements in terms of structure, color, and density. This invention achieves independent and precise control of fabric microscopic physical parameters through natural language interaction, and the generated fabric images significantly outperform existing methods in terms of visual realism and structural accuracy.

[0039] Furthermore, the method of this invention places both the model optimization and image generation process in the forward inference stage of the model. Users only need to modify the input text description to obtain new generation results in real time, without the need to retrain the model or perform complex post-processing. This greatly reduces the threshold and time cost of digital material creation, and provides an efficient and high-quality solution for virtual clothing design.

[0040] The beneficial effects of this invention are as follows: This invention proposes an interactive, multi-parameter fabric texture generation method and system. By combining a stable diffusion model with low-rank adaptive (LoRA) fine-tuning technology, it achieves precise mapping of multi-dimensional physical parameters in the latent space. This overcomes the shortcomings of traditional texture synthesis methods, such as mechanized generation results, and the lack of diversity and difficulty in editing of GAN-based methods, making the generated model interpretable and easy to control. Furthermore, by introducing a DSL Agent module, this invention significantly reduces the operational threshold while ensuring generation quality. It can automatically convert fuzzy natural language descriptions input by non-professional users into standardized fabric design instructions that can be recognized by the model, thereby expanding the system's applicability in popular design, education, and industrial applications. By constructing a professional dataset containing rich information on organizational structure, density, fineness, and color, and employing data augmentation and a multi-level quality assessment system, this invention significantly outperforms existing methods in image fidelity, color reproduction, and structural accuracy. In particular, it can accurately present organizational structures such as plain weave and twill weave, as well as professional parameters such as fly count. This provides a parameter-driven, low-threshold, low-cost, editable, and high-quality fabric texture generation solution for virtual clothing design, which can be widely applied in fields such as virtual clothing design, digital fashion, home textile product development, and intelligent manufacturing. Attached Figure Description

[0041] Figure 1 This is a schematic diagram illustrating the process of establishing an interactive multi-parameter fabric texture generation method and system according to an embodiment of the present invention. Figure 2 This invention provides a flowchart for fabric image prediction, illustrating the specific process by which user-input natural language or parameter descriptions are parsed by the DSL Agent module and then used to generate a model to complete fabric image inference and prediction. Figure 3 The flowchart provided in this embodiment of the invention illustrates the training process of a fabric image generation model based on a multi-parameter dataset. Detailed Implementation

[0042] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.

[0043] This invention provides a method for generating fabric images, such as... Figure 1 As shown, it includes the following steps: In step S101, in a laboratory environment, high-resolution digital images (resolution no less than 512×512) of a total of 300 real fabric samples are acquired using high-precision scanning equipment (such as a fiber fineness meter). The acquired samples cover three basic weave structures: plain weave, twill weave, and satin weave, including 152 plain weave samples, 85 twill weave samples, and 64 satin weave samples, to ensure that the model can learn diverse fabric structural features.

[0044] In step S102, for each fabric sample, multiple physical parameters are simultaneously measured and recorded, and these parameters are stored in a structured form, including fly count, number of yarns in complete fabric, yarn material, color, and mass per unit area (g / m²). 2 )wait.

[0045] To broaden data coverage and improve the model's generalization ability, data augmentation processing is performed on the fabric samples. This process includes operations such as rotation, flipping, brightness perturbation, color perturbation, and Gaussian noise injection, which can be formally represented as:

[0046] in, Represents rotation and flip operators. These are the brightness scaling and offset coefficients, respectively. The noise is Gaussian. Using this method, the original 300 samples are expanded to approximately 4000-4500 samples, of which plain weave accounts for approximately 48%, twill weave approximately 30%, and satin weave approximately 22%, ensuring that the dataset covers rich structural and color features.

[0047] In step S103, the present invention uses SD as the base pre-trained model. This model has a dual network structure of Base and Refiner, and can stably generate high-detail images at a resolution of 1024×1024.

[0048] Furthermore, to reduce training overhead and improve model adaptability, a low-rank adaptive (LoRA) technique is introduced. In the cross-attention layer of U-Net, low-rank matrices A and B are introduced to modify the original weight matrix W, resulting in:

[0049] in, Let be the rank of the low-rank matrix. This is the scaling factor. During the training phase, the original weights are frozen, and only parameters A and B are updated, thereby significantly reducing computational overhead and the risk of overfitting.

[0050] The model was trained on a regular high-performance GPU using the Prodigy optimization algorithm, with an initial learning rate of 1. The training dataset was divided into training, validation, and test sets in a 7:1.5:1.5 ratio. The training process iterated approximately 100,000 times and took about 12 hours.

[0051] In step S104, a natural language description containing the physical parameters of the target fabric is input. This description is converted into a semantic vector z by the CLIP text encoder and then compared with a randomly initialized noise tensor in the latent space. After fusion, the input is given to U-Net, and conditional iterative denoising is performed in the latent space. The process can be represented as follows:

[0052] in, For noise samples, For noise predictor, For the diffused noise figure, For conditional vectors, The variables are standard Gaussian. A latent representation is obtained through multiple iterations, and this latent representation is reconstructed into a pixel-level high-resolution fabric image using a VAE decoder. The entire generation process takes only about 3–5 seconds to obtain a fabric texture image with a resolution of 1024×1024.

[0053] Furthermore, such as Figure 2 As shown, the method of the present invention supports interactive editing and includes the following steps: In step S201, the fabric text description information input by the user is received. The description information can be a non-professional, vague instruction. In step S202, the fabric description information input by the user is input into the DSL Agent module, which parses, standardizes and completes the natural language description, and generates a fabric parameter description that conforms to the model input specification. In step S203, the consistency of the parameter description processed by the DSL Agent is checked. The input structured parameters include fabric structure, warp and weft density, yarn fineness, material and color. In step S204, after the parameter verification is passed, the standardized fabric parameter description is input into the fabric image generation model after LoRA fine-tuning. In step S205, the fabric image generation model performs conditional diffusion denoising operation in the latent space to generate the corresponding fabric latent representation; In step S206, the latent representation is decoded into a pixel-level fabric texture image by a VAE image decoder, and the generated result is output. The arrows in the diagram indicate the direction of information flow. The system can update the fabric image in real time based on the modified input parameters, thereby achieving interactive generation and locally controllable editing.

[0054] In steps S105 and S207, in order to objectively evaluate the quality of the generated image, this invention establishes a comprehensive evaluation system that includes multiple indicators, including: Peak signal-to-noise ratio (PSNR) is used to evaluate the pixel-level fidelity between the generated image and the real reference image. Its calculation formula is as follows:

[0055] A higher PSNR value indicates a smaller pixel-level error. Testing showed that the PSNR value of the image generated by this invention is consistently around 12dB, while the PSNR of the comparison model DressCode is approximately 10.7dB. The Structural Similarity Index (SSIM) simulates the human visual system to evaluate the structural similarity between a generated image and a real image. Its calculation formula is as follows:

[0056] The SSIM index of the images generated by this invention is significantly higher than that of the baseline model, indicating that it is superior in terms of visual perception quality.

[0057] Color similarity assessment uses the K-means clustering algorithm to extract the dominant color sets of both the generated and real images, and calculates the average of the minimum Euclidean distances between all color pairs in the two dominant color sets. The calculation formula is as follows:

[0058] The smaller this value, the more accurate the color reproduction. The value of this indicator in the method of this invention is much smaller than that of the comparison model, reflecting higher color reproduction accuracy; Fabric structure accuracy: The accuracy of the fabric structure (such as plain weave, twill weave, satin weave) and key parameters (such as fly count) in the generated image is evaluated by automatic algorithms or professional judgment.

[0059] Experimental results show that after data augmentation and LoRA fine-tuning, the accuracy of fabric structure recognition by the model of this invention is significantly improved. Specifically, in terms of Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM), the method of this invention improves by an average of approximately 2.45 dB and 0.05 compared to the unenhanced model, respectively; and in terms of color Euclidean distance, it reduces by more than 30%, indicating that the generated image is closer to the real sample in terms of color reproduction. Meanwhile, in terms of fabric structure recognition, the accuracy rates for plain weave, twill weave, and satin weave reach 100%, 98%, and 94%, respectively, a significant improvement compared to the unenhanced model (67%, 29%, and 65%, respectively); the fabric density accuracy also increases from approximately 89% in the original model to over 95%. These experimental data fully demonstrate that the model of this invention, after data augmentation and LoRA fine-tuning, can effectively improve the structural consistency and physical controllability of the generated fabric images, and has significant scientific application value.

[0060] The present invention provides a method for generating fabric images, the specific architecture and workflow of which are as follows: Figure 3 As shown, in the fabric sample data construction stage, high-resolution fabric images are obtained by collecting real fabric samples. Simultaneously, the physical parameters of the fabric samples are measured and recorded, including fabric structure, warp and weft density, yarn fineness, material, and color. Based on this, a corpus of fabric parameters with both non-professional fuzzy descriptions and professional precise descriptions is constructed to achieve a mapping between natural language descriptions and professional physical parameters. Image preprocessing and data augmentation operations are performed on the collected fabric images, and combined with the generated standardized text descriptions, a fabric image-text pair training dataset is constructed for model training.

[0061] During the model training and generative inference phases, a DSL Agent is constructed using a large language model based on the training dataset. This Agent parses and standardizes the natural language descriptions input by the user, generating fabric parameter descriptions that conform to the model input format. Subsequently, these parameter descriptions are input into a generative model based on the SDXL architecture. While keeping the original model parameters frozen, a low-rank adaptive module (LoRA) is introduced into the cross-attention layer of the U-Net network for targeted fine-tuning and parameter optimization. During the generation process, the semantic vectors obtained by the text encoder, randomly initialized latent noise, and conditional features adapted by LoRA are input into the U-Net network for multi-step denoising operations to generate corresponding latent representations. After decoding by the VAE decoder, the latent representations are output as high-resolution fabric texture images.

[0062] In the quality assessment and feedback optimization stage, the generated fabric texture image is subjected to multi-dimensional quality assessment, including peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), color similarity assessment (K-means clustering algorithm), and fabric structure accuracy. The assessment results are used as feedback information to guide further optimization of model parameters or assist in parameter adjustment in subsequent generation processes, thereby improving the realism and parameter consistency of the generated fabric image.

[0063] This method successfully achieves precise and independent control of the microscopic physical parameters of fabrics through natural language commands, generating high-definition fabric images with rich details, accurate structure, and realistic colors, effectively solving the problems of poor editability and insufficient realism in existing related technologies.

[0064] The above-described embodiments are merely one implementation of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of protection of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the appended claims.

Claims

1. An interactive multi-parameter fabric texture generation method, characterized in that, Includes the following steps: S1. Acquire high-resolution images of real fabric samples and measure multiple physical parameters corresponding to each sample; Data augmentation operations are performed on the acquired high-definition images to form a fabric image-physical parameter pair dataset for model training; S2. Based on the pre-trained stable diffusion model, a fabric image generation model is constructed. The stable diffusion model has a dual network structure of Base and Refiner, and the low-rank adaptive LoRA technology is used to fine-tune the fabric image generation model. The image-physical parameter pairs, i.e., image-text pairs, in the dataset constructed in step S1 are used as training samples, where the text is a normalized description including the physical parameters. Through training, the fine-tuned model learns the accurate mapping relationship from the text description of the fabric physical parameters to the corresponding high-definition fabric image. S3. Input a new text description containing the physical parameters of the target fabric into the fabric image generation model trained in step S2; the fabric image generation model encodes the text in the latent space, fuses the semantic vector of the text encoding with random noise, and performs multi-step iterative denoising through the U-Net architecture to generate denoised latent variables; finally, the latent variables are decoded into a higher resolution fabric image through an image decoder. S4. Use objective evaluation indicators to assess the quality of the fabric image generated in step S3; The construction of the fabric image-physical parameter dataset in step S1 specifically includes: S1.

1. Acquire original high-definition digital images of various fabric samples based on high-precision scanning equipment, and simultaneously record the physical parameters of each sample. The physical parameters should include five categories of physical parameters: fabric structure, warp and weft density, yarn fineness, material and color. The physical parameters are acquired by manual measurement or instrument measurement and stored in a structured manner. S1.

2. Perform one or more operations on the acquired raw high-resolution digital images, including rotation, flipping, brightness adjustment, color transformation, and noise addition, to expand the dataset size and improve the model's generalization ability; the data augmentation operations are formally represented as: in, Represents rotation and flip operators. These are the brightness scaling and offset coefficients, respectively. It is Gaussian noise; S1.

3. Write a corresponding normalized label description for each image. This normalized label description systematically includes the corresponding physical parameter information. The objective evaluation metrics mentioned in step S4 include: peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM). Color similarity evaluation based on K-means clustering measures the accuracy of color reproduction by calculating the Euclidean distance between the generated image and the dominant color set of the real image; the calculation formula is as follows: in, and Representing the first and second parts of the real image and the generated image respectively k The color vector of the primary color It is the number of clusters; For the first real image k A primary color, The k-th dominant color of the generated image; the smaller this distance, the more consistent the dominant color of the generated image is with the real image. Fabric structure accuracy is assessed by using automated algorithms or manual judgment to evaluate the correctness of the fabric structure and key parameters in the generated images. Step S2, which involves fine-tuning the fabric image generation model using a low-rank adaptive technique, specifically includes: In the U-Net architecture of the pre-trained stable diffusion model, a pair of trainable low-rank matrices with rank parameters r ranging from 4 to 8 are introduced at the linear projection weight matrix of each cross-attention layer. , , making the original weights It is replaced during forward computation. in Let be the rank of the low-rank matrix. This is the scaling factor; the original weights are frozen during fine-tuning. Only for , The parameters are updated via backpropagation; this enables the model to adapt to fabric texture generation tasks while maintaining the original image generation capability. The training process is carried out in multiple rounds, and the weights are dynamically adjusted based on the performance on the validation set. The U-Net architecture includes symmetrical downsampling and upsampling paths. The downsampling path consists of four downsampling stages, each consisting of two 3×3 convolutional layers, a downsampling operation with a stride of 2, and a progressively increasing number of channels from 64, 128, 256 to 512. The upsampling path consists of four upsampling stages, each consisting of a 2×2 deconvolution operation and two 3×3 convolutional layers, with the number of channels progressively decreasing to a final 64. The intermediate layer is a dual-channel cross-attention block that connects the downsampling and upsampling paths, enabling cross-scale feature fusion. During the training phase, the text containing fabric physical parameters in the image-text pair is converted into a semantic vector by a text encoder, and randomly initialized latent noise is input into the U-Net architecture along with the semantic vector. At each time step, the U-Net architecture performs conditional denoising on the noise tensor through a multi-scale feature fusion mechanism, outputting a denoised latent representation. This latent representation is then decoded into a pixel-level image by an image decoder. The denoising process is represented as follows: in, It is a time step. For noise samples, For noise predictor, For the diffused noise figure, It is the noise adjustment parameter, i.e., variance. For conditional vectors, Using standard Gaussian variables, U-Net gradually obtains denoised latent variables in the latent space; The loss function is calculated based on the generated image obtained from pixel-level image decoding and the corresponding target image. It includes two parts: pixel-level reconstruction loss and perceptual loss. This loss function is used... Only for the low-rank matrix , Perform gradient descent updates; iterate the above training phase process until the convergence condition is met. After training is completed, the components are fused proportionally during the inference phase. The original weights are dynamically controlled to fine-tune the effect, and the low-rank adaptive technique is injected into the Base subnetwork and Refiner subnetwork of the fabric image generation model, or into the cross-attention layer at multiple scales, to achieve hierarchical refinement and multi-scale control. The objective evaluation indicators mentioned in step S4 include: Peak Signal-to-Noise Ratio (PSNR) is used to evaluate the pixel-level fidelity between the generated image and the real image; its calculation formula is: in, It is the maximum value of the pixel. This is the mean squared error, where m and n are the height and width of the image, respectively. It is the pixel value of the actual image at position (i,j). It is the pixel value of the generated image at position (i,j); The Structural Similarity Index (SSIM) is used to simulate human visual perception to assess the structural similarity between generated and real images. in, These represent image patches at corresponding positions in the real image and the generated image, respectively. and The average brightness of the real image and the generated image are respectively. and The corresponding luminance variance, Let the covariance of the two be , To prevent constants with zero denominators, SSIM takes values ​​from 0 to 1, with values ​​closer to 1 indicating higher structural similarity. The interactive multi-parameter fabric texture generation method is executed by the following system: The dataset construction module is used to acquire original images of fabric samples using high-precision scanning equipment, and simultaneously measure and label the physical parameters of each sample, including fabric structure, warp and weft density, yarn fineness, material and color. It expands the sample size through various image enhancement techniques and generates standardized label descriptions containing various physical attributes for each image. The model fine-tuning training module, based on the pre-trained stable diffusion model, introduces LoRA technology and uses the training dataset to optimize the parameters of the cross-attention layer in the model, so that it can quickly adapt to the specific task requirements of fabric image generation without adjusting most of the weights of the original model. The image generation module receives the target fabric parameter description in natural language and performs conditional denoising in the latent space through the trained generation model to output a higher resolution fabric image that meets the parameter requirements. The parameter-driven generation method transforms the originally complex and uncertain fabric generation task into an appearance-controllable generation task, better matching the user's needs. The quality assessment module quantifies the generated image from multiple dimensions, including pixel fidelity, structural similarity, color reproduction, and tissue structure accuracy. The assessment metrics include peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), color distance based on K-means clustering, and fabric structure judgment accuracy, to comprehensively verify the quality of the generated results. The color fine control module uses RGB, Lab, or HSV color spaces to set the main or secondary color of the target fabric as a continuous color vector, and splices or weightedly fuses it with the semantic vector output by the text encoder in the latent space. During model training, the main color set of the generated image and the real sample is further constrained based on the color distance loss function to reduce the deviation between similar colors such as dark blue and blue-black, and beige and warm white, thereby improving the stability and consistency of color reproduction. The structural consistency discrimination module includes a directional feature extraction unit for identifying the warp and weft directions, and constructs a warp-weft differentiation model based on self-supervised contrastive learning to determine whether the arrangement direction of the warp and weft in the generated image conforms to the input parameters. For weave structures including plain weave, twill weave, and satin weave, the generated image is screened a second time by pattern template matching or convolutional kernel response features. When the recognition result does not match the target parameters, a structural constraint signal is applied to the generation model to improve the restoration accuracy at the weave structure level. The physical parameter feedback module constructs an estimation model from the image to the elastic coefficient or friction coefficient, and performs physical rationality verification on the generated texture results without actual fabric samples; when the estimation result deviates from the preset parameters by more than a threshold, mechanical constraints are applied to the latent variables of the generated model. The intelligent agent module, serving as the system's language intermediary layer, automatically converts user-input natural language descriptions or fuzzy semantic commands into professional fabric descriptions recognizable by fashion designers or fabric engineers. Based on a knowledge base in the apparel field and large language model technology, the intelligent agent module internally constructs a dedicated semantic graph and terminology mapping rules for the fabric design domain. It can automatically complete missing parameters, standardize terminology expressions, and correct non-professional expressions. A quality constraint mechanism is introduced during semantic parsing, ensuring that the generation model receives complete and professional input conditions through contextual consistency judgment and parameter rationality verification, thereby significantly reducing generation errors caused by fuzzy input from non-professional users.

2. An electronic device, characterized in that, An interactive multi-parameter fabric texture generation method as described in claim 1 is comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the interactive multi-parameter fabric texture generation method as described.

3. A computer-readable storage medium, characterized in that, The device is used to store the interactive multi-parameter fabric texture generation method as described in claim 1, wherein computer program instructions are stored thereon, and when the computer program instructions are executed by a processor, the interactive multi-parameter fabric texture generation method as described is implemented.

Citation Information

Patent Citations

  • A system and method for rapid generation of patterns for brocade satin fabric

    CN114622322B

  • A silk thread texture image intelligent generation method and system for decorative tapestry

    CN120655805B

  • Gray fabric defect generation method and system, medium and computer

    CN117541564A

  • Wallpaper display method and device based on vehicle-mounted information entertainment system

    CN118860546A

  • Fabric defect image denoising method based on improved generative adversarial network

    CN119048391A