Non-abandoned filament image generation method and system based on Stable Diffusion
By building an intangible cultural heritage filament image generation system based on Stable Diffusion, the problem of AI model insufficient understanding of filament process is solved, the generation quality and design efficiency are improved, and the modern transformation of traditional processes is achieved.
Patent Information
- Application Number
- CN202510505702.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-08-08
AI Technical Summary
It is difficult for existing AI models to accurately understand and express the cultural elements and process characteristics of the filament process, the production quality is poor, and it is difficult to provide design solutions that meet the actual process production requirements.
A non-legacy filament image generation system based on Stable Diffusion is built. By acquiring the visual visual features of the filament process, image and text data are collected, and data sets are constructed using keyword tag strategies. The generation model is fine-tuned through low-rank adaptive training strategies, and the text-to-image, image-to-image and local redrawing functions are integrated to realize user interaction.
It improves the AI model's understanding and generation quality of filament process, lowers the design threshold, promotes the integration of traditional crafts and modern technologies, and supports process inheritance and creative innovation.
Smart Images

Figure CN120451305A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of generative artificial intelligence, and more specifically, to a method and system for generating intangible cultural heritage filigree images based on StableDiffusion. Background Art
[0002] Filigree craftsmanship uses a variety of complex techniques, including pinching, filling, assembling, and welding, to shape precious metal wires such as gold and silver into exquisite patterns and three-dimensional forms, endowed with immense artistic value and cultural significance. However, the traditional filigree production process is extremely complex, requiring artisans to comprehensively consider multiple factors, including pattern selection, formal design, and decorative themes, and skillfully translate these design elements into detailed manual operations. This not only places high demands on practitioners' skills but also poses a major barrier to participation for new designers, artisans, and enthusiasts.
[0003] In recent years, artificial intelligence (AI) technology has made significant progress in the field of creative design, providing a new technological path for the modernization of traditional crafts. However, existing AIGC technology still has significant shortcomings when applied to traditional filigree craft design: on the one hand, existing AI models lack an in-depth understanding of filigree craft, a specific field of knowledge, and struggle to accurately capture and express the unique cultural elements and craft characteristics of filigree crafts. On the other hand, due to the extremely delicate structure and complex details of filigree crafts, current AI generation models produce poor quality when processing such high-precision visual content. Furthermore, existing AI systems struggle to provide feasible solutions that meet the actual production requirements of the craft. Summary of the Invention
[0004] The purpose of the present invention is to address the deficiencies of the existing technology and propose a method and system for generating intangible cultural heritage filigree images based on Stable Diffusion.
[0005] First, a method for generating intangible cultural heritage filigree images based on Stable Diffusion is provided, including:
[0006] S1. Obtain visual features of filigree craft;
[0007] S2. collecting image data and text data containing a variety of filigree patterns, and preprocessing the image data;
[0008] S3, annotating the image data and text data using a keyword-based labeling strategy to construct a paired filigree image-text dataset;
[0009] S4. Based on the generative diffusion model, the model is fine-tuned through a low-rank adaptive training strategy to obtain a dedicated generative model adapted to the characteristics of filigree craftsmanship;
[0010] S5. Build a user interaction interface that integrates text-to-image generation, image-to-image conversion, and local redrawing functions, and generates or adjusts filigree design images based on user input.
[0011] Preferably, in S4, the total number of training steps of the model is determined by the following formula:
[0012] Training Steps=Number of Images×Repeat×Epoch÷Batch Size
[0013] Among them, Training Steps is the number of training steps, Number of Images is the number of images, Repeat is the number of repetitions, Epoch is the training round, and Batch Size is the batch size.
[0014] Preferably, in S4, an adaptive learning rate optimizer is used during model training, and the model is automatically saved after each epoch.
[0015] Preferably, in S4, after the training is completed, the performance of the dedicated generation model is tested using the XYZ Plot script and the loss value, and the optimal model parameters are selected by comparing the generation results under different weight ranges.
[0016] Preferably, in S5, the text-to-image generation function is implemented by the following steps:
[0017] Enter keywords or select keywords from the prompt library; the keywords include product type, filigree pattern, decorative pattern and gem information;
[0018] Specify the desired image size and number;
[0019] Based on the keywords, image size and quantity, a design solution is generated.
[0020] Preferably, the image-to-image conversion function is implemented by the following steps:
[0021] Upload a reference image or sketch;
[0022] Based on the reference image or sketch, the generation result is controlled in combination with the ControlNet plug-in; and the similarity between the generated image and the uploaded image is controlled by adjusting the image weight value.
[0023] Preferably, in S5, the local redrawing function is implemented by the following steps:
[0024] Upload the image that needs to be modified;
[0025] Use the Smudge Tool to select and mark the area that needs to be modified;
[0026] Enter relevant keywords to describe the desired changes or adjustments;
[0027] Generates a modified image according to the input instructions.
[0028] In a second aspect, a system for generating an intangible cultural heritage filigree image based on Stable Diffusion is provided, for executing any of the methods described in the first aspect, including:
[0029] An acquisition module, used to obtain visual features of filigree craft;
[0030] An acquisition module, for acquiring image data and text data containing a variety of filigree patterns, and preprocessing the image data;
[0031] a labeling module, configured to label the image data and text data using a keyword-based labeling strategy to construct a paired filigree image-text dataset;
[0032] The fine-tuning module is used to fine-tune the generative diffusion model through a low-rank adaptive training strategy to obtain a specialized generative model adapted to the characteristics of filigree craftsmanship.
[0033] The construction module is used to build a user interaction interface, integrate text-to-image generation function, image-to-image conversion function and local redrawing function, and generate or adjust the filigree design image according to user input.
[0034] According to a third aspect, a computer storage medium is provided, wherein a computer program is stored in the computer storage medium; when the computer program is executed on a computer, the computer executes any one of the methods described in the first aspect.
[0035] In a fourth aspect, an electronic device is provided, including:
[0036] Memory, used to store computer programs;
[0037] A processor is used to execute the computer program to implement any method as described in the first aspect.
[0038] The beneficial effects of the present invention are:
[0039] 1. This paper deeply analyzes the characteristics of filigree craftsmanship, constructs a specialized dataset, and optimizes the labeling strategy, enabling the AI model to accurately understand and precisely express the fine structure and profound cultural connotations of filigree craftsmanship.
[0040] 2. This paper adopts the LoRA low-rank adaptive training method, which significantly improves the quality and accuracy of the model in generating filigree images, and effectively solves the generation quality problem of traditional AI models when processing high-precision visual content.
[0041] 3. This invention combines the ComfyUI system to build an AI-Filigree generation tool system, integrating three core functions: text-to-image, image-to-image, and partial redrawing. It not only lowers the design threshold but also improves design efficiency, allowing designers without professional skills to easily create high-quality filigree design images.
[0042] 4. This invention promotes the deep integration of traditional filigree craftsmanship and modern technology, and opens up an innovative path for the digital inheritance of traditional craftsmanship and its modern application. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 A flowchart of a method for generating an intangible cultural heritage filigree image based on Stable Diffusion provided by the present invention;
[0044] Figure 2 A schematic diagram of a traditional filigree visualization feature library provided by an embodiment of the present invention;
[0045] Figure 3 A schematic diagram of a system user interface provided by an embodiment of the present invention;
[0046] Figure 4 An example diagram of the text-to-image function provided by an embodiment of the present invention;
[0047] Figure 5 An example diagram of the image-to-image function provided by an embodiment of the present invention;
[0048] Figure 6 An example diagram of the local redraw function provided by an embodiment of the present invention;
[0049] Figure 7 A schematic diagram of the system operation workflow provided by an embodiment of the present invention;
[0050] Figure 8 Schematic diagram of a label comparison experiment provided by an embodiment of the present invention;
[0051] Figure 9 Schematic diagram of a model comparison experiment provided by an embodiment of the present invention;
[0052] Figure 10 A schematic diagram showing a comparison of the Loss function values provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0053] The present invention will be further described below with reference to the following examples. The following examples are provided only to facilitate understanding of the present invention. It should be noted that, without departing from the principles of the present invention, it is possible for a person skilled in the art to make various modifications to the present invention, and such improvements and modifications fall within the scope of the claims of the present invention.
[0054] Example 1:
[0055] As an example, this invention provides a method for generating images of intangible cultural heritage filigree based on Stable Diffusion. This method aims to accelerate the visualization process of filigree crafts through generative artificial intelligence (AI), improve the efficiency of filigree craft design, and lower the technical barriers to traditional craft design, while maintaining the core artistic characteristics of traditional filigree crafts. This invention not only effectively promotes the integration of traditional filigree crafts with modern digital technologies, but also supports craft inheritance and creative innovation.
[0056] Specifically, such as Figure 1 As shown, the method includes:
[0057] S1. Obtain visual features of filigree craftsmanship.
[0058] In S1, the present invention analyzes and organizes the visual features of filigree craft from various aspects such as filigree wire type, filigree pattern, modeling, and decorative pattern through comprehensive research and analysis of traditional Chinese filigree craft.
[0059] S2. Collecting image data and text data containing a variety of filigree patterns, and preprocessing the image data.
[0060] For example, through the systematic organization of the visual features of filigree craftsmanship, ten common filigree patterns were identified, including copper coin pattern, jujube flower pattern, spiral pattern, small six pattern, curling grass pattern, curling head pattern, straw mat pattern, arched silk pattern, bubble pattern and pine flower pattern; high-quality filigree image data and text data were collected from museum collections, auction websites and private collections.
[0061] Image data refers to images of filigree products using one or more patterns. Text data refers to the text descriptions corresponding to the filigree images. It is also necessary to ensure that the number of images for each of the ten filigree patterns is equal. All images must be pre-processed, including cropping, removing background interference, and resizing to a uniform 1024×1024 pixel resolution.
[0062] S3. Annotate the image data and text data using a keyword tagging strategy to construct a paired filigree image-text dataset.
[0063] Specifically, based on the content of the collected data, a keyword-based tagging strategy was employed, combining English and pinyin tags to describe traditional Chinese cultural elements. For example, the label for the "Xiaoliu filigree" pattern was "Xiaoliu filigree." Tags were sorted according to the desired model's focus, such as trigger words, filigree pattern, vessel shape, decorative pattern, gemstone, material, and other information. This ultimately formed a paired filigree image-text dataset for training the AI model.
[0064] S4. Based on the generative diffusion model, the model is fine-tuned through a low-rank adaptive training strategy to obtain a dedicated generative model adapted to the characteristics of filigree craftsmanship.
[0065] For example, a filigree image-text dataset built using S3 was used, with Stable Diffusion XL 1.0 as the basic training model. The LoRA low-rank adaptive training strategy was used to fine-tune the model. Adaptive learning rate optimizers such as DAdaptation were selected. A multi-round training strategy was adopted and the model parameters of each round were saved. Performance testing and evaluation were performed through XYZ Plot scripts and loss value comparisons. The optimal parameter configuration model was selected, and finally a dedicated LoRA model that can accurately understand and generate filigree craft features was obtained.
[0066] S5. Build a user interaction interface that integrates text-to-image generation, image-to-image conversion, and local redrawing functions, and generates or adjusts filigree design images based on user input.
[0067] In S5, an integrated design tool is developed based on the ComfyUI node-driven graphical user interface system, realizing three core functional modules: 1) Text-to-image function (T2I), where users can quickly generate a variety of filigree craft design solutions by inputting keywords (such as patterns, decorative patterns, product types, etc.) or selecting the prompt library recommended by the system; 2) Image-to-image conversion function (I2I), where users upload reference images or sketches, and combine ControlNet technology to precisely control the generated results, supporting style conversion or morphological adjustment of images; 3) Local redrawing function, where users can choose to upload existing design drawings, paint designated areas, and adjust local designs by inputting keywords to automatically generate the modified images.
[0068] Example 2:
[0069] Based on Example 1, Example 2 of the present invention provides a more specific method for generating an intangible cultural heritage filigree image based on Stable Diffusion, including:
[0070] S1. Obtain visual features of filigree craftsmanship.
[0071] S2. Collecting image data and text data containing a variety of filigree patterns, and preprocessing the image data.
[0072] S3. Annotate the image data and text data using a keyword tagging strategy to construct a paired filigree image-text dataset.
[0073] S4. Based on the generative diffusion model, the model is fine-tuned through a low-rank adaptive training strategy to obtain a dedicated generative model adapted to the characteristics of filigree craftsmanship.
[0074] In S4, the total number of training steps of the model is determined by the following formula:
[0075] Training Steps=Number of Images×Repeat×Epoch÷Batch Size
[0076] Among them, Training Steps refers to the number of training steps, which refers to the number of times the model updates parameters during the training process. Each time a batch of data training is completed, a parameter update is performed, which is counted as a training step. Number of Images refers to the number of images, which refers to the total number of images contained in the dataset, that is, the number of samples. Repeat refers to the number of repetitions, that is, the number of times the dataset is reused, usually referring to the situation where the dataset is fully traversed multiple times within an epoch. Epoch refers to the training round, which refers to the process in which the entire dataset is completely used for training once. One epoch means that each sample participates in one training. Batch Size refers to the batch size, which refers to the number of samples used in each training. Dividing the dataset into multiple small batches for training can improve computational efficiency.
[0077] In addition, an adaptive learning rate optimizer (such as DAdaptLion and DAdaptation) is used during model training, and the model is automatically saved after each epoch. This can effectively reduce the time and labor costs caused by repeatedly adjusting training parameters, and facilitate the subsequent comparison and selection of different versions of the model, thereby improving overall training efficiency and model quality.
[0078] Furthermore, after the training is completed, the performance of the dedicated generative model is tested using the XYZ Plot script and loss value, and the optimal model parameters are selected by comparing the generation results under different weight ranges.
[0079] S5. Build a user interaction interface that integrates text-to-image generation, image-to-image conversion, and local redrawing functions, and generates or adjusts filigree design images based on user input.
[0080] In S5, the text-to-image generation function is implemented by the following steps:
[0081] The user enters keywords, including product type, filigree pattern, decorative pattern and gemstone information, or selects from the system's recommended suggestion library; the user specifies the required image size and quantity, and the system generates a variety of design solutions based on the input keywords. For example, Figure 4 Shows a variety of filigree earring designs generated by inputting the prompt words "filigree,earrings,gold".
[0082] The image-to-image conversion function is implemented by the following steps:
[0083] Users upload a reference image or sketch and use the ControlNet plug-in to precisely control the generated results. This feature supports multiple control modes, such as edge detection and depth mapping. Users can adjust the image weight to control the similarity between the generated image and the uploaded image. Higher weights result in a more similar generated image to the uploaded image; lower weights result in a more distinct generated image from the original input image.
[0084] The local redraw function is implemented by the following steps:
[0085] Users upload the image they want to modify and use the smudge tool to select and mark the areas they want to edit. They then enter relevant keywords to describe the desired changes or adjustments, and the system generates the modified image based on the input instructions. Users can quickly preview the modified results to ensure they meet their design requirements.
[0086] In addition, the present invention also conducted label comparison experiments, such as Figure 8 As shown, Figure 8 (a1-a3) images generated using natural language format labels; Figure 8 (a4-a6) use keyword format labeling to label the generated images; Figure 8 Images (b1-b3) were generated using keyword-formatted labels with traditional filigree cultural elements expressed in English (a1, a4, b1), Pinyin (a2, a5, b2), and a combination of Pinyin and English (a3, a6, b3). The results show that the keyword-formatted group (a4-a6) significantly outperformed the natural language group (a1-a3) in terms of generation performance. Furthermore, using a combination of Pinyin and English (b3) to write traditional cultural labels was found to be most effective in reducing label confusion and achieving even better generation quality.
[0087] The present invention also carried out a model comparison experiment, such as Figure 9As shown, after training, the model can be compared using the Plot script. By evaluating the model's generated graphs (a), multiple sets of well-performing graphs can be generated and annotated (b). This allows the model to be judged for underfitting (failure to fully learn the data characteristics) or overfitting (performance on the training data but poor generalization). This ensures that the filigree design images generated by the system retain the characteristics of traditional craftsmanship while also being innovative and diverse, providing users with high-quality design references. Based on the training data set, the Filigree06-000007 model with a weight value between 0.8 and 1.0 is optimal.
[0088] The present invention also compares the loss function values, such as Figure 10 As shown in the figure, by comparing the loss function values at different training rounds, you can intuitively judge the model's convergence. A gradual decrease in the loss function value as training progresses indicates that the model is gradually optimizing and approaching the optimal solution. If the loss function value stabilizes and no longer decreases significantly, the model has reached convergence and training can be terminated at an appropriate time.
[0089] It should be noted that the parts in this embodiment that are the same or similar to those in Example 1 can be referenced to each other and will not be described in detail in this application.
[0090] Example 3:
[0091] Based on Example 2, Example 3 of the present application provides an intangible cultural heritage filigree image generation system based on Stable Diffusion, including:
[0092] An acquisition module, used to obtain visual features of filigree craft;
[0093] An acquisition module, for acquiring image data and text data containing a variety of filigree patterns, and preprocessing the image data;
[0094] a labeling module, configured to label the image data and text data using a keyword-based labeling strategy to construct a paired filigree image-text dataset;
[0095] The fine-tuning module is used to fine-tune the generative diffusion model through a low-rank adaptive training strategy to obtain a specialized generative model adapted to the characteristics of filigree craftsmanship.
[0096] The construction module is used to build a user interaction interface, integrate text-to-image generation function, image-to-image conversion function and local redrawing function, and generate or adjust the filigree design image according to user input.
[0097] It should be noted that the system provided in this embodiment is a system corresponding to the method provided in Example 2. Therefore, the parts in this embodiment that are the same or similar to those in Example 2 can be referenced to each other and will not be repeated in this application.
Claims
1. A method for generating intangible cultural heritage filigree images based on Stable Diffusion, characterized in that: include: S1. Obtain visual features of filigree craft; S2. collecting image data and text data containing a variety of filigree patterns, and preprocessing the image data; S3, annotating the image data and text data using a keyword-based labeling strategy to construct a paired filigree image-text dataset; S4. Based on the generative diffusion model, the model is fine-tuned through a low-rank adaptive training strategy to obtain a dedicated generative model adapted to the characteristics of filigree craftsmanship; S5. Build a user interaction interface that integrates text-to-image generation, image-to-image conversion, and local redrawing functions, and generates or adjusts filigree design images based on user input.
2. The method for generating intangible cultural heritage filigree images based on Stable Diffusion according to claim 1, characterized in that: In S4, the total number of training steps of the model is determined by the following formula: Training Steps=Number of Images×Repeat×Epoch÷Batch Size Among them, Training Steps is the number of training steps, Number of Images is the number of images, Repeat is the number of repetitions, Epoch is the training round, and Batch Size is the batch size.
3. The method for generating intangible cultural heritage filigree images based on Stable Diffusion according to claim 2, characterized in that: In S4, an adaptive learning rate optimizer is used during model training, and the model is automatically saved after each epoch.
4. The method for generating intangible cultural heritage filigree images based on Stable Diffusion according to claim 3, characterized in that: In S4, after the training is completed, the performance of the dedicated generative model is tested using the XYZ Plot script and the loss value, and the optimal model parameters are selected by comparing the generation results under different weight ranges.
5. The method for generating intangible cultural heritage filigree images based on Stable Diffusion according to claim 4, characterized in that: In S5, the text-to-image generation function is implemented by the following steps: Enter keywords or select keywords from the prompt library; the keywords include product type, filigree pattern, decorative pattern and gem information; Specify the desired image size and number; Based on the keywords, image size and quantity, a design solution is generated.
6. The method for generating intangible cultural heritage filigree images based on Stable Diffusion according to claim 5, characterized in that: The image-to-image conversion function is implemented by the following steps: Upload a reference image or sketch; Based on the reference image or sketch, the generation result is controlled in combination with the ControlNet plug-in; and the similarity between the generated image and the uploaded image is controlled by adjusting the image weight value.
7. The method for generating intangible cultural heritage filigree images based on Stable Diffusion according to claim 6, characterized in that: In S5, the local redrawing function is implemented by the following steps: Upload the image that needs to be modified; Use the Smudge Tool to select and mark the area that needs to be modified; Enter relevant keywords to describe the desired changes or adjustments; Generates a modified image according to the input instructions.
8. A non-legacy filigree image generation system based on Stable Diffusion, characterized by: Used to perform the method according to any one of claims 1 to 7, comprising: An acquisition module, used to obtain visual features of filigree craft; An acquisition module, for acquiring image data and text data containing a variety of filigree patterns, and preprocessing the image data; a labeling module, configured to label the image data and text data using a keyword-based labeling strategy to construct a paired filigree image-text dataset; The fine-tuning module is used to fine-tune the generative diffusion model through a low-rank adaptive training strategy to obtain a specialized generative model adapted to the characteristics of filigree craftsmanship. The construction module is used to build a user interaction interface, integrate text-to-image generation function, image-to-image conversion function and local redrawing function, and generate or adjust the filigree design image according to user input.
9. A computer storage medium, characterized in that The computer storage medium stores a computer program; when the computer program is run on a computer, the computer executes the method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the method according to any one of claims 1 to 7.
Citation Information
Cited By
Lake and Hunan woodcarving image generation method, device and equipment based on LoRA model and storage medium
CN121095382A