Ru porcelain image generation method fusing Ru porcelain knowledge graph and fine tuning control

By constructing Ru porcelain knowledge graph and LoRA fine-tuning control methods, the texture, glaze color and instrument shape problems in Ru porcelain image generation are solved, and the high-quality generation of Ru porcelain images is achieved, which is in line with traditional aesthetics.

CN120451310APending Publication Date: 2025-08-08GUANGDONG UNIV OF TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510536528.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing generative model has problems such as open-sheet texture distortion, glaze rendering deviation and device type specification in the generation of Ru porcelain. Due to data scarcity, the existing optimization strategy is not ideal.

Method used

Construct the Ru porcelain knowledge graph, combine the large language model to generate prompt words, and build a lightweight Stable Diffusion model suitable for Ru porcelain image generation through LoRA fine-tuning and ControlNet structure control, and combine the multimodal large model Janus for image evaluation and optimization.

Benefits of technology

It realizes the reproduction of precise texture and glaze color features of Ru porcelain images, ensures that the instrument shape conforms to the Song Dynasty standards, improves the stability and consistency of the generated images, and conforms to the traditional aesthetic paradigm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451310A_ABST
    Figure CN120451310A_ABST
Patent Text Reader

Abstract

The invention provides a Ru porcelain image generation method fusing a Ru porcelain knowledge graph and fine tuning control, and the method comprises the steps: constructing the Ru porcelain knowledge graph, and generating prompt words based on the Ru porcelain knowledge graph through employing a large language model; constructing an improved Ru porcelain image generation model, and training the improved Ru porcelain image generation model by using the Ru porcelain image data set; and the trained improved Ru porcelain image generation model generates a Ru porcelain image according to the cue word, and a multi-mode large model Janus is adopted to evaluate the Ru porcelain image generated by the improved model. According to the method, a lightweight model suitable for Ru porcelain image generation is constructed in combination with LoRA and ControlNet; the LoRA module is used for adjusting parameters of a cross attention layer and finely reproducing texture and glazing color characteristics of Ru porcelain; the ControlNet module is used for controlling the structure of the device type and the posture to ensure that the image form is accurate; and introducing a multi-modal large model to carry out image evaluation on the generated image, and optimizing the cue word according to an evaluation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image generation technology, and in particular to a Ru porcelain image generation method that integrates a Ru porcelain knowledge graph and fine-tuning control. Background Art

[0002] In today's rapidly developing digital technology landscape, generative artificial intelligence (AIGC) has been widely applied in the field of artistic creation. However, existing technologies still have significant shortcomings in generating high-precision cultural artifacts such as Ru porcelain. The digitization of Ru porcelain faces a conflict between data scarcity and model generalization capabilities. As one of the five famous kilns of the Song Dynasty, Ru porcelain holds a pivotal position in the global history of ceramic art. However, existing Ru porcelain pieces are extremely scarce, and most are national first-class cultural relics. Therefore, high-precision 3D scanning and spectral data collection are difficult.

[0003] Furthermore, existing public datasets are mostly two-dimensional images, lacking key process parameters such as glaze thickness, microscopic crazing distribution, and body composition. Consequently, the scarcity of training samples leads to "feature collapse" in the generative model's learning of Ru porcelain features. The generated results are unable to faithfully reproduce the complex craftsmanship characteristics of Song Dynasty Ru porcelain, especially in terms of glaze rendering and crazing texture restoration. The generated model's results differ significantly from those of actual Ru porcelain.

[0004] To address these challenges, numerous optimization strategies have been proposed. However, while existing generative models, such as StableDiffusion 3.5 and DALL-E 3, have achieved excellent performance in general image synthesis, their underlying architectures are not specifically optimized for the specific craftsmanship of Ru porcelain, leading to three core issues: distorted crackle textures, biased glaze rendering, and inaccurate vessel shapes. Craze textures are simplified into random cracks in general models, failing to reproduce the layering and directional patterns of the 18 classic crackle patterns found in Ru porcelain. Existing models struggle to accurately represent the microcrystallization and reflectivity of the glaze. Furthermore, the disproportionate shape of Ru porcelain vessels prevents the generated shapes from conforming to the strict specifications of Song Dynasty Ru porcelain. For example, the string pattern spacing error on a string-patterned tripod jar exceeds 15%, and the number of petals on a sunflower-shaped bowl deviates from historical data. Finally, the generated results deviate from traditional aesthetic norms.

[0005] To improve generation quality, researchers have also tried various optimization methods, including full-scale fine-tuning, prompt word engineering, optimization, and transfer learning. However, due to the extreme scarcity of Ru porcelain data (typically fewer than 100 sets of samples), these methods have been less than ideal. For example, full-scale fine-tuning can easily lead to overfitting, resulting in highly repetitive and stereotyped images. Transfer learning, on the other hand, fails to fully account for the unique characteristics of Ru porcelain, such as body and glaze composition and firing temperature, leading to significant deviations between the generated results and the actual Ru porcelain. Summary of the Invention

[0006] In response to the shortcomings of the existing technology, the present invention provides a Ru porcelain image generation method that integrates Ru porcelain knowledge graph and fine-tuning control. The present invention ensures that the model can accurately understand and express the typical key visual features of Ru porcelain, such as glaze color, body, crazing, etc., and achieves image quality improvement through a closed-loop feedback mechanism of prompt word-image-evaluation-optimization.

[0007] The technical solution of the present invention is: a Ru porcelain image generation method integrating Ru porcelain knowledge graph and fine-tuning control, comprising the following steps:

[0008] S1), building a knowledge graph of Ru porcelain;

[0009] S2) Generate prompt words using a large language model based on the Ru porcelain knowledge graph;

[0010] S3) Construct an improved Ru porcelain image generation model Stable Diffusion and train it using the Ru porcelain image dataset;

[0011] S4) The trained improved Ru porcelain image generation model Stable Diffusion generates Ru porcelain images according to the prompt words, and the multimodal large model Janus is used to evaluate the Ru porcelain images generated by the Stable Diffusion model.

[0012] Preferably, in step S1), the construction of the Ru porcelain knowledge graph is as follows:

[0013] S11) Obtain the shape, glaze color, body quality, crackle texture, process flow and artistic characteristics of Ru porcelain;

[0014] S12) Constructing the Ru porcelain knowledge graph in the Neo4j graph database using the “node-relationship-attribute” model;

[0015] S13) After the construction is completed, the nodes and edges in the Ru porcelain knowledge graph are managed and called through the Cypher query language.

[0016] Preferably, in step S2), structured semantic information is injected into the large language model through the API interface and combined with the Ru porcelain knowledge graph to generate prompt words containing Ru porcelain craft style elements and description details, and then the prompt words are optimized through the prompt word conversion module; the prompt word conversion module receives and processes the natural language prompt information generated by the large language model, extracts the key semantic elements therein and assigns weight labels.

[0017] Preferably, in step S2), the prompt word conversion module includes a keyword weighter, a prompt word visualizer, a structure graph binder, and a feedback optimizer; the keyword weighter identifies keywords in natural language based on the Ru porcelain knowledge graph and assigns weights; the prompt word visualizer is used to visualize the weighted keywords and their weights; the structure graph binder is used to bind structural keywords and graphic input; the feedback optimizer is used for image quality evaluation feedback and fine-tuning the keyword weights in the prompt words.

[0018] Preferably, in step S3), a Ru porcelain image dataset is obtained and the Ru porcelain images are annotated through the Ru porcelain knowledge graph; each Ru porcelain image after preprocessing includes two parts: a structured keyword combination generated based on the Ru porcelain knowledge graph and a natural language supplementary description.

[0019] Preferably, in step S3), the improved Ru porcelain generation model Stable Diffusion is constructed as follows:

[0020] The Stable Diffusion model adopts the U-Net network as the framework, and the U-Net network includes multiple MM-DiT blocks; the cross-attention layer in the MM-DiT block of the Stable Diffusion model is fine-tuned by introducing the LoRA module; and the pre-trained conditional control network ControlNet is added to the front end of the trained LoRA module; by applying the conditional control information of ControlNet to the Stable Diffusion model, the shape, texture and posture of Ru porcelain are controlled.

[0021] Preferably, in step S3), during the training process, the input Ru porcelain data set of the Stable Diffusion model includes Ru porcelain text and Ru porcelain images; wherein, the Ru porcelain text is subjected to text feature extraction through a text encoder; and the denoising task is performed by fusing the Ru porcelain text and the Ru porcelain image and inputting them into the MM-DIT block of the Stable Diffusion model.

[0022] Preferably, in step S3), LoRA parameter embedding is introduced into the Query, Key and Value calculation links of the cross attention layer in the MM-DiT block to perform low-rank adjustment on the cross attention weights. The pre-trained weight matrix is set to LoRA uses low-rank matrix decomposition to represent the parameter update ΔW, namely:

[0023] W′=W0+ΔW

[0024] ΔW=BA

[0025] W0+ΔW=W0+BA

[0026] Where W′ is the updated weight after training; and is a trainable low-rank matrix with rank r<<min(d,k);

[0027] During the training process, the original pre-trained weight matrix W0 is fixed, and only the matrix and inserted into the LoRA module are optimized. Therefore, the forward propagation calculation process of the improved Stable Diffusion model is expressed as:

[0028] h=W0x+ΔWx=W0x+BAx

[0029] Where h represents the feature expression output by the Stable Diffusion model; x is the feature after the fusion of Ru porcelain text and image.

[0030] Preferably, in step S3), the calculation formula for each cross attention layer is:

[0031]

[0032] Where Q = W Q x is the query matrix; K = W K x represents the bond matrix; V = W V x represents the value matrix; W Q ,W K ,W V is the pre-trained weight matrix; x is the input feature; d k is the feature dimension; T represents the transposition operation;

[0033] The Ru porcelain text features extracted by the text encoder from the Ru porcelain text information are converted into K and V matrices. The latent Ru porcelain image features obtained after Ru porcelain image compression are then processed by U-Net to generate a Q matrix. The similarity between Q and K is calculated through the cross-attention mechanism, and V is then adjusted to make the features processed by U-Net more consistent with the text semantic information.

[0034] The cross attention layer interactively fuses the information of Ru porcelain image and Ru porcelain text in this way to guide the Ru porcelain image generation process, and the W Q 、W K 、W V Fine-tune and finally achieve effective fusion of style features.

[0035] As a preference, in step S3), the loss function of the Stable Diffusion model of the conditional control network ControlNet is introduced Expressed as:

[0036]

[0037] Where, Represents the random variable x,t,∈,c t ,c f The mathematical expectation of v = α t x0+β t ∈ is a linear combination of x and ε; α t represents the weight coefficient of the time step t data; x0 represents the original image without noise; β t represents the weight coefficient of the noise at time step t; v θ represents the output prediction value under parameter θ; x t represents the image with noise added at time step t; ε is Gaussian noise; c t Represents the text condition input at time step t; c f represents the control condition at time step t;

[0038] ControlNet's conditional control information is fully applied to the Stable Diffusion model, enabling precise control over details such as the shape, texture, and posture of Ru porcelain, significantly improving the controllability and visual quality of the generated images.

[0039] Preferably, in step S4), the Ru porcelain image generated by the Stable Diffusion model is input into the multimodal large model Janus for evaluation, and the evaluation result is input into the large language model to guide it to perform semantic optimization on the original prompt word and generate a new improved prompt word.

[0040] The beneficial effects of the present invention are:

[0041] 1. This invention uses a prompt word conversion module to structure key terms in natural language and match them with semantic nodes in the graph, achieving consistency between prompt word specialization and image semantic control. This effectively improves the model's ability to recognize and express the semantic features of Ru porcelain, and solves the problem of generalized and inaccurate prompt words in existing generation models.

[0042] 2. This paper combines LoRA fine-tuning with ControlNet structural control to build a lightweight Stable Diffusion model suitable for Ru porcelain image generation. The LoRA module is used to adjust the Cross-Attention parameters in the Stable Diffusion model to accurately reproduce the texture and glaze characteristics of Ru porcelain. The ControlNet module is used for structural control of the shape and posture of the porcelain to ensure accurate image morphology.

[0043] 3. The present invention also designs an image evaluation and closed-loop optimization mechanism, introduces a multimodal large model to perform quality analysis on the generated image, and feeds the evaluation results back to the prompt word optimization module to realize the cyclic update of prompting, generation, evaluation, and re-optimization, effectively improving the stability and consistency of model generation, and making the generated image more in line with the craft characteristics of Ru porcelain. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 Schematic diagram of the process of the present invention;

[0045] Figure 2 A schematic diagram of the Ru porcelain knowledge graph constructed by the present invention;

[0046] Figure 3 A schematic diagram of generating prompt words for the large language model of the present invention;

[0047] Figure 4 This is the structural framework diagram of the improved Ru porcelain image generation model Stable Diffusion of the present invention;

[0048] Figure 5 This is a structural framework diagram of the MM-DiT block of the present invention;

[0049] Figure 6 Schematic diagram of the process of generating Ru porcelain image of the present invention. DETAILED DESCRIPTION

[0050] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings:

[0051] like Figure 1 As shown, this embodiment provides a Ru porcelain image generation method that integrates Ru porcelain knowledge graph and fine-tuning control, including the following steps:

[0052] S1) Construct a knowledge graph of Ru porcelain; the details are as follows:

[0053] S11) Obtain the shape of Ru porcelain (such as string-patterned tripod jars and sunflower-shaped bowls), glaze color (such as sky blue and powder blue, including Lab color difference standard ΔE<3), body (such as incense ash body and agate body), crackle texture (such as crab claw pattern and cicada wing pattern), process (throwing → drying → glazing → firing) and artistic characteristics (such as aesthetic images such as "like jade but not jade" and "ice skin and jade bones");

[0054] S12) Construct the Ru porcelain knowledge graph in the Neo4j graph database using the “node-relationship-attribute” model; the constructed Ru porcelain knowledge graph is as follows Figure 2 As shown;

[0055] S13) Once constructed, the nodes and edges in the Ru porcelain knowledge graph are managed and accessed using the Cypher query language. This embodiment presents key process knowledge and semantic relationships in the Ru porcelain field in a structured, queryable, and scalable graph data format, providing semantic support for large-scale model prompt word optimization, image generation structure control, and model fine-tuning data construction.

[0056] S2) Generate prompt words using a large language model based on the Ru porcelain knowledge graph;

[0057] In this embodiment, the structured semantic information is injected into the large language model through the API interface and combined with the Ru porcelain knowledge graph to generate prompt words containing Ru porcelain craft style elements and description details, and then the prompt words are optimized through the prompt word conversion module; the prompt word conversion module receives and processes the natural language prompt information generated by the large language model, extracts the key semantic elements therein and assigns weight labels, such as Figure 3 shown.

[0058] This embodiment combines a large language model with a Ru porcelain knowledge graph to generate prompt words that are more characteristic of the Ru porcelain style, thereby improving the accuracy and consistency of style, texture, and structure during the image generation process.

[0059] In this implementation, users only need to input a concise natural language description into the large language model, which then outputs preliminary prompts containing elements of Ru porcelain craftsmanship and descriptive details. After processing by the prompt conversion module, the final prompt content is generated with controllable weights and accurate semantics.

[0060] In this embodiment, the prompt word conversion module includes a keyword weighter, a prompt word visualizer, a structure graph binder, and a feedback optimizer; the keyword weighter identifies keywords (shape, texture, body, glaze color, etc.) in natural language based on the Ru porcelain knowledge graph and assigns weights; the prompt word visualizer is used to visualize the weighted keywords and their weights; the structure graph binder is used to bind structural keywords and graphic input; the feedback optimizer is used for image quality evaluation feedback and fine-tuning the keyword weights in the prompt words.

[0061] S3) Construct an improved Ru porcelain image generation model Stable Diffusion and train it using the Ru porcelain image dataset;

[0062] In this embodiment, a Ru porcelain image dataset is obtained and the Ru porcelain images are annotated through the Ru porcelain knowledge graph; each Ru porcelain image after preprocessing includes two parts: a structured keyword combination generated based on the Ru porcelain knowledge graph and a natural language supplementary description.

[0063] like Figure 4 and 5As shown, the construction of the improved Ru porcelain generation model Stable Diffusion is:

[0064] The Stable Diffusion model adopts the U-Net network as the framework, and the U-Net network includes multiple MM-DiT blocks; the cross-attention layer in the MM-DiT block of the Stable Diffusion model is fine-tuned by introducing the LoRA module; and the pre-trained conditional control network ControlNet is added to the front end of the trained LoRA module; by applying the conditional control information of ControlNet to the Stable Diffusion model, the shape, texture and posture of Ru porcelain are controlled.

[0065] During the training process, the input Ru porcelain dataset of the Stable Diffusion model includes Ru porcelain text and Ru porcelain images; wherein, the Ru porcelain text is subjected to text feature extraction through a text encoder; and the Ru porcelain text and Ru porcelain image are fused and input into the MM-DIT block of the Stable Diffusion model to perform the denoising task.

[0066] By introducing LoRA parameter embedding in the Query, Key, and Value calculation links of the cross attention layer in the MM-DiT block, the cross attention weight is low-rank adjusted, and the pre-trained weight matrix is set to LoRA uses low-rank matrix decomposition to represent the parameter update ΔW, namely:

[0067] W′=W0+ΔW

[0068] ΔW=BA

[0069] W0+ΔW=W0+BA

[0070] Where W′ is the updated weight after training; and is a trainable low-rank matrix with rank r<<min(d,k);

[0071] During the training process, the original pre-trained weight matrix W0 is fixed, and only the matrix and inserted into the LoRA module are optimized. Therefore, the forward propagation calculation process of the improved Stable Diffusion model is expressed as:

[0072] h=W0x+ΔWx=W0x+BAx

[0073] Where h represents the feature expression output by the Stable Diffusion model; x is the feature after the fusion of Ru porcelain text and image.

[0074] The calculation formula for each cross attention layer is:

[0075]

[0076] Where Q = W Q x is the query matrix; K = W K x represents the bond matrix; V = W V x represents the value matrix; W Q ,W K ,W V is the pre-trained weight matrix; x is the input feature; d k is the feature dimension; T represents the transposition operation;

[0077] The Ru porcelain text features extracted by the text encoder from the Ru porcelain text information are converted into K and V matrices. The latent Ru porcelain image features obtained after Ru porcelain image compression are then processed by U-Net to generate a Q matrix. The similarity between Q and K is calculated through the cross-attention mechanism, and V is then adjusted to make the features processed by U-Net more consistent with the text semantic information.

[0078] The cross attention layer interactively fuses the information of Ru porcelain image and Ru porcelain text in this way to guide the Ru porcelain image generation process, and the W Q 、W K 、W V Fine-tune and finally achieve effective fusion of style features.

[0079] This example significantly improves the performance of Ru porcelain image generation using the LoRA module. This method achieves style fine-tuning of the Stable Diffusion model at a very low cost, offering high flexibility and efficiency. It is well-suited for resource-constrained applications, particularly for generating Ru porcelain images in specific styles.

[0080] In this embodiment, the control network ControlNet is implemented by copying the structure of the StableDiffusion model and freezing the weights of the original model, and only updating the weights of the copied model. In order to accurately control the generation of Ru porcelain images, it is only necessary to add the pre-trained ControlNet to the front end of the trained LoRA module to achieve precise conditional control and significantly improve the generation effect of the model. Expressed as:

[0081]

[0082] Where, Represents the random variable x,t,∈,c t ,c f The mathematical expectation of v = α t x0+β t ∈ is a linear combination of x and ε; α t represents the weight coefficient of the time step t data; x0 represents the original image without noise; β t represents the weight coefficient of the noise at time step t; v θ represents the output prediction value under parameter θ; x t represents the image with noise added at time step t; ε is Gaussian noise; c t Represents the text condition input at time step t; c f represents the control condition at time step t;

[0083] ControlNet's conditional control information is fully applied to the Stable Diffusion model, enabling precise control over details such as the shape, texture, and posture of Ru porcelain, significantly improving the controllability and visual quality of the generated images.

[0084] In addition, during the training process of the improved Stable Diffusion model of this embodiment, the input of the Stable Diffusion model includes the Ru porcelain image and its text description; this embodiment uses the Adam optimizer with adaptive learning rate, which uses the first-order moment estimation and second-order moment estimation of the gradient to dynamically adjust the learning rate of each parameter. After bias correction, the learning rate of each iteration is within a stable range, which can make the update of model parameters more stable. The loss function L MSE Using MSE Loss (mean square error), its expression is as follows:

[0085]

[0086] Where y i Represents the target value of the i-th sample, that is, the true value; represents the model prediction value of the i-th sample, and N represents the total number of samples.

[0087] In addition, this embodiment constructs trigger words during the training process. The trigger words avoid ambiguity in the model when keywords are input, and make it clear that the model only learns about Ru porcelain, which helps to ensure the generalization performance of the model.

[0088] S4) The trained improved Ru porcelain image generation model Stable Diffusion generates Ru porcelain images according to the prompt words, and the multimodal large model Janus is used to evaluate the Ru porcelain images generated by the Stable Diffusion model.

[0089] like Figure 6 As shown in the figure, the Ru porcelain image generated by the Stable Diffusion model is input into the multimodal large model Janus for evaluation, and the evaluation results are input into the large language model to guide it to semantically optimize the original prompt words and generate new improved prompt words.

[0090] The above embodiments and descriptions are only for explaining the principles and best embodiments of the present invention. Without departing from the spirit and scope of the present invention, the present invention may be subject to various changes and improvements, which shall fall within the scope of the invention to be protected.

Claims

1. A Ru porcelain image generation method integrating Ru porcelain knowledge graph and fine-tuning control, characterized in that: The steps include: S1), building a knowledge graph of Ru porcelain; S2) Generate prompt words using a large language model based on the Ru porcelain knowledge graph; S3) Construct an improved Ru porcelain image generation model Stable Diffusion and train it using the Ru porcelain image dataset; S4) The trained improved Ru porcelain image generation model Stable Diffusion generates Ru porcelain images according to the prompt words, and the multimodal large model Janus is used to evaluate the Ru porcelain images generated by the improved Stable Diffusion model.

2. The Ru porcelain image generation method integrating Ru porcelain knowledge graph and fine-tuning control according to claim 1 is characterized by: In step S1), the construction of the Ru porcelain knowledge graph is as follows: S11) Obtain the shape, glaze color, body quality, crackle texture, process flow and artistic characteristics of Ru porcelain; S12) Constructing the Ru porcelain knowledge graph in the Neo4j graph database using the "node-relationship-attribute" model; S13) After the construction is completed, the nodes and edges in the Ru porcelain knowledge graph are managed and called through the Cypher query language.

3. The Ru porcelain image generation method integrating Ru porcelain knowledge graph and fine-tuning control according to claim 1 is characterized by: In step S2), structured semantic information is injected into the large language model through the API interface and combined with the Ru porcelain knowledge graph to generate prompt words containing Ru porcelain craft style elements and description details, and then the prompt words are optimized through the prompt word conversion module; the prompt word conversion module receives and processes the natural language prompt information generated by the large language model, extracts the key semantic elements therein and assigns weight labels.

4. The Ru porcelain image generation method integrating Ru porcelain knowledge graph and fine-tuning control according to claim 3 is characterized by: In step S2), the prompt word conversion module includes a keyword weighter, a prompt word visualizer, a structure graph binder, and a feedback optimizer; the keyword weighter identifies keywords in natural language based on the Ru porcelain knowledge graph and assigns weights; the prompt word visualizer is used to visualize the weighted keywords and their weights; The structure diagram binder is used to bind structure keywords and graphic input; The feedback optimizer is used for image quality evaluation feedback and fine-tuning the weights of keywords in the prompt words.

5. The Ru porcelain image generation method integrating Ru porcelain knowledge graph and fine-tuning control according to claim 1 is characterized by: In step S3), the improved Ru porcelain generation model Stable Diffusion is constructed as follows: The Stable Diffusion model adopts the U-Net network as the framework, and the U-Net network includes multiple MM-DiT blocks; the cross-attention layer in the MM-DiT block of the Stable Diffusion model is fine-tuned by introducing the LoRA module; and the pre-trained conditional control network ControlNet is added to the front end of the trained LoRA module; by applying the conditional control information of ControlNet to the Stable Diffusion model, the shape, texture and posture of Ru porcelain are controlled.

6. The Ru porcelain image generation method integrating Ru porcelain knowledge graph and fine-tuning control according to claim 5 is characterized by: In step S3), the LoRA parameter embedding is introduced into the query, key and value calculation links of the cross attention layer in the MM-DiT block to adjust the cross attention weight to a low rank. The pre-trained weight matrix is set to LoRA uses low-rank matrix decomposition to represent the parameter update ΔW, namely: W′=W0+ΔW ΔW=BA W0+ΔW=W0+BA Where W′ is the updated weight after training; and is a trainable low-rank matrix with rank r<<min(d,k); During the training process, the original pre-trained weight matrix W0 is fixed, and only the matrix and inserted into the LoRA module are optimized. Therefore, the forward propagation calculation process of the improved Stable Diffusion model is expressed as: h=E0x+ΔWx=W0x+BAx Where h represents the feature expression output by the Stable Diffusion model; x is the feature after the fusion of Ru porcelain text and image.

7. The Ru porcelain image generation method integrating Ru porcelain knowledge graph and fine-tuning control according to claim 6 is characterized by: In step S3), the calculation formula for each cross attention layer is: Where Q = W Q x is the query matrix; K = W K x represents the bond matrix; V = W V x represents the value matrix; W Q ,W K ,W V is the pre-trained weight matrix; x is the input feature; d k is the feature dimension; T represents the transpose operation.

8. The Ru porcelain image generation method integrating Ru porcelain knowledge graph and fine-tuning control according to claim 7 is characterized by: In step S3), the Ru porcelain text features extracted from the Ru porcelain text information by the text encoder are converted into K and V matrices; the potential Ru porcelain image features obtained after the Ru porcelain image is compressed are then processed by U-Net; and a Q matrix is generated; the similarity between Q and K is calculated through a cross-attention mechanism, and V is adjusted so that the features processed by U-Net can be more consistent with the text semantic information; The cross attention layer interactively fuses the information of Ru porcelain image and Ru porcelain text in this way to guide the Ru porcelain image generation process, and the W Q 、W K 、W V Fine-tune and finally achieve effective fusion of style features.

9. The Ru porcelain image generation method integrating Ru porcelain knowledge graph and fine-tuning control according to claim 5 is characterized by: In step S3), the loss function of the Stable Diffusion model of the conditional control network ControlNet is introduced Expressed as: Where, Represents the random variable x,t,∈,c t ,c f The mathematical expectation of v = α t x0+β t ∈ is a linear combination of x and ε; α t represents the weight coefficient of the time step t data; x0 represents the original image without noise; β t represents the weight coefficient of the noise at time step t; v θ represents the output prediction value under parameter θ; x t represents the image with noise added at time step t; ε is Gaussian noise; c t Represents the text condition input at time step t; c f represents the control condition at time step t; ControlNet's conditional control information is fully applied to the Stable Diffusion model, enabling precise control over details such as the shape, texture, and posture of Ru porcelain, significantly improving the controllability and visual quality of the generated images.

10. The Ru porcelain image generation method integrating Ru porcelain knowledge graph and fine-tuning control according to claim 1 is characterized by: In step S4), the Ru porcelain image generated by the Stable Diffusion model is input into the multimodal large model Janus for evaluation, and the evaluation results are input into the large language model to guide it to perform semantic optimization on the original prompt word and generate a new improved prompt word.

Citation Information

Cited By

  • Multi-conditional constraint fusion generation method and system for planning intention graph generation

    CN120745021A

  • Lake and Hunan woodcarving image generation method, device and equipment based on LoRA model and storage medium

    CN121095382A

  • Lake xiang wood carving image generation method, device and equipment based on LoRA model and storage medium

    CN121095382B