Diffusion model Thangka line draft graph coloring method and device

Through the diffusion model Thangka line drawing coloring method, the problems of Thangka line drawing coloring accuracy and training stability in the existing technology were solved, and high-quality Thangka line drawing coloring was achieved, maintaining artistic style and cultural characteristics.

CN120198529APending Publication Date: 2025-06-24QINGHAI NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510357509.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

It is difficult to accurately realize the coloring of Thangka line drawings in the prior art, especially when dealing with Thangka images with dense lines, rich semantic information and gradient color, there are problems of coloring errors and training difficulties.

Method used

The diffusion model Thangka line drawing coloring method is adopted. By obtaining Thangka images and extracting line drawings, the Thangka data set is constructed, and the Thangka line drawing coloring framework is trained, including initializing the coloring network and quality enhancement diffusion model, network training and denoising processing are performed to complete the coloring of the Thangka line drawing to be colored.

Benefits of technology

It realizes the coloring of Thangka line drafts accurately without designing special color gradients and semantic segmentation, simplifies the coloring process, while maintaining the artistic style and cultural characteristics of Thangka.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198529A_ABST
    Figure CN120198529A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to a diffusion model thangka draft image coloring method and equipment, and aims to realize primary coloring on a thangka draft by using an initial coloring network so as to ensure semantic alignment. And subsequently, the preliminary coloring image is input into diffusion for quality enhancement, and the Thangka color is further refined and smoothed, so that the coloring of the Thangka line draft can be realized under the condition that special color gradient and semantic segmentation are not involved, the coloring process is greatly simplified, and meanwhile, the artistic style and cultural characteristics of the Thangka are kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a method and device for coloring a line drawing of a Thangka using a diffusion model. Background Art

[0002] With the development of color transfer technology, coloring based on line drawings has attracted wide attention. Especially in combination with deep learning frameworks, great progress has been made in related fields. Currently, most of the existing line drawing coloring is for coloring anime faces, animated characters, etc. Although great progress has been made in the existing technology, there are still some difficulties and challenges: (1) Compared with other line drawing images, the lines of Thangka are dense and rich in semantic information, making it difficult to achieve accurate coloring and prone to incorrect color correspondence. (2) Thangka usually uses natural mineral pigments, which are rich and bright in color and have gradient colors, making it challenging to color Thangka line drawings. (3) Since it takes dozens of days or years for a professional Thangka painter to draw a Thangka, Thangka is very precious and there is no existing dataset for training and use.

[0003] Due to its dense lines, rich semantics, bright colors and color gradients, Thangka poses unique challenges to deep learning models. How to ensure the correct semantic correspondence of Thangka lines is a problem that deep learning needs to solve. At the same time, due to the unique colors of Thangka, including gradient colors, the model needs to be able to correctly color different gradient Thangka. To address this challenge, researchers have proposed various methods to enhance line drawing coloring. Mainly using semantic segmentation, designing color feature extractors, etc., in the hope of helping the deep neural network learn semantic correspondence and coloring. However, the existing technology still cannot achieve good coloring for Thangka images with large semantic differences. At the same time, the existing technology uses many conditional constraints and is difficult and unstable to train. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a method and device for coloring a Thangka line drawing using a diffusion model to overcome the problems existing in the current prior art.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions:

[0006] On the one hand, the present application provides a method for coloring a Thangka line drawing using a diffusion model, including:

[0007] Obtain a Thangka image, extract a line drawing from the Thangka image, and generate a line drawing image;

[0008] Construct a Thangka dataset according to the Thangka image and the line drawing image;

[0009] Construct a coloring framework for the Thangka line drawing;

[0010] Train the Thangka line drawing coloring framework through the Thangka dataset;

[0011] Obtain the Thangka line drawing to be colored;

[0012] Perform initial coloring on the Thangka line drawing to be colored through the trained Thangka line drawing coloring framework;

[0013] Enhance the quality of the Thangka line drawing to be colored through the Thangka line drawing coloring framework, and output the image and the Thangka line drawing to be colored Figure 1 Predict noise with the input noise prediction network;

[0014] Denoise according to the predicted noise through the denoising module to complete the coloring of the Thangka line drawing to be colored.

[0015] Further, for the method described above, the obtaining of the Thangka image, extracting the line drawing from the Thangka image, and generating a line drawing include:

[0016] Obtain the Thangka image collected by on-site shooting;

[0017] Crop the Thangka image into a square image of 256×256 pixels;

[0018] Extract the line drawing from the Thangka image to generate the corresponding line drawing.

[0019] Further, for the method described above, the Thangka line drawing coloring framework includes: an initial coloring network and a quality enhancement diffusion model.

[0020] Further, for the method described above, the training of the Thangka line drawing coloring framework through the Thangka dataset includes:

[0021] Divide the Thangka dataset into a training set and a test set;

[0022] Train the initial coloring network through the training set and test the initial coloring network through the test set.

[0023] Further, for the method described above, the performing of initial coloring on the Thangka line drawing to be colored through the trained Thangka line drawing coloring framework includes:

[0024] Perform initial coloring on the Thangka line drawing to be colored through the initial coloring network.

[0025] Further, for the method described above, the quality enhancement diffusion model is a DDIM model;

[0026] The quality enhancement of the Thangka line drawing to be colored through the Thangka line drawing coloring framework includes:

[0027] Performing quality enhancement on the Thangka line drawing to be colored through the DDIM model.

[0028] Furthermore, the above-mentioned method further includes:

[0029] Introducing color consistency loss, perceptual loss, style loss, and generation loss for the initialized coloring network.

[0030] Furthermore, for the above-mentioned method, the formula definition of the DDIM model is:

[0031]

[0032] Wherein, represents the noisy image at time t'+1, represents the real image, represents the noisy image at time t′, t' represents the sampling time series, and ε θ represents the noise prediction network, represents the cumulative signal retention ratio up to time step t'+1.

[0033] On the other hand, the present application provides a diffusion model Thangka line drawing coloring device, including a processor and a memory, and the processor is connected to the memory:

[0034] Wherein, the processor is used to call and execute the program stored in the memory;

[0035] The memory is used to store the program, and the program is at least used to execute the diffusion model Thangka line drawing coloring method described in any one of the above.

[0036] The beneficial effects of the present invention are:

[0037] The present application first obtains a Thangka image, extracts a line drawing from the Thangka image to generate a line drawing, constructs a Thangka dataset based on the Thangka image and the line drawing, constructs a Thangka line drawing coloring framework, trains the network of the Thangka line drawing coloring framework through the Thangka dataset, obtains the Thangka line drawing to be colored, performs initial coloring on the Thangka line drawing to be colored through the trained Thangka line drawing coloring framework, performs quality enhancement on the Thangka line drawing to be colored through the Thangka line drawing coloring framework, and combines the output image and the Thangka line drawing to be colored Figure 1The predicted noise is output by the input noise prediction network, and denoising is performed through the denoising module according to the predicted noise to complete the coloring of the line drawing of the thangka to be colored. In this application, an initial coloring network is used to perform a preliminary coloring on the thangka line drawing to ensure semantic alignment. Subsequently, the preliminary coloring diagram is input into the diffusion for quality enhancement to further refine and smooth the colors of the thangka. In this way, it is possible to achieve the coloring of the thangka line drawing without involving the design of specialized color gradients and semantic segmentation, which greatly simplifies the coloring process while maintaining the artistic style and cultural characteristics of the thangka. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0039] Figure 1 It is a flowchart provided by an embodiment of a method for coloring a thangka line drawing using a diffusion model of the present invention;

[0040] Figure 2 It is a schematic diagram of the coloring framework structure of a thangka line drawing provided by an embodiment of a method for coloring a thangka line drawing using a diffusion model of the present invention;

[0041] Figure 3 It is a visualization example diagram of a thangka provided by an embodiment of a method for coloring a thangka line drawing using a diffusion model of the present invention;

[0042] Figure 4 It is a step-by-step coloring diagram of a thangka provided by an embodiment of a method for coloring a thangka line drawing using a diffusion model of the present invention;

[0043] Figure 5 It is a schematic diagram of the structure provided by an embodiment of a device for coloring a thangka line drawing using a diffusion model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be described in detail below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0045] Figure 1 It is a flowchart provided by an embodiment of a method for coloring a thangka line drawing using a diffusion model of the present invention. Please refer to Figure 1, this embodiment may include the following steps:

[0046] S1. Obtain a Thangka image, extract the line drawing from the Thangka image, and generate a line drawing image;

[0047] S2. Construct a Thangka dataset based on the Thangka image and the line drawing image;

[0048] S3. Construct a coloring framework for the Thangka line drawing;

[0049] S4. Perform network training on the coloring framework for the Thangka line drawing through the Thangka dataset;

[0050] S5. Obtain the Thangka line drawing to be colored;

[0051] S6. Perform initial coloring on the Thangka line drawing to be colored through the trained coloring framework for the Thangka line drawing;

[0052] S7. Enhance the quality of the Thangka line drawing to be colored through the coloring framework for the Thangka line drawing, and use the output image and the Thangka line drawing to be colored Figure 1 to predict the prediction noise with the input noise prediction network;

[0053] S8. Denoise according to the prediction noise through the denoising module to complete the coloring of the Thangka line drawing to be colored.

[0054] It should be noted that the coloring framework for the Thangka line drawing is constructed using the PyTorch framework.

[0055] It can be understood that this application first obtains a Thangka image, extracts the line drawing from the Thangka image to generate a line drawing image, constructs a Thangka dataset based on the Thangka image and the line drawing image, constructs a coloring framework for the Thangka line drawing, performs network training on the coloring framework for the Thangka line drawing through the Thangka dataset, obtains the Thangka line drawing to be colored, performs initial coloring on the Thangka line drawing to be colored through the trained coloring framework for the Thangka line drawing, enhances the quality of the Thangka line drawing to be colored through the coloring framework for the Thangka line drawing, and uses the output image and the Thangka line drawing to be colored Figure 1 to predict the prediction noise with the input noise prediction network, and denoise according to the prediction noise through the denoising module to complete the coloring of the Thangka line drawing to be colored. In this application, using an initial coloring network to perform a preliminary coloring on the Thangka line drawing ensures semantic alignment. Subsequently, the preliminary coloring image is input into the diffusion for quality enhancement to further refine and smooth the Thangka colors. In this way, it is possible to achieve the coloring of the Thangka line drawing without involving the design of specialized color gradients and semantic segmentation, which greatly simplifies the coloring process while maintaining the artistic style and cultural characteristics of the Thangka.

[0056] Preferably, step S1 includes:

[0057] Obtain Thangka images collected through on-site shooting;

[0058] Crop the Thangka images into square images of 256×256 pixels;

[0059] Extract the line drawings from the Thangka images to generate corresponding line drawing images.

[0060] It can be understood that Thangka is famous for its delicate lines and rich colors. Due to the complex and time-consuming process of drawing Thangka, the number of high-quality Thangka works is limited. To solve this problem, we adopted the method of on-site shooting and collection to obtain more Thangka image resources. After shooting, for the convenience of subsequent computer vision processing and deep learning training, we adjusted the size of each Thangka image uniformly and cropped it into a square image of 256×256 pixels. Such a size can not only ensure that the details of the image are not overly compressed but also meet the input requirements of most existing deep learning models. Subsequently, we extracted the line drawings from the collected Thangka images, including 10,030 line drawing images and 10,030 color images each. These images will be used to construct a dedicated Thangka dataset for training and testing deep learning models, with the expectation of automatically identifying and analyzing the artistic features of Thangka, thereby promoting the digital protection and research of Thangka art. Through such a dataset, we hope to improve the accessibility of Thangka art and also provide a powerful tool for artists and researchers to better understand and appreciate this unique cultural heritage.

[0061] Preferably, the Thangka line drawing coloring framework includes: initializing the coloring network and the quality enhancement diffusion model.

[0062] It can be understood that

[0063] The Thangka line drawing coloring framework in this application is divided into two stages, and the specific process is as Figure 2 shown. Different from other line drawing coloring methods, due to the complexity of Thangka line drawings, we designed a two-stage coloring method. In the first stage, we extract the features of the line drawing and the reference image for cross-attention calculation, perform feature fusion to form a preliminarily colored Thangka, and then perform forward diffusion and noise addition on the Thangka reference image, connect it with the line drawing image, and input them into the noise prediction network to output the predicted noise, and finally reach the denoising module to output the final color image.

[0064] Preferably, step S4 includes:

[0065] Divide the Thangka dataset into a training set and a test set;

[0066] Train the initialized coloring network with the training set and test the initialized coloring network with the test set.

[0067] Preferably, step S6 includes:

[0068] Performing initial coloring on the Thangka line drawing to be colored through an initial coloring network.

[0069] Preferably, it further includes:

[0070] Introducing color consistency loss, perceptual loss, style loss, and generation loss into the initial coloring network.

[0071] It can be understood that the coloring of the Thangka line drawing has a great impact on coloring due to the dense lines and complex semantic information. Therefore, we use an initial coloring network to perform a preliminary coloring on the Thangka. We use a self-built Thangka dataset (including paired Thangka color pictures and line drawings) to train the network. Since the colors of the Thangka are rich, we introduce a color consistency loss to constrain the colors in the initial coloring stage, and its formula is defined as follows:

[0072]

[0073] B is the batch size, which we set to 8, H represents the image height, and W represents the image width. represents the pixel value of the i-th sample in the generated image at (j, k), and I ijk represents the pixel value of the real image at the i-th sample at the position (j, k).

[0074] At the same time, for the problem of dense lines and rich semantics of the Thangka, we introduce perceptual loss and style loss to constrain the coloring, and their formulas are defined as follows:

[0075]

[0076] Among them represents the generated image, y represents the real image, and Gram represents the Gram matrix of the image.

[0077]

[0078] φ l (x) is the feature extraction function of the l-th layer extracted from the pre-trained VGG network, and a l is the weight of each layer of features. is the generated image, and y is the reference image.

[0079] In addition, since our model is a generative model, we introduce generation loss to make the training process more stable, and its formula is defined as follows:

[0080]

[0081] Among them represents the discriminator's evaluation of the generated image The output aims to make the generated image pass through the discriminator, and it is hoped that the output result is close to 1.

[0082] Preferably, the quality enhancement diffusion model is the DDIM model;

[0083] Quality enhancement of the Thangka line drawing to be colored is performed through the Thangka line drawing coloring framework, including:

[0084] Quality enhancement of the Thangka line drawing to be colored is performed through the DDIM model.

[0085] Preferably, the formula definition of the DDIM model is:

[0086]

[0087] where, represents the noisy image at time t'+1, represents the real image, represents the noisy image at time t′, t' represents the sampling time series, and ε θ represents the noise prediction network, represents the cumulative signal retention ratio up to time step t'+1.

[0088] It can be understood that only the Thangka line drawing after initial coloring has problems such as uneven colors and insufficient learning of color gradients. Therefore, we introduce diffusion for quality enhancement. Among them, the forward diffusion is a Markov chain process that transforms the image I gt into the noisy I t , and Gaussian noise is added at each time step. The forward process can be expressed as:

[0089]

[0090] where I noise represents the noisy image at time t, and I gt represents the real image, ε represents the added noise, represents the cumulative signal retention ratio up to time step t.

[0091] DDPM (Denoising Diffusion Probabilistic Models) simulates the process from the real data distribution to a simple distribution (such as the standard Gaussian distribution) by gradually adding noise. The core idea of DDPM is that through a forward process of a Markov chain, the data is gradually transformed into noise, and then through a reverse process, the data is gradually recovered from the noise. Since the reverse process of DDPM requires gradual denoising and the network needs to predict the noise at each step, this process is relatively slow. Therefore, we use DDIM (Denoising Diffusion Implicit Models). The training process of DDIM is the same as that of DDPM, but it improves the sampling process. DDIM is a non-Markov process that allows skipping steps in the denoising process without having to access all previous states at the current state, thus accelerating the inference speed of the model. DDIM allows for faster sampling and can provide deterministic output. By improving the sampling process, DDIM provides a faster generation method while maintaining the determinacy of the generated results, which is very valuable for application scenarios that require fast generation and reproducible results. Its formula definition is as follows:

[0092]

[0093] Where represents the noisy image at time t'+1, represents the real image, represents the noisy image at time t′, t' represents the sampling time series, and ε θ represents the noise prediction network, represents the cumulative signal retention ratio up to time step t'+1.

[0094] In specific practice, we evaluated the performance of AttDiffusion in the task of coloring Thangka line drawings. We implemented our method using the pytorch framework, used the python3.8 environment on an NVDIA 3090 GPU, and selected the Adam optimizer to train our model. To ensure the consistency and comparability of the experiments, we uniformly processed all images to a size of 256×256 pixels. This decision not only ensures the image quality but also optimizes the use of computing resources. To reveal the performance of AttDiffusion, we conducted a series of tests, including qualitative and quantitative analyses. Finally, we provided a comprehensive analysis to illustrate the operating mechanism of AttDiffusion.

[0095] In order to evaluate its performance in the thangka line drawing coloring task, we show some examples of thangka line drawing coloring in the figure. We selected three mainstream frameworks for comparison: Petalica Paint, Attention-Aware, and AnimeDiffusion. Figure 3 shown.

[0096] In the examples shown, we paid special attention to the coloring effects of the human faces and flowers with gradients in thangka. The coloring results of Petalica Paint performed poorly, with serious color overflow, resulting in unsatisfactory overall results. In the face coloring of Attention-Aware, the color is somewhat different from the reference image, and in the second face image, there is a problem of inconsistent color in the face coloring, and the semantic correspondence is not accurate enough. In addition, in the coloring of flowers, Attention-Aware failed to achieve the color gradient effect, and the color of the upper right corner also appeared messy. The overall color of AnimeDiffusion coloring is grayish, the semantic correspondence is poor, the face coloring image is very different from the reference color, and the flower color overflow is serious and the color is messy.

[0097] In contrast, AttDiffusion performs better in these aspects and can better handle color gradients and semantic correspondence of details, thus providing better coloring effects overall.

[0098] To evaluate the performance of AttDiffusion in the online manuscript colorization task, we used two mainstream frameworks for comparison: Petalica Paint and Attention-Aware. We used Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), and Learned Perceptual Patch Similarity (LPIPS) as evaluation metrics.

[0099] The results show that as shown in Table 1, our AttDiffusion performs well in the image colorization task. Specifically, it achieves significant improvements in PSNR and LPIPS indicators compared to Petalica Paint, almost doubling the improvement, which shows that AttDiffusion has obvious advantages in image detail recovery and color consistency. Although it is slightly lower than Attention-Aware in the SSIM indicator, the gap is only 0.01, which shows that the two perform quite close in terms of image structure similarity, and the gap of AttDiffusion is negligible. Compared with AnimeDiffusion, our method is higher than AnimeDiffusion in all evaluation indicators, further highlighting the comprehensive advantages of AttDiffusion in image colorization quality. Taking all indicators into consideration, our colorization results perform better overall, which fully demonstrates that AttDiffusion has better performance and higher quality in image colorization tasks, and can provide users with better quality and more natural image colorization effects.

[0100] Table 1 Comparison of quantitative methods

[0101] PSNR SSIM LPIPS petalica paint 8.9317 0.5608 0.44 Attention-ware 13.5540 0.7707 0.18 AnimeDiffusion 11.1008 0.7496 0.46 Our method 15.0369 0.7671 0.16

[0102] In order to prove the effectiveness of our modules, we conduct an ablation experiment, mainly testing the coloring results of only the initialization module, only the diffusion enhancement module, and all the modules. Figure 4 As shown in the figure, when only the coloring is initialized, the color overflow is serious, and there is a large area of ​​blur in the first thangka face coloring picture, the headdress color of the second thangka face is uneven, and the color content in the flowers is not smooth. In summary, the coloring is the smoothest and most uniform among all modules, which proves the effectiveness of our method.

[0103] It should be noted that this application designs an end-to-end network, which for the first time combines the attention mechanism with the diffusion model for Thangka line drawing coloring and can generate Thangka with good semantic correspondence. For the problem of difficult semantic correspondence due to the complex lines of Thangka, an initialization coloring network is designed. The Thangka line drawing and color image are input into the encoder to extract features. Subsequently, the features of the line drawing and the reference image are fused, and the preliminary colored image is output by the decoder. This method can generate a Thangka coloring image with good semantic correspondence and provides a starting point that conforms to the artistic characteristics of Thangka for the subsequent diffusion process. A diffusion enhancement network is designed to further refine the Thangka. The initialized colored image and the line drawing are input into the noise prediction network. After connecting the features of the initialization coloring network and the attention block and inputting them into the decoder to output the predicted noise, it reaches the denoising module to output the final colored image. Through this method, the color gradient and smoothness are achieved to improve the naturalness and artistry of the coloring effect and generate higher-quality Thangka. Due to the preciousness and scarcity of Thangka, the Thangka dataset is scarce. After on-site shooting and collection, we built a Thangka dataset by ourselves, which contains 10,030 pairs of reference images and line drawings each, for the training and testing of Thangka line drawing coloring.

[0104] The present invention also provides a device for coloring Thangka line drawings using a diffusion model to implement the above method embodiments. Figure 5 It is a schematic structural diagram provided by an embodiment of a device for coloring Thangka line drawings using a diffusion model of the present invention. As Figure 5 shown, the device for coloring Thangka line drawings using a diffusion model in this embodiment includes a processor 21 and a memory 22, and the processor 21 is connected to the memory 22. Among them, the processor 21 is used to call and execute the program stored in the memory 22; the memory 22 is used to store the program, and the program is at least used to execute the method for coloring Thangka line drawings using a diffusion model in the above embodiments.

[0105] The specific implementation scheme of the device for coloring Thangka line drawings using a diffusion model provided by the embodiments of this application can refer to the implementation manner of the method for coloring Thangka line drawings using a diffusion model in any of the above embodiments, and will not be elaborated here.

[0106] It can be understood that the same or similar parts in the above embodiments can be referred to each other, and the content not detailed in some embodiments can be seen in the same or similar content in other embodiments.

[0107] It should be noted that in the description of the present invention, terms such as "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present invention, unless otherwise specified, the meaning of "a plurality" refers to at least two.

[0108] Any process or method description depicted in a flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code that includes one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present invention includes additional implementations where functions may be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.

[0109] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0110] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0111] In addition, in each embodiment of the present invention, the functional units can be integrated into one processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0112] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc.

[0113] In the description of this specification, the descriptions with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0114] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A diffusion model thangka line drawing coloring method, characterized in that: include: Acquire a thangka image, extract a line drawing from the thangka image, and generate a line drawing; Constructing a thangka dataset according to the thangka image and the line drawing; Construct a coloring framework for thangka line drawings; Performing network training on the thangka line drawing coloring framework through the thangka data set; Get the thangka line drawing to be colored; Initializing and coloring the thangka line drawing to be colored by using the trained thangka line drawing coloring framework; The quality of the thangka line drawing to be colored is enhanced by the thangka line drawing coloring framework, and the output image and the thangka line drawing to be colored are input into the noise prediction network to output the predicted noise; The predicted noise is denoised by a denoising module to complete the coloring of the thangka line drawing to be colored.

2. The method according to claim 1, characterized in that The step of obtaining a thangka image, extracting a line drawing from the thangka image, and generating a line drawing includes: Acquire thangka images collected through field photography; The thangka image is cropped into a square image of 256×256 pixels; The line drawing of the thangka image is extracted to generate a corresponding line drawing.

3. The method according to claim 2, characterized in that The thangka line drawing coloring framework includes: an initialization coloring network and a quality enhancement diffusion model.

4. The method according to claim 3, characterized in that: The network training of the thangka line drawing coloring framework using the thangka dataset includes: Dividing the thangka dataset into a training set and a test set; The initialized coloring network is trained by the training set, and the initialized coloring network is tested by the test set.

5. The method according to claim 4, characterized in that The step of initializing and coloring the thangka line drawing to be colored by using the trained thangka line drawing coloring framework includes: The thangka line drawing to be colored is initialized and colored by initializing the coloring network.

6. The method according to claim 5, characterized in that The mass enhanced diffusion model is a DDIM model; The quality enhancement of the thangka line drawing to be colored by using the thangka line drawing coloring framework includes: The quality of the thangka line drawing to be colored is enhanced by using the DDIM model.

7. The method according to claim 6, characterized in that Also includes: Color consistency loss, perceptual loss, style loss and generation loss are introduced for the initialization colorization network.

8. The method according to claim 7, characterized in that The formula of the DDIM model is defined as: in, represents the noisy image at time t'+1, represents the real image, represents the noisy image at time t, t' represents the sampling time series, ε θ represents the noise prediction network, Indicates the cumulative signal retention ratio up to time step t'+1.

9. A diffusion model thangka line drawing coloring device, characterized in that: The invention comprises a processor and a memory, wherein the processor is connected to the memory: Wherein, the processor is used to call and execute the program stored in the memory; The memory is used to store the program, and the program is at least used to execute the diffusion model thangka line drawing coloring method described in any one of claims 1-8.