Method and device for extracting line draft graph of Tibetan colored drawing image based on diffusion model
Through the multi-layer convolutional network structure based on the diffusion model, the quality and accuracy problems in the extraction of Tibetan-style colored images are solved, and high-quality and high-precision line drawings are achieved, which is suitable for the separation of complex textures and the elimination of background interference.
Patent Information
- Application Number
- CN202510022531.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-06-24
AI Technical Summary
The existing image line drawing methods cannot meet the needs of high quality and high precision when processing Tibetan-style painted images, especially in terms of separation of complex textures, high-precision extraction of line drawings, elimination of background interference and improvement of computing efficiency.
Using a diffusion model-based method, the diffusion model of a multi-layer convolution network structure is gradually denoised and reconstructed from the original image, simplifying the texture information into the main line structure, and optimizing the process to obtain high-quality line drawings.
This method can effectively preserve the details in the complex texture of the hidden colored image, suppress noise interference, and generate more continuous lines and smoother edges, which conform to the natural characteristics of the hand-painted effect, improving the quality and accuracy of the line drawing.
Smart Images

Figure CN120198520A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data security, and particularly to a method and device for extracting line drawings of Tibetan painted images based on a diffusion model. Background Art
[0002] Tibetan painted art is an important part of Tibetan culture, famous for its rich historical heritage and unique artistic creativity. This art form widely exists in traditional artworks, characterized by complex lines and rich colors. With the development of digital technology, the demand for digital processing and protection of Tibetan painted images is increasing day by day. As a key step in digital processing, line drawing extraction is of great significance for protecting and inheriting this cultural heritage.
[0003] In summary, the existing image line drawing extraction methods cannot meet the requirements of high-quality and high-precision line drawing extraction when dealing with Tibetan painted images. Summary of the Invention
[0004] In view of this, the embodiments of this application are committed to providing a method and device for extracting line drawings of Tibetan painted images based on a diffusion model to more comprehensively and effectively analyze data security.
[0005] This application provides a method for extracting line drawings of Tibetan painted images based on a diffusion model, including:
[0006] Obtain the original image;
[0007] Input the original image into a preset diffusion model to perform line drawing extraction and obtain an initial line drawing;
[0008] The diffusion model adopts a multi-layer convolutional network structure. The diffusion model adopts a multi-layer convolutional network structure, which is used to start from the original image, gradually simplify the texture information into the main line structure by gradually denoising and reconstructing the image, optimize the initial line drawing, and obtain the initial line drawing.
[0009] Optimize the initial line drawing to obtain the target line drawing.
[0010] In some embodiments, it further includes:
[0011] Obtain a preset sample image;
[0012] Build and train a diffusion model based on the sample image.
[0013] In some embodiments, it further includes:
[0014] Adaptive adjust the number of iteration steps, diffusion intensity, and post - processing parameters of the diffusion model according to the detail complexity of the original image.
[0015] In some embodiments, it further includes:
[0016] If there are multiple original images, after extracting the line drawing of one original image, obtain another original image again and perform line drawing extraction until the line drawings of all original images are completed.
[0017] In some embodiments, it further includes:
[0018] Perform quality detection on the target line drawing;
[0019] Adjust the diffusion model based on the quality detection result.
[0020] After obtaining the original image in some embodiments, it further includes:
[0021] Perform format conversion, grayscale processing, denoising, and enhancement on the original image.
[0022] In some embodiments, the optimizing the initial line drawing to obtain the target line drawing includes:
[0023] Remove background noise and redundant details in the initial line drawing;
[0024] Perform smoothing processing and thinning processing on the lines in the initial line drawing;
[0025] Adjust the line thickness to obtain the target line drawing.
[0026] This application also provides a Tibetan painted image line drawing extraction device based on a diffusion model, including:
[0027] An acquisition module, configured to acquire an original image;
[0028] An extraction module, configured to input the original image into a preset diffusion model to perform line drawing extraction to obtain an initial line drawing;
[0029] The diffusion model adopts a multi - layer convolutional network structure. The diffusion model adopts a multi - layer convolutional network structure, which is used to start from the original image, gradually simplify the texture information into the main line structure by gradually denoising and reconstructing the image, optimize the initial line drawing, and obtain the initial line drawing.
[0030] An optimization module, configured to optimize the initial line drawing to obtain the target line drawing.
[0031] This application also provides an electronic device, including:
[0032] A processor and a memory for storing a program executable by the processor;
[0033] The processor is configured to implement the method for extracting the line drawing of the Tibetan painted image based on the diffusion model as described above by running the program in the memory.
[0034] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the processor is caused to execute the method for extracting the line drawing of the Tibetan painted image based on the diffusion model as described above.
[0035] A method for extracting the line drawing of a Tibetan painted image based on a diffusion model provided by the present application includes obtaining an original image; inputting the original image into a preset diffusion model for line drawing extraction to obtain an initial line drawing; the diffusion model adopts a multi-layer convolutional network structure. The diffusion model adopts a multi-layer convolutional network structure, which is used to start from the original image, gradually denoise and reconstruct the image, and gradually simplify the texture information into the main line structure. Optimize the initial line drawing to obtain the initial line drawing. Optimize the initial line drawing to obtain the initial line drawing. Optimize the initial line drawing to obtain the target line drawing. With such a setting, the diffusion model adopts a multi-layer convolutional network structure, which can gradually denoise and reconstruct the image from the original image, and simplify the texture information into the main line structure. This process of gradually denoising can effectively retain the details in the complex texture of the Tibetan painted image while suppressing noise interference, which is very crucial for the complex line and color characteristics of the Tibetan painted image. Compared with the traditional edge detection method, when processing the original image, this method can make the generated lines more continuous, the edges smoother, and less likely to break through the diffusion model and subsequent optimization processing, and the generated line drawing is more in line with the natural characteristics of the hand-drawn effect, improving the quality of the line drawing. The diffusion model can learn the global and local features of the image during the training process, which makes the model have high stability. And by appropriately adjusting the model parameters, the control of the line thickness and clarity can be realized, so as to better meet the requirements for the quality of the target line drawing in different application scenarios. Description of the Drawings
[0036] By describing the embodiments of the present application in more detail in conjunction with the drawings, the above and other objects, features, and advantages of the present application will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application, and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0037] Figure 1It is a schematic flowchart of a method for extracting a line drawing of a Tibetan painted image based on a diffusion model provided by an embodiment of the present application.
[0038] Figure 2 It is a partial schematic flowchart of the method provided by an embodiment of the present application.
[0039] Figure 3 It is a partial schematic flowchart of the method provided by another embodiment of the present application.
[0040] Figure 4 It is a partial schematic flowchart of the method provided by yet another embodiment of the present application.
[0041] Figure 5 It is a partial schematic flowchart of the method provided by yet another embodiment of the present application.
[0042] Figure 6 It is a partial schematic flowchart of the method provided by an embodiment of the present application.
[0043] Figure 7 It is a schematic structural diagram of a device for extracting a line drawing of a Tibetan painted image based on a diffusion model provided by an embodiment of the present application.
[0044] Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0046] Application Overview
[0047] In early research, the extraction of the line drawing of an image was achieved by designing edge extraction operators. Such as the Sobel edge detection operator and the Canny edge detection operator. The Canny edge detection operator first smooths and denoises the grayscale image through Gaussian filtering, calculates the gradient magnitude and direction of the gray value of the pixel points in the image, then takes the positions with large gradient changes as the image edges and refines them using the non-maximum suppression method, and finally uses the double-threshold detection method to extract the image edges. This edge detection operator can capture as many edges in the image as possible and is still widely used at present, but its rough extraction performance is not suitable for Tibetan painted images with delicate lines. Felz-Hutt combined edge detection with image segmentation and extracted edge information by fusing global features and local features, transforming any contour information into hierarchical regions. However, the painting styles of Tibetan painted images are diverse and the colors are complex, making it difficult to extract the contours. The processing flow of the gPb-owt-ucm method is that the gPb contour detector extracts features in a multi-scale manner on three channels of brightness, color, and texture and converts them into soft boundary maps in all directions at each position. Generally, when dividing contours, each pixel point performs binary classification, that is, it is a contour or not a contour. SketchTokens converts the problem into multi-class classification, reducing the difficulty of the classifier to a certain extent, and using the manually labeled contour information as the class label for multi-class classification. Therefore, the labeling work is complex and heavy.
[0048] Recently, the basic tasks based on Transformer have been extended and applied to visual tasks. It is applied together with CNN in DETR and other variants. The Vision Transformer (ViT) directly uses Transformer to map into a sequence of image patches and achieves good results. However, experiments show that this method is only suitable for images with single colors and simple lines and is not suitable for complex Tibetan painted images. Tibetan painted images have complex lines, numerous colors, and diverse painting styles. Therefore, it is difficult to accurately extract the line drawing of Tibetan paintings. Although the above network models can roughly extract the line drawing of the image, the extraction effect for complex Tibetan painted images is poor.
[0049] This application aims to solve the technical problem of extracting the line drawing of Tibetan painted images, and the application object is Tibetan painted images with extremely complex lines and colors. It mainly includes the following four points:
[0050] (1) Separation of complex textures in Tibetan painted images. Tibetan painted art usually contains delicate and complex lines and colors. Existing line drawing extraction methods may have problems such as distortion, detail loss, or color interference when dealing with these complex textures.
[0051] (2) High-precision extraction of image line drawings. The lines and edge details of Tibetan painted sculptures are rich, and it is difficult for traditional methods to accurately extract the line drawing structure, resulting in unclear results or lack of artistry and aesthetics.
[0052] (3) Elimination of background interference. Tibetan painted sculpture images contain many background elements, and the commonly used edge detection methods are easily interfered by these elements when extracting line drawings, resulting in inaccurate line drawing extraction.
[0053] (4) Efficient calculation and model optimization. Since Tibetan painted sculpture images usually have high resolution and complex artistic designs, improving the calculation efficiency is also an important technical issue while ensuring accurate extraction.
[0054] This application aims to overcome the above problems by introducing diffusion model technology, achieve higher-quality and higher-precision line drawing image extraction, and provide technical support for the digital processing and application of Tibetan painted sculpture art. Based on the high-precision line drawing extraction ability of the diffusion model, it can effectively separate clear and complete lines in complex Tibetan painted sculpture images, and significantly reduce background interference and noise. Specific advantages include: (1) Retention of complex details and noise suppression. The diffusion model can gradually denoise, thus smoothly extracting the line structure in the image, retaining details and reducing noise interference. This is very effective in the processing of complex textures of Tibetan painted sculpture images. (2) Strong adaptability and adaptation to various image features. The diffusion model has good adaptability and can show high robustness in Tibetan painted sculpture images of different styles and resolutions. Regardless of the amount of image details, high-quality line drawings can be obtained. (3) Image refinement and enhanced continuity. Compared with traditional edge detection methods, the lines generated by the diffusion model are more continuous, the edges are smoother, and are not prone to breakage, and the generated line drawings are more in line with the natural characteristics of hand-drawn effects. (4) High model stability and controllable generation quality. The diffusion model can learn the global and local features of the image during training, with high and stable generation quality. By appropriately adjusting the model parameters, the control of line thickness and clarity can be achieved.
[0055] After introducing the basic principle of this application, various non-limiting embodiments of this application will be specifically introduced with reference to the accompanying drawings.
[0056] Exemplary Method
[0057] Figure 1 is a schematic flowchart of a method for extracting a line drawing of a Tibetan painted sculpture image based on a diffusion model provided by an embodiment of this application. As Figure 1 shown, this method includes the following content.
[0058] Step S110, obtain the original image;
[0059] This is the starting step of the whole process. In this step, the original images of Tibetan painted sculptures need to be obtained from the corresponding data sources. These original images can be obtained in various ways, such as scanning Tibetan painted sculptures using a scanner, or photographing the physical objects of Tibetan painted sculptures with a camera, etc. The obtained original images contain the complete information of Tibetan painted sculptures, including their complex lines, rich colors, and unique painting styles, etc., but may also contain some noise and unnecessary details.
[0060] Step S120: Input the original image into a preset diffusion model to extract a sketch map and obtain an initial sketch map.
[0061] The diffusion model adopts a multi-layer convolutional network structure. The diffusion model adopts a multi-layer convolutional network structure, which is used to start from the original image, gradually simplify the texture information into the main line structure by gradually denoising and reconstructing the image, optimize the initial sketch map, and obtain the initial sketch; optimize the initial sketch map to obtain the initial sketch map.
[0062] First, input the original image obtained in step S110 into a preset diffusion model. This diffusion model is specifically designed for processing Tibetan painted sculpture images and adopts a multi-layer convolutional network structure.
[0063] After the original image enters the diffusion model, the model starts to process it from the original image. It extracts the sketch information through the process of gradually denoising and reconstructing the image. Specifically, the model simulates the gradual change of pixel values in the image. In this process, on the one hand, the noise in the image is removed, and on the other hand, the line information is retained and highlighted. As the processing progresses step by step, the texture information of the image will be gradually simplified, and finally, the main line structure will be formed, thus obtaining the initial sketch map.
[0064] It should be emphasized here that the multi-layer convolutional network structure of the diffusion model can better adapt to the complex texture and line characteristics of Tibetan painted sculpture images. It can perform multi-level analysis and processing on the image, so as to more accurately extract the sketch information.
[0065] Step S130: Optimize the initial sketch to obtain a target sketch map.
[0066] After obtaining the initial sketch map, it is also necessary to perform further optimization processing on it. This step is to improve the quality of the sketch map and make it more in line with the actual application requirements. The optimization process may include operations such as adjusting the clarity of the lines, removing residual noise, further refining the lines, and enhancing the continuity of the lines. Through these optimization measures, the initial sketch map can be further improved, and finally, a target sketch map with higher quality and more in line with the requirements can be obtained.
[0067] Generally speaking, these steps together constitute a complete processing flow from the original image to a high-quality line drawing. Each step is to ensure the accuracy and aesthetics of the final result.
[0068] Furthermore, the solution provided by this application also includes: obtaining a preset sample image; constructing and training a diffusion model based on the sample image.
[0069] To construct and train an effective diffusion model, a set of preset sample images need to be obtained first. These sample images are crucial for the model to learn the characteristics of Tibetan painted images. The sample images should be representative, covering various styles, line complexities, and color variations of Tibetan painted works. They can be from existing Tibetan painted image databases or collected and sorted by professionals. The sample image set should not only include the original Tibetan painted images but also the corresponding accurate line drawings. These corresponding line drawings can be used as reference standards during the training process to help the model learn how to accurately extract line drawing information from the original image.
[0070] The diffusion model adopts a multi-layer convolutional network structure. This structural design can effectively process the complex textures and line information of Tibetan painted images. When constructing the model, it is necessary to determine the structural parameters of each layer of the model, such as the number of layers, filter size, stride, etc. The settings of these parameters should be adjusted according to the characteristics of Tibetan painted images to ensure that the model can fully learn the features of the images.
[0071] Input the obtained sample images into the constructed diffusion model for training. During the training process, the model will learn various features in the sample images, especially the relationship between the original image and the corresponding line drawing.
[0072] A specific loss function is used during the training process to optimize the model parameters. This loss function usually consists of two parts: denoising loss and structure reconstruction loss. The denoising loss is used to measure the effect of the model in removing image noise, and the structure reconstruction loss is used to evaluate the ability of the model to simplify the texture information of the original image into the main line structure. By continuously adjusting the model parameters to minimize the value of the loss function, the model can gradually improve the accuracy and effect of line drawing extraction. The model will repeatedly learn the sample image dataset. Through multiple iterative trainings, the model can better adapt to various features of Tibetan painted images, thereby improving its performance in line drawing extraction in practical applications.
[0073] In some embodiments, the method for extracting the line drawing of a Tibetan painted image based on a diffusion model further includes: adaptively adjusting the number of iteration steps, diffusion intensity, and post-processing parameters of the diffusion model according to the detail complexity of the original image.
[0074] Tibetan painted images vary greatly in detail. Some images may have relatively simple lines and fewer color levels, while others have extremely complex lines and rich color variations. Therefore, it is necessary to evaluate and analyze the detail complexity of the original images. This can be achieved in various ways, such as analyzing the texture features, line density, color distribution, etc. of the images. Through these analyses, it can be determined whether the original images belong to the categories of simple, medium, or complex detail levels.
[0075] For original images with low detail complexity, perhaps not many iteration steps are needed to extract a relatively accurate line drawing. Because the images themselves are relatively simple, the diffusion model can quickly identify and extract the main line structures. For images with high detail complexity, such as those Tibetan painted images with complex textures and dense lines, more iteration steps are required. This allows the diffusion model to have sufficient time and opportunities to gradually denoise and refine the lines to ensure that a high-quality line drawing can be extracted.
[0076] The diffusion intensity determines the degree of change of the model to the image during the denoising and line extraction processes. For original images with simple details, the diffusion intensity can be appropriately reduced. Because there are fewer interference factors in the images themselves, a good extraction effect can be achieved without too strong a diffusion process. On the contrary, for complex original images, the diffusion intensity may need to be increased to more effectively remove noise and highlight the line structures, but at the same time, attention should be paid to avoiding excessive diffusion resulting in the loss or distortion of line information.
[0077] The post-processing process includes operations such as edge enhancement, contrast adjustment, line smoothing, and thinning of the extracted line drawing. For original images with different detail complexities, these post-processing parameters also need to be adjusted accordingly. For example, for images with simple details, perhaps only slight edge enhancement and line smoothing are needed; while for complex images, stronger edge enhancement, higher contrast adjustment, and more refined line smoothing and thinning processing may be required to ensure that the quality of the final obtained line drawing meets the requirements.
[0078] By adaptively adjusting the iteration steps, diffusion intensity, and post-processing parameters of the diffusion model according to the detail complexity of the original images, the efficiency and quality of line drawing extraction can be improved. This adaptive adjustment enables the model to better adapt to different types of Tibetan painted images, avoiding the problem of poor extraction effects that may be caused by using unified parameter settings for all images. Whether it is a simple or complex original image, a more accurate and higher-quality line drawing can be obtained, thus enhancing the applicability and practicality of the entire method.
[0079] In some embodiments, the method for extracting line drawings of Tibetan painted images based on a diffusion model further includes: if there are multiple original images, after completing the extraction of the line drawing of one original image, another original image is retrieved and the line drawing is extracted until the extraction of the line drawings of all original images is completed.
[0080] When faced with multiple original images that need to have their line drawings extracted, the system will process each image sequentially. First, for the first original image, all the steps of the method for extracting line drawings based on the diffusion model will be fully executed. That is, starting from obtaining the original image, through inputting it into the diffusion model for line drawing extraction to obtain an initial line drawing, and then optimizing the initial line drawing to obtain the target line drawing. After the extraction of the line drawing of the first original image is completed, the system will retrieve the next original image and repeat the above line drawing extraction process. This process will continue until the line drawings of all original images are completed. By processing multiple original images sequentially, system resources can be fully utilized, avoiding the waste of time in separately setting and adjusting each image. During the processing, the system can maintain a relatively stable working state and quickly switch to the processing of the next image, thereby improving the overall processing efficiency. Using the same method for extracting line drawings based on the diffusion model for each original image ensures the consistency of the processing process and results. This is very important for situations where a large number of Tibetan painted images need to be uniformly processed and analyzed. For example, when digitizing a batch of Tibetan painted works for archiving, a unified processing method can ensure that the quality and style of the line drawings extracted from all images are consistent, facilitating subsequent management and application.
[0081] In some embodiments, the method for extracting line drawings of Tibetan painted images based on a diffusion model further includes: performing quality detection on the target line drawing; and adjusting the diffusion model based on the quality detection results.
[0082] When performing quality detection on the target line drawing, multiple indicators are involved. For example, the clarity of the lines is an important indicator, which measures whether the lines in the line drawing are clear and distinguishable, and whether there are any blurred or broken situations. Contrast is also one of the key indicators. Appropriate contrast can make the lines more prominent, facilitating observation and subsequent applications. In addition, the noise level is also something that needs to be detected. A lower noise level means that the line drawing is purer and there are no excessive interference factors. Multiple methods can be used to detect these indicators. For the clarity of the lines, it can be evaluated by calculating the sharpness of the line edges. Contrast can be determined by analyzing the gray histogram of the image, observing the distribution of different gray levels. The noise level can be measured by some noise estimation algorithms, such as algorithms based on the differences of neighboring pixels.
[0083] If the quality inspection results show that there are some problems with the line drawing, such as insufficient line clarity, then the diffusion model needs to be adjusted according to this result. If the line clarity decreases because the model is overly smoothed during the denoising process, then it may be necessary to adjust the denoising parameters of the diffusion model and appropriately reduce the denoising intensity. If the line refinement is insufficient, it may be necessary to adjust the structure or parameters related to line refinement in the model. The diffusion model can be adjusted in various ways. For example, the training parameters of the model, such as the learning rate and the number of iterations, can be readjusted, and then the model can be retrained. The structure of the model can also be fine-tuned, such as adjusting the filter size or number of convolutional layers. Through these adjustment measures, the diffusion model can better adapt to the characteristics of the original image, thereby improving the quality of line drawing extraction.
[0084] By performing quality inspection on the target line drawing and adjusting the diffusion model according to the results, the process and results of line drawing extraction can be continuously optimized. For each original image, this feedback mechanism can be used to improve the quality of the line drawing, making it more in line with the requirements of practical applications. As the diffusion model is continuously adjusted according to the quality inspection results, the model can better adapt to the characteristics of different original images. This helps to improve the performance of the model when processing various Tibetan painted images, enabling it to more accurately extract high-quality line drawings.
[0085] In some embodiments, in the method for extracting a line drawing of a Tibetan painted image based on a diffusion model, after obtaining the original image, the following steps are further included: performing format conversion, grayscale processing, denoising, and enhancement on the original image.
[0086] Specifically, a high-resolution color image of a Tibetan painting is obtained from an input device (such as a scanner or a camera) and converted into a standard format (such as PNG or JPEG) for subsequent processing. The color image is converted into a grayscale image to simplify the processing process and avoid interference from color information. An adaptive filtering method (such as Gaussian filtering or median filtering) is used to pre-denoise the image, and at the same time, the image details are enhanced by adjusting the contrast to ensure line clarity.
[0087] The optimization of the initial line drawing to obtain the target line drawing includes: removing background noise and redundant details from the initial line drawing; performing smoothing and refinement processing on the lines in the initial line drawing; adjusting the line thickness to obtain the target line drawing.
[0088] The initial line drawing may contain background noise from various sources. For example, device noise that may be introduced during image acquisition, or noise remaining due to imperfect algorithms during the previous line drawing extraction process. Such noise interferes with the clarity and accuracy of the line drawing. To remove the background noise, erosion operations in morphological operations can be employed. The erosion operation scans the image with a structuring element, removing noise points and isolated pixels smaller than the structuring element, thereby purifying the line drawing. Then, dilation operations can be carried out to fill the small holes generated during the erosion process, making the line drawing more complete. Filtering techniques such as Gaussian filtering can also be utilized. Gaussian filtering performs weighted average processing on the image based on the Gaussian function, effectively reducing the noise level and making the line drawing clearer.
[0089] There may be some redundant details in the initial line drawing. These details are not the core line structures of the line drawing but may affect the simplicity and overall effect of the line drawing. For example, some minor texture changes or overly fine branch lines. The redundant details can be identified and removed by setting thresholds. For regions whose gray values are below or above a certain threshold, if these regions do not belong to the main line structures, they can be removed. Image segmentation methods can also be adopted to divide the line drawing into different regions, and then identify and remove those regions that are not the main line structures, thereby highlighting the main line structures of the line drawing.
[0090] The lines in the initial line drawing may have problems such as discontinuity and jaggedness, affecting the aesthetics and quality of the line drawing. Bilateral filtering is an effective method for line smoothing. When calculating the pixel value, it not only considers its own pixel value but also the pixel values in the neighborhood, and realizes line smoothing by performing weighted average on the pixel values in the neighborhood. This method can better retain the edge information of the line while smoothing the line. In addition, morphological thinning algorithms can also be used for line smoothing. For example, template-based thinning algorithms perform specific operations on the pixels at the line edges to make the lines smoother.
[0091] Line thinning processing aims to make the lines more refined, meeting the requirements of artistic aesthetics and practical applications. Morphological thinning algorithms can be adopted to gradually thin the lines by continuously removing the redundant pixels at the line edges. For example, using the connectivity thinning algorithm, through iterative processing, the lines gradually become thinner to achieve a satisfactory thinning effect.
[0092] Different application scenarios have different requirements for the line thickness. For example, in digital displays, thinner lines may better reflect a sense of refinement; while in printing applications, thicker lines may be more conducive to ensuring printing quality and clarity. Therefore, it is necessary to adjust the line thickness according to the actual application scenario. During the process of line drawing extraction based on the diffusion model, the line thickness can be adjusted by modifying the relevant parameters in the model. For example, the parameters of the convolutional layer related to line generation, etc. In subsequent optimization processing, the line thickness can also be adjusted through image processing algorithms. For example, by changing the width value of the line or using scaling operations to change the line thickness. Through these methods, the requirements for line thickness in different application scenarios can be met, and the target line drawing can be obtained.
[0093] The following further illustrates the solution provided by this application with reference to Figures 2 to 6 specific embodiments:
[0094] Referring to Figure 2 , the upper left corner is the pixel space. From left to right, the three pictures are the real line drawing, the color image, and the extracted (predicted) line drawing (where part of the content of the picture is blocked). The lower left corner is the main line of the model, the upper right corner is the breakdown illustration of the Generative Adversarial Networks (GAN), the lower right corner is the breakdown illustration of the Adaptive Gaussian Filter (Adaptive GFT), and the illustration of the unlabeled elements in the figure.
[0095] Specifically, in the figure, ① is that the real line drawing is input into the Encoder for feature extraction and dimensionality reduction. The Encoder can convert image data into a format that can be further analyzed by the model. Dimensionality reduction helps reduce the storage and computational requirements of the data while retaining the main features of the image; ② is that the real line drawing is compressed into a low-dimensional vector z0 through a series of non-linear transformations by the autoencoder; ③ is the forward diffusion process, which transforms the data into pure noise zt by gradually adding noise to the data; ④ is the reverse generation process, which learns to gradually remove the noise from the noise and recover the original data; ⑤ is that the color image passes through a Condition encoder, which takes the input image and conditional information as inputs and outputs an encoded vector or representation through the processing of the neural network; ⑥ is that after passing through a UNet-based conditional encoder, it receives the noisy image and time step as inputs and outputs a denoised image. By iterating this process, high-quality images can finally be generated from the pure noise image; ⑦ is that the compressed vector of the color image after passing through the conditional encoder is merged with the compressed vector of the line drawing image after passing through the encoder and input into the decoder; ⑧ is that the compressed vector output by the encoder of the line drawing image passes through two adaptive Gaussian filters to filter out invalid information and retain valid information; ⑨ is that the valid information output by the adaptive Gaussian filter passes through two Decoders respectively; ⑩ is that after being processed by the decoder, it is converted into a vector highly related to the original line drawing image This conversion process aims to restore the information of the original image as accurately as possible. Is a vector highly related to the original line drawing image Passes through a decoder for restoring to a real image. Is the line drawing restored and generated after passing through the encoder. Is to input the generated line drawing into the Generative Adversarial Network (GAN). Is to input the real line drawing into the Generative Adversarial Network (GAN). Is to judge the generated line drawing as fake. Is to judge the real line drawing as real. Is to reward and punish the generated line drawing and the real line drawing through the UNet model, gradually reducing the loss. Is that the loss output after passing through the Generative Adversarial Network (GAN) acts on the vector highly related to the original line drawing image In which, loop optimization is performed until the optimal effect is achieved. It is a process of flattening and mapping in an Adaptive Gaussian Filter (Adaptive GFT), aiming to split an image into several parts and map them into an Adaptive Gaussian Filter (Adaptive GFT) to filter out invalid information and clear the way for subsequent processing.
[0096] Refer to Figure 3 In some embodiments, the method for extracting the line drawing of a Tibetan painted image in a diffusion model includes:
[0097] Image preprocessing
[0098] 1. Input image preparation: Obtain a high-resolution color image of a Tibetan painting from an input device (such as a scanner or a camera) and convert it into a standard format (such as PNG or JPEG) for subsequent processing.
[0099] 2. Grayscale processing: Convert the color image into a grayscale image to simplify the processing process and avoid interference from color information.
[0100] 3. Denoising and enhancement: Use an adaptive filtering method (such as Gaussian filtering or median filtering) to pre-denoise the image, and at the same time enhance the image details through contrast adjustment to ensure the clarity of the line drawing.
[0101] Constructing and training a diffusion model
[0102] Dataset preparation: Prepare a dataset of Tibetan painted images, including a large number of labeled line drawings and corresponding painted images, for model training.
[0103] 2. Diffusion model structure design: Use a multi-layer convolutional network structure to implement the diffusion model, and set the structures of each layer of the diffusion model (including the number of layers, filter size, stride, etc.) to adapt to the complex textures of Tibetan paintings.
[0104] 3. Model training: Through repeated iterative training, optimize the loss function of the model so that it can automatically extract line information from complex textures. The loss function can be designed to include denoising loss and structure reconstruction loss to balance image smoothness and line clarity.
[0105] 4. Model saving: Save the trained diffusion model as a model file for inference.
[0106] Diffusion process and line drawing extraction
[0107] 1. Initialize the diffusion process: Input the preprocessed image into the diffusion model for step-by-step denoising. The diffusion process starts from the original image and gradually simplifies the texture information into the main line structure by step-by-step denoising and image reconstruction.
[0108] 2. Gradually extract the line drawing: In each iteration step, the image extracted by the diffusion model will gradually approach a pure line structure. The model saves the generated image at each step to observe and adjust the extraction effect of each step.
[0109] 3. Line enhancement: Use edge enhancement algorithms (such as Sobel or Laplace edge detection) to further highlight the extracted lines and ensure the clarity and continuity of the lines.
[0110] Adaptive Gaussian filter post - processing and optimization based on GAN network
[0111] 1. Remove background and redundant details: Remove background noise and redundant details in the line drawing through morphological operations (such as erosion and dilation) to keep the pure lines.
[0112] 2. Smooth and thin the line drawing: Smooth the lines using bilateral filtering or morphological thinning algorithms to make the line edges smoother and the details clearer.
[0113] 3. Adjust the line thickness: Adjust the line thickness according to user requirements or output requirements to make the line drawing suitable for subsequent applications (such as printing or digital display).
[0114] Output and saving
[0115] 1. Format conversion: Convert the final line drawing image to common image formats (such as PNG, SVG) for easy storage and use.
[0116] 2. Save the output: Save the line drawing image to the specified path and provide a preview function for users to check and confirm the quality of the line drawing.
[0117] Furthermore, to improve the automation degree of the line drawing extraction process, the following refinement steps can be added: Adaptive parameter adjustment: Introduce an automated parameter optimization module to adaptively adjust the iteration steps, diffusion intensity, and post - processing parameters of the diffusion model according to the different detail complexities of the input image. The process is as Figure 3 shown. Figure 4
[0118] Furthermore, in some embodiments, a batch image input and processing function is also implemented, which is applicable to the scenario of large - scale line drawing extraction of Tibetan painted images. The process is as Figure 5 shown.
[0119] Furthermore, in some embodiments, a line drawing quality detection module can also be added. Using image clarity, contrast, and noise evaluation metrics, automatically judge whether the extraction effect meets the expected standard to achieve quality control and feedback regulation. The process is as Figure 6 shown.
[0120] The above refinement scheme ensures that high-quality line drawings can be efficiently and automatically extracted in different application scenarios, increasing the flexibility and adaptability of the method.
[0121] Exemplary Device
[0122] The device embodiment of the present application can be used to execute the method embodiment of the present application. For details not disclosed in the device embodiment of the present application, please refer to the method embodiment of the present application.
[0123] Figure 7 The block diagram of a Tibetan painted image line drawing extraction device based on a diffusion model provided by an embodiment of the present application is shown. As Figure 7 shown, the device includes:
[0124] An acquisition module 41, configured to acquire data lineage events; wherein, the data lineage events include events of data collection behavior, storage behavior, usage behavior, processing behavior, transmission behavior, provision behavior, public behavior, and destruction behavior;
[0125] An execution module 42, configured to execute all behaviors included in the data lineage events based on a preset virtual execution environment to obtain a data dictionary corresponding to the data lineage events; wherein, the data dictionary is used to characterize the classification and status of data related to the data lineage events;
[0126] When it is determined that the data dictionary includes illegal data, an alarm is given.
[0127] Exemplary Electronic Device
[0128] Next, an electronic device according to an embodiment of the present application will be described with reference to Figure 8 to describe an electronic device according to an embodiment of the present application. Figure 8 The block diagram of an electronic device according to an embodiment of the present application is illustrated.
[0129] As Figure 8 shown, the electronic device 800 includes one or more processors 810 and a memory 820.
[0130] The processor 810 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 800 to execute desired functions.
[0131] The memory 820 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor 810 may run the program instructions to implement the method for extracting the line drawing of the Tibetan painted image based on the diffusion model in various embodiments of the present application described above and / or other desired functions. Various contents such as category correspondence relationships may also be stored in the computer-readable storage media.
[0132] In one example, the electronic device 800 may further include: an input device 830 and an output device 840, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0133] In addition, the input device 830 may further include, for example, a keyboard, a mouse, an interface, etc. The output device 840 may output various information to the outside, including analysis results, etc. The output device 840 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0134] Of course, for simplicity, Figure 8 only some of the components related to the present application in the electronic device are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, according to specific application scenarios, the electronic device may further include any other appropriate components.
[0135] Exemplary Computer Program Product and Computer Readable Storage Medium
[0136] In addition to the above methods and devices, the embodiments of the present application may also be a computer program product, which includes computer program instructions that, when run by a processor, cause the processor to execute the steps in the method for extracting the line drawing of the Tibetan painted image based on the diffusion model according to various embodiments of the present application described in the "Exemplary Method" section of this specification.
[0137] The computer program product may be written in any combination of one or more programming languages for executing the program code of the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0138] In addition, an embodiment of the present application may also be a computer-readable storage medium having computer program instructions stored thereon, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the method for extracting a line drawing of a Tibetan painted image based on a diffusion model according to various embodiments of the present application described in the "Exemplary Method" section above of this specification.
[0139] The computer-readable storage medium may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0140] The foregoing description has been presented for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present application to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A method for extracting line drawings of Tibetan painted images based on a diffusion model, characterized in that: include: Get the original image; Inputting the original image into a preset diffusion model to extract a line drawing to obtain an initial line drawing; The diffusion model adopts a multi-layer convolutional network structure, which is used to start from the original image, gradually simplify the texture information into the main line structure by gradually denoising and reconstructing the image, and optimize the initial line draft to obtain the initial line draft; optimize the initial line draft to obtain the initial line draft; The initial line draft is optimized to obtain a target line draft diagram.
2. The method for extracting line drawings of Tibetan painted images based on a diffusion model according to claim 1, characterized in that: Also includes: Get the preset sample image; A diffusion model is constructed and trained based on the sample images.
3. The method for extracting line drawings of Tibetan painted images based on a diffusion model according to claim 1, characterized in that: Also includes: According to the detail complexity of the original image, the number of iteration steps, diffusion intensity and post-processing parameters of the diffusion model are adaptively adjusted.
4. The method for extracting line drawings of Tibetan painted images based on a diffusion model according to claim 1, characterized in that: Also includes: If there are multiple original images, after completing the line drawing extraction for one original image, another original image is reacquired and the line drawing extraction is performed until the line drawing extraction for all original images is completed.
5. The method for extracting line drawings of Tibetan painted images based on a diffusion model according to claim 1, characterized in that: Also includes: Performing quality inspection on the target line drawing; Based on the quality detection structure, the diffusion model is adjusted.
6. The method for extracting line drawings of Tibetan painted images based on a diffusion model according to claim 5, characterized in that: After obtaining the original image, the method further includes: The original image is format converted, grayed, denoised and enhanced.
7. The method for extracting line drawings of Tibetan painted images based on a diffusion model according to claim 5, characterized in that: The optimizing the initial line draft to obtain the target line draft includes: Removing background noise and redundant details from the initial line drawing; Smoothing and thinning the lines in the initial line draft; Adjust the line thickness to get the target line drawing.
8. A device for extracting line drawings of Tibetan painted images based on a diffusion model, characterized in that: include: An acquisition module, used for acquiring the original image; An extraction module, used for inputting the original image into a preset diffusion model to extract a line drawing to obtain an initial line drawing; The diffusion model adopts a multi-layer convolutional network structure. The diffusion model adopts a multi-layer convolutional network structure, which is used to start from the original image, gradually simplify the texture information into the main line structure by gradually denoising and reconstructing the image, and optimize the initial line draft to obtain the initial line draft; Optimizing the initial line draft to obtain an initial line draft; The optimization module is used to optimize the initial line draft to obtain a target line draft.
9. An electronic device, characterized in that: include: A processor, and a memory for storing a program executable by the processor; The processor is used to implement the Tibetan painted image line drawing extraction method based on the diffusion model as described in any one of claims 1 to 7 by running the program in the memory.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the processor executes the method for extracting line drawings of Tibetan painted images based on a diffusion model as described in any one of claims 1 to 7.