Tibetan thangka image restoration method and system based on GAN network

Through the adaptive channel and spatial attention mechanism of the OA-GAN network, combined with multiple loss functions, the problems of pigment feature loss and detail blurring in Tibetan thangka image restoration are solved, and the restoration effect of natural color gradient and accurate semantics is achieved.

CN120612256APending Publication Date: 2025-09-09QINGHAI NORMAL UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510708606.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

When restoring Tibetan thangka images, existing technologies are prone to losing the characteristics of mineral pigments, resulting in color scale breaks and semantic distortion, and inaccurate detail restoration.

Method used

The OA-GAN network is used to extract mineral pigment features through the adaptive channel attention mechanism, and the texture features are extracted by combining the dynamic spatial attention mechanism. In the discriminator, the L1 reconstruction loss function, color smoothing loss function, perceptual loss function and adversarial loss function are used to construct a hybrid loss combination function for image restoration.

Benefits of technology

The mineral pigment characteristics and texture details of the thangka are effectively preserved. The color of the restored image changes naturally, which conforms to the semantic characteristics of Tibetan thangka and avoids color level breaks and blurred details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612256A_ABST
    Figure CN120612256A_ABST
Patent Text Reader

Abstract

The invention relates to a Tibetan Thangka image restoration method and system based on a GAN network. The method comprises the following steps: constructing a Tibetan Thangka image sample data set; inputting the Tibetan Thangka image sample data set into a generator of the OA-GAN network; in the generator, mineral pigment features are extracted through a self-adaptive channel attention mechanism, and line features are extracted through a dynamic space attention mechanism; generating a Tibetan Thangka restoration image according to the extracted mineral pigment features and the extracted line features; inputting the Tibetan thangka restored image into a discriminator of the OA-GAN network; in the discriminator, constructing a mixed loss combination function through an L1 reconstruction loss function, a color smooth loss function, a perception loss function and an adversarial loss function, and discriminating the Tibetan Thangka restored image through the mixed loss combination function; and confrontation training is carried out according to a discrimination result of the discriminator, a trained image restoration model is obtained, and the image restoration model is used for outputting a restored Tibetan Thangka image when the to-be-restored Tibetan Thangka image is input.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a Tibetan thangka image restoration method and system based on a GAN network. Background Art

[0002] Tibetan Thangka, also known as "Tangga" and "Tangka", is a unique painting art form in Tibetan culture. Its subject matter covers many areas such as Tibetan history, politics, culture and social life, and has distinct national characteristics.

[0003] With the advancement of deep learning technology, color transfer methods based on generative adversarial networks (GANs) have made significant breakthroughs in the field of image restoration, especially in grayscale image colorization and historical image restoration, demonstrating excellent color reconstruction capabilities. However, when it comes to the unique art form of Tibetan thangka, existing technologies face severe challenges:

[0004] (1) Compared with conventional images, thangka images have dense and interlaced line features, and the direction of the lines is strongly associated with religious semantics. Traditional convolutional neural networks are prone to color correspondence errors in the intersection area due to their local perception characteristics.

[0005] (2) Thangkas are rendered using natural mineral pigments (such as cinnabar and azurite) to create a unique gradient effect.

[0006] In summary, existing image restoration algorithms rely on manually designed feature extraction rules, resulting in a rigid restoration model that is unable to adapt to degradation characteristics at varying degrees of aging. These algorithms employ single-channel grayscale processing, resulting in images that lose the unique mineral pigment characteristics of thangkas, leading to problems such as color scale discontinuity and semantic distortion. In recent years, deep learning models have resulted in blurred details and edge artifacts when restoring fine textures in thangka images (such as gold thread outlines and folds in Buddha clothing). Summary of the Invention

[0007] In order to at least to some extent overcome the problems that the image restoration algorithms in the related art may cause the restored images to lose the unique mineral pigment characteristics of thangkas, and have color scale breaks, semantic distortion, and other problems when performing Tibetan thangka image restoration, the present application provides a Tibetan thangka image restoration method and system based on a GAN network.

[0008] The scheme of this application is as follows:

[0009] According to a first aspect of an embodiment of the present application, a method for restoring a Tibetan thangka image based on a GAN network is provided, comprising:

[0010] Construct a Tibetan thangka image sample dataset;

[0011] Input the Tibetan thangka image sample dataset into the generator of the OA-GAN (Object-Attention Generative Adversarial Network) network;

[0012] In the generator, mineral pigment features are extracted through an adaptive channel attention mechanism, and texture features are extracted through a dynamic spatial attention mechanism;

[0013] Generate Tibetan thangka restoration images based on the extracted mineral pigment features and texture features;

[0014] Inputting the Tibetan thangka restoration image into the discriminator of the OA-GAN network;

[0015] In the discriminator, a hybrid loss combination function is constructed by using an L1 reconstruction loss function, a color smoothing loss function, a perceptual loss function, and an adversarial loss function, and the Tibetan thangka restoration image is discriminated by using the hybrid loss combination function;

[0016] Adversarial training is performed based on the discrimination results of the discriminator to obtain a trained image restoration model. The image restoration model is used to output a restored Tibetan thangka image when a Tibetan thangka image to be restored is input.

[0017] Preferably, constructing a Tibetan thangka image sample dataset includes:

[0018] Obtaining a photographic image of Tibetan thangka;

[0019] Adjusting the size of the Tibetan thangka image to a uniform size;

[0020] Performing degradation processing on the Tibetan thangka image;

[0021] The degraded images and original images of Tibetan thangka images are used as Tibetan thangka image sample datasets.

[0022] Preferably, the degradation processing of the Tibetan thangka image comprises:

[0023] The Tibetan thangka image is subjected to degradation processing in different degrees.

[0024] Preferably, in the generator, mineral pigment features are extracted by an adaptive channel attention mechanism, including:

[0025] Perform global average pooling on the input feature map:

[0026]

[0027] Among them, X cRepresents the feature map of the cth channel, X∈R B ×C×H×W;z c represents the global average pooling result of the c-th channel; B represents the batch size, which is set to 8; H represents the image height; W represents the image width;

[0028] Adaptive dimensionality reduction is performed on the global average pooling result:

[0029]

[0030] in, Represents the result of adaptive dimensionality reduction; z represents the global average pooling result; W1 represents the weight matrix of the first fully connected layer, which is used to achieve channel dimension compression; W2 represents the weight matrix of the second fully connected layer, which is used to achieve channel dimension recovery; ReLU is the ReLU activation function, which is used to introduce nonlinear expression capabilities;

[0031] Generate channel weights for adaptive dimensionality reduction results:

[0032]

[0033] Among them, α c is the attention weight of the c-th channel; represents the adaptive dimensionality reduction result of the c-th channel; σ represents the Sigmiod function;

[0034] Recalibrate the channel weights of each channel:

[0035] X out =X⊙(α1, α2, ..., α c );

[0036] Among them, ⊙ represents element-by-element multiplication, which is used to realize the dynamic weighting of channel features; X out Represents the feature map after feature recalibration.

[0037] Preferably, in the generator, texture features are extracted by a dynamic spatial attention mechanism, including:

[0038] Perform cross-channel aggregation on the input feature map to obtain the cross-channel maximum feature map and the cross-channel average feature map:

[0039]

[0040] Among them, X :c represents the c-th channel feature map of all samples in the Tibetan Thangka image sample dataset, M max Represents the cross-channel maximum feature map, which is used to highlight the texture structure; M avg Represents the cross-channel average feature map, which is used to preserve background gradient information;

[0041] Generate spatial weights based on the cross-channel maximum feature map and the cross-channel average feature map:

[0042]

[0043] Among them, β represents the spatial attention weight map, Represents the spatial filtering operation of the 7×7 convolution kernel;

[0044] Perform feature enhancement on the feature map after feature recalibration:

[0045] X out =X⊙β.

[0046] Preferably, the method further comprises:

[0047] The adaptive channel attention mechanism and the dynamic spatial attention mechanism are integrated into the upsampling module of the OA-GAN network.

[0048] Preferably, the method further comprises:

[0049] The adaptive channel attention mechanism and the dynamic spatial attention mechanism are integrated in the bottleneck layer of the OA-GAN network.

[0050] Preferably, in the discriminator, a hybrid loss combination function is constructed by using an L1 reconstruction loss function, a color smoothing loss function, a perceptual loss function, and an adversarial loss function, including:

[0051] Construct the L1 reconstruction loss function:

[0052]

[0053] Among them, X fake represents the Tibetan thangka restoration image generated by the generator; X real represents the real image in the Tibetan Thangka image sample dataset; L L1 represents the L1 reconstruction loss function;

[0054] The cross-channel penalty is determined based on the magnitude of the horizontal gradient map and the magnitude of the vertical gradient map:

[0055] P cross =0.3·max(|G X ∣,∣G y ∣);

[0056] Among them, P cross represents the value of cross-channel penalty; G X Represents the amplitude of the horizontal gradient map; G y Indicates the amplitude of the vertical gradient image; max(·) means taking the maximum value bit by bit;

[0057] Determine the saturation penalty based on the pixel value:

[0058] P sat =e 5|F|-4.5

[0059] Where F represents the pixel value;

[0060] Construct a color smoothing loss function based on cross-channel penalty and saturation penalty:

[0061] L color =0.5(G+P cross )+0.1P sat ;

[0062] Among them, L color represents the color smoothing loss function;

[0063] Construct the perceptual loss function:

[0064]

[0065] where φ j represents the feature extraction function of the jth layer of the VGG19 network; C j H j W j Respectively represent the number of channels, height and width of the feature map of the jth layer; L perc represents the color smoothing loss function;

[0066] Constructing adversarial loss function:

[0067]

[0068] Among them, B represents the batch size; D represents the discriminator’s discrimination result; L GAN represents the adversarial loss function;

[0069] Construct a hybrid loss combination function based on the color smoothing loss function, the perceptual loss function, and the adversarial loss function:

[0070] L total =100L L1 +0.1L color +0.5L perc +L GAN ;

[0071] Among them, L total represents the hybrid loss combination function.

[0072] Preferably, the method further comprises:

[0073] Perform resolution enhancement on the restored Tibetan thangka images.

[0074] According to a second aspect of an embodiment of the present application, a Tibetan thangka image restoration system based on a GAN network is provided, comprising:

[0075] processor and memory;

[0076] The processor and the memory are connected via a communication bus:

[0077] The processor is configured to call and execute the program stored in the memory;

[0078] The memory is used to store a program, and the program is at least used to execute a Tibetan thangka image restoration method based on a GAN network as described in any one of the above items.

[0079] The technical solution provided by this application may have the following beneficial effects:

[0080] In this technical solution, the image restoration model is trained through the OA-GAN network. The generator of the OA-GAN network extracts mineral pigment features through an adaptive channel attention mechanism, and extracts texture features through a dynamic spatial attention mechanism. This allows the trained image restoration model to handle different degrees of degradation features, retain the spectral characteristics of natural mineral pigments, avoid the color level breakage problem caused by traditional grayscale processing, and enhance the details of the gold thread and folds of Buddhist clothes. In addition, the discriminator of the OA-GAN network constructs a hybrid loss combination function through the L1 reconstruction loss function, color smoothing loss function, perceptual loss function, and adversarial loss function to identify the restored Tibetan thangka images. This allows the restored Tibetan thangka images to have the semantic features of real Tibetan thangka images, and to be more consistent with the colors of Tibetan thangka images and have a natural gradient.

[0081] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0083] Figure 1 This is a flow chart of a method for restoring Tibetan thangka images based on a GAN network, provided in one embodiment of the present application;

[0084] Figure 2 This is a comparison chart of the restoration results of various mainstream image restoration frameworks provided by an embodiment of the present application;

[0085] Figure 3 This is a framework diagram of an OA-GAN network provided by an embodiment of the present application;

[0086] Figure 4 This is a schematic diagram of a process for constructing a Tibetan thangka image sample dataset in a Tibetan thangka image restoration method based on a GAN network provided by an embodiment of the present application;

[0087] Figure 5 This is a comparison diagram of the Tibetan thangka image restored by the technical solution provided in one embodiment of the present application and the original image;

[0088] Figure 6 This is a structural diagram of a Tibetan thangka image restoration system based on a GAN network provided in one embodiment of the present application.

[0089] Reference numerals: processor-21; memory-22. DETAILED DESCRIPTION

[0090] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0091] Example 1

[0092] Figure 1 This is a flow chart of a Tibetan thangka image restoration method based on a GAN network provided by an embodiment of the present application, with reference to Figure 1 , a Tibetan thangka image restoration method based on GAN network, including:

[0093] S11: Construct a Tibetan thangka image sample dataset;

[0094] S12: Input the Tibetan thangka image sample dataset into the generator of the OA-GAN network;

[0095] S13: In the generator, mineral pigment features are extracted through adaptive channel attention mechanism, and texture features are extracted through dynamic spatial attention mechanism;

[0096] S14: Generate Tibetan thangka restoration images based on the extracted mineral pigment features and texture features;

[0097] S15: Input the Tibetan thangka restoration image into the discriminator of the OA-GAN network;

[0098] S16: In the discriminator, a hybrid loss combination function is constructed by using the L1 reconstruction loss function, the color smoothing loss function, the perceptual loss function, and the adversarial loss function. The hybrid loss combination function is used to discriminate the Tibetan thangka restoration image.

[0099] S17: Perform adversarial training based on the discrimination result of the discriminator to obtain a trained image restoration model. The image restoration model is used to output a restored Tibetan thangka image when a Tibetan thangka image to be restored is input.

[0100] It should be noted that the OA-GAN network has the ability to restore old art images. In order to evaluate its performance in the task of restoring old art images, Figure 2 Show some examples of Tibetan thangka image restoration. In this example, three mainstream frameworks are selected for comparison: DeOldify, CAP-VSTNET and Cyclegan, and the restoration results of the OA-GAN network are placed on Figure 2 The far right side is for comparison.

[0101] Figure 2 The visualization example of Tibetan thangka images in the video includes three categories: flowers, mountains, and clouds. It can be seen that the thangka background color semantics correspond correctly and the gradient color of the flowers is correct, which helps to intuitively understand the restoration of old thangka images.

[0102] The examples presented focused on color matching in thangkas and the coloring of flowers with gradients. DeOldify's restoration results were poor, with inconsistent color matching, resulting in a poor overall effect. CAP-VSTNET's overall restoration showed some discrepancies between the colors and the reference image. While the semantic correspondence was relatively accurate, the pattern background was noisy. Furthermore, CAP-VSTNET failed to achieve a gradient color effect in the flower coloring. CycleGAN achieved significantly better restoration results than the previous two methods, but some areas showed significant color bleeding and a chaotic color palette.

[0103] In contrast, the OA-GAN network performs better in these aspects and can better handle color gradients and semantic correspondence of details, thus providing better coloring effects overall.

[0104] The architecture of the OA-GAN network in this embodiment is as follows Figure 3 As shown in the figure, the framework adopts a two-stage restoration process, which combines generative adversarial networks with super-resolution reconstruction to achieve accurate restoration of old thangka images.

[0105] In this technical solution, the image restoration model is trained through the OA-GAN network. The generator of the OA-GAN network extracts mineral pigment features through an adaptive channel attention mechanism, and extracts texture features through a dynamic spatial attention mechanism. This allows the trained image restoration model to handle different degrees of degradation features, retain the spectral characteristics of natural mineral pigments, avoid the color level breakage problem caused by traditional grayscale processing, and enhance the details of the gold thread and folds of Buddhist clothes. In addition, the discriminator of the OA-GAN network constructs a hybrid loss combination function through the L1 reconstruction loss function, color smoothing loss function, perceptual loss function, and adversarial loss function to identify the restored Tibetan thangka images. This allows the restored Tibetan thangka images to have the semantic features of real Tibetan thangka images, and to be more consistent with the colors of Tibetan thangka images and have a natural gradient.

[0106] Example 2

[0107] Reference Figure 4 , construct a Tibetan thangka image sample dataset, including:

[0108] S21: Acquire a Tibetan thangka image;

[0109] S22: resizing the Tibetan thangka images;

[0110] S23: performing degradation processing on the Tibetan thangka image;

[0111] S24: The degraded image and the original image of the Tibetan thangka image camera image are used as a Tibetan thangka image sample dataset.

[0112] It should be noted that thangkas are renowned for their fine lines and rich colors. Due to the complex and time-consuming thangka painting process, high-quality thangka works are limited. To address this issue, this embodiment employs an on-site photography and collection method to obtain more thangka image resources. After filming, each thangka image was uniformly resized and cropped into a 256×256 pixel square image to facilitate subsequent computer vision processing and deep learning training. This size ensures that image details are not overly compressed while also meeting the input requirements of most existing deep learning models. The collected thangka images were subsequently subjected to varying degrees of degradation, resulting in a total of 1,576 degraded and color images. These images will be used to construct a specialized thangka dataset for training and testing deep learning models, enabling them to automatically identify and analyze the artistic characteristics of thangkas, thereby promoting the digital preservation and research of thangka art.

[0113] It should be noted that the Tibetan thangka images were subjected to varying degrees of degradation, artificially creating image degradation (such as blur, noise, and low resolution) to simulate low-quality images captured in real-world environments. Training the model to recognize these degraded images helps enhance its ability to analyze and identify thangkas in a variety of real-world conditions, such as low light, poor camera equipment, and aged and damaged thangkas. By training the model on images of varying quality levels, it can better capture the essential and representative artistic features of thangkas, rather than relying solely on high-quality texture or color information.

[0114] Example 3

[0115] In this embodiment, an adaptive channel attention module is configured for the generator, which focuses on the channel characteristics of mineral pigments to solve the problem of mixing multiple colors in thangka.

[0116] In the generator, mineral pigment features are extracted through an adaptive channel attention mechanism, including:

[0117] Perform global average pooling on the input feature map:

[0118]

[0119] Among them, X c Represents the feature map of the cth channel, X∈R B ×C×H×W;z c represents the global average pooling result of the c-th channel; B represents the batch size, which is set to 8; H represents the image height; W represents the image width;

[0120] Adaptive dimensionality reduction is performed on the global average pooling result:

[0121]

[0122] in, Represents the result of adaptive dimensionality reduction; z represents the global average pooling result; W1 represents the weight matrix of the first fully connected layer, which is used to achieve channel dimension compression; W2 represents the weight matrix of the second fully connected layer, which is used to achieve channel dimension recovery; ReLU is the ReLU activation function, which is used to introduce nonlinear expression capabilities;

[0123] Generate channel weights for adaptive dimensionality reduction results:

[0124]

[0125] Among them, α c is the attention weight of the c-th channel; represents the adaptive dimensionality reduction result of the c-th channel; σ represents the Sigmiod function;

[0126] Recalibrate the channel weights of each channel:

[0127] X out =X⊙(α1,α2,...,α c );

[0128] Among them, ⊙ represents element-by-element multiplication, which is used to realize the dynamic weighting of channel features; X out Represents the feature map after feature recalibration.

[0129] In this embodiment, a dynamic spatial attention module is also configured for the generator, which is used to better adapt to the texture features such as clothing folds and gold lines in thangka images.

[0130] In the generator, texture features are extracted through a dynamic spatial attention mechanism, including:

[0131] Perform cross-channel aggregation on the input feature map to obtain the cross-channel maximum feature map and the cross-channel average feature map:

[0132]

[0133]

[0134] Among them, X :c represents the c-th channel feature map of all samples in the Tibetan Thangka image sample dataset, M max Represents the cross-channel maximum feature map, which is used to highlight the texture structure; M avg Represents the cross-channel average feature map, which is used to preserve background gradient information;

[0135] Generate spatial weights based on the cross-channel maximum feature map and the cross-channel average feature map:

[0136]

[0137] Among them, β represents the spatial attention weight map, with a value of [0,1]; Represents the spatial filtering operation of the 7×7 convolution kernel;

[0138] Perform feature enhancement on the feature map after feature recalibration:

[0139] X out =X⊙β.

[0140] It should be noted that the method also includes:

[0141] Integrate adaptive channel attention mechanism and dynamic spatial attention mechanism into the upsampling module of OA-GAN network;

[0142] The adaptive channel attention mechanism and dynamic spatial attention mechanism are integrated in the bottleneck layer of the OA-GAN network.

[0143] The expression for integrating the adaptive channel attention mechanism and the dynamic spatial attention mechanism in the upsampling module of the OA-GAN network is:

[0144] F out =A spatial (ReLU(IN(A channel (Γ(F in )))));

[0145] Among them, F in Represents the input feature tensor, which comes from the output of the previous layer; Γ represents the operator representing transposed convolution upsampling; A spatial Indicates the channel attention module; IN indicates the operator represents instance normalization to eliminate differences between batches; ReLU indicates the nonlinear activation function to enhance the model's expressiveness; A channel represents the spatial attention module; F out Represents the output feature tensor with half the number of channels.

[0146] In order to solve the cross-region color correlation problem, the adaptive channel attention mechanism and dynamic spatial attention mechanism are also integrated in the bottleneck layer of the OA-GAN network, and their expressions are as follows:

[0147] F bottleneck =A spatial (A channel (F encoder ));

[0148] Among them, F encoder Represents the high-level semantic features output by the encoder, F bottleneck Represents the enhanced bottleneck feature, including the global semantic-color mapping relationship.

[0149] Example 4

[0150] To ensure that the thangka images generated by the generator more closely match the characteristics of the original thangka images, this embodiment designs a discriminator to constrain the generated graph. The discriminator uses a hybrid loss combination function, which primarily ensures a natural gradient of mineral pigments and prevents color deviations caused by simultaneous mutations in multiple channels, or overflow.

[0151] In the discriminator, a hybrid loss combination function is constructed through the L1 reconstruction loss function, color smoothing loss function, perceptual loss function and adversarial loss function, including:

[0152] Construct the L1 reconstruction loss function:

[0153]

[0154] Among them, X fake represents the Tibetan thangka restoration image generated by the generator; X real represents the real image in the Tibetan Thangka image sample dataset; L :1 represents the L1 reconstruction loss function; ||X fake -X real ||1 represents the L1 norm (sum of absolute values), which calculates the absolute difference of each pixel;

[0155] The cross-channel penalty is determined based on the magnitude of the horizontal gradient map and the magnitude of the vertical gradient map:

[0156] P cross =0.3·max(|G X ∣,∣G y ∣);

[0157] Among them, P cross represents the value of cross-channel penalty; G X Represents the amplitude of the horizontal gradient map; G y Indicates the amplitude of the vertical gradient image; max(·) means taking the maximum value bit by bit;

[0158] Determine the saturation penalty based on the pixel value:

[0159] P sat =e 5|F|-4.5

[0160] Where F represents the pixel value;

[0161] Construct a color smoothing loss function based on cross-channel penalty and saturation penalty:

[0162] L color =0.5(G+P cross )+0.1P sat ;

[0163] Among them, L color represents the color smoothing loss function;

[0164] Construct the perceptual loss function:

[0165]

[0166] where φ j represents the feature extraction function of the jth layer of the VGG19 network; C j H j W j Respectively represent the number of channels, height and width of the feature map of the jth layer; L perc represents the color smoothing loss function;

[0167] Constructing adversarial loss function:

[0168]

[0169] Among them, B represents the batch size; D represents the discriminator’s judgment result, judging whether the input is a real image or a generated image, 1 for correct and 0 for wrong; L GAN represents the adversarial loss function;

[0170] Construct a hybrid loss combination function based on the color smoothing loss function, the perceptual loss function, and the adversarial loss function:

[0171] L total =100L L1 +0.1L color +0.5L perc +L GAN ;

[0172] Among them, L total represents the hybrid loss combination function.

[0173] In order to verify the effectiveness of each module, this embodiment provides an ablation experiment example, which mainly tests the image restoration results of adding color-aware attention mechanism (ColorAwareAttention), adding color-aware attention mechanism and gradient-aware block (GradientAwareBlock), and adding all modules. Figure 5 As shown in the figure, when only color-aware attention is used for inpainting, color overflow is severe and the color gradient is poorly rendered. When the color-aware attention mechanism and gradient-aware blocks are added, the color content in the flower is not smooth and color overflow is severe. In summary, the image inpainting is the smoothest and most uniform among all modules, proving the effectiveness of this technical solution.

[0174] Example 5

[0175] The method also includes:

[0176] Perform resolution enhancement on the restored Tibetan thangka images.

[0177] In this embodiment, a super-resolution module is connected to the image generated by the generator, and the final repair image is output through this module, so that the image resolution is increased and the generated repaired image has higher quality.

[0178] Example 6

[0179] Figure 6 This is a schematic diagram of the structure of a Tibetan thangka image restoration system based on a GAN network provided by an embodiment of the present application, referring to Figure 6 , a Tibetan thangka image restoration system based on GAN network, including:

[0180] processing,31 and memory,32;

[0181] The processor 31 and the memory 32 are connected via a communication bus:

[0182] The processor 31 is used to call and execute the program stored in the memory 32;

[0183] The memory 32 is used to store a program, and the program is used to at least execute a Tibetan thangka image restoration method based on a GAN network in the above embodiment.

[0184] It can be understood that the same or similar parts of the above embodiments can be referenced to each other, and the contents not described in detail in some embodiments can refer to the same or similar contents in other embodiments.

[0185] It should be noted that, in the description of this application, the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, the meaning of "plurality" refers to at least two.

[0186] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.

[0187] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0188] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0189] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0190] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.

[0191] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present application. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0192] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. A Tibetan thangka image restoration method based on GAN network, characterized in that: include: Construct a Tibetan thangka image sample dataset; Input the Tibetan thangka image sample dataset into the generator of the OA-GAN network; In the generator, mineral pigment features are extracted through an adaptive channel attention mechanism, and texture features are extracted through a dynamic spatial attention mechanism; Generate Tibetan thangka restoration images based on the extracted mineral pigment features and texture features; Inputting the Tibetan thangka restoration image into the discriminator of the OA-GAN network; In the discriminator, a hybrid loss combination function is constructed by using an L1 reconstruction loss function, a color smoothing loss function, a perceptual loss function, and an adversarial loss function, and the Tibetan thangka restoration image is discriminated by using the hybrid loss combination function; Adversarial training is performed based on the discrimination results of the discriminator to obtain a trained image restoration model. The image restoration model is used to output a restored Tibetan thangka image when a Tibetan thangka image to be restored is input.

2. The method according to claim 1, characterized in that Construct a Tibetan thangka image sample dataset, including: Obtaining a photographic image of Tibetan thangka; Adjusting the size of the Tibetan thangka image to a uniform size; Performing degradation processing on the Tibetan thangka image; The degraded images and original images of Tibetan thangka images are used as Tibetan thangka image sample datasets.

3. The method according to claim 1, characterized in that Degradation processing is performed on the Tibetan thangka image, including: The Tibetan thangka image is subjected to degradation processing in different degrees.

4. The method according to claim 1, wherein In the generator, mineral pigment features are extracted through an adaptive channel attention mechanism, including: Perform global average pooling on the input feature map: Among them, X c Represents the feature map of the c-th channel, X∈R B×C×H×W ;z c represents the global average pooling result of the c-th channel; B represents the batch size, which is set to 8; H represents the image height; W represents the image width; Adaptive dimensionality reduction is performed on the global average pooling result: in, Represents the result of adaptive dimensionality reduction; z represents the global average pooling result; W1 represents the weight matrix of the first fully connected layer, which is used to achieve channel dimension compression; W2 represents the weight matrix of the second fully connected layer, which is used to achieve channel dimension recovery; ReLU is the ReLU activation function, which is used to introduce nonlinear expression capabilities; Generate channel weights for adaptive dimensionality reduction results: Among them, α c is the attention weight of the c-th channel; represents the adaptive dimensionality reduction result of the c-th channel; σ represents the Sigmiod function; Recalibrate the channel weights of each channel: X out =X⊙(α1,α2,...,α c ); Among them, ⊙ represents element-by-element multiplication, which is used to realize the dynamic weighting of channel features; X out Represents the feature map after feature recalibration.

5. The method according to claim 4, characterized in that In the generator, texture features are extracted through a dynamic spatial attention mechanism, including: Perform cross-channel aggregation on the input feature map to obtain the cross-channel maximum feature map and the cross-channel average feature map: Among them, X :c represents the c-th channel feature map of all samples in the Tibetan Thangka image sample dataset, M max Represents the cross-channel maximum feature map, which is used to highlight the texture structure; M avg Represents the cross-channel average feature map, which is used to preserve background gradient information; Generate spatial weights based on the cross-channel maximum feature map and the cross-channel average feature map: Among them, β represents the spatial attention weight map, Represents the spatial filtering operation of the 7×7 convolution kernel; Perform feature enhancement on the feature map after feature recalibration: X out =X⊙β。 6. The method according to claim 5, characterized in that The method further comprises: The adaptive channel attention mechanism and the dynamic spatial attention mechanism are integrated into the upsampling module of the OA-GAN network.

7. The method according to claim 5, characterized in that The method further comprises: The adaptive channel attention mechanism and the dynamic spatial attention mechanism are integrated in the bottleneck layer of the OA-GAN network.

8. The method according to claim 1, characterized in that In the discriminator, a hybrid loss combination function is constructed by using the L1 reconstruction loss function, the color smoothness loss function, the perceptual loss function, and the adversarial loss function, including: Construct the L1 reconstruction loss function: Among them, X fake represents the Tibetan thangka restoration image generated by the generator; X real represents the real image in the Tibetan Thangka image sample dataset; L L1 represents the L1 reconstruction loss function; The cross-channel penalty is determined based on the magnitude of the horizontal gradient map and the magnitude of the vertical gradient map: P cross =0.3·max(∣G X ∣,∣G y ∣); Among them, P cross represents the value of cross-channel penalty; G X Represents the amplitude of the horizontal gradient map; G y Indicates the amplitude of the vertical gradient image; max(·) means taking the maximum value bit by bit; Determine the saturation penalty based on the pixel value: P sat =and 5|F|-4.5 Where F represents the pixel value; Construct a color smoothing loss function based on cross-channel penalty and saturation penalty: L color =0.5(G+P cross )+0.1P sat ; Among them, L color represents the color smoothing loss function; Construct the perceptual loss function: where φ j represents the feature extraction function of the jth layer of the VGG19 network; C j H j W j Respectively represent the number of channels, height and width of the feature map of the jth layer; L perc represents the color smoothing loss function; Constructing adversarial loss function: Among them, B represents the batch size; D represents the discriminator’s discrimination result; L GAN represents the adversarial loss function; Construct a hybrid loss combination function based on the color smoothing loss function, the perceptual loss function, and the adversarial loss function: L total =100L L1 +0.1L color +0.5L perc +L GAN ; Among them, L total represents the hybrid loss combination function.

9. The method according to claim 1, characterized in that The method further comprises: Perform resolution enhancement on the restored Tibetan thangka images.

10. A Tibetan thangka image restoration system based on GAN network, characterized in that: include: processor and memory; The processor and the memory are connected via a communication bus: The processor is configured to call and execute the program stored in the memory; The memory is used to store a program, and the program is used at least to execute the Tibetan thangka image restoration method based on the GAN network as described in any one of claims 1 to 9.

Citation Information

Cited By

  • Thangka generation method based on structure and pattern double-channel constraint diffusion model

    CN121639855A

  • A method for generating a Thangka based on a structure and pattern double-channel constrained diffusion model

    CN121639855B