A Thangka image restoration method, system and storage medium

By using the method of using the Encoder-decoder structure, vector quantization codebook and parallel CSWin resolution Transformer module based on the Thangka image repair, the problems of low accuracy of prior information and difficult to maintain structural integrity in Thangka image repair are solved, and high-quality image repair effects are achieved.

CN118710552BActive Publication Date: 2025-06-20TIBET UNIVERSITY FOR NATIONALITIES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410843431.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-27
Publication Date
2025-06-20
Estimated Expiration
2044-06-27

AI Technical Summary

Technical Problem

The prior art, when repairing Thangka images with rich textures and complex structures, the predicted prior information is low, resulting in repair errors and it is difficult to maintain the structural integrity and local details of the image.

Method used

The encoder-decoder structure based on the Transformer model is adopted, combined with vector quantized codebooks and parallel CSWin resolution Transformer modules, and efficient image repair is achieved through multi-scale feature guidance modules and adaptive learning rate adjustment.

Benefits of technology

It improves the accuracy and quality of Thangka image repair, maintains the structural integrity and local details of the image, and significantly improves the visual effect of the repair results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118710552B_ABST
    Figure CN118710552B_ABST
Patent Text Reader

Abstract

The present invention provides a Thangka image restoration method, system and storage medium, which relates to the technical field of digital image processing. By dividing the input image into non-overlapping sub-regions and performing non-linear transformation, local information is effectively isolated. At the same time, a vector quantization codebook is introduced to better capture and retain the image structure information and details. The parallel CSWin resolution Transformer module enhances the context modeling ability through the cross-shaped window and local enhanced position encoding; the novel multi-scale feature guidance module adaptively learns the feature information of the non-defective region through local knowledge of different scales; experiments of the CDCT model on multiple data sets show that its restoration results are competitive; by adopting SSIM, PSNR and the comprehensive quality evaluation index QI, a significant improvement in the restoration quality is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital image processing, and specifically to a Thangka image restoration method, system and storage medium. Background Art

[0002] In the field of digital image processing, image restoration technology has always been a research hotspot. Its purpose is to restore the damaged image content through computational methods to restore the integrity and beauty of the image. This field started at the end of the 20th century. In the early days, it mainly relied on simple interpolation algorithms such as nearest neighbor interpolation and bilinear interpolation. In the 21st century, with the rapid development of computer vision and machine learning, image restoration technology has been significantly improved. Especially with the rise of deep learning, image restoration methods based on neural networks, such as convolutional neural networks (CNNs) and generative adversarial networks (GANs), have gradually become mainstream. They can learn complex mappings of a large amount of data and achieve highly complex image restoration effects. However, traditional methods and early deep learning methods have limitations in dealing with images with rich textures and complex structures. Different from natural images, the costumes of the Buddha statues in Thangka or the patterns of flower vines, clouds, and landscapes in the background are complex, delicate, and rich in details. When facing such Thangka images, the accuracy of the prior information predicted by these methods will be greatly reduced, resulting in a large number of restoration errors when filling internal textures. When dealing with art images such as Thangka, there are still huge challenges.

[0003] The deficiencies of the prior art are mainly reflected in the limitations in restoring highly structured and detailed image content. Especially for Thangka images, they usually contain fine lines and complex patterns. Once these contents are damaged, it is difficult to achieve satisfactory restoration effects using simple texture replication or basic learning models. In addition, the prior art often ignores the global consistency of the image and the restoration of local details during the restoration process, resulting in the restored image not being seamlessly docked with the original image visually. More critically, most methods fail to effectively utilize the internal structural information of the image and cannot maintain the structural integrity during the image restoration process. This problem is particularly prominent when there is a large area of damage. Although the generative adversarial network (GAN) performs well in providing realistic textures for restored images, its training stability and mode collapse are still problems that need to be solved. In addition, the quality evaluation indicators in the prior art are mostly single-index evaluations, lacking a multi-dimensional comprehensive evaluation of image quality.

[0004] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0005] The purpose of the present invention is to provide a Thangka image restoration method, system and storage medium to solve the problems raised in the above background technology.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A Thangka image restoration method, the specific steps include:

[0008] Step S1, collect Thangka images containing damaged parts and missing parts, preprocess the collected Thangka images, and construct the preprocessed Thangka images into an image dataset;

[0009] Step S2, construct an encoder-decoder structure based on the Transformer model and use it to jointly learn a discrete codebook. Input the Thangka images inside the image dataset into the constructed encoder. Use the encoder to divide the input Thangka images into non-overlapping sub-regions of a fixed size, and map the sub-regions to a continuous latent space representation through a non-linear transformation to obtain feature vectors;

[0010] Step S3, introduce a vector quantization codebook, perform vector quantization on the continuous latent feature vectors output by the encoder. The discrete codebook is constructed using a clustering algorithm. Each vector in the codebook represents the latent space representation of a piece of sub-region in the image. Finally, the obtained discrete codebook is used as the codebook prior knowledge;

[0011] Step S4, construct a parallel CSWin resolution Transformer module. This module adopts a cross-shaped window and a local enhanced position encoding design. Use the feature vectors in Step S3 as the input, and add additional learnable position embeddings to the feature vectors to retain spatial information. Then flatten the feature vectors along the spatial dimension to obtain the final input of this module to predict the probability distribution of the next index;

[0012] Step S5, use the parallel CSWin resolution Transformer module in Step S4 to accurately infer the indices of the missing tokens, and find the corresponding discrete vectors from the discrete codebook obtained in Step S3 through these indices for image restoration. After completing one repair attempt, the system will enter an iterative loop;

[0013] Step S6, after each iteration, collect the generated restored images, obtain the structural similarity index between the restored image and the reference image, the peak signal-to-noise ratio PSNR between the restored image and the reference image, and analyze and process the structural similarity index and the peak signal-to-noise ratio PSNR to generate a comprehensive quality evaluation index QI. This index is used to evaluate the image quality and restoration effect and generate a corresponding quadratic adjustment strategy for the learning rate;

[0014] Step S7. On the basis of adaptive learning rate adjustment, design a multi-scale feature guidance module, which utilizes the features of the non-damaged area to promote the consistency of the generated area and the undamaged area in terms of structure and texture, thereby improving the quality and fidelity of the repair result;

[0015] Step S8. During the repair process, according to the comprehensive quality evaluation index QI generated in Step S6 and the learning rate adjustment strategy, dynamically adjust the learning rate of the model to optimize the repair effect;

[0016] Step S9. After all iterations are completed, post-process the finally repaired Thangka image, including but not limited to image enhancement, color correction, and detail optimization, to improve the quality and visual effect of the repaired image, and finally output the repaired Thangka image.

[0017] Furthermore, after each iteration, collect the generated repaired image, obtain the structural similarity index and peak signal-to-noise ratio PSNR between the repaired image and the reference image, and analyze and process the structural similarity index and peak signal-to-noise ratio PSNR to generate a comprehensive quality evaluation index QI, which is used to evaluate the image quality and repair effect and generate a corresponding secondary learning rate adjustment strategy;

[0018] After each iteration, collect the generated repaired image and calculate the following quality parameters:

[0019] The structural similarity index SSIM is as follows:

[0020]

[0021] where x and y are the local windows of the reference image and the repaired image respectively, μ x 、μ y are the means, is the variance, σ xy is the covariance, and c1, c2 are constants used to stabilize the calculation;

[0022] The peak signal-to-noise ratio PSNR is as follows:

[0023]

[0024] where MAX I is the maximum value of the image pixels, and MSE(x,y) is the mean square error;

[0025] Combining SSIM and PSNR, generate a comprehensive quality evaluation index QI, and the calculation formula is as follows:

[0026]

[0027] Parameter explanation: ω iis a weight factor used to balance the impact of different quality assessment indicators, where ω1, ω2, ω3, and ω4 correspond to the weights of SSIM, PSNR, FSIM, and NIQE, respectively;

[0028] f(Metric i″′ ) is a complex function, i″′∈{1, 2, 3, 4}, where Metric i″′ When i″′ takes the values ​​of 1, 2, 3, and 4, it represents SSIM, PSNR, FSIM, and NIQE, respectively, and is used to perform nonlinear transformation on each quality assessment indicator. The formula is as follows:

[0029] f(SSIM)=log(1+SSIM)

[0030] f(PSNR) = exp(-PSNR / 100)

[0031]

[0032] g(NIQE) is a normalization function used to adjust the impact of NIQE. x∈NIQE;

[0033] The value range of QI is set to (0,1). When QI is close to 1, it means that the image quality is close to the original image and the restoration effect is good; when QI is close to 0, it means that the image quality is poor and the restoration effect is not good.

[0034] Furthermore, when the QI value increases, it means that the image restoration quality is improved and the image is closer to the visual and structural features of the original image; conversely, a decrease in the QI value indicates that the restoration effect is poor and the model parameters or training strategy need to be adjusted;

[0035] According to the changes in these indicators, the learning rates of the generator and discriminator are dynamically adjusted. If the quality evaluation indicators improve slowly or decrease, the learning rate is increased to explore new parameter spaces; if the quality evaluation indicators improve steadily, the learning rate is maintained or moderately reduced to stabilize the training; specifically, the following are included:

[0036] According to the generated quality evaluation index QI, the learning rate is adjusted twice:

[0037] 1. t+1 =lr t ·(1+β5·(QIt-QItarget))

[0038] Among them, QIt is the quality evaluation index at the tth iteration, QItarget is the target quality evaluation index, and β5 is the adjustment factor used to control the impact of the quality evaluation index on the learning rate;

[0039] When the value of QIt is in Interval 1 (0, 0.3), the image quality is poor. It is necessary to increase the learning rate to explore new parameters and quickly improve the image restoration effect. The threshold is set to 0.2. When the value is lower than this, the learning rate is urgently increased in order to achieve significant improvement;

[0040] When the value of QIt is in Interval 2 [0.3, 0.7), there is room for improvement in the image quality. An adjustment strategy is adopted to maintain or slightly increase the learning rate to steadily improve the image quality. The threshold is 0.5 to maintain the stability of training and continuous improvement;

[0041] When the value of QIt is in Interval 3 [0.7, 1), it indicates that the image quality is close to ideal. In this interval, the learning rate is reduced to stabilize the training and prevent overfitting; the threshold is set to 0.85. When the value exceeds this, the learning rate is further reduced to ensure continuous optimization and stability of the quality.

[0042] A Thangka image restoration system, and the system is used to execute the described method.

[0043] A storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the Thangka image restoration method are realized.

[0044] Compared with the prior art, the beneficial effects of the present invention are:

[0045] 1. A novel codebook learning framework is designed. Among them, the encoder divides the input image into non-overlapping sub-regions (patches) of a fixed size, and then non-linearly transforms them into latent feature vectors, ensuring the effective isolation of local information; introducing a vector quantization codebook can more effectively capture and retain the structural information and details of the image, so as to produce more realistic results in the reconstruction process;

[0046] 2. A parallel CSWinTransformer module is designed. The design of the cross-shaped window and local enhanced position encoding strengthens the context modeling ability while reducing the computational cost and improving the accuracy of index prediction;

[0047] 3. A multi-scale feature guidance module is innovatively designed. Among them, LKAs of different scales utilize local information and channel adaptability to better learn the feature information of non-defective regions;

[0048] 4. The CDCT model has conducted a large number of experiments with existing advanced methods on Celeba-HQ, Places2, and the self-made Thangka dataset; qualitative and quantitative experiments show that the restoration results of the CDCT model are competitive; it has opened up a novel technical path for the restoration work of cultural heritages such as Thangka images;

[0049] 5. SSIM and PSNR are adopted as quality evaluation indicators, and the comprehensive quality evaluation index QI is introduced. Through the multi-dimensional evaluation of the image restoration results, the secondary adjustment of the learning rate is realized, further improving the restoration quality.

[0050] In summary, in the codebook learning stage, a network framework based on vector quantization codebook is designed and improved to discretize and encode the intermediate features of the input image, obtaining a discrete codebook with rich context; in the second stage, a parallel Transformer module based on a cross-shaped window is proposed, which can accurately predict the index combination of the missing area of the image at a limited computational cost; in addition, a multi-scale feature guidance module is proposed to gradually fuse the features of the intact area and the texture features in the codebook, so as to better retain the local details of the intact area. Brief Description of the Drawings

[0051] Figure 1 Schematic diagram of the overall method flow of the present invention;

[0052] Figure 2 Schematic diagram of the overall framework of the CDCT model of the present invention;

[0053] Figure 3 Schematic diagram of the parallel CSWin resolution Transformer module of the present invention;

[0054] Figure 4 Schematic diagram of the multi-scale feature guidance module of the present invention;

[0055] Figure 5 Schematic diagram of the change curve of the loss value of the present invention with the number of iterations;

[0056] Figure 6 Qualitative comparison results of the present invention on the Celeba-HQ dataset;

[0057] Figure 7 Qualitative comparison results of the present invention on the places2 dataset;

[0058] Figure 8 Qualitative comparison results of the present invention on the self-made Thangka dataset;

[0059] Figure 9 Schematic diagram of the visual effect analysis of each component of the model of the present invention. Detailed Description of the Invention

[0060] To make the purpose, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to specific embodiments.

[0061] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meanings understood by those of ordinary skill in the field to which the present invention pertains. The "first", "second" and similar terms used in the present invention do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Upper", "lower", "left", "right", etc. are only used to represent relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0062] Embodiment 1:

[0063] Please refer to Figures 1 to 9 , the present invention provides a technical solution:

[0064] A Thangka image restoration method, the specific steps include:

[0065] Step S1, collect Thangka images containing damaged parts and missing parts, and preprocess the collected Thangka images, and construct the preprocessed Thangka images into an image data set;

[0066] Step S2, construct an encoder-decoder structure based on the Transformer model, and use it to jointly learn a discrete codebook. Input the Thangka images inside the image data set into the constructed encoder. Use the encoder to divide the input Thangka images into non-overlapping sub-regions of a fixed size, and map the sub-regions to a continuous latent space representation through a non-linear transformation to obtain feature vectors;

[0067] Step S3, introduce a vector quantization codebook, perform vector quantization on the continuous latent feature vectors output by the encoder. The discrete codebook is constructed using a clustering algorithm. Each vector in the codebook represents the latent space representation of a piece of sub-region in the image. Finally, the obtained discrete codebook is used as the codebook prior knowledge;

[0068] Step S4, construct a parallel CSWin resolution Transformer module. This module adopts a cross-shaped window and a local enhanced position encoding (LePE) design. Take the feature vectors in Step S3 as the input, and add additional learnable position embeddings to the feature vectors to retain spatial information. Subsequently, flatten the feature vectors along the spatial dimension to obtain the final input of this module, so as to predict the probability distribution of the next possible index;

[0069] Step S5: Use the parallel CSWin resolution Transformer module in Step S4 to accurately infer the indices of the missing tokens, and find the corresponding discrete vectors from the discrete codebook obtained in Step S3 using these indices for image inpainting. After completing one inpainting attempt, the system will enter an iterative loop;

[0070] Step S6: After each iteration, collect the generated inpainted image, and obtain the structural similarity index between the inpainted image and the reference image, the peak signal-to-noise ratio PSNR between the inpainted image and the reference image, and analyze and process the structural similarity index and the peak signal-to-noise ratio PSNR to generate a comprehensive quality evaluation index QI. This index is used to evaluate the image quality and inpainting effect, and generate a corresponding secondary learning rate adjustment strategy;

[0071] Step S7: On the basis of adaptive learning rate adjustment, design a multi-scale feature guidance module. This module utilizes the features of the non-damaged area to promote the consistency of the generated area and the undamaged area in terms of structure and texture, thereby improving the quality and fidelity of the inpainting result;

[0072] Step S8: During the inpainting process, dynamically adjust the learning rate of the model according to the comprehensive quality evaluation index QI and the learning rate adjustment strategy generated in Step S6 to optimize the inpainting effect;

[0073] Step S9: After completing all iterations, perform post-processing on the finally inpainted Thangka image, including but not limited to image enhancement, color correction, and detail optimization, to improve the quality and visual effect of the inpainted image, and finally output the inpainted Thangka image.

[0074] Example 2:

[0075] Based on Example 1, further explain that in the shared codebook learning stage, the system architecture includes three core components: a codebook encoder E, a codebook decoder G, and a codebook containing K discrete encodings This is a set containing K discrete encodings, and each c k is a codeword, representing a specific feature or pattern;

[0076] When processing the input image I, first the codebook encoder E converts the image I t into a latent representation Z in a high-dimensional space, representing an image with a height H, a width W, and 3 color channels;

[0077] That is where d represents the number of dimensions that make up this latent vector; m×n is the spatial resolution in the latent representation;

[0078] Subsequently, an element-wise quantization operation q(·) is adopted, which is a function for quantizing each vector in the latent representation Z to the closest codeword in the codebook C. This operation is performed element-wise, that is, quantization is performed separately for each element (i,j) in Z;

[0079] Quantize each vector in this spatial latent representation Z to the closest codeword c in the codebook k , thereby obtaining the output Z of vector quantization c and the corresponding code token sequence s ∈ {0, …, N-1} m′·n′ ::

[0080]

[0081] where each element is the closest codeword found for Z (i,j) in the codebook C; the quantization operation is implemented by calculating the distance ||Z (i,j) -c k || and selecting the codeword corresponding to the minimum distance;

[0082] Subsequently, the decoder G reconstructs a high-quality image I c given Z rec ; the formed m′·n′ code token sequence s represents a new latent discrete representation, clearly indicating the codeword indices at various positions in the learned codebook, that is, when s (i,j) =k, the overall reconstruction I rec ≈I t is formulated as:

[0083] I rec = G(Z c ) = G(q(E(I)))(2)

[0084] The encoder performs a mapping operation to convert image data of size H×W into a discrete coding form of scale H / m×W / n, where the parameters m,n identify the downsampling ratio;

[0085] This process is essentially to condense the information in each m×n region within the image I t into a single coding unit. Therefore, when referring to any coding element in Z c , m×n also symbolizes the corresponding coverage range of this coding in the original image I t in space;

[0086] The codebook and the model are trained end-to-end through the reconstruction loss. Four image-level reconstruction losses are adopted in this application, and the reconstruction losses include the L1 loss L1, the perceptual loss L per , and the adversarial loss L advThe harmony style loss L style ;

[0087] The specific loss function is defined as follows:

[0088]

[0089] Among them, Φ refers to the feature extractor in the VGG19 network. Since there is insufficient loss constraint at the image level when updating the codebook items,

[0090] I t represents the target image, that is, it is hoped that the image generated by the model can be close to the reference image;

[0091] I rec represents the reconstructed image, ‖I t -I rec ‖1 represents the L1 distance between the target image and the reconstructed image, that is, the sum of the absolute values of the differences between the corresponding pixel values of the two images;

[0092] Φ refers to the feature extractor in the VGG19 network;

[0093] Φ(I t ) and Φ(I rec ) respectively represent the feature representations of the target image and the reconstructed image extracted using the pre-trained deep neural network;

[0094] represents the square of the Euclidean distance between these two sets of feature representations, which is used to measure the perceptual difference between the two images;

[0095] D(I t ) and D(I rec ) respectively represent the discrimination results of the discriminator on the target image and the reconstructed image, where D is the discriminator network;

[0096] logD(I t ) and log(1 - D(I rec )) respectively represent the logarithms of the probabilities of the discriminator correctly identifying the target image and misidentifying the reconstructed image;

[0097] and respectively represent the Gram matrices of the target image and the reconstructed image on the k-th feature channel, where represents the parameter of the k-th feature channel;

[0098] represents the expected value of the L1 distance between the Gram matrices on all feature channels, which is used to measure the style difference between the two images, where M k is the number of elements in the k-th feature channel;

[0099] Therefore, this application also adopts the loss L at the intermediate code level quantize to reduce the difference between the codebook C and the embedded input feature Z;

[0100]

[0101] where sg(·) refers to the stop gradient operator, and the parameter β is set to 0.25 to balance the weights between the encoder and the codebook update speed;

[0102] Regarding the non-differentiability of the feature quantization process shown in formula (1), a direct transfer strategy is adopted, that is, during backpropagation, the gradient is mirrored from the decoding link to the encoding link to ensure the implementation of backpropagation;

[0103] To comprehensively guide the prior knowledge learning of the codebook, the comprehensive loss function L codebook is used as the optimization objective to drive the entire end-to-end training process;

[0104] L codebook = L1 + L per + L quantize + λ adv ·L adv + L style (5)

[0105] where, in the experiments of this application, λ adv is set to 0.8;

[0106] Although more codebook terms can simplify the reconstruction, redundant elements will cause ambiguity in subsequent code prediction. Therefore, in the CDCT method of this application, the number of terms N of the codebook is set to 1024, which is sufficient to achieve accurate image reconstruction; in addition, the codebook dimension d is set to 256.

[0107] Embodiment 3:

[0108] Based on Embodiment 2, it is further described that the specific design points of the codebook encoder E specifically include:

[0109] Traditional CNN-based encoders use several convolutional kernels to process the input image in a sliding window manner, which is not suitable for image inpainting because they introduce interference between the masked area and the unmasked area. Therefore, the encoder in the shared codebook learning stage is designed to process the input image in a non-overlapping patch manner and through multiple linear residual layers;

[0110] Specifically, the Token representation is extracted in 8 blocks using a linear residual structure. In each block, it includes two sets of GELU activation functions, linear layers, and residual connections;

[0111] First, perform an unfold operation on the input image to change the size to (3×m×n, L), where L refers to the number of blocks, and then transform the features to the size of (L, d) through an adjustment layer;

[0112] Then, in each block, perform transformations on the input features from 256 to 128 dimensions and from 128 to 256 dimensions; after feature extraction through eight linear residual layers, obtain the latent representation Z through a fold operation;

[0113] A large compression ratio of r = H / n = W / m = 32 is obtained, which enables good anti-degradation robustness and manageable computational costs for global modeling in the second stage;

[0114] The decoder G of this application consists of 3 transposed convolutions and 1 convolution layer for upsampling; the size of the transposed convolution kernel is 4×4, indicating that both the width and height of the transposed convolution kernel are 4; the stride is 2, indicating that the sliding stride of the transposed convolution kernel on the input image is 2;

[0115] The padding size is 1, indicating that one pixel is padded at the edge of the input image to maintain the spatial size of the output; first, upsample the features of dimension 256×32×32 to 64×256×256 through three transposed convolutions, and then adjust the output to 256×256×3 through 1 convolution with a convolution kernel size of 3×3, a reflection padding parameter of 1, and a stride of 1 to obtain the reconstructed image.

[0116] Example 4:

[0117] Based on Example 3, it is further illustrated that the image inpainting stage based on the codebook prior specifically includes the following content:

[0118] In the existing Transformer architectures for image inpainting and completion, the indices of the quantized pixels are used both as inputs and as prediction targets; although this strategy of predicting missing indices using context indices improves computational efficiency, this input type of Transformer will have a serious information loss problem, which is not conducive to index sequence prediction; therefore, the parallel CSWinTransformer module (PCT) designed in this application directly uses

[0119] the feature vectors of the codebook encoder as inputs, which helps to make more accurate predictions while reducing information loss;

[0120] The parallel CSWinTransformer module is as Figure 3 shown,

[0121] for the feature vectors Add additional learnable position embeddings to preserve spatial information, and then flatten the feature vectors along the spatial dimension to obtain the final input of this module;

[0122] The model uses 12 parallel CSWin Transformer blocks. Each block consists of a parallel multi-head self-attention block, a cross-shaped window attention block, and a feed-forward layer (Feedforward1);

[0123] The number of heads of self-attention is set to 8. Different from common Transformer modules, the PCT module combines multi-head and cross-shaped windows, greatly reducing the computational amount and achieving a better repair effect. In addition, the cross-shaped window attention block adds a position encoding mechanism LePE to the linearly projected value V to enhance the local inductive bias;

[0124] It should be noted that the cross-shaped window attention and full self-attention in the PCT module are trained from different receptive fields and are connected together through residual connections. Therefore, the standard self-attention block will not be affected by the CSWin attention block. The Swish function in Feedforward1 can better smooth the gradient while retaining the non-linear characteristics of ReLU;

[0125] The design of the cross-shaped window and local enhanced position encoding is as follows:

[0126] Different from axial attention, the cross-shaped window attention splits the channels into horizontal and vertical stripes. Half of the heads capture the horizontal stripe attention, and the other half of the heads capture the vertical stripe attention;

[0127] Specifically, taking the horizontal stripe self-attention as an example, the feature matrix S is equally spaced into a series of non-overlapping horizontal stripe segments [S 1 ,..,S N , where N = H / b, and each stripe segment contains elements with b columns and W rows. In addition, the hyperparameter b is flexibly adjusted to balance the learning ability and computational cost. Assuming that the dimensions of the query, key, and value vectors corresponding to each head are all d, the output of the horizontal stripe self-attention processed by each head is defined by the following expression:

[0128] S = [S 1 ,S 2 ,…,S N ,

[0129] Y i =Attention(S i W Q ,S i W K ,S i W V),(6)

[0130] Attention H (S)=[Y 1 ,Y 2 ,…,Y N

[0131] where S i ∈R (b×W)×C , i = 1, ..., N, W Q ∈R C×d , W K ∈R C×d , W V ∈R C×d respectively represent the query matrix, key matrix, and value matrix obtained by linearly transforming the input feature matrix by each head;

[0132] S = [S 1 , S 2 , …, S N means that the feature matrix S is evenly divided into a series of non - overlapping horizontal stripe segments with a width of b. Each stripe segment S i contains elements with b columns and W rows;

[0133] N = H / b means that N is the number of horizontal stripe segments, equal to the height H of the feature matrix divided by the width b of each stripe segment;

[0134] S i ∈R (b×W)×C means that each stripe segment S i is a b×W matrix, where b is the number of columns, W is the number of rows, and C is the feature dimension;

[0135] W Q ∈R C×d , W K ∈R C×d , W V ∈R C×d means that these are the weight matrices of the linear transformation used to convert the input feature matrix S i into the query matrix (Query), key matrix (Key), and value matrix (Value). C is the dimension of the input feature, and d is the dimension of each head;

[0136] Y i =Attention(S i W Q , S i W K , S i W V ) means that this is the calculation process of the attention mechanism, where S i W​Q , S i W K are the linear transformation results of the query, key, and value respectively, and the output Y calculated by the attention mechanism i is the attention output of the i-th horizontal stripe segment;

[0137] Attention H (S) = [Y 1 , Y 2 , …, Y N represents that this is the set of attention output results of all horizontal stripe segments, denoted as the attention output in the horizontal direction;

[0138] Attention V (S) represents that this is the output result of the local self-attention operation performed on the vertical stripe region, denoted as the attention output in the vertical direction;

[0139] Similarly, the local self-attention operation performed on the vertical stripe region can also be derived accordingly, and the output result corresponding to each head is represented by Attention V (S);

[0140] The output of the PCT module described above passes through a linear layer and uses the Softmax function to be mapped into a probability distribution, which corresponds to the probability distribution of K latent vectors in the codebook e; that is, the probability of the image patch corresponding to the features in the codebook;

[0141] To quantify the degree of consistency between the model prediction and the class label, the PCT module is trained to predict the probability distribution p(s i | s <i ); making the training objective equal to minimizing the negative log-likelihood of the data representation:

[0142] L Transformer = E x′~p(x′) [-logp(s)] (7)

[0143] where p(s) = ∏ i p(s i | s <i );

[0144] p(s i | s <i ) is a conditional probability distribution, which represents the probability distribution of predicting the i-th index s <i in the prediction sequence under the condition that all elements s i with indices less than i in the known sequence are known. This probability distribution is learned by the PCT module during the training process;

[0145] LTransformer It is the loss function of the Transformer model, which is used to quantify the degree of consistency between the model prediction and the true class label. During the training process, the goal is to minimize this loss function;

[0146] E x~p(x′) is the symbol of the expected value, indicating the average of the samples x' drawn from the data distribution p(x'). Here, it represents the average value of the loss function calculated for all possible data samples x';

[0147] -logp(s) is the negative log-likelihood, which is used to calculate the difference between the probability distribution p(s) predicted by the model and the true label s. The negative log-likelihood is a commonly used loss function for optimizing probability models;

[0148] p(s) = ∏ i p(s i |s <i ) is the product of the probability distributions of all indices in the sequence, which represents the prediction probability of the model for the entire sequence. The probability of each index s i is based on the conditional probability of all previous indices s i ;

[0149] In each iteration, the quality of the generated image is evaluated using SSIM and PSNR quality evaluation metrics;

[0150] Gradient information collection: In each iteration, the gradient information of the generator and the discriminator is collected where L represents the loss function;

[0151] Learning rate adjustment:

[0152] The Adam optimizer is used, and its learning rate adjustment formula is:

[0153]

[0154] where, lr t is the learning rate of the t-th iteration, and β1, β2 are the hyperparameters of the Adam optimizer, which are set to 0.9 and 0.95 respectively;

[0155] To further adaptively adjust the learning rate, the gradient change rate grad vart is introduced to dynamically adjust the learning rate:

[0156] lr t+1 = lr t ·(1 + α3·grad vart )

[0157] where, grad vartis the gradient variance in the t-th iteration, and α3 is an adjustment factor used to control the degree of influence of gradient changes on the learning rate;

[0158] Gradient variance calculation:

[0159] Calculate the gradient variance to reflect the stability of the gradient:

[0160]

[0161] where g t,i is the gradient of the i-th parameter in the t-th iteration, N2 is the total number of parameters, and μ t is the average value of the gradient;

[0162] Learning rate update: Update the learning rate according to the above formula and use the new learning rate for parameter update in the next iteration; By introducing the gradient variance, this scheme can dynamically adjust the learning rate, and compared with the traditional fixed learning rate or simple learning rate decay strategy, it can more finely control the training process and improve the adaptability and effect of the model on complex image restoration tasks;

[0163] After each iteration, collect the generated restored image, obtain the structural similarity index and peak signal-to-noise ratio PSNR between the restored image and the reference image, and analyze and process the structural similarity index and peak signal-to-noise ratio PSNR to generate a comprehensive quality evaluation index QI, which is used to evaluate the image quality and restoration effect and generate a corresponding secondary learning rate adjustment strategy; The specific content includes the following:

[0164] After each iteration, collect the generated restored image and calculate the following quality parameters:

[0165] The structural similarity index SSIM is as follows:

[0166]

[0167] where x and y are the local windows of the reference image and the restored image respectively, and μ x and μ y are the mean values, is the variance, σ xy is the covariance, and c1 and c2 are constants used to stabilize the calculation;

[0168] The peak signal-to-noise ratio PSNR is as follows:

[0169]

[0170] where MAX I is the maximum value of the image pixels, and MSE(x, y) is the mean square error;

[0171] Combining SSIM and PSNR, a comprehensive quality evaluation index QI is generated, and its calculation formula is as follows:

[0172]

[0173] Parameter explanation, ω i is the weight factor, which is used to balance the influence of different quality evaluation indicators. Among them, ω1, ω2, ω3, and ω4 correspond to the weights of SSIM, PSNR, FSIM, and NIQE respectively, and determine the influence degree of each evaluation indicator on the overall QI value, and

[0174] NIQE is a reference-free image quality assessment method, which evaluates the image quality based on the deviation of the natural scene statistics NSS; its mathematical expression can be simplified as:

[0175]

[0176] where v represents the feature vector of the test image, v0 is the average value of the feature vectors extracted from the reference image library, Σ is the covariance matrix of the feature vectors, and Σ -1 is the inverse matrix of the covariance matrix; this formula calculates the Mahalanobis distance between the feature vector of the test image and the feature vector of the reference image;

[0177] FSIM represents Feature Similarity Index; FSIM is a similarity index based on image features, which is used to evaluate the similarity between two images; its mathematical expression is simplified as:

[0178]

[0179] where S L (x, y) is the luminance similarity at position (x, y), and S C (x, y) is the contrast similarity at position (x, y), and S p (x, y) is the phase consistency at position (x, y); this formula calculates a comprehensive similarity metric value by comprehensively considering features such as the luminance, contrast, and phase consistency of the image;

[0180] f(Metric i″′ ) is a complex function, and i″′ ∈ {1, 2, 3, 4}, where when i″′ in Metric i″′ takes the values 1, 2, 3, and 4 respectively, it represents SSIM, PSNR, FSIM, and NIQE respectively, and is used for non-linear transformation of each quality evaluation indicator. The formula is as follows:

[0181] f(SSIM) = log(1 + SSIM)

[0182] f(PSNR) = exp(-PSNR / 100)

[0183]

[0184] g(NIQE) is a normalization function used to adjust the impact of NIQE. x∈NIQE;

[0185] The value range of QI is set to (0,1). When QI is close to 1, it means that the image quality is close to the original image and the restoration effect is good; when QI is close to 0, it means that the image quality is poor and the restoration effect is not good.

[0186] Auxiliary formula:

[0187] Used to compress the NIQE value into the range of (0,1);

[0188] When the QI value increases, it means that the image restoration quality is improved and the image is closer to the visual and structural features of the original image; conversely, a decrease in the QI value indicates that the restoration effect is poor and the model parameters or training strategy need to be adjusted;

[0189] This formula comprehensively considers the influence of multiple image quality evaluation indicators by introducing complex nonlinear transformation functions and normalization functions, and dynamically adjusts the importance of each indicator through weight factors, thus achieving a comprehensive evaluation of the image restoration quality. This design not only improves the accuracy of the evaluation, but also enhances the adaptability of the model to different image restoration tasks.

[0190] According to the changes in these indicators, the learning rates of the generator and discriminator are dynamically adjusted. If the quality evaluation indicators improve slowly or decrease, the learning rate is increased to explore new parameter spaces; if the quality evaluation indicators improve steadily, the learning rate is maintained or moderately reduced to stabilize the training; specifically, the following are included:

[0191] Secondary adjustment of learning rate:

[0192] According to the generated comprehensive quality evaluation index QI, the learning rate is adjusted twice:

[0193] 1. t+1 =lr t ·(1+β5·(QIt-QItarget))

[0194] Among them, QIt is the comprehensive quality evaluation index at the tth iteration, QItarget is the target quality evaluation index, and β5 is an adjustment factor used to control the impact of the quality evaluation index on the learning rate;

[0195] The value range of QIt (0, 1) is divided into three intervals, namely Interval 1 (0, 0.3); Interval 2 [0.3, 0.7); Interval 3 [0.7, 1);

[0196] In Interval 1 (0, 0.3), the image quality is poor, and the learning rate needs to be increased to explore new parameters and quickly improve the image restoration effect. The threshold is set to 0.2. When the value is lower than this, the learning rate is urgently increased in order to achieve significant improvement;

[0197] Quantitative content description: The image restoration quality in this interval is poor and the restoration effect is not good; at this time, both the SSIM and PSNR metrics show a significant decline, decreasing by 30% and 25% respectively; FSIM and NIQE also show a serious degradation of the image quality, with FSIM decreasing by 20% and NIQE increasing by 40%; and if SSIM and PSNR do not improve by more than 5% or show a decline in three consecutive iterations, the learning rate should be increased; in this case, the learning rate should be significantly increased by 50% to explore the new parameter space and try to improve the image restoration effect;

[0198] Judgment criteria and rules: The threshold is set to 0.2. When the QIt value is lower than 0.2, the emergency adjustment mechanism is activated; the specific rules are as follows: If the QIt value is lower than 0.2 in three consecutive iterations, the learning rate is increased by 50%, and the changes in SSIM and PSNR are re-evaluated; if SSIM and PSNR do not improve significantly in the next five iterations, the improvement is not more than 10%, then the learning rate is further increased to 75%;

[0199] Interaction rule description: In the interval where the QIt value is lower than 0.3, the interactive changes of each parameter are as follows:

[0200] When SSIM drops by 30%, PSNR drops by 25% accordingly, FSIM drops by 20%, and NIQE increases by 40%; this negative correlation change between parameters indicates the overall degradation of the image quality; the 50% increase in the learning rate aims to find new parameter combinations that can improve the image quality through the exploration of the parameter space;

[0201] Interval 2: [0.3, 0.7), there is room for improvement in the image quality; an adjustment strategy is adopted to maintain or slightly increase the learning rate to steadily improve the image quality, and the threshold is 0.5 to maintain the stability of training and continuous improvement;

[0202] Quantitative content description, when the QIt value is in the interval [0.3, 0.7), the image restoration quality has improved, but it still has not reached the ideal state; the SSIM and PSNR metrics show a slight improvement, increasing by 10% and 15% respectively; FSIM remains stable, and NIQE drops by 10%. At this time, the learning rate is adjusted and increased by 10% to maintain the stability of training and continue to observe the changes in the metrics;

[0203] Judgment criteria and rules, set the threshold value to 0.5. When the QIt value fluctuates around 0.5, adopt a conservative strategy. The specific rules are as follows: If the QIt value fluctuates between 0.45 and 0.55 for five consecutive iterations, keep the current learning rate unchanged; If the QIt value is lower than 0.45 or higher than 0.55 for five consecutive iterations, increase or decrease the learning rate by 5% accordingly.

[0204] Description of interaction rules. In the interval where the QIt value is in [0.3, 0.7), the interaction changes of each parameter are as follows: When SSIM increases by 10%, PSNR increases by 15%, FSIM remains unchanged, and NIQE decreases by 10%. This positive correlation change between parameters indicates the gradual improvement of image quality. A moderate increase in the learning rate of 10% aims to stabilize the current improvement trend and avoid unstable training caused by excessive adjustment.

[0205] Interval three [0.7, 1), indicating that the image quality is close to ideal. In this interval, reduce the learning rate to stabilize the training and prevent overfitting. Set the threshold value to 0.85. When exceeding this value, further reduce the learning rate to ensure continuous optimization and stability of the quality.

[0206] Description of quantization content:

[0207] In the interval where the QIt value is in [0.7, 1), the image inpainting quality is close to or reaches the ideal state. The SSIM and PSNR metrics show significant improvements, increasing by 20% and 25% respectively. FSIM and NIQE also show significant improvements in image quality, with FSIM increasing by 15% and NIQE decreasing by 30%. At this time, the learning rate should be reduced by 10% to stabilize the training and prevent overfitting.

[0208] Judgment criteria and rules, set the threshold value to 0.85. When the QIt value exceeds 0.85, start the stable strategy. The specific rules are as follows: If the QIt value exceeds 0.85 for three consecutive iterations, reduce the learning rate by 10%, and continue to monitor the changes in SSIM and PSNR. If SSIM and PSNR remain stable or continue to improve in the next five iterations, further reduce the learning rate to 15%.

[0209] Description of interaction rules. In the interval where the QIt value is higher than 0.7, the interaction changes of each parameter are as follows: When SSIM increases by 20%, PSNR increases by 25%, FSIM increases by 15%, and NIQE decreases by 30%. This positive correlation change between parameters indicates the significant improvement of image quality. The learning rate is reduced by 15% to stabilize the current high-quality inpainting state and prevent overfitting caused by too high a learning rate.

[0210] Example five:

[0211] Based on Embodiment 4, it is further illustrated that the designed multi-scale feature guidance module makes full use of the features of the non-damaged area to promote the coordination between the generated area and the undamaged area in terms of structure and texture, and improve the quality and fidelity of the repair result; specifically, it includes the following contents:

[0212] As Figure 4 shown, the designed multi-scale feature guidance module aims to retain the details of the non-masked area of the image; assuming that the input image is the masked input Y with the mask m, this module represents the masked image as a multi-layer feature map instead of compressing it into a single layer feature;

[0213] Injecting a large-kernel-based convolution into the multi-scale feature guidance module aims to integrate the advantages of CNN operations and attention mechanisms;

[0214] Specifically, the LKA (Large Kernel Attention) structure is used, and this structure uses a depthwise convolution (DW-Conv) with a dilation rate of d to extract local features. Then, a (2d - 1)×(2d - 1) depthwise dilated convolution (DW-D-Conv) is used to capture long-range dependencies. Finally, a 1×1 pointwise convolution is used to integrate information and adjust the number of channels, enhancing the interaction between channels;

[0215] Since LKA focuses on optimizing the feature expression of the occluded area with the help of a wide receptive field, it helps in the global learning of regular textures in the frequency domain; in addition, to ensure the generalization ability of LKA, a feed-forward network 2 is added after the LKA module; specifically, the feed-forward network 2 consists of RMS normalization, a 3×3 convolution, a Swish activation function, a 3×3 convolution, and Dropout;

[0216] Among them, the RMSNorm normalization function is used to improve the training stability; the Swish function can better smooth the gradient while retaining the non-linear characteristics of ReLU, and can also solve the problem that the ReLU function is not zero-centered and has a zero gradient in the negative part.

[0217] Embodiment 6:

[0218] Based on Example 5, it is further explained that this experiment trained and evaluated the model in this paper on three different datasets: Celeba-HQ is an extended version of the CelebA dataset, containing high-quality and high-resolution face images. 27,000 images were selected for training, 3,000 images were used for testing and validation, and 20 scene categories were selected for the experiment, among which 90,000 images were used for training and 10,000 images were used for quantitative evaluation; the self-made Tibetan Thangka dataset, including Buddhist Thangka, Esoteric and Exoteric Thangka, and family Thangka, etc., among which 2,500 were used for training and 500 were used for testing and validation, as shown in Table 1 specifically;

[0219] Table 1 Settings of Celeba, Facade and self-made Thangka datasets

[0220]

[0221]

[0222] For quantitative comparison, this embodiment uses various image quality metrics, including the traditional peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), mean absolute error (MAE), and the latest feature-based learned perceptual image patch similarity (LPIPS);

[0223] The implementation details are as follows:

[0224] For the first-stage shared codebook learning stage, the method in this paper uses the Adam optimizer (β1 = 0, β2 = 0.9) for optimization, and the batch size is 16;

[0225] For the second-stage image inpainting stage based on the codebook prior, use Adam (β1 = 0.9, β2 = 0.95) for optimization, the batch size is 4, set the learning rates of the two stages to 2e-4 and 3e-4 respectively, and use the cosine scheduler for decay; the method in this embodiment is implemented using the PyTorch framework and trained using 1 NVIDIA 3090 GPU;

[0226] All comparison models were compared in the Celeba-HQ and Places2 datasets. This application also retrained EC, CTSDG, ICT, PUT, and MAT on the self-made Thangka dataset to further discuss the inpainting effect;

[0227] By Figure 5It can be seen that the first-stage network of the proposed model is trained on the Places2 dataset. As the training progresses, the Quantize loss and Adv loss rise briefly, and then stabilize in the oscillation. The L1 loss, Perceptual loss, and Style loss on the right are gradually reduced through continuous training and tuning, and the quality of the generated images can be improved.

[0228] Embodiment seven:

[0229] Further explanation is given on the basis of Example 6. Figure 6 , Figure 7 and Figure 8 Visual comparison of restoration results of randomly selected test images from Celeba-HQ, Places2, and self-made thangka datasets are shown;

[0230] The proposed model is compared with existing advanced methods on the Celeba dataset. Figure 6 As shown in the figure, the EC and CTSDG restoration methods have incomplete structure prediction when facing large-area defective images, resulting in large-area distortion in the restoration results;

[0231] exist Figure 6 The cheeks and eyes of the characters in rows 3 and 5 in (b) and (c) are missing;

[0232] ICT uses Transformer to reconstruct visual priors. The overall structure of the restoration result is relatively reasonable, but the image restoration details are not perfect.

[0233] like Figure 6 (d) The frames of the repaired glasses in the first and second rows are deformed, and the human eye is asymmetrical. MAT is a large-area defect repair model guided by a mask. It does not work well for small missing areas in the image.

[0234] like Figure 6 In the restoration results of the 4th and 5th rows in (e), the hair and eyes of the characters do not conform to the facial features;

[0235] The P-VQVAE encoder in PUT converts the original resolution image into latent features in a non-overlapping manner, avoiding the cross-information influence, but the understanding of semantic features is insufficient;

[0236] Figure 6 The last two rows in (f) do not integrate the filled area with the surrounding pixels well, the color of the hat is not coordinated, and the details of the glasses are not perfect;

[0237] Compared with the above methods, the proposed algorithm combines the idea of ​​vector quantization and introduces the parallel CSWinTransformer module and the multi-scale feature guidance module, showing a restoration effect with clear edges and natural color transitions. Even in severely damaged areas, the restoration content is semantically reasonable without any inharmonious or abrupt parts.

[0238] Figure 7 The restoration effects of each model on the places2 dataset are shown. EC and CTSDG produce blurry and inconsistent boundary artifacts because they cannot capture long-distance features. ICT loses a lot of information during the downsampling process, so the restored horse legs have defects.

[0239] The repair results of MAT and PUT show semantic inconsistency and color differences;

[0240] like Figure 7 The third row of (e) generates a cabinet on the grassland;

[0241] Figure 7 In the fourth row of (f), the background restoration after removing the person results in unnatural reefs;

[0242] This method avoids the loss of image information by learning through shared codebook, thus obtaining richer semantic information and achieving high-fidelity image restoration.

[0243] Figure 8 Comparison chart of repairs of damaged areas of various thangkas. When the EC algorithm faces large-area defects, it is limited by a small receptive field, resulting in large-scale texture blur in the repair results and the inability to reconstruct the image structure.

[0244] When dealing with partially missing areas of a person, the CTSDG algorithm can reconstruct the basic outline of the person by taking advantage of edge information. However, its restoration level is not ideal in terms of material properties and microscopic details.

[0245] from Figure 8 (e) It can be seen that the MAT algorithm still shows strong repair ability when the defect area is large. It repairs the eye area in the second and fourth rows that the first two algorithms cannot repair. However, the eye position is still unreasonable and the face is distorted.

[0246] Visually, after the restoration of these five images by the proposed algorithm, both the structural coherence and the accuracy of texture details are consistent with the original images; therefore, it can be verified that the proposed algorithm is more suitable for images with complex textures and rich colors such as thangkas;

[0247] Due to the differences in the feelings and judgment criteria of different individuals, specific numerical comparisons can more accurately reflect the subtle differences in advantages and disadvantages, making the research results more verifiable and reproducible. Four evaluation criteria, namely PSNR, SSIM, MAE, and LPIPS, were selected, and experiments were conducted on the Celeba-HQ, Places2, and self-made Thangka datasets.

[0248] All test images were uniformly set to a resolution of 256×256, and irregular masks with the same masking ratio were imposed on these images. The existing mainstream algorithms such as EC, CTSDG, ICT, MAT, PUT, etc. were compared with the algorithm model proposed in this paper. On this basis, the specific numerical values of various evaluation indicators were statistically obtained, as shown in Table 2.

[0249] Analyzing the results in Table 2, it can be concluded that in the Places2 scene dataset and the self-made Thangka dataset, compared with other algorithms, the algorithm in this paper shows significant advantages in terms of similarity both at the pixel level and the structural level. In some cases, there are differences between the objective evaluation indicators and the intuitive visual observations, which exactly confirms the limitations of relying solely on a single objective or subjective evaluation method to measure the quality of image inpainting. At the same time, this also strongly proves the rationality and necessity of the comprehensive evaluation method combining the two evaluation methods in this paper.

[0250] The following is the objective quantitative comparison between the algorithm in this paper and EC, CTSDG, ICT, MAT, PUT on three datasets with different masking ratios.

[0251] Table 2

[0252]

[0253] 3.4 Ablation study:

[0254] In order to verify the effectiveness of the key components of the method proposed in this application, a series of ablation experiments were conducted on the self-made Thangka dataset in this application. The main experiments include the following:

[0255] (b) In the Encoder part of the CDCT model in this paper, the Linear layer is replaced with a Conv layer of the same size.

[0256] (c) The parallel CSWinTransformer module is replaced with a standard Transformer module of the same quantity.

[0257] (d) The parallel structure between the standard self-attention and CSWin attention in the PCT module is changed to a serial structure.

[0258] (e) Remove the multi-scale feature guidance module.

[0259] (f) Replace the LKA structure in the multi-scale feature guidance module with a Conv layer.

[0260] (g) Implement the complete network structure of this application.

[0261] Table 3 shows the objective evaluation results of the ablation study of different components; Variants 1 and 2 use the encoder from VQGAN and the standard Transformer module, resulting in excessive information compression and insufficient utilization of local details, which affects the performance of the model; Variant 3 changes the parallel manner of the internal structure of the PCT module to a sequential manner, and the metrics decrease slightly.

[0262] Variants 4 and 5 demonstrate that the multi-scale feature guidance module can fully utilize the features of the unmasked area while maintaining the decoding ability of the latent representation; The complete model with the addition of the linear residual encoder module, the parallel CSWinTransformer module, and the multi-scale feature guidance module has an average increase of 1.741 dB and 0.038 in PSNR value and SSIM value, and an average decrease of 0.0221 and 0.0053 in LPIPS value and MAE value compared to other replaced components; This indicates that these improved modules have a positive impact on the quality of the repair results.

[0263] The following Table 3 shows the quantitative ablation analysis of the method in this paper on the self-made Thangka dataset:

[0264] Table 3

[0265]

[0266]

[0267] Figure 9 Gives the visualization results of each component of the model in this paper.

[0268] As Figure 9 shown in (b), there is a lack of consistency between the damaged area and the surrounding area in Variant 1, and there are light and dark color differences in the skin color of the character's arms, face, and chest; It can be seen from Variants 2 and 3 that the beads in the character's hand are uneven in size, resulting in an artifact problem; As Figure 9 shown in (e), after removing the multi-scale feature guidance module, the local effective information of the repair result decreases, the character's fingers are affected by the surrounding blue background, and the edge structure of the image eyes is distorted and the transition is unnatural.

[0269] As Figure 9 shown in (g), it verifies the effectiveness and superiority of the CDCT algorithm proposed in this paper in dealing with the problem of color-complex images, and a more realistic and reasonable repair effect is obtained using this method.

[0270] In the first-stage network, the proposed model embeds continuous features into a discrete space of finite size, namely k code vectors; in this embodiment, an ablation study was conducted to understand the impact of the number of code vectors (k) in the codebook on the model performance; Table 4 shows that when the codebook size is 1024 on the Thangka dataset, better results are produced and it is more effective in improving the reconstruction quality, rather than the larger the codebook vector, the more reasonable the data compression;

[0271] Table 4 Impact of different codebook sizes on model performance:

[0272] Codebook size (k) PSNR / dB↑ SSIM↑ LPIPS↓ MAE↓ 512 26.033 0.839 0.0491 0.0250 1024 27.868 0.889 0.0311 0.0208 2048 26.889 0.868 0.0414 0.0216

[0273] In this embodiment, 5 groups of experiments were conducted to determine the optimal hyperparameter settings of Attention head and embedding dimension;

[0274] When the attention head of the PCT module is set to 8 and the embedding dimension is set to 512, the model can better capture the long-range dependencies in the input sequence, and the four evaluation metrics are significantly improved, while avoiding the excessive embedding dimension from increasing the computational burden of the model, as shown in Table 5;

[0275] Table 5 Performance of different hyperparameter combinations of the PCT module

[0276] Heads Embedding dims Params (M) PSNR / dB↑ SSIM↑ LPIPS↓ MAE↓ 4 512 53.05 26.684 0.892 0.0420 0.0218 8 512 53.05 27.752 0.908 0.0302 0.0200 16 512 53.05 26.158 0.874 0.0471 0.0231 8 256 20.95 23.487 0.822 0.0884 0.0292 8 768 106.13 26.174 0.875 0.0470 0.0230

[0277] This embodiment proposes an image inpainting method that combines a discrete codebook and Transformer, which has multiple new design features;

[0278] First, a linear encoder is used instead of convolutional downsampling, and the feature blocks are encoded independently to avoid the influence of information crossover. Different from conventional inpainting models, in this embodiment, a codebook is used to discretely encode the intermediate features of the model; second, to avoid information loss in Transformer, the input to Transformer is not discrete tokens, i.e., indices, but the features output by the encoder;

[0279] At the same time, the discrete tokens are only used as the output of Transformer; in addition, the design of the parallel CSwinTransformer module improves the accuracy of token prediction and also reduces the number of parameters; subsequently, an additional multi-scale feature guidance module is added to the decoder, which can better retain the local details of the non-defective area and recover the details from the quantized output of the encoder;

[0280] Through extensive experiments on multiple representative tasks, it is verified that the CDCT method can not only process thangka images with diverse colors and rich semantics, but also effectively repair various defects in natural images; through in-depth ablation research, the effectiveness of the proposed model design is demonstrated; it aims to accurately identify and repair the incomplete parts in thangka images, which serves as a new direction for optimizing image restoration work.

[0281] Embodiment eight:

[0282] A thangka image restoration system, the system is used to execute the method described.

[0283] Embodiment nine:

[0284] A storage medium stores a computer program, which implements the steps of the thangka image restoration method when executed by a processor.

[0285] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters in the formula are set by technicians in this field according to actual conditions.

[0286] The above embodiments may be implemented in whole or in part by software, hardware, firmware or any other combination thereof. When implemented by software, the above embodiments may be implemented in whole or in part in the form of a computer program product. Those skilled in the art may appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein may be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed by hardware or software methods depends on the specific application and design constraints of the technical solution.

[0287] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, and may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0288] The above description is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.

Claims

1. A thangka image restoration method, characterized in that: The specific steps include: Step S1, collecting thangka images containing damaged parts and missing parts, and preprocessing the collected thangka images, and constructing the preprocessed thangka images into an image data set; Step S2, constructing an encoder-decoder structure based on the Transformer model and using it to jointly learn a discrete codebook, inputting the thangka image in the image data set into the constructed encoder, using the encoder to divide the input thangka image into non-overlapping sub-regions of fixed size, and mapping the sub-regions to a continuous latent space representation through a nonlinear transformation to obtain a feature vector; Step S3, introducing a vector quantization codebook, and performing vector quantization on the continuous potential feature vector output by the encoder, wherein the discrete codebook is constructed by using a clustering algorithm, and each vector in the codebook represents a potential space representation of a sub-region in the image, and finally the obtained discrete codebook is used as codebook prior knowledge; Step S4, construct a parallel CSWin resolution Transformer module, which adopts the design of cross-shaped window and local enhanced position encoding, takes the feature vector in step S3 as input, adds additional learnable position embedding to the feature vector to preserve spatial information, and then flattens the feature vector along the spatial dimension to obtain the final input of the module to predict the probability distribution of the next index; Step S5: Use the parallel CSWin resolution Transformer module in step S4 to accurately infer the index of the missing token, and use these indexes to find the corresponding discrete vectors from the discrete codebook obtained in step S3 for image restoration. After completing one restoration attempt, the system will enter an iterative loop; Step S6, after each iteration, collect the generated repaired image, and obtain the structural similarity index between the repaired image and the reference image, and the peak signal-to-noise ratio PSNR between the repaired image and the reference image, and analyze and process the structural similarity index and the peak signal-to-noise ratio PSNR to generate a comprehensive quality evaluation index QI, which is used to evaluate the image quality and the repair effect, and generate a corresponding learning rate secondary adjustment strategy; Step S7, based on the adaptive learning rate adjustment, a multi-scale feature guidance module is designed, which utilizes the features of the non-damaged area to promote the consistency of the generated area with the non-damaged area in structure and texture, thereby improving the quality and fidelity of the repair result; Step S8: During the repair process, dynamically adjust the learning rate of the model according to the comprehensive quality evaluation index QI and the learning rate adjustment strategy generated in step S6 to optimize the repair effect; Step S9: After all iterations are completed, the final restored thangka image is post-processed, including but not limited to image enhancement, color correction and detail optimization, to improve the quality and visual effect of the restored image, and finally the restored thangka image is output.

2. A thangka image restoration method according to claim 1, characterized in that: In the shared codebook learning phase, the system architecture consists of three core components: the codebook encoder E, the codebook decoder G, and a codebook containing K discrete codes When processing the input image When the codebook encoder E first converts the image I t Transformed into a latent representation Z in a high-dimensional space, Represents an image with height H, width W and 3 color channels, where H and W represent the vertical and horizontal dimensions of the image respectively; H and W are used to describe the actual size of the image; Right now Here d represents the number of dimensions that make up this latent vector; where m×n is the spatial resolution in the latent representation, and m and n represent the height and width of the spatial dimension respectively; Quantize each vector in this spatial latent representation Z to the nearest codeword c in the codebook k , thus obtaining the following vector quantization output Z c and the corresponding code token sequence s∈{0,…,N-1} m′·n′ , and marked as formula (1): s (i,j) =argmin k ‖Z (i,j) -c k ‖2 Among them, each element It's Z (i,j) The closest codeword found in codebook C; the quantization operation is performed by calculating the distance ||Z (i,j) -c k ||And select the code word corresponding to the minimum distance to achieve; is the output of vector quantization, representing the potential representation Z at position i and resolution j (i,j) quantized to the closest codeword in codebook C; q(Z) is the vector quantization operation that maps each vector in the latent representation Z to the closest codeword in the codebook C; is to find the code in C that matches Z (i,j) The codeword c with the smallest distance k operation; argmin means finding the distance || Z (i,j) -c k ||The smallest codeword c k The index of s (i,j) is an element in the quantized code token sequence, representing the potential representation Z at position i and resolution j (i,j) The index of the codeword to be quantized; arg min k ‖Z (i,j) -c k ‖2 is found in codebook C with Z (i,j) The codeword c with the smallest Euclidean distance k The operation arg min represents finding the Euclidean distance ‖Z (i,j) -c k ‖2The smallest codeword c k The index of c k ∈C means that each codeword ck in the codebook C is a predefined vector used to quantize the potential representation Z; Z (i,j) is the vector of the potential representation Z at position i and resolution j; ||Z (i,j) -c k || is Z (i,j) Code word c k The distance between them is the Euclidean distance or other distance metric; ‖Z (i,j) -c k ‖2 is Z (i,j) Code word c k The Euclidean distance between s∈{0,…,N-1} m′·n′ is the quantized code token sequence, where m′·n′ represents the spatial size of the potential representation Z, that is, the total number of vectors contained in Z. Here, m′·n′ refers to the number of elements in the potential representation Z, each element corresponds to a potential vector, and N is the number of codewords in the codebook C; Then, the decoder G is given Z c Reconstruct high-quality image I rec ; The resulting m′·n′ code token sequence s represents a new potential discrete representation, That is, when (i,j) = k, Overall Reconstruction I rec ≈I t The formula is as follows and is marked as formula (2): I rec =G(Z c )=G(q(E(I))) The encoder performs a mapping operation to convert image data of size H×W into a discrete encoding form of scale H / m×W / n; and uses reconstruction loss to enable end-to-end training of the codebook and model.

3. A thangka image restoration method according to claim 2, characterized in that: The codebook encoder E design point includes that the encoder of the shared codebook learning stage is designed to process the input image in a non-overlapping patch manner and through multiple linear residual layers; The codebook prior-based image inpainting stage consists of adding additional learnable position embeddings to the feature vector to preserve spatial information, followed by flattening the feature vector along the spatial dimension to obtain the final input to this module; The model uses parallel CSWinTransformer blocks, where each block consists of parallel multi-head self-attention blocks and cross-window attention blocks, and feed-forward layers; the PCT module combines multi-head and cross-window, and the cross-window attention block adds a position encoding mechanism LePE on the linear projection value V to enhance the local inductive bias; The cross-shaped window attention splits the channel into horizontal and vertical stripes, half of the head captures the horizontal stripe attention, and the other half of the head captures the vertical stripe attention; The output of the PCT module passes through a linear layer and is mapped into a probability distribution using a Softmax function, which corresponds to the probability distribution of the K potential vectors in the codebook e.

4. A thangka image restoration method according to claim 3, characterized in that: The PCT module is trained to predict the probability distribution p(s) of the next index i |s <i ), making the training objective equal to minimizing the negative log-likelihood of the data representation; In each iteration, the quality of the generated image is evaluated using SSIM and PSNR quality evaluation indicators; Gradient information collection: In each iteration, the gradient information of the generator and the discriminator is collected. Where L represents the loss function; Learning rate adjustment: Using the Adam optimizer, the learning rate adjustment formula is: Among them, lr t is the learning rate of the tth iteration, β1 and β2 are the hyperparameters of the Adam optimizer, which are set to 0.9 and 0.95 respectively; In order to further adaptively adjust the learning rate, the gradient change rate grad is introduced vart To dynamically adjust the learning rate: lr t+1 =lr t ·(1+α3·grad vart ) Among them, grad vart is the gradient variance in the tth iteration, and α3 is an adjustment factor used to control the influence of gradient changes on the learning rate; Calculate the gradient variance to reflect the stability of the gradient: Among them, g t,i is the gradient of the i-th parameter at the t-th iteration, N2 is the total number of parameters, μ t is the average value of the gradient.

5. A thangka image restoration method according to claim 4, characterized in that: After each iteration, the generated repaired image is collected, and the structural similarity index and peak signal-to-noise ratio (PSNR) between the repaired image and the reference image are obtained. The structural similarity index and peak signal-to-noise ratio (PSNR) are analyzed and processed to generate a comprehensive quality evaluation index (QI). This index is used to evaluate the image quality and the repair effect, and to generate a corresponding secondary adjustment strategy for the learning rate. After each iteration, the resulting inpainted image is collected and the following quality parameters are calculated: The structural similarity index SSIM is as follows: In SSIM(x,y), x and y are the local windows of the reference image and the repaired image respectively, and μ x , μ y is the mean, is the variance, σ xy is the covariance, c1 and c2 are constants used to stabilize the calculation; The peak signal-to-noise ratio PSNR is as follows: In PSNR(x,y), x and y are the local windows of the reference image and the repaired image respectively, and MAX I is the maximum value of the image pixels, MSE(x,y) is the mean square error; Combining SSIM and PSNR, a comprehensive quality evaluation index QI is generated. The calculation formula is as follows: Parameter explanation, ω i is a weight factor used to balance the impact of different quality assessment indicators, where ω1, ω2, ω3, and ω4 correspond to the weights of SSIM, PSNR, FSIM, and NIQE, respectively; f(Metric i″′ ) is a complex function, i″′∈{1, 2, 3, 4}, where Metric i″′ When i″′ takes the values ​​of 1, 2, 3, and 4, it represents SSIM, PSNR, FSIM, and NIQE, respectively, and is used to perform nonlinear transformation on each quality assessment indicator. The formula is as follows: f(SSIM)=log(1+SSIM) f(PSNR) = exp(-PSNR / 100) g(NIQE) is a normalization function used to adjust the impact of NIQE. x∈NIQE; The value range of QI is set to (0,1). When QI is close to 1, it means that the image quality is close to the original image and the restoration effect is good; when QI is close to 0, it means that the image quality is poor and the restoration effect is not good.

6. A thangka image restoration method according to claim 5, characterized in that: When the QI value increases, it means that the image restoration quality is improved and the image is closer to the visual and structural features of the original image; On the contrary, a decrease in QI value indicates that the restoration effect is poor and the model parameters or training strategy need to be adjusted; According to the changes in these indicators, the learning rates of the generator and discriminator are dynamically adjusted. If the quality evaluation indicators improve slowly or decrease, the learning rate is increased to explore new parameter spaces; if the quality evaluation indicators improve steadily, the learning rate is maintained or moderately reduced to stabilize the training; specifically, the following are included: According to the generated comprehensive quality evaluation index QI, the learning rate is adjusted twice: lr t+1 =lr t ·(1+β5·(QIt-QItarget)) Among them, QIt is the comprehensive quality evaluation index at the tth iteration, QItarget is the target quality evaluation index, and β5 is an adjustment factor used to control the impact of the quality evaluation index on the learning rate; When QIt is in the interval 1 (0, 0.3), the image quality is poor, and the learning rate needs to be increased to explore new parameters and quickly improve the image restoration effect. The threshold is set to 0.

2. When it is lower than this value, the learning rate is urgently increased to achieve significant improvement; When the QIt value is in the interval [0.3, 0.7), there is room for improvement in image quality. An adjustment strategy is adopted to maintain or slightly increase the learning rate to steadily improve the image quality, and the threshold is 0.5 to maintain the stability and continuous improvement of training. When the QIt value is in the interval [0.7, 1), it means that the image quality is close to ideal. In this interval, the learning rate is reduced to stabilize the training and prevent overfitting. The threshold is set to 0.

85. When it exceeds this value, the learning rate is further reduced to ensure continuous optimization and stability of quality.

7. A thangka image restoration method according to claim 6, characterized in that: The multi-scale feature guidance module is designed, which utilizes the features of the non-damaged area to promote the consistency of the structure and texture between the generated area and the undamaged area, thereby improving the quality and fidelity of the repair result; specifically, it includes the following contents: Assuming the input image is a mask input Y with mask m, this module represents the mask image input as a multi-layer feature map, injects a large kernel-based convolution in the multi-scale feature guidance module, Use the LKA structure, which uses a dilation rate of d Deep convolution extracts local features, then captures long-distance dependencies through a (2d-1)×(2d-1) deep dilated convolution. Finally, 1×1 point-by-point convolution integrates information and adjusts the number of channels to enhance the interaction between channels. A feedforward network 2 is added after the LKA module, and the feedforward network 2 consists of RMS normalization, 3×3 convolution, Swish activation function, 3×3 convolution and Dropout.

8. A thangka image restoration system, characterized by: The system is used to execute the method according to any one of claims 1 to 7.

9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the thangka image restoration method described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Image generation system and method

    CN113449135A

  • SNAU-Net-based liver and tumor segmentation method

    CN118196113A