Dunhuang mural digital repairing method based on Mamba enhanced coding feature fusion network

By adopting a method based on Mamba enhanced coding feature fusion network in the digital restoration technology of murals, combined with the main repair network and the detailed enhancement network, the limitations of the existing technology in the recovery of complex textures and large-area defects are solved, and the high-quality mural restoration effect is achieved.

CN120163741APending Publication Date: 2025-06-17LANZHOU JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510255581.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing digital restoration technology of murals has limitations in restoring complex textures and large-area defects, and it is difficult to meet the needs of high-quality mural restoration.

Method used

The digital restoration method of Dunhuang murals based on Mamba enhanced coding feature fusion network is adopted. Through the combination of main repair networks and details enhancement networks, the initial restoration and detail restoration of murals are completed by fusion of texture and geometric features. The main repair network extracts features through dual encoders, learns texture and geometric prior information, and introduces gated encoding modules and prior encoding modules. Detail enhancement The network further improves the repair effect through the combination of jump connections and perceived loss and style loss.

Benefits of technology

High-quality digital restoration of murals has been achieved, significantly improving the restoration effect of texture and geometric features, reducing artifacts and texture blur, and improving the clarity and naturalness of the repaired image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163741A_ABST
    Figure CN120163741A_ABST
Patent Text Reader

Abstract

The invention discloses a Dunhuang mural digital restoration method based on a Mama enhanced coding feature fusion network, and relates to the technical field of mural digital restoration. The method comprises two sub-networks, namely a main repair network and a detail enhancement network. The main repair sub-network completes the preliminary repair task of the mural through the fusion of texture and geometric features, and the detail enhancement network performs secondary repair on the basis of the main repair network and introduces jump connection to enhance the detail capture capability. A contrast experiment result on a Dunhuang mural data set shows that compared with an existing popular method, the method disclosed by the invention obtains a better mural digital repairing effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital restoration of murals, and particularly to a method for digital restoration of Dunhuang murals based on a Mamba enhanced coding feature fusion network. Background Art

[0002] Early image restoration methods mainly included partial differential equations (PDEs) and texture synthesis techniques. PDEs filled small-scale defects by simulating pixel diffusion, but it was difficult to restore large missing areas and complex textures. Texture synthesis methods filled in the missing parts by replicating surrounding textures. With the development of technology, low-rank matrix factorization methods were introduced, which restored images by extracting self-similar features, but there were still limitations in restoring complex textures and large-scale defects, and it was difficult to meet the requirements of high-quality mural restoration.

[0003] In recent years, image restoration technologies based on deep learning have developed rapidly. By analyzing the patterns, colors, and geometric features of murals, algorithms are used to fill in damaged areas and restore the integrity and beauty of the murals. In particular, models such as convolutional neural networks (CNNs) and Transformers have been widely used in image restoration. For example, the CNN-based restoration model proposed by Iizuka et al. can handle small-scale defects well, but when faced with murals with large cracks, the generated content often lacks consistency and is prone to problems such as texture and structure distortion. Generative adversarial networks can generate more realistic image content through adversarial training of generators and discriminators, especially suitable for the restoration of complex scenes and large damaged areas. Kingma et al. proposed a VAE method that generates images by modeling the distribution in the latent space, but the generated images are often relatively blurred and difficult to meet the requirements in terms of details and sharpness. The GAN method proposed by Goodfellow et al. can generate high-quality images, but instability and mode collapse often occur during training, especially when the data distribution is complex or the loss function is designed unreasonably. The introduction of self-attention mechanisms and Transformer architectures enables the model to better capture global dependencies and perform well in maintaining the overall consistency of images and restoring details. For example, the Restormer model proposed by Zamir et al. performs well in capturing the global structure of images, but it has high hardware requirements in applications.

[0004] Aiming at the deficiencies of the above methods, the present invention proposes a Dunhuang mural restoration network that fuses texture and geometric features, including two sub-networks: a main restoration network and a detail enhancement network. The main restoration sub-network completes the preliminary restoration task of the mural through the fusion of texture and geometric features. The detail enhancement network performs secondary restoration on the basis of the main restoration network and introduces skip connections to enhance the detail capture ability. The comparative experimental results on the Dunhuang mural dataset show that the method of the present invention achieves better digital restoration effects of murals than existing popular methods. Summary of the Invention

[0005] The purpose of the present invention is to provide a digital restoration method for Dunhuang murals based on the Mamba enhanced coding feature fusion network to solve the above problems.

[0006] To achieve the above purpose, the technical solution adopted by the present invention is as follows: The Dunhuang mural restoration network consists of two sub-networks: a main restoration network and a detail enhancement network. The main restoration network uses dual encoders to extract the features of the damaged mural and the mural line drawing respectively, learns the mural texture and geometric prior information, and fills the missing area content. The encoder for extracting the features of the mural line drawing is composed of a gated coding module and a prior coding module. On this basis, the encoder for extracting the features of the damaged mural adds a MAMBA enhanced coding module, introduces a scanning mechanism to reduce the computational overhead, and adds a dynamic multi-scale semantic fusion module. Through multi-scale feature extraction, adaptive dynamic weight allocation, and gated attention mechanism, after the main restoration network completes the preliminary restoration task of the image, the detail enhancement network is used for the detail restoration of the image to achieve the delicate capture of geometric features, the accurate restoration of texture, and the high-precision reconstruction of the damaged part. The specific content is as follows: Step 1: Use the dual encoders of the main restoration network to extract the features of the damaged mural and the mural line drawing, learn the mural texture and geometric prior information, and fill the missing area content; The main restoration network effectively constrains the latent space of the image through the prior encoder module, and uses the local texture details extracted from the damaged mural and the global structure information obtained from the mural line drawing to preliminarily fill the missing area of the image; The main restoration network introduces gated convolution to adaptively control the action area of the convolution kernel, filter irrelevant information, and enable the network to focus on filling the missing area. L main The loss function is designed as follows: (1)

[0007] In the formula, I out1 is the output of the main restoration network, I g is the real image.M is a mask, and is the balance coefficient, and represent the KL loss functions of geometric features and texture features respectively, is the adversarial loss; Step 2: Based on the image initially generated by the main repair network, use the detail enhancement network to further repair the details of the image; For the subtle changes in complex textures and edges, the main repair network designs skip connections to enhance the ability to capture details. In addition, the detail enhancement network introduces perceptual loss and style loss , and the perceptual loss extracts feature maps through a pre-trained VGG network, comparing the repaired image with the real image from the perspective of high-level features to ensure their visual consistency in semantics: (2)

[0008] In the formula, F i represents the feature extraction result of the i th layer of the VGG network, I fuse represents the input processed by the main repair network, represents the output processed by the detail enhancement network. By calculating the covariance of features through the Gram matrix, an image with consistent generated style is constrained. The value of its style loss is: (3)

[0009] The above G I is the feature representation based on the Gram matrix. Finally, the total loss function of the detail enhancement network combines the reconstruction loss, perceptual loss and style loss, and are the weight coefficients of the pixel-level loss and style loss respectively: (4)

[0010] Through the above repair process, the main repair network initially restores the original structure and texture of the image, and the detail enhancement network further enriches the detail level, ensuring the clarity and naturalness of the repaired image.

[0011] Furthermore, in the field of image inpainting, there is a Mamba enhanced encoding module. The Mamba enhanced encoding module uses a 3×3 convolutional layer to extract local and global features of the input, and then sends these features to a bidirectional Mam2 gating module for scan-style feature extraction, and performs "width first and height second" and reverse "height first and width second" to further capture deep global dependency relationships and detailed information. The features are fused layer by layer in the encoder to gradually enrich multi-scale information, and the high-resolution features are refined and compressed through the Conv1 layer, and finally an optimized feature representation is generated.

[0012] Furthermore, the bidirectional Mam2 gating module includes Mam2 units for forward and backward propagation. The bidirectional Mam2 gating module projects and maps the features through a fully connected layer, and at the same time combines the gating mechanism to dynamically filter and enhance the features, so as to output high-quality fused features; in this module, the input features first pass through the Mam2 unit for forward propagation for convolution and linear transformation to extract preliminary expressive features; then, these features are passed to the Mam2 unit for backward propagation to further strengthen the modeling of global dependency relationships and the capture of detailed information; through the synergistic effect of bidirectional propagation, the module can more comprehensively mine the deep features of the input data and finally generate an optimized feature representation; Among them, forward propagation extracts the global context information of the input features through layer-by-layer convolution operations and feature enhancement, and captures important global dependency relationships through multi-layer linear transformation of the input features; Backward propagation processes the features in reverse order, captures additional context associations through reverse learning, so as to make up for the feature relationships that may be missed in forward propagation. The gating mechanism dynamically adjusts the weight distribution of the features through matrix multiplication and non-linear activation functions. Specifically, the input features are first reduced in dimension through linear projection, which reduces the computational complexity while retaining the main information: (5)

[0013] In the formula, W p is the weight matrix of linear projection, which is used to perform a linear transformation on the input feature X By multiplying with the input features, the adjustment of feature dimensions and the linear combination of information are realized; b p is the bias vector. Adding the bias vector in the linear transformation increases the expressive power of the model, so as to better fit the data. The projected feature X' is divided into multiple subspace features through the Split function X split for subsequent separate processing of the features in different subspaces: (6)

[0014] One of the subspace features X split First, add the bias, and then perform a non - linear adjustment through the Softplus function. The Softplus function maps the input value to a non - negative output value and has a smooth characteristic. Then, combined with the dynamic time step dt, adjust the feature weights to obtain the non - linear weighted subspace feature X b :[[]]END]] (7)

[0015] Divide into XBC Perform a one - dimensional convolution operation Conv1D on the subspace features, and then further refine the features through the LeakyReLU activation function to obtain the one - dimensional convolution refined features X, B, C , with a slope on the negative half - axis of LeakyReLU to avoid the problem of neuron death in the negative half - axis of the ReLU function and can better transmit gradients: (8) (9) (10)

[0016] Input the one - dimensional convolution refined features B , C into the SSD module. The SSD module models the context of the features by learning the non - linear relationship between the features and captures the dependency relationship between regions: (11)

[0017] After that, the output result of the SSD module y is added to the non - linear weighted subspace feature X b and the subspace feature Z to ensure the integrity of the input information: (12)

[0018] The fused features Z' are linearly projected, where W f is the weight matrix and b f is the bias vector, and remapped to the target space to generate the final output features: (13).

[0019] Furthermore, the Mamba enhanced encoding module also collaborates with the dynamic multi-scale semantic fusion module. The dynamic multi-scale semantic fusion module consists of three core parts: multi-scale feature extraction, adaptive dynamic weight allocation, and gated attention mechanism. Multi-scale convolution is used to generate query features and key features of different scales on texture features T s and geometric features F s to generate query features and key features of different scales. For the input feature map, queries are extracted from Q i and keys are extracted from K i , and the attention matrix for each scale T s is calculated, where Q i is the dimension of the vector. When calculating the attention score, dividing by F s is to scale the result, which helps to make the gradient more stable during training and avoid excessive gradient values. The formula is as follows: K i By calculating the cross-feature correlation matrix i , the dynamic multi-scale semantic fusion module introduces an adaptive dynamic weight allocation mechanism. Specifically, geometric features A i first extract global features through global average pooling d k , and the calculation formula is as follows: Subsequently, the dynamic multi-scale semantic fusion module processes the global feature (14) (15) (16)

[0020] through a feature weight generation function A i formed by two layers of 1×1 convolution and ReLU activation function to generate a set of values. These values are normalized through the Softmax operation to ensure that the sum of weights of different scales is 1, obtaining the dynamic weight F s : z The calculation formula is as follows: (17)

[0021] Subsequently, the dynamic multi-scale semantic fusion module processes the global feature h through a feature weight generation function z formed by two layers of 1×1 convolution and ReLU activation function to generate a set of values. These values are normalized through the Softmax operation to ensure that the sum of weights of different scales is 1, obtaining the dynamic weight w i : (18)

[0022] The dynamic weight mechanism can adaptively adjust features at different scales, enabling the dynamic multi-scale semantic fusion module to dynamically optimize feature fusion according to the characteristics of the input data and task requirements. After generating the dynamic weights, the module further uses the attention mechanism to fuse the features. By A i V weighting w i and summing the attention-enhanced features at each scale, the fused feature is generated F att : (19)

[0023] Among them, V i represents the eigenvalue extracted from F s The module further introduces learnable gating parameters g to dynamically adjust the intensity of the attention output. Finally, the fused feature is combined with the original geometric feature F s through a residual connection to obtain the final output F out : (20).

[0024] According to one aspect of the present invention, a method for digital restoration of Dunhuang murals based on the Mamba enhanced encoding feature fusion network is provided. Compared with the prior art, the present invention has the following beneficial effects: (1) The method for digital restoration of Dunhuang murals based on the Mamba enhanced encoding feature fusion network includes two sub-networks: the main restoration network and the detail enhancement network. The main restoration network completes the preliminary restoration task of the mural through the fusion of texture and geometric features, and adds a gating encoding module and a prior encoding module to solve the problems of artifacts and texture blur. The detail enhancement network performs secondary restoration on the basis of the main restoration network, further extracts details and introduces skip connections to achieve better results.

[0025] (2) Add the Mamba enhanced encoding module to the main restoration sub-network. The core functions of this module are as follows: First, use a 3×3 convolutional layer to capture the local details and global context information of the image, and gradually enrich the multi-scale features through layer-by-layer fusion. Combine a 1×1 convolution to further refine and compress the features, thereby generating high-resolution optimized features. Second, combine the bidirectional Mam2 gating module to capture global dependencies through forward and backward propagation, and at the same time accurately screen and strengthen key features through the dynamic gating mechanism, improving the depth and flexibility of feature expression. This design significantly improves the efficiency and quality of feature extraction, providing strong support for completing complex mural restoration tasks.

[0026] (3) Add the Dynamic Multi-Scale Semantic Fusion Module (DMSFM) to the main restoration sub-network. Through multi-scale feature extraction, adaptive dynamic weight allocation, and gated attention mechanism, it realizes the efficient fusion of texture and geometric features. This module can adaptively adjust the feature weights at different scales, and use the attention mechanism to strengthen key features and suppress redundant information, providing a more flexible and adaptable technical solution for semantic fusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 Schematic diagram of the method network of the present invention; Figure 2 The Mamba enhanced encoding module of the present invention; Figure 3 The Dynamic Multi-Scale Semantic Fusion Module of the present invention; Figure 4 Comparison of the method of the present invention with the method without the Mamba enhanced encoding module and the Dynamic Multi-Scale Semantic Fusion Module; Figure 5 Qualitative comparison of the method of the present invention with six other methods on narrative murals; Figure 6 Qualitative comparison of the method of the present invention with six other methods on Buddha statue murals; Figure 7 Qualitative comparison of the method of the present invention with six other methods on caisson murals. DETAILED DESCRIPTION OF THE INVENTION

[0028] In order to make the technical means, creative features, achieved purposes and effects of the present invention easy to understand, the present invention will be further described below in conjunction with specific embodiments.

[0029] This is as Figures 1 - 7As shown in the figure, the innovative Dunhuang mural digital restoration method based on the Mamba-enhanced coding feature fusion network of the present invention consists of two sub-networks: the main restoration network and the detail enhancement network. The main restoration network uses dual encoders to extract the features of the damaged mural and the mural line drawing respectively, learn the mural texture and geometric prior information, and fill in the missing area content. This network effectively establishes the dependency relationship between the restored mural and the original mural, overcoming the deficiencies of the poor generation quality of VAE and the unstable training of GAN.

[0030] The core design of the present invention is to add a gating coding module and a prior coding module in the main restoration network to solve the problems of artifacts and texture blur. Add the MAMBA enhanced coding module, introduce a scanning mechanism, reduce the computational overhead, use a 3×3 convolutional layer to capture the local details and global context information of the image across regions, extract the correlation, and ensure the rapid convergence of the training process. Add a dynamic multi-scale semantic fusion module (Dynamic Multi-Scale Semantic Fusion Module), and through multi-scale feature extraction, adaptive dynamic weight allocation, and gating attention mechanism, realize the efficient fusion of texture and geometric features. After using the main restoration network to complete the preliminary image restoration task, then use the detail enhancement network to perform detail restoration of the image, realizing the careful capture of geometric features, the accurate restoration of texture, and the high-precision reconstruction of the damaged part.

[0031] Implementation steps and key algorithms of the technical solution of the present invention: Figure 1 The figure shows the technical solution network diagram of the method of the present invention. This method consists of a primary restoration sub-network (Primary Restoration Network) and a refinement enhancement sub-network (Refinement Enhancement Network). The two sub-networks focus on the preliminary restoration and detail optimization of the damaged image respectively to improve the image restoration effect and overall quality.

[0032] Step 1: Use the dual encoders of the main restoration network to extract the features of the damaged mural and the mural line drawing, learn the mural texture and geometric prior information, and fill in the missing area content.

[0033] This network effectively constrains the latent space of the image through the prior encoder module, and uses the local texture details extracted from the damaged mural and the global structure information obtained from the mural line drawing to preliminarily fill in the missing area of the image. It can largely restore the geometric shape and color texture of the missing part of the image, and the restored image achieves a high visual consistency in terms of overall contour and local details.

[0034] To address the common artifacts and texture blurring issues in image inpainting, the main inpainting network introduces gated convolution, which adaptively controls the effective region of the convolution kernel, filters out irrelevant information, and enables the network to focus on filling in the missing regions, thereby reducing artifacts and enhancing the clarity of texture details to ensure that the restored image is realistic and natural. Its L main The loss function is designed as follows: (1)

[0035] In the formula, I out1 is the output of the main inpainting network, I g is the real image, M is the mask, and are the balance coefficients, and represent the KL loss functions of geometric features and texture features respectively, is the adversarial loss.

[0036] Step 2: Based on the image initially generated by the main inpainting network, the detail enhancement network further repairs the details of the image.

[0037] For the subtle changes in complex textures and edges, the network designs skip connections to enhance the ability to capture details. This network can capture multi-level global and local information to obtain a finer inpainting effect. In addition, to further improve the visual quality of the inpainted image and maintain style consistency, the detail enhancement network introduces perceptual loss and style loss . The perceptual loss extracts feature maps through a pre-trained VGG network and compares the inpainted image with the real image from the perspective of high-level features to ensure their visual consistency in semantics: (2)

[0038] In the formula, F i represents the feature extraction result of the i th layer of the VGG network, I fuse represents the input processed by the main inpainting network, represents the output processed by the detail enhancement network. To ensure that the inpainted image matches the real image in terms of overall visual style, we calculate the covariance of features through the Gram matrix to constrain the generation of images with consistent styles, and the value of its style loss is: (3)

[0039] The aboveG I is the feature representation based on the Gram matrix. Finally, the total loss function of the detail enhancement network combines the reconstruction loss, perceptual loss, and style loss, and are the weight coefficients of the pixel-level loss and style loss respectively: (4)

[0040] Through the above repair process, the dual network proposed by the present invention has obtained satisfactory results in both the overall repair effect and detail processing. The main repair network initially restores the original structure and texture of the image, and the detail enhancement network further enriches the detail level to ensure the clarity and naturalness of the repaired image.

[0041] The core network module and key algorithms are: In the field of image repair, feature extraction and fusion are the key links to improve the model performance. The Mamba enhanced encoding module is designed. This module uses a 3×3 convolutional layer to extract local and global features of the input, and then sends these features into the bidirectional Mam2 gating module for scanning feature extraction, and performs "width first and height second" and reverse "height first and width second" to further capture deep global dependency relationships and detail information. The features are fused layer by layer in the encoder to gradually enrich multi-scale information, and the high-resolution features are refined and compressed through the Conv1 layer to finally generate an optimized feature representation. The structure of the Mamba enhanced encoding module is as Figure 2 shown.

[0042] Figure 2 The bidirectional Mam2 gating module in

[0043] Among them, the forward propagation (Mam2 forward) extracts the global context information of the input features through layer-by-layer convolution operations and feature enhancement. It captures important global dependencies by performing multi-layer linear transformations on the input features. The backward propagation (Mam2 backward) processes the features in reverse order and captures additional context associations through reverse-order learning, thereby making up for the feature relationships that may be missed in the forward propagation. This unit can supplement the global dependencies and significantly improve the overall expression ability of the network. The gating mechanism dynamically adjusts the weight distribution of the features through matrix multiplication and non-linear activation functions. It can effectively control which features should be strengthened or suppressed, thereby improving the flexibility of feature selection and ensuring that key information is fully utilized.

[0044] The present invention innovatively proposes a Mam2 unit for feature dynamic enhancement and local context modeling in image inpainting tasks. This module realizes the efficient extraction and fusion of features by combining dynamic adjustment, local feature modeling, and skip connection mechanisms. Its core design includes linear projection, Softplus bias adjustment, structured state space duality (SSD), convolutional activation, and skip connection operations. Specifically, the input features are first reduced in dimension through linear projection, which retains the main information while reducing the computational complexity: (5)

[0045] In the formula, W p is the weight matrix of the linear projection, which is used to perform a linear transformation on the input feature X By multiplying with the input feature, it realizes the adjustment of the feature dimension and the linear combination of information; b p is the bias vector. Adding a bias vector in the linear transformation can increase the expression ability of the model, making the linear transformation not only a scaling and translation of the input feature, but also a certain offset, so as to better fit the data. The projected feature X' is divided into multiple subspace features X split through the Split function for subsequent separate processing of the features in different subspaces: (6)

[0046] One of the subspace features X split is first added with the bias Bias, and then non-linearly adjusted through the Softplus function. The Softplus function can map the input value to a non-negative output value and has a smooth characteristic. Then, combined with the dynamic time step dt to adjust the feature weight, the non-linear weighted subspace feature X b is obtained: (7)

[0047] Perform a one-dimensional convolutional operation Conv1D on the subspace features divided into XBC , and then further refine the features through the LeakyReLU activation function to obtain one-dimensional convolutional refined features X, B, C . LeakyReLU is an activation function that has a small slope on the negative half-axis, avoiding the problem of neuron death in the ReLU function on the negative half-axis and being able to better transmit gradients: (8) (9) (10)

[0048] Input the one-dimensional convolutional refined features B , C into the SSD module, which models the context of the features by learning the non-linear relationships between the features and captures the dependencies between regions: (11)

[0049] After that, add the output result y of the SSD module to the non-linear weighted subspace features X b and the subspace features Z to ensure the integrity of the input information: (12)

[0050] The fused features Z' are linearly projected, where W f is the weight matrix and b f is the bias vector, and remapped to the target space to generate the final output features: (13)

[0051] The Mamba enhanced encoding module significantly improves the efficiency and quality of feature extraction, providing strong support for complex repair tasks. To further enhance the model's adaptability in handling complex tasks, a Dynamic Multi-Scale Semantic Fusion Module (DMSFM) is proposed. This module further enhances the model's performance by efficiently fusing texture features and geometric features and working in cooperation with the Mamba enhanced encoding module. Its structure is as Figure 3 shown.

[0052] This module consists of three core parts: multi-scale feature extraction, adaptive dynamic weight allocation, and gated attention mechanism. Through the close combination of these three parts, efficient feature fusion and precise expression are achieved. The design of the module starts from multi-scale feature modeling with the goal of realizing the complementarity of global and local information. Multi-scale convolutions are used to generate query features T s and key features F s at different scales on texture features Q i and geometric features K i . For the input feature map (texture features T s and geometric features F s ), we extract queries T s from Q i , extract keys F s from K i , and calculate the attention matrix i at each scale A i , where d k is the dimension of the vector. When calculating the attention scores, dividing by is to scale the results, which helps to make the gradients more stable during training and avoid overly large gradient values. The formula is as follows: (14) (15) (16)

[0053] By calculating the cross-feature correlation matrix A i , the module can effectively capture local detail information and global structural associations simultaneously. Parallel computing at multiple scales enables the module to extract features from different receptive fields and comprehensively capture the semantic information of multi-scale features. To further improve the efficiency of feature fusion, the module introduces an adaptive dynamic weight allocation mechanism to address the differences in importance among features at different scales. Specifically, geometric features F s first extract global features z through global average pooling (GAP). The calculation formula is as follows: (17)

[0054] Subsequently, the module generates a set of values by processing the global features through a feature weight generation function composed of two layers of 1×1 convolution and ReLU activation functions h for the global features z to generate a set of values. These values are normalized through a Softmax operation to ensure that the sum of weights at different scales is 1, resulting in dynamic weights w i : (18)

[0055] The dynamic weight mechanism can adaptively adjust features at different scales, enabling the module to dynamically optimize feature fusion according to the characteristics of the input data and the requirements of the task. This design improves the adaptability and feature expression ability of the module. Especially in complex task scenarios, it can more flexibly process multi-dimensional features. After generating the dynamic weights, the module further uses the attention mechanism to fuse the features. By weighted summing the attention-enhanced features at each scale A i V by the weights w i a fused feature is generated F att : (19)

[0056] where V i represents the feature value extracted from F s . In this way, the module not only synthesizes multi-scale semantic information but also highlights the important feature expressions for mural restoration. On this basis, the module introduces learnable gating parameters g to dynamically adjust the intensity of the attention output. Finally, the fused feature is combined with the original geometric feature F s through a residual connection to obtain the final output F out : (20)

[0057] The design of DMSFM is innovative. First, the multi-scale feature extraction part comprehensively models local details and global structures by capturing features in parallel at different scales. Second, the adaptive dynamic weight allocation mechanism enables the module to meet different task requirements by flexibly adjusting the weights of feature fusion. Finally, the introduction of the gated attention mechanism enhances the module's ability to express important features while suppressing the interference of irrelevant features. Through the organic combination of multi-scale feature extraction, adaptive weight allocation, and gated attention mechanism, DMSFM achieves the efficient fusion of texture features and geometric features, significantly improving the flexibility and adaptability of feature expression.

[0058] The combination of the Mamba enhanced encoding module and DMSFM can significantly improve the effect of the image inpainting task. The multi-scale features extracted by the Mamba module provide richer basic information for DMSFM, enabling the latter to optimize the fusion of texture and geometric features from a more comprehensive perspective. Specifically, the local and global features output by the Mamba enhanced encoding module can effectively serve as part of the texture feature TS and geometric feature FS of DMSFM after appropriate processing, thus providing more accurate guidance for feature fusion. The synergistic effect of the two not only enhances the depth and breadth of feature extraction but also improves the efficiency and quality of feature fusion. Through this close combination, the Mamba enhanced encoding module and DMSFM jointly promote the development of image inpainting technology, enabling complex inpainting tasks to be completed more efficiently and accurately.

[0059] The Dunhuang Mogao Grottoes mural dataset used in the experiments of this invention contains mural images from more than 200 grottoes, totaling 4,564 images. The themes are rich and diverse, including Buddha mural paintings, portrait mural paintings, narrative mural paintings, caisson ceiling mural paintings, etc. We use data augmentation methods such as cropping and random rotation to expand the collected mural images to 23,486 images. The dataset also includes the mural line drawings automatically generated by the DexiNed method, adding additional geometric information. The actual damaged murals are hand-drawn by Dunhuang scholars, containing prior knowledge and experience. Three types of murals, namely narrative, Buddha, and caisson ceiling, are randomly selected from the Dunhuang mural dataset for experimental analysis. The validation mask uses an irregular random mask, and the mask sizes are different for each group of tests.

[0060] In this experiment, the method of the present invention was compared with three classic methods, PIC, RFR, and Repaint, and three state-of-the-art methods, STNet, TSBGNet, and StrDiffusion. As described above, we conducted comparative experiments and analyses on narrative murals, Buddha statues murals, and caisson ceiling murals in the Dunhuang mural dataset. The evaluation metrics used included mean squared error (MSE), peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), contrast improvement index (CII), natural image quality evaluation (NIQE), and Frechet distance (FID). Each selected evaluation metric focuses on different aspects of image quality evaluation. For example, MSE and PSNR mainly evaluate pixel-level differences, while SSIM, CII, NIQE, and FID pay more attention to perceptual quality, structure, contrast, and the similarity and diversity between the generated image and the real image.

[0061] Table 1: Comparative test results for narrative murals in the Dunhuang mural dataset.

[0062]

[0063] Table 2: Comparative test results for Buddha statues murals in the Dunhuang mural dataset.

[0064]

[0065] Table 3: Comparative test results for caisson ceiling murals in the Dunhuang mural dataset

[0066] It can be clearly seen from Tables 1 - 3 that the method of the present invention has obvious advantages in image restoration. Specifically, the overall restoration effect of the method of the present invention on narrative murals is the best, and it is significantly better than other methods in terms of pixel-level differences. The method of the present invention also has good performance in terms of improving image structure, perceptual quality, and contrast. For Buddha statues murals, the method of the present invention is similar to the original image in terms of perceptual quality, with almost no pixel-level error during the image restoration process and less image distortion. Although the restoration effect of the method of the present invention on caisson ceiling murals does not perform optimally in all evaluation metrics, it still shows strong comprehensive competitiveness compared with other algorithms.

[0067] Such as Figure 4As shown in the figure, the Mamba Enhanced Coding Module and the Dynamic Multiscale Semantic Fusion Module (DMSFM) each play an important role in the image restoration task. First, the Mamba Enhanced Coding Module effectively improves the efficiency and quality of feature extraction through the scanning mechanism and the bidirectional Mam2 gating module, reduces the artifacts after restoration, and makes the mural restoration more complete and natural. Subsequently, the DMSFM successfully fuses texture and geometric features through multi-scale feature extraction and dynamic weight allocation mechanism, improves the line restoration effect, and avoids contextual semantic errors. When DMSFM is not used, although the approximate shape can be restored, the details are poorly processed and redundant lines are prone to appear. After using DMSFM, the restoration results are more accurate and the details are better. Combining these two modules, the Mamba Enhanced Coding Module and DMSFM complement each other, improve the overall effect of image restoration, and enhance the accuracy and efficiency of the restoration task.

[0068] like Figure 5 As shown in the figure, for narrative murals, PIC can restore basic structural features well, but has obvious defects in detail and texture restoration, especially in the continuity of the background and the presentation of subtle changes; RFR performs well in processing simple textures, but has poor restoration effects on complex scenes, especially when processing small objects in the image, it is easy to be blurred; although RePaint has a good performance in detail restoration, it has certain artifacts in the restoration of facial expressions and body postures of characters. STNet and TSBGNet can fill complex blank areas well, and are better than the previous methods in detail restoration and semantic consistency, but they are still lacking in mural color completion. The image obtained by StrDiffusion also shows more artifacts, especially in the restoration of details in the edge area, there is a certain degree of image distortion. In comparison, the method of the present invention performs best in narrative mural restoration. It can not only accurately restore details, but also maintain the overall semantic consistency of the image, and has the best restoration effect on complex backgrounds and character details.

[0069] like Figure 6As shown, for the Buddha statue mural with relatively simple structure, PIC and RFR can restore the basic shape and color of the Buddha statue. However, RFR still lacks in restoring the fine textures of the Buddha's face and clothing, and produces a large area of artifacts. STNet and TSBGNet can better restore the facial details and the folds of the clothing. Nevertheless, there are still some image areas showing certain color distortion or local blurring. The repair effect of StrDiffusion on this dataset is relatively average. Especially when dealing with objects with strong metallic and shiny textures, unnatural reflections and distortions occur. In comparison, the method of the present invention performs better in restoring the facial details, clothing textures and overall structure of the Buddha statue, and can accurately restore the fine features of the Buddha statue. Especially in the aspect of complex detail repair, it shows obvious advantages compared with other repair algorithms and demonstrates stronger repair capabilities.

[0070] Figure 7 As shown, for the caisson mural containing a large number of geometric shapes and patterns, PIC and RFR are difficult to restore the fine patterns on this dataset. Especially in terms of the continuity and symmetry of the patterns, incoherent results are easily produced. Although RePaint performs well in restoring details, it has a poor effect when dealing with the complex texture patterns of the caisson, and geometric shape distortions and discolorations often occur. STNet and TSBGNet can better restore complex patterns, especially with improvements in symmetry and detail preservation. However, there are still certain defects in dealing with relatively fine local areas. StrDiffusion performs poorly on this dataset. Especially when dealing with highly complex patterns, the continuity of the patterns is easily lost. In comparison, the method of the present invention shows obvious advantages in the repair of the caisson mural. It can not only accurately restore complex geometric patterns, but also maintain the symmetry and continuity of the details of the patterns, making the repair result more natural and realistic.

[0071] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.

[0072] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A digital restoration method for Dunhuang murals based on Mamba enhanced coding feature fusion network, characterized by: The Dunhuang mural restoration network consists of two sub-networks: a main restoration network and a detail enhancement network. The main restoration network uses a dual encoder to extract the features of the defective murals and the mural line drawings respectively. The encoder for extracting the features of the mural line drawings is composed of a gated coding module and a priori coding module. The encoder for extracting the features of the defective murals adds a MAMBA enhancement coding module on this basis, introduces a scanning mechanism, and adds a dynamic multi-scale semantic fusion module. Through multi-scale feature extraction, adaptive dynamic weight allocation and gated attention mechanism, the main restoration network is used to complete the preliminary image restoration task, and then the detail enhancement network is used to perform image detail restoration. The specific contents are as follows: Step 1: Use the main restoration network dual encoder to extract the features of the defective murals and mural line drawings, learn the mural texture and geometric prior information, and fill in the missing area content; The main restoration network effectively constrains the latent space of the image through the prior encoder module, and uses the local texture details extracted from the defective murals and the global structural information obtained from the mural line drawing to preliminarily fill in the missing areas of the image; The main repair network introduces gated convolution to adaptively control the area of ​​action of the convolution kernel, filter out irrelevant information, and make the network focus on filling in the missing area. L main The loss function is designed as follows: (1) In the formula, I out1 is the output of the main repair network, I g is a real image, M It's a mask. and is the balance coefficient, and Represent the KL loss function of geometric features and texture features respectively, To combat losses; Step 2: Based on the image initially generated by the main restoration network, the detail enhancement network is used to further restore the image details; In order to deal with the subtle changes in complex textures and edges, the main restoration network designs jump connections to enhance the ability to capture details. In addition, the detail enhancement network introduces perceptual loss and style loss , the perceptual loss extracts feature maps through the pre-trained VGG network, compares the restored image with the real image from the perspective of high-level features, and ensures the semantic visual consistency between the two: (2) In the formula, F i Represents the VGG network i The feature extraction results of the layer, I fuse represents the input after being processed by the main inpainting network, Represents the output of the detail enhancement network processing. The covariance of the features is calculated through the Gram matrix to constrain the generation of images with consistent style. The value of the style loss is: (3) Above G I is the feature representation based on the Gram matrix. Finally, the total loss function of the detail enhancement network is Combining reconstruction loss, perception loss and style loss, and They are the weight coefficients for pixel-level loss and style loss respectively: (4) Through the above restoration process, the main restoration network preliminarily restores the original structure and texture of the image, and the detail enhancement network further enriches the detail level to ensure the clarity and naturalness of the restored image.

2. The method for digital restoration of Dunhuang murals based on the Mamba enhanced coding feature fusion network according to claim 1 is characterized in that: In the field of image restoration, a Mamba enhanced coding module is provided. The Mamba enhanced coding module uses a 3×3 convolutional layer to extract local and global features of the input, and then sends these features to the bidirectional Mam2 gating module for scanning feature extraction, and performs "width first, height later" and the reverse "height first, width later" to further capture deep global dependencies and detail information. The features are fused layer by layer in the encoder to gradually enrich multi-scale information, and the high-resolution features are refined and compressed through the Conv1 layer to finally generate an optimized feature representation.

3. The method for digital restoration of Dunhuang murals based on Mamba enhanced coding feature fusion network according to claim 2 is characterized in that: The bidirectional Mam2 gating module includes Mam2 units for forward and backward propagation. The bidirectional Mam2 gating module projects and maps features through a fully connected layer, and dynamically filters and enhances features in combination with a gating mechanism. In the bidirectional Mam2 gating, the input features are first convolved and linearly transformed by the Mam2 unit for forward propagation to extract preliminary expression features; then, these features are passed to the Mam2 unit for backward propagation to further strengthen the modeling of global dependencies and the capture of detailed information; through the synergistic effect of bidirectional propagation, the module can more comprehensively mine the deep features of the input data and ultimately generate an optimized feature representation; Among them, the forward propagation extracts the global context information of the input features through layer-by-layer convolution operations and feature enhancement, and captures important global dependencies by performing multi-layer linear transformations on the input features; Backward propagation processes features in reverse order and captures additional contextual associations through reverse learning, thereby making up for the feature relationships that may be missed in forward propagation. The gating mechanism dynamically adjusts the weight distribution of features through matrix multiplication and nonlinear activation functions. Specifically, the input features are first reduced in dimension through linear projection to reduce the computational complexity while retaining the main information: (5) In the formula, W p is the weight matrix of the linear projection, which is used to X Perform linear transformation and adjust the feature dimension and linearly combine the information by multiplying the input features. b p The bias vector is added to the linear transformation to increase the expressiveness of the model, so as to better fit the data. The features after projection X' Divided into multiple subspace features through the Split function X split So that the features of different subspaces can be processed separately later: (6) One of the subspace features X split First add the bias Bias, then use the Softplus function for nonlinear adjustment. The Softplus function maps the input value to a non-negative output value and has a smooth characteristic. Then, the dynamic time step dt is combined to adjust the feature weight to obtain the nonlinear weighted subspace feature. X b : (7) Divide into XBC The subspace features of the one-dimensional convolution operation Conv1D are then further refined through the LeakyReLU activation function to obtain the one-dimensional convolution refinement features. X, B, C , a slope is set on the negative semi-axis of LeakyReLU to avoid the problem of ReLU function neuron death on the negative semi-axis, and to better transfer the gradient: (8) (9) (10) Refine the features using one-dimensional convolution B , C The input is sent to the SSD module, which performs context modeling on the features by learning the nonlinear relationship between the features and capturing the dependencies between regions: (11) After that, the output results of the SSD module will be y With nonlinear weighted subspace features X b And the subspace characteristics Z Add together to ensure the integrity of the input information: (12) The fused features Z' By linear projection, W f is the weight matrix, b f The bias vector is remapped to the target space to generate the final output features: (13)。 4. The method for digital restoration of Dunhuang murals based on Mamba enhanced coding feature fusion network according to claim 1 is characterized in that: The Mamba enhanced coding module also cooperates with the dynamic multi-scale semantic fusion module, which includes three core parts: multi-scale feature extraction, adaptive dynamic weight allocation and gated attention mechanism. Multi-scale convolution is used to extract texture features. T s and geometric features F s Generate query features of different scales Q i and key features K i , for the input feature map, from T s Extract query from Q i ,from F s Extract the key K i , and calculate each scale i The attention matrix A i ,in d k is the dimension of the vector. When calculating the attention score, divide by This is to scale the results, which helps to make the gradient more stable during training and avoid excessive gradient values. The formula is as follows: (14) (15) (16) By calculating the cross-feature correlation matrix A i The dynamic multi-scale semantic fusion module introduces an adaptive dynamic weight allocation mechanism. Specifically, the geometric features F s First, global features are extracted by global average pooling z , the calculation formula is as follows: (17) Subsequently, the Mamba enhanced coding module generates a feature weight function consisting of two layers of 1×1 convolution and ReLU activation function. h For global features z Processing is performed to generate a set of values, which are normalized by the Softmax operation to ensure that the sum of weights at different scales is 1, and the dynamic weights are obtained. w i : (18) The dynamic weight mechanism can adaptively adjust features of different scales, so that the dynamic multi-scale semantic fusion module can dynamically optimize feature fusion according to the characteristics of the input data and task requirements. After completing the dynamic weight generation, the dynamic multi-scale semantic fusion module further uses the attention mechanism to fuse the features. By focusing on each scale to enhance the features A i V By weight w i Weighted summation to generate fusion features F att : (19) in, V i Indicates from F s The module further introduces learnable gating parameters g , which is used to dynamically adjust the intensity of the attention output. The final fusion feature is connected with the original geometric features through residual connection F s Combined to get the final output F out : (20)。