Depth feature multi-dimensional perception guided Thangka image reflection iteration removal method
Through the iterative prediction network of deep features multi-dimensional perception guided, combined with the strategies of frequency domain information separation and mask guidance, the problem of reflection interference in Regon Thangka's digitization is solved, efficient and accurate reflection removal is achieved, and image quality and scientific research value is improved.
Patent Information
- Application Number
- CN202510359032.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-24
AI Technical Summary
During the digitalization process of Regontangka, the reflection interference caused by the glass cover seriously reduced the image quality, and the existing technology is difficult to effectively remove reflections, resulting in a decline in image visual quality and hindering scientific research.
The iterative prediction network with deep features multi-dimensional perception guidance is adopted, and reflection removal is achieved by combining the depth multi-dimensional features of the reflection layer, the transmission layer and the original input image, combined with the image repair strategy of frequency domain information separation and mask guidance.
It significantly improves the robustness and accuracy of reflection removal, can effectively remove the highlight reflection of Thangka images, maintaining the visual quality of the image and the needs of scientific research.
Smart Images

Figure CN120198331A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image reflection removal, and particularly to an image reflection removal method based on deep learning, specifically a method for removing reflections from Thangka images of Regong Art Background Art
[0002] In practical scenarios such as the digitalization of museum collections and the shooting of commercial window advertisements, the reflection interference caused by glass media seriously reduces the image quality. In the reflection removal task, the line features of ordinary images are mostly natural edges, with obvious pixels, and the colors are mainly linearly superimposed by the RGB three channels. Their textures are random and mainly dominated by spatial low frequencies. However, the images of art treasures are mostly lines of different thicknesses and sizes, with multiple layers of pigment shading. And different from the diffuse reflection in nature, the exhibits in glass covers are mostly directly irradiated by spotlights and contain multiple layers of composite reflections. Thangka images of Regong Art have fine brushwork lines with millimeter-level integrity, special color gamuts of mineral pigments, metal reflections in the gold powder areas, and diffuse reflections of coarse mineral particles. Most Thangka artworks are stored in museums, temples, and personal exhibition halls, and are covered physically with glass covers to prevent dust, touch, insects, etc. However, during the process of photographing and digitalizing artworks, the light reflected by the glass and the light transmitted by the objects behind the glass will both be captured by the camera. Clear details and compositions of the image cannot be observed in the reflection area, which will undoubtedly greatly reduce the visual quality of the image and hinder subsequent visual applications and scientific research. Figure 1 As shown in (a) in Figure 1 the Thangka image shown in (b) in
[0003] During the reflection removal process, the observed value image Ι with reflection is usually regarded as the weighted sum of the transmission layer T and the reflection layer R, that is where W is the alpha blending mask, Denotes element-wise multiplication. The research goal of single-image reflection removal is to recover the transmission layer T from the input image I. However, due to the inherent uncertainty of single-image reflection removal itself, it is difficult for existing models to directly learn the accurate mapping relationship of the transmission layer T from the input image I. During the digital acquisition process of Thangka paintings in Regong area, the gold foil is prone to specular highlights under a fixed light source, and local strong reflections may obscure the details of the embossed relief. In addition, the coarse granularity of mineral pigments can cause the coupling effect of scattering, metallic reflection, and the periodic reflection caused by the weaving texture of silk. Currently, most studies mostly adopt cascaded deep models to reduce the uncertainty of transmission layer estimation. For example, CEILNet first estimates the edge map and then estimates the transmission map. BDN, IBCLN, and RAGNet predict the reflection layer in the early stage and predict the transmission layer in the subsequent stage of the network model. However, for art treasures such as Thangka paintings in Regong area with intricate lines and colorful colors, the pixel deviation in edge preservation is too large, the color is distorted based on the layering and color display effect of pigments, and the structural continuity of religious symbols is also difficult to constrain, making it impossible to achieve effective reflection removal and image reconstruction.
[0004] Traditional reflection removal methods rely on image prior knowledge to reduce the impact of reflection, such as using features like gradient sparsity, smoothness, and ghosting. For example, the relative smoothness prior model, based on the assumption of reflection blurriness, penalizes large gradients to reduce reflection. Using the "ghost" clue combined with Gaussian mixture model regularization can automatically suppress reflection. Some scholars introduce user annotations to guide the hierarchical separation of reflection. There are also methods that set Laplacian fidelity and gradient sparsity as optimization objectives and use wavelet transform regularization to separate ghosting, as well as separate reflection and transmission based on multi-scale depth of field analysis. However, the thresholds of these methods are easily interfered by noise, are based on specific assumptions, and are slow to process, making it difficult to meet the requirements of efficiency and accuracy.
[0005] With the development of deep learning, the performance of learning-based single-image reflection removal methods has been significantly improved. For example, CEILNet uses a two-stage network to first predict the edge map of the transmission map and then separate the reflection and reconstruct the transmission layer through edge guidance. ERRNet uses high-level context features to reduce uncertainty, introduces unaligned real-world paired images, and constrains the output with an invariant loss. Some scholars adjust intensity attenuation, apply random wavelets to adapt to different reflection scenarios, and also design loss functions to utilize reflection characteristics. BDNet incorporates the reflection layer into guiding the recovery of the transmission layer. Some scholars develop a two-stream framework to process features in parallel for efficient reflection separation. The cyclic network-based scheme emphasizes strong reflection boundaries to improve the results. RAGNet uses a reflection-aware guidance module to guide the prediction of the transmission layer. However, these methods do not fully utilize the features of each dimension of the input image, cannot mine the fine features of Thangka images, and are difficult to remove the high-brightness reflections in Thangka images.
[0006] In view of the problem of reflection removal in the current digitalization process of Regong Thangka, there is an urgent need for an effective method for reflection removal to effectively digitally protect and inherit Thangka. The present invention not only utilizes the information of the reflection layer, but also mines the initially predicted transmission layer and the information of the original input layer, providing richer features for the model. At the same time, an iterative refinement strategy is adopted to achieve more accurate and stable performance in the reflection removal task. Summary of the Invention
[0007] To solve the problem of image quality degradation caused by glass cover reflection in the digitalization process of Regong Thangka in the prior art, the object of the present invention is to provide a method for iteratively removing reflections from Thangka images guided by multi-dimensional perception of deep features. The present invention significantly improves the robustness of reflection removal by jointly utilizing the deep multi-dimensional features of the initially predicted reflection layer, transmission layer, and the original input image, combined with the frequency domain information separation and mask-guided image restoration strategy.
[0008] To achieve the above object, the technical solution adopted by the present invention is: a method for iteratively removing reflections from Thangka images guided by multi-dimensional perception of deep features, which is implemented based on an iterative prediction network for reflection perception guidance in multi-dimensions of deep features, and specifically includes the following steps:
[0009] Step 1: Generate a reflection layer prior through a reflection feature enhancement module: Optimize the reflection layer prediction by combining the U-Net model with a squeeze-and-excitation module and a residual mechanism;
[0010] Step 2: Construct a multi-dimensional feature fusion mechanism to optimize the transmission layer: Based on a deep feature pyramid network, refine the capture of Thangka texture features through a spatial feature interaction block and a deep feature extraction block, and fuse high and low frequency features and mask constraints to achieve refined restoration of the transmission layer.
[0011] As a further improvement of the present invention, the specific steps of Step 1 include the following:
[0012] Step 1.1: The U-net model takes the original image Ι as the input to predict the reflection layer R';
[0013] Step 1.2: Add a squeeze-and-excitation module SE-Block behind each layer of the U-net model to reshape the weights of the feature map according to the global average pooling result of the input feature map, perform feature re-weighting on the highlighted areas of Thangka, and impose constraints based on the non-uniform specular reflection characteristics of the gold powder painting areas; Expand the image receptive field through dilated convolution to capture feature information at different scales and long-term dependency information in the sequence, and jointly retain the line texture with the residual structure of ResNet to transmit reflection features during the iteration process.
[0014] As a further improvement of the present invention, in step 2, the spatial feature interaction block includes a token mixer and two layers of linear layers connected in sequence, and there is a skip connection between the input end of the token mixer and the output end of the last layer of linear layer.
[0015] As a further improvement of the present invention, in step 2, the depth feature extraction block includes two layers of token mixers and two layers of linear layers connected in sequence, and skip connections are used between the input ends and the output ends of the two layers of token mixers, and there is a skip connection between the input end of the first layer of token mixer and the output end of the last layer of linear layer.
[0016] As a further improvement of the present invention, step 2 is specifically as follows:
[0017] In each layer of the iteration process, the iterative prediction network uses the original image Ι, the preliminary predicted reflection layer R', and the preliminary predicted transmission layer T' as the original input and predicts the final transmission layer T, where T' = I - R; R' and T' are used to extract features through the VGG model to obtain F R and F T , Ι obtains F through the depth pyramid network DFPN I ;
[0018] The frequency domain feature fusion module FFB uses octave convolution to extract the high-frequency branch and the low-frequency branch from the ordinary features, that is, the glass interference fringes and the image body texture, effectively maintaining the continuity of the millimeter-level meticulous brushstrokes, and then deeply fusing them with the original features to form a richer and more effective feature representation, and feeding it into the next layer of the iterative prediction network, that is, the multi-dimensional feature perception guidance module MFPG;
[0019] During the operation of the iterative prediction network, the input features are processed by the decoder to obtain the decoded feature F dec , on this basis, addition and subtraction operations are performed using the observed feature, the reflection feature, and the transmission feature to generate a preliminary feature result; subsequently, through the further processing of the channel attention block and the spatial attention block CBAM, the edge gradient is enhanced in the cotton lint area, and the differential feature F diff is obtained; in addition, the iterative prediction network uses F I , F R and F dec The result after concat is used as the input, and through the linear transformation of two 1×1 convolutional layers and the non-linear mapping of the Sigmoid activation function, the mask M is obtained based on the brightness threshold.
[0020] As a further improvement of the present invention, the loss function of the iterative prediction network includes a mask loss, a reconstruction loss, a perceptual loss, an exclusion loss, an adversarial loss, and a color loss, where:
[0021] The mask loss is:
[0022]
[0023] where M diff is the mask of F diff , i represents the i-th layer, ‖·‖1 represents the L1 norm, and [] represents the conditions to be satisfied. is the threshold for dividing the region with heavier reflection, and ξ is the threshold for dividing the region with lighter reflection;
[0024] The reconstruction loss is used to minimize the difference between the network output images T', R' and the corresponding ground truths T, R. The reconstruction loss is:
[0025]
[0026] Based on the pre-trained VGG-19 model, minimize the L1 difference between φ(T'), φ(R') and φ(T), φ(R) in the selected feature layers, where l is the index of the conv1_2, conv2_2, conv3_2, conv4_2 and conv5_2 layers, and the weight κ l is used to balance different layers; then the perceptual loss is:
[0027]
[0028] The exclusion loss is:
[0029]
[0030] where, λ T and λ R are normalization factors, and are the gradients of T and R, and ‖·‖ F is the Frobenius norm; T ↓n and R ↓n represent n downsampled versions of T and R, where T ↓0 and R ↓0 are the original inputs; during training, set N = 2.
[0031] In the adversarial loss, the entire network is used as the generator G, and a 4-layer discriminator D is constructed. Its parameters are updated by: The parameters of the generator G are optimized as:
[0032] In the color loss, the difference between the output image T' and the corresponding ground truth T is minimized pixel by pixel for the R, G, B channels, which is expressed as follows:
[0033]
[0034] Where \(N = n\times h\times w\), \(n\) is the batch number, \(h\) is the height of the image, \(w\) is the width of the image, \(c\in\{0,1,2\}\) is the RGB channel, \(i\in\{1,2,\ldots,n\}\), \(j\in\{1,2,\ldots,h\}\), \(k\in\{1,2,\ldots,w\}\);
[0035] Then, the final total loss is:
[0036] \(L=\lambda_1L\) rec +\(\lambda_2L\) percep +\(\lambda_3L\) excl +\(\lambda_4L\) adv +\(\lambda_5L\) mask +\(\lambda_6L\) color
[0037] Where \(\lambda_1 = \lambda_2 = \lambda_5 = \lambda_6 = 1\), \(\lambda_3 = 0.01\), \(\lambda_4 = 0.2\).
[0038] The present invention proposes an iterative prediction network model MPGINet (Multi-dimensional Perception-Guided Iterative Reflection Removal Network) for extracting image depth features and performing multi-dimensional reflection perception guidance according to a mask. Specifically, at the initial stage, the network model makes a preliminary prediction on the reflection layer and the transmission layer, and feeds the information of the reflection layer \(R'\), the transmission layer and the original input image \(I\) obtained from the preliminary prediction into the next stage of the network for joint estimation to predict the transmission layer \(T\). During the whole process, the output image of each iteration is used as the input image of the next iteration. In this way of iterative refinement, the gold foil reflection signal is dynamically separated, the highlight overflow is suppressed, and the quality of the final output image is gradually optimized. At the same time, the network strengthens the feature weights of the gold foil area through a squeeze-and-excitation module, retains details by combining a residual structure, and significantly improves the performance of single-image reflection removal by using cross-level information.
[0039] In images with heavier reflection regions, the linear combination assumption is no longer applicable to the mapping learning of the network. To address this issue, the present invention uses a mask repair mechanism through a reflection perception module to effectively recover the transmission layer in the manner of image repair, and repair the texture lost due to dynamic range overflow. Specifically, on the one hand, the DFPN module is used to perform depth feature extraction and spatial information perception on the original input image, fully mining the feature information in the original image, fusing multi-scale features through the spatial feature interaction block and the depth feature extraction block, and decoupling the complex reflection signals brought by mineral pigments, silk cotton, and glass interference, providing strong support for the subsequent image repair process. On the other hand, the information of the initially predicted reflection layer and transmission layer is jointly estimated, and the interference fringes of glass reflection, i.e., high-frequency noise and the low-frequency information of the transmission layer, are separated through the frequency domain module. This operation aims to remove the redundant information in the initially predicted image, effectively fuse it with the frequency domain features, and then generate a mask through the features of the decoder and encoder, as well as the multi-dimensional features in the MFPG module. This mask works in cooperation with partial convolution to effectively reduce the adverse effects caused by deviating from the linear combination assumption.
[0040] The main contributions of the present invention are summarized as follows:
[0041] (1) Innovatively proposed a single-image reflection removal network MPGINet. This network jointly estimates the multi-dimensional features of the reflection layer, transmission layer, and the original input image, and adopts an iterative refinement strategy to optimize the final output, significantly improving the accuracy and stability of reflection removal.
[0042] (2) Aiming at the rich texture details of Thangka images, a depth feature extraction module DFPN is designed. Among them, the depth feature extraction block and the spatial feature interaction block perform multi-scale feature extraction and fusion of image information, which can accurately capture the key information in the image and provide effective guidance for the estimation of the transmission layer during the image repair process.
[0043] (3) Aiming at the highlight overflow and strong reflection interference regions, a multi-dimensional feature perception guidance module MFPG is designed, which uses the features of different levels of the image and reconstructs the image transmission layer in the manner of image patching according to the mask.
[0044] (4) A large number of experiments conducted on multiple natural datasets and Thangka datasets show that the method proposed in the present invention surpasses the current state-of-the-art methods in both quantitative and qualitative dimensions, fully verifying the superiority and effectiveness of this method.
[0045] The beneficial effects of the present invention are:
[0046] The present invention proposes a multi-dimensional joint inference iterative network (MPGINet) guided by deep features, aiming to effectively solve the complex reflection artifact problem caused by high-intensity museum lighting in the digital protection of Regong Thangka. The present invention constructs a two-stage feature processing framework: in the first stage, the reflection layer is jointly predicted through a squeeze-and-excitation module; in the second stage, a multi-scale feature representation space is established through a deep feature guidance mechanism to achieve precise modeling of the underlying structure of the image. Specifically, first, the separation and suppression of reflection interference signals are completed in the frequency domain space. A multi-dimensional feature fusion strategy based on the attention mechanism is adopted for the texture breakage problem caused by high-reflection regions, and semantic consistency reconstruction of damaged regions is achieved by constructing a mask-guided adversarial repair module. At the same time, a multi-channel constraint function in the color space is introduced to effectively maintain the mineral pigment color characteristics unique to Thangka art. Description of the Drawings
[0047] Figure 1 Schematic diagrams of a pure Thangka picture and a Thangka picture contaminated by reflection;
[0048] Figure 2 Overall network structure diagram of an embodiment of the present invention;
[0049] Figure 3 Structure diagram of the squeeze-and-excitation module in an embodiment of the present invention;
[0050] Figure 4 Structure diagram of the deep feature pyramid network in an embodiment of the present invention;
[0051] Figure 5 Schematic diagram of the visualization comparison results of the reflection removal effects of different models on Thangka images in an embodiment of the present invention;
[0052] Figure 6 Schematic diagram of the visualization comparison results of the reflection removal effects of different models on natural images in an embodiment of the present invention;
[0053] Figure 7 Schematic diagram of the influence of the DFPN module on the texture details of the image in an embodiment of the present invention;
[0054] Figure 8 Schematic diagram of the influence of the color loss function on the image color in an embodiment of the present invention;
[0055] Figure 9 Schematic diagram of the visualization results of the ablation experiment based on the natural dataset in an embodiment of the present invention. Detailed Description of the Embodiment
[0056] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0057] Embodiment
[0058] A method for iterative removal of reflections in Thangka images guided by multi-dimensional perception of deep features, which is implemented based on an iterative prediction network that guides reflection perception based on multi-dimensional deep features. Its overall architecture is as follows Figure 2 As shown, this embodiment proposes a deep feature and multi-dimensional reflection-guided perception network based on two-stage collaborative learning, namely MPGINet. Different from traditional single-stage reflection removal methods, MPGINet generates a reflection layer prior in the first stage through a reflection feature enhancement module, and constructs a multi-dimensional feature fusion mechanism in the second stage to optimize the transmission layer. The cascaded reflection-transmission dual network can dynamically separate the gold foil specular reflection and substrate transmission signals through multiple iterative optimizations, gradually repair the high-light and strong reflection areas, and strengthen the learning of the Thangka gold foil edge features. Specifically, in each iteration N, the network takes the original degraded image Ι and the transmission layer estimate T predicted in the previous iteration N-1 as the input to predict the transmission layer T N and continues the iteration. Different from them, this embodiment emphasizes the reflection features in the first stage for preliminary reflection prediction. In the second stage, it combines the deep features and frequency domain features of the multi-dimensional images of the original image Ι, the predicted reflection image R', and the predicted transmission image T N-1 and realizes the reconstruction of the transmission layer T in the form of image inpainting through a mask N Reconstruction.
[0059] The following further explains this embodiment:
[0060] 1. Reflection layer prediction:
[0061] When deeply studying the regions with significant reflection phenomena, a two-stage model architecture is also constructed. In the first stage, a U-net model with multi-scale perception ability is selected as the network foundation, which takes the original image Ι as the input to predict the reflection layer R'. As Figure 2 shown, a squeeze-and-excitation block (SE-Block) is added behind each layer of the U-net, as Figure 3 shown. The SE layer is a structure for channel attention mechanism. It can reshape the weights of the feature map according to the global average pooling result of the input feature map, re-weight the features of the Thangka highlight regions, and perform constraints based on the non-uniform specular reflection characteristics of the gold powder painting regions, so as to emphasize important features and suppress unimportant features. And further expand the image receptive field through dilated convolution, capture feature information at different scales and long-term dependency information in the sequence, and jointly retain the line texture with the residual structure of ResNet to accurately and effectively transmit the reflection features during the two-stage iterative process.
[0062] 2. Multi-scale feature fusion and extraction:
[0063] In the current research work, this embodiment innovatively proposes a deep feature extraction network, namely DFPN (Deep Feature Pyramid Networks) derived from the feature pyramid network architecture. As Figure 4 shown, FPN was initially a key technical means in the fields of object detection and instance segmentation. The core design concept lies in effectively integrating the multi-scale features contained in the pre-designed classification network, providing a solid feature foundation for complex visual tasks.
[0064] The DFPN proposed in this embodiment is designed to be applied to the reflection removal task, and realizes the extraction and integration process of image depth features for the areas with severe reflection in Thangka images. DFPN can more finely extract the rich texture details and spatial position information in Thangka images, providing high-value data support for the subsequent image processing process.
[0065] DFPN mainly includes two key components: the spatial feature interaction block SFB (Spatial feature interaction block) and the depth feature extraction block DFB (Depth feature extraction block), which respectively replace the 1×1 convolution and 3×3 convolution in the traditional feature pyramid network. The two work together to achieve efficient feature processing, showing stronger adaptability and expressiveness when facing multi-scale feature extraction and depth feature perception tasks. Specifically, the spatial feature interaction block SFB integrates a token mixer, two linear layers, and a skip connection. The detailed structure can be referred to Figure 4 . Among them, the token mixer uses a 7×7 depth convolution to capture the scattering features brought by the coarse granularity of mineral pigments and suppress the periodic reflection of silk and cotton fabrics. The depth feature extraction block DFB adds a 7×7 dilated convolution on the basis of the spatial feature interaction block SFB and enhances the feature extraction ability through two additional skip connections, so as to focus on the metal reflection features brought by gold foil and provide guidance for subsequent reflection signal decoupling.
[0066] DFPN starts multi-scale feature fusion at the early stage of transmission layer prediction. This strategy ensures that high-level features in the early stage are integrated to guide the learning of low-level features in subsequent stages, can capture the long-term dependencies between spatial features, and can effectively decouple the separation of complex signals and reduce the amplitude of glass interference fringes through the pyramid structure, while the parameters and computational overhead only increase slightly.
[0067] 3. Multi-dimensional feature perception guidance:
[0068] In the second stage of the model, the original image Ι, the preliminary predicted reflection layer R', and the preliminary predicted transmission layer T' are used as the original inputs to predict the final transmission layer T, where T' = I - R. R' and T' are used to extract features through the VGG model to obtain F R and F T , and Ι is used to obtain F through DFPN I .
[0069] Through in-depth analysis, it is found that for the reflection layer and the transmission layer, frequency information is undoubtedly a crucial feature dimension. Specifically, high-frequency information can usually acutely indicate the characteristic attributes of the reflection layer, while low-frequency information has a close corresponding relationship with the transmission layer. To achieve an effective conversion from ordinary features to multi-frequency feature representations, as Figure 2 shown, a frequency-domain feature fusion module FFB (Frequency Fusion Block) is designed, in which octave convolution is used to extract high-frequency and low-frequency branches from ordinary features, namely glass interference fringes and the texture of the image body, effectively maintaining the continuity of millimeter-level meticulous brushstrokes, and then deeply fusing them with the original features to form a richer and more effective feature representation, and feeding it into the next layer of the model, namely the multi-dimensional feature perception guidance module MFPG (Multi dimensional feature perception guidance module). Through multi-dimensional feature guidance, the boundaries of lines and pigments can be effectively located, and the narrow-band color gamut features formed by the multi-layer superposition of mineral pigments can be effectively utilized, thereby constraining the prediction process of the transmission layer and performing high-quality image reconstruction on Thangka paintings of Regong art
[0070] During the operation of the model, as Figure 2 shown, the input features are processed by the decoder to obtain the decoded feature F dec . On this basis, addition and subtraction operations are performed using the observed features, reflection features, and transmission features to generate a preliminary feature result. Subsequently, through further processing by the channel attention block and the spatial attention block CBAM, the edge gradient is enhanced in the cotton fiber area, and the differential feature F diff is obtained. In addition, the module takes as input and passes through the linear transformation of two 1×1 convolutional layers and the non-linear mapping of the Sigmoid activation function, and based on the brightness threshold, the mask M is obtained. With the help of the mask, the model can accurately perceive the regions that deviate from the linear combination hypothesis, and use the surrounding decoder features to restore the transmission layer of these regions in the way of image inpainting, perform local contrast enhancement, and restore the texture of the overflow area, effectively optimizing the reflection removal in the areas with severe reflection
[0071] Considering the similarity between the repair of the double reflection region and reflection removal, specific weight and offset operations are performed on the surrounding regions where the mask M is greater than 0 to make full use of the feature information in this region; for other regions, their operation results are set to 0, so as to achieve the optimal utilization of the mask information. Aiming at the similarity problem of features between different channels, partial convolution is used to reduce redundant calculations. Only partial input channels are convolved, and the remaining channels remain unchanged, so as to effectively reduce redundant calculations and improve the operation efficiency of the model on the premise of ensuring the model performance.
[0072] 4. Loss function:
[0073] (1) Mask loss:
[0074] To better utilize the mask M during training, an additional mask loss is used for optimization to obtain better transmission map prediction. For regions with heavy reflection maps, the feature map F diff is difficult to provide information features, so its mask value is approximated to 0. In regions with less reflection, the mask value is assumed to be 1 to avoid trivial solutions as a regularization term and force the partial convolution layer to utilize both F diff and F dec . Then the total mask loss is:
[0075]
[0076] where i represents the i-th layer, ‖·‖1 represents the L1 norm, [] represents the conditions to be satisfied, is the threshold for dividing regions with heavy reflection, and ξ is the threshold for dividing regions with less reflection.
[0077] (2) Reconstruction loss:
[0078] During training, a large number of synthetic image pairs are used. The reconstruction loss can minimize the difference between the network output images T', R' and the corresponding ground truths T, R.
[0079]
[0080] (3) Perceptual loss:
[0081] Based on the pre-trained VGG-19 model, the L1 difference between φ(T'), φ(R') and φ(T), φ(R) in the selected feature layers is also minimized, where l is the index of the conv1_2, conv2_2, conv3_2, conv4_2 and conv5_2 layers, and the weight κ l is used to balance different layers.
[0082]
[0083] (4) Excluding losses:
[0084]
[0085] Among them, λ T and λ R are normalization factors, and are the gradients of T and R, ‖·‖ F is the Frobenius norm. T ↓n and R ↓n represent n downsampled versions of T and R, where T ↓0 and R ↓0 are the original inputs. During training, set N = 2,
[0086] (5) Adversarial loss:
[0087] To further improve the visual quality of the output image, the entire network is used as the generator G, and a 4-layer discriminator D is also constructed. Its parameter update is through:
[0088]
[0089] The parameter optimization of the generator G is:
[0090]
[0091] (6) Color loss:
[0092] There are a large number of colors and color variations in Thangka images. To reduce problems such as color distortion and color cast in the image after removing reflections, based on the color constancy loss, a color loss function is proposed to effectively constrain the image color. Minimize the difference between the output image T' and the corresponding ground truth T at the pixel level for the R, G, and B channels, which is expressed as follows:
[0093]
[0094] where N = n × h × w, n is the batch number, h is the height of the image, w is the width of the image, c ∈ {0, 1, 2} is the RGB channel, i ∈ {1, 2,..., n}, j ∈ {1, 2,..., h}, k ∈ {1, 2,..., w}.
[0095] (7) Total loss:
[0096] In summary, the final total loss is:
[0097] L = λ1L rec + λ2L percep + λ3L excl + λ4Ladv +λ5L mask +λ6L color
[0098] where λ1 = λ2 = λ5 = λ6 = 1, λ3 = 0.01, and λ4 = 0.2.
[0099] The following is a further illustration of this embodiment through experiments:
[0100] 1. Experimental details:
[0101] This embodiment uses the Thangka dataset and the natural dataset as the training set. The image pairs in the Thangka dataset are sourced from museums, personal art studios, etc. in Tongren City, Huangnan Tibetan Autonomous Prefecture, Qinghai Province. For these original data, preprocessing operations were first carried out, and then more than 7,000 groups of image pairs were selected from them. The natural dataset selected the same number (more than 7,000 groups) of image pairs from the PASCAL VOC dataset. Each image pair is used to generate a transmission map and a reflection map respectively. In each iteration, a group of image pairs consisting of two images is randomly selected, and its short side is randomly scaled to the interval [224, 448]. Then, image patches with a size of 224×224 are cropped from these two images respectively to synthesize the input image. The synthesis process follows the same data synthesis protocol as ERRNet. In the training stage, the ratio of the synthetic dataset to the real dataset is 8:2.
[0102] The test set consists of two parts. One part is the Thangka test dataset specially constructed for this embodiment, which contains a total of 336 pictures. The other part is the commonly used real-world dataset, including 20 pictures in real20 and three subsets in SIR. SIR 2 Wild contains 55 wild scene images, SIR 2 Solid contains 200 controlled scene images of daily life objects, and SIR 2 postcard contains 199 postcard images generated in a specific way, which is to use one postcard as the transmission image and another postcard as the reflection image for synthesis.
[0103] The model is trained based on the Adam optimizer, where the hyperparameters are set as β1 = 0.9, β2 = 0.999, and the learning rate is fixed at 1×10 -4 . The training process is executed for 150 epochs in total, and the number of samples in each batch (Batch size) is set to 1. The system environment used in the experiment is Ubuntu 18.04, equipped with an NVIDIA RTX 3090 24GB GPU. The software environment relied on by the experiment is Python 3.8, and the deep learning framework uses PyTorch 1.8.
[0104] 2. Comparative Experiments:
[0105] In this embodiment, a systematic experiment was carried out for the single - image reflection removal task, and six representative deep - learning methods were selected as the benchmark comparison models. The quantitative evaluation results shown in Table 1 indicate that the model proposed in this embodiment exhibits significant advantages on the Thangka dataset, and its peak signal - to - noise ratio (PSNR) and structural similarity index (SSIM) both reach the optimal level. Specifically, in terms of the PSNR index, this model improves by 1.88 dB compared with the sub - optimal method, the SSIM value increases by 0.027, and in terms of the average quantitative performance, it improves by more than 37% compared with the existing methods, fully verifying the effectiveness of the proposed method.
[0106] Figure 5 The visualization processing results of different methods on complex Thangka images are shown. Through comparative analysis, it can be found that: traditional methods tend to over - denoise, resulting in a decrease in the overall brightness of the image and obvious color distortion; the output images generated by the BDN and IBCLN models have significantly higher brightness than the input images, producing unnatural over - exposure effects in the dark regions; although DSRNet can effectively suppress strong reflection regions, its ability to retain high - frequency details is insufficient, resulting in blurred processing results; while RAGNet and RDNet fail to fully eliminate reflection interference in Thangka samples with rich color textures. In contrast, in this embodiment, while completely removing reflection artifacts, the unique mineral pigment color characteristics of Thangka images are effectively maintained, with detail fidelity in complex texture regions, and its processing results are significantly superior to existing methods in terms of visual perception quality and color restoration.
[0107] To systematically verify the generalization performance and robustness of the model proposed in this embodiment, this embodiment conducted extended experiments on two public benchmark datasets, SIR 2 and real20. As shown in Table 2, the model of this embodiment reaches the optimal or sub - optimal level in most evaluation indicators. Compared with the Thangka dataset, the public datasets mainly contain natural - scene images and the proportion of high - light reflection samples is relatively low. By constructing a multi - scale depth feature extraction mechanism and a multi - dimensional feature joint iterative prediction strategy, this embodiment demonstrates strong generalization ability in the natural - image domain. Figure 6 The comparison results show that the model of this embodiment effectively realizes artifact suppression and high - light reflection elimination in natural - image processing, and by introducing a color consistency constraint term, it maintains the original color - gamut characteristics of the image during the enhancement process.
[0108] Table 1 Quantitative Experimental Results Based on the Thangka Dataset (Optimal in Bold, Sub - optimal Underlined)
[0109]
[0110] Table 2 Quantitative experimental results based on public natural data sets (the best is bold, the suboptimal is underlined)
[0111]
[0112] 3. Ablation experiment:
[0113] In order to further verify the effectiveness of the reflection removal method proposed in this embodiment, an ablation experiment was first conducted on the Regong Thangka dataset. The experimental design included five control conditions: removing the DFPN module, removing the FFB module, removing the MFPG module, removing the SE-Block module, and removing the color loss function. All control experiments were conducted under unified parameters, hardware, and software conditions. As shown in Table 3, the quantitative evaluation of the Regong Thangka dataset shows that the complete model has achieved significant improvements in PSNR and SSIM indicators compared to the conditions where each module is missing.
[0114] Table 3 Ablation experiments on the Tangka dataset and natural dataset
[0115]
[0116] Specifically, the absence of the DFPN module leads to a 2.63dB decrease in the peak signal-to-noise ratio, which confirms the key role of the deep semantic feature extraction and fusion mechanism implemented by the multi-scale feature pyramid network in detail preservation. Figure 7 It can be seen that in the scenario of strong reflection interference, the missing DFPN module causes a lot of loss in image details and texture, while the complete model can retain the fine features of the image and effectively maintain the precise line structure of the Rekongthang Thangka.
[0117] In addition to Tangka's unique meticulous lines, the second is the rich colors based on natural mineral and plant pigments. Figure 8 From the analysis, it can be seen that due to the lack of color loss constraints, the model has a certain deviation in understanding the color detail information, and there are areas with abnormal colors in the output image. After introducing the color loss function proposed in this embodiment, the quality of thangka images has been greatly improved. From the subjective visual effect, the color abnormality area is significantly reduced, and the color transition is more natural and harmonious; from the analysis of image quality evaluation indicators, all indicators have been significantly improved, which strongly verifies the effectiveness and superiority of the color loss function in improving the quality of thangka images.
[0118] In addition to the Thangka dataset, this embodiment also conducts ablation experiments on natural datasets. The quantitative experimental results are shown in Table 3, from which it can be clearly observed that the various modules of the method proposed in this embodiment also show effectiveness on natural datasets.
[0119] Figure 9The image output results when different modules are missing in the natural dataset are shown. When the SE-Block module and the MFPG module are missing, there are obvious reflection regions in the output image that cannot be effectively removed. This is because the SE-Block module can enhance the attention to key features by adaptively adjusting the channel feature responses, while the MFPG module weights the features through multi-dimensional feature perception. The absence of both leads to the model's inability to accurately focus on the important information in the image, making it difficult to eliminate the interference of the reflection regions. The overall image output without the FFB module loses image details and overall sharpness because the high and low frequency information cannot be reasonably integrated.
[0120] The image output by the complete method proposed in this embodiment successfully overcomes the above various defects, which fully demonstrates the effectiveness and necessity of each module in improving image quality and dealing with specific problems. This effectiveness is not only reflected in the subjective visual evaluation but also strongly supported by the quantitative experimental results.
[0121] The MPGINet in this embodiment adopts a two-stage iterative architecture. In the initial stage, the reflection layer prediction is optimized by combining the U-Net with the compression excitation module and the residual mechanism; in the second stage, the newly proposed deep pyramid network (DFPN) is adopted, and the fine-grained capture of Thangka texture features is realized through the spatial feature interaction block (SFB) and the deep feature extraction block (DFB), and the fine-grained restoration of the transmission layer is achieved by fusing the high and low frequency features and the mask constraint. Experimental results show that on the Thangka dataset, the PSNR and SSIM of MPGINet reach 28.90dB and 0.962 respectively, an improvement of 1.88dB and 0.027 compared with the existing best method; in the natural scene datasets (SIR 2 , Real20), the average PSNR is improved by 1.64dB, verifying the generalization ability of the method. Ablation experiments further show that the DFPN module and the color loss function play a key role in detail preservation and color restoration.
[0122] The above-described embodiments only represent the specific implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. A method for iteratively removing reflection from thangka images guided by multi-dimensional perception of deep features, characterized in that: The iterative prediction network implementation based on multi-dimensional reflection perception guidance based on deep features specifically includes the following steps: Step 1: Generate reflection layer prior through reflection feature enhancement module: Optimize reflection layer prediction through U-Net model and compression excitation module combined with residual mechanism; Step 2: Construct a multi-dimensional feature fusion mechanism to optimize the transmission layer: Based on the deep feature pyramid network, the spatial feature interaction block and the deep feature extraction block are used to capture the thangka texture features in a refined manner, and the high- and low-frequency features are integrated with the mask constraints to achieve refined restoration of the transmission layer.
2. The method for iteratively removing reflection from thangka images guided by multi-dimensional perception of deep features according to claim 1 is characterized in that: The step 1 specifically comprises the following steps: Step 1.1, the U-net model uses the original image I as input to predict the reflection layer R'; Step 1.2, add a compression excitation module SE-Block after each layer of the U-net model, reshape the weight of the feature map according to the global average pooling result of the input feature map, implement feature reweighting on the highlighted area of the thangka, and constrain it based on the non-uniform specular reflection characteristics of the gold powder depiction area; expand the image receptive field through extended convolution, capture feature information of different scales and long-term dependency information in the sequence, and combine the residual structure of ResNet to retain the line texture and transfer the reflection feature in the iterative process.
3. The method for iteratively removing reflection from thangka images guided by multi-dimensional perception of deep features according to claim 1 is characterized in that: In step 2, the spatial feature interaction block includes a token mixer and two linear layers connected in sequence, and a jump connection is formed between the input end of the token mixer and the output end of the last linear layer.
4. The method for iteratively removing reflection from thangka images guided by multi-dimensional perception of deep features according to claim 3 is characterized in that: In step 2, the deep feature extraction block includes two layers of token mixers and two layers of linear layers connected in sequence, and a jump connection is used between the input and output ends of the two layers of token mixers, and a jump connection is used between the input end of the first layer of token mixer and the output end of the last layer of linear layer.
5. The method for iteratively removing reflection from thangka images guided by multi-dimensional perception of deep features according to claim 4 is characterized in that: The step 2 is specifically as follows: In each layer iteration process, the iterative prediction network takes the original image Ι, the preliminary predicted reflection layer R', and the preliminary predicted transmission layer T' as the original input and predicts the final transmission layer T, where T' = IR; R' and T' are extracted by the VGG model to obtain F R and F T , Ι obtains F through the deep pyramid network DFPN I ; The frequency domain feature fusion module FFB uses octave convolution to extract high-frequency branches and low-frequency branches from common features, namely glass interference fringes and image body texture, effectively maintaining the continuity of millimeter-level fine brushwork lines, and then deeply fuses them with the original features to form a richer and more effective feature representation, which is fed into the next layer of the iterative prediction network, namely the multi-dimensional feature perception guidance module MFPG; During the operation of the iterative prediction network, the input features are processed by the decoder to obtain the decoded features F dec On this basis, the observation features, reflection features and transmission features are added and subtracted to generate preliminary feature results; then, through further processing by the channel attention block and the spatial attention block CBAM, the edge gradient is enhanced in the cotton hairiness area, and the differential feature F is obtained. diff ; In addition, the iterative prediction network is based on F I 、F R and F dec The result after concat is used as input through the linear transformation of two 1×1 convolutional layers and the nonlinear mapping of the Sigmoid activation function to obtain the mask M based on the brightness threshold.
6. The method for iteratively removing reflections from thangka images guided by multi-dimensional perception of deep features according to claim 5 is characterized in that: The loss function of the iterative prediction network includes mask loss, reconstruction loss, perceptual loss, exclusion loss, adversarial loss and color loss, where: The mask loss is: Among them, M diff Yes F diff The mask of , i represents the i-th layer, ‖·‖1 represents the L1 norm, [] represents the conditions that need to be met, is the threshold for dividing the area with heavier reflection, and ξ is the threshold for dividing the area with lighter reflection; The reconstruction loss is used to minimize the difference between the network output images T', R' and the corresponding true values T, R. The reconstruction loss is: Based on the pre-trained VGG-19 model, minimize the L1 difference between φ(T'), φ(R') and φ(T), φ(R) in the selected feature layer, where l is the index of conv1_2, conv2_2, conv3_2, conv4_2 and conv5_2 layers, and the weight κ l Used to balance different layers; then the perceptual loss is: The excluded losses are: in, λ T and λ R is the normalization factor, and is the gradient of T and R, ‖·‖ F is the Frobenius norm; T ↓n and R ↓n represents n downsampled versions of T and R, where T ↓0 and R ↓0 is the original input; during the training process, set N = 2, In the adversarial loss, the entire network is used as the generator G, and a 4-layer discriminator D is constructed, whose parameters are updated by: The parameters of the generator G are optimized as follows: In the color loss, the difference between the output image T' and the corresponding true value T is minimized at the pixel level for the three channels R, G, and B, as follows: Where N = n × h × w, n is the batch size, h is the height of the image, w is the width of the image, c ∈ {0, 1, 2} is the RGB channel, i ∈ {1, 2, ..., n}, j ∈ {1, 2, ..., h}, k ∈ {1, 2, ..., w}; Then, the final total loss is: L=λ1L rec +λ2L percep +λ3L excl +λ4L adv +λ5L mask +λ6L color Among them, λ1 = λ2 = λ5 = λ6 = 1, λ3 = 0.01, and λ4 = 0.2.