RAW-RGB collaborative color moire removing method based on Mama architecture

Through the RAW-RGB collaborative color moiré removal method based on the Mamba architecture, combining RAW and RGB data and adopting the dual-path adaptive feature fusion module DAFF, the problem of high computational complexity in the existing technology is solved, and efficient color moiré removal is achieved on resource-constrained devices.

CN120807305APending Publication Date: 2025-10-17UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510673487.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies usually only process color moiré removal in a single domain, resulting in high computational complexity, making it difficult to deploy on resource-constrained devices, and the removal performance needs to be improved.

Method used

A RAW-RGB collaborative color moiré removal method based on the Mamba architecture is adopted. Through the shallow demoiré and RAW feature hierarchy stages and the deep fusion and color-aware demoiré stages, RAW and RGB data are combined, and the dual-path adaptive feature fusion module DAFF is used for cross-modal information fusion. The network architecture is optimized to reduce computational complexity.

Benefits of technology

While maintaining high performance, it significantly reduces computational complexity, effectively removes color moiré, and improves image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807305A_ABST
    Figure CN120807305A_ABST
Patent Text Reader

Abstract

The invention discloses an RAW-RGB collaborative color moire removal method based on a Mama framework. The method adopts a two-stage framework. In the first stage, shallow moire removal is carried out in an RAW domain through a simple moire removal block SDB, and multi-scale RAW features are extracted; and in the second stage, the multi-scale RAW features and the RGB features are fused by using a dual-path adaptive feature fusion (DAFF) module, and deep moire removal and accurate color restoration are realized by enhancing a moire removal block (DMB). The DMB and DAFF modules are specifically designed to reduce the Mama design of computational complexity while maintaining the de-moire performance. By utilizing complementary information of RAW and RGB domains, moire patterns are effectively eliminated, and meanwhile accurate color calibration is kept. Experiments on a TMM22 data set show that the method has the most advanced performance in quantitative index and qualitative visual comparison, and meanwhile, the low calculation complexity is kept.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical fields of image processing and machine learning, and particularly relates to a technique for removing color moire based on machine learning. BACKGROUND

[0002] Moire is a kind of visual interference caused by frequency aliasing effect. With the rapid development of mobile phones and camera technology, it has become a convenient and efficient way to record information by shooting electronic screens. However, such photos are usually accompanied by complex high-frequency moire due to the frequency similarity between the color filter array CFA of the imaging device and the display pixels. This complex high-frequency moire belongs to the spatial aliasing type moire, which is manifested as red, green, blue and other color mottling stripes, i.e. color moire. Color moire significantly reduces the image quality. The complexity and high-frequency characteristics of color moire make it an important challenge in the field of image processing to remove it.

[0003] The generation of color moire is closely related to the color filter array CFA of the imaging device and the pixel arrangement of the display device. The frequency domain range of moire pollution is wide, covering almost all scales, which means that color moire not only affects the high-frequency part of the image, but also penetrates into the medium-frequency and low-frequency parts. In addition, the frequency, shape and color of color moire will change with the type of shooting device, shooting angle and characteristics of the display device, which further increases the difficulty of moire removal. The effects of color moire in RGB domain and RAW domain are shown in (a) and (b) of FIG. 1, respectively. Figure 1

[0004] At present, the color moire removal methods based on deep learning mainly focus on the RGB domain, and feature decomposition and reconstruction are carried out in the spatial domain or frequency domain through multi-scale convolutional neural network CNN. Although these methods can remove color moire to some extent, the moire repair performance needs to be improved.

[0005] RAW domain data is the original data output by the image sensor without demosaicing, white balance, noise reduction, compression and other processing, which directly records the photoelectric signal captured by the sensor. Compared with RGB domain image moire removal, RAW domain data contains deeper data depth and more original image information, which makes RAW domain have unique advantages in moire removal. SUMMARY

[0006] ​The technical problem to be solved by the present application is that the current color moire removal is usually processed in a single domain, such as a single RGB domain or a single RAW domain, and in order to ensure processing performance, a network architecture with high computational complexity is usually relied on, which is difficult to deploy on limited devices, and the applicant considers the complementarity between RAW and RGB data, and proposes a method for removing color moire combining RAW and RGB data and optimizing the network architecture, which can be applied to resource-limited devices.

[0007] The technical solution adopted by the present application to solve the above technical problem is a RAW-RGB collaborative color moire removal method based on a Mamba architecture, which obtains a color moire removed RGB image through a shallow moire removal and RAW feature hierarchy stage and a deep fusion and color perception moire removal stage corresponding to a given RAW image and RGB image pair.

[0008] (1) The shallow moire removal and RAW feature hierarchy stage includes:

[0009] The preprocessing step is to preprocess the RAW image into an RGGB format to match the Bayer color filter array layout.

[0010] The shallow feature extraction step is to input the preprocessed RAW image into a shallow convolution module to extract shallow features.

[0011] The deep feature extraction step is to input the shallow features into a simple moire removal block SDB coding and decoding network, the SDB coding and decoding network adopts a multi-scale down-sampling encoder and up-sampling decoder structure, and the SDB coding and decoding network outputs multi-size deep RAW features extracted in a hierarchical manner.

[0012] (2) The deep fusion and color perception moire removal stage includes:

[0013] The RGB feature generation step is to extract RGB features matching the size of the shallow features from the input RGB image, and output the RGB features as the first-level RGB input features to an enhanced moire removal block DMB coding and decoding network and an input end of a DAFF module connected to the first-level down-sampling encoding input of the DMB coding and decoding network.

[0014] The feature fusion step is to run the DAFF module in the encoding stage of the DMB coding and decoding network, gradually fuse each scale of RAW features and the corresponding RGB input features through cross-modal interaction, and output the fused features to the same-level input of the down-sampling encoding of the DMB coding and decoding network, and output each level of fused features to the input end of the corresponding scale of up-sampling decoding.

[0015] RGB image reconstruction step: each level of upsampling decoding utilizes the fused feature output to reconstruct multi-scale RGB output features, generating multi-scale RGB features; each scale-specific convolution module generates a reconstructed RGB image through the corresponding RGB output feature, completing the collaborative color moire removal using the complementarity of RAW and RGB data.

[0016] RAW data provides more original image information, while RGB data provides device-dependent color calibration information. By jointly using these two types of data, effectively utilizing RAW and RGB data, fully utilizing cross-modal information, and improving moire removal effect.

[0017] The DAFF module is based on the Mamba design of the efficient scanning algorithm, and the DMB is based on the Mamba design of the selective 8-direction scanning algorithm. By optimizing the network architecture, the computational complexity can be significantly reduced while maintaining high performance.

[0018] In the training process, the loss function L RAW and the loss function L RGB of the deep layer fusion and color perception moire removal stage are respectively:

[0019] L RAW =||G RAW -O RAW ||1

[0020]

[0021] The total loss function L is:

[0022] Total

[0023] L Total =λ3L RAW +λ4L RGB

[0024] where O RAW represents the output image of the shallow moire removal and RAW feature hierarchy stage, represents the i-th level output image of the deep layer fusion and color perception moire removal stage, G represents the corresponding clean image, ||·||1 represents the L1 norm, λ1, λ2, λ3, λ4 are respectively 1.0, 1.0, 1.0, 0.5, M represents the total number of multi-scale, i represents the i-th level, ψ j represents the feature extraction of VGG at the j-th layer.

[0025] The present application innovatively designs a two-stage architecture based on RAW data and RGB data, which can fully extract and fuse RAW data and RGB data to realize moire removal. The first stage is a simple moire removal block SDB, which is based on a convolutional encoder-decoder, and roughly suppresses moire artifacts in the RAW domain while preserving the original local features. In the second stage, an 8-direction scanning enhanced moire removal Mamba block DMB is specially designed to utilize RAW and RGB data for deep multi-scale artifact removal. Cross-stage feature fusion is achieved through a dual-path adaptive feature fusion DAFF module with an efficient scanning algorithm, which gradually integrates dual-domain information. A hybrid loss function is proposed to supervise both stages simultaneously: a single-scale L1 loss is used in the first stage to guide rough artifact removal in the RAW domain, while a multi-scale VGG perceptual loss and an L1 reconstruction loss are combined in the second stage to eliminate residual moire textures across frequency bands.

[0026] The present application effectively combines shallow RAW data processing and deep cross-modal fusion to solve the moire problem. The dual supervision mechanism uniquely coordinates global structure preservation and fine texture restoration through end-to-end optimization. Extensive experiments conducted on the TMM22 dataset demonstrate the rationality of the network design, achieving the most advanced qualitative and quantitative results compared to existing methods. BRIEF DESCRIPTION OF DRAWINGS

[0027] Figure 1 (a) is a moire RGB image, and (b) is a moire RAW image.

[0028] Figure 2 (a) is the overall architecture of MPFNet, (b) is the SDB module schematic diagram, and (b) is the DAFF module schematic diagram.

[0029] Figure 3 (a) is the detailed structure of the moire removal Mamba block DMB, and (b) is the adaptive visual state space model AVSSM in the DMB.

[0030] Figure 4 is a specific schematic diagram of the dynamic feature enhancement module DFFM in the dual-path adaptive feature fusion DAFF module and the cross-modal fusion Mamba module CMFM. DETAILED DESCRIPTION

[0031] An image dataset is prepared for training the model of removing color moire. The dataset is from a downloaded public dataset TMM22, which includes 540 pairs of high-resolution RAW and RGB images with a resolution of 1800x1300, covering three categories: natural images, document images, and webpage images. During the training process, these images are further processed into overlapping image blocks, generating 63180 pairs of 256x256 image blocks. In addition, the dataset also provides 408 pairs of 512x512 images for testing.

[0032] A RAW-RGB collaborative moire removal network model based on the Mamba architecture is constructed, as shown in (a) in FIG. 1, including a shallow convolution module (Shallow Convolution), an encoder-decoder network composed of a simple moire removal block SDB, a convolutional module (Convolution Layers) of RGB, an encoder-decoder network composed of an enhanced moire removal block DMB, and several double-path adaptive feature fusion DAFF modules. Figure 2

[0033] Hereinafter, the encoder-decoder network composed of the simple moire removal block SDB is referred to as the SDB encoding-decoding network, and the encoder-decoder network composed of the enhanced moire removal block DMB is referred to as the DMB encoding-decoding network.

[0034] The model is divided into a shallow moire removal and RAW feature hierarchy and a deep fusion and color perception moire structure.

[0035] Among them, the shallow moire removal and RAW feature hierarchy includes a shallow convolution module and an SDB encoding-decoding network; the deep fusion and color perception moire structure includes an RGB convolution module and a DMB encoding-decoding network.

[0036] The RAW image of the image to be processed is input to the shallow convolution module after preprocessing, and the shallow convolution module outputs shallow features to the SDB encoding-decoding network. The multi-scale RAW features output by the SDB encoding-decoding network are output to the DAFF module of the corresponding scale, which also receives the RGB input features of the same scale. Each scale DAFF module fuses RAW features and RGB input features. The received RGB input features of the same scale come from the output of the downsampling encoder of the previous level of the DMB encoding-decoding network or the output of the RGB convolution module. Each scale DAFF module outputs the fusion features of each scale to the DMB encoding-decoding network, and the DMB encoding-decoding network outputs the reconstructed RGB images of each scale, completing the removal of color moire.

[0037] ​The SDB coding network and the DMB coding network adopt the same multi-scale down-sampling encoder and up-sampling decoder structure. The output of each level of the up-sampling decoder in the SDB coding network corresponds to the input of a DAFF module, the output of the DAFF module is connected to the input of the corresponding level of the down-sampling encoder in the DMB coding network of the same scale, and the DAFF module also has an input connected to the encoding output of the previous level in the DMB coding network.

[0038] In the SDB coding network and the DMB coding network, a skip connection is usually arranged between the down-sampling encoder and the up-sampling decoder. The embodiment is to connect the third-level encoding output and the first-level decoding output to the second-level decoding input through a DAFF module in the SDB coding network and the DMB coding network.

[0039] In the embodiment, the SDB coding network and the DMB coding network both adopt a four-level down-sampling encoder and a four-level up-sampling decoder.

[0040] Based on the model architecture shown in (a) in FIG. 1, the following steps are performed for the color moire removal, including (1) a shallow moire removal and RAW feature hierarchy stage and (2) a deep fusion and color perception moire removal stage. Figure 2

[0041] The shallow moire removal and RAW feature hierarchy stage Stage1 includes:

[0042] 1-1 Preprocessing step: receiving a given pair of RAW format image I RAWImage ∈R 1×H×W and RGB format image I RGB ∈R 3×H×W , pre-process I RAWImage into a 4-channel RGGB format image I RAW ∈R 4×H / 2×W / 2 ; where R is the real number field, indicating that each pixel value in the image is a real number, and the superscript CxHxW represents the image dimension, C, H and W representing the channel number, height and width respectively. The RGGB format is the most common arrangement in the Bayer color filter array of the imaging sensor. Pre-processing I RAWImage into the RGGB format is used to match the Bayer color filter array layout.

[0043] 1-2 Shallow feature extraction step: pre-processed I RAW is input into a shallow convolution module (Shallow Convolution) composed of convolution and GELU activation function, and the shallow convolution module outputs the extracted shallow feature F RAW ∈R C ×H / 2×W / 2 .​

[0044] 1-3 Deep feature extraction step: shallow feature F RAW Input into SDB codec network, SDB codec network outputs deep RAW features F extracted hierarchically RAWi , i = 1, 2, 3, 4, represents the i-th level. In the training process, the last scale F RAW4 A dedicated convolution module is used to generate reconstructed RAW image O RAW . Image O RAW The loss calculation for the training process.

[0045] (2) Deep fusion and color-aware demoising stage Stage2

[0046] 2-1 RAW feature generation step: for input RGB image I RGB ∈R 3×H×W Down-sampling and convolution processing to match the size of the shallow feature F RAW RGB feature, which is output as the first level RGB input feature to the DMB codec network and the input end of the DMB codec network connected to the first level down-sampling encoding input of the DMB codec network;

[0047] 2-2 Feature fusion step: DMB codec network runs in the encoding stage of the DMB codec network, through cross-modal interaction fusion, that is, each scale RAW feature F RAWi The corresponding i-th level RGB input feature is fused and output to the i-th level down-sampling encoding input of the DMB codec network, and each level of down-sampling encoding outputs each level of fusion feature to the input end of the corresponding scale up-sampling decoding;

[0048] 2-3 RGB image reconstruction step:

[0049] Each level of up-sampling decoding uses the fusion feature to output multi-scale RGB output features, and generates multi-scale RGB features. Each scale-specific convolution module processes these features to generate reconstructed RGB image O RGBi , i = 1, 2, 3, 4, from which the collaborative demoising processing using the complementarity of RAW and RGB data is completed.

[0050] As Figure 2(b) shown, each SDB consists of a convolutional layer Conv, an activation function GELU and a shortcut connection, sequentially connected in data processing order, SDB input X, 5x5 convolutional layer Conv5, GELU, 3x3 convolutional layer Conv3, GELU, 1x1 convolutional layer Conv1, an adder and SDB output Y, the other input end of the adder is connected to SDB input X to form a shortcut connection. The adder is used for element-wise addition.

[0051] Further, the DAFF module is based on the Mamba design of the efficient scanning algorithm, and the DMB is based on the Mamba design of the selective 8-direction scanning algorithm, which can effectively model long-distance dependencies while maintaining linear computational complexity. The 8-directions include 4-directions in the horizontal and vertical directions and 4-directions in the diagonal direction. As shown in Figure 2 (c) shown, is a structural schematic diagram of the DAFF module, including two DFFM modules and a CMFM module connected thereto.

[0052] Further, as shown in (a) of Figure 3 The specific structure of the enhanced demoising block DMB is shown in (a) of FIG. 6. The input mixed feature map X is first normalized by layer normalization LN to stabilize the training process. The normalized feature map LN(X) enters the AVSSM module for processing to extract the moire pattern information. The output AVSSM(LN(X)) of the AVSSM module is weighted by a coefficient a and then fused with the input feature X through element-wise addition. The fused feature X1 is further normalized by layer normalization LN and convolved by Conv1, and then activated by the GELU activation function to introduce nonlinearity. The activated feature GELU(Conv1(LN(X1))) is input into the channel attention CA module to enhance the weight of important features. The output CA(GELU(Conv1(LN(X1)))) of the CA module is added to the first fused feature X1 weighted by a coefficient b to obtain the feature X2. X2 is added to the local feature LDC(X) enhanced by the local dynamic convolution LDC module with the input X, and finally the output feature Y is obtained, which is expressed as follows:

[0053] X1=AVSSM(LN(X))+aX

[0054] X2=CA(GELU(Conv1(LN(X1))))+bX1

[0055] Y=LDC(X)+X2

[0056] As shown in (b) of Figure 3 The specific structure of the adaptive visual state space model AVSSM is shown in (b) of FIG. 6. The input feature map X nFirst, the feature dimension is expanded by a linear layer to achieve a more rich representation, and the expanded features are processed by a SiLU activation function to emphasize nonlinearity, generating high-level feature representations for short connection, while the expanded features X n are processed in parallel by 5x5, 3x3 and 1x1 convolution layers Conv1(X n ), Conv3(X n ), Conv5(X n ) and then spliced to capture information of different scales of moire, and the obtained multi-scale features X cat are sequentially processed by a 1x1 convolution layer Conv1, a SiLU activation function and a selective state space model SSSM, a global scanning mechanism is used to establish long-distance dependency, and the output SSSM(SiLU(Conv1(X cat )) of the SSSM module is normalized by layer normalization LN, the normalized features Fa and the output Fb of the short connection branch are fused by element multiplication, and then the fused features are processed by layer normalization LN, and finally the feature map F g is output, which is expressed as follows:

[0057] X n '=Linear(X n )

[0058] X cat =car(Conv1(X n '),Conv3(X n '),Conv5(X n '))

[0059] F a =LN(SSSM(SiLU(Conv1(X cat ))))

[0060] F b =SiLU(X n ')

[0061] F g =LN(F a *F b )

[0062] Specifically, as Figure 4As shown, the double-path adaptive feature enhancement module DAFF structure is as follows: it contains two branches, the input two different modal feature maps F1 and F2 needing to be fused are first subjected to element addition fusion to obtain preliminary fusion features, two shallow features and the initial fusion features are respectively input into the dynamic feature enhancement modules DFEM in the two branches for feature enhancement to obtain two enhanced features, the two enhanced features are simultaneously input into the cross-modal fusion Mamba module CMFM for feature fusion, and finally the fusion features H are output f .

[0063] The DAFF module is innovatively used at the skip connection between the encoder and the decoder. The skip connection of the existing network using the codec structure to remove moire is either directly connected or processed by other modules. The DAFF module can process the processed data and the pre-processed data from the skip connection, thereby better identifying and removing moire.

[0064] The structure of the dynamic feature enhancement module DFEM is as follows: the features F1 and F2 are first subjected to weighting based on LDC to obtain texture features T1 and T2 of the enhanced F1 and F2. T1 and T2 are subjected to element-wise subtraction by the subtracter, thereby obtaining feature differences, and the feature differences are subjected to global average pooling GAP and the activation function Sigmod to generate an adaptive weight map F f is obtained by element addition of different modal features. These learned weights are then applied to the fusion feature map and the original feature map F1 and F2 to amplify the cross-feature difference, while efficiently extracting adaptive weight maps and enhancing the complementary features and texture details existing in the input images T1 and T2 to obtain features D1 and D2. The expression is as follows:

[0065]

[0066] The structure of the cross-modal fusion Mamba module CMFM is as follows: the input features D1 and D2 are subjected to normalization LN, LN(D1) and LN(D2) are respectively passed through the linear layer Linear and the deep convolution Dwc to obtain C1 and C2, then C1 and C2 are subjected to element-wise multiplication by the multiplier, and the multiplication results C1, C2 and The mixed feature H and the mixed enhanced feature H are added by the adder to generate, and then are input into the selective state space model SSSM to capture the long-term spatial dependence relationship. The SSSM output SSSM(H) is multiplied with the linear layer Linear output Linear(LN(D1) and Linear(LN(D2) to obtain the features H1 and H2, respectively. Next, H1 and H2 are added and then input into the linear layer Linear. The linear layer Linear output Linear(H1⊕H2) feature is added with H1 and H2 after the efficient channel attention CA operation to reduce the channel redundancy, as follows:

[0067] C1 = Dwc(Linear(LN(D1)))

[0068] C2 = Dwc(Linear(LN(D2)))

[0069]

[0070] The final fusion feature map X of each layer is:

[0071] wherein, Dwc is a deep convolution operation, and are element multiplication and addition operations, respectively.

[0072] The network module is designed using Mamba. The linear calculation complexity of Mamba itself can reduce the calculation complexity. The selectivity of Mamba mainly lies in that the parameters thereof are learnable, and some parts in the input sequence can be selectively remembered or ignored. The selective state space model SSSM used in the above module is an 8-direction scanning algorithm-based Mamba module, and ES2D is an efficient scanning algorithm-based Mamba module.

[0073] Preferably, the loss function adopts a hybrid loss function. In the first stage, a RAW single-scale L1 norm loss is used, and in the second stage, a multi-scale deep convolutional neural network VGG perceptual loss and an L1 norm reconstruction loss are combined to optimize the multi-scale moire. The VGG network is used herein only for calculating the loss.

[0074] The first stage loss function is:

[0075] L RAW = ||G RAW -O RAW ||1

[0076] The second stage loss function is:

[0077]

[0078] The total loss function is:

[0079] L Total = lambda3L RAW + lambda4L RGB

[0080] Where O represents the network output image, O RAW represents the shallow layer demoistening and RAW feature hierarchy stage output image, represents the deep layer fusion and color perception demoistening stage i-level output image, G represents the corresponding clean image of the output, ||.||1 represents the L1 norm, lambda1, lambda2, lambda3, lambda4 are respectively 1.0, 1.0, 1.0, 0.5, M represents the total number of multi-scale levels, i represents the i-level, represents After inputting the VGG network, the feature extraction of the jth layer output by the VGG network is obtained.

[0081] The trained image demoistening network receives the verification data set image which needs to be demoistened, and outputs the image after completing the demoistening, and the specific method is as follows:

[0082] Load the image demoistening network model weight trained, update the parameters in the model. Secondly, the verification data set moire image is input into the network model as input, and the processed picture is output after network processing. Note that only the maximum image output by the model needs to be compared with the real clear image during model verification.

[0083]

[0084] Table 1

[0085] The method in the application is compared with other related work on the TMM22 data set, and the experimental results are shown in Table 1. From the data results in Table 1, it can be seen that the demoistening effect of the application on the TMM22 data set is better than that of other existing methods, which shows the effectiveness of the application.

Claims

1. A RAW-RGB collaborative color moiré removal method based on the Mamba architecture, characterized in that: For a given pair of RAW and RGB images, the RGB image with color moiré removed is obtained through the shallow demoiré and RAW feature hierarchy stages and the deep fusion and color-aware demoiré stages. (1) Shallow de-moiré and RAW feature hierarchy stages include: Preprocessing step: preprocess the RAW image into RGGB format to match the Bayer color filter array layout; Shallow feature extraction step: The preprocessed RAW image is input into the shallow convolution module to extract shallow features; Deep feature extraction step: Shallow features are input to the simple demoiré block SDB codec network. The SDB codec network uses a multi-scale downsampling encoder and upsampling decoder structure. The SDB codec network outputs hierarchically extracted multi-scale deep RAW features. (2) Deep fusion and color perception de-moiré stage: RGB feature generation step: Extract RGB features that match the size of shallow features from the input RGB image, and output this RGB feature as the first-level RGB input feature to the enhanced de-moiré block DMB codec network and one input end of the DAFF module connected to the first-level downsampling encoding input in the DMB codec network; Feature fusion step: The DAFF module runs in the encoding stage of the DMB codec network. Through cross-modal interaction, it gradually fuses the RAW features of each scale with the corresponding RGB input features and outputs them to the downsampling encoding level input of the DMB codec network. Each level of downsampling encoding outputs the fused features of each level to the input of the upsampling decoding of the corresponding scale. RGB image reconstruction steps: Upsampling decoding at each level uses the fused feature output to reconstruct multi-scale RGB output features to generate multi-scale RGB features; each scale-specific convolution module generates a reconstructed RGB image through the corresponding RGB output features, completing the collaborative de-coloring moiré processing using the complementarity of RAW and RGB data.

2. The method according to claim 1, characterized in that In the simple de-moiré block SDB encoding and decoding network, each SDB consists of a 1×1 convolutional layer Conv1, a 3×3 convolutional layer Conv3, a 5×5 convolutional layer Conv5, an activation function GELU and a shortcut connection. In the order of data processing, the SDB input, the 5×5 convolutional layer Conv5, the activation function GELU, the 3×3 convolutional layer Conv3, the activation function GELU, the 1×1 convolutional layer Conv1, the adder and the SDB output are connected in sequence. The other input of the adder is connected to the SDB input to form a shortcut connection.

3. The method according to claim 1, characterized in that The DAFF module is based on the Mamba design of the efficient scanning algorithm, and the DMB is based on the Mamba design of the selective 8-way scanning algorithm.

4. The method according to claim 1, characterized in that The processing flow of enhanced de-moiré block DMB is as follows: The input fusion feature X is first normalized, and the normalized feature map enters the AVSSM module to extract moiré pattern information. The output of the adaptive visual state space model AVSSM module is weighted by the preset coefficient a of the terminal, and then fused with the input fusion feature X by element-by-element addition. The fused feature X1 is further normalized and convolved, and then activated by the GELU activation function. The activated feature passes through the channel attention CA module. The output of the CA module is added to the first fused feature X1 weighted by the coefficient b to obtain feature X2. X2 is then combined with the input fusion feature X through the feature output of the local dynamic convolution LDC module to obtain the DMB output feature.

5. The method according to claim 4, characterized in that: The processing flow of the adaptive visual state space model AVSSM is as follows: Input feature map X n First, the feature dimension is expanded through the linear layer, and the expanded feature X n 'Generate the output F of the short-circuit branch through the SiLU activation function b , and the expanded feature X n ′The multi-scale feature X is obtained by parallel processing of 5×5, 3×3 and 1×1 convolutional layers and then concatenating them. cat , X cat Then it is processed in sequence by 1×1 convolution layer, SiLU activation function and selective state space model SSSM. The output of SSSM module is processed by layer normalization LN. The normalized feature F a With the short-circuit branch output F b After fusion through element multiplication, the output feature map of the AVSSM module is obtained through layer normalization LN.

6. The method according to claim 1, characterized in that During the training process, the loss function L of the shallow de-moiré and RAW feature hierarchy stage is RAW And the loss function L in the deep fusion and color-aware de-moiré stages RGB They are: L RAW =||G RAW -O RAW ||1 Total loss function L Total for: L Total =λ3L RAW +λ4L RGB Among them O RAW Represents the output image of the shallow de-moiré and RAW feature hierarchy stage, represents the output image of the i-th level in the deep fusion and color-aware de-moiré stage, G represents the clean image corresponding to the output, ||·||1 represents the L1 norm, λ1, λ2, λ3, and λ4 are 1.0, 1.0, 1.0, and 0.5 respectively, M represents the total number of multi-scale levels, i represents the i-th level, and ψ j Represents the feature extraction of VGG at the jth layer.