Ultrahigh-definition image moire removing method based on pyramid learnable band-pass filter
By introducing a pyramid-learning bandpass filter module and a cross-layer feature fusion module in ultra-high-definition image processing, the shortcomings of the prior art in processing ultra-high-definition image molars are solved, and effective modeling and removal of complex molars are achieved, and image quality is significantly improved.
Patent Information
- Application Number
- CN202510016610.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-06
AI Technical Summary
The prior art has shortcomings in processing molar patterns in ultra-high-definition images, such as ignoring the connection between molar components at different scales, insufficient feature extraction, and difficult to deal with complex molar frequency coupling conditions.
A method for demolarization of ultra-high-definition image based on pyramid-learning bandpass filter is proposed. Through the pyramid-learning bandpass filter module (P-LBF) decoupling and removing molar patterns, and combining with the cross-layer feature fusion module (CLF) for feature alignment and fusion, effectively modeling and removing complex molar patterns.
Effective modeling and removal of coupled molar patterns in ultra-high-definition images is achieved, significantly improving image quality, and achieving the best demolar patterns on the UHDM dataset.
Smart Images

Figure CN120070260A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for removing moiré patterns from ultra-high-definition images based on a learnable pyramid band-pass filter, and particularly relates to the field of image moiré pattern effect processing based on deep learning technology. Background Art
[0002] With the development of the times, electronic devices can be seen everywhere in life, and the scenario of using electronic devices such as mobile phones to record information by taking pictures of the screen is becoming more and more common. However, due to the mismatch between the color filter array of the camera sensor and the sub-pixels of the screen or the interference between the color filter array and the high-frequency information, color interference signals of various colors and shapes often appear in the captured images. These interference signals usually appear in the form of stripes, meshes or ripples, with diverse and irregular colors. Such interference signals are called moiré patterns. Moiré patterns are almost distributed throughout the entire space of the image, and their frequencies are very complex, seriously reducing the quality of the image and having a great impact on downstream computer vision tasks such as object detection and image classification. Therefore, it is worth studying to restore a clean image from a moiré pattern image. Theoretically, moiré pattern removal can be regarded as a traditional image restoration process, whose purpose is to eliminate the moiré pattern noise in the photographed image and restore the color space information. However, different from other image restoration tasks such as image denoising, image deblurring, and image de-raining, since the mixing range of the moiré pattern and the original image information in the spatial and frequency domains is very wide, rather than simply existing in the high-frequency or low-frequency regions of the image, this makes the restoration of moiré pattern images more challenging than other image restoration tasks.
[0003] With the development of deep learning, it has been found that deep learning performs excellently in the field of image processing, and its application in image moiré pattern removal is also possible. Deep learning can use its powerful feature learning ability and pattern recognition ability to reduce or eliminate the influence of moiré patterns. At present, many researchers have also used deep learning methods to study the task of image moiré pattern removal and achieved certain results. However, there are still some deficiencies in the existing research on the processing of ultra-high-definition images. For example, DMCNN ignores the connection between moiré components at different scales, and at the same time, the features extracted at each scale are not sufficient. FHDeNet removes moiré patterns from both global and local scales, but ignores the connection between different scales, and it is difficult to handle the case where the moiré frequencies of ultra-high-definition images are coupled with each other. MopNet classifies moiré patterns, but the moiré patterns of ultra-high-definition images are very complex and cannot be well classified. MBCNN processes at a single scale at the same semantic level, has difficulty in modeling the complex moiré patterns of ultra-high-definition images, and lacks information interaction between features of different depths. Although ESDNet is designed for ultra-high-definition images, it lacks explicit modeling of moiré patterns and has limited ability in dealing with uncommon moiré patterns.
[0004] To solve the above problems, the present invention proposes a method for removing moiré from ultra-high-definition images based on a pyramid learnable band-pass filter, which can effectively model the coupled moiré in ultra-high-definition images to restore clear images. Summary of the Invention
[0005] Technical problems to be solved: Aiming at the problem that the existing moiré removal technology cannot effectively remove moiré from ultra-high-definition images, the present invention starts from the perspective of moiré frequency-domain aliasing in ultra-high-definition images and proposes a pyramid learnable band-pass filter module (P-LBF) with the same semantics. This module decouples and filters the coupled moiré at different scales, and then uses a feature alignment block (FA) and a feature fusion block (FF) respectively to perform feature alignment and effective fusion on the filtered features at multiple scales, so as to effectively model the coupled moiré. Secondly, the networks for processing ultra-high-definition images often have a relatively deep depth, and most of the existing methods operate on features at different depths by simple pixel-level addition or channel connection. To explore a more effective way of connecting features at different depths, the present invention proposes a cross-layer feature fusion module (CLF).
[0006] Implementation steps: The present invention proposes a method for removing moiré from ultra-high-definition images based on a pyramid learnable band-pass filter. The basic steps are as follows:
[0007] Step 1: Construct an image moiré removal network model based on a pyramid learnable band-pass filter. The network model is mainly based on an encoder-decoder network, and a cross-layer feature fusion module (CLF) is used to strengthen the connection between the encoder and the decoder.
[0008] Step 2: Given a moiré image, first reduce the computational cost through a pixel shuffle, and then extract shallow features through a convolutional layer.
[0009] Step 3: Input the shallow features into the encoder-decoder network.
[0010] Both the encoder module and the decoder module of the encoder-decoder network (each layer of the encoder module and the decoder module has the same structure) are three-layer structures, which are used to gradually remove moiré from the moiré-contaminated image under receptive fields of different scales. Bilinear interpolation downsampling is used between every two encoder levels to halve the feature size and gradually increase the receptive field. Similarly, bilinear interpolation is also used between every two decoders to double the feature resolution and gradually change it to the same size as the input image.
[0011] Step 4: In each layer of the encoder and decoder modules, first, the dilated residual dense block (DRDB) is used to further extract the input features. Then, the pyramid learnable band-pass filter module (P-LBF) is used to decouple and filter the coupled moiré patterns. Finally, the output features F of the DRDB R are added to the output features of the P-LBF to obtain the final output features of the current layer of the encoder / decoder module.
[0012] Step 5: To achieve more effective information interaction between the shallow features in the encoder and the deep features in the decoder, through the cross-layer feature fusion module (CLF), the network can better retain image details while removing moiré patterns.
[0013] Step 6: Train the constructed image de-moiré network model.
[0014] Step 7: The trained image de-moiré network model receives the images of the validation dataset that need to be de-moiré processed, and outputs the images after completing the de-moiré processing.
[0015] Furthermore, the specific structure of the pyramid learnable band-pass filter module P-LBF is as follows:
[0016] The P-LBF module consists of three parts: the pyramid filter block PF, the feature alignment block FA, and the feature fusion block FF. The PF is a pyramid structure composed of basic filter unit BU blocks, which decouples and filters the coupled moiré patterns at different scales. The multi-scale features are aligned through the FA block. The aligned features are effectively fused through the FF block.
[0017] The specific structure of the BU block is as follows: The input features are connected with the features before convolution in the channel dimension after passing through the dilated convolution with a convolution kernel size of 3*3 and the ReLU activation function, and then the above-mentioned repeated process is carried out 4 times. The dilation rates of the 5 dilated convolutions are 1, 2, 3, 2, 1 in sequence. After being processed by 5 dilated convolutions plus the ReLU activation function, the number of channels is changed to 64 through a convolution with a convolution kernel of 3*3 to obtain a 64-channel feature map. Then, the learnable band-pass weights and the 64-channel feature map are input together into the block frequency-domain inverse transform (FDIT) to learn the band-pass weights, and then through a 3*3 convolution and a feature scaling layer (FSL). The feature scaling layer linearly constrains the output of the convolutional layer to prevent the generation of excessive gradients. The FSL contains a learnable parameter s initialized to 0.1, which will be updated together with other learnable parameters in the network during the training phase. The features after passing through the FSL are added to the original input features of the BU block as the output features of the BU.
[0018] The specific structure of the PF block is as follows: The input features of the PF block first go through 2x downsampling and 4x downsampling respectively, and together with the original-scale features, three features F of different scales are obtained x1 , F x2 , F x3 . Then the three features of different scales are respectively fed into the BU block for decoupling filtering to remove the coupled moiré patterns. Finally, bilinear upsampling is used to restore the downsampled features to a scale comparable to the original scale and connect them in the channel dimension to obtain the pyramid-filtered feature F c .
[0019] The specific structure of the FA block is as follows: FA takes the pyramid-filtered feature F c as the input feature. First, global average pooling is used in the channel dimension to obtain a 1D vector v c , and then three fully connected layers and a sigmoid layer are used to learn the weight distribution in the channel dimension. Finally, the learned weight distribution is multiplied by the feature F c to obtain the aligned feature F c′ .
[0020] The specific structure of the FF block: First, average pooling is used in the H dimension and the W dimension respectively to obtain one-dimensional features v h , v w , and then a sigmoid layer is used to obtain the weight distribution in the H and W dimensions. Immediately afterwards, F c′ is multiplied by the weights in different dimensions respectively to obtain the features F h and F w . Finally, F h , F w and F c′ are added together and fused using a 1x1 convolution to obtain the output feature of FF
[0021] Furthermore, the specific structure of the CLF module is as follows:
[0022] CLF accepts two input features, which are respectively the output features F of the encoder of the i-th (i = 1, 2) layer i E and the feature obtained by bilinearly upsampling the output feature of the decoder of the (i + 1)-th (i = 1, 2) layer by two times Input the feature F of the i-th level encoder i E into a 1×1 convolutional layer, an LReLU layer, and a 1×1 convolutional layer to generate a scaling factor γ, and input the feature F i E into a 1×1 convolutional layer, an LReLU layer, and a 1×1 convolutional layer to generate a translation factor β. The dimensions of the parameters γ and β are the same as have the same dimension and are used to guide the decoder features transformation. Specifically, the features are scaled and translated through the following formula:
[0023]
[0024] to obtain the adjusted features To better retain image details, channel-level connection is performed between the output features F i E of the encoder and the output features of the modulated decoder . Finally, 3×3 convolution is used to fuse the connected features and reduce the feature dimension, so as to obtain the output features of the CLF. The output features of the CLF are used as the input features of the decoder of the i-th (i = 1, 2) layer for further processing.
[0025] Furthermore, the specific method of step 6 is as follows:
[0026] The training method of the network model is to first input the moire image I in the dataset moire ; then, the required moire-removed image I is obtained through the designed network model pred ; finally, the moire-removed image I output by the model is continuously optimized using the loss function pred , so that it gradually resembles the real clear image I in the prepared dataset clear .
[0027] During the training process, the loss function Loss uses the L1 function combined with the perceptual loss Lp as the loss function. The loss function is specifically expressed as:
[0028] Loss = L1(I pred , I clear ) + k * Lp(I pred , I clear )
[0029] where Lp(I pred , I clear ) = ||μ(I pred ) - μ(I clear )||, μ is the pre-trained VGG-16 network. k is a model hyperparameter used to balance the L1 loss and the perceptual loss. The training uses a multi-supervision strategy to supervise the step-by-step output of the decoder module, specifically expressed as:
[0030] Loss_total = Loss(I pred1 , I clear1 ) + Loss(I pred2 , Iclear2 ) + Loss(I pred3 , I clear3 )
[0031] where I pred1 , I pred2 , I pred3 are the predicted images of the three decoder layers of the decoder module respectively, and I clear1 , I clear2 , I clear3 are the real clear images of the corresponding sizes respectively.
[0032] Preferably, the dataset during training uses the publicly available UHDM dataset. The dataset contains 5000 groups of 4K images of different scenes. Each group of images contains two images, namely the moiré image I moire contaminated by moiré and the original clear image I clear . Among the 5000 groups of images, 4500 groups of images are used as the training dataset, and the other 500 groups of images are used as the validation dataset. Among them, the moiré image I moire is used as the input image data during the training process of the network model, while the original clear image I clear is used as the reference image for comparison with the model predicted image I pred during the training process of the network model.
[0033] Furthermore, the specific method of step 7 is as follows:
[0034] Load the weights of the image moiré removal network model trained in step 6 and update the parameters in the model. Secondly, use the moiré image I moise of the validation dataset as the input data and input it into the network model. The input data passes through the encoder and decoder in sequence to obtain the model output image I pred3 after moiré removal processing. Note that only the maximum image output by the model needs to be obtained for comparison with the real clear image during model validation, that is, I pred3 .
[0035] The beneficial effects of the present invention are as follows:
[0036] The present invention innovatively proposes a pyramid learnable band-pass filter (P-LBF) to perform multi-scale frequency domain filtering in a learnable manner to effectively model and remove complex and coupled moiré patterns in ultra-high definition images. It innovatively proposes a cross-layer feature fusion module (CLF) to achieve more effective feature information interaction. The present invention has achieved the best moiré removal effect on the ultra-high definition image dataset UHDM. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 This is the overall network architecture of the present invention and the structural diagram of the CLF module.
[0038] Figure 2 This is the specific structure of the P-LBF module proposed by the present invention. Detailed implementation manners
[0039] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0040] The present invention first makes the following definitions and explanations:
[0041] I moire : Moiré pattern image
[0042] I clear : Real clear image
[0043] I pred : Output image of the moiré removal network model
[0044] Implementation steps: The present invention proposes a super-high-definition image moiré removal method based on a pyramid learnable band-pass filter, and its basic steps are as follows:
[0045] Step 1: Dataset preparation;
[0046] Obtain the image dataset required for network training in Step 2. Download the publicly available UHDM dataset. The dataset contains 5000 groups of 4K images of different scenes. Each group of images contains two images, namely the moiré pattern image I moire contaminated by moiré and the original clear image I clear . Among the 5000 groups of images, 4500 groups of images are used as the training dataset, and the other 500 groups of images are used as the validation dataset. Among them, the moiré pattern image I moire is used as the input image data during the network model training process, and the original clear image I clear is used as the reference image for comparison with the model prediction image I pred during the network model training process.
[0047] Step 2: Construct an image moiré removal network model based on a pyramid learnable band-pass filter.
[0048] The network model is mainly based on the encoder-decoder network, and the cross-layer feature fusion module (CLF) is used to strengthen the connection between the encoder and the decoder.
[0049] Given a moiré image, we first perform a pixel shuffle to reduce computational cost, and then extract shallow features through a convolutional layer. The shallow features are then fed into the encoder-decoder network.
[0050] The encoder and decoder modules of the encoder-decoder network (each layer of the encoder and decoder modules uses the same structure) are both 3-layer structures, which are used to gradually remove moiré from images contaminated by moiré at different scales of receptive fields. Bilinear interpolation downsampling is used between every two encoder levels to halve the feature size to gradually increase the receptive field. Similarly, bilinear interpolation is also used between every two decoders to double the feature resolution and gradually become the same size as the input image.
[0051] In each layer of the encoder and decoder modules, the dilated residual dense block (DRDB) is first used to further extract the input features, and then the pyramid learnable bandpass filter module (P-LBF) is used to decouple and filter the coupled moiré patterns. Finally, the output feature F R The output features of P-LBF are added to obtain the final output features of the current layer of the encoder / decoder module.
[0052] In order to achieve more effective information interaction between shallow features in the encoder and deep features in the decoder, the cross-layer feature fusion module (CLF) is used to enable the network to better retain image details while removing moiré patterns.
[0053] The specific structure of the P-LBF module is as follows: The P-LBF module consists of three parts: a pyramid filter block (PF), a feature alignment block (FA), and a feature fusion block (FF). The PF is a pyramid structure composed of basic filter unit (BU) blocks, which performs decoupled filtering on coupled moiré patterns at different scales. Since the features after frequency domain filtering at different scales do not have explicit domain consistency constraints and may have domain deviations, we then designed an FA block to align multi-scale features. In order to strengthen the correlation between the features after decoupled filtering at different scales, we designed an FF block to effectively fuse the aligned features.
[0054] The specific structure of the BU block is as follows: The input features are connected with the features before convolution in the channel dimension after passing through the dilated convolution with a kernel size of 3*3 and the ReLU activation function, and then go through the above-mentioned repeated process 4 times. The dilation rates of the 5 dilated convolutions are 1, 2, 3, 2, 1 in sequence. After passing through the 5 dilated convolutions plus the ReLU activation function, the number of channels is changed to 64 through a convolution with a kernel of 3*3, that is, a feature map with 64 channels is obtained. Then, the learnable band-pass weights and the 64-channel feature map are input into the block frequency-domain inverse transform (FDIT) to learn the band-pass weights, and then through a 3*3 convolution and a feature scaling layer (FSL). The feature scaling layer performs a linear constraint on the output of the convolutional layer to prevent the generation of excessive gradients. FSL contains a learnable parameter s initialized to 0.1, and this parameter will be updated together with other learnable parameters in the network during the training phase. The features after passing through FSL are added to the original input features of the BU block as the output features of the BU block.
[0055] The specific structure of the PF block is as follows: The input features of the PF block are first downsampled by 2 times and 4 times respectively and then together with the original-scale features to obtain 3 features F of different scales x1 , F x2 , F x3 . Then, the 3 features of different scales are respectively given to the BU block for decoupling filtering to remove the coupled moiré patterns. Finally, bilinear upsampling is used to restore the downsampled features to a scale comparable to the original scale and connect them in the channel dimension to obtain the pyramid-filtered feature F c .
[0056] The specific structure of the FA block is as follows: FA uses the pyramid-filtered feature F c as the input feature. First, global average pooling is used in the channel dimension to obtain a 1D vector v c , and then 3 fully connected layers and a sigmoid layer are used to learn the weight distribution in the channel dimension. Finally, the learned weight distribution is multiplied by the feature F c to obtain the aligned feature F c′ .
[0057] The specific structure of the FF block is as follows: Since the DCT transform used in the BU block is separable in H and W, we calculate the spatial attention separately in the H and W dimensions to ensure that the calculated results match the filtering results of the BU block. We first use average pooling in the H dimension and the W dimension respectively to obtain one-dimensional features v h , v w , and then use the sigmoid layer to obtain the weight distributions in the H and W dimensions. Immediately afterwards, F c′ is multiplied by the weights in different dimensions respectively to obtain the features F h and Fw Finally, we add F h , F w and F c′ together and use a 1x1 convolution to fuse them to obtain the output feature of FF.
[0058] Specific structure of the CLF module: When processing ultra-high-definition moiré images, the network generally needs to have a relatively deep depth to gradually process complex moiré patterns. However, increasing the network depth will also lead to a gradual loss of image details and unstable network gradient transmission. Therefore, we propose to gradually modulate the deep features with shallow features to solve this problem.
[0059] Specifically, CLF accepts two input features, which are the output features F i E of the encoder at the i-th (i = 1, 2) layer and the features obtained by bilinearly upsampling the output features of the decoder at the (i + 1)-th (i = 1, 2) layer by a factor of two. Input the feature F i E of the i-th level encoder into a 1×1 convolutional layer, an LReLU layer, and a 1×1 convolutional layer to generate a scaling factor γ, and input the feature F i E into a 1×1 convolutional layer, an LReLU layer, and a 1×1 convolutional layer to generate a translation factor β. The dimensions of the parameters γ and β are the same as those of the feature and are used to guide the transformation of the decoder feature Specifically, the feature is scaled and translated through the following formula:
[0060]
[0061] to obtain the adjusted feature To better preserve image details, we perform channel-level connection between the output feature F i E of the encoder and the output feature of the modulated decoder Finally, we use a 3×3 convolution to fuse the connected features and reduce the feature dimension, thereby obtaining the output feature of CLF. The output feature of CLF is used as the input feature of the decoder at the i-th (i = 1, 2) layer for further processing.
[0062] Specific structure of the complete network: Given a moiré image, first, a pixel shuffle is used to reduce the computational cost, and then a convolutional layer is used to extract shallow features. Then, the shallow features are input into an encoder-decoder network. Our main network framework consists of 3 encoders and 3 decoders. Bilinear interpolation downsampling is used between every two encoder levels to halve the feature size and gradually increase the receptive field. Similarly, bilinear interpolation is also used between every two decoders to double the feature resolution and gradually change it to the same size as the input image.
[0063] Inside each encoder / decoder, we effectively process complex and coupled moiré patterns through the P-LBF module. To achieve more effective information interaction between the shallow features in the encoder and the deep features in the decoder, we designed a CLF module to enable the network to better retain image details while removing moiré patterns.
[0064] Step 3: Train the constructed image de-moiré network model using the image dataset obtained in Step 1;
[0065] The training method of the network model is to first input the moiré image I prepared in Step 1 moire ; then, the required de-moiré processed image I is obtained through the designed network model pred ; finally, the loss function is used to continuously optimize the de-moiré processed image I output by the model pred , making it gradually similar to the real clear image I in the dataset prepared in Step 1 clear .
[0066] During the training process, the loss function Loss uses the L1 function combined with the perceptual loss Lp as the loss function. The loss function is specifically expressed as:
[0067] Loss = L1(I pred , I clear ) + k * Lp(I pred , I clear )
[0068] where Lp(I pred , I clear ) = ||μ(I pred ) - μ(I clear )||, μ is the pre-trained VGG-16 network. k is a model hyperparameter used to balance the L1 loss and the perceptual loss. Training uses a multi-supervision strategy to supervise the gradual output of the decoder module, specifically expressed as:
[0069] Loss_total = Loss(I pred1 , I clear1 ) + Loss(I pred2 , I clear2 ) + Loss(I pred3 , I clear3 )
[0070] where I pred1 , I pred2 , and I pred3 are the predicted images of the three decoder levels of the decoder module respectively, and I clear1 , I clear2 , and I clear3 are the real clear images of the corresponding sizes respectively.
[0071] Step 4: The trained moiré removal network model for images receives the validation dataset images that need to be processed for moiré removal in Step 1, and outputs the images after completing the moiré removal process.
[0072] Load the weights of the moiré removal network model for images trained in Step 3 to update the parameters in the model. Secondly, use the moiré images I moire in the validation dataset in Step 1 as input data and input it into the network model. The input data passes through the encoder and decoder in sequence to obtain the model output image I pred3 after moiré removal. Note that when validating the model, only the maximum image output by the model needs to be compared with the real clear image, that is, I pred3 .
[0073]
[0074] Table 1
[0075] We conducted a comparative experiment on the UHDM dataset for the method in the present invention and other related works. The experimental results are shown in Table 1. It can be seen from the data results in Table 1 that the moiré removal effect of the present invention on the UHDM dataset is better than other existing methods, demonstrating the effectiveness of the present invention.
[0076] The above content is a further detailed description of the present invention in combination with specific / preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, they can still make several substitutions or modifications to these described embodiments, and these substitution or modification methods should all be regarded as belonging to the protection scope of the present invention.
[0077] The parts not detailed in the present invention belong to the well-known technology in the art.
Claims
1. An ultra-high-definition image de-moiré method based on a pyramid learnable bandpass filter, characterized in that: The steps include: Step 1: Construct an image de-moiré network model based on a pyramid learnable bandpass filter. The network model is based on the encoder-decoder network as the main framework, and the cross-layer feature fusion module CLF is used to strengthen the connection between the encoder and the decoder. Step 2: Given a moiré image, first use a pixel reorganization to reduce the computational cost, and then use a convolutional layer to extract shallow features; Step 3: Input shallow features into the encoder-decoder network; The encoder and decoder modules of the encoder-decoder network are both three-layer structures, which are used to gradually remove moiré from images contaminated by moiré at different scales of receptive fields. Bilinear interpolation downsampling is used between every two encoder levels to halve the feature size to gradually increase the receptive field. Similarly, bilinear interpolation is also used between every two decoders to double the feature resolution and gradually become the same size as the input image. Step 4: In each layer of the encoder and decoder modules, the dilated residual dense block DRDB is first used to further extract the input features, and then the pyramid learnable bandpass filter module P-LBF is used to decouple and filter the coupled moiré patterns. Finally, the output feature F R Add the output features of P-LBF to get the final output features of the current layer of the encoder / decoder module; Step 5: In order to achieve more effective information interaction between shallow features in the encoder and deep features in the decoder, the cross-layer feature fusion module CLF is used to enable the network to better retain image details while removing moiré patterns. Step 6: Train the constructed image de-moiré network model; Step 7: The trained image demoiré network model receives the verification dataset image that needs to be demoiré processed, and outputs the image after completing the demoiré processing.
2. The method for removing moiré from ultra-high-definition images based on a pyramid learnable bandpass filter according to claim 1, characterized in that: The specific structure of the pyramid learnable bandpass filter module P-LBF is as follows: The P-LBF module consists of three parts: pyramid filter block PF, feature alignment block FA and feature fusion block FF. PF is a pyramid structure composed of basic filter units BU blocks, which performs decoupling filtering on coupled moiré patterns at different scales. The multi-scale features are aligned through the FA block. The aligned features are effectively fused through the FF block. The specific structure of the BU block is as follows: the input features are connected to the features before convolution in the channel dimension after undergoing a dilated convolution with a convolution kernel size of 3*3 and a ReLU activation function, and then the above process is repeated 4 times. The dilation rates of the 5 dilated convolutions are 1, 2, 3, 2, and 1 respectively. After 5 dilated convolutions and ReLU activation function processing, a convolution with a convolution kernel of 3*3 is used to change the number of channels to 64, that is, a 64-channel feature map is obtained. Then the learnable bandpass weights and the 64-channel feature map are input together into the block frequency domain inverse transform FDIT to learn the bandpass weights, and then a 3*3 convolution and feature scaling layer FSL are passed. The output of the convolution layer is linearly constrained by the feature scaling layer to prevent excessive gradients. FSL contains a learnable parameter s initialized to 0.1, which will be updated together with other learnable parameters in the network during the training phase. The features after FSL are added to the original input features of the BU block as the output features of the BU. The specific structure of the PF block is as follows: The input features of the PF block are first downsampled by 2 times and 4 times respectively, and combined with the original scale features to obtain three different scale features F x1 , F x2 , F x3 ; Then the features of three different scales are respectively sent to the BU block for decoupling filtering to remove the coupled moiré; Finally, bilinear upsampling is used to restore the downsampled features to a scale equivalent to the original scale and connected in the channel dimension to obtain the pyramid filtered features F c ; The specific structure of the FA block is as follows: FA takes the pyramid-filtered feature F c As input features; first, global average pooling is used on the channel dimension to obtain a 1-dimensional vector v c , then use 3 fully connected layers and sigmoid layers to learn the weight distribution on the channel dimension, and finally combine the learned weight distribution with the feature F c Multiply to get the aligned feature F c′ ; The specific structure of the FF block: First, use average pooling on the H dimension and W dimension to obtain the one-dimensional feature v h ,v w , and then use the sigmoid layer to obtain the weight distribution on the H and W dimensions, and then use F c′ Multiply the weights on different dimensions to get feature F h and F w ; Finally, F h ,F w and F c′ The output features of FF are obtained by adding and fusing them using a 1x1 convolution.
3. The method for removing moiré from ultra-high-definition images based on a pyramid learnable bandpass filter according to claim 1, characterized in that: The specific structure of the CLF module is as follows: CLF accepts two input features, namely the output feature F of the i-th layer encoder encoder i E The output features of the i+1th layer decoder are bilinearly upsampled by twice. Where i = 1, 2; the feature F of the i-th level encoder i E Input to the 1×1 convolution layer, LReLU layer and 1×1 convolution layer to generate the scaling factor γ, and transform the feature F i E Input to the 1×1 convolution layer, LReLU layer and 1×1 convolution layer to generate the translation factor β; the dimensions of parameters γ and β are the same as The same dimension as that used to guide the decoder features Specifically, the feature is transformed by the following formula To zoom and pan: Get the adjusted features In order to better preserve image details, the output feature F of the encoder i E and the modulated decoder output characteristics Channel-level connections are performed between them; finally, 3×3 convolution is used to fuse the connected features and reduce the feature dimension to obtain the output features of CLF, which are further processed as the input features of the i-th layer decoder.
4. The method for removing moiré from ultra-high-definition images based on a pyramid learnable bandpass filter according to claim 1, characterized in that: Step 6: The training method of the network model is to first input the mole image I in the dataset moire ; Then, the desired image I after de-moiré processing is obtained through the designed network model pred ; Finally, the loss function is used to continuously optimize the model output after the de-moiré processed image I pred , making it gradually similar to the real clear image I in the prepared dataset clear ; During the training process, the loss function Loss uses the L1 function combined with the perceptual loss Lp as the loss function; the loss function is specifically expressed as: Loss=L1(I pred ,I clear )+k*Lp(I pred ,I clear ) Among them, Lp(I pred ,I clear )=||μ(I pred )-μ(I clear )||, μ is the pre-trained VGG-16 network; k is the model hyperparameter used to balance L1 loss and perceptual loss; the training adopts a multi-supervision strategy to supervise the progressive output of the decoder module, which is specifically expressed as: Loss_total=Loss(I pred1 ,I clear1 )+Loss(I pred2 ,I clear2 )+Loss(I pred3 ,I clear3 ) Among them I pred1 , I pred2 , I pred3 are the predicted images of the three decoder layers of the decoder module, I clear1 , I clear2 , I clear3 They are real clear images of corresponding sizes respectively.
5. The method for removing moiré from ultra-high-definition images based on a pyramid learnable bandpass filter according to claim 4, characterized in that: The training dataset uses the public dataset UHDM dataset, which contains 5000 sets of 4K images of different scenes. Each set of images contains two images, namely, the moiré image I contaminated by moiré and the moiré image I contaminated by moiré. moire and the original clear image I clear ; 4500 of the 5000 images are used as training data sets, and the other 500 images are used as verification data sets; among them, the moiré image I moire As the input image data in the network model training process, the original clear image I clear As the network model training process, it is used to predict the image I pred Reference image for comparison.
6. The method for removing moiré from ultra-high-definition images based on a pyramid learnable bandpass filter according to claim 4 or 5, characterized in that: Step 7: Load the image de-moiré network model weights trained in step 6 and update the parameters in the model; secondly, validate the moiré image I of the dataset moire The input data is passed into the network model as input data. The input data passes through the encoder and decoder in turn to obtain the model output image I after de-moiré processing. pred3 , Note that when verifying the model, it is only necessary to obtain the maximum image output by the model and compare it with the real clear image, that is, I pred3 .
Citation Information
Patent Citations
Image restoration method based on deep learning image moire elimination
CN112184591A
Moire removing system and method based on learnable frequency domain prior
CN113554566A
Medical image segmentation system
CN117593275A
Image moire removing method based on multi-scale band-pass filter
CN118195932A
Image moire removing method based on RGB channel trilateral compensation
CN118570075A