Network model training method for removing image banding noise
By adopting a network model with a hybrid structure of CNN and Transformer, combining deep feature extraction, feature fusion and upsampling operations, the problems of low training efficiency and low image clarity in the prior art are solved, and a more efficient image denoising effect is achieved.
Patent Information
- Application Number
- CN202311789764.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-25
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is inefficient and the processed image sharpness is not high when constructing and training a network model for removing image band noise.
A network model with a hybrid structure of CNN and Transformer is used to train the model through multiple deep feature extraction, feature fusion and upsampling operations, combined with the calculation of the loss function. The specific steps include preprocessing the noise image and the clear image, establishing a training set, and performing multiple rounds of feature extraction and fusion through the codec network, and finally outputting the denoised clear image.
The image denoising effect is improved, and the network model's ability to remove strip noise in the image is enhanced. At the same time, the image edges and detailed structure are retained, achieving higher visual performance and peak signal-to-noise ratio.
Smart Images

Figure CN120219209A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of digital image processing, and particularly relates to a method for training a network model for removing strip noise in images. Background Art
[0002] In the era of big data, image information has become an essential information source in daily life. In practical applications, images are often contaminated with noise due to the influence of the surrounding shooting environment and imaging devices. To solve the problem of image noise pollution, intelligent image denoising technology has emerged and developed rapidly. In particular, using deep learning network models to quickly process noisy images has become a current development hotspot.
[0003] The denoising algorithm based on deep learning regards the image denoising problem as a regression problem, makes full use of the network architecture to automatically learn more essential features in the data from the training data, has a greater representation ability than traditional model-based methods, and can well protect the edges and details in the image. Therefore, using the denoising algorithm of deep learning can achieve better denoising effects. However, the difficulty of this technology lies in constructing an efficient neural network model and training it effectively. Currently, the industry has made many attempts in this field, but the current network model training methods and effects are not ideal. The training efficiency of the network model is low, and the image processing ability is not satisfactory, which is an urgent problem to be solved. Summary of the Invention
[0004] In view of the above analysis, an embodiment of the present invention aims to provide a method for training a network model for removing strip noise in images, so as to solve the problems of low training efficiency of the network model for strip noise images and still very low clarity of the processed images after model training in the prior art.
[0005] On the one hand, an embodiment of the present invention provides a method for training a network model for removing strip noise in images, including the following steps:
[0006] Preprocess the noisy image and the corresponding clear image to establish a training set;
[0007] Establish a network model, input the noisy images in the training set into the network model, and after multiple deep feature extraction, feature fusion and upsampling operations of the network model, output the denoised clear image, where
[0008] The network model adopts a hybrid structure of CNN and Transformer, including an input layer, an encoding and decoding network, and an output layer; the encoding and decoding network includes multiple encoding and decoding units; in each encoding and decoding unit, the encoder and the decoder are both provided with corresponding convolutional layers, which are used to perform convolutional operations on the data input to the corresponding encoder / decoder and then connect with the output residuals of the corresponding encoder / decoder as the input of the subsequent decoder / encoder;
[0009] Calculate the loss function based on the denoised clear image and the corresponding clear image in the training set, and complete the training of the network model when the value of the loss function reaches the threshold.
[0010] Based on a further improvement of the above method, the preprocessing of the noisy image and the corresponding clear image and the establishment of the training set include:
[0011] Select the same number of images with strip noise and the corresponding clear images;
[0012] Randomly crop the noisy images into noisy image patches of the same size, and crop the corresponding clear images at the same position to obtain clear image patches;
[0013] Randomly perform horizontal flipping or vertical flipping operations on the noisy image patches and the clear image patches;
[0014] Use the noisy image patches as the training input images and the clear image patches as the training control labels to complete the construction of the training set.
[0015] Based on a further improvement of the above method, the input layer is used to perform shallow feature extraction on the noisy image to obtain a shallow feature image;
[0016] The encoding and decoding network is used to perform multiple rounds of deep feature extraction, feature fusion, and upsampling on the shallow feature image and then output;
[0017] The output layer is used to fine-tune the image output by the encoding and decoding network and then output the denoised clear image.
[0018] Based on a further improvement of the above method, the feature fusion module includes a CAT dimension splicing module, a convolutional layer, and a CBAM convolutional attention module connected in sequence, where,
[0019] The CAT dimension splicing module receives the images output in parallel by the Transformer processing branch and the CFEB processing branch of the encoder, and performs channel dimension splicing;
[0020] The convolutional layer performs convolutional operations on the image to complete feature extraction;
[0021] The CBAM convolutional attention module includes channel attention processing and spatial attention processing, and outputs the processed image.
[0022] Based on the further improvement of the above method, the decoder is a network composed of multiple convolutional layers connected by residual connections, specifically including:
[0023] 4m convolutional layers and m - 1 upsampling layers;
[0024] One upsampling layer is connected after every 4 convolutional layers;
[0025] The upsampling layer is used to reduce the dimension of the image and enlarge the size of the image;
[0026] The first convolutional layer reduces the dimension of the noise image and reduces the number of channels;
[0027] For the second to the penultimate convolutional layer, the input and output of each layer are summed and then input to the next layer;
[0028] After the convolution operation of the last convolutional layer, the processed image is output.
[0029] Based on the further improvement of the above method, the encoding and decoding unit includes an encoder, a feature fusion module, a decoder connected in sequence, a convolutional layer parallel to the encoder and the feature fusion module, and a transposed convolutional layer parallel to the decoder;
[0030] When there are n encoding and decoding units in the network, let the encoder be e i , the feature fusion module be c i , the decoder be d i , the convolutional layer be l i , the transposed convolutional layer be t i , i = 1, 2,... n, where
[0031] The data input to the encoder e i is synchronously input to the convolutional layer l i . After the serial processing of the encoder e i and the feature fusion module c i , the output is summed with the output of the convolutional layer l i to obtain the first summation data, and the first summation data is simultaneously used as the input of the decoder d i and the transposed convolutional layer t i ; the first summation data is also summed with the output of the feature fusion module c n-i+1 and the output of the convolutional layer l n-i+1 to obtain the second summation data, and the second summation data is simultaneously used as the input of the decoder d n-i+1 and the transposed convolutional layer l n-i+1 ;
[0032] Decoder di The output and the transposed convolutional layer t i The outputs are summed to obtain the third summed data, and the third summed data is simultaneously used as the input of the encoder e i+1 and the convolutional layer l i+1 ; The third summed data is also input to the decoder d n-i and the transposed convolutional layer t n-i The outputs are summed to obtain the fourth summed data, and the fourth summed data is simultaneously used as the input of the encoder e n-i+1 and the convolutional layer l n-i+1 ;
[0033] The output of the decoder d n is summed with the output of the transposed convolutional layer t n as the output data of the encoding and decoding network.
[0034] Based on a further improvement of the above method, in the encoding and decoding unit,
[0035] The convolutional layer is used to transform the dimension of the input image that has not been encoded and feature fused into the same dimension as the image output by the feature fusion module;
[0036] The transposed convolutional layer is used to transform the dimension of the input image that has not been decoded into the same dimension as the image output by the decoder.
[0037] Based on a further improvement of the above method, the encoder includes parallel Transformer processing branches and CFEB processing branches;
[0038] The Transformer processing branch is used to extract global features and deep semantic features; the CFEB processing branch is used to extract local image features and image detail features;
[0039] The two processing branches process the input image to the encoder in parallel and output synchronously.
[0040] Based on a further improvement of the above method, the Transformer processing branch includes an optimized network structure composed of multiple sequentially connected Transformer basic units, where
[0041] The first Transformer basic unit includes a first convolutional layer, a dimension compression transposed layer, a first normalization layer, multi-head self-attention, a second normalization layer, and a multi-layer perceptron connected in sequence;
[0042] The second to the last Transformer basic units all include a first normalization layer, multi-head self-attention, a second normalization layer, and a multi-layer perceptron connected in sequence;
[0043] The last Transformer basic unit further includes: a dimensionality compression transpose layer, a second convolutional layer, and a third convolutional layer;
[0044] For each Transformer basic unit, residual operations are introduced at the end of the multi-head attention and at the end of the multi-layer perceptron respectively.
[0045] Based on a further improvement of the above method, the CFEB processing branch includes multiple convolutional layers and downsampling layers. Among them, residual connections are used between the convolutional layers. Specifically:
[0046] 4m convolutional layers and m - 1 downsampling layers;
[0047] After every 4 convolutional layers, a downsampling layer is connected. The downsampling layer is used to increase the number of channels and reduce the image size;
[0048] The first convolutional layer increases the number of channels of the noise image, while the length and width of the image remain unchanged;
[0049] For the second to the penultimate convolutional layer, the sum of the input and output of each layer is calculated and then input into the next layer;
[0050] After the convolutional operation of the last convolutional layer, the processed image is output.
[0051] The purpose of the present invention is to provide an image denoising method and system, which can improve the visual performance of the reconstructed image, as well as the peak signal-to-noise ratio and structural similarity. The present invention provides a network model training method for removing image banding noise, which can construct an effective training set to effectively train the network model, enhance the ability of the network model to remove strip noise in the image while retaining the image edges and fine texture structures, and obtain clear images. The training method adopted by the present invention can train a more general denoising image model, improve the simplicity of image denoising, and has practical application value.
[0052] In the present invention, the above technical solutions can also be combined with each other to achieve more preferred combination schemes. Other features and advantages of the present invention will be described in the subsequent specification. Moreover, some advantages can be made obvious from the specification, or can be understood by implementing the present invention. The purpose and other advantages of the present invention can be realized and obtained through the content specifically pointed out in the specification and the drawings. Description of the Drawings
[0053] The accompanying drawings are only for the purpose of showing specific embodiments and are not considered to be a limitation of the present invention. To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0054] Figure 1 It is a flowchart of a network model training method provided by an embodiment of the present invention;
[0055] Figure 2 It is a flowchart for establishing a noise image training set in an embodiment of the present invention;
[0056] Figure 3 It is the overall structure diagram of a residual connection image denoising neural network model provided by an embodiment of the present invention;
[0057] Figure 4 It is the hierarchical structure diagram of the Transformer unit provided by an embodiment of the present invention;
[0058] Figure 5 It is a schematic diagram of network optimization including 2 Transformer units provided by an embodiment of the present invention;
[0059] Figure 6 It is the neural network structure diagram of the encoder CFEB module provided by an embodiment of the present invention;
[0060] Figure 7 It is the structure diagram of the feature fusion module with two parallel branches of the encoder as inputs provided by an embodiment of the present invention;
[0061] Figure 8 It is the structure diagram of the channel attention module in an embodiment of the present invention;
[0062] Figure 9 It is the structure diagram of the spatial attention module in an embodiment of the present invention;
[0063] Figure 10 It is the structure diagram of the CBAM branch in an embodiment of the present invention;
[0064] Figure 11 It is the neural network structure diagram of the decoder provided by an embodiment of the present invention. Specific Embodiments
[0065] The preferred embodiments of the present invention will be specifically described below in conjunction with the accompanying drawings. The accompanying drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, but are not used to limit the scope of the present invention.
[0066] In one embodiment of the present invention, the dataset used is an express label image dataset with strip noise, and the dataset contains 3,500 pairs of noisy images and clear images.
[0067] A network model training method for removing image strip noise includes the following steps:
[0068] Preprocess the noisy image and the corresponding clear image to establish a training set;
[0069] Establish a network model, input the noisy images in the training set into the network model, and after multiple deep feature extraction, feature fusion and upsampling operations of the network model, output the denoised clear image, where
[0070] The network model adopts a hybrid structure of CNN and Transformer, including an input layer, an encoding and decoding network, and an output layer; the encoding and decoding network includes multiple encoding and decoding units; in each encoding and decoding unit, the encoder and decoder are both provided with corresponding convolutional layers, which are used to perform convolution operations on the data input to the corresponding encoder / decoder and then connect the output residuals of the corresponding encoder / decoder as the input of the subsequent decoder / encoder;
[0071] Calculate the loss function based on the denoised clear image and the corresponding clear image in the training set, and when the value of the loss function reaches the threshold, the training of the network model is completed.
[0072] Specifically, as Figure 1 shown in the main flow chart of the present invention, it includes the following steps:
[0073] Step 1: Preprocess the noisy image and the corresponding clear image to establish a training set.
[0074] Further, the construction process includes:
[0075] Select the same number of images with strip noise and the corresponding clear images;
[0076] Randomly crop the noisy images into noisy image patches of the same size, and crop the corresponding clear images at the same position to obtain clear image patches;
[0077] Randomly perform horizontal flipping or vertical flipping operations on the noisy image patches and the clear image patches;
[0078] Use the noisy image patches as training input images and the clear image patches as training control labels to complete the construction of the training set.
[0079] In this embodiment, 3,000 express label images with strip noise and 3,000 corresponding clear images are selected;
[0080] The preprocessing process is as follows Figure 2 and includes:
[0081] S11. Randomly crop the noisy image into image patches of the same size, and crop the corresponding clear image at the same position to obtain clear image patches;
[0082] The image is represented by three dimensions: channel, length, and width. In this embodiment, the image is cropped to a size of (3, 160, 160), that is, the cropped image patch has 3 channels, a length of 160 pixels, and a width of 160 pixels.
[0083] S12. Randomly perform horizontal flipping or vertical flipping operations on the noisy image patches and clear image patches for data augmentation;
[0084] S13. Use the noisy image patches as input images and the clear image patches as labels to construct a training set.
[0085] Step 2: Establish a network model. Input the strip-noisy images in the training set into the network model. After multiple deep feature extractions, feature fusions, and upsampling operations by the network model, output the denoised clear image. Among them,
[0086] The network model adopts a hybrid structure of CNN and Transformer, including an input layer, an encoding-decoding network, and an output layer; the encoding-decoding network includes multiple encoding-decoding units; in each encoding-decoding unit, the encoder and decoder are both provided with corresponding convolutional layers for performing convolution operations on the data input to the corresponding encoder / decoder and connecting the output residuals of the corresponding encoder / decoder as the input of the subsequent decoder / encoder.
[0087] In this embodiment, during the model training stage, every 24 images are used as a processing batch for batch training.
[0088] Figure 3 shows the overall structure of the network model adopting a hybrid structure of CNN and Transformer. Specifically, the network model includes: an input layer for performing shallow feature extraction on the noisy image to obtain a shallow feature image;
[0089] an encoding-decoding network for performing multiple rounds of deep feature extraction, feature fusion, and upsampling on the shallow feature image and then outputting;
[0090] an output layer for fine-tuning the image output by the encoding-decoding network and then outputting the clear image with noise removed.
[0091] The network model structure is introduced in detail below.
[0092] Specifically, the input layer is used to extract shallow features from the noise image to obtain a shallow feature image.
[0093] In this embodiment, the input layer is composed of a first convolutional operation layer with a convolutional kernel of 3×3, which is used to extract shallow features from the noise image and output an image with output channels, length, and width of (3, 160, 160) to the next layer.
[0094] The encoding and decoding network is used to perform multiple rounds of deep feature extraction, feature fusion, and upsampling on the shallow feature image and then output.
[0095] Furthermore, each layer of the encoding and decoding units is connected in sequence, the structures of each layer of the encoding and decoding units are the same, and the layers of the encoding and decoding units are also connected in a residual manner.
[0096] Furthermore, the encoding and decoding unit includes an encoder, a feature fusion module, and a decoder connected in sequence, a convolutional layer parallel to the encoder and the feature fusion module, and a transposed convolutional layer parallel to the decoder;
[0097] When there are n layers of encoding and decoding units in the network, let the encoder be e i , the feature fusion module be c i , the decoder be d i , the convolutional layer be l i , the transposed convolutional layer be t i , i = 1, 2,... n, where
[0098] The data input to the encoder e i is synchronously input to the convolutional layer l i . After being serially processed by the encoder e i and the feature fusion module c i , the output is summed with the output of the convolutional layer l i to obtain the first summation data, and the first summation data is simultaneously used as the input of the decoder d i and the transposed convolutional layer t i ; the first summation data is also summed with the output of the feature fusion module c n-i+1 and the output of the convolutional layer l n-i+1 to obtain the second summation data, and the second summation data is simultaneously used as the input of the decoder d n-i+1 and the transposed convolutional layer l n-i+1 ;
[0099] The output of the decoder d i is summed with the output of the transposed convolutional layer t i to obtain the third summation data, and the third summation data is simultaneously used as the input of the encoder e i+1 and the convolutional layer l i+1 ; the third summation data is also summed with the output of the decoder d n-i, the transposed convolutional layer t n-i The outputs are summed to obtain the fourth sum data, and the fourth sum data is simultaneously used as the input of the encoder e n-i+1 , the convolutional layer l n-i+1 ;
[0100] The output of the decoder d n , and the output of the transposed convolutional layer t n are summed as the output data of the encoding and decoding network.
[0101] Specifically, the network contains multiple encoding and decoding units composed of an encoder, a feature fusion module, a decoder, a convolutional layer, and a transposed convolutional layer. It is necessary to comprehensively consider network parameters, inference speed, and image denoising effect, and analyze the experimental results to determine the number of encoding and decoding units in the network. If too many encoding and decoding units are set, the network parameters will be too large and the inference speed will decrease. If too few are set, the image features will not be extracted sufficiently and too much information will be lost, which is not conducive to the subsequent decoder for image detail reconstruction, and ultimately leads to a blurred result image output by the network.
[0102] An important indicator to measure the denoising effect of a neural network is PSNR. PSNR stands for "Peak Signal-to-Noise Ratio", and its Chinese meaning is peak signal-to-noise ratio. PSNR is defined based on MSE (mean squared error). For a given original image I of size m×n and its noisy image K after adding noise, its MSE can be defined as:
[0103]
[0104] where m and n respectively represent the length and width of the picture, i and j represent the abscissa and ordinate of a pixel point, and I(i,j) represents the pixel value at the position (i, j) in the image I. Then PSNR can be defined as:
[0105]
[0106] where Max I is the maximum pixel value of the image, and the unit of PSNR is dB. If each pixel is represented by 8-bit binary, its value is 2 8 -1 = 255. If the image is a grayscale image, the PSNR value can be directly calculated using the above formula; if the image is a color image, calculate the MSE value of each of the three channels of the RGB image and then find the average value MSE1, and then substitute MSE1 into the PSNR formula to calculate the PSNR of the color picture. In the experiment, by analyzing the PSNR values of the output result images of neural networks with different repetition times of the encoder, feature fusion module, and decoder, it is found that setting the repetition times to 3 is more appropriate. If the repetition times are increased further, the effect on improving the PSNR value is very limited.
[0107] Preferably, the number of layers of the encoding and decoding unit is set to 3 layers.
[0108] Specifically, taking the number of layers of the encoding and decoding unit being set to 3 layers as an example, the processing process of data within the encoding and decoding unit will be described in detail below.
[0109] From Figure 3 It can be seen that in this embodiment, the image (3, 160, 160) after being processed by the input layer is input to the encoder e1 and the convolutional layer l1. After being serially processed by the encoder e1 and the feature fusion module c1, an image with the number of channels, length, and width of (128, 40, 40) is output, which is summed with the output of the convolutional layer l1. The number of channels, length, and width of the image after being processed by the convolutional layer l1 is also (128, 40, 40); moreover, the image after the summation processing of the two is not only output to the decoder d1 and the transposed convolutional layer t1 for processing, but also summed with the result output after being processed by the feature fusion module c3 and the output result after convolution with the convolutional layer l3, and then input to the decoder d3 and the transposed convolutional layer t3.
[0110] The output image after being processed by the decoder d1 is an image with the number of channels, length, and width of (3, 160, 160), which is summed with the output image after the convolutional operation of the transposed convolutional layer t1. The number of channels, length, and width of the image after being processed by the transposed convolutional layer t1 is also (3, 160, 160), and it is not only output to the encoder e2 and processed with the convolutional layer l2, but also summed with the output results of the decoder d2 and the transposed convolutional layer t2 as the input to the encoder e3 and the convolutional layer l3.
[0111] Furthermore, in the encoding and decoding unit,
[0112] The convolutional layer is used to transform the dimension of the input image that has not been encoded and feature - fused into the same dimension as the image output by the feature fusion module;
[0113] The transposed convolutional layer is used to transform the dimension of the input image that has not been decoded into the same dimension as the image output by the decoder.
[0114] The dimension of the image that has not been encoded or decoded is different from the dimension of the image that has been encoded or decoded. Therefore, it is necessary to transform the dimension of the image that has not been encoded or decoded into the same number of channels, length, and width as the image after encoding and decoding through convolution or transposed convolution so that the image features can be added.
[0115] Specifically, in this embodiment, the input image dimension of the convolutional layer is (3, 160, 160), the output image dimension is (128, 40, 40), the convolutional kernel size is (5, 5), the padding number is 1, and the stride is 4; the input image dimension of the transposed convolutional layer is (128, 40, 40), the output image dimension is (3, 160, 160), the convolutional kernel size is (4, 4), the padding number is 0, and the stride is 4.
[0116] In this embodiment, by setting a residual connection between the encoder and the decoder in each encoding and decoding unit, and setting parallel convolutional layers in each encoder and decoder, through this innovative network structure, the relatively original image features extracted by the previous-stage encoder can be transmitted to the non-adjacent decoder in the subsequent stage, which can help the decoder better and more fully reconstruct the image features. By analyzing the experimental results, it can be seen that using this method helps the network make full use of all the features of the input picture, and the denoising training effect of the obtained result map is better.
[0117] Further, the encoder includes a parallel Transformer processing branch and a CFEB processing branch;
[0118] The Transformer processing branch is used to extract global features and deep semantic features; the CFEB processing branch is used to extract local image features and image detail features;
[0119] The two processing branches process the input image to the encoder in parallel and output synchronously.
[0120] Further, the Transformer processing branch includes a plurality of sequentially connected Transformer basic units, where
[0121] Each Transformer basic unit includes a first convolutional layer, a first dimension compression transposed layer, a first normalization layer, a multi-head self-attention, a second normalization layer, a multi-layer perceptron, a second dimension compression transposed layer, a second convolutional layer, and a third convolutional layer connected in sequence;
[0122] Residual operations are introduced at the end of the multi-head attention and the end of the multi-layer perceptron of each Transformer basic unit;
[0123] The image processing process of the Transformer processing branch of the encoder module is as follows:
[0124] Perform the following processing in each Transformer basic unit:
[0125] The first convolutional layer adjusts the input image dimension;
[0126] The first - dimension compression transpose layer adjusts the input image format;
[0127] The first normalization layer performs normal distribution standardization on the input image,
[0128] The multi - head attention layer marks different levels of content, regions, and long - range dependencies of the input image;
[0129] The sum of the input of the first normalization layer and the output of the multi - head attention layer is output to the second normalization layer;
[0130] The second normalization layer performs normal distribution standardization on the image again;
[0131] The multi - layer perceptron layer extracts deep - level features of the image;
[0132] The sum of the input of the second normalization layer and the output of the multi - layer perceptron layer is input to the second - dimension compression transpose layer to restore the image format;
[0133] The second convolutional layer enhances the features of the input image and adjusts the image dimension once;
[0134] The third convolutional layer enhances the features of the input image, adjusts the image dimension, and then outputs to the next Transformer basic unit,
[0135] The input image is output after being processed by each Transformer basic unit in sequence.
[0136] Specifically, Figure 4 shows the standard structure of the Transformer unit in an embodiment of the present invention.
[0137] Furthermore, the Transformer processing branch includes an optimized network structure composed of multiple sequentially connected Transformer basic units, where,
[0138] The first Transformer basic unit includes a first convolutional layer, a dimension compression transpose layer, a first normalization layer, multi - head self - attention, a second normalization layer, and a multi - layer perceptron connected in sequence;
[0139] The second to the last Transformer basic units all include a first normalization layer, multi - head self - attention, a second normalization layer, and a multi - layer perceptron connected in sequence;
[0140] The last Transformer basic unit further includes: a dimension compression transpose layer, a second convolutional layer, and a third convolutional layer;
[0141] A residual operation is introduced at the end of the multi-head attention and the end of the multi-layer perceptron of each Transformer basic unit respectively.
[0142] Specifically, in this embodiment, each Transformer processing branch of the encoder is composed of two connected Transformer units.
[0143] Preferably, as Figure 5 shown in the upper figure in, in a preferred embodiment of the present invention, the network structure of the Transformer processing branch composed of two Transformer units, in order to optimize the network structure, remove redundancy, and improve performance, the second-dimensional compression transpose layer, the second and third convolutional layers of the previous Transformer unit, and the first convolutional layer and the first-dimensional compression transpose layer of the subsequent Transformer unit are removed, obtaining Figure 5 the network structure shown in the lower figure in.
[0144] Specifically, in this embodiment, the first convolutional layer transforms the image dimension into (180, 160, 160).
[0145] The first-dimensional compression transpose layer transforms the image format into (25,600, 180) to meet the format requirements and facilitate the processing of subsequent modules.
[0146] The first normalization layer (LayerNorm1) first normalizes the image into a standard normal distribution matrix mapped to a mean of 0 and a variance of 1, making the data distribution more stable and beneficial to the stability and generalization ability of network training.
[0147] The image is input into the multi-head self-attention layer (MSA1: Multi-head Self Attention), and the neural network automatically learns and selectively focuses on the important information in the input to establish associations, adaptively capturing the long-range dependencies between elements by calculating the relative importance between elements.
[0148] Specifically, in this embodiment, the number of attention heads of the multi-head self-attention layer (MSA) is 6, and the window size is 8.
[0149] The result after the MSA layer is processed is connected to the input of the first normalization layer with residual, and then input into the second normalization layer (LayerNorm2), which performs the same operation as the first normalization layer.
[0150] The image is input into the multi-layer perceptron (MLP1) layer, introducing a more complex non-linear transformation model to extract features through multiple hidden layers.
[0151] The processed image is connected to the input residual of the second normalization layer, and the output image is input into the second Transformer unit.
[0152] Then, the image is processed by the first normalization layer, the multi-head attention layer, the second normalization layer, and the multi-layer perceptron layer of the second Transformer unit. The processing method and network structure are the same as those of the first Transformer unit.
[0153] Then, it is input into the second-dimensional compression transpose layer to transform the image format into (180, 160, 160).
[0154] Then, connect the second convolutional layer to transform the image dimension into (128, 80, 80).
[0155] Finally, innovatively connect the third convolutional layer for image enhancement, adjust the brightness, contrast, color, etc. of the image, so that the image is clearer, brighter, has more distinct colors and visual effects, and transform the image dimension into (128, 40, 40).
[0156] The convolutional kernel size of the convolutional layer included in the Transformer unit is (3, 3), the padding number is 1, and the stride is 1.
[0157] Furthermore, the CFEB processing branch includes multiple convolutional layers and downsampling layers. Among them, the convolutional layers are connected in a residual connection manner, specifically:
[0158] 4m convolutional layers and m - 1 downsampling layers;
[0159] Connect a downsampling layer after every 4 convolutional layers. The downsampling layer is used to increase the number of image channels and reduce the image size;
[0160] The first convolutional layer increases the number of channels of the noise image;
[0161] For the second to the penultimate convolutional layer, the sum of the input and output of each layer is calculated and then input into the next layer;
[0162] After the convolutional operation of the last convolutional layer, the processed image is output.
[0163] Specifically, in this embodiment, as Figure 6 shown, the CFEB processing branch convolutional neural network of each encoder is a network composed of 12 convolutional layers and 2 upsampling layers connected by residual connections. For example, the image with a dimension of (3, 160, 160) input by the input layer is processed in the network as follows:
[0164] The input image of convolutional layer 1 is named A, and the dimension of A is (3, 160, 160). The output image of convolutional layer 1 is named B, and the dimension of B is (32, 160, 160).
[0165] The input image of convolutional layer 2 is B, and the output image of convolutional layer 2 is named C, and the dimension of C is (32, 160, 160).
[0166] The input of convolutional layer 3 is named D, D = B + C, the dimension of D is (32, 160, 160), and the output image of convolutional layer 3 is named E, and the dimension of E is (32, 160, 160).
[0167] The input image of convolutional layer 4 is named F, F = D + E, the dimension of F is (32, 160, 160), and the output image of convolutional layer 4 is named G, and the dimension of G is (32, 160, 160).
[0168] Then, it passes through a downsampling layer 1, which is also composed of convolutional operations. The input image is named H, H = F + G, the dimension of H is (32, 160, 160), and the output image is named I, and the dimension of I is (64, 80, 80).
[0169] The input image of convolutional layer 5 is I, and the output image is named J, and the dimension of J is (64, 80, 80).
[0170] The input of convolutional layer 6 is named K, K = J + I, the dimension of K is (64, 80, 80), and the output image is named L, and the dimension of L is (64, 80, 80).
[0171] The input of convolutional layer 7 is named M, M = K + L, the dimension of M is (64, 80, 80), and the output image is named N, and the dimension of N is (64, 80, 80).
[0172] The input of convolutional layer 8 is named O, O = N + M, the dimension of O is (64, 80, 80), and the output image is named P, and the dimension of P is (64, 80, 80).
[0173] Then, it passes through a downsampling layer 2, which is also composed of convolutional operations. The input image is named G, G = P + O, the dimension of G is (64, 80, 80), and the output image is named R, and the dimension of R is (128, 40, 40).
[0174] The input image of convolutional layer 9 is R, the dimension of R is (128, 40, 40), and the output image is named S, and the dimension of S is (128, 40, 40).
[0175] The input image of convolutional layer 10 is named T, T = S + R, the dimension of T is (128, 40, 40), and the output image is named U, the dimension of U is (128, 40, 40);
[0176] The input image of convolutional layer 11 is named V, V = U + T, the dimension of V is (128, 40, 40), and the output image is named W, the dimension of W is (128, 40, 40);
[0177] The input image of convolutional layer 12 is named X, X = W + V, the dimension of X is (128, 40, 40), and the output image is named Y, the dimension of Y is (128, 40, 40).
[0178] Furthermore, the feature fusion module includes a CAT dimension concatenation module, a convolutional layer, and a CBAM convolutional attention module connected in sequence, where
[0179] The CAT dimension concatenation module receives the images output in parallel by the Transformer processing branch and the CFEB processing branch of the encoder, and performs channel dimension concatenation;
[0180] The convolutional layer performs a convolution operation on the image to complete feature extraction;
[0181] The CBAM convolutional attention module includes channel attention processing and spatial attention processing, and outputs the processed image.
[0182] As Figure 7 shown, in an embodiment of the present invention, the feature fusion module includes a CAT dimension concatenation module, a convolutional layer, and a CBAM convolutional attention module (Convolutional Block Attention Module) layer. Specifically,
[0183] The CAT dimension concatenation module connects the image with a dimension of (128, 40, 40) output by the Transformer processing branch of adjacent encoder modules, and the image with a dimension of (128, 40, 40) output in parallel by the CFEB module processing branch, and performs channel dimension concatenation into an image with a dimension of (256, 40, 40);
[0184] The convolutional layer performs a convolution operation on the concatenated image, changing the number of image channels, length, and width to (128, 40, 40), and completing feature extraction;
[0185] The CBAM layer includes channel attention processing and spatial attention processing, assigns weights to image channels and image pixels, and outputs the processed image.
[0186] As Figure 8As shown, the Channel Attention module contains two branches. Among them,
[0187] The first branch inputs an input image with dimensions (128, 40, 40) into an adaptive average pooling operation (avg_pool) to capture the average features of each channel in the entire feature map, reducing the length and width of the feature map to 1×1 while keeping the number of channels unchanged, resulting in a feature map of (128, 1, 1);
[0188] The feature map is input into the first fully connected layer (Fc_layer1), and the dimension of the feature map is converted to (8, 1, 1) through a 1×1 convolution without using a bias (bias = False);
[0189] The feature map output by the first fully connected layer undergoes a ReLU activation function operation (Relu1); The input of the ReLU activation function is a real number or a vector of real numbers, representing the weighted input of the neuron. The ReLU activation function changes the input less than zero to zero, while the input greater than or equal to zero remains unchanged. This makes it a very simple but effective non - linear activation function. The ReLU activation function operation is computationally simple, does not involve complex mathematical operations, and is fast. It can reduce the vanishing gradient, contribute to the training of deep neural networks; introduce non - linearity, enabling the neural network to learn complex mapping relationships;
[0190] The image processed by the ReLU function is input into the second fully connected layer (Fc_layer2), and an output image with dimensions (128, 1, 1) is obtained through a 1x1 convolution without using a bias (bias = False).
[0191] The second branch performs a row - adaptive max - pooling operation (max_pool) on an image with the same dimensions (128, 40, 40) as the input to the first branch, to capture the maximum features of each channel in the entire feature map, and reduce the image length and width to 1×1 while keeping the number of channels unchanged, outputting an image with dimensions (128, 1, 1);
[0192] The feature map is input into the third fully connected layer, and the image dimension is converted to (8, 1, 1) through a 1×1 convolution without using a bias (bias = False);
[0193] Perform a ReLU activation function operation (Relu2) on the output image of the third fully connected layer; the input of the ReLU activation function is a real number or a vector of real numbers, representing the weighted input of a neuron. The ReLU activation function changes the input less than zero to zero, while the input greater than or equal to zero remains unchanged. This makes it a very simple but effective non-linear activation function. The ReLU activation function operation is computationally simple, does not involve complex mathematical operations, and is fast. It can reduce the vanishing gradient, which helps to train deep neural networks; it introduces non-linearity, enabling the neural network to learn complex mapping relationships.
[0194] Input the image processed by the ReLU function into the fourth fully connected layer (Fc_layer4), and then perform a fully connected operation through a 1×1 convolution to obtain an image with a dimension of (128, 1, 1), without using a bias (bias = False).
[0195] Add the image features obtained from the first branch and the second branch element-wise, and then normalize the output image and its features through the Sigmoid activation function (sigmoid), restricting the output to between 0 and 1, and outputting an image with a dimension of (128, 1, 1). This final output can be regarded as the attention weight for each channel in the input feature map, used to adjust the contributions of different channels to improve the feature representation ability.
[0196] Multiply the image and features output by the channel attention module with the image input to the channel attention, and output an image with a dimension of (128, 40, 40) to the spatial attention module.
[0197] As Figure 9 shown, for the said spatial attention module (Spatial Attention), after performing average pooling (mean layer) and max pooling (max layer) operations on the input image, it outputs images with a dimension of (1, 40, 40) respectively;
[0198] Input the two images output in the previous step into the CAT layer for concatenation in the channel dimension, and the output dimension becomes (2, 40, 40);
[0199] Then, input the image into the convolutional layer (conv1), and output an image with a dimension of (1, 40, 40). The convolutional kernel size of this convolutional layer is (7, 7), the padding number is (3, 3), and the stride is 1;
[0200] Finally, input the image output by the convolutional layer into the sigmoid layer, and perform sigmoid activation through the sigmoid activation function, without changing the dimension, and the dimension remains (1, 40, 40). The Sigmoid function is often used as a threshold function for neural networks, and its formula is
[0201]
[0202] Where x is the pixel value in the feature map, and the sigmoid activation function maps the variable to between 0 and 1. This function is monotonically increasing and symmetric about (0, 0.5), with a slower change rate at both ends.
[0203] Therefore, the weight matrix image of the final output of the model is a feature map with 1 channel, and its spatial dimension is the same as that of the input, i.e., (1, 40, 40). This feature map represents the attention distribution of the input image under the spatial attention mechanism.
[0204] Furthermore, the input image and the output image of the spatial attention module are multiplied to obtain the output image of the CBAM layer.
[0205] As Figure 10 shown, the input image and the output image of the channel attention module are multiplied to obtain the output image marked with channel attention; the output image obtained after further inputting it into the spatial attention module is multiplied by the output image marked with channel attention, and the output image is the image marked with channel attention and spatial attention. The dimension of the image will not be changed during this process, only the size of the image pixel value will be changed.
[0206] The purpose of both channel attention and spatial attention is to assign different weights to different channels or pixels / regions according to the content of the input feature map. They can make the network pay more attention to important features, suppress noise or irrelevant information, thereby improving the performance and robustness of the model. The attention mechanism introduced in this way can effectively enhance the expression ability and discriminability of features.
[0207] Furthermore, the decoder is a network composed of multiple convolutional layers connected by residual connections, specifically including:
[0208] 4m convolutional layers and m - 1 upsampling layers;
[0209] One upsampling layer is connected after every 4 convolutional layers;
[0210] The upsampling layer is used to reduce the number of image channels and enlarge the length and width of the image;
[0211] The first convolutional layer reduces the number of channels of the noise image;
[0212] For the second to the penultimate convolutional layer, the sum of the input and the output of each layer is input into the next layer;
[0213] After the convolution operation of the last convolutional layer, the processed image is output.
[0214] Specifically,
[0215] As Figure 11 shown, in this embodiment, each decoder adopts a network composed of 12 convolutional layers and 2 upsampling layers connected by residual connections. For example, the processing process of an image with an input dimension of (128, 40, 40) in the input layer in the network is as follows:
[0216] The input image of convolutional layer 1 is named A_d, the dimension of A_d is (128, 40, 40), the output image of convolutional layer 1 is named B_d, and the dimension of B_d is (128, 40, 40);
[0217] The input image of convolutional layer 2 is named C_d, C_d = B_d, the dimension of C_d is (128, 40, 40), the output image of convolutional layer 2 is named D_d, and the dimension of D_d is (128, 40, 40);
[0218] The input image of convolutional layer 3 is named E_d, E_d = D_d + C_d, the dimension of E_d is (128, 40, 40), the output image of convolutional layer 3 is named F_d, and the dimension of F_d is (128, 40, 40);
[0219] The input image of convolutional layer 4 is named G_d, G_d = E_d + F_d, the dimension of G_d is (128, 40, 40), the output image of convolutional layer 4 is named H_d, and the dimension of H_d is (128, 40, 40);
[0220] Then, it passes through an upsampling layer 1, which is also composed of convolutional operations. Its input image is named I_d, I_d = H_d + G_d, the dimension of I_d is (128, 40, 40), and the output image is named J_d, and the dimension of J_d is (64, 80, 80);
[0221] The input image of convolutional layer 5 is J_d, and the output image of convolutional layer 5 is named K_d, and the dimension of K_d is (64, 80, 80);
[0222] The input image of convolutional layer 6 is named L_d, L_d = K_d + J_d, the dimension of L_d is (64, 80, 80), and the output image of convolutional layer 6 is named M_d, and the dimension of M_d is (64, 80, 80);
[0223] The input image of convolutional layer 7 is named N_d, N_d = M_d + L_d, the dimension of N_d is (64, 80, 80), and the output image of convolutional layer 7 is named O_d, and the dimension of O_d is (64, 80, 80);
[0224] The input image of convolutional layer 8 is named P_d, P_d = O_d + N_d, the dimension of P_d is (64, 80, 80), the output image of convolutional layer 8 is named R_d, and the dimension of R_d is (64, 80, 80);
[0225] Then, it passes through an upsampling layer 2, which is also composed of convolutional operations. Its input image is named S_d, S_d = R_d + P_d, the dimension of S_d is (64, 80, 80), and the output image is named T_d, and the dimension of T_d is (32, 160, 160);
[0226] The input image of convolutional layer 9 is named T_d, the dimension of T_d is (32, 160, 160), the output image of convolutional layer 9 is named U_d, and the dimension of U_d is (32, 160, 160);
[0227] The input image of convolutional layer 10 is named V_d, V_d = U_d + T_d, the dimension of V_d is (32, 160, 160), the output image of convolutional layer 10 is named W_d, and the dimension of W_d is (32, 160, 160);
[0228] The input image of convolutional layer 11 is named X_d, X_d = W_d + V_d, the dimension of X_d is (32, 160, 160), the output image of convolutional layer 11 is named Y_d, and the dimension of Y_d is (32, 160, 160);
[0229] The input image of convolutional layer 12 is named Z_d, Z_d = Y_d + X_d, the dimension of Z_d is (32, 160, 160), the output image of convolutional layer 12 is named de_out, and the dimension of de_out is (3, 160, 160).
[0230] Furthermore, for the convolutional layer of the output layer, after fine-tuning the pixels of the input feature map, a clear image with the same size and channels as the initial input noise image is output.
[0231] The output layer consists of a convolution, and the size of the convolution kernel is 3×3. After passing through the output layer, a denoised image is obtained, and the training is completed.
[0232] Step 3: Calculate the loss function for the denoised clear image and the corresponding clear image in the training set. When the threshold is reached, the training of the network model is completed.
[0233] Specifically, during the training process, the loss function is calculated for the generated denoised image and the corresponding clear image in the training set, so that the network can perform backpropagation. The loss function is where I RHQI is the denoised image output after the neural network processes the noisy image. HQ IGT is the ground truth label, that is, the clear image, where ∈ is a constant. In this embodiment, the empirical value is set to 10. -3 The selected optimizer is Adam, with parameters being the default parameters, and the initial learning rate is 2×10 -4 .
[0234] When the loss function converges, a trained network model is obtained.
[0235] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the embodiments of the present disclosure.
[0236] Those skilled in the art can understand that all or part of the processes of implementing the above embodiment methods can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disc, a read-only memory, or a random access memory, etc.
[0237] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.
Claims
1. A method for training a network model for removing image banding noise, characterized in that, The method includes the following steps: Preprocess the noisy image and the corresponding clear image to establish a training set; Establish a network model, input the noisy images in the training set into the network model, and after multiple deep feature extraction, feature fusion, and upsampling operations of the network model, output the denoised clear image, where The network model adopts a hybrid structure of CNN and Transformer, including an input layer, an encoding-decoding network, and an output layer; the encoding-decoding network includes multiple encoding-decoding units; in each encoding-decoding unit, the encoder and the decoder are both provided with corresponding convolutional layers for performing convolutional operations on the data input to the corresponding encoder / decoder and then connecting the output residuals of the corresponding encoder / decoder as the input of the subsequent decoder / encoder; Calculate the loss function based on the denoised clear image and the corresponding clear image in the training set, and complete the training of the network model when the value of the loss function reaches the threshold.
2. A method for training a network model for removing image banding noise according to claim 1, characterized in that, The preprocessing of the noisy image and the corresponding clear image to establish a training set includes: Select an equal number of images with strip noise and the corresponding clear images; Randomly crop the noisy images into noisy image patches of the same size, and crop the corresponding clear images at the same position to obtain clear image patches; Randomly perform horizontal flipping or vertical flipping operations on the noisy image patches and the clear image patches; Use the noisy image patches as training input images and the clear image patches as training reference labels to complete the construction of the training set.
3. The method for training a network model for removing image band noise according to claim 2, wherein The input layer is used for performing shallow feature extraction on the noisy image to obtain a shallow feature image; The encoding-decoding network is used for performing multiple rounds of deep feature extraction, feature fusion, and upsampling on the shallow feature image and then outputting; The output layer is used for fine-tuning the image output by the encoding-decoding network and then outputting the denoised clear image.
4. A method for training a network model for removing image banding noise according to claim 3, characterized in that, The feature fusion module includes a CAT dimension concatenation module, a convolutional layer, and a CBAM convolutional attention module connected in sequence, where The CAT dimension concatenation module receives the images output in parallel by the Transformer processing branch and the CFEB processing branch of the encoder, and performs channel dimension concatenation; The convolutional layer performs convolutional operations on the image to complete feature extraction; The CBAM convolutional attention module includes channel attention processing and spatial attention processing, and outputs the processed image.
5. A method for training a network model for removing image banding noise according to claim 4, characterized in that, The decoder is a network composed of multiple convolutional layers connected by residual connections, specifically including: 4m convolutional layers and m - 1 upsampling layers; One upsampling layer connected after every 4 convolutional layers; The upsampling layer is used for reducing the dimension of the image and enlarging the size of the image; The first convolutional layer reduces the dimension of the noisy image and reduces the number of channels; For the second to the second-to-last convolutional layers, the sum of the input and the output of each layer is input to the next layer; After the convolutional operation of the last convolutional layer, the processed image is output.
6. A method for training a network model for removing image banding noise according to claim 5, characterized in that, The encoding-decoding unit includes an encoder, a feature fusion module, a decoder, a convolutional layer parallel to the encoder and the feature fusion module, and a transposed convolutional layer parallel to the decoder; When there are n layers of encoding and decoding units in the network, let the encoder be e i , the feature fusion module be c i , the decoder be d i , the convolutional layer be l i , the transposed convolutional layer be t i , i = 1, 2,... n, where Input to the encoder e i The data synchronization input is input to the convolutional layer l i , and after passing through the encoder e i and the feature fusion module c i The output after serial processing is summed with the output of the convolutional layer l i to obtain the first summation data, and the first summation data is simultaneously used as the input to the decoder d i and the transposed convolutional layer t i ; the first summation data is also summed with the output of the feature fusion module c n-i+1 and the output of the convolutional layer l n-i+1 to obtain the second summation data, and the second summation data is simultaneously used as the input to the decoder d n-i+1 and the transposed convolutional layer l n-i+1 ; Decoder d i 's output is summed with the output of the transposed convolutional layer t i to obtain the third summation data, and the third summation data is simultaneously used as the input of the encoder e i+1 and the convolutional layer l i+1 ; the third summation data is also summed with the output of the decoder d n-i and the transposed convolutional layer t n-i to obtain the fourth summation data, and the fourth summation data is simultaneously used as the input of the encoder e n-i+1 and the convolutional layer l n-i+1 . Decoder d n 's output is summed with the output of the transposed convolutional layer t n to be the output data of the encoder-decoder network.
7. A method for training a network model for removing image banding noise according to claim 6, characterized in that, In the encoding and decoding unit, The convolutional layer is used to change the dimension of the input image that has not been encoded and feature-fused to be the same as the dimension of the image output by the feature fusion module; The transposed convolutional layer is used to change the dimension of the input image that has not been decoded to be the same as the dimension of the image output by the decoder.
8. A method for training a network model for removing image banding noise according to claim 7, characterized in that, The encoder includes a parallel Transformer processing branch and a CFEB processing branch; The Transformer processing branch is used to extract global features and deep semantic features; the CFEB processing branch is used to extract local image features and image detail features; The two processing branches process the image input to the encoder in parallel and output synchronously.
9. The method for training a network model for removing image banding noise according to claim 8, wherein The Transformer processing branch includes an optimized network structure composed of multiple sequentially connected Transformer basic units, wherein The first Transformer basic unit includes a first convolutional layer, a dimension compression transposed layer, a first normalization layer, multi-head self-attention, a second normalization layer, and a multi-layer perceptron connected in sequence; The second to the penultimate Transformer basic units each include a first normalization layer, multi-head self-attention, a second normalization layer, and a multi-layer perceptron connected in sequence; The last Transformer basic unit includes a first normalization layer, multi-head self-attention, a second normalization layer, a multi-layer perceptron, a dimension compression transposed layer, a second convolutional layer, and a third convolutional layer connected in sequence; Residual operations are introduced at the end of the multi-head attention and the end of the multi-layer perceptron of each Transformer basic unit.
10. The method for training a network model for removing image banding noise according to claim 9, wherein The CFEB processing branch includes multiple convolutional layers and downsampling layers, wherein the convolutional layers are connected in a residual connection manner, specifically: 4m convolutional layers and m - 1 downsampling layers; One downsampling layer is connected after every 4 convolutional layers. The downsampling layer is used to increase the number of channels and reduce the image size; The first convolutional layer increases the number of channels of the noise image, and the length and width of the image remain unchanged; For the second to the penultimate convolutional layer, the input and output of each layer are summed and then input to the next layer; After the convolution operation of the last convolutional layer, the processed image is output.