Image strip noise removing method based on CNN and Transform mixed structure

Through the image denoising method based on the mixed structure of CNN and Transformer, the problems of low image denoising efficiency and low clarity in the prior art are solved, and strip noise in the image is efficiently removed and image clarity is improved, and the recognition performance of the intelligent logistics sorting system is improved.

CN120219208APending Publication Date: 2025-06-27INST OF MICROELECTRONICS CHINESE ACAD OF SCI LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311789763.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is inefficient when removing strip noise in the image and the image clarity is not high after denoising, which affects the recognition speed and accuracy of the intelligent logistics sorting system.

Method used

Using an image denoising method based on a hybrid structure of CNN and Transformer, deep feature extraction, feature fusion and upsampling are performed through multi-layer codec units, and a clear image is finally output.

Benefits of technology

It improves the efficiency and clarity of image denoising, can effectively remove strip noise in the image, while retaining image edges and details, and improves the recognition speed and accuracy of the intelligent logistics sorting system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219208A_ABST
    Figure CN120219208A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of digital image processing, and particularly relates to an image banded noise removing method based on a CNN and Transform mixed structure, which comprises the following steps: preprocessing a noise image and a corresponding clear image, and establishing a training set; the method comprises the following steps: establishing a network model adopting a CNN and Transform mixed structure, wherein the network model comprises an input layer, a coding and decoding convolutional neural network in multi-layer residual connection and an output layer; and after the network model is trained by using the training set, inputting an image with banded noise into the network model, and finally outputting a clear image through cross superposition operation of multiple times of feature extraction, fusion and convolution of the network model. Experimental results show that when an image with strip-shaped noise is given, the random strip-shaped noise in the image can be efficiently removed, edges and details in the image can be well reserved, and a clear image without strip-shaped noise corresponding to the image can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of digital image processing, and particularly relates to a method for removing strip noise in an image by using deep learning to improve the image quality. Background Art

[0002] With the advent of the big data era, image information has the advantages of being intuitive, vivid, and easy to understand, and has become an essential information source in daily life. In reality, due to the influence of the surrounding environment of picture taking and imaging equipment, the collected images often contain certain noise pollution, which brings various obstacles to subsequent information analysis and seriously affects the efficiency of image information recognition. In the actual application scenario of the intelligent logistics sorting system for sorting express packages, the system first needs to take pictures and identify the express labels. The express label images with strip noise will cause the system to fail to recognize, resulting in the inability to sort the packages. Therefore, the express label images containing noise will seriously reduce the sorting rate of the express packages by the system. Therefore, it is necessary to invent a method to remove the stripe noise in the express labels, so as to improve the accuracy and efficiency of express label recognition in the intelligent logistics sorting process.

[0003] The reasons for the appearance of strip noise in the express label pictures obtained by shooting can be divided into two points. The first point is that the lens of the shooting photo is dirty, and the second point is that there are strip stains on the express label pictures themselves. During the transportation of express packages through the conveyor belt, an industrial camera shoots the express labels on the express packages from the bottom. Since the express packages pass above the lens of the industrial camera, as the number of transported packages increases, it is inevitable that the lens of the industrial camera accumulates dust, resulting in the express label pictures taken having serious strip noise. If the strip noise generated in this case is to be eliminated before imaging, it is necessary to promptly detect that the lens is dirty and then notify the relevant staff to wipe the lens stain. A large number of express package labels are pasted at the interfaces or gaps of the packages, which will cause the express labels at the interfaces or gaps to bulge and protrude, resulting in strip stains on the labels at the interfaces or gaps of the express packages. In this case, only image denoising processing can be performed through the subsequent software side.

[0004] Currently, there are mainly two types of image denoising technologies, namely traditional image denoising technologies and image denoising technologies based on neural networks. Traditional denoising methods, such as filtering, sparse methods, non-local mean algorithms, and non-adaptive methods, are used for image denoising and have achieved good results. However, these traditional model methods, on the one hand, require prior information of the image, and the artificially defined prior information of the image is difficult to capture all the features of the image, and manual parameter adjustment is also required, which easily leads to low model efficiency; on the other hand, probability modeling often involves complex optimization problems, resulting in complex calculations and long processing times, and a large amount of time and computational costs are required for each processed image.

[0005] The denoising algorithm based on deep learning regards the image denoising problem as a regression problem, fully utilizes the network architecture to automatically learn more essential features in the data from the training data, has a greater representation ability than traditional model-based methods, and can well protect the edges and details in the image. Therefore, using the deep learning denoising algorithm can achieve better denoising effects.

[0006] The strip noise in the express label image will seriously interfere with the subsequent device's recognition of barcodes and QR codes in the label, thus greatly reducing the recognition speed and recognition rate of the intelligent logistics sorting system. In view of this, it is indeed necessary to propose a method for removing strip noise from the label image to solve the above problems, thereby improving the recognition speed and recognition rate of the intelligent logistics sorting system. Summary of the Invention

[0007] In view of the above analysis, the embodiments of the present invention aim to provide a method for removing strip noise from images based on a hybrid structure of CNN and Transformer to solve the problems of low denoising efficiency for banded noise images and low clarity of the images after denoising in the prior art.

[0008] On the one hand, the embodiments of the present invention provide a method for removing strip noise from images based on a hybrid structure of CNN and Transformer, including the following steps:

[0009] Preprocess the noise image and the corresponding clear image to establish a training set;

[0010] Establish a network model, train the network model through the training set to obtain a trained network model; the network model adopts a hybrid structure of CNN and Transformer, including an input layer, an encoding and decoding network, and an output layer; the encoding and decoding network includes multiple encoding and decoding units; in each encoding and decoding unit, the encoder and decoder are both provided with corresponding convolutional layers for performing convolutional operations on the data input to the corresponding encoder / decoder and using the output of the corresponding encoder / decoder as the input of the subsequent decoder / encoder together.

[0011] Input the image with banded noise into the network model, and after multiple deep feature extraction, feature fusion, and upsampling operations of the network model, finally output a clear image.

[0012] Based on a further improvement of the above method, the input layer is used to perform shallow feature extraction on the noise image to obtain a shallow feature image;

[0013] The encoding and decoding network is used to perform multiple rounds of deep feature extraction, feature fusion, and upsampling on the shallow feature image and then output.

[0014] The output layer is used to output a clear image with noise removed after fine-tuning the image output by the encoding and decoding network.

[0015] Based on the further improvement of the above method, each layer of the encoding and decoding unit is connected in sequence, the structures of each layer of the encoding and decoding unit are the same, and the layers of the encoding and decoding unit are also connected in a residual manner.

[0016] Based on the further improvement of the above method, the encoding and decoding unit includes an encoder, a feature fusion module, and a decoder connected in sequence, a convolutional layer parallel to the encoder and the feature fusion module, and a transposed convolutional layer parallel to the decoder;

[0017] When there are n layers of encoding and decoding units in the network, let the encoder be e i , the feature fusion module be c i , the decoder be d i , the convolutional layer be l i , the transposed convolutional layer be t i , i = 1, 2,... n, where,

[0018] The data input to the encoder e i is synchronously input to the convolutional layer l i . After being serially processed by the encoder e i and the feature fusion module c i , the output is summed with the output of the convolutional layer l i to obtain the first summation data. The first summation data is simultaneously used as the input of the decoder d i and the transposed convolutional layer t i ; the first summation data is also summed with the output of the feature fusion module c n-i+1 and the output of the convolutional layer l n-i+1 to obtain the second summation data. The second summation data is simultaneously used as the input of the decoder d n-i+1 and the transposed convolutional layer l n-i+1 ;

[0019] The output of the decoder d i is summed with the output of the transposed convolutional layer t i to obtain the third summation data. The third summation data is simultaneously used as the input of the encoder e i+1 and the convolutional layer l i+1 ; the third summation data is also summed with the output of the decoder d n-i and the output of the transposed convolutional layer t n-i to obtain the fourth summation data. The fourth summation data is simultaneously used as the input of the encoder e n-i+1 and the convolutional layer l n-i+1 ;

[0020] The output of the decoder d n , and the transposed convolutional layer tn Sum the outputs as the output data of the encoding and decoding network.

[0021] Based on further improvement of the above method, in the encoding and decoding unit,

[0022] The convolutional layer is used to transform the dimension of the input image that has not been encoded and feature-fused into the same dimension as the image output by the feature fusion module;

[0023] The transposed convolutional layer is used to transform the dimension of the input image that has not been decoded into the same dimension as the image output by the decoder.

[0024] Based on further improvement of the above method, the encoder includes a parallel Transformer processing branch and a CFEB processing branch;

[0025] The Transformer processing branch is used to extract global features and deep semantic features; the CFEB processing branch is used to extract local image features and image detail features;

[0026] The two processing branches process the input image to the encoder in parallel and output synchronously.

[0027] Based on further improvement of the above method, the Transformer processing branch includes a plurality of sequentially connected Transformer basic units, where

[0028] Each Transformer basic unit includes a first convolutional layer, a first dimension compression transposed layer, a first normalization layer, multi-head self-attention, a second normalization layer, a multi-layer perceptron, a second dimension compression transposed layer, a second convolutional layer, and a third convolutional layer connected in sequence;

[0029] Residual operations are introduced at the end of the multi-head attention and the end of the multi-layer perceptron of each Transformer basic unit.

[0030] Based on further improvement of the above method, the CFEB processing branch includes a plurality of convolutional layers and downsampling layers, where residual connections are used between the convolutional layers.

[0031] Based on further improvement of the above method, the CFEB processing branch includes:

[0032] 4m convolutional layers and m - 1 downsampling layers;

[0033] One downsampling layer is connected after every 4 convolutional layers. The downsampling layer is used to increase the number of channels of the image and reduce the image size;

[0034] The first convolutional layer increases the number of channels of the noise image, and the length and width of the image remain unchanged;

[0035] The second to the second-to-last convolutional layers sum the input and output of each layer and then input them into the next layer;

[0036] After the convolutional operation of the last convolutional layer, the processed image is output.

[0037] Based on the further improvement of the above method, the image processing process of the Transformer processing branch of the encoder module is as follows:

[0038] The following processing is performed in each Transformer basic unit:

[0039] The first convolutional layer adjusts the dimension of the input image;

[0040] The first dimension compression transpose layer adjusts the format of the input image;

[0041] The first normalization layer performs normal distribution standardization processing on the input image,

[0042] The multi-head attention layer marks the content, regions, and long-range dependence relationships at different levels of the input image;

[0043] The sum of the input of the first normalization layer and the output of the multi-head attention layer is output to the second normalization layer;

[0044] The second normalization layer performs normal distribution standardization processing on the image again;

[0045] The multi-layer perceptron layer extracts the deep features of the image;

[0046] The sum of the input of the second normalization layer and the output of the multi-layer perceptron layer is input to the second dimension compression transpose layer to restore the image format;

[0047] The second convolutional layer enhances the features of the input image and adjusts the dimension of the image once;

[0048] The third convolutional layer enhances the features of the input image, adjusts the dimension of the image, and then outputs it to the next Transformer basic unit,

[0049] The input image is output after being processed by each Transformer basic unit in turn.

[0050] The object of the present invention is to provide an image denoising method and system, which can improve the visual performance of the reconstructed image, as well as the peak signal-to-noise ratio and structural similarity. The present invention provides a new strip noise removal algorithm for labeled images based on a hybrid structure of CNN and Transformer, which can remove strip noise in the image while retaining the image edges and fine texture structures, so as to achieve the purpose of obtaining a clear image. The network designed by the present invention can train a more universal denoising image model, improve the simplicity of image denoising, and has practical application value.

[0051] In the present invention, the above technical solutions can also be combined with each other to realize more preferred combination schemes. Other features and advantages of the present invention will be described in the following specification, and some advantages can be made obvious from the specification, or understood by implementing the present invention. The object and other advantages of the present invention can be realized and obtained through the content specifically pointed out in the specification and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The drawings are only for the purpose of showing specific embodiments, and are not considered as a limitation to the present invention. In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the description of the embodiments will be briefly introduced below. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0053] Figure 1 It is a flowchart of an image denoising method provided by an embodiment of the present invention;

[0054] Figure 2 It is a flowchart for establishing a noise image training set according to an embodiment of the present invention;

[0055] Figure 3 It is an overall structure diagram of a residual connection image denoising neural network model provided by an embodiment of the present invention;

[0056] Figure 4 It is a hierarchical structure diagram of a Transformer unit provided by an embodiment of the present invention;

[0057] Figure 5 It is a schematic diagram of network optimization including 2 Transformer units provided by an embodiment of the present invention;

[0058] Figure 6 It is a neural network structure diagram of an encoder CFEB module provided by an embodiment of the present invention;

[0059] Figure 7The structural diagram of the feature fusion module with two parallel branches of the encoder provided in the embodiment of the present invention as inputs;

[0060] Figure 8 The structural diagram of the channel attention module in the embodiment of the present invention;

[0061] Figure 9 The structural diagram of the spatial attention module in the embodiment of the present invention;

[0062] Figure 10 The structural diagram of the CBAM branch in the embodiment of the present invention;

[0063] Figure 11 The neural network structural diagram of the decoder provided in the embodiment of the present invention;

[0064] Figure 12(a) is an image of one group before strip noise removal in the embodiment of the present invention;

[0065] Figure 12(b) is an image of one group after strip noise removal in the embodiment of the present invention;

[0066] Figure 12(c) is an image of one group before strip noise removal in the embodiment of the present invention;

[0067] Figure 12(d) is an image of one group after strip noise removal in the embodiment of the present invention; Detailed implementation manners

[0068] The following combines the drawings to specifically describe the preferred embodiments of the present invention. Among them, the drawings form a part of this application and are used together with the embodiments of the present invention to explain the principle of the present invention, rather than to limit the scope of the present invention.

[0069] The dataset used in an embodiment of the present invention is an express label image dataset with strip noise, and the dataset contains 3,500 pairs of noisy images and clear image pairs.

[0070] A method for removing strip noise from label images based on a hybrid structure of CNN and Transformer, as Figure 1 shown in the main flowchart of the present invention, includes the following steps:

[0071] Step 1: Preprocess the noisy image and the corresponding clear image to establish a training set.

[0072] In this embodiment, 3,000 express label images with strip noise and 3,000 corresponding clear images are selected;

[0073] The preprocessing process, as Figure 2 shown, includes:

[0074] S11. Randomly crop the noisy image into image patches of the same size, and crop the corresponding clear image at the same position to obtain clear image patches.

[0075] The image is represented by three dimensions: channel, length, and width. In this embodiment, the image is cropped to a size of (3, 160, 160), that is, the cropped image patch has 3 channels, a length of 160 pixels, and a width of 160 pixels.

[0076] S12. Randomly perform horizontal flipping or vertical flipping operations on the noisy image patches and clear image patches for data augmentation.

[0077] S13. Use the noisy image patches as input images and the clear image patches as labels to construct a training set.

[0078] Step 2: Establish a network model, and train the network model through the training set to obtain a trained network model.

[0079] In this embodiment, during the model training stage, every 24 images are used as a processing batch for batch training.

[0080] Figure 3 The overall structure of the network model using a hybrid structure of CNN and Transformer is shown. The network model adopts a hybrid structure of CNN and Transformer, including an input layer, an encoder-decoder network, and an output layer; the encoder-decoder network includes multiple encoder-decoder units; in each layer of encoder-decoder units, the encoder and decoder are both provided with corresponding convolutional layers, which are used to perform convolutional operations on the data input to the corresponding encoder / decoder and then, together with the output of the corresponding encoder / decoder, serve as the input for the subsequent decoder / encoder.

[0081] The following details the network model structure.

[0082] Specifically, the input layer is used to perform shallow feature extraction on the noisy image to obtain a shallow feature image.

[0083] In this embodiment, the input layer consists of a first convolutional operation layer with a convolution kernel of 3×3, which is used to perform shallow feature extraction on the noisy image and output an image with the number of output channels, length, and width of (3, 160, 160) to the next layer.

[0084] The encoder-decoder network is used to perform multiple rounds of deep feature extraction, feature fusion, and upsampling on the shallow feature image and then output.

[0085] Furthermore, each layer of encoder-decoder units is connected in sequence, the structures of each layer of encoder-decoder units are the same, and each layer of encoder-decoder units is also connected in a residual manner.

[0086] Further, the encoding and decoding unit includes an encoder, a feature fusion module, and a decoder connected in sequence, a convolutional layer parallel to the encoder and the feature fusion module, and a transposed convolutional layer parallel to the decoder;

[0087] When there are n encoding and decoding units in the network, let the encoder be e i , the feature fusion module be c i , the decoder be d i , the convolutional layer be l i , the transposed convolutional layer be t i , i = 1, 2,... n, where

[0088] The data input to the encoder e i is synchronously input to the convolutional layer l i , and after being serially processed by the encoder e i and the feature fusion module c i , the output is summed with the output of the convolutional layer l i to obtain the first summation data, and the first summation data is simultaneously used as the input of the decoder d i and the transposed convolutional layer t i ; the first summation data is also summed with the output of the feature fusion module c n-i+1 and the output of the convolutional layer l n-i+1 to obtain the second summation data, and the second summation data is simultaneously used as the input of the decoder d n-i+1 , the transposed convolutional layer l n-i+1 ;

[0089] The output of the decoder d i is summed with the output of the transposed convolutional layer t i to obtain the third summation data, and the third summation data is simultaneously used as the input of the encoder e i+1 , the convolutional layer l i+1 ; the third summation data is also summed with the output of the decoder d n-i , the transposed convolutional layer t n-i to obtain the fourth summation data, and the fourth summation data is simultaneously used as the input of the encoder e n-i+1 , the convolutional layer l n-i+1 ;

[0090] The output of the decoder d n is summed with the output of the transposed convolutional layer t n as the output data of the encoding and decoding network.

[0091] The network contains multiple encoding and decoding units composed of an encoder, a feature fusion module, a decoder, a convolutional layer, and a transposed convolutional layer. It is necessary to comprehensively consider network parameters, inference speed, and image denoising effect, and analyze the experimental results to determine the number of encoding and decoding units in the network. If too many encoding and decoding units are set, the network parameters will be too large and the inference speed will decrease. If too few are set, the image features will not be extracted sufficiently, resulting in too much information loss, which is not conducive to the subsequent decoder for image detail reconstruction, and ultimately leads to a blurred result image output by the network.

[0092] An important indicator to measure the denoising effect of a neural network is PSNR. PSNR stands for "Peak Signal-to-Noise Ratio", and its Chinese meaning is peak signal-to-noise ratio. PSNR is defined based on MSE (mean squared error). For a given original image I of size m×n and its noisy image K after adding noise, its MSE can be defined as:

[0093]

[0094] where m and n respectively represent the length and width of the picture, i and j represent the abscissa and ordinate of a pixel point, and I(i,j) represents the pixel value at the position (i, j) in the image I. Then PSNR can be defined as:

[0095]

[0096] where Max I is the maximum pixel value of the image, and the unit of PSNR is dB. If each pixel is represented by 8-bit binary, its value is 2 8 -1 = 255. If the image is a grayscale image, the PSNR value can be directly calculated using the above formula; if the image is a color image, calculate the MSE value of each of the three channels of the RGB image and then find the average value MSE1, and then substitute MSE1 into the PSNR formula to calculate the PSNR of the color picture. In the experiment, by analyzing the PSNR values of the output result images of neural networks with different repetition times of the encoder, feature fusion module, and decoder, it is found that setting the repetition time to 3 is more appropriate. If the repetition time is increased further, the effect on improving the PSNR value is very limited.

[0097] Preferably, the number of layers of the encoding and decoding unit is set to 3 layers.

[0098] Specifically, taking the number of layers of the encoding and decoding unit set to 3 layers as an example, the processing process of data within the encoding and decoding unit will be described in detail below.

[0099] From Figure 3It can be seen that in this embodiment, the image (3, 160, 160) after being processed by the input layer is input into the encoder e1 and the convolutional layer l1. After being serially processed by the encoder e1 and the feature fusion module c1, an image with the number of channels, length, and width of (128, 40, 40) is output, which is summed with the output of the convolutional layer l1. The number of channels, length, and width of the image after being processed by the convolutional layer l1 are also (128, 40, 40). Moreover, the image after the summation processing of the two is not only output to the decoder d1 and the transposed convolutional layer t1 for processing, but also summed with the result output after being processed by the feature fusion module c3 and the output result after convolution with the convolutional layer l3, and then input into the decoder d3 and the transposed convolutional layer t3.

[0100] The output image after being processed by the decoder d1 is an image with the number of channels, length, and width of (3, 160, 160), which is summed with the output image after the convolutional operation of the transposed convolutional layer t1. The number of channels, length, and width of the image after being processed by the transposed convolutional layer t1 are also (3, 160, 160). It is not only output to the encoder e2 and processed with the convolutional layer l2, but also summed with the output results of the decoder d2 and the transposed convolutional layer t2 as the input of the encoder e3 and the convolutional layer l3.

[0101] Further, in the encoding and decoding unit,

[0102] The convolutional layer is used to transform the dimension of the input image that has not been encoded and feature - fused into the same dimension as the image output by the feature fusion module.

[0103] The transposed convolutional layer is used to transform the dimension of the input image that has not been decoded into the same dimension as the image output by the decoder.

[0104] The dimensions of the image that has not been encoded or decoded are different from those of the image that has been encoded or decoded. Therefore, it is necessary to use convolution or transposed convolution to transform the dimensions of the image that has not been encoded or decoded into the same number of channels, length, and width as the image after encoding and decoding, so that the image features can be added.

[0105] Specifically, in this embodiment, the input image dimension of the convolutional layer is (3, 160, 160), the output image dimension is (128, 40, 40), the convolutional kernel size is (5, 5), the padding number is 1, and the stride is 4; the input image dimension of the transposed convolutional layer is (128, 40, 40), the output image dimension is (3, 160, 160), the convolutional kernel size is (4, 4), the padding number is 0, and the stride is 4.

[0106] In the present invention, a residual connection is set between the encoder and the decoder in each encoding and decoding unit, and parallel convolutional layers are set in each encoder and decoder. Through this innovative network structure, relatively original image features extracted by the previous-stage encoder can be transmitted to the non-adjacent decoder in the subsequent stage, which can help the decoder better and more fully reconstruct the image features. From the analysis of the experimental results, it can be seen that using this method helps the network make full use of all the features of the input picture, and the resulting image has a better denoising effect.

[0107] Further, the encoder includes a parallel Transformer processing branch and a CFEB processing branch;

[0108] The Transformer processing branch is used to extract global features and deep semantic features; the CFEB processing branch is used to extract local image features and image detail features;

[0109] The two processing branches process the input image to the encoder in parallel and output synchronously.

[0110] Further, the Transformer processing branch includes a plurality of sequentially connected basic Transformer units, where,

[0111] Each basic Transformer unit includes a first convolutional layer, a first dimension compression transpose layer, a first normalization layer, multi-head self-attention, a second normalization layer, a multi-layer perceptron, a second dimension compression transpose layer, a second convolutional layer, and a third convolutional layer connected in sequence;

[0112] Residual operations are introduced at the end of the multi-head attention and the end of the multi-layer perceptron of each basic Transformer unit.

[0113] Specifically, Figure 4 shows the standard structure of the Transformer unit in an embodiment of the present invention.

[0114] Specifically,

[0115] In one embodiment, the Transformer processing branch of each encoder is composed of 2 connected Transformer units.

[0116] Further, the image processing process of the Transformer processing branch of the encoder module is as follows:

[0117] The following processing is performed in each basic Transformer unit:

[0118] The first convolutional layer adjusts the dimension of the input image;

[0119] The first - dimensional compression transpose layer adjusts the input image format;

[0120] The first normalization layer performs normal distribution standardization on the input image,

[0121] The multi - head attention layer marks different - level contents, regions, and long - range dependency relationships of the input image;

[0122] The sum of the input of the first normalization layer and the output of the multi - head attention layer is output to the second normalization layer;

[0123] The second normalization layer performs normal distribution standardization on the image again;

[0124] The multi - layer perceptron layer extracts deep - level features of the image;

[0125] The sum of the input of the second normalization layer and the output of the multi - layer perceptron layer is input to the second - dimensional compression transpose layer to restore the image format;

[0126] The second convolutional layer enhances the features of the input image and adjusts the image dimension once;

[0127] The third convolutional layer enhances the features of the input image, adjusts the image dimension, and then outputs to the next Transformer basic unit,

[0128] The input image is output after being processed by each Transformer basic unit in sequence.

[0129] Preferably, as Figure 5 shown in the upper - middle figure, in a preferred embodiment of the present invention, the Transformer processing branch is a network structure composed of 2 Transformer units. To optimize the network structure, remove redundancy, and improve performance, the second - dimensional compression transpose layer, the second and third convolutional layers of the previous Transformer unit, and the first convolutional layer and the first - dimensional compression transpose layer of the subsequent Transformer unit are removed, obtaining Figure 5 the network structure shown in the lower - middle figure.

[0130] Specifically,

[0131] In this embodiment, the first convolutional layer transforms the image dimension to (180, 160, 160).

[0132] The first - dimensional compression transpose layer transforms the image format to (25600, 180) to meet the format requirements and facilitate subsequent module processing.

[0133] The first normalization layer (LayerNorm1) first normalizes the image into a standard normal distribution matrix mapped to a mean of 0 and a variance of 1, making the data distribution more stable and facilitating the stability and generalization ability of network training.

[0134] The image is input into the multi-head self-attention layer (MSA1: Multi-head Self Attention), and the neural network automatically learns and selectively focuses on the important information in the input to establish associations, adaptively capturing the long-range dependencies between elements by calculating the relative importance between elements.

[0135] Specifically, in this embodiment, the number of attention heads of the multi-head self-attention layer (MSA) is 6, and the window size is 8.

[0136] The result after the MSA layer is processed is connected to the input residual of the first normalization layer and input into the second normalization layer (LayerNorm2), which performs the same operation as the first normalization layer.

[0137] The image is input into the multi-layer perceptron (MLP1) layer, introducing a more complex non-linear transformation model for feature extraction through multiple hidden layers.

[0138] The processed image is connected to the input residual of the second normalization layer, and the output image is input into the second Transformer unit.

[0139] Then, the image is processed by the first normalization layer, multi-head attention layer, second normalization layer, and multi-layer perceptron layer of the second Transformer unit, and the processing method and network structure are the same as those of the first Transformer unit.

[0140] Then, it is input into the second dimensionality compression transpose layer to transform the image format into (180, 160, 160).

[0141] Then, it is connected to the second convolutional layer to transform the image dimension into (128, 80, 80).

[0142] Finally, it is innovatively connected to the third convolutional layer for image enhancement, adjusting the brightness, contrast, color, etc. of the image, so that the image is clearer, brighter, has more distinct colors and visual effects, and the image dimension is transformed into (128, 40, 40).

[0143] The convolutional kernel size of the convolutional layer included in the Transformer unit is (3, 3), the padding number is 1, and the stride is 1.

[0144] Furthermore, the CFEB processing branch includes multiple convolutional layers and downsampling layers, and among them, the convolutional layers are connected in a residual connection manner.

[0145] Further, the CFEB processing branch includes:

[0146] 4m convolutional layers and m - 1 downsampling layers;

[0147] One downsampling layer is connected after every 4 convolutional layers. The downsampling layer is used to increase the number of image channels and reduce the image size;

[0148] The first convolutional layer increases the number of channels of the noise image;

[0149] For the second to the second - last convolutional layers, the sum of the input and output of each layer is input into the next layer;

[0150] After the convolution operation of the last convolutional layer, the processed image is output.

[0151] Specifically,

[0152] In this embodiment, as Figure 6 shown, the CFEB processing branch convolutional neural network of each encoder is a network composed of 12 convolutional layers and 2 upsampling layers connected by residual connections. For example, the processing process of an image with an input dimension of (3, 160, 160) in the network is as follows:

[0153] The input image of convolutional layer 1 is named A, with a dimension of (3, 160, 160). The output image of convolutional layer 1 is named B, with a dimension of (32, 160, 160);

[0154] The input image of convolutional layer 2 is B, and the output image of convolutional layer 2 is named C, with a dimension of (32, 160, 160);

[0155] The input of convolutional layer 3 is named D, D = B + C, with a dimension of (32, 160, 160). The output image of convolutional layer 3 is named E, with a dimension of (32, 160, 160);

[0156] The input image of convolutional layer 4 is named F, F = D + E, with a dimension of (32, 160, 160). The output image of convolutional layer 4 is named G, with a dimension of (32, 160, 160);

[0157] Then, it passes through a downsampling layer 1, which is also composed of a convolution operation. Its input image is named H, H = F + G, with a dimension of (32, 160, 160). The output image is named I, with a dimension of (64, 80, 80);

[0158] The input image of convolutional layer 5 is I, and the output image is named J, with a dimension of (64, 80, 80);

[0159] The input image of convolutional layer 6 is named K, K = J + I, the dimension of K is (64, 80, 80), and the output image is named L, the dimension of L is (64, 80, 80);

[0160] The input image of convolutional layer 7 is named M, M = K + L, the dimension of M is (64, 80, 80), and the output image is named N, the dimension of N is (64, 80, 80);

[0161] The input image of convolutional layer 8 is named O, O = N + M, the dimension of O is (64, 80, 80), and the output image is named P, the dimension of P is (64, 80, 80);

[0162] Then, it passes through a downsampling layer 2, which is also composed of convolutional operations. Its input image is named G, G = P + O, the dimension of G is (64, 80, 80), and the output image is named R, the dimension of R is (128, 40, 40);

[0163] The input image of convolutional layer 9 is R, the dimension of R is (128, 40, 40), and the output image is named S, the dimension of S is (128, 40, 40);

[0164] The input image of convolutional layer 10 is named T, T = S + R, the dimension of T is (128, 40, 40), and the output image is named U, the dimension of U is (128, 40, 40);

[0165] The input image of convolutional layer 11 is named V, V = U + T, the dimension of V is (128, 40, 40), and the output image is named W, the dimension of W is (128, 40, 40);

[0166] The input image of convolutional layer 12 is named X, X = W + V, the dimension of X is (128, 40, 40), and the output image is named Y, the dimension of Y is (128, 40, 40).

[0167] Furthermore, the feature fusion module includes a CAT dimension concatenation module, a convolutional layer, and a CBAM convolutional attention module connected in sequence, where

[0168] The CAT dimension concatenation module receives the images output in parallel by the Transformer processing branch and the CFEB processing branch of the encoder, and performs channel dimension concatenation;

[0169] The convolutional layer performs convolutional operations on the image to complete feature extraction;

[0170] The CBAM convolutional attention module includes channel attention processing and spatial attention processing, and outputs the processed image.

[0171] Such asFigure 7 As shown in the figure, in one embodiment of the present invention, the feature fusion module includes a CAT dimension splicing module, a convolutional layer, and a CBAM (Convolutional Block Attention Module) layer. Specifically,

[0172] The CAT dimension splicing module connects the image with a dimension of (128, 40, 40) output from the Transformer processing branch of adjacent encoder modules, and the image with a dimension of (128, 40, 40) output in parallel from the CFEB module processing branch, and performs channel dimension splicing to form an image with a dimension of (256, 40, 40);

[0173] The convolutional layer performs a convolution operation on the spliced image, changing the number of image channels, length, and width to (128, 40, 40) to complete feature extraction;

[0174] The CBAM layer includes channel attention processing and spatial attention processing, assigns weights to image channels and image pixels, and outputs the processed image.

[0175] As Figure 8 shown, the channel attention module (Channel Attention) includes two branches, where

[0176] The first branch inputs the input image with a dimension of (128, 40, 40) into an adaptive average pooling operation (avg_pool) to capture the average features of each channel in the entire feature map, reduces the length and width of the feature map to 1×1, and keeps the number of channels unchanged, obtaining a feature map with a dimension of (128, 1, 1);

[0177] The feature map is input into the first fully connected layer (Fc_layer1), and the dimension of the feature map is converted to (8, 1, 1) through a 1×1 convolution, and at the same time, no bias is used (bias = False);

[0178] The feature map output by the first fully connected layer is subjected to a ReLU activation function operation (Relu1); the input of the ReLU activation function is a real number or a real number vector representing the weighted input of the neuron. The ReLU activation function changes the input less than zero to zero, while the input greater than or equal to zero remains unchanged. This makes it a very simple but effective non-linear activation function. The ReLU activation function operation is computationally simple, does not involve complex mathematical operations, and is fast. It can reduce gradient disappearance, help train deep neural networks; introduce non-linearity, enabling the neural network to learn complex mapping relationships;

[0179] The image processed by the ReLU function is input into the second fully connected layer (Fc_layer2), and the output image with a dimension of (128, 1, 1) is obtained through 1x1 convolution, while no bias is used (bias = False).

[0180] For the second branch, the image with a dimension of (128, 40, 40) that is the same as the input of the first branch undergoes a row adaptive max pooling operation (max_pool) to capture the maximum feature of each channel in the entire feature map, and reduces the image length and width to 1×1 while keeping the number of channels unchanged, outputting an image with a dimension of (128, 1, 1);

[0181] The feature map is input into the third fully connected layer, and the image dimension is converted to (8, 1, 1) through 1×1 convolution, while no bias is used (bias = False);

[0182] The output image of the third fully connected layer undergoes a ReLU activation function operation (Relu2); the input of the ReLU activation function is a real number or a vector of real numbers, representing the weighted input of the neuron. The ReLU activation function changes the input less than zero to zero, while the input greater than or equal to zero remains unchanged. This makes it a very simple but effective non - linear activation function. The ReLU activation function operation is computationally simple, does not involve complex mathematical operations, and is fast. It can reduce the vanishing gradient, which helps in training deep neural networks; it introduces non - linearity, enabling the neural network to learn complex mapping relationships.

[0183] The image processed by the ReLU function is input into the fourth fully connected layer (Fc_layer4), and then a fully connected operation is performed through 1×1 convolution to obtain an image with a dimension of (128, 1, 1), while no bias is used (bias = False).

[0184] The image features obtained from the first branch and the second branch are added element - by - element, and then the output image and its features are normalized through the Sigmoid activation function (sigmoid), restricting the output to between 0 and 1, and outputting an image with a dimension of (128, 1, 1). This final output can be regarded as the attention weight for each channel in the input feature map, used to adjust the contributions of different channels to improve the feature representation ability.

[0185] The image and features output by the channel attention module are multiplied by the image input to the channel attention, and an image with a dimension of (128, 40, 40) is output to the spatial attention module.

[0186] As Figure 9As shown, for the Spatial Attention module, after the input image undergoes average pooling (mean layer) and max pooling (max layer) operations, images with dimensions of (1, 40, 40) are output respectively;

[0187] The two images output in the previous step are input into the CAT layer and concatenated in the channel dimension, and the output image dimension becomes (2, 40, 40);

[0188] Then, the image is input into the convolutional layer (conv1), and an image with dimensions of (1, 40, 40) is output. The convolutional kernel size of this convolutional layer is (7, 7), the padding number is (3, 3), and the stride is 1;

[0189] Finally, the image output by the convolutional layer is input into the sigmoid layer, and sigmoid activation is performed through the sigmoid activation function without changing the dimension, and the dimension remains (1, 40, 40). The sigmoid function is often used as the threshold function of a neural network, and its formula is

[0190]

[0191] where x is the pixel value in the feature map. The sigmoid activation function maps the variable to between 0 and 1. This function is monotonically increasing and symmetric about (0, 0.5), and changes slowly at both ends.

[0192] Therefore, the weight matrix image of the final output of the model is a feature map with 1 channel, and its spatial dimension is the same as the input, that is, (1, 40, 40). This feature map represents the attention distribution of the input image under the spatial attention mechanism.

[0193] Furthermore, the input image and the output image of the spatial attention module are multiplied to obtain the output image of the CBAM layer.

[0194] As Figure 10 shown, the input image and the output image of the channel attention module are multiplied to obtain the output image marked with channel attention; the output image obtained by further inputting it into the spatial attention module and processing is multiplied by the output image marked with channel attention, and the output image is the image marked with both channel attention and spatial attention. In this process, the dimension of the image will not be changed, only the size of the image pixel value will be changed.

[0195] The purpose of both channel attention and spatial attention is to assign different weights to different channels or pixels / regions according to the content of the input feature map. They can make the network pay more attention to important features, suppress noise or irrelevant information, thereby improving the performance and robustness of the model. The attention mechanism introduced in this way can effectively enhance the expressive power and discriminability of features.

[0196] Furthermore, the decoder is a network composed of multiple convolutional layers connected by residual connections, specifically including:

[0197] 4m convolutional layers and m - 1 upsampling layers;

[0198] One upsampling layer is connected after every 4 convolutional layers;

[0199] The upsampling layer is used to reduce the number of image channels and enlarge the length and width of the image;

[0200] The first convolutional layer reduces the number of channels of the noise image;

[0201] For the second to the second - last convolutional layers, the sum of the input and output of each layer is calculated and then input into the next layer;

[0202] After the convolution operation of the last convolutional layer, the processed image is output.

[0203] Specifically,

[0204] As Figure 11 shown, in this embodiment, each decoder is a network composed of 12 convolutional layers and 2 upsampling layers connected by residual connections. For example, an image with an input dimension of (128, 40, 40) is processed in the network as follows:

[0205] The input image of convolutional layer 1 is named A_d, and the dimension of A_d is (128, 40, 40). The output image of convolutional layer 1 is named B_d, and the dimension of B_d is (128, 40, 40);

[0206] The input image of convolutional layer 2 is named C_d, C_d = B_d, and the dimension of C_d is (128, 40, 40). The output image of convolutional layer 2 is named D_d, and the dimension of D_d is (128, 40, 40);

[0207] The input image of convolutional layer 3 is named E_d, E_d = D_d + C_d, and the dimension of E_d is (128, 40, 40). The output image of convolutional layer 3 is named F_d, and the dimension of F_d is (128, 40, 40);

[0208] The input image of convolutional layer 4 is named \(G_d\), \(G_d = E_d+F_d\), the dimension of \(G_d\) is \((128, 40, 40)\), the output image of convolutional layer 4 is named \(H_d\), and the dimension of \(H_d\) is \((128, 40, 40)\);

[0209] Then, it passes through an upsampling layer 1 which is also composed of convolutional operations. Its input image is named \(I_d\), \(I_d = H_d+G_d\), the dimension of \(I_d\) is \((128, 40, 40)\), and the output image is named \(J_d\), and the dimension of \(J_d\) is \((64, 80, 80)\);

[0210] The input image of convolutional layer 5 is \(J_d\), and the output image of convolutional layer 5 is named \(K_d\), and the dimension of \(K_d\) is \((64, 80, 80)\);

[0211] The input image of convolutional layer 6 is named \(L_d\), \(L_d = K_d+J_d\), the dimension of \(L_d\) is \((64, 80, 80)\), and the output image of convolutional layer 6 is named \(M_d\), and the dimension of \(M_d\) is \((64, 80, 80)\);

[0212] The input image of convolutional layer 7 is named \(N_d\), \(N_d = M_d+L_d\), the dimension of \(N_d\) is \((64, 80, 80)\), and the output image of convolutional layer 7 is named \(O_d\), and the dimension of \(O_d\) is \((64, 80, 80)\);

[0213] The input image of convolutional layer 8 is named \(P_d\), \(P_d = O_d+N_d\), the dimension of \(P_d\) is \((64, 80, 80)\), and the output image of convolutional layer 8 is named \(R_d\), and the dimension of \(R_d\) is \((64, 80, 80)\);

[0214] Then, it passes through an upsampling layer 2 which is also composed of convolutional operations. Its input image is named \(S_d\), \(S_d = R_d+P_d\), the dimension of \(S_d\) is \((64, 80, 80)\), and the output image is named \(T_d\), and the dimension of \(T_d\) is \((32, 160, 160)\);

[0215] The input image of convolutional layer 9 is named \(T_d\), the dimension of \(T_d\) is \((32, 160, 160)\), and the output image of convolutional layer 9 is named \(U_d\), and the dimension of \(U_d\) is \((32, 160, 160)\);

[0216] The input image of convolutional layer 10 is named \(V_d\), \(V_d = U_d+T_d\), the dimension of \(V_d\) is \((32, 160, 160)\), and the output image of convolutional layer 10 is named \(W_d\), and the dimension of \(W_d\) is \((32, 160, 160)\);

[0217] The input image of convolutional layer 11 is named X_d, X_d = W_d + V_d, the dimension of X_d is (32, 160, 160), the output image of convolutional layer 11 is named Y_d, and the dimension of Y_d is (32, 160, 160);

[0218] The input image of convolutional layer 12 is named Z_d, Z_d = Y_d + X_d, the dimension of Z_d is (32, 160, 160), the output image of convolutional layer 12 is named de_out, and the dimension of de_out is (3, 160, 160).

[0219] Furthermore, the convolutional layer of the output layer fine-tunes the pixels of the input feature map and then outputs a clear image with the same size and channels as the initial input noisy image.

[0220] The output layer consists of a convolution with a kernel size of 3×3. After passing through the output layer, a denoised image is obtained, and the training is completed.

[0221] During the training process, the loss function is calculated for the generated denoised image and the corresponding clear image in the training set so that the network can perform backpropagation. The loss function is where I RHQ is the denoised image output by the neural network processing of the noisy image, I HQ is the label truth value, that is, the clear image, where ∈ is a constant, and in this embodiment, the empirical value is set to 10 -3 . The selected optimizer is Adam, the parameters are default parameters, and the initial learning rate is 2×10 -4 .

[0222] When the loss function converges, a trained network model is obtained.

[0223] Step 3, input the strip-shaped noisy image into the network model, and finally output a clear image through multiple cross-overlapping operations of feature extraction, fusion, and convolution of the network model.

[0224] In the experimental results section, Figure 12(a) and 12(c)Two images with different strip noises are selected. After being processed by the strip noise removal method for labeled images based on the hybrid structure of CNN and Transformer proposed by the present invention, the strip noises in the pictures can be effectively removed, and clear images 12(b) and 12(d) are output. The experimental results show that the present invention can obtain clear images. After the noisy pictures are denoised by the denoising method of the present invention, the objective index peak signal-to-noise ratio PSNR reaches a relatively high 33.28 dB, and the structural similarity SSIM reaches a relatively high 0.9009. It can be seen from the experimental results that the image denoising algorithm of the present invention can remove strip noises in any direction relative to the barcodes in the express labels. Especially in terms of retaining the texture and details of the pictures, the denoising method of the present invention shows excellent denoising results.

[0225] The above description is only the preferred embodiments of the present disclosure and the explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the embodiments of the present disclosure that have similar functions.

[0226] Those skilled in the art can understand that all or part of the processes for implementing the methods of the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disc, a read-only memory or a random access memory, etc.

[0227] The above is only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.

Claims

1. An image de-band noise method based on a hybrid structure of CNN and Transformer, characterized in that including the following steps, Preprocess the noisy image and the corresponding clear image to establish a training set; Build a network model, and train the network model through the training set to obtain a trained network model; the network model adopts a hybrid structure of CNN and Transformer, including an input layer, an encoding and decoding network, and an output layer; the encoding and decoding network includes multiple layers of encoding and decoding units; in each layer of encoding and decoding unit, the encoder and the decoder are both provided with corresponding convolutional layers, which are used to perform convolutional operations on the data input to the corresponding encoder / decoder and then, together with the output of the corresponding encoder / decoder, serve as the input of the subsequent decoder / encoder; Input the image with banded noise into the network model, and after multiple deep feature extraction, feature fusion and upsampling operations of the network model, finally output a clear image.

2. The method for removing banded noise from an image based on a hybrid structure of CNN and Transformer according to claim 1, wherein, The input layer is used to perform shallow feature extraction on the noisy image to obtain a shallow feature image; The encoding and decoding network is used to perform multiple rounds of deep feature extraction, feature fusion and upsampling on the shallow feature image and then output; The output layer is used to fine-tune the image output by the encoding and decoding network and then output a clear image with noise removed.

3. A method for image denoising based on a hybrid structure of CNN and Transformer according to claim 1, characterized in that, Each layer of encoding and decoding units is connected in sequence, the structures of each layer of encoding and decoding units are the same, and each layer of encoding and decoding units is also connected in a residual manner.

4. A method for image denoising based on a hybrid structure of CNN and Transformer according to claim 3, characterized in that, The encoding and decoding unit includes an encoder, a feature fusion module, and a decoder connected in sequence, a convolutional layer parallel to the encoder and the feature fusion module, and a transposed convolutional layer parallel to the decoder; When there are n layers of encoding and decoding units in the network, let the encoder be e i , the feature fusion module be c i , the decoder be d i , the convolutional layer be l i , the transposed convolutional layer be t i , i = 1, 2,... n, where Input to the encoder e i The data synchronization input is input to the convolutional layer l i , and after passing through the encoder e i and the feature fusion module c i The output after serial processing is summed with the output of the convolutional layer l i to obtain the first summation data, and the first summation data is simultaneously used as the input to the decoder d i and the transposed convolutional layer t i ; the first summation data is also summed with the output of the feature fusion module c n-i+1 and the output of the convolutional layer l n-i+1 to obtain the second summation data, and the second summation data is simultaneously used as the input to the decoder d n-i+1 and the transposed convolutional layer l n-i+1 ; Decoder d i 's output is summed with the output of the transposed convolutional layer t i to obtain the third summation data, and the third summation data is simultaneously used as the input of the encoder e i+1 and the convolutional layer l i+1 ; the third summation data is also summed with the output of the decoder d n-i and the transposed convolutional layer t n-i to obtain the fourth summation data, and the fourth summation data is simultaneously used as the input of the encoder e n-i+1 and the convolutional layer l n-i+1 . Decoder d n The output of n the transposed convolutional layer t is summed with the output of the decoder d to be the output data of the encoder-decoder network.

5. A method for image denoising based on a hybrid structure of CNN and Transformer according to claim 4, characterized in that In the encoding and decoding unit, The convolutional layer is used to change the dimension of the input image that has not been encoded and feature-fused to be the same as the dimension of the image output by the feature fusion module; The transposed convolutional layer is used to change the dimension of the input image that has not been decoded to be the same as the dimension of the image output by the decoder.

6. A method for image denoising based on a hybrid structure of CNN and Transformer according to claim 4, characterized in that, The encoder includes a parallel Transformer processing branch and a CFEB processing branch; The Transformer processing branch is used to extract global features and deep semantic features; the CFEB processing branch is used to extract local features and detail features of the image; The two processing branches process the image input to the encoder in parallel and output synchronously.

7. The method for removing noise from an image based on a hybrid structure of CNN and Transformer according to claim 6, wherein, The Transformer processing branch includes multiple sequentially connected Transformer basic units, wherein, Each Transformer basic unit includes a first convolutional layer, a first dimension compression transposed layer, a first normalization layer, a multi-head self-attention, a second normalization layer, a multi-layer perceptron, a second dimension compression transposed layer, a second convolutional layer, and a third convolutional layer connected in sequence; Residual operations are respectively introduced at the end of the multi-head attention and the end of the multi-layer perceptron of each Transformer basic unit.

8. A method for image denoising based on a hybrid structure of CNN and Transformer according to claim 6, characterized in that the CFEB processing branch includes a plurality of convolutional layers and downsampling layers, wherein residual connections are used between the convolutional layers.

9. A method for image denoising based on a hybrid structure of CNN and Transformer according to claim 8, characterized in that, The CFEB processing branch includes: 4m convolutional layers and m - 1 downsampling layers; One downsampling layer is connected after every 4 convolutional layers. The downsampling layer is used to increase the number of channels and reduce the image size; The first convolutional layer increases the number of channels of the noise image while keeping the length and width of the image unchanged; For the second to the second - last convolutional layer, the sum of the input and output of each layer is input to the next layer; After the convolution operation of the last convolutional layer, the processed image is output.

10. A method for image denoising based on a hybrid structure of CNN and Transformer according to claim 7, characterized in that, The image processing process of the Transformer processing branch of the encoder module is as follows: The following processing is performed in each Transformer basic unit: The first convolutional layer adjusts the dimension of the input image; The first dimension compression transpose layer adjusts the format of the input image; The first normalization layer performs normal distribution standardization processing on the input image, The multi - head attention layer marks different levels of content, regions, and long - range dependency relationships of the input image; The sum of the input of the first normalization layer and the output of the multi - head attention layer is output to the second normalization layer; The second normalization layer performs normal distribution standardization processing on the image again; The multi - layer perceptron layer extracts deep - level features of the image; The sum of the input of the second normalization layer and the output of the multi - layer perceptron layer is input to the second dimension compression transpose layer to restore the image format; The second convolutional layer enhances the features of the input image and adjusts the image dimension once; The third convolutional layer enhances the features of the input image, adjusts the image dimension, and then outputs it to the next Transformer basic unit, The input image is output after being processed by each Transformer basic unit in sequence.