A method for infrared image denoising based on convolutional transposed self-attention

By introducing the infrared image denoising method of convolutional transposed self-attention and convolutional gated linear unit, the problem of difficulty in capturing global features in infrared image denoising is solved, and efficient infrared image denoising effect is achieved, which is better than traditional methods.

CN118229572BActive Publication Date: 2025-09-09HUAIAN KUNBO INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410416198.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-08
Publication Date
2025-09-09
Estimated Expiration
2044-04-08

AI Technical Summary

Technical Problem

Existing infrared image denoising methods are difficult to effectively capture global features and semantic information, have high computational costs, and traditional convolutional neural networks perform poorly in infrared image denoising tasks.

Method used

A convolution-transposed self-attention method is used to denoise infrared images. Combining the convolutional gated linear unit and the channel coordinate attention module, the self-attention mechanism is used to model global dependencies, dynamically assign feature map weights, suppress irrelevant information, and restore image details.

Benefits of technology

While reducing the amount of calculation, the spatial feature information of the infrared image is fully extracted, the denoising quality is improved, and it is better than the existing methods, especially the denoising effect of infrared images under different noise intensities is significant.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118229572B_ABST
    Figure CN118229572B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for infrared image denoising based on convolutional transposed self-attention. First, a dataset is constructed by pairing clean images and noisy images; a network model based on convolutional transposed self-attention is constructed, and the network model adopts a symmetrical encoder-decoder structure; the network model based on convolutional transposed self-attention is trained based on the constructed infrared image training dataset; and finally, the trained network model is tested using data from a test dataset. The present invention introduces a self-attention mechanism, which has not been deeply studied in the field of infrared image denoising, to model the global dependency of infrared images. Based on this mechanism, an improved convolutional transposed self-attention is designed. While reducing the amount of computation, it can still fully extract the spatial feature information of infrared images, effectively improving the denoising quality of infrared images and filling the research gap in infrared image denoising.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and in particular relates to an infrared image denoising method based on convolution transposed self-attention. Background Art

[0002] Infrared imaging, as an important non-visible light imaging technology, is widely used in military, medical, security, and other fields due to its anti-interference properties, high penetration, high sensitivity, and high concealment. However, due to the inherent characteristics of infrared imaging equipment and the influence of environmental noise, infrared images are often interfered with by various noises, such as Gaussian noise. This noise not only degrades the visual quality of infrared images but also adversely affects subsequent image analysis and processing tasks. However, overcoming the manufacturing process and physical limitations of infrared imaging system hardware is difficult, and improving image quality through hardware is time-consuming and labor-intensive. Image denoising, a research direction in computer vision, uses algorithms to reconstruct low-quality noisy input images into high-quality clean output images. It plays a vital role in tasks such as small object detection and medical image segmentation. Therefore, it can also be used as a reliable and effective method to improve infrared image quality.

[0003] In past studies, many methods have been proposed in the field of image denoising, which can be mainly divided into three categories: spatial domain-based methods, frequency domain-based methods, and deep learning-based methods.

[0004] Spatial-domain methods directly process the grayscale values ​​corresponding to each pixel in the image in two-dimensional space to achieve noise removal. Typically, the entire image is divided into multiple sub-blocks of fixed size. Based on the statistical characteristics of the pixels in each sub-block and the similarity between adjacent pixels as a theoretical basis, methods such as weighted averaging are used to restore the original image information. These methods are simple and intuitive, often directly operating on image pixels, making them easy to understand and implement. However, they struggle to effectively distinguish image features from noise features. In contrast, frequency-domain methods convert the image to the frequency domain through methods such as Fourier transforms and wavelet transforms for processing. They selectively remove noise in the frequency domain, resulting in a more sparse representation of the signal. The actual image signal is replaced by a sparse representation with fewer coefficients than the original image data, thereby better preserving image details and edges, achieving better denoising and enhancement results. However, these methods may introduce ringing artifacts during the denoising process, causing some ring-like artifacts to appear in the image.

[0005] In recent years, deep learning-based methods have been widely studied due to their significant performance improvements over traditional methods in various computer vision tasks. Deep learning-based methods construct neural network models, perform end-to-end training using pairs of clean and noisy images, and ultimately use the trained network model to denoise and reconstruct the image. Compared to visible light image denoising, relatively few studies have focused on infrared image denoising, and these studies have primarily focused on the application of convolutional neural networks (CNNs). CNN-based infrared denoising methods extract local features of the image through operations such as convolution and pooling. However, their receptive field is limited, and obtaining global features requires stacking multiple convolutional layers or using convolution kernels with larger windows, which incurs high computational costs. Therefore, it is difficult to model the global dependencies of infrared images and cannot fully capture the semantic information in infrared images. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this paper provides an infrared image denoising method based on convolutional transposed self-attention. This method uses convolutional transposed self-attention to model global dependencies in infrared images. It also introduces convolutional gated linear units to improve information flow within the network, transferring useful information and suppressing irrelevant information. Finally, during feature fusion, a channel-coordinate attention module dynamically assigns weights to different parts of the feature map along different dimensions, thereby achieving denoising and reconstruction of low-quality noisy infrared images.

[0007] This paper introduces the Transformer self-attention mechanism to model the global dependencies of infrared images. Based on this, it improves and designs convolutional transposed self-attention, which can fully extract the spatial feature information of infrared images while reducing the amount of computation. Furthermore, a convolutional gated linear unit is used to improve the traditional feedforward network. The gating mechanism improves the information flow in the network, transmits useful information, and suppresses irrelevant information, thereby better restoring the details of the reconstructed image. Finally, a channel coordinate attention module is used for feature fusion, dynamically assigning weights to different parts of the feature map to better represent the pattern characteristics in the infrared image.

[0008] The technical solution adopted by the present invention to solve the technical problem includes the following steps:

[0009] Step 1: First, the obtained open-source infrared image dataset is divided into training and test data, and the data is preprocessed, including image cropping, image noise addition, and image enhancement. This preprocessing generates noisy images that correspond one-to-one with the original clean images. These paired clean and noisy images are used to construct the training and test datasets for the network model.

[0010] Step 2: Construct a network model based on convolutional transposed self-attention. The network model adopts a symmetrical encoder-decoder structure.

[0011] Step 3: Based on the constructed infrared image training dataset, train the network model based on convolutional transposed self-attention.

[0012] Step 4: Use the test data set to test the trained network model and calculate the Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM) of the reconstructed image.

[0013] Furthermore, the open-source infrared image datasets include five: the IR700 dataset, the DLS-NUC-100 dataset, the IR100 dataset, the ESPOLFIR dataset, and the Flir dataset. 600 images from the IR700 dataset were randomly selected, and 64×64 regions were randomly cropped from the images. These images were then further enhanced by horizontal flipping and random rotations of 90, 180, and 270 degrees, forming the training dataset. The remaining 100 images from the IR700 dataset and all images from the other four datasets formed the test dataset. A method for simulating infrared noisy images was used, namely, adding Gaussian white noise with a mean of 0 and standard deviations of 15, 25, and 50 to the images to produce paired clean and noisy images.

[0014] Furthermore, the network model based on convolutional transposed self-attention adopts a symmetrical four-layer encoder-decoder structure, consisting of four encoders and three decoders. The first four encoders gradually reduce the size of the feature map while increasing the channel depth, while the last three decoders gradually recover the low-resolution latent features. The first three encoder-decoder layers consist of L1, L2, and L3 IDTransformer modules, respectively. The fourth encoder layer consists of L4 IDTransformer modules. The number of IDTransformer modules gradually increases with the depth of the layer. Each IDTransformer module contains a convolutional transposed self-attention module and a convolutional gated linear unit for further feature extraction.

[0015] For the input noisy infrared image Where h, w and 1 represent the height, width and number of channels of the input image respectively. Before inputting to the first encoder, a convolution layer with a convolution kernel size of 3×3 is used to extract the features of the input image to obtain a shallow feature map. The input is mapped from a simple low-dimensional space to a more abstract high-dimensional space, so that low-frequency information in the image can be extracted at an early stage. The shallow feature map is then input into the first encoder to start further feature extraction, and a deep feature map is obtained through a four-layer encoder-decoder structure. In this process, the encoder reduces the size of the high-resolution feature map layer by layer through pixel-unshuffle. The channel is expanded to 8c, and the decoder uses pixel-shuffle to restore the shape of the low-resolution feature map to h×w×c layer by layer. In addition, in order to fully extract features, the encoding features and the decoding features are spliced ​​in the channel dimension through jump connections, and the spliced ​​features are fused using the Channel Coordinate Attention Block (CCAB) to achieve channel dimensionality reduction and feature refinement. Subsequently, the obtained deep feature map is passed through L r An IDTransformer module is used to extract features and further enrich feature diversity. Finally, a convolution layer with a convolution kernel size of 3×3 is used to generate a residual image, which is then residually connected with the input image to obtain the final denoised image.

[0016]

[0017] Furthermore, the convolutional transposed self-attention module implicitly encodes global contextual information by calculating self-attention over the channel dimension rather than spatially. Unlike using a simple mapping matrix to generate the query vector Q, key vector K, and value vector V, the convolutional transposed self-attention module introduces point-wise convolution and depth-wise convolution to generate these variables. Point-wise convolution uses a 1×1 kernel size and independently performs convolution operations on each input channel at each spatial position. Depth-wise convolution, on the other hand, uses a kernel equal to the number of input channels. The kernel for each channel focuses solely on the features of that channel and is not mixed with features from other channels. This reduces the number of parameters and computations while more flexibly capturing inter-channel feature relationships. This aggregates pixel-by-pixel cross-channel context and emphasizes channel-by-channel spatial context, compensating for spatial information that may have been lost previously.

[0018] For the intermediate feature map input to the convolutional transposed self-attention module First, layer normalization is used to keep the data distribution stable, and then a convolution layer with a convolution kernel size of 1×1 and a depth-wise convolution layer with a convolution kernel size of 3×3 are applied to generate the query vector Q, key vector K and value vector It can be expressed as:

[0019] Q = DConv3×3 (Conv 1×1 (X ctasb )),

[0020] K=DConv 3×3 (Conv 1×1 (X ctasb )),

[0021] V=DConv 3×3 (Conv 1×1 (X ctasb )),

[0022] Among them, Conv 1×1 Represents a convolution layer with a convolution kernel size of 1×1, DConv 3×3 represents a depth-wise convolution layer with a convolution kernel size of 3×3. In order to perform the transposed self-attention calculation, Q, K, and V need to be transformed first. The calculation process is expressed as follows:

[0023]

[0024] in, and are Q, K and V after deformation respectively, Indicates that the dot product operation is performed to obtain the transposed feature map And use the learnable parameter α to control the size of the transposed feature map, and then use softmax to generate the weight distribution pair Perform weighted operations. Like traditional self-attention calculations, the convolution-transposed self-attention module also uses a multi-head approach to learn independent attention maps in parallel. In addition, to maintain feature diversity, the convolution-transposed self-attention module introduces an additional depth-wise convolution with a kernel size of 3×3 to extract features from the value vector V. This is then added to the result of the self-attention calculation. Finally, it is processed through a convolution layer with a kernel size of 1×1 and a residual connection is made with the original input. This can be expressed as:

[0025]

[0026] in It is the feature map obtained after the convolution transposition self-attention module processing.

[0027] Furthermore, the convolutional gated linear unit improves the information flow in the network through a gating mechanism, controlling which features should be passed forward, suppressing features containing useless information, and allowing only useful information to be passed further through the network layers. This allows subsequent layers of the network to focus on finer image details that complement other layers, thereby producing high-quality output results. In addition, depth-wise convolution is introduced to encode information about spatially adjacent pixel positions, which helps learn local image structure for effective restoration.

[0028] For the intermediate feature map input to the convolutional gated linear unit After processing, the output is a new feature map The intermediate calculation process can be expressed as:

[0029]

[0030] Among them, W i represents linear mapping, σ represents GELU activation function, DConv 3×3 represents a depth-wise convolution layer with a convolution kernel size of 3×3. Represents element-wise multiplication.

[0031] Furthermore, the encoding features and decoding features of the corresponding layers are concatenated in the channel dimension through skip connections, and the channel coordinate attention module is used to perform channel dimensionality reduction and feature refinement on the concatenated features. First, a convolutional layer with a convolution kernel size of 1×1 is used to reduce the channel dimension of the concatenated features, and channel attention and coordinate attention are combined to effectively integrate spatial information while considering the encoding of inter-channel information. Channel attention uses global maximum pooling (GMP) and global average pooling (GAP) to obtain global information of input features. Coordinate attention uses horizontal global average pooling (HGAP) and vertical global average pooling (VGAP) to aggregate input features in the horizontal and vertical directions, respectively, to obtain two independent direction-aware feature maps.

[0032] For the intermediate feature map obtained by concatenating the encoding features and decoding features in the channel dimension First, a convolution layer with a convolution kernel size of 1×1 is used to reduce the channel dimension, and we get It can be expressed as:

[0033]

[0034] After that, CCAB first sequentially converts the intermediate feature maps after dimensionality reduction into After channel attention, a channel attention map is calculated right Each channel of is weightedly fused, so that the network can adaptively focus on the information of different channels, which can be expressed as:

[0035]

[0036] in Indicates element-by-element multiplication. Represents the intermediate feature map obtained after channel attention. The horizontal attention map is calculated by coordinate attention and vertical attention map right Each direction in space is weighted and fused, so that the network can adaptively focus on information in different directions in space and obtain the final output features. It can be expressed as:

[0037]

[0038] Furthermore, in channel attention, GMP and GAP are used to aggregate the global spatial information of the feature map to generate a pair of different spatial context descriptors. and Represent the global maximum pooling feature and the global average pooling feature respectively. Then, both descriptors are input into a weight sharing network for feature refinement. This weight sharing network consists of two convolutional layers with a convolution kernel size of 1×1 and a ReLU activation function. In order to reduce parameter overhead, the number of channels of the descriptors in the weight sharing network is compressed by r times and restored before merging. After processing the two descriptors, the output feature vectors are merged using element-wise summation and the channel attention map is generated using the Sigmoid function. It can be expressed as:

[0039]

[0040] Where σ represents the sigmoid function, and CRC represents a weight-sharing network consisting of two convolutional layers with a convolution kernel size of 1×1 and a ReLU activation function.

[0041] Furthermore, in coordinate attention, in order to enable the attention module to capture spatial long-distance dependencies with precise location information, the global pooling is improved, and the global pooling operation is decomposed into a pair of one-dimensional feature encoding operations. HGAP and VGAP are used to aggregate the spatial information of the feature map in the horizontal and vertical directions, respectively, to obtain a pair of direction-aware descriptors. and Unlike channel attention, this opponent descriptor passes through two networks with non-shared weights. These two networks also consist of two convolutional layers with a convolution kernel size of 1×1 and a ReLU activation function. After compressing the number of channels of the descriptor in the weight-sharing network by r times, the number of channels is restored before the Sigmoid function is processed. Finally, the Sigmoid function is used to generate a horizontal attention map. and vertical attention map It can be expressed as:

[0042]

[0043]

[0044] Each attention map captures the long-range dependencies of the input feature map along a spatial direction, so that the position information can be preserved in the generated attention map. The two attention maps are applied to the input feature map by multiplication to emphasize the representation of interest.

[0045] Furthermore, the specific method of step 3 is as follows: use 600 pairs of low-quality noisy infrared images I after preprocessing N and high-quality clean infrared images I C Perform model training and set I N After inputting the model, the output is the denoised reconstructed image I R , using the output image I R and the original high quality clean infrared image I C Calculate the Charbonnier loss. The calculation process can be expressed as:

[0046]

[0047] Where N is the number of image pairs in each training iteration, i.e., the batch size, which is set to 8 in the experiment. μ is a small positive number used to smooth the loss function, which is set to 1e-9 in the experiment. By minimizing the loss and optimizing the network model parameters using the Adam optimizer, the total number of training iterations is set to 3,000,000, and the learning rate is initialized to 2e-4 and reduced by half at 150,000, 200,000, 250,000, and 275,000 iterations.

[0048] Compared with the existing technology, the present invention has made the following contributions in research and innovation:

[0049] 1. This invention introduces the self-attention mechanism, which has not been deeply studied in the field of infrared image denoising, to model the global dependency of infrared images. On this basis, it improves and designs the convolution transposed self-attention, which can fully extract the spatial feature information of infrared images while reducing the amount of calculation, effectively improving the denoising quality of infrared images, and filling the research gap in infrared image denoising.

[0050] 2. Through the deployment of convolutional transposed self-attention module, convolutional gated linear unit and bidirectional information interaction module, the present invention has achieved advanced level in comparative experiments with cutting-edge methods in the field of image denoising. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments.

[0052] Figure 1 It is a flow chart of a method according to an embodiment of the present invention.

[0053] Figure 2 The infrared images used for training and testing in this invention come from five datasets, including different backgrounds, scenes, and objects both outdoor and indoor.

[0054] Figure 3 It is a schematic diagram of the overall architecture of the model of an embodiment of the present invention.

[0055] Figure 4 2 is a schematic diagram of the structure of the convolutional transposed self-attention module according to an embodiment of the present invention.

[0056] Figure 5 2 is a schematic diagram of the structure of a convolutional gated linear unit according to an embodiment of the present invention.

[0057] Figure 6 2 is a schematic diagram of the channel coordinate attention module structure of an embodiment of the present invention.

[0058] Figure 7 This is a comparison chart of the visual effects of the images reconstructed by the embodiment of the present invention and other mainstream models when the noise intensity is 15.

[0059] Figure 8 This is a comparison chart of the visual effects of the images reconstructed by the embodiment of the present invention and other mainstream models when the noise intensity is 25.

[0060] Figure 9 This is a comparison chart of the visual effects of images reconstructed by the embodiment of the present invention and other mainstream models when the noise intensity is 50. DETAILED DESCRIPTION

[0061] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0062] like Figure 1-9 As shown in the figure, a method for infrared image denoising based on convolution transposed self-attention has the following specific steps:

[0063] Step 1: First, the obtained open-source infrared image dataset is divided into training and test data, and the data is preprocessed, including image cropping, image noise addition, and image enhancement. This preprocessing generates noisy images that correspond one-to-one with the original clean images. These paired clean and noisy images are used to construct the training and test datasets for the network model.

[0064] There are five open-source infrared image datasets: the IR700 dataset (700 images), the DLS-NUC-100 dataset (100 images), the IR100 dataset (100 images), the ESPOL FIR dataset (101 images), and the Flir dataset (50 images). Table 1 further describes each dataset, including the number of images, resolution, and objects and scenes included. Figure 2 The following are some sample images from these datasets. 600 images from the IR700 dataset were randomly selected, and 64×64 regions were randomly cropped from these images. Further image enhancement was performed by horizontal flipping and random rotations of 90, 180, and 270 degrees to form the training dataset. The remaining 100 images from the IR700 dataset and all images from the other four datasets constituted the test dataset. To compare with existing image denoising algorithms, a method of simulating infrared noisy images was used, that is, adding Gaussian white noise with a mean of 0 and standard deviations of 15, 25, and 50 to the images. This resulted in pairs of clean and noisy images, which can be used to train and test the network model under different noise intensities.

[0065]

[0066] Table 1

[0067] Step 2: Build a network model based on convolutional transposition self-attention, such as Figure 3As shown in the figure, the network model adopts a symmetrical four-layer encoder-decoder structure, including four encoders and three decoders. The first four encoders reduce the size of the feature map layer by layer while increasing the channel depth, and the last three decoders restore the low-resolution potential features layer by layer. The first three layers of encoder-decoders are composed of L1, L2, and L3 IDTransformer modules respectively, and the fourth layer of encoder is composed of L4 IDTransformer modules. The number of IDTransformer modules gradually increases with the depth of the layer. Each IDTransformer module contains a Convolutional Transposed Self-Attention Block (CTSAB) and a Convolutional Gated Linear Unit (CGLU) for further feature extraction.

[0068] For the input noisy infrared image Where h, w and 1 represent the height, width and number of channels of the input image respectively. Before inputting to the first encoder, a convolution layer with a convolution kernel size of 3×3 is used to extract the features of the input image to obtain a shallow feature map. The input is mapped from a simple low-dimensional space to a more abstract high-dimensional space, so that low-frequency information in the image can be extracted at an early stage. The shallow feature map is then input into the first encoder to start further feature extraction, and a deep feature map is obtained through a four-layer encoder-decoder structure. In this process, the encoder reduces the size of the high-resolution feature map layer by layer through pixel-unshuffle. The channel is expanded to 8c, and the decoder uses pixel-shuffle to restore the shape of the low-resolution feature map to h×w×c layer by layer. In addition, in order to fully extract features, the encoding features and the decoding features are spliced ​​in the channel dimension through jump connections, and the spliced ​​features are fused using the Channel Coordinate Attention Block (CCAB) to achieve channel dimensionality reduction and feature refinement. Subsequently, the obtained deep feature map is passed through L r An IDTransformer module is used to extract features and further enrich feature diversity. Finally, a convolution layer with a convolution kernel size of 3×3 is used to generate a residual image, which is then residually connected with the input image to obtain the final denoised image.

[0069]

[0070] Since the computational cost in Transformer mainly comes from self-attention calculation, when traditional self-attention calculation is applied to computer vision tasks, the input sequence length depends on the input image resolution, and the computational complexity increases exponentially. In order to alleviate this problem, Figure 4 As shown in the figure, the convolutional transposed self-attention module implicitly encodes global context information by calculating self-attention along the channel dimension rather than spatially. Unlike using a simple mapping matrix to generate the query vector Q, key vector K, and value vector V, the convolutional transposed self-attention module introduces point-wise convolution and depth-wise convolution to generate these variables. Point-wise convolution has a kernel size of 1×1 and performs convolution operations on each input channel independently at each spatial position. Depth-wise convolution, on the other hand, uses a kernel equal to the number of input channels. The kernel for each channel focuses only on the features of that channel and is not mixed with features from other channels. This reduces the number of parameters and computations while more flexibly capturing feature relationships between channels. This aggregates pixel-by-pixel cross-channel context and emphasizes channel-by-channel spatial context, compensating for spatial information that may have been lost previously.

[0071] For the intermediate feature map input to the convolutional transposed self-attention module First, layer normalization (LN) is used to keep the data distribution stable, and then a convolution layer with a convolution kernel size of 1×1 and a depth-wise convolution layer with a convolution kernel size of 3×3 are applied to generate the query vector Q, key vector K and value vector It can be expressed as:

[0072] Q = DConv 3×3 (Conv 1×1 (X ctasb )),

[0073] K=DConv 3×3 (Conv 1×1 (X ctasb )),

[0074] V=DConv 3×3 (Conv 1×1 (X ctasb )),

[0075] Among them, Conv 1×1 Represents a convolution layer with a convolution kernel size of 1×1, DConv 3×3 represents a depth-wise convolution layer with a convolution kernel size of 3×3. In order to perform the transposed self-attention calculation, Q, K, and V need to be transformed first. The calculation process is expressed as follows:

[0076]

[0077] in, and are Q, K and V after deformation respectively, Indicates that the dot product operation is performed to obtain the transposed feature map And use the learnable parameter α to control the size of the transposed feature map, and then use softmax to generate the weight distribution pair Perform weighted operations. Like traditional self-attention calculations, the convolution-transposed self-attention module also uses a multi-head approach to learn independent attention maps in parallel. In addition, to maintain feature diversity, the convolution-transposed self-attention module introduces an additional depth-wise convolution with a kernel size of 3×3 to extract features from the value vector V. This is then added to the result of the self-attention calculation. Finally, it is processed through a convolution layer with a kernel size of 1×1 and a residual connection is made with the original input. This can be expressed as:

[0078]

[0079] in It is the feature map obtained after the convolution transposition self-attention module processing.

[0080] In the Transformer model, in addition to self-attention calculation, the feedforward network is also an important component. It is applied after each self-attention layer as a nonlinear mapping, usually consisting of two linear layers sandwiched by a nonlinear activation function, to transform features and enhance the representation ability of the model during encoding and decoding. However, the feedforward network of this structure ignores the modeling of spatial information, and the redundant information in the channel hinders the feature expression ability, which may affect the performance and representation ability of the model. To solve this problem, such as Figure 5 As shown in Figure 1, the convolutional gated linear unit improves the information flow in the network through a gating mechanism, controlling which features should be passed forward, suppressing features containing useless information, and allowing only useful information to be passed further through the network layers. This allows subsequent layers of the network to focus on finer image details that complement other layers, thereby producing high-quality output results. In addition, depth-wise convolution is introduced to encode information about spatially adjacent pixel positions, which helps learn local image structure for effective restoration.

[0081] For the intermediate feature map input to the convolutional gated linear unit After processing, the output is a new feature map The intermediate calculation process can be expressed as:

[0082]

[0083] Among them, W i represents linear mapping, σ represents GELU activation function, DConv 3×3 represents a depth-wise convolution layer with a convolution kernel size of 3×3. Represents element-wise multiplication.

[0084] The encoding features and decoding features of the corresponding layers are spliced ​​in the channel dimension through jump connections, and the channel coordinate attention module is used to perform channel dimension reduction and feature refinement on the spliced ​​features. The channel coordinate attention module structure is as follows: Figure 6 As shown. First, a convolutional layer with a convolution kernel size of 1×1 is used to reduce the channel dimension of the spliced ​​features, and channel attention and coordinate attention are combined to effectively integrate spatial information while considering the encoding of inter-channel information. Channel attention uses global maximum pooling (GMP) and global average pooling (GAP) to obtain the global information of the input features. Coordinate attention uses horizontal global average pooling (HGAP) and vertical global average pooling (VGAP) to aggregate the input features in the horizontal and vertical directions respectively, thereby obtaining two independent direction-aware feature maps, which can be complementary applied to the input feature map to enhance the representation of the object of interest.

[0085] For the intermediate feature map obtained by concatenating the encoding features and decoding features in the channel dimension First, a convolution layer with a convolution kernel size of 1×1 is used to reduce the channel dimension, and we get It can be expressed as:

[0086]

[0087] After that, CCAB first sequentially converts the intermediate feature maps after dimensionality reduction into After channel attention, a channel attention map is calculated right Each channel of is weightedly fused, so that the network can adaptively focus on the information of different channels, which can be expressed as:

[0088]

[0089] in Indicates element-by-element multiplication. Represents the intermediate feature map obtained after channel attention. The horizontal attention map is calculated by coordinate attention and vertical attention map right Each direction in space is weighted and fused, so that the network can adaptively focus on information in different directions in space and obtain the final output features. It can be expressed as:

[0090]

[0091] In channel attention, GMP and GAP are used to aggregate the global spatial information of the feature map and generate a pair of different spatial context descriptors. and Represent the global maximum pooling feature and the global average pooling feature respectively. Then, both descriptors are input into a weight sharing network for feature refinement. This weight sharing network consists of two convolutional layers with a convolution kernel size of 1×1 and a ReLU activation function. In order to reduce parameter overhead, the number of channels of the descriptors in the weight sharing network is compressed by r times and restored before merging. After processing the two descriptors, the output feature vectors are merged using element-wise summation and the channel attention map is generated using the Sigmoid function. It can be expressed as:

[0092]

[0093] Where σ represents the sigmoid function, and CRC represents a weight-sharing network consisting of two convolutional layers with a convolution kernel size of 1×1 and a ReLU activation function.

[0094] In coordinate attention, in order to enable the attention module to capture spatial long-distance dependencies with precise location information, the global pooling is improved and the global pooling operation is decomposed into a pair of one-dimensional feature encoding operations. HGAP and VGAP are used to aggregate the spatial information of the feature map in the horizontal and vertical directions respectively, and a pair of direction-aware descriptors are obtained. and Unlike channel attention, this pair of descriptors passes through two networks that do not share weights. This is because channel attention uses the maximum and average values ​​to describe global spatial information, while coordinate attention describes spatial information in two directions. The two networks are also composed of two convolution layers with a convolution kernel size of 1×1 and a ReLU activation function. After compressing the number of channels of the descriptor in the weight-sharing network by r times, the number of channels is restored before the Sigmoid function is processed, and finally the Sigmoid function is used to generate the horizontal attention map. and vertical attention map It can be expressed as:

[0095]

[0096]

[0097] Each attention map captures the long-range dependencies of the input feature map along a spatial direction, so that the position information can be preserved in the generated attention map. The two attention maps are applied to the input feature map by multiplication to emphasize the representation of interest.

[0098] Step 3: Based on the constructed infrared image training dataset, train the network model based on convolutional transposed self-attention.

[0099] In the embodiment of the present invention, the network model adopts a four-layer encoder-decoder structure. The number of IDTransformer blocks from the first to the fourth layer are 4, 6, 6 and 8 respectively. The number of channels in each layer is 48, 96, 192 and 384 respectively. The number of heads for multi-head self-attention calculation in the layer is 1, 2, 4 and 8 respectively. The number of IDTransformer blocks used for feature extraction is set to 4. The experimental environment is based on the PyTorch framework, running on Ubuntu 20.04, and using two Nvidia GeForce 3090Ti graphics cards for parallel accelerated computing. 600 pairs of low-quality noisy infrared images I after preprocessing are used. N and high-quality clean infrared images I C Perform model training and set I N After inputting the model, the output is the denoised reconstructed image I R , using the output image I R and the original high quality clean infrared image I C Calculate the Charbonnier loss. The calculation process can be expressed as:

[0100]

[0101] Where N is the number of image pairs in each training iteration, i.e., the batch size, which is set to 8 in the experiment. μ is a small positive number used to smooth the loss function, which is set to 1e-9 in the experiment. By minimizing the loss and optimizing the network model parameters using the Adam optimizer, the total number of training iterations is set to 3,000,000, and the learning rate is initialized to 2e-4 and reduced by half at 150,000, 200,000, 250,000, and 275,000 iterations.

[0102] Step 4. Following the process of step 3, use the same training method to train several mainstream image denoising network models proposed in recent years, including the denoising methods REDNet, DnCNN, MemNet, MWCNN and DRUNet for visible light images, and the denoising methods MLFAN and SMNet for infrared images.

[0103] Step 5: Test each trained model using test data. The Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM) of the reconstructed images are calculated and compared, and the performance of the models is verified using these metrics. PSNR measures the difference between the reconstructed image after algorithmic denoising and the original clean image. It reflects the overall characteristics of the image and is measured in decibels (dB). A higher PSNR value indicates that the reconstructed image is closer to the original clean image, and the denoising effect is better. In contrast, SSIM focuses more on the texture features of the image and has a value range of 0 to 1. The closer the SSIM value is to 1, the better the texture features of the reconstructed image are restored relative to the original clean image. The test data includes the remaining 100 images from the IR700 dataset, 100 images from the DLS-NUC-100 dataset, 100 images from the IR100 dataset, 101 images from the ESPOL FIR dataset, and 50 images from the Flir dataset. The infrared images in these datasets include a variety of outdoor and indoor backgrounds, scenes, and objects, which can better verify the performance and generalization ability of the model. Tables 2, 3, and 4 respectively show the quantitative results of each model at noise intensities of 15, 25, and 50. It can be seen that the proposed IDTransformer achieves the best results for different test sets under different noise intensities. When tested on the IR700_test dataset, at each noise intensity, the PSNR improved by 0.08dB to 0.12dB and the SSIM improved by 0.0007 to 0.0016 compared with the best method; when tested on the DLS-NUC-100 dataset, at each noise intensity, the PSNR improved by 0.06dB to 0.13dB and the SSIM improved by 0.0005 to 0.0017 compared with the best method; when tested on the IR100 dataset, at each noise intensity, the PSNR improved by 0.08dB to 0.12dB and the SSIM improved by 0.0007 to 0.0016 compared with the best method The method was improved by 0.02dB~0.05dB, and the SSIM was improved by 0.0002~0.0006; when tested on the ESPOLFIR dataset, under various noise intensities, the PSNR was improved by 0.03dB~0.05dB, and the SSIM was improved by 0.0001~0.0012 compared with the best method; when tested on the Flir dataset, under various noise intensities, the PSNR was improved by 0.07dB~0.008dB, and the SSIM was improved by 0.0009~0.0011 compared with the best method.At the same time, on five test datasets, when the noise intensity is 15, the PSNR is improved by an average of 0.08dB to 0.57dB, and the SSIM is improved by an average of 0.0010 to 0.0058; when the noise intensity is 25, the PSNR is improved by an average of 0.06dB to 0.85dB, and the SSIM is improved by an average of 0.0007 to 0.0126; when the noise intensity is 50, the PSNR is improved by an average of 0.06dB to 1.34dB, and the SSIM is improved by an average of 0.0007 to 0.0339. The improvement of the proposed model on the IR700_test dataset is the most obvious, probably because it shares the same data distribution with the training dataset. In comparison, although the improvement on other datasets is not significant, it is still better than other denoising methods, which shows the effectiveness and robustness of the proposed model.

[0104]

[0105]

[0106] Table 2

[0107]

[0108] Table 3

[0109]

[0110]

[0111] Table 4

[0112] To verify the effectiveness of CTSAB, CGLU, and CCAB, as well as the internal designs within CTSAB and CCAB, we conducted ablation experiments on these modules. Considering that the impact of each module on the network performance is similar under different noise intensities, we chose to test under a noise intensity of 50. The results shown in Table 5 illustrate the effectiveness of each module in the proposed model. The CGLU in the IDTransformer is replaced by a feedforward network consisting of two linear mapping layers and a ReLU activation function, while CCAB is replaced by a convolutional layer with a convolution kernel size of 1×1. The main body of the network is now composed of CTSAB. CTSAB introduces the self-attention mechanism into the infrared image denoising task, models the global dependencies of the image, and performs better than most models on various datasets. For CGLU and CCAB, from the five test data sets, without CGLU, the PSNR is reduced by 0.02dB on average, and the SSIM is reduced by 0.0002 on average. Without CCAB, the PSNR is reduced by 0.05dB on average, and the SSIM is reduced by 0.0007 on average. When both CGLU and CCAB are removed, the PSNR is reduced by 0.07dB on average, and the SSIM is reduced by 0.0009 on average. The results in Table 6 illustrate the effectiveness of the product-transposed self-attention module design in the proposed model, where w / o means not used, w / means used, and 3DConv p Indicates the use of depth-wise convolution with a convolution kernel size of 3×3 for feature mapping of Q, K, and V, 3DConv vDenotes the use of depth-wise convolution with a kernel size of 3×3 for feature extraction of V. Experimental results show that when depth-wise convolution is not used for feature mapping of Q, K, and V, the network model's performance degrades significantly. This indicates that the spatial context information supplemented by depth-wise convolution is essential for CTASB when calculating channel-wise self-attention, and the feature extraction of V further ensures feature diversity. Across five test datasets, when depth-wise convolution is not used for feature mapping of Q, K, and V, the average PSNR decreases by 0.12dB and the average SSIM decreases by 0.0022. When depth-wise convolution is not used for feature extraction of V, the average PSNR decreases by 0.02dB and the average SSIM decreases by 0.0002. When depth-wise convolution is not used at all, the average PSNR decreases by 0.14dB and the average SSIM decreases by 0.0029. These experimental results demonstrate the effectiveness of CTSAB's internal design. The results shown in Table 7 illustrate the effectiveness of the channel coordinate attention module design in the proposed model. w / o indicates not used, w / indicates used, CH indicates channel attention, and CO indicates coordinate attention. When the module does not include channel attention and coordinate attention, only one convolutional layer with a convolution kernel size of 1×1 is used to reduce the feature dimension. From the five test datasets, without channel attention, the PSNR decreased by an average of 0.02dB and the SSIM decreased by an average of 0.0002. Without coordinate attention, the PSNR decreased by an average of 0.03dB and the SSIM decreased by an average of 0.0005. When both channel attention and coordinate attention were removed, the PSNR decreased by an average of 0.05dB and the SSIM decreased by an average of 0.0007. The experimental results demonstrate that CTSAB, CGLU, and CCAB, as well as the internal designs in CTSAB and CCAB, are effective in improving model performance.

[0113]

[0114]

[0115] Table 5

[0116]

[0117] Table 6

[0118]

[0119] Table 7

[0120] Step 6: Compare the visual effects of the images reconstructed by each model to intuitively verify the model performance. Figure 7 、8 As shown in Figures 9 and 10, representative images from each test dataset were selected for illustration at varying noise intensities. The white box in the left image is a magnified area used for detail comparison. The images on the right, from top to bottom and from left to right, show the original clear image, the noisy image after adding noise, the image denoised using DnCNN, the image denoised using DRUNet, the image denoised using SMNet, and the image denoised using IDTransformer. In Figure 171.png, a metal nameplate from the rear of a car is selected for comparison. Compared to images restored by other methods, the image restored by IDTransformer more clearly discerns the "Volkswagen" logo. In Figure 0037.png, a guardrail around a house is selected for comparison. Compared to other image denoising methods, IDTransformer can more clearly restore the original structure of the guardrail while removing noise. In 34.png, the fan grille of an air conditioner outdoor unit was selected for comparison. In the clear image, the fan grille is composed of straight lines and curves. In the images restored by other methods, the straight lines are blurred and the curves are irregular. In contrast, the image restored by IDTransformer has more regular curves. In thermal_050-S.png, the car logo on the front of a vehicle was selected for comparison. Compared with the images restored by other methods, IDTransformer can better restore the original outline and internal shape of the car logo. In 419.png, the floor structure of a residential building was selected for comparison. In the clear image, the windows and exterior walls between floors make the structure appear grid-like. Although IDTransformer cannot clearly restore the vertical structure of the floors, the horizontal structure is easy to observe. The images restored by other methods are relatively blurry, and the original texture structure cannot be seen. In 23.png, the wire mesh around the stadium is selected for comparison. The images restored by other methods often have discontinuous lines or uneven thickness. In contrast, the wire mesh restored by IDTransformer is more natural in visual effect.

[0121] The above description is a further detailed description of the present invention in conjunction with specific / preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. Those skilled in the art of the present invention may make various substitutions or modifications to the described embodiments without departing from the scope of the present invention, and such substitutions or modifications should be considered to fall within the scope of protection of the present invention.

[0122] Parts of the present invention that are not described in detail belong to the common knowledge of those skilled in the art.

Claims

1. A method for infrared image denoising based on convolutional transposed self-attention, characterized in that: The steps include: Step 1: First, the obtained open source infrared image dataset is divided into training data and test data, and the data is preprocessed. The preprocessing includes image cropping, image noise addition, and image enhancement. Through preprocessing, we obtain noisy images that correspond one-to-one to the original clean images. We then use pairs of clean and noisy images to construct training and test datasets for the network model. Step 2: Build a network model based on convolutional transposed self-attention, which adopts a symmetrical encoder-decoder structure; Step 3: Based on the constructed infrared image training dataset, train the network model based on convolutional transposed self-attention; Step 4: Use the test data set to test the trained network model and calculate the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) of the reconstructed image. The network model based on convolutional transposed self-attention adopts a symmetrical four-layer encoder-decoder structure, including four encoders and three decoders. The first four encoders gradually reduce the size of the feature map while increasing the channel depth, and the last three decoders gradually restore the low-resolution latent features. The first three encoder-decoder layers consist of L1, L2, and L3 IDTransformer modules, respectively. The fourth encoder layer consists of L4 IDTransformer modules. The number of IDTransformer modules increases with the depth of the layers. Each IDTransformer module contains a convolutional transposed self-attention module and a convolutional gated linear unit for further feature extraction. For the input noisy infrared image Where h, w and 1 represent the height, width and number of channels of the input image respectively. Before inputting to the first encoder, a convolution layer with a convolution kernel size of 3×3 is used to extract the features of the input image to obtain a shallow feature map. The input is mapped from a simple low-dimensional space to a more abstract high-dimensional space, thereby extracting low-frequency information in the image at an early stage; the shallow feature map is then input into the first encoder to start further feature extraction, and a deep feature map is obtained through a four-layer encoder-decoder structure. In this process, the encoder reduces the size of the high-resolution feature map layer by layer through pixel-unshuffle. The channel is expanded to 8c, and the decoder uses pixel-shuffle to restore the shape of the low-resolution feature map to h×w×c layer by layer. In addition, in order to fully extract features, the encoding features and the decoding features are spliced ​​in the channel dimension through jump connections, and the channel coordinate attention module is used to fuse the spliced ​​features to achieve channel dimensionality reduction and feature refinement. Subsequently, the obtained deep feature map is passed through L r An IDTransformer module is used to extract features and further enrich feature diversity. Finally, a convolution layer with a convolution kernel size of 3×3 is used to generate a residual image, which is then residually connected with the input image to obtain the final denoised image. The convolutional transposed self-attention module implicitly encodes global contextual information by calculating self-attention in the channel dimension rather than in space. Unlike using a simple mapping matrix to generate the query vector Q, key vector K, and value vector V, the convolutional transposed self-attention module introduces point-wise convolution and depth-wise convolution to generate these variables. The convolution kernel size of point-wise convolution is 1×1, and convolution operations are performed independently on each channel of the input at each spatial position. Depth-wise convolution uses a convolution kernel equal to the number of input channels. The convolution kernel of each channel focuses only on the features of that channel and is not mixed with features of other channels. While reducing the number of parameters and computations, it can more flexibly capture the feature relationship between channels, thereby aggregating pixel-by-pixel cross-channel context and emphasizing channel-by-channel spatial context, compensating for spatial information that may have been lost before. For the intermediate feature map input to the convolutional transposed self-attention module First, layer normalization is used to keep the data distribution stable, and then a convolution layer with a convolution kernel size of 1×1 and a depth-wise convolution layer with a convolution kernel size of 3×3 are applied to generate the query vector Q, key vector K and value vector It can be expressed as: Q=DConv 3×3 (Conv 1×1 (X ctasb )), K=DConv 3×3 (Conv 1×1 (X ctasb )), V=DConv 3×3 (Conv 1×1 (X ctasb )), Among them, Conv 1×1 Represents a convolution layer with a convolution kernel size of 1×1, DConv 3×3 represents a depth-wise convolution layer with a convolution kernel size of 3×3. In order to perform the transposed self-attention calculation, Q, K, and V need to be deformed first. The calculation process is expressed as follows: in, and are Q, K and V after deformation respectively, Indicates that the dot product operation is performed to obtain the transposed feature map And use the learnable parameter α to control the size of the transposed feature map, and then use softmax to generate the weight distribution pair Perform weighted operations; Like traditional self-attention calculations, the convolution-transposed self-attention module also uses a multi-head approach to group and parallelly learn independent attention maps; In addition, to maintain feature diversity, the convolution-transposed self-attention module introduces an additional depth-wise convolution with a convolution kernel size of 3×3 to extract features from the value vector V, and adds it to the result of the self-attention calculation. Finally, it is processed by a convolution layer with a convolution kernel size of 1×1 and residually connected with the original input, which can be expressed as follows: in It is the feature map obtained after the convolution transposition self-attention module processing.

2. The infrared image denoising method based on convolutional transposed self-attention according to claim 1, characterized in that: The open-source infrared image datasets include five: the IR700 dataset, the DLS-NUC-100 dataset, the IR100 dataset, the ESPOLFIR dataset, and the Flir dataset. 600 images from the IR700 dataset were randomly selected, and 64×64 regions were randomly cropped from the images. These images were then further enhanced by horizontal flipping and random rotations of 90, 180, and 270 degrees, thus forming a training dataset. The remaining 100 images from the IR700 dataset and all images from the other four datasets constituted a test dataset. A method of simulating infrared noise images was used, i.e., Gaussian white noise with a mean of 0 and standard deviations of 15, 25, and 50, respectively, was added to the images to obtain paired clean and noisy images.

3. The infrared image denoising method based on convolutional transposed self-attention according to claim 1, characterized in that: The convolutional gated linear unit improves the information flow in the network through a gating mechanism, controlling which features should be passed forward, suppressing features containing useless information, and allowing only useful information to be passed further through the network layers. This allows subsequent layers of the network to focus on finer image details that complement other layers, thereby producing high-quality output results. In addition, depth-wise convolution is introduced to encode information about spatially adjacent pixel positions, which helps learn local image structures for effective restoration. For the intermediate feature map input to the convolutional gated linear unit After processing, the output is a new feature map The intermediate calculation process can be expressed as: Among them, W i represents linear mapping, σ represents GELU activation function, DConv 3×3 represents a depth-wise convolution layer with a convolution kernel size of 3×3. Represents element-wise multiplication.

4. The infrared image denoising method based on convolutional transposed self-attention according to claim 1, characterized in that: The encoding features and decoding features of the corresponding layers are concatenated in the channel dimension through skip connections, and the channel coordinate attention module is used to reduce the channel dimension and refine the features after concatenation. First, a convolutional layer with a convolution kernel size of 1×1 is used to reduce the channel dimension of the concatenated features. Then, channel attention and coordinate attention are combined to effectively integrate spatial information while considering the encoding of inter-channel information. Channel attention uses global maximum pooling GMP and global average pooling GAP to obtain the global information of input features; Coordinate attention uses horizontal global average pooling (HGAP) and vertical global average pooling (VGAP) to aggregate input features in the horizontal and vertical directions respectively, thereby obtaining two independent direction-aware feature maps. For the intermediate feature map obtained by concatenating the encoding features and decoding features in the channel dimension First, a convolution layer with a convolution kernel size of 1×1 is used to reduce the channel dimension, and we get It can be expressed as: After that, CCAB first sequentially converts the intermediate feature maps after dimensionality reduction into After channel attention, a channel attention map is calculated right Each channel of is weightedly fused, so that the network can adaptively focus on the information of different channels, which can be expressed as: in Indicates element-by-element multiplication. Represents the intermediate feature map obtained after channel attention; then The horizontal attention map is calculated by coordinate attention and vertical attention map right Each direction in space is weighted and fused, so that the network can adaptively focus on information in different directions in space and obtain the final output features. It can be expressed as:

5. The infrared image denoising method based on convolutional transposed self-attention according to claim 4, characterized in that: In channel attention, GMP and GAP are used to aggregate the global spatial information of the feature map and generate a pair of different spatial context descriptors. and Represent the global maximum pooling feature and the global average pooling feature respectively; then, both descriptors are input into a weight sharing network for feature refinement. This weight sharing network consists of two convolutional layers with a convolution kernel size of 1×1 and a ReLU activation function. In order to reduce parameter overhead, the number of channels of the descriptors in the weight sharing network is compressed by a factor of r and restored before merging; After processing these two descriptors, the output feature vectors are merged using element-wise summation and the channel attention map is generated using the Sigmoid function. It can be expressed as: Where σ represents the sigmoid function, and CRC represents a weight-sharing network consisting of two convolutional layers with a convolution kernel size of 1×1 and a ReLU activation function.

6. The infrared image denoising method based on convolutional transposed self-attention according to claim 5, characterized in that: In coordinate attention, in order to enable the attention module to capture spatial long-distance dependencies with precise location information, the global pooling is improved and the global pooling operation is decomposed into a pair of one-dimensional feature encoding operations. HGAP and VGAP are used to aggregate the spatial information of the feature map in the horizontal and vertical directions respectively, and a pair of direction-aware descriptors are obtained. and Unlike channel attention, this opponent descriptor passes through two networks with non-shared weights. These two networks also consist of two convolutional layers with a convolution kernel size of 1×1 and a ReLU activation function. After compressing the number of channels of the descriptor in the weight-sharing network by r times, the number of channels is restored before the Sigmoid function is processed. Finally, the Sigmoid function is used to generate a horizontal attention map. and vertical attention map It can be expressed as: Each attention map captures the long-range dependencies of the input feature map along a spatial direction, so that the position information can be preserved in the generated attention map. The two attention maps are applied to the input feature map by multiplication to emphasize the representation of interest.

7. The infrared image denoising method based on convolution-transposed self-attention according to any one of claims 1 to 6, characterized in that: The specific method of step 3 is as follows: use 600 pairs of low-quality noisy infrared images I after preprocessing N and high-quality clean infrared images I C Perform model training and set I N After inputting the model, the output is the denoised reconstructed image I R , using the output image I R and the original high quality clean infrared image I C Calculate the Charbonnier loss, the calculation process is expressed as: Where N is the number of image pairs in each iterative training, i.e., the batch size, which is set to 8 in the experiment, and μ is a small positive number used to smooth the loss function, which is set to 1e-9 in the experiment. By minimizing the loss and optimizing the network model parameters using the Adam optimizer, the total number of training iterations is set to 3,000,000, and the learning rate is initialized to 2e-4 and reduced by half at 150,000, 200,000, 250,000, and 275,000 iterations.

Citation Information

Patent Citations

  • Deep learning infrared image denoising method and system based on multi-head self-attention mechanism

    CN114399433A

  • Expression recognition method based on attention-modulated contextual spatial information

    WO2023185243A1