Terahertz image super-resolution reconstruction method and system based on multi-dimensional attention fusion network
By adopting a multi-dimensional attention fusion network in terahertz image super-resolution reconstruction, combining ESRGAN and multi-scale fusion modules, the problems of multi-scale feature extraction and complex degradation adaptability in the prior art are solved, and efficient and accurate image reconstruction is achieved.
Patent Information
- Application Number
- CN202510460102.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The prior art is difficult to take into account multi-scale feature extraction, complex degradation adaptability and lightweight calculations in super-resolution reconstruction of terahertz images, resulting in insufficient reconstruction accuracy and speed.
Using a multi-dimensional attention fusion network method, a network is constructed to realize super-resolution reconstruction of terahertz images through ESRGAN super-segment network, multi-scale fusion module and decomposition channel attention mechanism.
It improves the accuracy and robustness of super-resolution reconstruction of terahertz images, while reducing network complexity, achieving fast and lightweight computing efficiency, and is suitable for real-time applications.
Smart Images

Figure CN120147136A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image super-resolution, and particularly relates to a terahertz image super-resolution reconstruction method and system based on a multi-dimensional attention fusion network. Background Art
[0002] With its unique penetrability and non-ionizing characteristics, terahertz imaging technology has shown important application values in fields such as security inspection and biomedicine. However, limited by the diffraction limit of the hardware system and environmental noise interference, the originally acquired terahertz images generally have problems such as low resolution and blurred textures, seriously affecting the accuracy of key tasks such as dangerous goods identification and pathological diagnosis.
[0003] In recent years, image super-resolution methods based on deep learning have made remarkable progress in the visible light field. However, traditional interpolation-based super-resolution methods will generate artifacts when reconstructing high-frequency details, and existing deep learning solutions have certain limitations. Existing methods rely on single-scale convolutional kernels and are difficult to capture multi-band features, resulting in the loss of key details such as the edges of metal knives; they are insufficient in modeling complex degradations in actual imaging (such as blur, noise, and downsampling aliasing), and when the degradation factor σ>0.3, the PSNR index drops sharply by more than 3dB; directly using large-size convolutional kernels can expand the receptive field, but it leads to a quadratic growth in the number of model parameters, making it difficult to meet the real-time requirements of security inspection equipment. Therefore, it is urgent to construct a dedicated reconstruction framework that takes into account multi-scale feature extraction capabilities, complex degradation adaptability, and lightweight calculations. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides a terahertz image super-resolution reconstruction method based on a multi-dimensional attention fusion network. By using a super-resolution network based on ESRGAN, a multi-scale fusion module based on large kernel convolution decomposition, and a decomposed channel attention mechanism to construct a network, a terahertz image of the person to be inspected is taken by a terahertz imaging device, and super-resolution processing of the target image is performed by a corresponding multiple. The method includes:
[0005] Step S1: Use a terahertz image acquisition device to collect terahertz images of the person to be detected to form a terahertz data set;
[0006] Step S2: Use a degradation model to degrade the features of the terahertz data set to generate a low-resolution image training data set with extensive degradation features;
[0007] Step S3: Construct a multi-dimensional attention fusion network based on multi-scale fusion and channel attention mechanism, and use the low-resolution image training data set with extensive degradation features to train the multi-dimensional attention fusion network to obtain a trained multi-dimensional attention fusion network;
[0008] Step S4: Input the real-time acquired terahertz image into the trained multi-dimensional attention fusion network to obtain the super-resolution reconstructed image of the terahertz image.
[0009] Optionally, the content of step S2 specifically includes:
[0010] Input the terahertz image I in the terahertz dataset collected in step S1 HR into the blurring convolution for spatial domain degradation processing to obtain the spatially degraded image;
[0011] Convolve the spatially degraded image with the terahertz image I HR to obtain the blurred feature map;
[0012] Overlay composite noise on the blurred feature map to obtain the noisy blurred feature map;
[0013] Use bicubic interpolation to downsample the blurred feature map to obtain the sampled blurred feature map;
[0014] Add the noisy blurred feature map and the sampled blurred feature map to obtain a low-resolution image with extensive degradation features;
[0015] Generate low-resolution images for all terahertz images in the terahertz dataset to obtain a low-resolution image training dataset with extensive degradation features.
[0016] Optionally, the content of the spatially degraded image specifically is:
[0017]
[0018] where σ a is the standard deviation of the Gaussian kernel controlling the x direction respectively, and σ b is the standard deviation of the Gaussian kernel controlling the y direction respectively, and (a, b) is the kernel coordinate.
[0019] Optionally, in step S3, the specific content of the multi-dimensional attention fusion network includes:
[0020] Perform shallow feature extraction on the low-resolution terahertz image I in the low-resolution image training dataset LR to obtain the shallow feature map;
[0021] After expanding the channel dimension of the shallow feature map using the convolutional layer, perform convolution and channel segmentation to obtain the segmented feature maps κ 1 and κ 2 ;
[0022] Use point convolution on the segmented feature maps κ 1 and κ 2Process it to obtain a locally enhanced feature map;
[0023] Concatenate the locally enhanced feature maps in the channel dimension and then adjust the number of channels to obtain an output feature map;
[0024] Obtain shallow and middle-level features based on the output feature map;
[0025] Input the shallow and middle-level features into a multi-scale fusion block to obtain a multi-scale fusion output map.
[0026] Optionally, the process of obtaining the locally enhanced feature map is specifically as follows:
[0027] κ(κ 1 , κ 2 ) = split(f c (f c (κ) 3×3 ) 1×1 );
[0028] Where split(·) is a feature segmentation function used to segment the feature map in the channel dimension; κ 1 , κ 2 ∈ R c×h×w are the segmented feature maps; f c (·) 1×1 is a convolutional layer with a kernel size of 1×1.
[0029] Optionally, the process of obtaining the shallow and middle-level features is specifically as follows:
[0030] Generate channel statistics by global average pooling of the output feature map, and obtain the channel attention output feature map of any element in the channel statistics through a gated unit;
[0031] Output the feature map and the channel attention output feature map through a long skip connection to obtain the shallow and middle-level features of this element;
[0032] Add the feature maps element by element to obtain the shallow and middle-level features.
[0033] Optionally, the content of the multi-scale fusion block specifically includes:
[0034] Input the shallow and middle-level features into two convolutional blocks to obtain a feature map;
[0035] Extract multi-dimensional feature information χ i1 , χ i2 from the feature map through a large kernel self-attention structure;
[0036] Obtain the output result χ i1 of LKA for the multi-dimensional feature information χ iΔ1 through a large kernel attention function;
[0037] Multiply the output result χ iΔ1 element - by - element with the multi - dimensional feature information χ i2 to obtain the output feature map of the LSA;
[0038] Convolve and adjust the number of channels of the output feature map of the LSA, and then sum them element - by - element to obtain the output result of the scale fusion block.
[0039] The present invention also discloses a terahertz image super - resolution reconstruction system based on a multi - dimensional attention fusion network. The system includes:
[0040] A data acquisition module, configured to collect terahertz images of the person to be detected using a terahertz image acquisition device to form a terahertz data set;
[0041] A training data set construction module, configured to use a degradation model to perform feature degradation on the terahertz data set to generate a low - resolution image training data set with a wide range of degradation features;
[0042] A network construction module, configured to construct a multi - dimensional attention fusion network based on multi - scale fusion and channel attention mechanisms, and use the low - resolution image training data set with a wide range of degradation features to train the multi - dimensional attention fusion network to obtain a trained multi - dimensional attention fusion network;
[0043] An image reconstruction module, configured to input the terahertz image collected in real - time into the trained multi - dimensional attention fusion network to obtain a terahertz image super - resolution reconstructed image.
[0044] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0045] The present invention provides a terahertz image super - resolution reconstruction method and system based on a multi - dimensional attention fusion network. General super - resolution algorithms are difficult to balance reconstruction speed and accuracy. The present invention effectively captures the detailed features in terahertz images by introducing a multi - scale fusion module (MSFM) and a decomposed channel attention module (DCAM), thereby improving the accuracy and robustness of super - resolution reconstruction. At the same time, the designed MSFM uses decomposed large - kernel convolution, and DCAM uses channel attention, which can improve the computational efficiency while reducing the network complexity. This method has the advantages of fast speed, light weight and accuracy, and is suitable for real - time applications and various scenarios of terahertz image super - resolution. It has broad application prospects in the fields of terahertz imaging, security inspection, non - destructive testing, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] To more clearly illustrate the technical solution of the present invention, the accompanying drawings required in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0047] Figure 1 Schematic diagram of collecting a low-resolution terahertz image dataset in an embodiment of the present invention;
[0048] Figure 2 Schematic diagram of a terahertz image degradation model in an embodiment of the present invention;
[0049] Figure 3 Schematic diagram of the overall architecture of a super-resolution network in an embodiment of the present invention;
[0050] Figure 4 Schematic diagram of the decomposed channel attention mechanism in the overall architecture of a super-resolution network in an embodiment of the present invention;
[0051] Figure 5 Schematic diagram of the network architecture of the multi-scale fusion module in the overall architecture of a super-resolution network in an embodiment of the present invention;
[0052] Figure 6 Schematic diagram of the large kernel self-attention structure in the network architecture of the multi-scale fusion module in an embodiment of the present invention;
[0053] Figure 7 Schematic diagram of the method steps of a terahertz image super-resolution reconstruction method based on a multi-dimensional attention fusion network in an embodiment of the present invention. Detailed implementation manners
[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0055] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0056] Embodiment 1
[0057] A terahertz image super-resolution reconstruction method based on a multi-dimensional attention fusion network, as Figure 7 shown, the method includes:
[0058] Step S1: Use a terahertz image acquisition device to collect terahertz images of the person to be detected, forming a terahertz dataset.
[0059] The terahertz image acquisition device includes a computer 103, a terahertz imaging device 102, a terahertz band transceiver, and the person to be detected carrying items; the computer 103 is connected to the terahertz imaging device 102 and at the same time connected to the terahertz transceiver to obtain a terahertz image dataset. In the dataset image, a high-resolution and clear-detail terahertz image 104 of the person to be detected 101 is obtained through an enhanced super-resolution generative adversarial network integrated with a new module, which can facilitate the detection of dangerous goods more clearly.
[0060] As Figure 1 shown, the terahertz imaging device 102 takes an image of the person to be detected 101 to obtain a terahertz image 104. On the computer 103, the obtained low-resolution terahertz images are combined into a dataset [Image1 Thz , Image2 Thz ,... ImageK Thz , where the total number of images in the dataset ImageK Thz is K, the size of the original image is H×W, H is the height of the original image, and W is the width of the original image.
[0061] Step S2: Use a degradation model to degrade the features of the terahertz dataset to generate a low-resolution image training dataset with a wide range of degradation features.
[0062] Construct a degradation model as Figure 2 shown, which consists of a blur kernel convolution module, a noise addition module, and a downsampling module, and simulates the degradation process of the terahertz imaging system through physical mechanism driving.
[0063] Taking a high-resolution terahertz image I HR with a size of H×W as the input, the input image first enters the blur kernel convolution module for spatial domain degradation processing. An anisotropic Gaussian kernel k is used to simulate the diffraction effect of the terahertz beam, and its mathematical expression is:
[0064]
[0065] where σ a are the standard deviations of the Gaussian kernel in the x direction respectively, σ b are the standard deviations of the Gaussian kernel in the y direction respectively, and (a, b) is the kernel coordinate. The degraded feature map is generated through convolution operation:
[0066]
[0067] where, Denotes a convolution operation with a kernel size of K×K. Experiments show that when K≥15, it can effectively simulate the characteristics of the point spread function of the terahertz system. Add composite noise N to the blurred feature map B s , and its expression is:
[0068] N s =α g N g +α t N t +α p N p ;
[0069] Among them, N g is Gaussian noise, N t is thermal noise, N p is Poisson noise, and α g , α t , α p are noise intensity coefficients. Use bicubic interpolation f s (·) to downsample the degraded feature map:
[0070] I LR =f s (B)+N s ;
[0071] Generate a low-resolution image training dataset with a wide range of degradation characteristics [Image1 Thz1 , Image1 Thz2 ,...Image1 Thzn ,...ImageK Thzn .
[0072] Step S3: Construct a multi-dimensional attention fusion network based on multi-scale fusion and channel attention mechanism, and use the low-resolution image training dataset with the wide range of degradation characteristics to train the multi-dimensional attention fusion network to obtain a trained multi-dimensional attention fusion network.
[0073] Build the MAFNet network model as shown in Figure 3 . The deep learning network model includes a shallow feature extraction module, a residual dense block group, a multi-scale fusion module (MSFM), a decomposed channel attention module (DCAM), and a super-resolution reconstruction module.
[0074] Use a low-resolution terahertz image I LR with a size of h×w as the input to enter the shallow feature extraction module of the network model, and successively pass through 1 3×3 convolutional layer, 4 residual dense blocks (RRDB), and 15 multi-scale fusion modules (MSFM) to output multi-level feature maps F 1 , F 2 , F3 , and their number of channels are 64, 128, and 256 respectively. By reducing the number of original RRDB modules from 6 to 4 and adopting a channel decomposition strategy, the number of network parameters is reduced by 37.5%. In the depth feature fusion stage of the network model, a multi-scale fusion module (MSFM) is designed to achieve cross-channel and cross-space information interaction between feature maps by decomposing large kernel convolution and cross-scale attention mechanisms.
[0075] The decomposed channel attention module is as Figure 4 shown. First, given the input feature κ ∈ R c×h×w use the convolutional layer f c (·) 3×3 to perform channel dimension expansion. Then, through a 3×3 convolutional layer and channel splitting, the feature map is split along the channel dimension. The split feature maps are processed by the point convolution f c (·) 1×1 , which contains a convolutional layer with a kernel size of 1×1 to enhance the learning of local features. This process can be expressed as:
[0076] κ(κ 1 , κ 2 ) = split(f c (f c (κ) 3×3 ) 1×1 );
[0077] where split(·) represents the feature splitting function for splitting the feature map along the channel dimension. κ 1 , κ 2 ∈ R c×h×w are the split feature maps. Subsequently, the obtained feature maps κ 1 , κ 2 are concatenated along the channel dimension. Then, the number of channels is adjusted through a convolutional layer with a kernel size of 1×1 f c (·) 1×1 . This process can be expressed as:
[0078] κ out1 = f c (Concate(κ 1 , κ 2 )) 1×1 ;
[0079] where Concate(·) is the feature information fusion function, and κ out1 ∈ R c×h×w is the output feature map of the convolutional layer. Taking the feature map κ out1 = [κ 1 , κ 2 , … κ CAs the input, first generate the channel statistic κ through global average pooling z ∈R C , where the C-th element of κ z can be obtained by the following formula:
[0080]
[0081] where κ C (i, j) is the value of the C-th feature κ C at the position (i, j), and F GP (·) represents the global average pooling function. Then through a simple gating unit, the obtained channel statistic is used to scale the input This process can be expressed as:
[0082]
[0083] where W d (·) represents the set of weights of the convolutional layer, which is used to reduce the channel scale. W u (·) represents the convolutional layer, which is used to increase the channel scale. represents the output feature map of the channel attention. Connect the input and output feature maps through a long skip connection. Add the feature maps element-wise to generate the shallow and middle-level feature κ out . This process can be expressed as:
[0084] κ out = κ + κ out2 ;
[0085] The multi-scale fusion module is as shown in Figure 5 Assume that the given input feature map κ out ∈R C×H×W . First, the input feature map κ out enters two convolutional blocks to obtain the feature map χ 1 . The convolutional block consists of convolution (Conv), batch normalization (BN), and ReLU. A short skip connection between the head and the tail enables the elements of the two feature maps to be added to obtain the feature map χ α . This process can be described as:
[0086] χ α = ConvBlock(κ out ) + ConvBlock(κ out ) + κ out ;
[0087] where ConvBlock(·) represents the convolutional block composed of Conv, BN, and ReLU. Secondly, as shown in Figure 6 the feature map χα Through the large kernel self-attention (LSA) structure to further extract multi-dimensional feature information. In LSA, the feature map χ α is split along the channel dimension into multiple small feature maps χ i1 , χ i2 , where i = 1, 2, 3 and the splitting ratio is 1:1. This process can be described as:
[0088] χ i1 , χ i2 = split(χ a );
[0089] Subsequently, these split feature maps χ i1 are respectively processed through the large kernel attention (LKA) function to obtain χ iΔ1 , that is, χ iΔ1 = Λ LKA (χ i1 ), and the specific operation of the LKA function is:
[0090] Λ LKA (·) f c [DW_D_Conv(DW_Conv(·))] 1×1 ;
[0091] which includes depthwise dilated convolution DW_Conv, depthwise convolution DW_D_Conv, and pointwise convolution f c (·) 1×1 . Then, χ iΔ1 is multiplied element-wise with the corresponding χ i2 to obtain the output feature map of LSA This process can be described as:
[0092]
[0093] Finally, the three are respectively adjusted in channel number through a 1×1 convolutional layer and added element-wise to obtain the final output feature map Out of MSFM. This process can be described as:
[0094]
[0095] The training method includes using the l 1 (Θ) loss function to guide the training of MAFNet. This process can be expressed as:
[0096]
[0097] In the above formula, F MAFNet(·) is a network function and Θ is a network parameter. The loss function is optimized using the stochastic gradient descent method. Set the network training parameters: learning rate lr, batchsize, division of the training set and validation set, optimizer, and number of training epochs.
[0098] Step S4: Input the real-time acquired terahertz image into the trained multi-dimensional attention fusion network to obtain a super-resolution reconstructed image of the terahertz image.
[0099] Use the trained network for prediction. Input the low-resolution (LR) terahertz image and output the super-resolution (SR) terahertz image. First, input the low-resolution image I to be measured LR into the network. The image size is H×W. After network inference, obtain the SR image I SR . The scale of the output feature map is H×W. MAFNet effectively extracts the detailed texture features of the terahertz image through the multi-scale fusion module (MSFM) and the decomposed channel attention module (DCAM), improving the accuracy and robustness of super-resolution reconstruction. At the same time, by adjusting the parameters of the upsampling module (PixelShuffle), it supports ×2, ×4, and ×8 super-resolution tasks.
[0100] Embodiment 2
[0101] A terahertz image super-resolution reconstruction system based on a multi-dimensional attention fusion network, the system includes:
[0102] A data acquisition module, which is used to collect terahertz images of the person to be detected using a terahertz image acquisition device to form a terahertz data set.
[0103] The terahertz image acquisition device includes a computer 103, a terahertz imaging device 102, a terahertz band transceiver, and a person to be detected carrying items; the computer 103 is connected to the terahertz imaging device 102 and is also connected to the terahertz transceiver to obtain a terahertz image data set. In the data set images, a high-resolution and detailed terahertz image 104 of the person to be detected 101 is obtained through an enhanced super-resolution generative adversarial network integrated with a new module, which can more clearly and conveniently perform dangerous goods detection.
[0104] Take a picture of the person to be detected 101 through the terahertz imaging device 102 to obtain the terahertz image 104. On the computer 103, form a data set [Image1 Thz , Image2 Thz ,...ImageK Thz from the obtained low-resolution terahertz images. Among them, the total number of images in the data set ImageK Thz is K, the original image size is H×W, H is the height of the original image, and W is the width of the original image.
[0105] A training dataset construction module, which is used to degrade the features of the terahertz dataset using a degradation model to generate a low-resolution image training dataset with a wide range of degradation features.
[0106] Construct as Figure 2 The shown degradation model, which consists of a blur kernel convolution module, a noise superposition module, and a downsampling module, and drives the simulation of the degradation process of the terahertz imaging system through physical mechanisms.
[0107] Using a high-resolution terahertz image I with dimensions H×W HR as the input, the input image first enters the blur kernel convolution module for spatial domain degradation processing. An anisotropic Gaussian kernel k is used to simulate the diffraction effect of the terahertz beam, and its mathematical expression is:
[0108]
[0109] where σ a are the standard deviations of the Gaussian kernel in the x direction respectively, and σ b are the standard deviations of the Gaussian kernel in the y direction respectively, and (a, b) are the kernel coordinates. The degraded feature map is generated through convolution operation:
[0110]
[0111] where represents the convolution operation, and the kernel size is K×K. Experiments show that when K≥15, the characteristics of the point spread function of the terahertz system can be effectively simulated. Composite noise N s is superimposed on the blurred feature map B, and its expression is:
[0112] N s =α g N g +α t N t +α p N p ;
[0113] where N g is Gaussian noise, N t is thermal noise, N p is Poisson noise, and α g , α t , α p are noise intensity coefficients. The degraded feature map is downsampled using bicubic interpolation f s (·):
[0114] I LR =f s (B)+N s ;
[0115] Generate a low-resolution image training dataset with extensive degradation characteristics [Image1 Thz1 , Image1 Thz2 ,...Image1 Thzn ,...ImageK Thzn .
[0116] A network construction module for constructing a multi-dimensional attention fusion network based on multi-scale fusion and channel attention mechanism, and training the multi-dimensional attention fusion network with the low-resolution image training dataset having extensive degradation characteristics to obtain a trained multi-dimensional attention fusion network.
[0117] Build the MAFNet network model as shown in Figure 3 . The deep learning network model includes a shallow feature extraction module, a residual dense block group, a multi-scale fusion module (MSFM), a decomposed channel attention module (DCAM), and a super-resolution reconstruction module.
[0118] Use the low-resolution terahertz image I with size h×w LR as the input to enter the shallow feature extraction module of the network model, and sequentially pass through 1 3×3 convolutional layer, 4 residual dense blocks (RRDB), and 15 multi-scale fusion modules (MSFM) to output multi-level feature maps F 1 , F 2 , F 3 , whose number of channels are 64, 128, and 256 respectively. By reducing the number of original RRDB modules from 6 to 4 and adopting a channel decomposition strategy, the number of network parameters is reduced by 37.5%. In the deep feature fusion stage of the network model, a multi-scale fusion module (MSFM) is designed to achieve cross-channel and cross-space information interaction between feature maps through decomposed large-kernel convolution and cross-scale attention mechanism.
[0119] The decomposed channel attention module is as shown in Figure 4 . First, given the input feature κ∈R c×h×w use the convolutional layer f c (·) 3×3 to perform channel dimension expansion. Then, through a 3×3 convolutional layer and channel splitting, the feature map is split along the channel dimension. The split feature maps are processed by the point convolution f c (·) 1×1 , and this point convolution contains a convolutional layer with a kernel size of 1×1 for enhancing the learning of local features. This process can be expressed as:
[0120] κ(κ 1 , κ 2 ) = split(f c (f c (κ)3×3 ) 1×1 );
[0121] Among them, split(·) represents the feature segmentation function, which is used to segment the feature map along the channel dimension. κ 1 , κ 2 ∈R c×h×w is the segmented feature map. Subsequently, the obtained feature maps κ 1 , κ 2 are concatenated along the channel dimension. Then, through a convolutional layer f c (·) 1×1 with a kernel size of 1×1 to adjust the number of channels. This process can be expressed as:
[0122] κ out1 = f c (Concate(κ 1 , κ 2 )) 1×1 ;
[0123] Among them, Concate(·) is the feature information fusion function, and κ out1 ∈R c×h×w is the output feature map of the convolutional layer. Taking the feature map κ out1 = [κ 1 , κ 2 , … κ C as the input, first generate the channel statistics κ z ∈R C through global average pooling. Among them, the C-th element of κ z can be obtained through the following formula:
[0124]
[0125] Among them, κ C (i, j) is the value of the C-th feature κ C at the position (i, j), and F GP (·) represents the global average pooling function. Then through a simple gating unit, the obtained channel statistics are used to scale the input This process can be expressed as:
[0126]
[0127] Among them, W d (·) represents the weight set of the convolutional layer, which is used to reduce the channel scale. W u (·) represents the convolutional layer, which is used to increase the channel scale. Represents the output feature map of channel attention. Connects the input and output feature maps through a long skip connection. Adds the feature maps element-wise to generate the shallow and middle-level feature κ out . This process can be expressed as:
[0128] κ out = κ + κ out2 ;
[0129] The multi-scale fusion module is as shown in Figure 5 . Assuming the given input feature map κ out ∈ R C×H×W . First, the input feature map κ out enters two convolutional blocks to obtain the feature map χ 1 . The convolutional block consists of convolution (Conv), batch normalization (BN), and ReLU. A short skip connection connects the head and tail to add the elements of the two feature maps to obtain the feature map χ α . This process can be described as:
[0130] χ α = ConvBlock(κ out ) + ConvBlock(κ out ) + κ out ;
[0131] where ConvBlock(·) represents the convolutional block composed of Conv, BN, and ReLU. Second, as shown in Figure 6 , the feature map χ α passes through the large kernel self-attention (LSA) structure to further extract multi-dimensional feature information. In LSA, the feature map χ α is split into multiple small feature maps χ i1 , χ i2 along the channel dimension, where i = 1, 2, 3 and the splitting ratio is 1:1. This process can be described as:
[0132] χ i1 , χ i2 = split(χ a );
[0133] Subsequently, these split feature maps χ i1 are respectively processed through the large kernel attention (LKA) function to obtain χ iΔ1 , that is, χ iΔ1 = Λ LKA (χ i1 ). The specific operation of the LKA function is:
[0134] Λ LKA (·) = f c [DW_D_Conv(DW_Conv(·))]1×1 ;
[0135] It includes depthwise dilated convolution DW_Conv, depthwise convolution DW_D_Conv and pointwise convolution f c (·) 1×1 Then, χ iΔ1 is multiplied element-wise with the corresponding χ i2 to obtain the output feature map of LSA This process can be described as:
[0136]
[0137] Finally, the three are respectively passed through a 1×1 convolutional layer to adjust the number of channels and then added element-wise to obtain the final output feature map Out of MSFM. This process can be described as:
[0138]
[0139] The training method includes using the l 1 (Θ) loss function to guide the training of MAFNet. This process can be expressed as:
[0140]
[0141] In the above formula, F MAFNet (·) is the network function and Θ is the network parameter. The loss function is optimized using the stochastic gradient descent method. Set the network training parameters: learning rate lr, batchsize, division of training set and validation set, optimizer and number of training epochs.
[0142] The image reconstruction module is used to input the real-time acquired terahertz image into the trained multi-dimensional attention fusion network to obtain the super-resolution reconstructed image of the terahertz image.
[0143] Use the trained network for prediction. Input the low-resolution (LR) terahertz image and output the super-resolution (SR) terahertz image. First, input the low-resolution image I to be measured LR into the network. The image size is H×W. After network inference, obtain the SR image I SR . The scale of the output feature map is H×W. MAFNet effectively extracts the detailed texture features of the terahertz image through the multi-scale fusion module (MSFM) and the decomposed channel attention module (DCAM), improving the accuracy and robustness of super-resolution reconstruction. At the same time, by adjusting the parameters of the upsampling module (PixelShuffle), it supports ×2, ×4, ×8 super-resolution tasks.
[0144] The embodiments described above are only descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A terahertz image super-resolution reconstruction method based on a multi-dimensional attention fusion network, characterized in that: The method comprises: Step S1, using a terahertz image acquisition device to acquire a terahertz image of the detected person to form a terahertz data set; Step S2: using a degradation model to perform feature degradation on the terahertz data set to generate a low-resolution image training data set with extensive degradation features; Step S3, constructing a multidimensional attention fusion network based on multi-scale fusion and channel attention mechanism, and using the low-resolution image training data set with extensive degradation characteristics to train the multidimensional attention fusion network to obtain a trained multidimensional attention fusion network; Step S4: input the real-time collected terahertz image into the trained multi-dimensional attention fusion network to obtain a terahertz image super-resolution reconstructed image.
2. According to claim 1, the terahertz image super-resolution reconstruction method based on multi-dimensional attention fusion network is characterized in that: The content of step S2 specifically includes: The terahertz image I in the terahertz data set collected in step S1 HR The spatial domain degradation processing is performed in the input fuzzy convolution to obtain the spatial domain degraded image; The spatial domain degraded image and the terahertz image I HR Perform convolution to obtain a blurred feature map; Superimposing composite noise on the fuzzy feature map to obtain a noise fuzzy feature map; Downsampling the fuzzy feature map using a bicubic interpolation method to obtain a sampled fuzzy feature map; Adding the noise blur feature map to the sample blur feature map to obtain a low-resolution image with extensive degradation features; All terahertz images in the terahertz data set are processed into low-resolution images to obtain a low-resolution image training data set with extensive degradation characteristics.
3. The terahertz image super-resolution reconstruction method based on a multi-dimensional attention fusion network according to claim 2 is characterized in that: The content of the spatial domain degraded image is specifically: Among them, σ a are the Gaussian kernel standard deviation controlling the x direction, σ b are the standard deviations of the Gaussian kernel controlling the y direction, and (a, b) are the kernel coordinates.
4. The terahertz image super-resolution reconstruction method based on a multi-dimensional attention fusion network according to claim 1, characterized in that: In step S3, the specific contents of the multi-dimensional attention fusion network include: The low-resolution terahertz image I in the low-resolution image training data set LR Perform shallow feature extraction to obtain a shallow feature map; After the shallow feature map is expanded in channel dimension by using a convolution layer, convolution and channel segmentation are performed to obtain segmented feature maps κ1 and κ2; Using point convolution to process the segmented feature maps κ1 and κ2 to obtain a local enhanced feature map; The local enhanced feature maps are spliced in the channel dimension and the number of channels is adjusted to obtain an output feature map; Obtain shallow-middle-layer features based on the output feature map; Based on the shallow and middle layer features, the multi-scale fusion block is input to obtain a multi-scale fusion output image.
5. The terahertz image super-resolution reconstruction method based on a multi-dimensional attention fusion network according to claim 4 is characterized in that: The process of obtaining the local enhanced feature map is specifically as follows: κ(κ1,κ2)=split(f c (f c (k) 3×3 ) 1×1 ); Among them, split(·) is the feature segmentation function, which is used to split the feature map according to the channel dimension; κ1,κ2∈R c×h×w is the feature map after segmentation; f c (·) 1×1 is a convolutional layer with a kernel size of 1×1.
6. The terahertz image super-resolution reconstruction method based on a multi-dimensional attention fusion network according to claim 5 is characterized in that: The process of obtaining the shallow-middle layer features is specifically as follows: Generate channel statistics by global average pooling of the output feature map, and obtain the channel attention output feature map of any element in the channel statistics by a gating unit; Through the long jump connection output feature map and the channel attention output feature map, the shallow middle-level features of this element are obtained; Add the feature maps element by element to get the shallow mid-level features.
7. The terahertz image super-resolution reconstruction method based on a multi-dimensional attention fusion network according to claim 6 is characterized in that: The content of the multi-scale fusion block specifically includes: Input the shallow middle layer features into two convolution blocks to obtain a feature map; The feature map is passed through a large kernel self-attention structure to extract multi-dimensional feature information χ i1 , χ i2 ; The multi-dimensional feature information x i1 The output of LKA is obtained by using a large kernel attention function. iΔ1 ; The output result χ iΔ1 With multi-dimensional feature information χ i2 Multiply element by element to get the output feature map of LSA; The output feature map of the LSA is convolved to adjust the number of channels and then added element by element to obtain the output result of the scale fusion block.
8. A terahertz image super-resolution reconstruction system based on a multi-dimensional attention fusion network, the system being used to implement the terahertz image super-resolution reconstruction method according to any one of claims 1 to 7, characterized in that the system include: A data acquisition module, used for using a terahertz image acquisition device to acquire terahertz images of the detected person to form a terahertz data set; A training data set construction module, used to perform feature degradation on the terahertz data set using a degradation model to generate a low-resolution image training data set with a wide range of degradation features; A network construction module, used for constructing a multidimensional attention fusion network based on multi-scale fusion and channel attention mechanism, and training the multidimensional attention fusion network using the low-resolution image training data set with extensive degradation characteristics to obtain a trained multidimensional attention fusion network; The image reconstruction module is used to input the real-time collected terahertz image into the trained multi-dimensional attention fusion network to obtain a terahertz image super-resolution reconstructed image.
Citation Information
Patent Citations
Super-resolution reconstruction method, device, equipment and readable storage medium
CN108665509A
Image super-resolution reconstruction method based on fused attention mechanism residual network
CN111192200A
Classification method based on deep fusion of hyperspectrum and terahertz data
CN111539447A
Image super-resolution reconstruction method based on back projection attention network
CN112215755A
Terahertz image super-resolution reconstruction method based on generative adversarial network
CN115358922A
Cited By
Unsupervised parallax tolerant terahertz image splicing method and system
CN120563317A
Terahertz polyethylene pipeline hot melting joint defect super-resolution imaging method based on CBAM-ESRGAN
CN121169689A