A method for constructing an attention mechanism network

By building an attention mechanism network, optimizing the parameters of the sampling network and reconstructing network, the problems of high imaging complexity and complex signal reconstruction algorithms in traditional technologies are solved, and the low-complexity and high-precision signal reconstruction effect is achieved.

CN115438769BActive Publication Date: 2025-06-10SUZHOU JIAOSHI INTELLIGENCE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110627516.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-04
Publication Date
2025-06-10
Estimated Expiration
2041-06-04

AI Technical Summary

Technical Problem

When traditional correlation imaging technology imaging in complex and changing environments, there are problems such as many sampling times, long imaging time, and complex system structure. In addition, compression sensing technology has high algorithm complexity, long reconstruction time, and many iterations in signal reconstruction.

Method used

Build an attention mechanism network, including sampling networks and reconstruction networks, and optimize network parameters to achieve efficient signal compression and reconstruction through technical means such as data amplification, convolutional layers and residual blocks.

Benefits of technology

Signal reconstruction with low computing complexity and high reconstruction accuracy is realized, which can better preserve image structure information and achieve better reconstruction effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115438769B_ABST
    Figure CN115438769B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for constructing an attention mechanism network. First, sample data is expanded through data augmentation technology to form training samples, and the obtained training samples are divided into a training set, a validation set, and a test set. Secondly, the training images are compressed by a sampling network to obtain compressed signals, and the compressed signals are reconstructed by a reconstruction network to obtain a reconstruction result. Then, the error between the reconstruction result and the training images is calculated, and based on this error, an optimizer is used to train the sampling network and the reconstruction network, and the trained sampling network and reconstruction network are used to reconstruct the test set, and the peak signal-to-noise ratio of signal reconstruction is statistically calculated. Finally, the sampling network and the reconstruction network are optimized according to the statistically obtained peak signal-to-noise ratio to obtain the final sampling network and reconstruction network. The sampling network obtained through continuous training and optimization can retain more image structure information and achieve better signal reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of signal reconstruction, and particularly to a method for constructing an attention mechanism network for reconstructing compressive sensing signals. Background Art

[0002] Correlated imaging, also known as ghost imaging, is a new imaging technology that can non-locally obtain target image information based on the quantum or classical correlation characteristics of optical field fluctuations. By performing intensity correlation operations between the reference optical field and the target detection optical field, traditional correlated imaging has problems such as a large number of sampling times, long imaging time, and complex system structure, and is not suitable for imaging in complex and changeable environments such as water bodies.

[0003] Compressive sensing is an emerging technology applied to signal sampling and processing. It makes full use of two important properties of signals: sparsity and compressibility, and can effectively reduce the number of samplings and storage. Therefore, compressive sensing technology has replaced traditional signal compression methods and has been widely applied in various fields. However, in terms of signal reconstruction, traditional compressive sensing technology still faces problems such as high algorithm complexity, long reconstruction time, and many iterations. Summary of the Invention

[0004] The present invention provides a method for constructing an attention mechanism network to solve the problems existing in the prior art. The network model has a simple structure, few parameters, low computational complexity, and better reconstruction accuracy.

[0005] To solve the above technical problems, the technical solution of the present invention is: A method for constructing an attention mechanism network, the attention mechanism network includes a sampling network and a reconstruction network connected in sequence for reconstructing compressive sensing signals, including the following steps:

[0006] S1: Expand the sample data through data augmentation technology to form training samples, and divide the obtained training samples into a training set, a validation set, and a test set;

[0007] S2: Construct a sampling network, and compress the training images in the training set through the sampling network to obtain compressed signals;

[0008] S3: Construct a reconstruction network, and reconstruct the compressed signal through the reconstruction network to obtain a reconstruction result;

[0009] S4: Calculate the error between the reconstruction result and the training image;

[0010] S5: Use the optimizer to train the sampling network and the reconstruction network according to the calculated error, and use the trained sampling network and reconstruction network to reconstruct the images in the test set, and calculate the peak signal-to-noise ratio of the signal reconstruction results;

[0011] S6: According to the peak signal-to-noise ratio obtained by statistics, adjust the training parameters in step S5 to optimize the sampling network and the reconstruction network, and obtain the final sampling network and reconstruction network.

[0012] Further, the sampling network includes a first convolutional layer, a second convolutional layer, a deconvolutional layer, and a summation layer. The second convolutional layer and the deconvolutional layer are connected to the first convolutional layer, and the summation layer is connected to the deconvolutional layer and the second convolutional layer.

[0013] Further, the sampling network includes a first convolutional layer and a second convolutional layer connected in sequence.

[0014] Further, the reconstruction network includes a plurality of residual blocks. Each residual block includes a third convolutional layer, a fourth convolutional layer, and an attention mechanism module connected in sequence. The attention mechanism module includes a global average pooling layer and a fully connected layer connected in sequence.

[0015] Further, the reconstruction network outputs a feature map, and this feature map is additively fused with the output result of the summation layer of the sampling network to obtain the reconstruction result of the attention mechanism network.

[0016] Further, the reconstruction network outputs a feature map, and this feature map is additively fused with the output result of the second convolutional layer of the sampling network to obtain the reconstruction result of the attention mechanism network.

[0017] Further, in step S4, the following function is used to calculate the error between the reconstruction result and the training image:

[0018]

[0019] where N represents the number of pixels of the training image, and X and Y represent the training image input to the sampling network and the reconstruction result output by the reconstruction network respectively.

[0020] Further, in step S4, the following formula is used to calculate the error between the reconstruction result and the training image:

[0021]

[0022]

[0023] where N represents the number of pixels of the training image, E i represents the boundary energy at each pixel position in the image, and d iRepresents the Euclidean distance from the i-th pixel to the boundary.

[0024] Further, in the step S4, the error between the reconstruction result and the training image is calculated using the following formula:

[0025]

[0026] Where N represents that the training image contains N pixels, and X and Y respectively represent the training image input to the sampling network and the reconstruction result output by the reconstruction network.

[0027] Further, in the step S4, the error between the reconstruction result and the training image is calculated using the following formula:

[0028]

[0029] V(Y) = ∑ i,j Y i+1,j -Y i,j | + |Y i,j+1 -Y i,j |

[0030] Where N represents that the training image contains N pixels, X and Y respectively represent the training image input to the sampling network and the reconstruction result output by the reconstruction network, and λ is a hyperparameter.

[0031] The method for constructing an attention mechanism network provided by the present invention, the constructed attention mechanism network includes two parts: a sampling network and a reconstruction network. The present invention continuously trains and optimizes the sampling network and the reconstruction network to obtain the final sampling network and reconstruction network, so that the final sampling network can autonomously learn the measurement matrix to compress the signal. This method can not only retain more image structure information, but also achieve better reconstruction. In addition, the network model proposed by the present invention has a simple structure, few parameters, low computational complexity and better reconstruction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is a flowchart of the method for constructing an attention mechanism network in a specific embodiment of the present invention;

[0033] Figure 2 is a schematic diagram of the frameworks of the sampling network and the reconstruction network in a specific embodiment of the present invention;

[0034] Figure 3 is a schematic diagram of the framework of the reconstruction network in a specific embodiment of the present invention.

[0035] As shown in the figure: 10, sampling network; 11, first convolutional layer; 12, second convolutional layer; 13, deconvolutional layer; 14, summation layer; 20, reconstruction network; 21, third convolutional layer; 22, fourth convolutional layer; 23, global average pooling layer; 24, fully connected layer. Detailed implementation

[0036] The present invention will be described in detail below with reference to the accompanying drawings:

[0037] As Figure 1 shown, the present invention provides a method for constructing an attention mechanism network. The attention mechanism network includes a sampling network 10 and a reconstruction network 20 connected in sequence, and is used for the reconstruction of compressed sensing signals. The method includes the following steps. It should be noted that the initially constructed sampling network is the initial sampling network, and the final sampling network is obtained after training and optimization. The initially constructed reconstruction network is the initial reconstruction network, and the final reconstruction network is obtained after training and optimization.

[0038] S1: The sample data is expanded through data augmentation technology to form training samples, and the obtained training samples are divided into a training set, a validation set, and a test set. In this embodiment, the training samples are divided into a training set, a validation set, and a test set according to 8:1:1. In this embodiment, data augmentation technologies such as rotation, translation, and flipping can be used to expand the sample data to form training samples.

[0039] S2: Construct the sampling network 10, and compress the training images in the training set through the sampling network 10 to obtain compressed signals. As Figure 2 shown, in this embodiment, the sampling network 10 includes a first convolutional layer 11, a second convolutional layer 12, a deconvolutional layer 13, and a summation layer 14. The deconvolutional layer 14 is connected to the first convolutional layer 11, and the summation layer 14 is connected to the second convolutional layer 12 and the deconvolutional layer 14. The first convolutional layer 11 uses a convolutional kernel of size k*k and compresses the image signal according to a stride of k. Here, k is an integer and can be divisible by the width and height of the input image, that is, divisible by the width and height of the training image. For example, k can be 8, so that the convolutional parameters can be reduced and more image details can be retained. Of course, k can also be other values. It should be noted that the larger the convolutional kernel, the more parameters the network needs to be trained and optimized. The number of output feature channels of the first convolutional layer is selected according to the sampling rate. For an image signal containing N pixels, if a compression rate of 0.1 is used, the feature map obtained after the first layer of convolution contains M≈0.1*N pixels, where M must be an integer. The second convolutional layer 12 uses a convolutional kernel of size 1*1. After the training image passes through the second convolutional layer 12, k*k feature maps will be obtained, and the number of pixels contained in the feature maps is exactly equal to the number of pixels N of the input image. Therefore, the feature maps are converted into the image resolution size here.

[0040] The deconvolution layer 14 is connected after the first convolution layer 11 and is used to sample the result of the first convolution layer 11. It uses a convolution kernel of size k*k and a stride of size k to generate one feature map, and the dimension of this feature map is the same as that of the input image.

[0041] The summation layer 14 is located at the end of the sampling network 10 and is connected to the second convolution layer 12 and the deconvolution layer 13. It performs weighted summation on the results of the second convolution layer 12 and the deconvolution layer 13 to generate one feature map, which is the compressed signal output by the sampling network 10.

[0042] S3: Construct a reconstruction network 20 to reconstruct the compressed signal through the reconstruction network 20 to obtain a reconstruction result; in this embodiment, the reconstruction network 20 is connected after the entire sampling network 10 to reconstruct the sampling result. In this embodiment, the reconstruction network 20 is composed of multiple residual blocks. In this embodiment, the reconstruction network 20 includes 5 residual blocks, and each residual block includes a third convolution layer 21, a fourth convolution layer 22, and an attention mechanism module. The attention mechanism module includes a global average pooling layer 23 and a fully connected layer 24 connected in sequence.

[0043] In this embodiment, both the third convolution layer and the fourth convolution layer of each residual block use a convolution kernel of size 3*3, with a stride of 3 and a padding of 1. Each convolution layer generates 64 feature maps, and a ReLU activation function is connected after each convolution layer.

[0044] The attention mechanism module in each residual block is connected after the fourth convolution layer. First, the result of the fourth convolution layer is used by the global average pooling layer to obtain a 64-dimensional vector. A fully connected layer is connected after the global average pooling layer, and the number of neurons in the fully connected layer is 64, which is responsible for learning the spatial relationship between the 64 feature channels output by the convolution layer. The result of the fully connected layer is first transformed through a Sigmoid function layer into a probability between 0 and 1. The Sigmoid function is expressed as follows:

[0045]

[0046] x and y represent the input and output respectively. Then, the result obtained by the Sigmoid function is multiplied channel by channel with the feature map obtained by the fourth convolution layer.

[0047] The fourth convolution layer of the reconstruction network 20 uses a convolution of size 3*3 to output one feature map, and this feature map is additively fused with the output result of the summation layer 14 at the end of the sampling network 10, that is, the compressed signal, as the reconstruction result of the attention mechanism network.

[0048] Of course, the sampling network 10 may also only include the above-mentioned first convolutional layer 11 and second convolutional layer 12. Finally, the feature map output by the fourth convolutional layer 22 of the reconstruction network 20 is fused with the output result of the last summation layer of the sampling network 10, that is, the compressed signal, through addition as the reconstruction result of the attention mechanism network.

[0049] S4: Calculate the error between the reconstruction result and the training image;

[0050] Assume that the training image X is compressed into an M-dimensional tensor by a convolutional block of size k*k through the sampling network 10 This process can be mathematically expressed as where C is the measurement matrix with a dimension of M×N, and N is the number of pixels of the input image. After passing through the reconstruction network 20, the tensor will be reconstructed into Y, which is an N-dimensional tensor representing the reconstructed result of the image compression signal and can be expressed as Calculate the loss between the output result Y of the reconstruction network and the input X of the network. The following methods can be adopted:

[0051] (1) The loss function can use the mean squared error (MSE) function, which can be expressed as:

[0052]

[0053] where N represents the number of pixels of the image, and X and Y represent the training image input by the sampling network and the reconstructed result output by the reconstruction network, respectively.

[0054] (2) The loss function can add the boundary energy weight to the mean squared error function, and the specific calculation formula is expressed as:

[0055]

[0056]

[0057] where N represents the number of pixels of the training image, E i represents the boundary energy at each pixel position in the image, d i represents the Euclidean distance from the i-th pixel to the boundary, which is mainly calculated through the boundary operator. The boundary energy calculated for each image will be sent into the network together with the image to participate in the calculation of the loss function. The purpose of using this loss function is to punish the loss value near the boundary and make the reconstructed result better retain the boundary information.

[0058] (3) The loss function can add the total variation of the reconstructed image to the mean squared error function, and the specific calculation formula is expressed as:

[0059]

[0060] V(Y) = ∑ i,j |Y i+1,j -Y i,j | + |Y i,j+1 -Y i,j | (5)

[0061] Add the total variation (TV) of the reconstructed image to the loss function to remove the noise in the reconstructed image and make the reconstructed image smoother. The image is a discrete two-dimensional signal. Equation (4) is the MSE loss between the reconstructed image and the training image, and Equation (5) is the TV loss of the reconstructed image. Here, λ is a hyperparameter that needs to be set manually and is used to control the ratio between the two losses.

[0062] (4) Add the MSE loss between the reconstructed image and the gradients of the input image to the loss function. Since there are two directions for the image gradient, namely the x-direction and the y-direction, the specific calculation formula for the gradient loss is expressed as:

[0063]

[0064] The first term of this formula (6) is the MSE loss between the reconstructed image and the input image, the second term is the MSE loss between the reconstructed image and the x-direction gradient of the input image, and the third term is the MSE loss between the reconstructed image and the y-direction gradient of the input image.

[0065] (5) Add the L1 loss function to the loss function. At the same time, during the training stage, use the strategy of alternating training with the MSE loss and the L1 loss. For example, use the MSE loss for the first 50,000 iterations and then switch to the L1 loss for the subsequent iterations to continue training. The main purpose is to enable the network to better converge to the minimum value and achieve a better reconstruction effect. The expression of the L1 loss is:

[0066]

[0067] where N represents the number of pixels in the image, and X and Y respectively represent the training image input to the sampling network and the reconstruction result output by the reconstruction network.

[0068] S5: Train the sampling network and the reconstruction network using an optimizer according to the calculated error, and use the trained sampling network and reconstruction network to reconstruct the test set, and statistically calculate the peak signal-to-noise ratio of the signal reconstruction result; it should be noted that the purpose of training the sampling network is to learn the measurement matrix and then use the measurement matrix to perform compressive sampling on the image. In this embodiment, the Adam optimizer can be used for training, and the parameters are β 1 = 0.9, β 2= 0.99; Other hyperparameter settings during training are as follows: the initial learning rate is set to 0.001, the total number of iterations is 500,000 times, and the learning rate decay rule is: the learning rate is reduced to 0.1 of the current learning rate at iteration numbers 250,000, 400,000, and 500,000 respectively, and the weight decay rate is 0.0005. During the training process of the network, the weight parameters need to be iteratively updated, and the update of these weight parameters requires the loss calculated in step S4. The specific method is: calculate the partial derivative of the loss with respect to each layer's parameters in the network respectively, and then update the weight parameter values of each layer according to the set learning rate. In addition, during the training process of the network, overfitting or non-convergence states may occur. The validation set is used to judge / verify a certain convergence state reached by the network during the training process, so as to determine whether to fine-tune the parameters and retrain.

[0069] S6: According to the statistically obtained peak signal-to-noise ratio, adjust the training parameters, optimize the sampling network and the reconstruction network, and obtain the final sampling network and reconstruction network.

[0070] The method for constructing an attention mechanism network provided by the present invention, the constructed attention mechanism network includes two parts: a sampling network and a reconstruction network. The present invention continuously trains and optimizes the sampling network and the reconstruction network to obtain the final sampling network and reconstruction network, so that the final sampling network can autonomously learn the measurement matrix to compress the signal. This method can not only retain more image structure information, but also achieve better reconstruction. In addition, the network model structure proposed by the present invention is simple, has fewer parameters, has a lower computational complexity, and has better reconstruction accuracy.

[0071] Although the embodiments of the present invention are described in the specification, these embodiments are only for reference and should not limit the protection scope of the present invention. All omissions, substitutions, and changes made within the scope of the purpose of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for constructing an attention mechanism network, the attention mechanism network including a sampling network and a reconstruction network connected in sequence for reconstructing compressed sensing signals, comprising the following steps: S1: Expand the sample data through data augmentation technology to form training samples, and divide the obtained training samples into a training set, a validation set, and a test set; S2: Construct a sampling network, and compress the training images in the training set through the sampling network to obtain compressed signals; S3: Construct a reconstruction network, and reconstruct the compressed signals through the reconstruction network to obtain a reconstruction result; S4: Calculate the error between the reconstruction result and the training images; S5: Use an optimizer to train the sampling network and the reconstruction network according to the calculated error, and use the trained sampling network and reconstruction network to reconstruct the images in the test set, and statistically calculate the peak signal-to-noise ratio of the signal reconstruction result; S6: According to the statistically obtained peak signal-to-noise ratio, adjust the training parameters in step S5 to optimize the sampling network and the reconstruction network to obtain the final sampling network and reconstruction network.

2. The method for constructing an attention mechanism network according to claim 1, characterized in that the sampling network includes a first convolutional layer, a second convolutional layer, a deconvolutional layer, and a summation layer, the second convolutional layer and the deconvolutional layer are connected to the first convolutional layer, and the summation layer is connected to the deconvolutional layer and the second convolutional layer.

3. The method for constructing an attention mechanism network according to claim 1, characterized in that the sampling network includes a first convolutional layer and a second convolutional layer connected in sequence.

4. The method for constructing an attention mechanism network according to claim 1, characterized in that the reconstruction network includes a plurality of residual blocks, each residual block including a third convolutional layer, a fourth convolutional layer, and an attention mechanism module connected in sequence, and the attention mechanism module includes a global average pooling layer and a fully connected layer connected in sequence.

5. The method for constructing an attention mechanism network according to claim 1, characterized in that the reconstruction network outputs a feature map, and this feature map is additively fused with the output result of the summation layer of the sampling network to obtain the reconstruction result of the attention mechanism network.

6. The method for constructing an attention mechanism network according to claim 1, characterized in that the reconstruction network outputs a feature map, and this feature map is additively fused with the output result of the second convolutional layer of the sampling network to obtain the reconstruction result of the attention mechanism network.

7. The method for constructing an attention mechanism network according to claim 1, characterized in that in step S4, the following function is used to calculate the error between the reconstruction result and the training images: where N represents the number of pixels of the training images, and X and Y respectively represent the training images input to the sampling network and the reconstruction result output by the reconstruction network.

8. The method for constructing an attention mechanism network according to claim 1, characterized in that in step S4, the following formula is used to calculate the error between the reconstruction result and the training images: where N represents the number of pixels of the training image, and E i represents the boundary energy at each pixel position in the image, and d i represents the Euclidean distance from the i-th pixel to the boundary.

9. The method for constructing an attention mechanism network according to claim 1, characterized in that In the step S4, the error between the reconstruction result and the training image is calculated using the following formula: where N represents that the training image contains N pixels, and X and Y respectively represent the training image input by the sampling network and the reconstruction result output by the reconstruction network.

10. The method for constructing an attention mechanism network according to claim 1, characterized in that in the step S4, the error between the reconstruction result and the training image is calculated using the following formula: V(Y) = ∑ i,j |Y i+1,j - Y i,j | + |Y i,j+1 - Y i,j | where N represents that the training image contains N pixels, X and Y respectively represent the training image input by the sampling network and the reconstruction result output by the reconstruction network, and λ is a hyperparameter.

Citation Information

Patent Citations

  • Image enhancement method and device and terminal equipment

    CN111047512A

  • Compressed sensing sampling reconstruction method and system based on linear sampling network and generative adversarial residual network

    CN112116601A