A multi-scale image compression and reconstruction method combined with attention mechanism
Through the multi-scale image compression reconstruction method combined with attention mechanism, the problems of poor reconstruction performance and too long reconstruction time in the prior art are solved, and a more efficient image reconstruction effect is achieved.
Patent Information
- Application Number
- CN202210688641.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-06-17
AI Technical Summary
The existing compression perception methods and deep learning-based compression perception reconstruction methods have problems such as poor reconstruction performance and too long reconstruction time, which leads to inability to effectively apply in actual scenarios.
Using a multi-scale image compression reconstruction method combined with attention mechanism, the measured values are initially reconstructed, multi-scale reconstruction and feature fusion, and the measured value residuals are further used to compensate for reconstruction to comprehensively obtain the reconstructed image.
Through multi-step processing of feature extraction and information utilization, the quality and efficiency of image reconstruction are significantly improved and better reconstruction effects are obtained.
Smart Images

Figure CN114926557B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image compression sensing, and more specifically, to a multi-scale image compression and reconstruction method combined with an attention mechanism. Background Art
[0002] The famous Nyquist Sampling Theory (NST) points out that in order for the sampled digital signal to completely retain the information in the original signal, the sampling frequency must be greater than twice the highest frequency in the signal. Because if the sampling frequency is lower than twice the highest frequency of the signal, the signal will be aliased after the frequency domain spectrum is moved. However, in some actual signal and image processing applications, sampling the signal in accordance with the Nyquist sampling theorem will lead to problems such as too high sampling frequency, too much sampled data that is not conducive to storage and transmission, and low sampling rate. In addition, oversampling caused by too high sampling frequency often damages the equipment. With the birth of Compressed Sensing (CS) theory, this problem has been solved.
[0003] Compressed sensing, as a novel sampling theory, can effectively utilize the sparse characteristics of signals to save the cost of data acquisition, transmission and storage, and can therefore be used in some fields such as wireless communications and signal acquisition. Compressed sensing is a theory that breaks through the Nyquist sampling theorem to compress the original signal sampling into a low-dimensional observation signal. With only a very small number of sampling points, the original signal can be accurately restored through a suitable reconstruction method. Currently, the commonly used reconstruction algorithms in compressed sensing include Basis Pursuit (BP) and Orthogonal Matching Pursuit (OMP). However, they are all based on iteration or convex relaxation, and it is difficult to have a good reconstruction performance when the sampling rate is low.
[0004] Therefore, although compressed sensing is basically perfected and has been applied to some fields, the existing compressed sensing methods and compressed sensing reconstruction methods based on deep learning still have shortcomings such as poor reconstruction performance and long reconstruction time, which makes it difficult to apply it well in actual scenarios. Summary of the invention
[0005] In order to overcome the technical defect of poor reconstruction performance of existing compressed sensing methods, the present invention provides a multi-scale image compression and reconstruction method combined with an attention mechanism.
[0006] In order to solve the above technical problems, the technical solution of the present invention is as follows:
[0007] A multi-scale image compression and reconstruction method combined with an attention mechanism comprises the following steps:
[0008] S1: Obtain the image to be reconstructed and build a multi-scale reconstruction model;
[0009] The multi-scale reconstruction model includes a sampling module, an initial reconstruction module and an enhanced reconstruction module;
[0010] S2: In the sampling module, convolution sampling is performed on the image to be reconstructed to obtain a measurement value, and the measurement value is adjusted;
[0011] S3: Performing initial reconstruction on the adjusted measurement value in the initial reconstruction module to obtain an initial reconstruction value and calculating the measurement value residual;
[0012] S4: in the enhanced reconstruction module, the initial reconstruction value is reconstructed at multiple scales to obtain an enhanced reconstruction value, and a compensated reconstruction value is calculated according to the measurement value residual;
[0013] S5: Obtain a reconstructed image according to the initial reconstruction value, the enhanced reconstruction value and the compensated reconstruction value.
[0014] In the above scheme, the measurement values are initially reconstructed, and then the features of different dimensions are extracted respectively through multi-scale reconstruction and then feature fusion is performed, so as to better extract the features of the image and reduce the feature loss; and in the reconstruction, the measurement value residual is further calculated and reconstructed separately, so as to make fuller use of the measurement value information and reduce the measurement value loss; finally, the initial reconstruction value, the enhanced reconstruction value and the compensated reconstruction value are combined to obtain the reconstructed image and obtain a better reconstruction effect.
[0015] Preferably, before performing convolution sampling on the image to be reconstructed, the method further includes dividing the image to be reconstructed into a plurality of image blocks with a pixel size of 33×33;
[0016] The convolution sampling is specifically:
[0017] y=S(x)=W M *x=Wx
[0018] Among them, y represents the measured value obtained by sampling, S() represents the mapping process of sampling, x represents the input image, and W M Represents the convolution kernel group, m = r × N 2 represents the number of convolution kernels, r represents the sampling rate, and N represents the image block size;
[0019] W represents the network parameter matrix composed of convolution kernel groups, specifically:
[0020]
[0021]
[0022] w1,w2,w3,w mThey respectively represent the first convolution kernel, the second convolution kernel, the third convolution kernel, and the mth convolution kernel in the network parameter matrix, and R represents a real number.
[0023] In the above scheme, the sampling module is composed of a convolution layer with a convolution kernel size of 33×33, which replaces the traditional measurement matrix for sampling. The original input is convolved with m convolution kernels respectively, and the m outputs obtained constitute the required measurement value y, which is equivalent to mapping the measurement matrix to the convolution layer parameters. During the network training process, along with the update of the convolution layer parameters, the measurement matrix is also updated at the same time, making the sampling adaptive.
[0024] Preferably, in step S2, the measurement value adjustment comprises the following steps:
[0025] S2.1: Perform global average pooling on the measured values:
[0026] y1=avgpolling(y)
[0027] S2.2: Measurement information exchange through one-dimensional convolution:
[0028] y2=conv1d(y1)
[0029] S2.3: Use sigmoid function for weight distribution:
[0030] β=sigmoid(y2)
[0031] S2.4: Perform normalization:
[0032] β1=softmax(β)
[0033] S2.5: Adjust the measured value:
[0034]
[0035] Among them, y1 represents the measured value after global average pooling, y2 represents the value of y1 after one-dimensional convolution, β represents the weight, β1 represents the normalized weight, and y f Represents the adjusted measurement value.
[0036] Preferably, the initial reconstruction is expressed as:
[0037] x1=IR(y f )=W int *y f
[0038] Among them, x1 is the initial reconstruction value, IR() is the initial reconstruction function, W int is the convolutional layer used for initial reconstruction.
[0039] Preferably, the measurement residual is calculated by the following formula:
[0040] △y=y f -W*x1.
[0041] Preferably, the initial reconstruction value is reconstructed at multiple scales by the following steps to obtain an enhanced reconstruction value:
[0042] S4.1: Extract features of different dimensions from the initial reconstruction value by strengthening the upper channel and the lower channel of the reconstruction module, where:
[0043] The upper channel is used to extract low-dimensional receptive field channel features, which specifically includes the following steps:
[0044] S4.1.1.1: Extract preliminary features of the upper channel through convolution:
[0045] f1(x1)=R(W1*x1)
[0046] Among them, x1 is the initial reconstruction value, W1 is the convolution layer for extracting the preliminary features of the upper channel, R() represents the ReLU activation function, and f1(x1) represents the preliminary features of the upper channel obtained by convolution extraction;
[0047] S4.1.1.2: Let f1(x1) be x 0 , for x 0 Perform the stacked residual shrinkage convolution operation to obtain the upper channel shrinkage feature:
[0048]
[0049] Among them, Shrinkage represents the residual shrinkage operation, W2 represents the first convolution layer in the residual shrinkage operation of the upper channel, and W3 represents the second convolution layer in the residual shrinkage operation of the upper channel. represents the nth residual contraction operation of the upper channel, represents the n-1th residual contraction operation of the upper channel, x n represents the output of the nth residual contraction operation in the upper channel, x n-2 represents the output of the n-2th residual shrinkage operation in the upper channel;
[0050] S4.1.1.3: Perform convolution operation on the upper channel contraction feature to obtain the low-dimensional receptive field channel feature:
[0051] f3(x n )=R(W4*x n )
[0052] Among them, W4 represents the final convolutional layer of the upper channel;
[0053] The lower channel is used to extract high-dimensional receptive field channel features, which specifically includes the following steps:
[0054] S4.1.2.1: Extract preliminary features of the lower channel through dilated convolution:
[0055] df1(x1)=R(W 11 *x1)
[0056] Among them, W 11 represents the convolution layer for extracting preliminary features of the lower channel, and df1(x1) represents the preliminary features of the lower channel obtained by convolution extraction;
[0057] S4.1.2.2: Let df1(x1) be x 00 , for x 00 Perform the stacked residual shrinkage convolution operation to obtain the shrinkage feature of the lower channel:
[0058]
[0059] Among them, W 22 represents the first convolutional layer in the residual contraction operation of the lower channel, W 33 represents the second convolutional layer in the residual contraction operation of the lower channel, represents the nth residual contraction operation of the lower channel, represents the n-1th residual contraction operation of the lower channel, x nn The output of the nth residual shrinkage operation in the lower channel, x (n-2)(n-2) represents the output of the n-2th residual contraction operation in the lower channel;
[0060] S4.1.2.3: Perform an expansion convolution operation on the contraction feature of the lower channel to obtain the high-dimensional receptive field channel feature:
[0061] df3(x nn )=R(W 44 *x nn )
[0062] Among them, W 44 represents the final convolutional layer of the lower channel;
[0063] S4.2: Fuse the low-dimensional receptive field channel features with the high-dimensional receptive field channel features, and use the super channel attention mechanism to enhance the fused features:
[0064] x c =ECA(Concat(f3(x2),df3(x2)))
[0065] Among them, x cis the enhanced feature, Concat(·) represents the feature fusion operation, and ECA(·) represents the attention mechanism processing;
[0066] S4.3: The enhanced features are convolved to obtain enhanced reconstruction values:
[0067] x2=W f *x c
[0068] Among them, W f Represents a convolutional layer for feature aggregation.
[0069] Preferably, the receptive field calculation of the dilated convolution is expressed as:
[0070] RF=k+(dilation-1)(k-1)=k·dilation-dilation+1
[0071] Among them, RF represents the receptive field size, k represents the convolution kernel size, and dilation represents the expansion rate.
[0072] Preferably, the residual shrinkage operation mainly includes:
[0073] A1: Perform global average pooling:
[0074] avg = Avgpolling(·)
[0075] Among them, avg represents the vector after average pooling, and Avgpolling() represents the average pooling operation;
[0076] A2: Information interaction through one-dimensional convolution:
[0077] x 1d =conv1d(avg)
[0078] Among them, conv1d() represents a one-dimensional convolution operation, x 1d Represents the value of avg after one-dimensional convolution;
[0079] A3: Use the sigmoid function to get the weight vector:
[0080] α=sigmoid(x 1d )
[0081] A4: Multiply avg by the weight vector to get the threshold vector:
[0082]
[0083] A5: Perform soft threshold shrinkage operation according to the threshold vector.
[0084] Preferably, the measurement value residual is reconstructed through three steps of convolution, residual feature extraction, and convolution in sequence to obtain a measurement value compensation reconstruction value.
[0085] Preferably, the reconstructed image is obtained by calculating the residual sum of the initial reconstruction value x1, the enhanced reconstruction value x2, and the compensated reconstruction value x3:
[0086] x f =x1+x2+x3.
[0087] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:
[0088] The present invention provides a multi-scale image compression and reconstruction method combined with an attention mechanism, which first performs an initial reconstruction on the measurement value, and then extracts different dimensional features through multi-scale reconstruction and then performs feature fusion, thereby better extracting features from the image and reducing feature loss; and in the reconstruction, further reconstructing separately by calculating the measurement value residual, which makes fuller use of the measurement value information and reduces the measurement value loss; finally, the initial reconstruction value, the enhanced reconstruction value and the compensated reconstruction value are combined to obtain a reconstructed image, thereby obtaining a better reconstruction effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] Figure 1 It is a flowchart of the implementation steps of the technical solution of the present invention;
[0090] Figure 2 It is a structural schematic diagram of the multi-scale reconstruction model in the present invention;
[0091] Figure 3 It is a schematic diagram of the process of adjusting the measurement value in the present invention;
[0092] Figure 4 Schematic diagram of the structure of the residual feature extraction module in the present invention;
[0093] Figure 5 Schematic diagram of the upper channel structure of the multi-scale module in the present invention;
[0094] Figure 6 Schematic diagram of the lower channel structure of the multi-scale module in the present invention;
[0095] Figure 7 It is a structural schematic diagram of a novel superposition residual shrinkage module in the present invention;
[0096] Figure 8 A line graph comparing the reconstruction performance of the present invention and the existing reconstruction method;
[0097] Fig. 9 It is a comparison chart of the intuitive reconstruction effects of the present invention and the existing reconstruction methods. DETAILED DESCRIPTION
[0098] The drawings are for illustrative purposes only and should not be construed as limiting the present patent;
[0099] In order to better illustrate the present embodiment, some parts in the drawings may be omitted, enlarged or reduced, and do not represent the size of the actual product;
[0100] It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0101] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0102] Example 1
[0103] like Figure 1 As shown, a multi-scale image compression and reconstruction method combined with an attention mechanism includes the following steps:
[0104] S1: Obtain the image to be reconstructed and build a multi-scale reconstruction model;
[0105] The multi-scale reconstruction model includes a sampling module, an initial reconstruction module and an enhanced reconstruction module;
[0106] S2: In the sampling module, convolution sampling is performed on the image to be reconstructed to obtain a measurement value, and the measurement value is adjusted;
[0107] S3: Performing initial reconstruction on the adjusted measurement value in the initial reconstruction module to obtain an initial reconstruction value and calculating the measurement value residual;
[0108] S4: in the enhanced reconstruction module, the initial reconstruction value is reconstructed at multiple scales to obtain an enhanced reconstruction value, and a compensated reconstruction value is calculated according to the measurement value residual;
[0109] S5: Obtain a reconstructed image according to the initial reconstruction value, the enhanced reconstruction value and the compensated reconstruction value.
[0110] In the specific implementation process, the measurement values are first reconstructed initially, and then the features of different dimensions are extracted respectively through multi-scale reconstruction and then feature fusion is performed, so as to better extract the features of the image and reduce the feature loss; and in the reconstruction, the measurement value residual is further calculated and reconstructed separately, so as to make fuller use of the measurement value information and reduce the measurement value loss; finally, the initial reconstruction value, the enhanced reconstruction value and the compensated reconstruction value are combined to obtain the reconstructed image and obtain a better reconstruction effect.
[0111] Example 2
[0112] A multi-scale image compression and reconstruction method combined with an attention mechanism comprises the following steps:
[0113] S1: Obtain the image to be reconstructed and build a multi-scale reconstruction model;
[0114] The multi-scale reconstruction model includes a sampling module, an initial reconstruction module and an enhanced reconstruction module. Figure 2 As shown;
[0115] In the specific implementation process, the training process of the multi-scale reconstruction model requires first training the sub-model consisting of the sampling module and the initial reconstruction module separately before training the overall model.
[0116] The loss function for sub-model training is:
[0117]
[0118] The loss function of the overall model training is:
[0119]
[0120] x is the image to be reconstructed, x f represents the reconstructed image, ε represents the regularization parameter, and the default setting is 0.1.
[0121] S2: In the sampling module, convolution sampling is performed on the image to be reconstructed to obtain a measurement value, and the measurement value is adjusted;
[0122] More specifically, before performing convolution sampling on the image to be reconstructed, the image to be reconstructed is divided into blocks, which are cut into a number of image blocks with a pixel size of 33×33; and the default image channel number l=1, and the sampling rate is set to r;
[0123] The convolution sampling is specifically:
[0124] y=S(x)=W M *x=Wx
[0125] Among them, y represents the measured value obtained by sampling, S() represents the mapping process of sampling, x represents the input image, and W M Represents the convolution kernel group, m = r × N 2 represents the number of convolution kernels, r represents the sampling rate, and N represents the image block size;
[0126] W represents the network parameter matrix composed of convolution kernel groups, specifically:
[0127]
[0128]
[0129] w1,w2,w3,w m They respectively represent the first convolution kernel, the second convolution kernel, the third convolution kernel, and the mth convolution kernel in the network parameter matrix, and R represents a real number.
[0130] In the specific implementation process, the sampling module consists of a convolution layer with a convolution kernel size of 33×33, which replaces the traditional measurement matrix for sampling. The original input is convolved with m convolution kernels respectively, and the m outputs obtained constitute the required measurement value y, which is equivalent to mapping the measurement matrix to the convolution layer parameters. During the network training process, along with the update of the convolution layer parameters, the measurement matrix is also updated at the same time, making the sampling adaptive.
[0131] During the training process, the parameters of the convolutional layer are initialized using the Xavier method. The models of each sampling rate are trained in 50 batches (Epoches), each batch size is 64 (Batch), and the learning rate α is set to 10 -4 , the parameters in the Adam optimizer are all set by default.
[0132] More specifically, Figure 3 As shown, in step S2, the measurement value adjustment includes the following steps:
[0133] S2.1: Perform global average pooling on the measured values:
[0134] y1=avgpolling(y)
[0135] S2.2: Measurement information exchange through one-dimensional convolution:
[0136] y2=conv1d(y1)
[0137] S2.3: Use sigmoid function for weight distribution:
[0138] β=sigmoid(y2)
[0139] S2.4: Perform normalization:
[0140] β1=softmax(β)
[0141] S2.5: Adjust the measured value:
[0142]
[0143] Among them, y1 represents the measured value after global average pooling, y2 represents the value of y1 after one-dimensional convolution, β represents the weight, β1 represents the normalized weight, and y f Represents the adjusted measurement value.
[0144] In the specific implementation process, the measurement value is adjusted in combination with the attention mechanism.
[0145] S3: Performing initial reconstruction on the adjusted measurement value in the initial reconstruction module to obtain an initial reconstruction value and calculating the measurement value residual;
[0146] More specifically, the initial reconstruction is expressed as:
[0147] x1=IR(y f )=W int *y f
[0148] Among them, x1 is the initial reconstruction value, IR() is the initial reconstruction function, W int is the convolutional layer used for initial reconstruction.
[0149] In the specific implementation process, the initial reconstruction module includes a convolution layer with a convolution kernel size of 1×1.
[0150] More specifically, the measurement residual is calculated using the following formula:
[0151] △y=y f -W*x1.
[0152] S4: in the enhanced reconstruction module, the initial reconstruction value is reconstructed at multiple scales to obtain an enhanced reconstruction value, and a compensated reconstruction value is calculated according to the measurement value residual;
[0153] More specifically, the initial reconstruction value is reconstructed at multiple scales through the following steps to obtain an enhanced reconstruction value:
[0154] S4.1: Extract features of different dimensions from the initial reconstruction value by strengthening the upper channel and the lower channel of the reconstruction module, where:
[0155] The upper channel is used to extract low-dimensional receptive field channel features, which specifically includes the following steps:
[0156] S4.1.1.1: Extract preliminary features of the upper channel through convolution:
[0157] f1(x1)=R(W1*x1)
[0158] Among them, x1 is the initial reconstruction value, W1 is the convolution layer for extracting the preliminary features of the upper channel, R() represents the ReLU activation function, and f1(x1) represents the preliminary features of the upper channel obtained by convolution extraction;
[0159] S4.1.1.2: Let f1(x1) be x 0 , for x 0 Perform the stacked residual shrinkage convolution operation to obtain the upper channel shrinkage feature:
[0160]
[0161] Among them, Shrinkage represents the residual shrinkage operation, W2 represents the first convolution layer in the residual shrinkage operation of the upper channel, and W3 represents the second convolution layer in the residual shrinkage operation of the upper channel. represents the nth residual contraction operation of the upper channel, represents the n-1th residual contraction operation of the upper channel, x n represents the output of the nth residual contraction operation in the upper channel, x n-2 represents the output of the n-2th residual shrinkage operation in the upper channel;
[0162] S4.1.1.3: Perform convolution operation on the upper channel contraction feature to obtain the low-dimensional receptive field channel feature:
[0163] f3(x n )=R(W4*x n )
[0164] Among them, W4 represents the final convolutional layer of the upper channel;
[0165] The lower channel is used to extract high-dimensional receptive field channel features, which specifically includes the following steps:
[0166] S4.1.2.1: Extract preliminary features of the lower channel through dilated convolution:
[0167] df1(x1)=R(W 11 *x1)
[0168] Among them, W 11 represents the convolution layer for extracting preliminary features of the lower channel, and df1(x1) represents the preliminary features of the lower channel obtained by convolution extraction;
[0169] S4.1.2.2: Let df1(x1) be x 00 , for x 00 Perform the stacked residual shrinkage convolution operation to obtain the shrinkage feature of the lower channel:
[0170]
[0171] Among them, W 22 represents the first convolutional layer in the residual contraction operation of the lower channel, W 33 represents the second convolutional layer in the residual contraction operation of the lower channel, represents the nth residual contraction operation of the lower channel, represents the n-1th residual contraction operation of the lower channel, x nn The output of the nth residual shrinkage operation in the lower channel, x (n-2)(n-2) represents the output of the n-2th residual contraction operation in the lower channel;
[0172] S4.1.2.3: Perform an expansion convolution operation on the contraction feature of the lower channel to obtain the high-dimensional receptive field channel feature:
[0173] df3(x nn )=R(W 44 *x nn )
[0174] Among them, W 44 represents the final convolutional layer of the lower channel;
[0175] S4.2: Fuse the low-dimensional receptive field channel features with the high-dimensional receptive field channel features, and use the super channel attention mechanism to enhance the fused features:
[0176] x c =ECA(Concat(f3(x2),df3(x2)))
[0177] Among them, x c is the enhanced feature, Concat(·) represents the feature fusion operation, and ECA(·) represents the attention mechanism processing;
[0178] S4.3: The enhanced features are convolved to obtain enhanced reconstruction values:
[0179] x2=W f *x c
[0180] Among them, W f Represents a convolutional layer for feature aggregation.
[0181] In the specific implementation process, the convolution kernel size of the convolution layer in the enhanced reconstruction module is set to 3, the step size is 1, and the default n in the stacked residual shrinkage convolution operation is 3. The enhanced reconstruction module includes a multi-scale module, a residual feature extraction module (such as Figure 4 As shown in ), a super channel attention module and multiple convolutional layers, where the multi-scale module includes an upper channel and a lower channel, as shown in Figure 5-6 As shown, the upper channel includes 2 convolution blocks and 1 new stacking residual shrinkage module, and the lower channel includes 2 dilated convolution blocks and 1 new stacking residual shrinkage module. The new stacking residual shrinkage module is shown in Figure 7 shown.
[0182] More specifically, the receptive field calculation of the dilated convolution is expressed as:
[0183] RF=k+(dilation-1)(k-1)=k·dilation-dilation+1
[0184] Among them, RF represents the receptive field size, k represents the convolution kernel size, and dilation represents the expansion rate.
[0185] More specifically, the residual shrinkage operation mainly includes:
[0186] A1: Perform global average pooling:
[0187] avg = Avgpolling(·)
[0188] Among them, avg represents the vector after average pooling, and Avgpolling() represents the average pooling operation;
[0189] A2: Information interaction through one-dimensional convolution:
[0190] x 1d =conv1d(avg)
[0191] Among them, conv1d() represents a one-dimensional convolution operation, x 1d Represents the value of avg after one-dimensional convolution;
[0192] A3: Use the sigmoid function to get the weight vector:
[0193] α=sigmoid(x 1d )
[0194] A4: Multiply avg by the weight vector to get the threshold vector:
[0195]
[0196] A5: Perform soft threshold shrinkage operation according to the threshold vector.
[0197] More specifically, the measurement value residual is reconstructed through three steps of convolution, residual feature extraction, and convolution to obtain a measurement value compensation reconstruction value.
[0198] S5: Obtain a reconstructed image according to the initial reconstruction value, the enhanced reconstruction value and the compensated reconstruction value.
[0199] More specifically, the reconstructed image is obtained by calculating the residual sum of the initial reconstruction value x1, the enhanced reconstruction value x2, and the compensated reconstruction value x3:
[0200] x f =x1+x2+x3.
[0201] Example 3
[0202] In this embodiment, a multi-scale image compression and reconstruction method combined with an attention mechanism is simulated with other existing methods under exactly the same conditions. The experimental platform used in the experiment is shown in Table 1. The reconstruction performance of this method is judged by comparing the average peak signal-to-noise ratio (PSNR) of 11 images in the Set11 test set.
[0203] Table 1
[0204]
[0205]
[0206] Table 2 shows the reconstruction performance comparison of a multi-scale image compression reconstruction method (Ours) combined with an attention mechanism at four sampling rates of 0.01, 0.04, 0.10, and 0.25 with other traditional compressed sensing reconstruction algorithms and reconstruction methods based on deep learning.
[0207] Table 2
[0208]
[0209] It can be seen that the reconstruction performance of the multi-scale image compression and reconstruction method (Ours) combined with the attention mechanism is better than that of other methods in Table 2 at any sampling rate. Figure 8 A line graph showing the reconstruction performance comparison between the multi-scale image compression and reconstruction method (Ours) combined with the attention mechanism and other different methods.
[0210] Table 3 shows the reconstruction time comparison of a multi-scale image compression reconstruction method (Ours) combined with an attention mechanism at four sampling rates of 0.01, 0.04, 0.10, and 0.25, as well as other traditional compressed sensing reconstruction algorithms and reconstruction methods based on deep learning.
[0211] Table 3
[0212]
[0213] Compared with SDA and ReconNet, the reconstruction time of the multi-scale image compression reconstruction method (Ours) combined with the attention mechanism is higher, but the reconstruction performance far exceeds these two reconstruction methods. Compared with DR2-Net and ISTA-Net, it not only has advantages in reconstruction performance, but also has a lower reconstruction time.
[0214] Fig. 9An intuitive comparison between the multi-scale image compression and reconstruction method (Ours) combined with the attention mechanism and other reconstruction methods is shown. It can be seen that the image reconstructed by the multi-scale image compression and reconstruction method (Ours) combined with the attention mechanism is clearer.
[0215] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the embodiments here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the claims of the present invention.
Claims
1. A multi-scale image compression and reconstruction method combined with an attention mechanism, characterized in that: The following steps are involved: S1: Obtain the image to be reconstructed and build a multi-scale reconstruction model; The multi-scale reconstruction model includes a sampling module, an initial reconstruction module and an enhanced reconstruction module; S2: In the sampling module, convolution sampling is performed on the image to be reconstructed to obtain a measurement value, and the measurement value is adjusted; S3: Performing initial reconstruction on the adjusted measurement value in the initial reconstruction module to obtain an initial reconstruction value and calculating the measurement value residual; S4: in the enhanced reconstruction module, the initial reconstruction value is reconstructed at multiple scales to obtain an enhanced reconstruction value, and a compensated reconstruction value is calculated according to the measurement value residual; The initial reconstruction value is reconstructed at multiple scales through the following steps to obtain an enhanced reconstruction value: S4.1: Extract features of different dimensions from the initial reconstruction value by strengthening the upper channel and the lower channel of the reconstruction module, where: The upper channel is used to extract low-dimensional receptive field channel features, which specifically includes the following steps: S4.1.1.1: Extract preliminary features of the upper channel through convolution: f1(x1)=R(W1*x1) Among them, x1 is the initial reconstruction value, W1 is the convolution layer for extracting the preliminary features of the upper channel, R() represents the ReLU activation function, and f1(x1) represents the preliminary features of the upper channel obtained by convolution extraction; S4.1.1.2: Let f1(x1) be x 0 , for x 0 Perform the stacked residual shrinkage convolution operation to obtain the upper channel shrinkage feature: Among them, Shrinkage represents the residual shrinkage operation, W2 represents the first convolution layer in the residual shrinkage operation of the upper channel, and W3 represents the second convolution layer in the residual shrinkage operation of the upper channel. represents the nth residual contraction operation of the upper channel, represents the n-1th residual contraction operation of the upper channel, x n represents the output of the nth residual contraction operation in the upper channel, x n-2 represents the output of the n-2th residual shrinkage operation in the upper channel; S4.1.1.3: Perform convolution operation on the upper channel contraction feature to obtain the low-dimensional receptive field channel feature: f3(x n )=R(W4*x n ) Among them, W4 represents the final convolutional layer of the upper channel; The lower channel is used to extract high-dimensional receptive field channel features, which specifically includes the following steps: S4.1.2.1: Extract preliminary features of the lower channel through dilated convolution: df1(x1)=R(W 11 *x1) Among them, W 11 represents the convolution layer for extracting preliminary features of the lower channel, and df1(x1) represents the preliminary features of the lower channel obtained by convolution extraction; S4.1.2.2: Let df1(x1) be x 00 , for x 00 Perform the stacked residual shrinkage convolution operation to obtain the shrinkage feature of the lower channel: Among them, W 22 represents the first convolutional layer in the residual contraction operation of the lower channel, W 33 represents the second convolutional layer in the residual contraction operation of the lower channel, represents the nth residual contraction operation of the lower channel, represents the n-1th residual contraction operation of the lower channel, x nn The output of the nth residual shrinkage operation in the lower channel, x (n-2)(n-2) represents the output of the n-2th residual contraction operation in the lower channel; S4.1.2.3: Perform an expansion convolution operation on the contraction feature of the lower channel to obtain the high-dimensional receptive field channel feature: df3(x nn )=R(W 44 *x nn ) Among them, W 44 represents the final convolutional layer of the lower channel; S4.2: Fuse the low-dimensional receptive field channel features with the high-dimensional receptive field channel features, and use the super channel attention mechanism to enhance the fused features: x c =ECA(Concat(f3(x2),df3(x2))) Among them, x c is the enhanced feature, Concat(·) represents the feature fusion operation, and ECA(·) represents the attention mechanism processing; S4.3: The enhanced features are convolved to obtain enhanced reconstruction values: x2=W f *x c Among them, W f Represents the convolutional layer used for feature aggregation; S5: Obtain a reconstructed image according to the initial reconstruction value, the enhanced reconstruction value and the compensated reconstruction value.
2. According to the multi-scale image compression and reconstruction method combined with the attention mechanism of claim 1, it is characterized in that: Before performing convolution sampling on the image to be reconstructed, the image to be reconstructed is divided into blocks, which are cut into a number of image blocks with a pixel size of 33×33; The convolution sampling is specifically: y=S(x)=W M *x=Wx Among them, y represents the measured value obtained by sampling, S() represents the mapping process of sampling, x represents the input image, and W M Represents the convolution kernel group, m = r × N 2 represents the number of convolution kernels, r represents the sampling rate, and N represents the image block size; W represents the network parameter matrix composed of convolution kernel groups, specifically: w1,w2,w3,w m They respectively represent the first convolution kernel, the second convolution kernel, the third convolution kernel, and the mth convolution kernel in the network parameter matrix, and R represents a real number.
3. According to claim 2, a multi-scale image compression and reconstruction method combined with an attention mechanism is characterized in that: In step S2, the measurement value adjustment includes the following steps: S2.1: Perform global average pooling on the measured values: y1=avgpolling(y) S2.2: Measurement information exchange through one-dimensional convolution: y2=conv1d(y1) S2.3: Use sigmoid function for weight distribution: β=sigmoid(y2) S2.4: Perform normalization: β1=softmax(β) S2.5: Adjust the measured value: Among them, y1 represents the measured value after global average pooling, y2 represents the value of y1 after one-dimensional convolution, β represents the weight, β1 represents the normalized weight, and y f Represents the adjusted measurement value.
4. According to claim 3, a multi-scale image compression and reconstruction method combined with an attention mechanism is characterized in that: The initial reconstruction is represented as: x1=IR(y f )=W int *and f Among them, x1 is the initial reconstruction value, IR() is the initial reconstruction function, W int is the convolutional layer used for initial reconstruction.
5. According to claim 4, a multi-scale image compression and reconstruction method combined with an attention mechanism is characterized in that: The measurement residual is calculated using the following formula: △y=y f -W*x1.
6. The multi-scale image compression and reconstruction method combined with the attention mechanism according to claim 1, characterized in that: The receptive field calculation of the dilated convolution is expressed as: RF=k+(dilation-1)(k-1)=k·dilation-dilation+1 Among them, RF represents the receptive field size, k represents the convolution kernel size, and dilation represents the expansion rate.
7. The multi-scale image compression and reconstruction method combined with the attention mechanism according to claim 1, characterized in that: The residual shrinkage operation mainly includes: A1: Perform global average pooling: avg = Avgpolling(·) Among them, avg represents the vector after average pooling, and Avgpolling() represents the average pooling operation; A2: Information interaction through one-dimensional convolution: x 1d =conv1d(avg) Among them, conv1d() represents a one-dimensional convolution operation, x 1d Represents the value of avg after one-dimensional convolution; A3: Use the sigmoid function to get the weight vector: α=sigmoid(x 1d ) A4: Multiply avg by the weight vector to get the threshold vector: A5: Perform soft threshold shrinkage operation according to the threshold vector.
8. The multi-scale image compression and reconstruction method combined with the attention mechanism according to claim 7, characterized in that: The measurement value residual is reconstructed through three steps of convolution, residual feature extraction, and convolution to obtain the measurement value compensation reconstruction value.
9. The multi-scale image compression and reconstruction method combined with the attention mechanism according to claim 8, characterized in that: The reconstructed image is obtained by calculating the residual sum of the initial reconstruction value x1, the enhanced reconstruction value x2, and the compensated reconstruction value x3: x f =x1+x2+x3。
Citation Information
Patent Citations
Content-guide Residual Network for Image Super-Resolution
AU2020100200A4
Super-resolution reconstruction method based on multi-scale residual attention
CN114331830A