A lightweight image super-resolution reconstruction method and system

By constructing a lightweight super-resolution neural network based on partial convolution and multi-scale attention, the problem of excessive parameter size in super-resolution networks is solved, achieving high-quality image restoration results with small model size and computational cost, making it suitable for industrial scenarios.

CN119540055BActive Publication Date: 2025-11-11NANJING FORESTRY UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411372260.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-29
Publication Date
2025-11-11
Estimated Expiration
2044-09-29

AI Technical Summary

Technical Problem

Existing super-resolution networks have too many parameters, making them difficult to deploy on mobile devices or under conditions where device performance is not ideal, and it is also difficult to achieve a good balance between the number of parameters and network performance.

Method used

A lightweight image super-resolution reconstruction method is adopted. A lightweight super-resolution neural network based on a partially convolutional information distillation mechanism module and a multi-scale attention mechanism module is constructed, including a shallow extraction module, a deep feature extraction module, and a reconstruction module. The multi-scale partially convolutional residual distillation module and the multi-scale enhanced spatial attention module are used to extract and fuse image features.

Benefits of technology

It achieves enhanced image resolution and high-quality image restoration with relatively small model size and computational cost, solving the problems of excessive parameter size and performance balance, and is suitable for practical industrial scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119540055B_ABST
    Figure CN119540055B_ABST
Patent Text Reader

Abstract

This invention discloses a lightweight image super-resolution reconstruction method and system, constructing a lightweight super-resolution neural network based on a partially convolutional information distillation mechanism module and a multi-scale attention mechanism module. The lightweight super-resolution neural network includes a shallow feature extraction module, a deep feature extraction module, and a reconstruction module. The deep feature extraction module includes five multi-scale partially convolutional residual distillation modules, a connection fusion layer, a second convolutional layer, a pixel attention module, and a second partially convolutional layer. The reconstruction module includes a first sub-pixel convolutional module, a second sub-pixel convolutional module, and a third convolutional layer. This invention not only effectively avoids channel redundancy and simplifies the feature extraction process, but also obtains more accurate spatial information distribution. Furthermore, by employing partially convolutional mechanisms to reduce module parameters, it achieves a good balance between network parameters and performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image processing technology, specifically a lightweight image super-resolution method and system based on information distillation and multi-scale attention. Background Technology

[0002] Single-image super-resolution is an important task in image processing, aiming to reconstruct high-resolution images from low-resolution images while simultaneously optimizing details and textures to improve visual perception quality. It is currently widely used in various real-world scenarios, including security monitoring, remote sensing, video enhancement, and image segmentation, and has therefore attracted widespread attention from both academia and industry.

[0003] In recent years, with the gradual development of deep learning technology, super-resolution technology has also matured, and many powerful super-resolution networks can achieve high-quality image restoration. However, in pursuit of high-quality image restoration results, these networks often have a very large number of parameters and high model complexity, making them difficult to deploy on mobile devices or under conditions where device performance is not ideal. To address this issue, researchers have begun to study efficient super-resolution reconstruction algorithms. Summary of the Invention

[0004] Purpose of the invention: To address the problems of excessively large parameter counts in existing super-resolution networks and the difficulty in achieving a good balance between parameter count and network performance, this invention provides a lightweight image super-resolution reconstruction method. This method can enhance image resolution with a smaller model size and computational load, and can therefore be applied in practical industrial applications.

[0005] Technical solution: To achieve the above objectives, the technical solution adopted by this invention is as follows:

[0006] A lightweight image super-resolution reconstruction method includes the following steps:

[0007] Step 1: Obtain lightweight images and preprocess them as the training set.

[0008] Step 2: Construct a lightweight super-resolution neural network based on a partially convolutional information distillation mechanism module and a multi-scale attention mechanism module.

[0009] The lightweight super-resolution neural network includes a shallow feature extraction module, a deep feature extraction module, and a reconstruction module.

[0010] The shallow extraction module includes a first convolutional layer.

[0011] The deep feature extraction module includes five multi-scale partial convolutional recalculation modules, a connection fusion layer, a second convolutional layer, a pixel attention module, and a second partial convolutional layer. The five multi-scale partial convolutional recalculation modules are connected in sequence, and the output of the first convolutional layer enters the first multi-scale partial convolutional recalculation module among the five multi-scale partial convolutional recalculation modules. At the same time, the outputs of the five multi-scale partial convolutional recalculation modules respectively enter the connection fusion layer. The connection fusion layer, the second convolutional layer, the pixel attention module, and the second partial convolutional layer are connected in sequence.

[0012] The reconstruction module includes a first subpixel convolutional module, a second subpixel convolutional module, and a third convolutional layer. The first subpixel convolutional module includes a first upsampling operation and a first summing layer. The second subpixel convolutional module includes a second summing layer and a second upsampling operation. The output of the first convolutional layer is fed into the third convolutional layer and the second summing layer, respectively. The third convolutional layer, the first upsampling operation, and the first summing layer are connected in sequence, as are the second summing layer, the second upsampling operation, and the first summing layer.

[0013] Step 3: Train the constructed lightweight super-resolution neural network using the training set to obtain the trained lightweight super-resolution neural network.

[0014] Step 4: Reconstruct the low-resolution image using the trained lightweight super-resolution neural network to obtain the reconstructed super-resolution image.

[0015] Preferably, the multi-scale partial convolutional residual distillation module includes four fourth-part residual modules, four fourth convolutional layers, a fourth connection fusion layer, a channel attention module, a multi-scale enhanced spatial attention module, and a fourth additive layer. The four fourth-part residual modules are fourth-part residual module one, fourth-part residual module two, fourth-part residual module three, and fourth-part residual module four. The four fourth convolutional layers are fourth convolutional layer one, fourth convolutional layer two, fourth convolutional layer three, and fourth convolutional layer four. The fourth-part residual modules one, four-part residual modules two, four-part residual modules three, four-part residual modules four, the fourth connection fusion layer, the fourth convolutional layer four, the channel attention module, the multi-scale enhanced spatial attention module, and the fourth additive layer one are connected sequentially. The input of the multi-scale partial convolutional residual distillation module is fed into the fourth-part residual module one, the fourth convolutional layer one, and the fourth additive layer one, respectively. The output of the fourth convolutional layer one is connected to the fourth connection fusion layer. The input of the fourth convolutional layer 2 is connected to the output of the fourth residual module 1, and the output of the fourth convolutional layer 2 is connected to the fourth connection fusion layer. The input of the fourth convolutional layer 3 is connected to the output of the fourth residual module 2, and the output of the fourth convolutional layer 3 is connected to the fourth connection fusion layer.

[0016] Preferably, the multi-scale enhanced spatial attention module includes a fifth convolutional layer 1, a fifth compressed convolutional layer 1, a fifth max pooling layer 1, a fifth compressed convolutional layer 2, a fifth max pooling layer 2, a fifth compressed convolutional layer 3, a fifth max pooling layer 3, a fifth compressed convolutional layer 4, a fifth local variance module, a fifth summing layer 1, a fifth partial convolutional group 1, a fifth partial convolutional group 2, a fifth partial convolutional group 3, a fifth upsampling module, a fifth summing layer 2, a fifth convolutional layer 2, a fifth activation function module, an element-wise multiplication operation module, and a fifth summing layer 3. The inputs of the multi-scale enhanced spatial attention module are respectively fed into the fifth convolutional layer 1, the element-wise multiplication operation module, and the fifth summing layer 3. The outputs of the fifth convolutional layer 1 are respectively fed into the fifth compressed convolutional layer 1, the fifth compressed convolutional layer 2, the fifth compressed convolutional layer 3, and the fifth compressed convolutional layer 4. The fifth compressed convolutional layer 1, the fifth max pooling layer 1, and the fifth partial convolutional group 1 are connected sequentially. The fifth compressed convolutional layer 2, the fifth max pooling layer 2, and the fifth partial convolutional group 2 are connected sequentially. The fifth compressed convolutional layer 3, the fifth max pooling layer 3, the fifth summing layer 1, and the fifth partial convolutional group 3 are connected sequentially. The fifth compressed convolutional layer 4, the fifth local variance module, and the fifth summing layer 1 are connected sequentially. The outputs of the fifth partial convolutional group 1, the fifth partial convolutional group 2, and the fifth partial convolutional group 3 respectively enter the fifth upsampling module. The fifth upsampling module, the fifth summing layer 2, the fifth convolutional layer 2, the fifth activation function module, the element-wise multiplication operation module, and the fifth summing layer 3 are connected sequentially. The output of the fifth convolutional layer 1 enters the fifth summing layer 2.

[0017] Preferably, the fourth residual module comprises a fourth convolutional layer, a second fourth summing layer, and a Gelu activation layer connected in sequence. The inputs of the fourth residual module are respectively fed into the fourth convolutional layer and the second fourth summing layer. The fourth convolutional layer comprises a fourth input layer and a fourth output layer. The input of the fourth input layer is fed into the fourth output layer through a convolution operation, and the input of the fourth input layer is fed into the fourth output layer through direct mapping.

[0018] Preferably, the information distillation method of the multi-scale partial convolution residual distillation module is as follows:

[0019] Step S101, the input F of the multi-scale partial convolutional residual distillation module in The input is divided into two paths: one path inputs to the fourth residual module 1, and the other path inputs to the fourth convolutional layer 1 for feature refinement, denoted as:

[0020] F c1 =PRB1(F in ),

[0021] F d1 =C1(F in )

[0022] In the formula, C1 represents a convolution operation in the fourth convolutional layer, PRB1 represents a residual operation in the fourth residual module, and F d1 F represents the features extracted by the fourth convolutional layer. c1 This indicates that the fourth residual module is sent to the next operation.

[0023] Step S102, obtain F c1 To further refine this process, it can be represented as follows:

[0024] F c2 =PRB2(PRB1(F in ))

[0025] F d2 =C2(PRB1(F in ))

[0026] In the formula, C2 represents the second convolution operation of the fourth convolutional layer, PRB2 represents the second residual operation of the fourth residual module, and F d2 F represents the features extracted by the second convolutional layer of the fourth convolutional layer. c2 This indicates that the fourth residual module 2 is sent to the next operation.

[0027] Step S103, obtain F c2 Further refinement is needed; this process can be represented as follows:

[0028] F c3 =PRB3(PRB2(PRB1(F in ))),

[0029] F d3 =C3(PRB2(PRB1(F in )))

[0030] In the formula, C3 represents the third convolution operation of the fourth convolutional layer, PRB3 represents the third residual operation of the fourth residual module, and F d3 F represents the features extracted by the third convolutional layer. c3 This indicates that the fourth part, residual module three, is sent to the next operation.

[0031] Step S104, obtain F c3 The data is then fed into the fourth part, residual module four, for the final refinement step. This process is represented as follows:

[0032] F c4 =PRB4(PRB3(PRB2(PRB1(F in ))))

[0033] PRB4 indicates the fourth part of the residual module, representing the four parts of residual operations.c4 This indicates that the fourth residual module is sent to the next operation.

[0034] Step S105: The features extracted in the above steps are fused. This process is represented as follows:

[0035] F fused

[0036] =Concat(C1(F in ), C2(PRB1(F in )), C3(PRB2(PRB1(F in ))), PRB4(PRB3(PRB2(PRB1(F in )))))

[0037] In the formula, F fused The term "Concat" indicates the fusion feature, and "Concat" represents the fusion operation of the fourth connection fusion layer.

[0038] Preferably, the multi-scale attention method of the multi-scale partial convolutional residual distillation module is as follows:

[0039] Step S201: First, input F of the multi-scale attention mechanism is... s The number of channels is reduced by using a fifth convolutional layer to reduce the number of network parameters. This process can be represented as:

[0040] F in,1 =C 1×1 (F S )

[0041] Among them, C 1×1 F represents the convolution operation in the fifth convolutional layer. S F represents the input to the multi-scale attention mechanism. in,1 This represents the result of the convolution operation in the fifth convolutional layer.

[0042] Step S202, obtain F in,1 Compression across three scales and the introduction of local variance can be represented as follows:

[0043]

[0044] in, B3 indicates compressed local features, where B3 represents a compression operation followed by a max pooling operation following the fifth compressed convolutional layer 1; B5 represents a compression operation followed by a max pooling operation following the fifth compressed convolutional layer 2; and B7 represents a compression operation followed by a max pooling operation following the fifth compressed convolutional layer 3. L-vatThis indicates that the operation is performed after the fourth compression convolution operation in the fifth compression convolution layer, followed by the operation of the fifth local variance module.

[0045] Step S203, obtain The features are further refined by feeding them into a group of convolutions constructed from partial convolutions. This process can be represented as follows:

[0046]

[0047] in, This indicates that the result is obtained after partial convolution groups, B g This indicates a partial convolution group operation.

[0048] Step S204, obtain Upsampling is performed to obtain a high-resolution feature map, which is then added to the result obtained in step S201. This process is represented as follows:

[0049]

[0050] in, B represents the obtained high-resolution feature map. up This indicates an upsampling operation.

[0051] Step S205: The high-resolution feature map obtained in step S204 is... Adjusted to F via the fifth convolutional layer S The number of channels is the same, then it is processed by the Sigmoid function and then... S Add them together. This process can be represented as:

[0052]

[0053] in, H represents element-wise multiplication. conv1-1 This indicates the second convolution operation in the fifth convolutional layer, and Sigmoid represents the Sigmoid function.

[0054] Preferably, the shallow extraction process of the shallow extraction module is as follows:

[0055] F0 = B SF (I LR )

[0056] Where F0 represents the result of shallow feature extraction, B SF I represents the shallow feature extraction process. LR This represents the input image.

[0057] The feature extraction process of the deep feature extraction module is as follows:

[0058] F1 = B DF (BSF (I LR ))

[0059] Where F1 represents the deep feature extraction result, B DF B represents the deep feature extraction process. SF I represents the shallow feature extraction process. LR This represents the input image.

[0060] The reconstruction process of the reconstruction module is as follows:

[0061] I SR =B SR (B DF (B SF (I LR )))+B up (I LR )

[0062] Among them, I SR B represents the final result. SR B represents the second subpixel convolution process. DF B represents the deep feature extraction process. SF I represents the shallow feature extraction process. LR B represents the input image. up This represents the first subpixel convolution process.

[0063] Preferably, the channel attention module fuses the feature F by comparing channel attention. fused The input F for the multi-scale attention mechanism is obtained through processing. s .

[0064] Another object of the present invention is to provide a lightweight image super-resolution reconstruction system for implementing the aforementioned lightweight image super-resolution reconstruction method, comprising an input unit, a lightweight super-resolution neural network unit, and an output unit, wherein:

[0065] The input unit is used to input the low-resolution image to be reconstructed.

[0066] The lightweight super-resolution neural network unit is used to store the trained lightweight super-resolution neural network. The trained lightweight super-resolution neural network is used to reconstruct the low-resolution image to obtain the reconstructed super-resolution image.

[0067] The output unit is used to output the reconstructed super-resolution image.

[0068] Another object of the present invention is to provide a computer system including a memory and a processor, wherein the memory is used to store computer programs / instructions. The processor is used to execute the computer programs / instructions to implement the lightweight image super-resolution reconstruction method described above.

[0069] Compared with the prior art, the present invention has the following advantages:

[0070] The information distillation mechanism based on partial convolution can effectively reduce the number of network parameters. The multi-scale enhanced spatial attention mechanism can effectively extract and fuse features of different scales by introducing multi-scale information to obtain more useful image features. Therefore, this invention has a lightweight effect, and the number of model parameters is only 550K. At the same time, the designed super-resolution network achieves a good balance between network parameters and performance. It solves the problems of excessive number of parameters in the original super-resolution network and the difficulty in achieving a good balance between parameter number and network performance. Attached Figure Description

[0071] Figure 1 This is a schematic diagram of the super-resolution network structure proposed in this invention;

[0072] Figure 2 This is a schematic diagram of a partial convolution structure;

[0073] Figure 3 This is a schematic diagram of some of the residual modules mentioned in this invention;

[0074] Figure 4 This is a schematic diagram of the information distillation mechanism based on partial convolution proposed in this invention;

[0075] Figure 5 This is a schematic diagram of the multi-scale enhanced spatial attention mechanism proposed in this invention. Detailed Implementation

[0076] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that these examples are for illustrative purposes only and are not intended to limit the scope of the invention. After reading this invention, any modifications of the invention in various equivalent forms by those skilled in the art will fall within the scope defined by the appended claims.

[0077] This invention designs a lightweight image super-resolution reconstruction method, such as... Figure 1-5 As shown, it includes the following steps:

[0078] Step 1: Obtain lightweight images and preprocess them as the training set.

[0079] In another embodiment, the selected training dataset is specifically the DF2K dataset, which includes 3550 images. The test sets used are the commonly used Set5, Set14, Urban100, B100, and Manga109. Low-resolution images from DF2K are cropped into image patches of 96×96, 144×144, and 192×192 sizes, and then randomly rotated by one of the following operations: 90°, 180°, 270°, or horizontally flipped, respectively, and used as network inputs at ×2, ×3, and ×4 scales.

[0080] Step 2: Construct a lightweight super-resolution neural network based on a partially convolutional information distillation mechanism module and a multi-scale attention mechanism module.

[0081] The lightweight super-resolution neural network includes a shallow feature extraction module, a deep feature extraction module, and a reconstruction module.

[0082] The shallow extraction module includes a first convolutional layer, which is a 3×3 convolutional layer, with the purpose of expanding the number of channels of the input image from the original RGB three channels to a feature map with 64 channels.

[0083] Given an input I LR B SF F0 and F0 represent the shallow feature extraction process and the result of shallow feature extraction, respectively. The shallow feature extraction module's shallow feature extraction process is as follows:

[0084] F0 = B SF (I LR )

[0085] Where F0 represents the result of shallow feature extraction, B SF I represents the shallow feature extraction process. LR This represents the input image.

[0086] The purpose of the deep feature extraction module is to further process the features extracted by the shallow feature extraction module. The deep feature extraction module includes five multi-scale partial convolutional recalculation modules, a connection fusion layer, a second convolutional layer, a pixel attention module, and a second partial convolutional layer. The five multi-scale partial convolutional recalculation modules are connected sequentially, and the output of the first convolutional layer enters the first multi-scale partial convolutional recalculation module. Simultaneously, the outputs of the five multi-scale partial convolutional recalculation modules respectively enter the connection fusion layer. The connection fusion layer, the second convolutional layer, the pixel attention module, and the second partial convolutional layer are connected sequentially. The second convolutional layer is a 1×1 convolution, and the second partial convolutional layer is a 3×3 partial convolution with a ratio of 1 / 4.

[0087] Assuming the deep feature extraction process is B DFIf the input feature map is F0 and the output result is F1, then the feature extraction process of the deep feature extraction module is as follows:

[0088] F1 = B DF (B SF (I LR ))

[0089] Where F1 represents the deep feature extraction result, B DF B represents the deep feature extraction process. SF I represents the shallow feature extraction process. LR This represents the input image.

[0090] In another embodiment, the multi-scale partial convolutional residual distillation module includes four fourth partial residual modules, four fourth convolutional layers, a fourth connection fusion layer, a channel attention module, a multi-scale enhanced spatial attention module, and a fourth addition layer. The four fourth partial residual modules are designated as fourth partial residual module one, fourth partial residual module two, fourth partial residual module three, and fourth partial residual module four. From front to back, the four fourth partial residual modules are a 3×3 partial residual module with a ratio of 1, a 3×3 partial residual module with a ratio of 1 / 2, a 3×3 partial residual module with a ratio of 1 / 4, and a 3×3 partial residual module with a ratio of 1. The four fourth convolutional layers are designated as Convolutional Layer 1, Convolutional Layer 2, Convolutional Layer 3, and Convolutional Layer 4, each a 1×1 convolutional layer. The fourth partial residual module 1, Fourth partial residual module 2, Fourth partial residual module 3, Fourth partial residual module 4, Fourth connection fusion layer, Fourth convolutional layer 4, Channel attention module, Multi-scale enhanced spatial attention module, and Fourth addition layer 1 are sequentially connected. The inputs of the multi-scale partial convolutional residual module are fed into Fourth partial residual module 1, Fourth convolutional layer 1, and Fourth addition layer 1, respectively. The output of Fourth convolutional layer 1 is connected to the Fourth connection fusion layer. The input of Fourth convolutional layer 2 is connected to the output of Fourth partial residual module 1, and its output is connected to the Fourth connection fusion layer. The input of Fourth convolutional layer 3 is connected to the output of Fourth partial residual module 2, and its output is connected to the Fourth connection fusion layer.

[0091] The fourth residual module comprises a fourth convolutional layer, a second fourth summing layer, and a Gelu activation layer connected in sequence. The inputs to the fourth residual module are fed into the fourth convolutional layer and the second fourth summing layer, respectively. The fourth convolutional layer comprises a fourth input layer and a fourth output layer. The input of the fourth input layer is fed into the fourth output layer through a convolution operation, and the input of the fourth input layer is fed into the fourth output layer through direct mapping. The fourth convolutional layer is a 3×3 partial convolutional layer.

[0092] In another embodiment, the information distillation method of the multi-scale partial convolution residual distillation module is as follows:

[0093] Step S101, the input F of the multi-scale partial convolutional residual distillation module in The input is divided into two paths: one path inputs to the fourth residual module 1, and the other path inputs to the fourth convolutional layer 1 for feature refinement, denoted as:

[0094] F c1 =PRB1(F in ),

[0095] F d1 =C1(F in )

[0096] In the formula, C1 represents the first convolutional operation of the fourth convolutional layer, used to refine features. The first convolutional layer is a 1×1 convolutional layer. PRB1 represents a part of the residual operation of the fourth residual module. F d1 F represents the features extracted by the fourth convolutional layer. c1 This indicates that the fourth residual module is sent to the next operation.

[0097] Step S102, obtain F c1 To further refine this process, it can be represented as follows:

[0098] F c2 =PRB2(PRB1(F in ))

[0099] F d2 =C2(PRB1(F in ))

[0100] In the formula, C2 represents the second convolution operation of the fourth convolutional layer, which is a 1×1 convolutional layer used to refine features; PRB2 represents the second residual operation of the fourth residual module; F d2 F represents the features extracted by the second convolutional layer of the fourth convolutional layer. c2 This indicates that the fourth residual module 2 is sent to the next operation.

[0101] Step S103, obtain F c2 Further refinement is needed; this process can be represented as follows:

[0102] F c3 =PRB3(PRB2(PRB1(F in ))),

[0103] F d3 =C3(PRB2(PRB1(F in )))

[0104] In the formula, C3 represents the third convolution operation of the fourth convolutional layer, which is a 1×1 convolutional layer used to refine features; PRB3 represents the third residual operation of the fourth residual module; F d3 F represents the features extracted by the third convolutional layer. c3 This indicates that the fourth part, residual module three, is sent to the next operation.

[0105] Step S104, obtain F c3 The data is then fed into the fourth part, residual module four, for the final refinement step. This process is represented as follows:

[0106] F c4 =PRB4(PRB3(PRB2(PRB1(F in ))))

[0107] PRB4 indicates the fourth part of the residual module, representing the four parts of residual operations. c4 This indicates that the fourth residual module is sent to the next operation.

[0108] Step S105: The features extracted in the above steps are fused. This process is represented as follows:

[0109] F fused

[0110] =Concat(C1(F in ), C2(PRB1(F in )), C3(PRB2(PRB1(F in ))), PRB4(PRB3(PRB2(PRB1(F in )))))

[0111] In the formula, F fused The term "Concat" indicates the fusion feature, and "Concat" represents the fusion operation of the fourth connection fusion layer.

[0112] In another embodiment, the multi-scale enhanced spatial attention module includes a fifth convolutional layer 1, a fifth compressed convolutional layer 1, a fifth max pooling layer 1, a fifth compressed convolutional layer 2, a fifth max pooling layer 2, a fifth compressed convolutional layer 3, a fifth max pooling layer 3, a fifth compressed convolutional layer 4, a fifth local variance module, a fifth summing layer 1, a fifth partial convolutional group 1, a fifth partial convolutional group 2, a fifth partial convolutional group 3, a fifth upsampling module, a fifth summing layer 2, a fifth convolutional layer 2, a fifth activation function module, an element-wise multiplication operation module, and a fifth summing layer 3. The inputs of the multi-scale enhanced spatial attention module are respectively fed into the fifth convolutional layer 1, the element-wise multiplication operation module, and the fifth summing layer 3. The outputs of the fifth convolutional layer 1 are respectively fed into the fifth compressed convolutional layer 1, the fifth compressed convolutional layer 2, the fifth compressed convolutional layer 3, and the fifth compressed convolutional layer 4. The fifth compressed convolutional layer 1, the fifth max pooling layer 1, and the fifth partial convolutional group 1 are connected sequentially. The fifth compressed convolutional layer 2, the fifth max pooling layer 2, and the fifth partial convolutional group 2 are connected sequentially. The fifth compressed convolutional layer 3, the fifth max pooling layer 3, the fifth summing layer 1, and the fifth partial convolutional group 3 are connected sequentially. The fifth compressed convolutional layer 4, the fifth local variance module, and the fifth summing layer 1 are connected sequentially. The outputs of the fifth partial convolutional group 1, the fifth partial convolutional group 2, and the fifth partial convolutional group 3 respectively enter the fifth upsampling module. The fifth upsampling module, the fifth summing layer 2, the fifth convolutional layer 2, the fifth activation function module, the element-wise multiplication operation module, and the fifth summing layer 3 are connected sequentially. The output of the fifth convolutional layer 1 enters the fifth summing layer 2.

[0113] In another embodiment, the channel attention module fuses the feature F by comparing channel attention. fused The input F for the multi-scale attention mechanism is obtained through processing. s .

[0114] In another embodiment, the multi-scale attention method of the multi-scale partial convolutional residual distillation module is as follows:

[0115] Step S201: First, input F of the multi-scale attention mechanism is... s The number of channels is reduced by using a fifth convolutional layer to reduce the number of network parameters. This process can be represented as:

[0116] F in,1 =C 1×1 (F S )

[0117] Among them, C 1×1 This represents the convolution operation of the fifth convolutional layer, which is a 1×1 convolutional layer. F S F represents the input to the multi-scale attention mechanism. in,1 This represents the result of the convolution operation in the fifth convolutional layer.

[0118] Step S202, obtain F in,1 Compression across three scales and the introduction of local variance can be represented as follows:

[0119]

[0120] in, The symbols represent compressed local features. B3 indicates a compression operation followed by a max pooling operation after the fifth compressed convolutional layer 1; B5 indicates a compression operation followed by a max pooling operation after the fifth compressed convolutional layer 2; and B7 indicates a compression operation followed by a max pooling operation after the fifth compressed convolutional layer 3. B3, B5, and B7 represent sequentially passing through a 3×3 convolution with a stride of 2 and max pooling operations with strides of 3, 5, and 7, respectively. L-vat This indicates that after the fifth compressed convolutional layer performs a four-compact convolution operation, and then undergoes the operation of the fifth local variance module, B L-vat This indicates a local variance operation with a scale of 7.

[0121] Step S203, obtain The features are further refined by feeding them into a group of convolutions constructed from partial convolutions. This process can be represented as follows:

[0122]

[0123] in, This indicates that the result is obtained after a partial convolution group, which consists of two partial convolutions, B. g This indicates a partial convolution group operation.

[0124] Step S204, the obtained low-resolution feature map Upsampling is performed to obtain a high-resolution feature map, which is then added to the result obtained in step S201. This process is represented as follows:

[0125]

[0126] in, B represents the obtained high-resolution feature map. up This indicates an upsampling operation.

[0127] Step S205: The high-resolution feature map obtained in step S204 is... Adjusted to F via the fifth convolutional layer S The number of channels is the same, then it is processed by the Sigmoid function and then... S Add them together. This process can be represented as:

[0128]

[0129] Among them, F out Indicates the output. H represents element-wise multiplication. conv1-1 This indicates the second convolution operation in the fifth convolutional layer, which is a 1×1 convolutional layer. Sigmoid represents the Sigmoid function.

[0130] The reconstruction module includes a first subpixel convolutional module, a second subpixel convolutional module, and a third convolutional layer. The first subpixel convolutional module includes a first upsampling operation and a first summing layer. The second subpixel convolutional module includes a second summing layer and a second upsampling operation. The output of the first convolutional layer is fed into the third convolutional layer and the second summing layer, respectively. The third convolutional layer, the first upsampling operation, and the first summing layer are connected in sequence, as are the second summing layer, the second upsampling operation, and the first summing layer. The third convolutional layer is a 3×3 convolution.

[0131] Assume the process of the subpixel convolution module is B SR The initial input feature map undergoes a 3×3 convolution and a subpixel convolution process to form B. up The input feature map is F1, and the final result is I. SR The reconstruction process of the reconstruction module is as follows:

[0132] I SR =B SR (B DF (B SF (I LR )))+B up (I LR )

[0133] Among them, I SR B represents the final result. SR B represents the second subpixel convolution process. DF B represents the deep feature extraction process. SF I represents the shallow feature extraction process. LR B represents the input image. up This represents the first subpixel convolution process.

[0134] The loss function is chosen as follows:

[0135] I loss =∥I SR -I HR ∥1+α∥F(I SR )-F(I HR )∥

[0136] Among them, I loss I represents the loss value. SR It means, IHR It means, ∥I SR -I HR ∥1 represents, F(.) represents, ∥F(I) represents SR )-F(I HR ∥ represents the frequency domain loss, and α is the adaptive weight. Gradient descent is used to iterate the model parameters during training. The optimizer is Adam, with the following parameter settings: batch size of 16, exponential decay rate β1 for first-order moment estimation of 0.9, exponential decay rate β2 for second-order moment estimation of 0.999, and short-float type value ε for maintaining numerical stability of 10. -8 The initial learning rate is set to 1e. -3 The cosine annealing training strategy was adopted, and the number of training iterations was 1×10. 6 .

[0137] Step 3: Train the constructed lightweight super-resolution neural network using the training set to obtain the trained lightweight super-resolution neural network.

[0138] Step 4: Reconstruct the low-resolution image using the trained lightweight super-resolution neural network to obtain the reconstructed super-resolution image.

[0139] During the reconstruction process, assume the input image is F in The output image is F out The multi-scale partial convolution residual distillation module process involved can then be described as follows:

[0140] (a) Input F in The process is divided into two paths. One path is extracted through a 1×1 convolution to obtain F. di The other path was further refined to obtain F. ci .

[0141] (b) Extract the F di Perform a Concat operation to obtain F. fused .

[0142] (c) The fused feature F fused After refinement via 1×1 convolution, the input F for the multi-scale enhanced spatial attention mechanism is obtained through contrastive channel attention. s .

[0143] (d) F s Perform compression at scales of 3, 5, and 7, along with a local variance operation at scale 7, and sum them to obtain...

[0144] (e) Upsample the low-resolution attention map obtained in (d) to obtain a high-resolution attention map and perform a Sigmoid activation operation to obtain the weights.

[0145] (f) Integrating the weights with the input F of the multi-scale enhanced spatial attention mechanism s Multiply them, then add them together to get the output F. out .

[0146] In another embodiment, a lightweight image super-resolution reconstruction system is provided for implementing the aforementioned lightweight image super-resolution reconstruction method, comprising an input unit, a lightweight super-resolution neural network unit, and an output unit, wherein:

[0147] The input unit is used to input the low-resolution image to be reconstructed.

[0148] The lightweight super-resolution neural network unit is used to store the trained lightweight super-resolution neural network. The trained lightweight super-resolution neural network is used to reconstruct the low-resolution image to obtain the reconstructed super-resolution image.

[0149] The output unit is used to output the reconstructed super-resolution image.

[0150] In another embodiment, a computer system is provided, including a memory and a processor, the memory for storing computer programs / instructions. The processor is used to execute the computer programs / instructions to implement the lightweight image super-resolution reconstruction method described above.

[0151] In this embodiment, we propose a Multi-Scale Partial Convolutional Residual Distillation Super-Resolution Network (MPCRDN), with the Multi-Scale Partial Convolutional Residual Distillation Module (MPCRDB) serving as the foundational block. This module includes our multi-level information distillation mechanism based on partial convolution and a multi-scale enhanced spatial attention mechanism. The multi-level information distillation mechanism based on partial convolution effectively avoids channel redundancy and simplifies the feature extraction process. The multi-scale enhanced spatial attention introduces multi-scale information and local variance information to obtain a more accurate spatial information distribution, and uses partial convolution to reduce module parameters. This achieves a good balance between network parameters and performance.

[0152] To demonstrate the effectiveness of our proposed module, we designed several ablation experiments.

[0153] As shown in Table 1, the information distillation mechanism with weights of 1×2×4 reduces the average PSNR by 0.05 dB compared to the information distillation mechanism with weights of 1×1×1, while reducing the number of parameters by nearly half. This is mainly because redundant feature channels appear during convolution. These redundant channels do not significantly improve performance but introduce a considerable number of parameters. Compared to the information distillation mechanisms with weights of 2×2×2 and 4×4×4, although these two weights result in fewer network model parameters, the PSNR and SSIM results are too low. This is mainly because the feature extraction capability is not strong enough, and the refinement of features is insufficient. Therefore, we use a weight ratio of 1×2×4 as the weight of the information distillation module in the final MPCRDN.

[0154] Furthermore, we designed ablation experiments on the attention mechanism. We used MPCRDN as the base network and replaced our proposed MESA with ESA, which is commonly used in efficient super-resolution. The results are shown in Table 2, demonstrating that our proposed MESA performs better than ESA. Compared to ESA, MESA can better utilize the spatial information in the image, achieving better performance with fewer parameters. Specifically, the combination of MESA and CCA increases the number of parameters by 2K compared to the combination of ESA and CCA, but improves the average PSNR of Set5 by 0.07dB.

[0155] Table 1. Experimental data on information distillation mechanisms with different weights.

[0156]

[0157] Table 2 Ablation experiments based on attention mechanisms

[0158]

[0159] In this embodiment, neural networks, gradient descent, and the Adam optimizer are all technical terms in the field, existing technologies, and not the main improvement points of this invention, so they will not be described in detail.

[0160] This example demonstrates a lightweight image super-resolution method based on information distillation and multi-scale attention, achieving super-resolution reconstruction of low-resolution images. The resulting image improves resolution and increases image size while restoring high-frequency details and basic texture shapes. Simultaneously, the network model has a small number of parameters, making it suitable for industrial applications. The use of partially convolutional information distillation and multi-scale enhanced spatial attention avoids channel redundancy, effectively reducing the computational cost and number of parameters, thus ensuring model efficiency.

[0161] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A lightweight image super-resolution reconstruction method, characterized in that, Includes the following steps: Step 1: Obtain lightweight images and preprocess them as the training set; Step 2: Construct a lightweight super-resolution neural network based on a partially convolutional information distillation mechanism module and a multi-scale attention mechanism module; The lightweight super-resolution neural network includes a shallow extraction module, a deep feature extraction module, and a reconstruction module. The shallow extraction module includes a first convolutional layer. The deep feature extraction module includes five multi-scale partial convolutional distillation modules, a connection fusion layer, a second convolutional layer, a pixel attention module, and a second partial convolutional layer. The five multi-scale partial convolutional distillation modules are connected in sequence, and the output of the first convolutional layer enters the first multi-scale partial convolutional distillation module among the five multi-scale partial convolutional distillation modules. At the same time, the outputs of the five multi-scale partial convolutional distillation modules enter the connection fusion layer respectively. The connection fusion layer, the second convolutional layer, the pixel attention module, and the second partial convolutional layer are connected in sequence. The reconstruction module includes a first subpixel convolution module, a second subpixel convolution module, and a third convolutional layer. The first subpixel convolution module includes a first upsampling operation and a first summing layer. The second subpixel convolution module includes a second summing layer and a second upsampling operation. The output of the first convolutional layer is fed into the third convolutional layer and the second summing layer, respectively. The third convolutional layer, the first upsampling operation, and the first summing layer are connected in sequence. The second summing layer, the second upsampling operation, and the first summing layer are also connected in sequence. Step 3: Train the constructed lightweight super-resolution neural network using the training set to obtain the trained lightweight super-resolution neural network. Step 4: Reconstruct the low-resolution image using the trained lightweight super-resolution neural network to obtain the reconstructed super-resolution image.

2. The lightweight image super-resolution reconstruction method according to claim 1, characterized in that: The multi-scale partial convolutional residual distillation module includes four fourth-part residual modules, four fourth convolutional layers, a fourth connection fusion layer, a channel attention module, a multi-scale enhanced spatial attention module, and a fourth addition layer. The four fourth-part residual modules are fourth-part residual module one, fourth-part residual module two, fourth-part residual module three, and fourth-part residual module four. The four fourth convolutional layers are fourth convolutional layer one, fourth convolutional layer two, fourth convolutional layer three, and fourth convolutional layer four. The fourth-part residual module one, fourth-part residual module two, fourth-part residual module three, and fourth-part residual module four, and the fourth connection fusion layer... The fusion layer, the fourth convolutional layer, the channel attention module, the multi-scale enhanced spatial attention module, and the fourth additive layer are connected in sequence. The input of the multi-scale partial convolutional residual distillation module is respectively fed into the fourth partial residual module, the fourth convolutional layer, and the fourth additive layer. The output of the fourth convolutional layer is connected to the fourth connection fusion layer. The input of the fourth convolutional layer is connected to the output of the fourth partial residual module, and the output of the fourth convolutional layer is connected to the fourth connection fusion layer. The input of the fourth convolutional layer is connected to the output of the fourth partial residual module, and the output of the fourth convolutional layer is connected to the fourth connection fusion layer.

3. The lightweight image super-resolution reconstruction method according to claim 2, characterized in that: The multi-scale enhanced spatial attention module includes a fifth convolutional layer 1, a fifth compressed convolutional layer 1, a fifth max pooling layer 1, a fifth compressed convolutional layer 2, a fifth max pooling layer 2, a fifth compressed convolutional layer 3, a fifth max pooling layer 3, a fifth compressed convolutional layer 4, a fifth local variance module, a fifth summing layer 1, a fifth partial convolutional group 1, a fifth partial convolutional group 2, a fifth partial convolutional group 3, a fifth upsampling module, a fifth summing layer 2, a fifth convolutional layer 2, a fifth activation function module, an element-wise multiplication operation module, and a fifth summing layer 3. The inputs of the multi-scale enhanced spatial attention module are respectively fed into the fifth convolutional layer 1, the element-wise multiplication operation module, and the fifth summing layer 3, and the outputs of the fifth convolutional layer 1 are respectively fed into the fifth compressed convolutional layer 1, the fifth compressed convolutional layer 2, and the fifth compressed convolutional layer 3. The fifth compressed convolutional layer four is formed by sequentially connecting the fifth compressed convolutional layer one, the fifth max pooling layer one, and the fifth partial convolutional group one; sequentially connecting the fifth compressed convolutional layer two, the fifth max pooling layer two, and the fifth partial convolutional group two; sequentially connecting the fifth compressed convolutional layer three, the fifth max pooling layer three, the fifth summing layer one, and the fifth partial convolutional group three; sequentially connecting the fifth compressed convolutional layer four, the fifth local variance module, and the fifth summing layer one; the outputs of the fifth partial convolutional group one, the fifth partial convolutional group two, and the fifth partial convolutional group three respectively enter the fifth upsampling module; sequentially connecting the fifth upsampling module, the fifth summing layer two, the fifth convolutional layer two, the fifth activation function module, the element-wise multiplication operation module, and the fifth summing layer three; and the output of the fifth convolutional layer one enters the fifth summing layer two.

4. The lightweight image super-resolution reconstruction method according to claim 3, characterized in that: The fourth residual module includes a fourth convolutional layer, a second fourth summing layer, and a Gelu activation layer connected in sequence. The inputs of the fourth residual module enter the fourth convolutional layer and the second fourth summing layer, respectively. The fourth convolutional layer includes a fourth input layer and a fourth output layer. The input of the fourth input layer enters the fourth output layer through a convolution operation, and the input of the fourth input layer enters the fourth output layer through direct mapping.

5. The lightweight image super-resolution reconstruction method according to claim 4, characterized in that: The information distillation method of the multi-scale partial convolution residual distillation module is as follows: Step S101, the input F of the multi-scale partial convolutional residual distillation module in The input is divided into two paths: one path inputs to the fourth residual module 1, and the other path inputs to the fourth convolutional layer 1 for feature refinement, denoted as: F c1 =PRB1(F in ), F d1 =C1(F in ) In the formula, C1 represents a convolution operation in the fourth convolutional layer, PRB1 represents a residual operation in the fourth residual module, and F d1 F represents the features extracted by the fourth convolutional layer. c1 This indicates that the fourth residual module is sent to the next operation; Step S102, obtain F c1 To further refine this process, it can be represented as follows: F c2 =PRB2(PRB1(F in )) F d2 =C2(PRB1(F in )) In the formula, C2 represents the second convolution operation of the fourth convolutional layer, PRB2 represents the second residual operation of the fourth residual module, and F d2 F represents the features extracted by the second convolutional layer of the fourth convolutional layer. c2 This indicates that the fourth residual module 2 is sent to the next operation; Step S103, obtain F c2 Further refinement is needed; this process can be represented as follows: F c3 =PRB3(PRB2(PRB1(F in ))), F d3 =C3(PRB2(PRB1(F in ))) In the formula, C3 represents the third convolution operation of the fourth convolutional layer, PRB3 represents the third residual operation of the fourth residual module, and F d3 F represents the features extracted by the third convolutional layer. c3 This indicates that the third residual module in the fourth part is sent to the next operation; Step S104, obtain F c3 The data is then fed into the fourth part, residual module four, for the final refinement step. This process is represented as follows: F c4 =PRB4(PRB3(PRB2(PRB1(F in )))) PRB4 indicates the fourth part of the residual module, representing the four parts of residual operations. c4 This indicates that the fourth residual module is sent to the next operation; Step S105: The features extracted in the above steps are fused. This process is represented as follows: F fused =Concat(C1(F in ), C2(PRB1(F in )), C3(PRB2(PRB1(F in ))), PRB4(PRB3(PRB2(PRB1(F in In the formula, F fused The term "Concat" indicates the fusion feature, and "Concat" represents the fusion operation of the fourth connection fusion layer.

6. The lightweight image super-resolution reconstruction method according to claim 5, characterized in that: The multi-scale attention method of the multi-scale partial convolutional residual distillation module is as follows: Step S201: First, input F of the multi-scale attention mechanism is... s The number of channels is reduced by using a fifth convolutional layer to reduce the number of network parameters. This process can be represented as: F in,1 =C 1×1 (F S ) Among them, C 1×1 F represents the convolution operation in the fifth convolutional layer. S F represents the input to the multi-scale attention mechanism. in,1 This represents the result of the convolution operation in the first layer of the fifth convolutional layer; Step S202, obtain F in,1 Compression across three scales and the introduction of local variance can be represented as follows: F s 3 in =B3(F in,1 )+B5(F in,1 )+B7(F in,1 )+B L-vat (C 1×1 (F S )) Among them, F s 3 in B3 indicates compressed local features, where B3 represents a compression operation followed by a max pooling operation following the fifth compressed convolutional layer 1; B5 represents a compression operation followed by a max pooling operation following the fifth compressed convolutional layer 2; and B7 represents a compression operation followed by a max pooling operation following the fifth compressed convolutional layer 3. L-vat This indicates that the operation is performed after the fourth compression convolution operation in the fifth compression convolution layer, followed by the operation of the fifth local variance module. Step S203, obtain The features are further refined by feeding them into a group of convolutions constructed from partial convolutions. This process can be represented as follows: in, This indicates that the result is obtained after partial convolution groups, B g Indicates partial convolution group operations; Step S204, obtain Upsampling is performed to obtain a high-resolution feature map, which is then added to the result obtained in step S201. This process is represented as follows: in, B represents the obtained high-resolution feature map. up Indicates an upsampling operation; Step S205: The high-resolution feature map obtained in step S204 is... Adjusted to F via the fifth convolutional layer S The number of channels is the same, then it is processed by the Sigmoid function and then... S Addition; this process is represented as: Among them, F out Indicates the output. H represents element-wise multiplication. conv1-1 This indicates the second convolution operation in the fifth convolutional layer, and Sigmoid represents the Sigmoid function.

7. The lightweight image super-resolution reconstruction method according to claim 6, characterized in that: The shallow extraction process of the shallow extraction module is as follows: F0=B SF (IN LR ) Where F0 represents the result of shallow feature extraction, B SF I represents the shallow feature extraction process. LR Indicates the input image; The feature extraction process of the deep feature extraction module is as follows: F1=B DF (B SF (I LR )) Where F1 represents the deep feature extraction result, B DF B represents the deep feature extraction process. SF I represents the shallow feature extraction process. LR Indicates the input image; The reconstruction process of the reconstruction module is as follows: I SR =B SR (B DF (B SF (I LR )))+B up (I LR ) Among them, I SR B represents the final result. SR B represents the second subpixel convolution process. DF B represents the deep feature extraction process. SF I represents the shallow feature extraction process. LR B represents the input image. up This represents the first subpixel convolution process.

8. The lightweight image super-resolution reconstruction method according to claim 7, characterized in that: The channel attention module compares the channel attention to fuse the feature F. fused The input F for the multi-scale attention mechanism is obtained through processing. s .

9. A lightweight image super-resolution reconstruction system, characterized in that: A lightweight image super-resolution reconstruction method according to any one of claims 1-8 is provided, comprising an input unit, a lightweight super-resolution neural network unit, and an output unit, wherein: The input unit is used to input the low-resolution image to be reconstructed. The lightweight super-resolution neural network unit is used to store the trained lightweight super-resolution neural network; the trained lightweight super-resolution neural network is used to reconstruct the low-resolution image to obtain the reconstructed super-resolution image; The output unit is used to output the reconstructed super-resolution image.

10. A computer system, characterized in that, It includes a memory and a processor, wherein the memory is used to store computer programs / instructions; and the processor is used to execute the computer programs / instructions to implement the lightweight image super-resolution reconstruction method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Remote sensing image super-resolution reconstruction method based on feature information distillation network

    CN115601236A

  • Blueprint separable residual balance distillation super-resolution reconstruction model and method of single image

    CN117036171A