Super-resolution image reconstruction method based on multi-scale large-kernel convolution double-residual neural network

Through the multi-scale large-core convolution dual-residual neural network and the super-resolution image reconstruction method optimized by visual attention mechanism, the problems of low image reconstruction quality and high computational complexity in the prior art are solved, and efficient, clear and real super-resolution image reconstruction is achieved.

CN120298218APending Publication Date: 2025-07-11NANJING TECH UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510360617.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing single-frame super-resolution image reconstruction method has problems in detail blurring and artifact increase, and the calculation complexity is high and the model generalization ability is insufficient, making it difficult to meet the practical application needs.

Method used

A super-resolution image reconstruction method based on a multi-scale large-core convolution dual-residual neural network is adopted, combined with a visual attention mechanism, image features are extracted through multi-scale large-core convolution and dual-residual structure, and image quality is optimized using PSNR and SSIM loss functions, and super-resolution images are generated by upsampling by sub-pixel convolution.

Benefits of technology

It improves the quality and efficiency of image reconstruction, retains more detailed information, reduces computational complexity, enhances the generalization ability and visual effects of the model, and is closer to human visual perception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298218A_ABST
    Figure CN120298218A_ABST
Patent Text Reader

Abstract

The invention discloses a super-resolution image reconstruction method based on a multi-scale large-kernel convolution double-residual neural network, which is suitable for the field of image processing, and comprises the following steps: cutting a data set, inputting a cut original low-resolution image into a preprocessing module, carrying out image normalization and data enhancement operation, and carrying out image reconstruction; generating a preprocessed low-resolution image; the preprocessed low-resolution images form a distorted image block data set, and a training set, a verification set and a test set are formed; according to an existing distorted image block data set, a super-resolution image reconstruction method based on a multi-scale large-kernel convolution double-residual neural network is constructed; and inputting the data set into the constructed multi-scale large-kernel convolution double-residual neural network to extract semantic features, and amplifying a feature map by using an up-sampling module of the model to generate a super-resolution image. According to the method, a multi-scale large-kernel convolution and double-residual structure is introduced, a visual attention mechanism is used in the neural network, the extracted features better conform to human visual perception features, and super-resolution image reconstruction is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a super-resolution image reconstruction method based on a multi-scale large kernel convolutional double residual neural network, and belongs to the field of image processing. Background Art

[0002] Super-Resolution (SR) technology aims to recover high-resolution (HR) images from low-resolution (LR) images to improve the visual quality and information content of the images. According to different input and reference information, super-resolution methods can be divided into single-image super-resolution (SISR) and multi-frame super-resolution (MFSR). Among them, SISR has become the focus of research due to its wide range of application scenarios, such as medical imaging, remote sensing image processing, and video enhancement.

[0003] Traditional SISR methods mainly rely on interpolation algorithms (such as bicubic interpolation), sparse representation-based methods, or dictionary learning, etc. These methods have improved the resolution of images to a certain extent, but due to the lack of effective prior knowledge modeling, they often lead to problems such as blurred details and increased artifacts. In recent years, the development of deep learning technology has provided stronger feature extraction and modeling capabilities for SISR. Super-resolution methods based on convolutional neural networks (CNNs), such as SRCNN, VDSR, ESRGAN, etc., can automatically learn the mapping relationship from low resolution to high resolution through a large amount of data training, thus making significant progress in restoring high-frequency details.

[0004] The super-resolution task faces challenges such as high computational complexity and insufficient model generalization ability. In order to improve the practical application value of super-resolution methods, it is necessary to further study lightweight network structures, optimize the feature extraction process by combining attention mechanisms, and explore loss functions that are more in line with human visual perception to improve the reconstruction quality and efficiency of the model. Summary of the Invention

[0005] The purpose of the present invention is to disclose a super-resolution image reconstruction method based on a multi-scale large kernel convolutional double residual neural network to improve the quality of the reconstructed image. This method combines a visual attention mechanism, fully considers the structural information and color performance of the image, and more accurately predicts the objective quality of the super-resolution reconstructed image. In an embodiment according to the present disclosure, the super-resolution image reconstruction method based on a multi-scale large kernel convolutional double residual neural network includes the following steps:

[0006] Step 1: Crop the dataset, input the original low-resolution image after cropping into the preprocessing module, perform image normalization and data augmentation operations, and generate the preprocessed low-resolution image.

[0007] Step 2: Compose the preprocessed low-resolution images into an image patch dataset, and form a training set, a validation set, and a test set;

[0008] Step 3: Based on the existing image patch dataset, construct a super-resolution image reconstruction method based on a multi-scale large kernel convolutional double residual neural network;

[0009] Step 4: Input the images in Step 2 into the multi-scale large kernel convolutional double residual neural network constructed in Step 3 to extract semantic features, and use the upsampling module of the model to magnify the feature map to generate a super-resolution image.

[0010] For a further technical solution, the specific process of cropping and preprocessing the dataset in Step 1 includes the following steps:

[0011] Step 1.1: Randomly crop the low-resolution image, and the size of the cropped image is H×W, where H and W are fixed sizes.

[0012] Step 1.2: Perform normalization processing on the cropped image, and the normalization formula is as follows:

[0013]

[0014] where I is the original pixel value, I max and I min are the maximum and minimum pixel values of the image, respectively.

[0015] Step 1.3: Perform data augmentation on the normalized image, including operations such as random flipping, rotation, and cropping, to increase the diversity of the data.

[0016] For a further technical solution, Step 2 specifically includes the following steps:

[0017] Step 2.1: Crop the preprocessed low-resolution image without overlapping according to a fixed size to generate small block images, and construct an image patch dataset.

[0018] Step 2.2: Divide the dataset according to the ratio of 60% for the training set, 20% for the validation set, and 20% for the test set to ensure the generalization ability of the model.

[0019] For a further technical solution, Step 3 specifically includes the following steps:

[0020] Step 3.1: The network model consists of multiple modules, mainly including modules such as GroupGLKA, LKAT, and CCM. Each module extracts different features of the image through convolution operations and attention mechanisms. Among them, the GroupGLKA module implements multi-scale large-kernel convolution operations on the input features through multiple convolutional layers, and extracts multi-scale features through convolutional kernels of different sizes. The LKAT module introduces a mechanism based on channel self-attention, which can effectively capture long-range dependence features in the image. The CCM module combines residual connections and convolution operations to further optimize feature extraction and expression.

[0021] Step 3.2: In the GroupGLKA module, the input features are processed by three different convolutional kernels (3×3, 5×5, and 7×7 respectively), and the features of different scales are weighted and fused through residual connections. These convolutional layers contain dilated convolution, which expands the receptive field through dilated convolution while maintaining computational efficiency. The core formula of this module is:

[0022] x = ProjLast(x·a)·Scale + 2·Shortcut

[0023] where x is the input feature, a is the multi-scale feature after convolution, ProjLast is the projection convolution operation, Scale is the learned scaling factor, and Shortcut is the residual connection.

[0024] Step 3.3: In the LKAT module, first, the input features are processed by a 1×1 convolution and non-linearly transformed through the GELU activation function. Then, large-kernel convolutions (7×7, 9×9) are used for feature extraction, and the output is obtained through the last 1×1 convolution. The core formula of this module is:

[0025] x = Conv0(x)·Att(x)·Conv1(x)

[0026] where Conv0 and Conv1 are 1×1 convolutions, and Att is the attention module based on dilated convolution and 1×1 convolution.

[0027] Step 3.4: In the CCM module, the input features are first spatially processed by the first convolutional layer (3×3 convolution), and then non-linearly transformed through the GELU activation function. Then, the processed features are output through the second convolutional layer (1×1 convolution). The core formula of this module is:

[0028] x = Conv2(GELU(Conv1(a))) + Shortcut(a)

[0029] Among them, Conv1 is a 3×3 convolution operation, GELU is an activation function, Conv2 is a 1×1 convolution operation, and Shortcut is a residual connection.

[0030] Step 3.5: The loss function of the super-resolution image reconstruction method includes PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index), which are used to measure the quality difference between the predicted reconstructed image and the real image. PSNR is used to evaluate the reconstruction quality of the image and measures the difference by calculating the maximum signal-to-noise ratio between the predicted image and the real image; SSIM reflects the structural similarity of the image by comparing the structure, brightness, and contrast. The combined use of the two can effectively evaluate the image quality and guide the network optimization. The calculation formulas of PSNR and SSIM are as follows:

[0031] PSNR calculation formula:

[0032]

[0033] Among them, MAXI is the maximum possible pixel value of the image, and MSE is the mean square error.

[0034] SSIM calculation formula:

[0035]

[0036] Among them, μ x , μ y is the average value of the image, σ x , σ y is the variance of the image, σ xy is the covariance, and C1, C2 are constants to avoid the denominator being zero.

[0037] For a further technical solution, step 4 specifically includes the following steps:

[0038] Step 4.1: Use the low-resolution image input in step 2 as the input of the multi-scale large-kernel convolutional double-residual neural network constructed in step 3;

[0039] Step 4.2: The low-resolution image passes through the feature extraction part of the network, which contains multiple multi-scale large-kernel convolutional modules and double-residual blocks. Among them, the multi-scale large-kernel convolutional module extracts different levels of semantic features through large-kernel convolutions of different scales, and the double-residual block enhances the feature transmission ability and alleviates the problem of gradient disappearance;

[0040] Step 4.3: At the output end of the feature extraction part, multiple feature maps are obtained, and these feature maps are weighted by the channel attention and spatial attention mechanisms to further improve the network's attention to key features;

[0041] Step 4.4: The feature map weighted by the attention mechanism enters the upsampling module and is enlarged using sub-pixel convolution operation to generate a super-resolution image;

[0042] Step 4.5: During the network training process, PSNR and SSIM are used as loss functions to measure the similarity between the reconstructed image and the high-resolution image. The PSNR loss is used to minimize the mean square error and improve the pixel-level reconstruction accuracy, while the SSIM loss optimizes the image quality from the perspective of structural information. The network adopts the Adam optimization algorithm, combined with adaptive learning rate and first-order and second-order moment estimates of gradients, to accelerate convergence and improve stability until the loss function converges or reaches the maximum number of iterations;

[0043] Step 4.6: After training is completed, the test set obtained in Step 2 is input into the trained super-resolution model, and a super-resolution image is generated through forward propagation, and evaluation metrics such as PSNR and SSIM are calculated to measure the reconstruction effect of the model.

[0044] The super-resolution image reconstruction method based on a multi-scale large kernel convolution dual residual neural network provided by the present invention has the following beneficial effects:

[0045] 1) The present invention is an end-to-end optimized network model algorithm for super-resolution image reconstruction, which directly learns the mapping relationship between low-resolution images and high-resolution images using a deep learning model. Compared with traditional interpolation algorithms and super-resolution methods based on handcrafted features, the super-resolution model based on deep learning has stronger feature extraction ability and generalization ability, and can achieve high-quality reconstruction effects on different types of image data. Through this optimized network model, super-resolution images can be reconstructed more accurately, more detailed information can be retained, and the clarity and authenticity of the images can be improved.

[0046] 2) The present invention adopts a multi-scale large kernel convolution structure, which effectively expands the receptive field, enabling the network to capture local details and global information of the image at different scales. Traditional super-resolution methods usually rely on deep networks to increase the receptive field, resulting in a large amount of computation. However, the present invention combines large kernel convolution with multi-scale information extraction, reduces the computational complexity while maintaining efficient information fusion, improves the expression ability of the model, and makes the reconstructed image have richer details and more natural visual effects.

[0047] 3) The present invention adopts a dual residual neural network structure, which enhances the gradient flow through residual connections, improves the training stability and feature reconstruction ability of the model. Compared with a single residual network, the dual residual structure can more effectively transmit the information of low-resolution images, reduce information loss, and thus improve the quality of super-resolution reconstruction. At the same time, this structure can better restore high-frequency details, reduce artifacts and blurring phenomena, and make the super-resolved image clearer and more real.

[0048] 4) The present invention adopts a multi-scale large kernel convolution structure, which effectively expands the receptive field, reduces the computational complexity at the same time, and improves the running efficiency of the model. In traditional super-resolution methods, it is usually necessary to increase the receptive field through a deep network, while the present invention combines large kernel convolution with residual learning to achieve efficient information fusion, enabling the network to capture more global and local information at a shallower level, thereby improving the quality of super-resolution images, reducing the consumption of computing resources at the same time, and enhancing the practicality of the model.

[0049] 5) The present invention draws on the characteristics of the human visual system and applies the visual attention mechanism to the convolutional neural network. Extract features sensitive to image quality from the image, and these features can better reflect the quality information of the image. Adopt the saliency mapping algorithm, select the salient image patches as the key areas for evaluation, evaluate the image quality more pertinently, without being interfered by irrelevant areas, and make the image quality evaluation closer to human subjective perception. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 is a flowchart of a super-resolution image reconstruction method based on a multi-scale large kernel convolution double residual neural network according to the present invention;

[0051] Figure 2 is a schematic diagram of the network structure of a super-resolution image reconstruction method based on a multi-scale large kernel convolution double residual neural network according to the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0052] The present invention proposes a super-resolution image reconstruction method, which uses a multi-scale large kernel convolution double residual neural network to achieve image super-resolution reconstruction. In order to more clearly elaborate the purpose, technical solution and advantages of the present invention, the technical solution of the present invention will be described in detail below with reference to the accompanying drawings. It should be noted that the specific embodiments described below are only used to explain the principle of the present invention and are not used to limit the scope of implementation of the present invention.

[0053] The present invention first provides a super-resolution image reconstruction method based on a multi-scale large kernel convolution double residual neural network, specifically referring to Figure 1 , including the following steps:

[0054] Step 1: Crop the data set, input the cropped original low-resolution image into the preprocessing module, perform image normalization and data augmentation operations, and generate the preprocessed low-resolution image.

[0055] Step 2: Compose the preprocessed low-resolution images into an image patch data set to form a training set, a validation set and a test set;

[0056] Step 3: Construct a super-resolution image reconstruction method based on a multi-scale large kernel convolutional double residual neural network according to the existing image patch dataset;

[0057] Step 4: Input the image in Step 2 into the multi-scale large kernel convolutional double residual neural network constructed in Step 3 to extract semantic features, and use the upsampling module of the model to magnify the feature map to generate a super-resolution image.

[0058] As a preferred embodiment of the present invention, in Step 1, the dataset is cropped and preprocessed. Cropping can help remove irrelevant background information in the image, enabling the model to focus more on key regions, thereby improving the efficiency and accuracy of feature extraction. Image normalization can scale the pixel values to a unified range (such as [0, 1] or [-1, 1]), reduce the difference in data distribution, accelerate model convergence, and improve training stability. Data augmentation (such as rotation, flipping, scaling, etc.) can increase the diversity of data, simulate more possible scenarios, effectively prevent model overfitting, and enhance its generalization ability. The combined effect of these operations can significantly improve the performance and robustness of the model.

[0059] Among them, Step 1 specifically includes the following steps:

[0060] Step 1.1: Randomly crop the low-resolution image, and the size of the cropped image is H×W, where H and W are fixed sizes.

[0061] Step 1.2: Perform normalization processing on the cropped image. The normalization formula is as follows:

[0062]

[0063] Among them, I is the original pixel value, I max and I min are the maximum and minimum pixel values of the image, respectively.

[0064] Step 1.3: Perform data augmentation on the normalized image, including operations such as random flipping, rotation, and cropping to increase the diversity of the data.

[0065] As a preferred embodiment of the present invention, in Step 2, the preprocessed low-resolution images are used to form an image patch dataset, which is divided into a training set, a validation set, and a test set, which can accelerate the model training speed;

[0066] Among them, Step 2 specifically includes the following steps:

[0067] Step 2.1: Crop the preprocessed low-resolution images without overlapping according to a fixed size to generate small block images, and construct an image patch dataset.

[0068] Step 2.2: Divide the dataset according to the ratio of 60% for the training set, 20% for the validation set, and 20% for the test set to ensure the generalization ability of the model.

[0069] As a preferred embodiment of the present invention, in step 3, a no-reference color image quality evaluation model based on a multi-scale large-kernel convolutional double residual neural network is constructed. First, consider the structure of each layer of the model. Secondly, select a suitable loss function and optimization algorithm for parameter optimization. Finally, according to the evaluation results of the model, perform hyperparameter tuning, such as adjusting the number of network layers, the number of neurons, etc., to further optimize the performance of the model.

[0070] Among them, step 3 specifically includes the following steps:

[0071] Step 3.1: The network model consists of multiple modules, mainly including modules such as GroupGLKA, LKAT, and CCM. Each module extracts different features of the image through convolutional operations and attention mechanisms. Among them, the GroupGLKA module realizes multi-scale large-kernel convolutional operations on the input features through multiple convolutional layers, and extracts multi-scale features through convolutional kernels of different sizes. The LKAT module introduces a mechanism based on channel self-attention, which can effectively capture long-range dependent features in the image. The CCM module combines residual connections and convolutional operations to further optimize the extraction and expression of features.

[0072] Step 3.2: In the GroupGLKA module, the input features are processed by three different convolutional kernels (3×3, 5×5, and 7×7) respectively, and the features of different scales are weighted and fused through residual connections. These convolutional layers contain dilated convolution to expand the receptive field while maintaining computational efficiency. The core formula of this module is:

[0073] x = ProjLast(x·a)·Scale + 2·Shortcut

[0074] Among them, x is the input feature, a is the multi-scale feature after convolution, ProjLast is the projection convolution operation, Scale is the learned scaling factor, and Shortcut is the residual connection.

[0075] Step 3.3: In the LKAT module, first perform 1×1 convolution processing on the input features and perform non-linear transformation through the GELU activation function. Then, use large-kernel convolution (7×7, 9×9) for feature extraction and output through the last 1×1 convolution. The core formula of this module is:

[0076] x = Conv0(x)·Att(x)·Conv1(x)

[0077] Among them, Conv0 and Conv1 are 1×1 convolutions, and Att is an attention module based on dilated convolution and 1×1 convolution.

[0078] Step 3.4: In the CCM module, the input features are first spatially processed by the first convolutional layer (3×3 convolution), and then non-linearly transformed through the GELU activation function. Subsequently, the processed features are output through the second convolutional layer (1×1 convolution). The core formula of this module is:

[0079] x = Conv2(GELU(Conv1(a))) + Shortcut(a)

[0080] Among them, Conv1 is a 3×3 convolution operation, GELU is an activation function, Conv2 is a 1×1 convolution operation, and Shortcut is a residual connection.

[0081] Step 3.5: The loss functions of the super-resolution image reconstruction method include PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index), which are used to measure the quality difference between the predicted reconstructed image and the real image. PSNR is used to evaluate the reconstruction quality of the image, and its difference is measured by calculating the maximum signal-to-noise ratio between the predicted image and the real image; SSIM reflects the structural similarity of the image through the comparison of structure, brightness, and contrast. The combined use of the two can effectively evaluate the image quality and guide network optimization. The calculation formulas of PSNR and SSIM are as follows:

[0082] PSNR calculation formula:

[0083]

[0084] Among them, MAX I is the maximum possible pixel value of the image, and MSE is the mean square error.

[0085] SSIM calculation formula:

[0086]

[0087] Among them, μ x , μ y is the average value of the image, σ x , σ y is the variance of the image, σ xy is the covariance, and C1, C2 are constants to avoid the denominator being zero.

[0088] As a preferred embodiment of the present invention, in step 4, the image in step 2 is input into the multi-scale large kernel convolutional double residual neural network constructed in step 3 to extract semantic features, and the upsampling module of the model is used to magnify the feature map to generate a super-resolution image.

[0089] Among them, step 4 specifically includes the following steps:

[0090] Step 4.1: Use the low-resolution image input in step 2 as the input of the multi-scale large-kernel convolutional double residual neural network constructed in step 3;

[0091] Step 4.2: The low-resolution image passes through the feature extraction part of the network, which contains multiple multi-scale large-kernel convolutional modules and double residual blocks. Among them, the multi-scale large-kernel convolutional modules extract semantic features at different levels through large-kernel convolutions of different scales, and the double residual blocks enhance the feature transfer ability and alleviate the problem of gradient disappearance;

[0092] Step 4.3: At the output end of the feature extraction part, multiple feature maps are obtained, and these feature maps are weighted by channel attention and spatial attention mechanisms to further improve the network's attention to key features;

[0093] Step 4.4: The feature maps weighted by the attention mechanism enter the upsampling module and are enlarged by sub-pixel convolution operation to generate a super-resolution image;

[0094] During the network training process, PSNR and SSIM are used as loss functions to measure the similarity between the reconstructed image and the high-resolution image. The PSNR loss is used to minimize the mean square error and improve the pixel-level reconstruction accuracy, while the SSIM loss optimizes the image quality from the perspective of structural information; The network adopts the Adam optimization algorithm, combined with adaptive learning rate and first-order and second-order moment estimates of gradients, to accelerate convergence and improve stability until the loss function converges or reaches the maximum number of iterations;

[0095] After training is completed, the test set obtained in step 2 is input into the trained super-resolution model, and a super-resolution image is generated through forward propagation, and evaluation indexes such as PSNR and SSIM are calculated to measure the reconstruction effect of the model.

[0096] In the above description, combination methods of various technical features are mentioned, but all possible combinations are not exhaustively listed. However, it should be emphasized that as long as the combination of these technical features is reasonable in implementation and there are no contradictions between them, they should be regarded as within the scope covered by this specification.

Claims

1. A super-resolution image reconstruction method based on a multi-scale large kernel convolutional double residual neural network, characterized in that The method includes the following steps: Step 1: Crop the dataset, and input the cropped original low-resolution image into a preprocessing module for image normalization and data augmentation operations to generate a preprocessed low-resolution image. Step 2: Compose the preprocessed low-resolution images into an image patch dataset to form a training set, a validation set, and a test set. Step 3: Based on the existing image patch dataset, construct a super-resolution image reconstruction method based on a multi-scale large kernel convolutional double residual neural network. Step 4: Input the images in Step 2 into the multi-scale large kernel convolutional double residual neural network constructed in Step 3 to extract semantic features, and use the upsampling module of the model to magnify the feature map to generate a super-resolution image.

2. The super-resolution image reconstruction method based on a multi-scale large kernel convolutional double residual neural network according to claim 1, wherein The specific process of cropping and preprocessing the dataset in Step 1 includes the following steps: Step 1.1: Randomly crop the low-resolution image, and the size of the cropped image is H×W, where H and W are fixed sizes. Step 1.2: Perform normalization processing on the cropped image, and the normalization formula is as follows: where I is the original pixel value, I max and I min are the maximum and minimum pixel values of the image, respectively. Step 1.3: Perform data augmentation on the normalized image, including operations such as random flipping, rotation, and cropping to increase data diversity.

3. The super-resolution image reconstruction method based on the multi-scale large kernel convolution double residual neural network according to claim 1, wherein, Step 2 specifically includes the following steps: Step 2.1: Crop the preprocessed low-resolution images non-overlappingly according to a fixed size to generate small images, and construct an image patch dataset. Step 2.2: Divide the dataset according to the ratio of 60% for the training set, 20% for the validation set, and 20% for the test set to ensure the generalization ability of the model.

4. The image processing network based on the multi-scale large kernel convolution attention mechanism according to claim 1, wherein, Step 3 specifically includes the following steps: Step 3.1: The network model consists of multiple modules, mainly including modules such as GroupGLKA, LKAT, and CCM. Each module extracts different features of the image through convolutional operations and attention mechanisms. Among them, the GroupGLKA module realizes multi-scale large kernel convolutional operations on the input features through multiple convolutional layers, and extracts multi-scale features through convolutional kernels of different sizes. The LKAT module introduces a mechanism based on channel self-attention, which can effectively capture long-range dependent features in the image. The CCM module combines residual connections and convolutional operations to further optimize feature extraction and expression. Step 3.2: In the GroupGLKA module, the input features are processed by three different convolutional kernels (3×3, 5×5, and 7×7 respectively), and different-scale features are weighted and fused through residual connections. These convolutional layers contain dilated convolution to expand the receptive field while maintaining computational efficiency. The core formula of this module is: x = ProjLast(x·a)·Scale + 2·Shortcut where x is the input feature, a is the multi-scale feature after convolution, ProjLast is the projection convolution operation, Scale is the learned scaling factor, and Shortcut is the residual connection. Step 3.3: In the LKAT module, first, the input features are processed by 1×1 convolution and undergo non-linear transformation through the GELU activation function. Then, large kernel convolutions (7×7, 9×9) are used for feature extraction, and the output is obtained through the last 1×1 convolution. The core formula of this module is: x = Conv0(x)·Att(x)·Conv1(x). Among them, Conv0 and Conv1 are 1×1 convolutions, and Att is an attention module based on dilated convolution and 1×1 convolution. Step 3.4: In the CCM module, the input features are first spatially processed by the first convolutional layer (3×3 convolution), and then undergo non-linear transformation through the GELU activation function. Subsequently, the processed features are output through the second convolutional layer (1×1 convolution). The core formula of this module is: x = Conv2(GELU(Conv1(a))) + Shortcut(a) Among them, Conv1 is a 3×3 convolution operation, GELU is the activation function, Conv2 is a 1×1 convolution operation, and Shortcut is the residual connection. The loss function of the super-resolution image reconstruction method includes PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index), which are used to measure the quality difference between the predicted reconstructed image and the real image. PSNR is used to evaluate the reconstruction quality of the image, and its difference is measured by calculating the maximum signal-to-noise ratio between the predicted image and the real image; SSIM reflects the structural similarity of the image by comparing structure, brightness, and contrast. The combined use of the two can effectively evaluate the image quality and guide network optimization. The calculation formulas of PSNR and SSIM are as follows: PSNR calculation formula: where MAX I is the maximum possible pixel value of the image, and MSE is the mean squared error. SSIM calculation formula: where μ x , μ y is the average value of the image, σ x , σ y is the variance of the image, σ xy is the covariance, and C1, C2 are constants to avoid a zero denominator.

5. The super-resolution method based on the multi-scale large kernel convolutional double residual neural network according to claim 1, wherein The specific steps of Step 4 are as follows: Step 4.1: Use the low-resolution image input in Step 2 as the input of the multi-scale large kernel convolutional double residual neural network constructed in Step 3; Step 4.2: The low-resolution image passes through the feature extraction part of the network, which includes multiple multi-scale large kernel convolutional modules and double residual blocks. Among them, the multi-scale large kernel convolutional module extracts semantic features at different levels through large kernel convolutions of different scales, and the double residual block enhances the feature transmission ability and alleviates the problem of gradient disappearance; Step 4.3: At the output end of the feature extraction part, multiple feature maps are obtained, and these feature maps are weighted by the channel attention and spatial attention mechanisms to further improve the network's attention to key features; Step 4.4: The feature maps weighted by the attention mechanism enter the upsampling module and are magnified using sub-pixel convolution operation to generate the super-resolution image; Step 4.5: During the network training process, PSNR and SSIM are used as loss functions to measure the similarity between the reconstructed image and the high-resolution image. The PSNR loss is used to minimize the mean square error and improve the pixel-level reconstruction accuracy, while the SSIM loss optimizes the image quality from the perspective of structural information. The network adopts the Adam optimization algorithm, combines the adaptive learning rate and the first-order and second-order moment estimates of the gradient, accelerates convergence and improves stability until the loss function converges or reaches the maximum number of iterations. Step 4.6: After training is completed, the test set obtained in Step 2 is input into the trained super-resolution model, and the super-resolution image is generated through forward propagation, and evaluation metrics such as PSNR and SSIM are calculated to measure the reconstruction effect of the model.

Citation Information

Cited By

  • Image super-resolution reconstruction method based on multi-scale macronuclear distillation attention network

    CN120510037A

  • Neural network optimization method based on visual state space model for bridge diseases

    CN121052293A

  • Terahertz polyethylene pipeline hot melting joint defect super-resolution imaging method based on CBAM-ESRGAN

    CN121169689A

  • Super-resolution high-definition image generation method based on optimized generative neural network

    CN121639837A