Super-resolution image reconstruction method based on multi-scale large separable kernel convolutional neural network
By optimizing super-resolution image reconstruction using multi-scale large separable kernel convolutional neural networks and visual attention mechanisms, the problems of detail blurring and artifact increase in traditional methods are solved, achieving efficient and clear image reconstruction results.
Patent Information
- Application Number
- CN202510863573.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-10-28
AI Technical Summary
Existing single-frame super-resolution image reconstruction methods have problems such as blurred details and increased artifacts, and they are computationally complex and lack model generalization ability, making it difficult to meet the needs of practical applications.
A super-resolution image reconstruction method based on multi-scale large separable kernel convolutional neural network is adopted. Combined with the visual attention mechanism, image features are extracted through multi-scale large separable kernel convolution and blueprint separable convolution structure, and image quality is optimized using PSNR and SSIM loss functions. Subpixel convolution is then used to generate super-resolution images.
It improves the quality and efficiency of image reconstruction, preserves more detailed information, reduces computational complexity, and generates clearer and more realistic super-resolution images that conform to human visual perception.
Smart Images

Figure CN120852167A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a super-resolution image reconstruction method based on a multi-scale large-kernel convolutional neural network, belonging to the field of image processing. Background Technology
[0002] Super-resolution (SR) technology aims to recover high-resolution (HR) images from low-resolution (LR) images, thereby improving the visual quality and information content of the images. Depending on the input and reference information, super-resolution methods can be divided into single-image super-resolution (SISR) and multi-frame super-resolution (MFSR). Among these, SISR has become a research focus due to its wide range of applications, such as medical imaging, remote sensing image processing, and video enhancement.
[0003] Traditional SISR methods primarily rely on interpolation algorithms (such as bicubic interpolation), sparse representation-based methods, or dictionary learning. While these methods improve image resolution to some extent, they often lead to problems such as blurred details and increased artifacts due to a lack of effective prior knowledge modeling. In recent years, the development of deep learning technology has provided SISR with stronger feature extraction and modeling capabilities. Super-resolution methods based on convolutional neural networks (CNNs), such as SRCNN, VDSR, and ESRGAN, can automatically learn the mapping relationship from low resolution to high resolution through training on large amounts of data, thus achieving significant progress in restoring high-frequency details.
[0004] Super-resolution tasks face challenges such as high computational complexity and insufficient model generalization ability. To improve the practical application value of super-resolution methods, further research is needed on lightweight network structures, optimizing the feature extraction process by incorporating attention mechanisms, and exploring loss functions that better align with human visual perception, in order to improve the reconstruction quality and efficiency of the models. Summary of the Invention
[0005] The purpose of this invention is to disclose a super-resolution image reconstruction method based on a multi-scale large separable kernel convolutional neural network to improve the quality of reconstructed images. This method incorporates a visual attention mechanism, fully considering the structural information and color representation of the image to more accurately predict the objective quality of the super-resolution reconstructed image. In embodiments according to this disclosure, the super-resolution image reconstruction method based on a multi-scale large separable kernel convolutional neural network includes the following steps:
[0006] Step 1: Crop the dataset, input the cropped original low-resolution image into the preprocessing module, perform image normalization and data augmentation operations to generate the preprocessed low-resolution image.
[0007] Step 2: Compile the preprocessed low-resolution images into an image patch dataset, which forms the training set, validation set, and test set;
[0008] Step 3: Based on the existing image patch dataset, construct a super-resolution image reconstruction method based on a multi-scale large separable kernel convolutional neural network;
[0009] Step 4: Input the image from Step 2 into the multi-scale large separable kernel convolutional neural network constructed in Step 3 to extract semantic features, and use the upsampling module of the model to enlarge the feature map to generate a super-resolution image.
[0010] A further technical solution involves the following steps in step 1: cropping and preprocessing the dataset.
[0011] Step 1.1: Randomly crop the low-resolution image. The size of the cropped image is H×W, where H and W are fixed dimensions.
[0012] Step 1.2: Normalize the cropped image. The normalization formula is as follows:
[0013]
[0014] Where I is the original pixel value, I max and I min These are the maximum and minimum pixel values of the image, respectively.
[0015] Step 1.3: Perform data augmentation on the normalized image, including random flipping, rotation, cropping, etc., to increase the diversity of the data.
[0016] Further technical solutions, step 2 specifically includes the following steps:
[0017] Step 2.1: Crop the preprocessed low-resolution image into small patches of images by a fixed size without overlap, and construct an image patch dataset.
[0018] Step 2.2: Divide the dataset into three parts: 60% for training, 20% for validation, and 20% for testing, to ensure the model's generalization ability.
[0019] Further technical solutions, step 3 specifically includes the following steps:
[0020] Step 3.1: This network model consists of multiple modules, mainly including GroupGLSKA, LKAT, and BSCCM. Each module extracts different features from the image through convolutional operations and attention mechanisms. Specifically, the GroupGLSKA module implements multi-scale, large separable kernel convolutional operations on the input features through multiple convolutional layers, extracting multi-scale features using convolutional kernels of different sizes. The LKAT module introduces a channel-based self-attention mechanism, which can effectively capture long-range dependent features in the image. The BSCCM module combines blueprint-based separable convolutional operations to further optimize feature extraction and representation.
[0021] Step 3.2: In the GroupGLSKA module, the input features are processed by three different convolutional kernels (3×3, 5×5, and 7×7, respectively), and features of different scales are weighted and fused through residual connections. These convolutional layers include dilated convolution, which expands the receptive field through dilated convolution while maintaining computational efficiency. The core formula of this module is:
[0022] x=ProjLast(x·a)·Scale+2·Shortcut
[0023] Where x is the input feature, a is the multi-scale feature after convolution, ProjLast is the projective convolution operation, Scale is the learning scaling factor, and Shortcut is the residual connection.
[0024] Step 3.3: In the LKAT module, the input features are first processed by a 1×1 convolution, followed by a non-linear transformation using the GELU activation function. Then, large kernel convolutions (7×7, 9×9) are used for feature extraction, and the output is obtained through a final 1×1 convolution. The core formula of this module is:
[0025] x=Conv0(x)·Att(x)·Conv1(x)
[0026] Here, Conv0 and Conv1 are 1×1 convolutions, and Att is an attention module based on dilated convolution and 1×1 convolution.
[0027] Step 3.4: In the BSCCM module, the input features are first convolved using BSConv, followed by the Gelu activation function. Then, a 1×1 convolution is used to output the processed features. The core formula of this module is:
[0028] x = Conv 1×1 (GELU(BSConvU(z)))
[0029] Where BSConv is the Blueprint Convolution operation, and Conv... 1x1It is a 1×1 convolution.
[0030] Step 3.5: The loss function of the super-resolution image reconstruction method includes PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index), used to measure the quality difference between the predicted reconstructed image and the real image. PSNR is used to evaluate the reconstruction quality of the image by calculating the maximum signal-to-noise ratio between the predicted and real images; SSIM reflects the structural similarity of the images by comparing structure, brightness, and contrast. The combined use of both can effectively evaluate image quality and guide network optimization. The calculation formulas for PSNR and SSIM are as follows:
[0031] PSNR calculation formula:
[0032]
[0033] Among them, MAX I is the maximum possible pixel value of the image, and MSE is the mean square error.
[0034] SSIM calculation formula:
[0035]
[0036] where μ x μ y σ is the average value of the image. x , σ y Let σ be the variance of the image. xy For covariance, C1 and C2 are constants to avoid denominators of 0.
[0037] Further technical solutions, step 4 specifically includes the following steps:
[0038] Step 4.1: Use the low-resolution image input in Step 2 as the input to the multi-scale large separable kernel convolutional neural network constructed in Step 3;
[0039] Step 4.2: The low-resolution image passes through the feature extraction part of the network. This part includes multiple multi-scale large separable kernel convolutional modules and blueprint convolutional blocks. The multi-scale large separable kernel convolutional modules extract semantic features at different levels through large separable kernel convolution at different scales, while the blueprint convolutional blocks enhance feature transfer capability and alleviate the gradient vanishing problem.
[0040] Step 4.3: At the output of the feature extraction part, multiple feature maps are obtained, and these feature maps are weighted by channel attention and spatial attention mechanisms to further enhance the network's attention to key features;
[0041] Step 4.4: The attention-weighted feature map enters the upsampling module and is enlarged using sub-pixel convolution to generate a super-resolution image;
[0042] Step 4.5: During network training, PSNR and SSIM are used as loss functions to measure the similarity between the reconstructed image and the high-resolution image. PSNR loss is used to minimize the mean square error and improve pixel-level reconstruction accuracy, while SSIM loss optimizes image quality from the perspective of structural information. The network adopts the Adam optimization algorithm, which combines adaptive learning rate and gradient first and second moment estimation to accelerate convergence and improve stability until the loss function converges or the maximum number of iterations is reached.
[0043] Step 4.6: After training is completed, the test set obtained in step 2 is input into the trained super-resolution model, and super-resolution images are generated through forward propagation. Evaluation metrics such as PSNR and SSIM are calculated to measure the reconstruction effect of the model.
[0044] The present invention provides a super-resolution image reconstruction method based on a multi-scale large separable kernel convolutional neural network, which has the following beneficial effects:
[0045] 1) This invention is an end-to-end optimized network model algorithm for super-resolution image reconstruction, which utilizes a deep learning model to directly learn the mapping relationship between low-resolution and high-resolution images. Compared to traditional interpolation algorithms and super-resolution methods based on handcrafted features, the deep learning-based super-resolution model has stronger feature extraction and generalization capabilities, enabling high-quality reconstruction results on different types of image data. Through this optimized network model, super-resolution images can be reconstructed more accurately, preserving more detail information and improving image clarity and realism.
[0046] 2) This invention employs a multi-scale, large separable kernel convolutional structure, effectively expanding the receptive field and enabling the network to capture local details and global information of images at different scales. Traditional super-resolution methods typically rely on deep networks to increase the receptive field, resulting in high computational costs. In contrast, this invention combines large separable kernel convolution with multi-scale information extraction, reducing computational complexity while maintaining efficient information fusion, thus improving the model's expressive power and resulting in reconstructed images with richer details and more natural visual effects.
[0047] 3) This invention employs a blueprint-separable convolutional network structure. By dividing the convolution operation into two stages—inter-channel modeling and spatial modeling—it effectively reduces the number of model parameters and computational complexity while maintaining strong feature representation capabilities. Compared to traditional convolutional structures, blueprint-separable convolution can more efficiently extract and fuse key information from low-resolution images, reducing redundant computation and information loss, thereby improving the quality of super-resolution reconstruction. Furthermore, this structure helps restore high-frequency details in images, reduces artifacts and blurring, and makes the generated super-resolution image clearer and more realistic.
[0048] 4) This invention employs a multi-scale, large separable kernel convolutional structure, effectively expanding the receptive field while reducing computational complexity and improving model efficiency. Traditional super-resolution methods typically require deep networks to increase the receptive field. This invention, however, combines large separable kernel convolution with residual learning to achieve efficient information fusion, enabling the network to capture more global and local information at a shallower level. This improves the quality of super-resolution images while reducing computational resource consumption and enhancing the model's practicality.
[0049] 5) This invention draws on the characteristics of the human visual system, applying the visual attention mechanism to convolutional neural networks. It extracts image quality-sensitive features from images, which better reflect image quality information. Employing a saliency mapping algorithm, it selects salient image patches as key evaluation areas, enabling more targeted image quality assessment, free from interference from irrelevant areas, and making image quality evaluation closer to human subjective perception. Attached Figure Description
[0050] Figure 1 The flowchart is a super-resolution image reconstruction method based on a multi-scale large separable kernel convolutional neural network, which is the subject of this invention.
[0051] Figure 2 This is a schematic diagram of the network structure of the super-resolution image reconstruction method based on a multi-scale large separable kernel convolutional neural network, which is involved in this invention. Detailed Implementation
[0052] This invention proposes a super-resolution image reconstruction method that utilizes a multi-scale, large-separability kernel convolutional neural network to achieve image super-resolution reconstruction. To more clearly illustrate the purpose, technical solution, and advantages of this invention, the technical solution will be described in detail below with reference to the accompanying drawings. It should be noted that the specific embodiments described below are only for explaining the principles of this invention and are not intended to limit the scope of implementation of this invention.
[0053] This invention first provides a super-resolution image reconstruction method based on a multi-scale large separable kernel convolutional neural network, specifically referring to... Figure 1 This includes the following steps:
[0054] Step 1: Crop the dataset, input the cropped original low-resolution image into the preprocessing module, perform image normalization and data augmentation operations to generate the preprocessed low-resolution image.
[0055] Step 2: Compile the preprocessed low-resolution images into an image patch dataset, which forms the training set, validation set, and test set;
[0056] Step 3: Based on the existing image patch dataset, construct a super-resolution image reconstruction method based on a multi-scale large separable kernel convolutional neural network;
[0057] Step 4: Input the image from Step 2 into the multi-scale large separable kernel convolutional neural network constructed in Step 3 to extract semantic features, and use the upsampling module of the model to enlarge the feature map to generate a super-resolution image.
[0058] In a preferred embodiment of the present invention, step 1 involves cropping and preprocessing the dataset. Cropping helps remove irrelevant background information from the image, allowing the model to focus more on key regions, thereby improving the efficiency and accuracy of feature extraction. Image normalization scales pixel values to a uniform range (e.g., [0, 1] or [-1, 1]), reducing data distribution differences, accelerating model convergence, and improving training stability. Data augmentation (e.g., rotation, flipping, scaling, etc.) increases data diversity, simulates more possible scenarios, effectively prevents model overfitting, and improves its generalization ability. These operations work together to significantly improve the model's performance and robustness.
[0059] Step 1 specifically includes the following steps:
[0060] Step 1.1: Randomly crop the low-resolution image. The size of the cropped image is H×W, where H and W are fixed dimensions.
[0061] Step 1.2: Normalize the cropped image. The normalization formula is as follows:
[0062]
[0063] Where I is the original pixel value, I max and I min These are the maximum and minimum pixel values of the image, respectively.
[0064] Step 1.3: Perform data augmentation on the normalized image, including random flipping, rotation, cropping, etc., to increase the diversity of the data.
[0065] In a preferred embodiment of the present invention, in step 2, the preprocessed low-resolution images are used to form an image patch dataset, which constitutes a training set, a validation set, and a test set, thereby accelerating the model training speed.
[0066] Step 2 specifically includes the following steps:
[0067] Step 2.1: Crop the preprocessed low-resolution image into small patches of images by a fixed size without overlap, and construct an image patch dataset.
[0068] Step 2.2: Divide the dataset into three parts: 60% for training, 20% for validation, and 20% for testing, to ensure the model's generalization ability.
[0069] In a preferred embodiment of the present invention, in step 3, a no-reference color image quality assessment model based on a multi-scale large separable kernel convolutional neural network is constructed. First, the structure of each layer of the model is considered. Then, a suitable loss function and optimization algorithm are selected for parameter optimization. Finally, based on the evaluation results of the model, hyperparameter tuning is performed, such as adjusting the number of network layers and the number of neurons, to further optimize the performance of the model.
[0070] Step 3 specifically includes the following steps:
[0071] Step 3.1: This network model consists of multiple modules, mainly including GroupGLSKA, LKAT, and BSCCM. Each module extracts different features from the image through convolutional operations and attention mechanisms. Specifically, the GroupGLSKA module implements multi-scale, large separable kernel convolutional operations on the input features through multiple convolutional layers, extracting multi-scale features using convolutional kernels of different sizes. The LKAT module introduces a channel-based self-attention mechanism, which can effectively capture long-range dependent features in the image. The BSCCM module combines blueprint-based separable convolutional operations to further optimize feature extraction and representation.
[0072] Step 3.2: In the GroupGLSKA module, the input features are processed by three different convolutional kernels (3×3, 5×5, and 7×7, respectively), and features of different scales are weighted and fused through residual connections. These convolutional layers include dilated convolution, which expands the receptive field through dilated convolution while maintaining computational efficiency. The core formula of this module is:
[0073] x=ProjLast(x·a)·Scale+2·Shortcut
[0074] Where x is the input feature, a is the multi-scale feature after convolution, ProjLast is the projective convolution operation, Scale is the learning scaling factor, and Shortcut is the residual connection.
[0075] Step 3.3: In the LKAT module, the input features are first processed by a 1×1 convolution, followed by a non-linear transformation using the GELU activation function. Then, large kernel convolutions (7×7, 9×9) are used for feature extraction, and the output is obtained through a final 1×1 convolution. The core formula of this module is:
[0076] x=Conv0(x)·Att(x)·Conv1(x)
[0077] Here, Conv0 and Conv1 are 1×1 convolutions, and Att is an attention module based on dilated convolution and 1×1 convolution.
[0078] Step 3.4: In the BSCCM module, the input features are first convolved using BSConv, followed by the Gelu activation function. Then, a 1×1 convolution is used to output the processed features. The core formula of this module is:
[0079] x = Conv 1×1 (GELU(BSConvU(x)))
[0080] Where BSConv is the Blueprint Convolution operation, and Conv... 1x1 It is a 1×1 convolution.
[0081] Step 3.5: The loss function of the super-resolution image reconstruction method includes PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index), used to measure the quality difference between the predicted reconstructed image and the real image. PSNR is used to evaluate the reconstruction quality of the image by calculating the maximum signal-to-noise ratio between the predicted and real images; SSIM reflects the structural similarity of the images by comparing structure, brightness, and contrast. The combined use of both can effectively evaluate image quality and guide network optimization. The calculation formulas for PSNR and SSIM are as follows:
[0082] PSNR calculation formula:
[0083]
[0084] Where MAX1 is the maximum possible pixel value of the image, and MSE is the mean square error.
[0085] SSIM calculation formula:
[0086]
[0087] where μx μ y σ is the average value of the image. x , σ y Let σ be the variance of the image. xy For covariance, C1 and C2 are constants to avoid denominators of 0.
[0088] In a preferred embodiment of the present invention, in step 4, the image from step 2 is input into the multi-scale large separable kernel convolutional neural network constructed in step 3 to extract semantic features, and the feature map is enlarged by the upsampling module of the model to generate a super-resolution image.
[0089] Step 4 specifically includes the following steps:
[0090] Step 4.1: Use the low-resolution image input in Step 2 as the input to the multi-scale large separable kernel convolutional neural network constructed in Step 3;
[0091] Step 4.2: The low-resolution image passes through the feature extraction part of the network. This part includes multiple multi-scale large separable kernel convolutional modules and blueprint convolutional blocks. The multi-scale large separable kernel convolutional modules extract semantic features at different levels through large separable kernel convolution at different scales, while the blueprint convolutional blocks enhance feature transfer capability and alleviate the gradient vanishing problem.
[0092] Step 4.3: At the output of the feature extraction part, multiple feature maps are obtained, and these feature maps are weighted by channel attention and spatial attention mechanisms to further enhance the network's attention to key features;
[0093] Step 4.4: The attention-weighted feature map enters the upsampling module and is enlarged using sub-pixel convolution to generate a super-resolution image;
[0094] Step 4.5: During network training, PSNR and SSIM are used as loss functions to measure the similarity between the reconstructed image and the high-resolution image. PSNR loss is used to minimize the mean square error and improve pixel-level reconstruction accuracy, while SSIM loss optimizes image quality from the perspective of structural information. The network adopts the Adam optimization algorithm, which combines adaptive learning rate and gradient first and second moment estimation to accelerate convergence and improve stability until the loss function converges or the maximum number of iterations is reached.
[0095] Step 4.6: After training is completed, the test set obtained in step 2 is input into the trained super-resolution model, and super-resolution images are generated through forward propagation. Evaluation metrics such as PSNR and SSIM are calculated to measure the reconstruction effect of the model.
[0096] The above description mentions various methods of combining technical features, but does not list all possible combinations. However, it should be emphasized that as long as the combination of these technical features is reasonable in implementation and does not contradict each other, it should be considered within the scope of this specification.
Claims
1. A super-resolution image reconstruction method based on a multi-scale, large-separable kernel convolutional neural network, characterized in that, The method includes the following steps: Step 1: Crop the dataset, input the cropped original low-resolution image into the preprocessing module, perform image normalization and data augmentation operations to generate the preprocessed low-resolution image. Step 2: Compile the preprocessed low-resolution images into an image patch dataset, which forms the training set, validation set, and test set; Step 3: Based on the existing image patch dataset, construct a super-resolution image reconstruction method based on a multi-scale large separable kernel convolutional neural network; Step 4: Input the image from Step 2 into the multi-scale large separable kernel convolutional neural network constructed in Step 3 to extract semantic features, and use the upsampling module of the model to enlarge the feature map to generate a super-resolution image.
2. The super-resolution image reconstruction method based on a multi-scale large separable kernel convolutional neural network according to claim 1, characterized in that, The specific process of cropping and preprocessing the dataset in step 1 includes the following steps: Step 1.1: Randomly crop the low-resolution image. The size of the cropped image is H×W, where H and W are fixed dimensions. Step 1.2: Normalize the cropped image. The normalization formula is as follows: Where I is the original pixel value, I max and I min These are the maximum and minimum pixel values of the image, respectively. Step 1.3: Perform data augmentation on the normalized image, including random flipping, rotation, cropping, etc., to increase the diversity of the data.
3. The super-resolution image reconstruction method based on a multi-scale large separable kernel convolutional neural network according to claim 1, characterized in that, Step 2 specifically includes the following steps: Step 2.1: Crop the preprocessed low-resolution image into small patches of images by a fixed size without overlap, and construct an image patch dataset. Step 2.2: Divide the dataset into three parts: 60% for training, 20% for validation, and 20% for testing, to ensure the model's generalization ability.
4. The image processing network based on the multi-scale large separable kernel convolutional attention mechanism according to claim 1, characterized in that, Step 3 specifically includes the following steps: Step 3.1: This network model consists of multiple modules, mainly including GroupGLSKA, LKAT, and BSCCM. Each module extracts different features from the image through convolution operations and attention mechanisms. Specifically, the GroupGLSKA module implements multi-scale, large separable kernel convolution operations on the input features through multiple convolutional layers, extracting multi-scale features using convolutional kernels of different sizes. The LKAT module introduces a channel-based self-attention mechanism, effectively capturing long-range dependent features in the image. The BSCCM module combines blueprint separable convolution operations to further optimize feature extraction and representation. Step 3.2: In the GroupGLSKA module, the input features are processed by three different convolutional kernels (3×3, 5×5, and 7×7, respectively), and features of different scales are weighted and fused through residual connections. These convolutional layers include dilated convolution, which expands the receptive field through dilated convolution while maintaining computational efficiency. The core formula of this module is: x=ProjLast(x·a)·Scale+2·Shortcut Where x is the input feature, a is the multi-scale feature after convolution, ProjLast is the projective convolution operation, Scale is the learning scaling factor, and Shortcut is the residual connection. Step 3.3: In the LKAT module, the input features are first processed by a 1×1 convolution, followed by a non-linear transformation using the GELU activation function. Then, large kernel convolutions (7×7, 9×9) are used for feature extraction, and the output is obtained through a final 1×1 convolution. The core formula of this module is: x=Conv0(x)·Att(x)·Conv1(x) Here, Conv0 and Conv1 are 1×1 convolutions, and Att is an attention module based on dilated convolution and 1×1 convolution. Step 3.4: In the BSCCM module, the input features are first convolved using BSConv, followed by the Gelu activation function. Then, a 1×1 convolution is used to output the processed features. The core formula of this module is: x=Conv 1×1 (GELU(BSConvU(x))) Here, BSConv is the blueprint convolution operation, and Conv1x1 is a 1×1 convolution. Step 3.5: The loss function of the super-resolution image reconstruction method includes PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index), used to measure the quality difference between the predicted reconstructed image and the real image. PSNR is used to evaluate the reconstruction quality of the image by calculating the maximum signal-to-noise ratio between the predicted and real images; SSIM reflects the structural similarity of the images by comparing structure, brightness, and contrast. The combined use of both can effectively evaluate image quality and guide network optimization. The calculation formulas for PSNR and SSIM are as follows: PSNR calculation formula: Where MAXI is the maximum possible pixel value of the image, and MSE is the mean square error. SSIM calculation formula: where μ x μ y σ is the average value of the image. x , σ y Let σ be the variance of the image. xy For covariance, C1 and C2 are constants to avoid denominators of 0.
5. The super-resolution method based on a multi-scale, highly separable kernel convolutional neural network according to claim 1, characterized in that, Step 4 specifically includes the following steps: Step 4.1: Use the low-resolution image input in Step 2 as the input to the multi-scale large separable kernel convolutional neural network constructed in Step 3; Step 4.2: The low-resolution image passes through the feature extraction part of the network. This part includes multiple multi-scale large separable kernel convolutional modules and blueprint convolutional blocks. The multi-scale large separable kernel convolutional modules extract semantic features at different levels through large separable kernel convolution at different scales, while the blueprint convolutional blocks enhance feature transfer capability and alleviate the gradient vanishing problem. Step 4.3: At the output of the feature extraction part, multiple feature maps are obtained, and these feature maps are weighted by channel attention and spatial attention mechanisms to further enhance the network's attention to key features; Step 4.4: The feature map weighted by the attention mechanism enters the upsampling module, and is enlarged by sub-pixel convolution to generate a super-resolution image; Step 4.5: During network training, PSNR and SSIM are used as loss functions to measure the similarity between the reconstructed image and the high-resolution image. PSNR loss is used to minimize the mean square error and improve pixel-level reconstruction accuracy, while SSIM loss optimizes image quality from the perspective of structural information. The network adopts the Adam optimization algorithm, which combines adaptive learning rate and gradient first and second moment estimation to accelerate convergence and improve stability until the loss function converges or the maximum number of iterations is reached. Step 4.6: After training is completed, the test set obtained in step 2 is input into the trained super-resolution model, and super-resolution images are generated through forward propagation. Evaluation metrics such as PSNR and SSIM are calculated to measure the reconstruction effect of the model.
Citation Information
Cited By
Sagolian environment video super-resolution method and system, electronic equipment and computer readable storage medium
CN121788357A