Image super-resolution reconstruction method based on attention mechanism and two-channel network
By adopting attention mechanism and dual-channel network in image super-resolution reconstruction, combining information cascade module, improved residual module and non-local cavity convolution block, the problems of information loss and local receptive field limitation in traditional methods are solved, and higher quality image reconstruction is achieved.
Patent Information
- Application Number
- CN202311675848.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-07
- Publication Date
- 2025-06-10
AI Technical Summary
Traditional image super-resolution reconstruction methods tend to lose smooth information during the calculation process, and are limited by local receptive fields, making it difficult to obtain global information of the image.
The image super-resolution reconstruction method based on attention mechanism and dual-channel network is adopted. By introducing information cascading module, improved residual module and non-local cavity convolution block, the global information of the image is captured and the reconstruction accuracy is improved.
It effectively solves the problems of information loss and local receptive field limitation, improves the quality and accuracy of image reconstruction, and can generate higher quality and more realistic high-definition images.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Background Art
[0001] Image super-resolution reconstruction technology is a method that generates high-quality, high-resolution images by using a set of low-quality, low-resolution images (or motion sequences). This technology has a wide range of applications in the field of computer vision, such as in surveillance imaging, medical imaging, and object recognition. In real life, due to factors such as the cost of image acquisition equipment, the video image transmission bandwidth, and the limitations of imaging modality technology, we do not always obtain large-size high-definition images with sharp edges and no blocky blur. In such cases, super-resolution reconstruction technology becomes particularly important. Its main goal is to process low-resolution images and restore higher-quality, clearer high-resolution images. The traditional image super-resolution (SR) reconstruction problem involves recovering high-resolution (HR) images from low-resolution (LR) images. And the learning-based image super-resolution reconstruction technology has been a research hotspot in recent years. By using the end-to-end mapping and convolutional neural network technology in deep learning, this method can learn the mapping relationship between low-resolution images and high-resolution images, and then infer the missing high-frequency details in the low-resolution images, so as to obtain high-quality images with higher quality and richer details. In recent years, the attention mechanism has become an important technology in image super-resolution reconstruction. By embedding the attention mechanism into the model, the algorithm can allocate attention more accurately, enabling the model to pay more attention to important features in the image. Such processing can further enhance the quality, clarity, and detail performance of the reconstructed image. In summary, image super-resolution reconstruction technology has broad application prospects, especially in scenarios where image data is limited or the resolution requirements are high. It can provide effective solutions, making the recognition and analysis of images more accurate and precise.
[0002] In traditional image reconstruction methods, the CNN (Convolutional Neural Network) is usually adopted for calculation. However, during the calculation process, these methods often gradually lose a lot of smooth information. Since the main goal of image super-resolution reconstruction is to obtain high-precision high-definition images, any loss of information will affect the final reconstruction accuracy. In addition, the calculation of the CNN adopts the method of local receptive fields, which means that limited by the size of the convolution kernel, the information obtained is only local information of the image. However, in an image, there are often correlations between pixel points, and the dependency relationships between pixel points in the image may involve relatively long distances. These long-distance dependency information is also crucial for image reconstruction, but it is difficult for traditional methods to obtain this global information. Therefore, in image super-resolution reconstruction, we need to find more comprehensive image information and solve the limitation of local receptive fields. One solution is to introduce an attention mechanism. Through the attention mechanism, the model can focus on the most important regions and features in the image, thereby more accurately restoring the details and structure of the image. At the same time, to solve the problems of information loss and obtaining global information, some novel super-resolution reconstruction methods adopt non-local dilated convolution blocks. This block combines dilated convolution layers with different dilation parameters and ordinary convolutional neural network layers, and can obtain the long-distance dependency information of the image within a larger receptive field, further improving the quality of the reconstructed image. In summary, by introducing technologies such as the attention mechanism and non-local dilated convolution blocks, the method of image super-resolution reconstruction can better capture the global information of the image, improve the reconstruction accuracy and quality, and meet the requirements of high-definition images. These innovative technologies have brought new breakthroughs and progress to the field of image reconstruction. Summary of the Invention
[0003] To overcome the problems existing in the prior art, the present invention proposes an image super-resolution reconstruction method based on an attention mechanism and a dual-channel network. The method includes a series of key steps: First, the image to be detected is obtained in real time and preprocessed to ensure the accuracy and adaptability of the input data; then, the preprocessed image is input into a pre-trained image super-resolution reconstruction model to obtain a high-definition reconstructed image. Here, the dual-channel network means that the model uses two parallel channels to extract high-frequency and low-frequency features respectively to more fully capture the details and structural information of the image. To accurately evaluate the quality of the reconstructed image, the present invention introduces metrics such as peak signal-to-noise ratio and structural similarity, which are commonly used to measure the similarity and accuracy between the image reconstruction result and the original high-definition image. By evaluating the reconstructed image, it is possible to better guide the optimization and improvement of the model, thereby obtaining a higher-quality super-resolution image. It is worth emphasizing that the image super-resolution reconstruction model adopted by the present invention is based on a convolutional neural network (CNN), which has been widely used in the field of computer vision and has performed excellently in image processing tasks. The powerful feature extraction ability of CNN helps to improve the efficiency and performance of the reconstruction model, thereby achieving more accurate and detailed image super-resolution reconstruction. Generally speaking, the present invention comprehensively applies advanced technologies such as attention mechanism, dual-channel network and CNN to achieve more accurate and reliable image super-resolution reconstruction. This method has important application prospects in the fields of image processing and computer vision, and can be used in fields such as surveillance imaging, medical imaging, object recognition, etc., providing higher-quality image data support for related applications.
[0004] The process of training the image super-resolution reconstruction model includes:
[0005] S1: First, collect a dataset of original high-definition pictures, which contains high-resolution image samples.
[0006] Next, we use a bicubic interpolation degradation model to scale each picture in the dataset. This interpolation technique can effectively downsample high-resolution images to low resolution, simulating the image degradation situation in actual application scenarios. Through this step, we obtain a set of downsampled low-resolution images, which will become the input for the image super-resolution reconstruction algorithm in subsequent research. By reconstructing these low-resolution images, our goal is to restore them to their high-definition state, making them visually similar to the original high-definition images. This preprocessing process is a key step in image super-resolution reconstruction research, providing a basis for the development and performance evaluation of subsequent algorithms. By reasonably downsampling the original high-definition image dataset, we can simulate a more realistic image input situation in actual applications, thus more accurately evaluating the performance of the image super-resolution reconstruction algorithm in actual applications.
[0007] S2: During the training process, we process each image data in the training dataset and input it into the image super-resolution reconstruction model. Inside the model, the image data is subjected to feature extraction through two different channels: the shallow feature channel and the deep feature channel. The shallow feature channel is responsible for extracting low-level features of the image, such as edges and textures. These features are crucial for retaining the details of the reconstructed image as they can capture the minute changes and subtle structures in the image; while the deep feature channel focuses on extracting high-level features of the image, which are more powerful in representing the semantic information of the image. Through the deep feature channel, the model can understand the more complex structures and contents in the image, thus better guiding the super-resolution reconstruction process; combining the shallow feature channel and the deep feature channel enables the image super-resolution reconstruction model to obtain richer and more accurate semantic information while retaining details, thereby generating high-quality high-definition reconstructed images. This feature extraction process is continuously optimized during the training process so that the model can learn more effective feature representations and improve the performance and effect of image super-resolution reconstruction.
[0008] S3: The first convolutional layer is used to extract the initial features of the input image, which are crucial for the initial characterization of the image. Next, we feed these initial features into the information cascading module. The role of the information cascading module is to interact and aggregate information between different convolutional layers. In this way, the model can obtain more comprehensive and rich information from features at different levels. In the information cascading module, we use appropriate algorithms and strategies to ensure the organic fusion of features from each convolutional layer, avoiding information loss and redundancy. Such a design enables our model to better understand the input image and utilize feature information from different levels during the reconstruction process. This information cascading process plays a crucial role in the accuracy and detail retention of image super-resolution reconstruction. It allows the model to obtain feature information of the image from multiple hierarchical perspectives and achieve higher-quality results when reconstructing the image. Therefore, by introducing this information cascading module, our image super-resolution reconstruction method is further optimized, providing stronger support for the generation of high-definition images.
[0009] S4: The hierarchical feature information aggregated by the information cascading module is input into the improved residual module, so that we can obtain the correlation information on the channels and the dependency information on the global space. In this process, the residual module plays a key role and can learn more effective feature representations from multiple feature channels, thereby further improving the performance and quality of image super-resolution reconstruction.
[0010] S5: Use a non-local dilated convolution block for global feature extraction to obtain the final deep feature map. This non-local dilated convolution block can perform information interaction across the entire image range, helping the model capture more global and richer features, thus providing a higher-quality deep feature representation for image super-resolution reconstruction.
[0011] S6: Use the second convolutional layer to extract the initial features of the input image. These initial features are then fed into the improved VGG network to extract the shallow features of the image, thus obtaining the shallow feature map. This process can help us capture features such as the underlying details and textures of the image, providing an important information basis for subsequent image super-resolution reconstruction.
[0012] S7: Fuse the deep feature map and the shallow feature map. Through this feature fusion, we effectively combine high-level semantic information and underlying detail features. Then, perform an upsampling operation on the fused feature map to restore it to the size of the original image. This step is to convert the low-resolution feature map into a high-resolution image, thus achieving image super-resolution reconstruction. Finally, we obtain a high-definition reconstructed image, which retains the semantic information of the original image while adding more details, significantly improving the image quality. This process can effectively restore the details and realism of the image,
[0013] bringing higher-quality results to the field of image processing.
[0014] Adopt the bicubic interpolation degradation model and perform scaling operations on each image by 2 times, 3 times, 4 times, and 8 times. Such a design fully considers the application scenarios of different scaling multiples, covering various requirements for image super-resolution reconstruction from small to large. By adopting the bicubic interpolation degradation model, we maintain a high image quality during the process of generating low-resolution images. In this way, we can more accurately restore the details and features of the image in the subsequent image reconstruction stage, making the results of image super-resolution reconstruction clearer and more realistic. The careful consideration of different scaling multiples of this method makes it more flexible and applicable in practical applications. At the same time, by combining the advantages of the attention mechanism and the dual-channel network, our image super-resolution reconstruction method brings higher-quality image reconstruction results to the field of image processing.
[0015] An information cascade module is introduced, which is composed of 10 stacked feature aggregation structures. Each feature aggregation structure includes at least three convolutional neural network layers, a feature channel merging layer, a channel attention layer, and a channel number transformation layer. These layers are connected together in sequence. Except for the last convolutional neural network layer, the output ends of other convolutional neural network layers are also connected to the feature channel merging layer. The feature channel merging layer, the channel attention layer, and the channel number transformation layer are connected together in sequence to form the information cascade module. During the process of image data processing, each information cascade module first extracts feature information of the input image through each convolutional neural network layer in sequence. Then, the feature information extracted by each layer of convolution is fused on the feature channel merging layer. After feature fusion, a channel attention mechanism is used to distinguish the importance of the merged information, so as to focus on those features that make important contributions to the reconstructed image. Finally, the channel number is reduced to the size of the input channel number through the channel number transformation layer. Repeat the above feature aggregation process 10 times, and we obtain the hierarchical feature information of the aggregated convolutional layer. The design of this information cascade module enables our method to capture multi-level information of images more effectively and improve the quality of image super-resolution reconstruction. By stacking the feature aggregation structures multiple times, our method enhances the depth and breadth of feature representation, further improving the image reconstruction effect.
[0016] The improved residual module includes three key parts: a residual network structure, a channel attention mechanism layer, and a spatial attention mechanism layer. First, the residual network structure is composed of two convolutional neural network layers and a non-linear activation layer. Such a design allows the model to extract multi-scale features of the image at different levels, enhancing the model's ability to understand image information. During the process of image data processing, the hierarchical feature information is first input into the residual network structure for feature information extraction. Then, the extracted feature information passes through the channel attention mechanism layer, and the channel attention mechanism is used to obtain the correlation between feature channels. Such processing enables the model to focus on the feature channels that are most critical for image reconstruction in a targeted manner, improving the model's attention focus. Subsequently, the processing continues and the feature information is sent to the spatial attention mechanism layer to obtain global spatial dependency information. The spatial attention mechanism helps the model better understand the global structure and semantic information of the image, thus more effectively guiding the image reconstruction process. Through such a residual module design, combined with the channel attention mechanism and the spatial attention mechanism, our image super-resolution reconstruction method can more accurately learn the important features in the image and make full use of the global context information during the reconstruction process. The application of this comprehensive feature extraction and attention mechanism enables our method to generate higher-quality and more realistic super-resolution images.
[0017] The non-local dilated convolution block consists of four parallel dilated convolution layers with dilation parameters of 1, 2, 4, and 6 respectively, and three ordinary convolutional neural network layers. During the image data processing, first, we use four dilated convolutions with different dilation parameters and two ordinary convolutional neural networks to extract feature information from the dependency information input to the improved residual network respectively. Such a design allows the model to capture multi-scale features of the image in parallel in multiple dilated convolution layers, and extract more global and high-level features in the ordinary convolutional neural network. Then, we fuse the feature information obtained from the four dilated convolutions on the feature channels, while the feature information extracted by the ordinary convolutional neural network is fused according to the values of the pixel matrix. This fusion method makes full use of the feature information extracted by the dilated convolution and the ordinary convolution at different levels, enhancing the model's comprehensive expression ability for image details and global information. Finally, we add these two fused feature information to obtain the global feature information. This process enables the model to understand the content of the image more comprehensively and provides a more accurate and rich feature representation for image super-resolution reconstruction. By introducing such a non-local dilated convolution block, our method can more effectively capture multi-scale information in the image, improving the quality and realism of image reconstruction.
[0018] The improved VGG network structure includes 10 ordinary convolutional layers and 3 pooling layers. We embed these pooling layers into the ordinary convolutional layers to form the VGG network structure. During the image data processing, first, we use 2 convolutional layers and 1 pooling operation to extract feature information of 64 channels. Then, we use another 2 convolutional layers and 1 pooling operation to extract feature information of 128 channels. Next, we extract feature information of 512 channels through 3 convolutional layers and 1 pooling operation. Finally, we use 3 convolutional layers to convert the 512-channel information back to 64 channels. It should be noted that when performing the pooling operation, we use padding to keep the feature scale unchanged. Such a design ensures that important information of the image is not lost during the pooling process, and at the same time allows the VGG network to extract multi-scale features of the image at different levels. By introducing such an improved VGG network structure, our image super-resolution reconstruction method can make more full use of the hierarchical feature representation ability of the VGG network, improving the abstraction and understanding of the image content. This improvement helps to obtain a richer and more accurate image feature expression, thus enhancing the effect and quality of image super-resolution reconstruction.
[0019] The expression of the loss function of the image super-resolution reconstruction model is:
[0020]
[0021] The Peak Signal-to-Noise Ratio (PSNR) and the Structural Similarity Index (SSIM) are used as the formulas for evaluating the reconstructed image:
[0022]
[0023]
[0024] Advantages of the present invention:
[0025] 1. The present invention introduces a dual-channel network. One of the networks adopts an improved residual structure to specifically extract valuable high-frequency features, also known as advanced features. The other network adopts an improved VGG structure. Here, we fine-tuned the parameters of the VGG convolutional layer and pooling layer to ensure that the input and output image sizes are consistent, and removed the last fully connected layer in the VGG structure. The purpose of this is to better adapt to the image super-resolution reconstruction task. Through such a dual-channel design, we can fully capture the high-frequency and low-frequency features in the image. High-frequency features mainly cover the details and edge information of the image, while low-frequency features contain the overall structure and background information of the image. After obtaining the features extracted by these two networks, we perform the step of feature fusion. This process aims to organically combine the high-frequency and low-frequency features to retain the detail information of the image while maintaining the accuracy of the overall structure.
[0026] 2. The present invention adopts a dense connection method at specific positions of the model (2 information cascading modules at the head and the tail). This means that the output of each convolutional layer is connected to the input of the subsequent layer, thus realizing the full utilization and transmission of feature information. Through this dense connection, the model can better utilize the convolutional features at different levels, enhance the correlation between features, and contribute to improving the accuracy of image reconstruction and the ability to retain details. In addition, in the final processing stage, we introduce a channel attention mechanism to weight the features. The channel attention mechanism can calculate the importance of each channel and weight the features according to the importance, rather than simply reducing the number of channels. By accurately calculating the channel weights, we can more accurately fuse the feature information of different channels, thereby further improving the quality and effect of image super-resolution reconstruction. Making full use of the advantages of dense connection and channel attention mechanism enables the image super-resolution reconstruction model to more comprehensively and accurately capture the features and details of the image. This comprehensive application brings new technological breakthroughs to the field of image processing and broader development prospects to the research and application of image super-resolution reconstruction. Through the method of the present invention, higher-quality and more realistic high-definition image reconstruction can be achieved, providing better image processing solutions for various application scenarios.
[0027] 3. The present invention introduces a spatial attention mechanism and combines it with the existing channel attention mechanism to enhance the extraction of global information and the comprehensive utilization of features. The spatial attention mechanism can accurately capture the global information of the image, thereby enhancing the model's perception ability of the overall structure. At the same time, through the channel attention mechanism, we can perform weighted processing on the features of different channels, enabling the model to pay more attention to important feature channels, thereby further optimizing the expression and utilization of features. In addition, before upsampling, we also adopt the non-local dilated convolution technique, and such an operation can perform global-dependent feature extraction on the previously obtained feature information. This global-dependent feature extraction helps to capture the long-range dependent information in the image, making the output result more closely related to the input image and increasing the detailed information of the output image. Through the spatial attention mechanism and non-local dilated convolution technique adopted in the present invention, we have achieved more accurate and comprehensive feature extraction, providing a better solution for the image super-resolution reconstruction task. This comprehensive application brings new technological breakthroughs and application prospects to the field of image processing. Through the method of the present invention, we can obtain higher-quality and more realistic high-definition image reconstruction results, providing stronger support for various image processing tasks. Description of the Drawings
[0028] Figure 1 It is the overall structure diagram of an image super-resolution reconstruction model provided by this patent;
[0029] Figure 2 It is the information cascade structure diagram provided by this patent;
[0030] Figure 3 It is the residual structure diagram provided by this patent;
[0031] Figure 4 It is the channel attention and spatial attention structure diagram provided by this patent;
[0032] Figure 5 It is the non-local dilated convolution diagram provided by this patent; Detailed Embodiment
[0033] The present invention proposes an image super-resolution reconstruction model structure, as Figure 1As shown. The model structure includes a deep feature channel, a shallow feature channel, an upsampling layer, and a third convolutional layer. The deep feature channel consists of a first convolutional layer, an information cascading module, an improved residual module, and a non-local dilated convolution block. During the processing, the input image first undergoes initial feature extraction through the first convolutional layer, and then the outputs of each layer of convolution are connected layer by layer through the information cascading module to achieve dense feature transmission. Next, an improved residual module is adopted, enabling the network to extract high-frequency features more deeply, thereby capturing the details and texture information in the image. Finally, a non-local dilated convolution block is introduced. Through global-dependent feature extraction, the feature representation is further optimized, making the output result more closely related to the input image and improving the quality of image reconstruction. The shallow feature channel consists of a second convolutional layer and an improved VGG network. After the input image is processed by the second convolutional layer, feature extraction is performed through the improved VGG network. Here, the parameters of the VGG network are fine-tuned to ensure that the sizes of the input and output images are the same, and the last fully connected layer is removed to better adapt to the super-resolution reconstruction task, obtaining a shallow feature map. In the fusion stage, the deep feature map and the shallow feature map are fused. Then, the fused image is upsampled through the upsampling layer, and the upsampled image is convolved using the third convolutional layer, finally obtaining a high-definition reconstructed image. This image super-resolution reconstruction model structure comprehensively utilizes the information of deep features and shallow features. Through fusion and upsampling processing, it can effectively improve the resolution and quality of the image. The method of the present invention brings new technological breakthroughs and application prospects to the field of image super-resolution reconstruction, providing an efficient and accurate solution for image processing tasks, and having broad application potential especially in the fields of medical imaging, surveillance imaging, and object recognition.
[0034] In the deep feature channel of the present invention, n information cascading modules and m improved residual modules are included. These information cascading modules are connected in series in sequence to form a group of information cascading modules for realizing layer-by-layer feature transmission and aggregation. Similarly, the improved residual modules are also connected in series together to form a group of improved residual modules for effectively extracting high-frequency features and optimizing feature representation. Through these cascaded module groups, we can make full use of deep features to achieve a more accurate and comprehensive image super-resolution reconstruction effect.
[0035] The process of training the image super-resolution reconstruction model includes:
[0036] S1: First, collect the original high-definition picture dataset, which contains high-resolution image samples.
[0037] Then, we use the bicubic interpolation degradation model to scale each picture in the dataset.
[0038] This interpolation technique can effectively downsample high-resolution images to low-resolution, simulating the image degradation in actual application scenarios. Through this step, we obtain a set of downsampled low-resolution images, which will serve as the input for the image super-resolution reconstruction algorithm in subsequent research. By reconstructing these low-resolution images, our goal is to restore them to their high-definition state, making them visually similar to the original high-definition images. This preprocessing process is a crucial step in image super-resolution reconstruction research, providing a basis for the development and performance evaluation of subsequent algorithms.
[0039] By reasonably downsampling the original high-definition image dataset, we can simulate a more realistic image input situation in actual applications, thus more accurately evaluating the performance of the image super-resolution reconstruction algorithm in practical applications.
[0040] S2: During the training process, we process each image data in the training dataset and input it into the image super-resolution reconstruction model. Inside the model, the image data undergoes feature extraction through two different channels: the shallow feature channel and the deep feature channel. The shallow feature channel is responsible for extracting low-level features of the image, such as edges and textures. These features are crucial for retaining the details of the reconstructed image as they can capture the subtle changes and fine structures in the image; while the deep feature channel focuses on extracting high-level features of the image, which are more powerful in representing the semantic information of the image. Through the deep feature channel, the model can understand the more complex structures and contents in the image, thus better guiding the super-resolution reconstruction process; combining the shallow feature channel and the deep feature channel enables the image super-resolution reconstruction model to obtain richer and more accurate semantic information while retaining details, thereby generating high-quality high-definition reconstructed images. This feature extraction process is continuously optimized during the training process to enable the model to learn more effective feature representations and improve the performance and effect of image super-resolution reconstruction.
[0041] S3: The first convolutional layer is used to extract the initial features of the input image, which are crucial for the preliminary characterization of the image. Next, we feed these initial features into the information cascading module. The role of the information cascading module is to interact and aggregate information between different convolutional layers. In this way, the model can obtain more comprehensive and rich information from features at different levels. In the information cascading module, we use appropriate algorithms and strategies to ensure the organic fusion of features from each convolutional layer, avoiding information loss and redundancy. Such a design enables our model to better understand the input image and utilize feature information from different levels during the reconstruction process. This information cascading process plays a crucial role in the accuracy and detail retention of image super-resolution reconstruction. It allows the model to obtain feature information of the image from multiple level perspectives and achieve higher-quality results when reconstructing the image. Therefore, through the introduction of this information cascading module, our image super-resolution reconstruction method is further optimized, providing stronger support for the generation of high-definition images.
[0042] S4: The hierarchical feature information aggregated by the information cascading module is input into the improved residual module, so that we can obtain the correlation information on the channels and the dependence information in the global space. In this process, the residual module plays a key role and can learn more effective feature representations from multiple feature channels, thereby further improving the performance and quality of image super-resolution reconstruction.
[0043] S5: The non-local dilated convolution block is used to perform global feature extraction to obtain the final deep feature map. This non-local dilated convolution block can perform information interaction across the entire image range, helping the model capture more global and rich features, thereby providing a higher-quality deep feature representation for image super-resolution reconstruction.
[0044] S6: The second convolutional layer is used to extract the initial features of the input image. These initial features are then input into the improved VGG network to extract the shallow features of the image, thereby obtaining the shallow feature map. This process can help us capture the underlying details and textures of the image, providing an important information basis for subsequent image super-resolution reconstruction.
[0045] S7: The deep feature map and the shallow feature map are fused. Through this feature fusion, we effectively combine the high-level semantic information and the underlying detail features. Then, an upsampling operation is performed on the fused feature map to restore it to the size of the original image. This step is to convert the low-resolution feature map into a high-resolution image, thereby achieving image super-resolution reconstruction. Finally, we obtain the high-definition reconstructed image, which retains the semantic information of the original image while adding more details, significantly improving the image quality. This process can effectively restore the details and realism of the image.
[0046] Bring better results to the field of image processing.
[0047] Dataset We used the DIV2K dataset as the training dataset, which contains 800 high-resolution (HR) images and the corresponding low-resolution (LR) images obtained through a degradation model (bicubic interpolation degradation). Such a data setting can fully cover the image variations and complexities in different scenarios, helping to improve the generalization performance of our model. From the DIV2K dataset, we carefully selected five images as the validation set to ensure the performance and generalization ability of the model during the training process. To comprehensively evaluate our image super-resolution reconstruction model, we selected some widely used test datasets, including Set5, Set14, Urban100, Manga109, and BSD100. These test datasets are known for their rich textures and diversity. After degradation, a large amount of detail information is lost in the low-resolution images, so the accuracy and quality requirements for image super-resolution reconstruction are very high. By testing on these datasets, we can more comprehensively evaluate and compare the performance differences between our method and other algorithms. When evaluating the quality of the image super-resolution reconstruction model, we adopted the traditional evaluation metrics PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index). PSNR is used to quantify the noise level between the image reconstruction and the original high-resolution image, while SSIM is used to measure the structural similarity between two images. These evaluation metrics can provide us with an objective measurement standard, helping to more accurately evaluate the performance of the image super-resolution reconstruction model and compare it with other methods. Through the above detailed dataset selection and evaluation metric settings, we can fully verify and analyze the performance of the image super-resolution reconstruction method proposed in this invention in diverse scenarios, as well as its feasibility and effectiveness in practical applications. These experimental results will provide strong support for the further improvement and application of our method.
[0048] During this training process, we adopted an iterative optimization method to train the neural network model. Each piece of data in the training set would go through one forward propagation and one backward propagation, forming one iteration. Through such a backward propagation process, we were able to update the model's parameters according to the error of the training data to continuously optimize the model's performance. To ensure the sufficient training of the model, we set the maximum number of iterations to 1000. During the training process, we updated the learning rate every 200 iterations. Such a strategy helped maintain the stability of the training process and enabled the model to converge better step by step. Throughout the 1000 - iteration training process, we particularly focused on the performance of the model on the test dataset. We continuously monitored the effect of the model on the test set and saved the model with the best performance on the test dataset and its corresponding parameters. Such a saving mechanism could ensure that we obtained the model with the optimal performance on the test dataset to guarantee that our method could achieve the best image super - resolution reconstruction effect in practical applications.
[0049] In our original high - definition image dataset, we adopted a bicubic interpolation degradation model to scale each image in the dataset by multiple factors. Specifically, we used scaling ratios of 2x, 3x, 4x, and 8x to perform the scaling operation on each high - definition image. Such processing simulated different degrees of image down - resolution in practical application scenarios, thus covering various application requirements for super - resolution reconstruction. In this way, we could include various scales and degrees of image degradation in the training dataset, enabling our image super - resolution reconstruction method to learn and understand image features at different scales more comprehensively. Such a data augmentation strategy helped improve the generalization ability of the model, enabling our model to better adapt to image super - resolution reconstruction tasks of various scales in practical applications, and thus achieving a higher - quality reconstruction effect.
[0050] Preprocessing the scaled dataset is an important step in this method, which includes multiple enhancement operations. First, we perform translation, horizontal and vertical flipping on the images to increase data diversity and robustness, thereby improving the generalization ability of the model. Next, we segment the enhanced data and cut each image into different small image patches. Such segmentation operations help to retain the detailed information in the images and can introduce more local features during the training process, enabling the model to learn and understand the structure of the images more comprehensively. Finally, we assemble the segmented small image patches to construct the final training dataset. Through this preprocessing process, we obtain a rich and diverse training dataset that contains various enhanced image samples and local image patches, providing sufficient samples and information for the training of the model. This data processing method helps to improve the training effect, enabling our image super-resolution reconstruction model to perform well in practical applications and handle the image super-resolution challenges in various complex scenarios.
[0051] As Figure 2 shown, the information cascade module is a key component of this method and consists of the following structure: stacking the following structure 10 times, which sequentially includes three layers of convolutional neural networks, a feature channel merging layer, a channel attention layer, and a channel number transformation layer. The processing process of this module is as follows: First, we use each layer of convolutional neural networks to sequentially extract feature information from the input image, gradually capturing the local features of the image through these convolutional layers. Next, the feature information extracted by each layer of convolution is merged on the feature channel merging layer. The purpose of this is to fuse the feature information at different levels together to form a more rich and comprehensive feature representation. Then, we introduce a channel attention mechanism to distinguish the importance of the merged feature information, which helps the model to pay more attention to the feature channels that contribute to image super-resolution reconstruction and improve the reconstruction accuracy of the model. Finally, the number of channels is reduced to the size of the input channels. This step can reduce the computational complexity while maintaining an effective feature representation. The entire information cascade module repeats the above processing steps 10 times to obtain the hierarchical feature information of the aggregated convolutional layer. Such a design enables our model to aggregate feature information layer by layer and comprehensively consider the feature representations at different levels, thereby obtaining more rich and accurate image features and providing strong support for image super-resolution reconstruction.
[0052] By using the information cascading module to aggregate image information, we can effectively retain the information of each convolutional layer. In a convolutional neural network, when an image first enters the network, the low-frequency information is usually sufficient and rich, and this information contains features such as the details and textures of the image. However, as the number of network layers increases, the model pays more attention to the extraction of abstract features, and some edge textures and smooth information may gradually be lost. In this case, by adopting the information cascading module, we can well capture more low-frequency information and integrate it into the model. The advantage of doing so is that we can comprehensively utilize the feature representations of each level, including those relatively detailed low-frequency information, enabling the model to more comprehensively understand the structure and content of the image. The introduction of the information cascading module can effectively enhance the model's representational ability for images, thereby improving the accuracy and quality of image super-resolution reconstruction. By fully retaining the information of each convolutional layer, we can better capture the details and textures of the image, making the reconstructed image clearer and more delicate, thus meeting the requirements for high-quality images in practical applications.
[0053] The structure of the improved residual module is as Figure 3 shown. It consists of three main components: a residual network structure, a channel attention mechanism layer, and a spatial attention mechanism layer. Among them, the residual network structure is composed of a convolutional neural network layer, a non-linear activation layer, and a convolutional neural network layer. In the process of image super-resolution reconstruction, the improved residual module plays a key role. Its processing process is as follows: First, the hierarchical feature information is input into the residual network structure, and the convolutional neural network layer and the non-linear activation layer are used to extract the feature information of the image. Then, through the channel attention mechanism layer, the correlation calculation of the extracted feature information is carried out on the channels to capture the correlation and importance between different channels. On this basis, the spatial attention mechanism layer is further used to obtain the global spatial dependency information, thereby strengthening the model's understanding of the overall structure of the image. The design of this improved residual module enables the model to more accurately learn the feature and structure information in the image, and when performing super-resolution reconstruction, it can better maintain the details and textures of the image, thus generating higher-quality reconstructed images. By combining the mechanisms of channel attention and spatial attention, the model can comprehensively utilize image information and further improve the performance and accuracy of image reconstruction.
[0054] We use the output of the cascaded module as the input to the improved residual module. After each ResNetBlock, a channel attention mechanism and a spatial attention mechanism are introduced to capture the correlation between feature channels and global spatial dependency information. Such a structure allows us to effectively fuse feature information at different levels throughout the network, thereby improving the feature expression ability. The introduction of the channel attention mechanism helps to enhance the modeling of the importance of feature channels, enabling the network to better learn the correlation between different channels. The application of the spatial attention mechanism can capture the global context information of the image within a wider receptive field range, thereby improving the perception ability of the global structure. Through the combination of these attention mechanisms, we can process the features of the image more meticulously, making the network more targeted and adaptable, and thus obtaining more accurate and clear results in the image super-resolution reconstruction task. This design can also help alleviate the vanishing gradient problem in the training of deep networks, providing a more stable gradient signal for model optimization. Overall, our method makes the image reconstruction process more robust and efficient by effectively integrating attention mechanisms and improved residual structures.
[0055] In the image super-resolution reconstruction model, we introduce a channel attention structure and a spatial attention structure, as Figure 4 shown. The channel attention structure consists of the following: First, the feature map of each channel is pooled through a global average pooling layer to obtain the weight representation of each channel; then, a 1x1 convolutional layer is used to reduce the number of channels to reduce computational complexity; then, a non-linear activation layer is introduced to increase the non-linearity of the network; finally, another 1x1 convolutional layer is used to restore the number of channels to the original size, and the obtained weight information is multiplied by the original input feature information to achieve the modeling of the correlation on the feature channels. The spatial attention structure consists of the following: First, the input feature information CxHxW is converted into a global feature map of HWx1x1 through a 1x1 convolutional layer for global spatial modeling; then, the softmax function is used to normalize the global feature map to ensure that the sum of the weights at each position is 1; then, the normalized feature map is multiplied back to the original input information to highlight the important spatial position information; finally, another 1x1 convolutional layer and a non-linear activation layer are used to further strengthen the modeling of the global spatial dependency. The introduction of these attention structures makes our model more flexible and accurate in processing image features. By fully exploring the correlation between feature channels and the global spatial dependency, our model can better capture the detailed information and context information of the image, and thus achieve better results in the image super-resolution reconstruction task. The addition of these detailed information makes our model more robust and adaptable in practical applications, bringing new breakthroughs to the field of image super-resolution reconstruction.
[0056] As shownFigure 5 As shown, we propose a non-local dilated convolution block, which consists of four parallel dilated convolution layers with dilation factors of 1, 2, 4, and 6 respectively, and three ordinary convolutional neural network layers. In the image super-resolution reconstruction model, the role of this non-local dilated convolution block is very crucial. When processing image data, first, the feature information will pass through four dilated convolution layers with different dilation factors and two ordinary convolutional neural network layers simultaneously, so as to extract feature information of different scales respectively. Then, we perform the operation of feature fusion. On the one hand, the feature information obtained from the four dilated convolutions is fused on the feature channels, which can retain the feature information at different scales, enabling the model to better adapt to feature structures of different sizes. On the other hand, the feature information extracted by the ordinary convolutional neural network is fused according to the values of the pixel matrix, which can make full use of the spatial information of the image, making the feature fusion more comprehensive and detailed. Finally, we add these two fused feature information together to obtain the global feature information. Such a design enables our model to better capture the global information in the image while taking into account the feature details at different scales. The introduction of the non-local dilated convolution block greatly improves the performance of the image super-resolution reconstruction model, enabling our model to more accurately restore high-definition details during the reconstruction process and obtain clearer and finer high-resolution images.
[0057] The improved VGG network structure uses 10 ordinary convolutional layers and 3 pooling layers, where the pooling layers are embedded in the ordinary convolutional layers to form the VGG network structure. The process of this module for processing image data is as follows: First, 2 convolutional layers and 1 pooling operation are used to extract the feature information of 64 channels, and then another 2 convolutional layers and 1 pooling operation are performed to further extract the feature information of 128 channels. Then, through 3 convolutional layers and 1 pooling operation, the feature information of 512 channels is extracted. Finally, 3 convolutional layers are used to process this 512-channel feature information and restore it to the original 64 channels. It should be particularly noted that the pooling layer uses padding during the processing to keep the feature scale unchanged, which is to ensure that important information of the image is not lost during the pooling operation. Through the above processing process, the improved VGG network structure can effectively extract the feature information of different channels and gradually process and restore these features. Such a design enables the model to better understand the content of the image and play an important role in the image super-resolution reconstruction process. The introduction of the improved VGG network structure enables our image super-resolution reconstruction model to more accurately restore the details and textures of the image and provide clearer and more realistic high-resolution images.
[0058] The extraction of global features is achieved using a non-local dilated convolution block. After the convolution block extracts the feature information, upsampling is performed to expand the feature map to the target size we need, thereby obtaining the final output result. By setting the dilation rate, non-local dilated convolution can expand the receptive field without increasing the number of parameters. Embedding it in non-local convolution effectively reduces the computational cost, and at the same time, it can obtain global information from different scales, making the feature extraction more comprehensive. This design enables our image super-resolution reconstruction model to fully utilize global information, effectively capture long-range dependencies in the image, and accurately restore the details and textures of the image. The introduction of the non-local dilated convolution block improves the model performance while avoiding additional computational burdens, making the image super-resolution reconstruction process more efficient and accurate.
[0059] F NLHC = H NLHC (F SA )
[0060] Among them, H NLHC represents the convolution operation of non-local dilated convolution, and F NLHC represents the feature information obtained after non-local dilated convolution. After upsampling the final feature information, we output the corresponding high-definition reconstructed image. That is, the formula for the reconstructed image is:
[0061] F UP = H UP (F NLHC )
[0062] Among them, H UP represents the convolution operation of upsampling, and F UP represents the output feature of upsampling.
[0063] The expression of the loss function of the image super-resolution reconstruction model is:
[0064]
[0065] Among them, θ represents the number of model parameters, CHR represents the super-resolution calculation equation, I i LR and I i HR represent the i-th low-resolution image and the i-th corresponding high-resolution image respectively, N represents the number of images in the dataset, HR represents high resolution, and LR represents low resolution. This loss function is a mathematical expression used when training the image super-resolution reconstruction model to measure the difference between the reconstructed image and the original high-definition image. By optimizing this loss function, our method can better guide the learning process of the model, making the generated super-resolution image closer to the real high-definition image, thereby improving the quality and accuracy of image reconstruction.
[0066] The expression of the super-resolution calculation equation is as follows:
[0067] C HR = F UP (F NLHC (F SA (F CA (F RBC (F IC (I LR ))))))
[0068] Among them, F UP represents the output information after upsampling, F NLHC represents the information extracted by non-local dilated convolution, F SA represents the information extracted by the spatial attention mechanism, F CA represents the information extracted by the channel attention mechanism, F RBC represents the information extracted by the residual block, F IC represents the information output by the cascade module.
Claims
1. This image super-resolution reconstruction method adopts the unique features of the attention mechanism and the dual-channel network to improve the clarity and quality of images. Specifically, this method includes the following key steps: obtaining the image to be detected in real time and performing preprocessing before processing. The preprocessing process covers operations such as denoising and contrast enhancement, aiming to ensure the accuracy and quality of the input image; the trained image super-resolution reconstruction model takes the preprocessed image as input. This model uses the attention mechanism to focus on and extract key details in the image, thus realizing the generation of high-definition reconstructed images; in order to objectively evaluate the quality of the reconstructed images, evaluation metrics such as peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) are used to comprehensively evaluate the reconstructed images. According to the evaluation results, the high-definition reconstructed images are labeled for subsequent analysis and processing; the dual-channel network in this image super-resolution reconstruction method is its key highlight. Among them, one network adopts an improved residual structure to specifically extract valuable high-frequency features (advanced features), which helps to retain the detail information of the image. And the other network uses an improved VGG network to ensure the consistency of the input and output image sizes and can effectively extract rich low-frequency features. Finally, the features of these two networks are fused to further enhance the clarity and realism of image reconstruction. By means of this method that cleverly combines the attention mechanism and the dual-channel network, the problems in image super-resolution reconstruction are successfully solved, and the clarity and quality of images are significantly improved. This method has broad application prospects in the fields of high-definition image reconstruction, surveillance image enhancement, medical image processing, etc., and has brought remarkable progress to the field of image processing. The process of training the image super-resolution reconstruction model includes: S1: First, collect the original high-definition picture dataset, which contains high-resolution image samples. Then, we use the bicubic interpolation degradation model to scale each picture in the dataset. This interpolation technique can effectively downsample high-resolution images to low resolution, simulating the image degradation situation in actual application scenarios. Through this step, we obtain a set of downsampled low-resolution images, which will become the input for the image super-resolution reconstruction algorithm in subsequent research. By reconstructing these low-resolution images, our goal is to restore them to the high-definition state and make them visually similar to the original high-definition images. This preprocessing process is a key step in image super-resolution reconstruction research, providing a basis for the development and performance evaluation of subsequent algorithms. By reasonably downsampling the original high-definition image dataset, we can simulate a more realistic image input situation in actual applications, thus more accurately evaluating the performance of the image super-resolution reconstruction algorithm in actual applications. S2: During the training process, we process each image data in the training dataset and input it into the image super-resolution reconstruction model. Inside the model, the image data undergoes feature extraction through two different channels: the shallow feature channel and the deep feature channel. The shallow feature channel is responsible for extracting low-level features of the image, such as edges and textures. These features are crucial for retaining the details of the reconstructed image because they can capture the tiny changes and subtle structures in the image; while the deep feature channel focuses on extracting high-level features, which are more powerful in representing the semantic information of the image. Through the deep feature channel, the model can understand the more complex structures and contents in the image, thus better guiding the super-resolution reconstruction process; combining the shallow feature channel and the deep feature channel enables the image super-resolution reconstruction model to obtain richer and more accurate semantic information while retaining details, thereby generating high-quality high-definition reconstructed images. This feature extraction process is continuously optimized during the training process to enable the model to learn more effective feature representations and improve the performance and effect of image super-resolution reconstruction. S3: The first convolutional layer is used to extract the initial features of the input image, which are crucial for the initial characterization of the image. Next, we input these initial features into the information cascading module. The role of the information cascading module is to interact and aggregate information between different convolutional layers. In this way, the model can obtain more comprehensive and rich information from features at different levels. In the information cascading module, we use appropriate algorithms and strategies to ensure the organic fusion of features from each convolutional layer, avoiding information loss and redundancy. Such a design enables our model to better understand the input image and utilize feature information from different levels during the reconstruction process. This information cascading process plays a crucial role in the accuracy and detail retention of image super-resolution reconstruction. It allows the model to obtain feature information of the image from multiple levels and achieve higher-quality results when reconstructing the image. Therefore, through the introduction of this information cascading module, our image super-resolution reconstruction method is further optimized, providing stronger support for the generation of high-definition images. S4: The hierarchical feature information aggregated by the information cascading module is input into the improved residual module, so that we can obtain the correlation information on the channels and the dependency information in the global space. In this process, the residual module plays a key role and can learn more effective feature representations from multiple feature channels, thereby further improving the performance and quality of image super-resolution reconstruction. S5: The non-local dilated convolution block is used to perform global feature extraction to obtain the final deep feature map. This non-local dilated convolution block can perform information interaction across the entire image range, helping the model capture more global and richer features, thereby providing a higher-quality deep feature representation for image super-resolution reconstruction. S6: Use the second convolutional layer to extract the initial features of the input image. These initial features are then fed into the improved VGG network to extract the shallow features of the image, thus obtaining the shallow feature map. This process can help us capture features such as the underlying details and textures of the image, providing an important information basis for subsequent image super-resolution reconstruction. S7: Fuse the deep feature map and the shallow feature map. Through this feature fusion, we effectively combine the high-level semantic information and the underlying detail features. Then, perform an upsampling operation on the fused feature map to restore it to the size of the original image. This step is to convert the low-resolution feature map into a high-resolution image, thus achieving image super-resolution reconstruction. Finally, we obtain the high-definition reconstructed image, which retains the semantic information of the original image while adding more details, significantly improving the image quality. This process can effectively restore the details and realism of the image, bringing higher-quality results to the field of image processing.
2. As described in claim 1, this image super-resolution reconstruction method based on the attention mechanism and the dual-channel network has unique features. When processing the pictures in the dataset, we adopted the bicubic interpolation degradation model and performed scaling operations on each picture by 2 times, 3 times, 4 times, and 8 times. Such a design fully considers the application scenarios of different scaling multiples, covering various requirements for image super-resolution reconstruction from small to large. By adopting the bicubic interpolation degradation model, we maintained a high image quality during the process of generating low-resolution images. In this way, we can more accurately restore the details and features of the image in the subsequent image reconstruction stage, making the results of image super-resolution reconstruction clearer and more realistic. The characteristic of carefully considering different scaling multiples of this method makes it more flexible and applicable in practical applications. At the same time, combining the advantages of the attention mechanism and the dual-channel network, our image super-resolution reconstruction method brings higher-quality image reconstruction effects to the field of image processing.
3. The image super-resolution reconstruction method based on the attention mechanism and the dual-channel network according to claim 1 is characterized in that the formula of the bicubic interpolation degradation model adopted is: I LR = H dn I HR + n, Among them, I LR represents the low-resolution image, H dn represents the degradation model, I HR represents the original high-resolution image, and n represents additional noise. This formula is a mathematical expression used during the image scaling process to reduce the image resolution and provide an appropriate low-resolution input for subsequent image reconstruction steps. Such a design enables our method to achieve high-quality image super-resolution reconstruction at different scaling factors according to different application requirements.
4. The image super-resolution reconstruction method based on the attention mechanism and the dual-channel network as described in claim 1 is characterized by its preprocessing process. This process performs enhancement processing on the scaled dataset, including operations such as translation, horizontal and vertical flipping, etc. Through these enhancement processes, we can expand the dataset, increase the diversity and richness of the data, thereby improving the generalization ability of the model. Then, we divide the enhanced data into different small image patches and combine these divided image patches into a training dataset. Such a processing method helps to provide richer image samples and more image segments of different sizes and contents, providing more comprehensive data support for model training. Through such a preprocessing process, we provide more diverse and challenging training data for the image super-resolution reconstruction method, enabling our model to better adapt to image reconstruction tasks in different scenarios and achieve higher performance and quality levels.
5. The image super-resolution reconstruction method based on the attention mechanism and dual-channel network according to claim 1 is characterized in that an information cascade module is introduced, and this module is composed of 10 stacked feature aggregation structures. Each feature aggregation structure includes at least three convolutional neural networks, a feature channel merging layer, a channel attention layer, and a channel number transformation layer. These layers are connected together in sequence. Except for the last convolutional neural network, the output ends of other convolutional neural networks are also connected to the feature channel merging layer. The feature channel merging layer, the channel attention layer, and the channel number transformation layer are connected in sequence to form the information cascade module. During the image data processing, each information cascade module first extracts feature information of the input image through each convolutional neural network in sequence. Then, the feature information extracted by each layer of convolution is fused on the feature channel merging layer. After feature fusion, the channel attention mechanism is used to distinguish the importance of the merged information, so as to focus on the features that make important contributions to the reconstructed image. Finally, the channel number is reduced to the size of the input channel number through the channel number transformation layer. Repeat the above feature aggregation process 10 times, and we obtain the hierarchical feature information of the aggregated convolutional layer. The design of this information cascade module enables our method to more effectively capture the multi-level information of the image and improve the quality of image super-resolution reconstruction. By stacking the feature aggregation structures multiple times, our method enhances the depth and breadth of feature representation, further improving the image reconstruction effect.
6. The image super-resolution reconstruction method based on the attention mechanism and dual-channel network according to claim 1, is characterized in that an improved residual module is introduced. This module includes three key parts: a residual network structure, a channel attention mechanism layer, and a spatial attention mechanism layer. First, the residual network structure is composed of two convolutional neural network layers and a non-linear activation layer. Such a design allows the model to extract multi-scale features of the image at different levels, enhancing the model's ability to understand image information. During the image data processing, the hierarchical feature information is first input into the residual network structure for feature information extraction. Then, the extracted feature information passes through the channel attention mechanism layer, and the channel attention mechanism is used to obtain the correlation between feature channels. Such processing enables the model to focus on the feature channels that are most critical for image reconstruction in a targeted manner, improving the attention focus of the model. Subsequently, the processing continues and the feature information is sent to the spatial attention mechanism layer to obtain the global spatial dependency information. The spatial attention mechanism helps the model better understand the global structure and semantic information of the image, thus more effectively guiding the image reconstruction process. Through such a residual module design, combined with the channel attention mechanism and the spatial attention mechanism, our image super-resolution reconstruction method can more accurately learn the important features in the image and make full use of the global context information during the reconstruction process. This comprehensive application of feature extraction and attention mechanism enables our method to generate higher-quality and more realistic super-resolution images.
7. The image super-resolution method based on attention mechanism and dual-channel network according to claim 1, characterized in that a non-local dilated convolution block is introduced. This module consists of four parallel dilated convolution layers with dilation parameters of 1, 2, 4, and 6 respectively and three ordinary convolutional neural network layers. During the image data processing, first, we use four dilated convolutions with different dilation parameters and two ordinary convolutional neural networks to extract feature information from the dependency information input to the improved residual network respectively. Such a design allows the model to capture multi-scale features of the image in parallel in multiple dilated convolution layers and extract more global and high-level features in the ordinary convolutional neural network. Then, we fuse the feature information obtained from the four dilated convolutions on the feature channels, while the feature information extracted by the ordinary convolutional neural network is fused according to the values of the pixel matrix. Such a fusion method makes full use of the feature information extracted by dilated convolution and ordinary convolution at different levels, enhancing the model's comprehensive expression ability for image details and global information. Finally, we add these two fused feature information to obtain global feature information. This process enables the model to understand the content of the image more comprehensively and provides a more accurate and rich feature representation for image super-resolution reconstruction. By introducing such a non-local dilated convolution block, our method can more effectively capture multi-scale information in the image, improving the quality and realism of image reconstruction.
8. The image super-resolution reconstruction method based on attention mechanism and dual-channel network according to claim 1, characterized in that it introduces an improved VGG network structure. This improved VGG network structure includes 10 ordinary convolutional layers and 3 pooling layers. We embed these pooling layers into the ordinary convolutional layers to form the VGG network structure. During the image data processing, first, we use 2 convolutional layers and 1 pooling operation to extract feature information of 64 channels. Then, we use another 2 convolutional layers and 1 pooling operation to extract feature information of 128 channels. Then, we extract feature information of 512 channels through 3 convolutional layers and 1 pooling operation. Finally, we use 3 convolutional layers to convert the 512-channel information back to 64 channels. It should be particularly noted that when performing the pooling operation, we use padding to keep the feature scale unchanged. Such a design ensures that important information of the image will not be lost during the pooling process, and at the same time allows the VGG network to extract multi-scale features of the image at different levels. By introducing such an improved VGG network structure, our image super-resolution reconstruction method can make more full use of the hierarchical feature representation ability of the VGG network, improving the abstraction and understanding of the image content. This improvement helps to obtain a richer and more accurate image feature expression, thus enhancing the effect and quality of image super-resolution reconstruction.
9. The image super-resolution reconstruction method based on the attention mechanism and the dual-channel network according to claim 1, characterized in that the expression of the loss function of the image super-resolution reconstruction model is: Among them, θ represents the number of parameters of the model, and C HR represents the super-resolution calculation equation, and I i LR and I i HR respectively represent the i-th low-resolution image and the corresponding i-th high-resolution image. N represents the number of images in the dataset. HR represents high resolution, and LR represents low resolution. This loss function is a mathematical expression used when training an image super-resolution reconstruction model to measure the difference between the reconstructed image and the original high-definition image. By optimizing this loss function, our method can better guide the learning process of the model, making the generated super-resolution image closer to the real high-definition image, thereby improving the quality and accuracy of image reconstruction.
10. The image super-resolution reconstruction method based on the attention mechanism and the dual-channel network according to claim 1, characterized in that, we use Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM) as the formulas for evaluating the reconstructed image. These evaluation metrics are mathematical expressions used to measure the quality and similarity between the reconstructed image and the original high-definition image. By using PSNR and SSIM evaluations, we can objectively evaluate the accuracy and fidelity of the image super-resolution reconstruction results, further guide the optimization and improvement of the algorithm, and ensure higher-quality super-resolution images.