Pixel-level image super-resolution reconstruction method, system, equipment and medium

The shallow feature map is generated through convolution and nonlinear processing, combined with global local feature extraction and multi-branch dual-path fusion, the information imbalance problem of image super-resolution reconstruction methods in the prior art is solved, and efficient high-resolution image reconstruction is achieved.

CN120339078AActive Publication Date: 2025-07-18YANTAI UNIV
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202510828533.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-18
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

Existing image super-resolution reconstruction methods lack long-distance modeling capabilities, resulting in local and global information imbalance, insufficient texture and high-frequency details recovery, and poor image reconstruction effect.

Method used

Convolution and nonlinear processing are used to generate shallow feature maps, and the global local feature extraction fusion and nonlinear feature map fusion are combined with pixel-level non-local self-attention processing and multi-branch dual-path fusion to achieve cross-scale feature fusion, and finally generate high-resolution images through subpixel reconstruction.

Benefits of technology

Improve image detail recovery and resolution, significantly improve reconstruction effect and image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339078A_ABST
    Figure CN120339078A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image reconstruction, in particular to a pixel-level image super-resolution reconstruction method, system and device and a medium. In order to solve the technical problem of poor image reconstruction effect in the prior art, the method comprises the following steps: firstly, performing convolution processing and nonlinear processing on a to-be-processed low-resolution image to obtain a shallow feature map; then, performing global local feature extraction fusion processing on the shallow feature map for several times to obtain a global local fusion feature map; carrying out convolution processing on the global and local fusion feature map, and fusing the global and local fusion feature map with the shallow feature map to obtain a deep feature map; finally, the deep feature map is subjected to sub-pixel-based image reconstruction processing, and a high-resolution reconstructed image is obtained.The image super-resolution reconstruction method is used in the field of image super-resolution reconstruction, image detail restoration is good, and image reconstruction quality is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image reconstruction, and specifically to a pixel-level image super-resolution reconstruction method, system, device, and medium. Background Art

[0002] High-resolution images have important application values in many fields such as medical imaging, remote sensing, security monitoring, and autonomous driving. However, directly obtaining high-resolution images through sensors often faces challenges, restricted by various factors such as the physical size of the sensor, pixel density, optical aberration, motion blur, transmission bandwidth, storage limitation, and cost. For example, the commonly used sensors in miniaturized devices have a small size, resulting in a decrease in the ability of pixels to capture light and a reduction in the signal-to-noise ratio; increasing pixel density can theoretically improve the resolution, but it causes an aggravated noise problem; the inherent defects of the optical system and motion blur in dynamic scenes further limit the image quality; in addition, high-resolution data also poses higher requirements for transmission and storage. These hardware-level limitations make it an attractive solution to improve the image resolution through software algorithms.

[0003] Image super-resolution reconstruction technology aims to reconstruct high-resolution images from low-resolution observed images, providing an economical and efficient way to solve the above-mentioned hardware limitations. Currently, image super-resolution reconstruction methods include regularization methods, convolutional neural network methods, and Transformer methods. However, these existing methods have limited ability to model long-range dependencies and excessively high global computational costs, resulting in an imbalance between local and global information, limited enhancement effects on high-frequency components such as textures and edges, insufficient recovery of high-frequency details, and poor image reconstruction effects, making the reconstructed images lack a sense of reality. Summary of the Invention

[0004] The purpose of the present invention is to provide a pixel-level image super-resolution reconstruction method, system, device, and medium.

[0005] The technical solution of the present invention is as follows: A pixel-level image super-resolution reconstruction method includes the following operations: S1. The low-resolution image to be processed undergoes convolution and nonlinear processing to obtain a shallow feature map; S2. The shallow feature map undergoes several global-local feature extraction and fusion processes to obtain a global-local fusion feature map; after the global-local fusion feature map undergoes convolution processing, it is fused with the shallow feature map to obtain a deep feature map; The operation of global-local feature extraction and fusion processing is as follows: The input feature map is processed by convolution, batch normalization, and activation function to obtain a non-linear feature map; the non-linear feature map is processed by pixel-level non-local self-attention to obtain an attention feature map; the attention feature map is fused with the input feature map to obtain an attention residual feature map; the attention residual feature map is processed by layer normalization to obtain an attention normalized feature map; the attention normalized feature map is processed by multi-branch dual-path fusion to obtain a cross-branch multi-scale fusion feature; the cross-branch multi-scale fusion feature is fused with the attention normalized feature map to obtain an output feature map, which is used to perform global-local feature extraction and fusion processing or convolution processing; The operation of pixel-level non-local self-attention processing is as follows: Each non-overlapping sub-feature map of the non-linear feature map is respectively processed by multi-head attention based on pixel similarity and then fused with its respective sub-feature map to obtain a number of multi-head attention sub-feature maps; Each multi-head attention sub-feature map is sequentially processed by feature enhancement, feature information interaction, and non-linear feature extraction, and then concatenated to obtain an attention feature map; S3. The deep feature map is processed by sub-pixel based image reconstruction to obtain a high-resolution reconstructed image.

[0006] In S2, the operation of multi-head attention processing based on pixel similarity is as follows: The pixel information within the neighborhood range of each pixel point in the sub-feature map forms its respective neighborhood pixel block; Select a pixel point from the sub-feature map as the reference pixel point, and the corresponding neighborhood pixel block as the reference neighborhood pixel block; Obtain the similarity between the reference neighborhood pixel block and each neighborhood pixel block under the current attention head in the sub-feature map as the similarity between the reference pixel point of the current head and each pixel point; The similarity between the reference pixel point of the current head and each pixel point is multiplied by the respective corresponding pixel embedding vector and then summed to obtain the attention feature of the current head; The attention features of all heads are weighted and summed and then regularized to obtain the initial attention feature map of the sub-feature map, which is used to perform the operation of fusing with the sub-feature map.

[0007] The similarity between the reference neighborhood pixel block and the neighborhood pixel block is calculated by the following formula: , For the i th attention head, the similarity between the reference neighborhood pixel block and the p th neighborhood pixel block , , are the embedding vectors of the reference neighborhood pixel block and the neighborhood pixel block respectively, , are the feature dimension and the total number of attention heads respectively, is the normalization function.

[0008] The operation of cross-branch multi-scale fusion feature in S2 is as follows: the attention-normalized feature map is subjected to feature extraction with different dilation rates to obtain several attention branch feature maps; the current attention branch feature map is respectively subjected to average pooling and max pooling to obtain the current branch average pooling feature map and the current branch max pooling feature map; the current branch average pooling feature map and the current branch max pooling feature map are subjected to weight generation processing based on parallel convolution to obtain the current first-path weight matrix; the current branch average pooling feature map and the current branch max pooling feature map are subjected to weight generation processing based on fusion convolution to obtain the current second-path weight matrix; the current first-path weight matrix and the current second-path weight matrix are multiplied element by element to obtain the current comprehensive weight matrix; each attention branch feature map is respectively weighted with its respective comprehensive weight matrix and then subjected to convolution processing to obtain the cross-branch multi-scale fusion feature.

[0009] The second-path weight matrix is calculated by the following formula: , is the m second-path weight matrix of the th attention branch feature map, m and are respectively the branch average pooling feature map and the branch max pooling feature map of the th n attention branch feature map, is the concatenation operation, and are the first convolution of the second path and the second convolution of the second path respectively, N is the total number of attention branch feature maps.

[0010] The operation of feature enhancement in S2 is specifically as follows: the multi-head attention sub-feature map is subjected to layer normalization, fully connected feed-forward network processing and regularization processing, and then added element by element to the multi-head attention sub-feature map to obtain the multi-head attention enhanced sub-feature map for performing the operation of feature information interaction.

[0011] The sub-pixel-based image reconstruction processing in S3 can be realized through convolution, sub-pixel convolution, activation function processing and convolution processing.

[0012] A pixel-level image super-resolution reconstruction system for implementing the above-mentioned pixel-level image super-resolution reconstruction method includes: A shallow feature map generation module, which is used to obtain a shallow feature map by performing convolution and non-linear processing on the to-be-processed low-resolution image; A deep feature map generation module, which is used to obtain a global-local fusion feature map by performing several times of global-local feature extraction and fusion processing on the shallow feature map; after the global-local fusion feature map is processed by convolution, it is fused with the shallow feature map to obtain a deep feature map; the operation of the global-local feature extraction and fusion processing is as follows: the input feature map is processed by convolution, batch normalization and an activation function to obtain a non-linear feature map; the non-linear feature map is processed by pixel-level non-local self-attention to obtain an attention feature map; the attention feature map is fused with the input feature map to obtain an attention residual feature map; the attention residual feature map is processed by layer normalization to obtain an attention-normalized feature map; the attention-normalized feature map is processed by multi-branch dual-path fusion to obtain a cross-branch multi-scale fusion feature; the cross-branch multi-scale fusion feature is fused with the attention-normalized feature map to obtain an output feature map, which is used to perform global-local feature extraction and fusion processing or convolution processing; the operation of the pixel-level non-local self-attention processing is as follows: each non-overlapping sub-feature map of the non-linear feature map is respectively processed by multi-head attention based on pixel similarity and then fused with its respective sub-feature map to obtain several multi-head attention sub-feature maps; each multi-head attention sub-feature map is respectively processed by feature enhancement, feature information interaction and non-linear feature extraction in sequence and then stitched together to obtain an attention feature map; A high-resolution reconstructed image generation module, which is used to obtain a high-resolution reconstructed image by performing sub-pixel-based image reconstruction processing on the deep feature map.

[0013] An image super-resolution reconstruction device at the pixel level, including a processor and a memory. Among them, when the processor executes the computer program stored in the memory, the above-mentioned image super-resolution reconstruction method at the pixel level is implemented.

[0014] A computer-readable storage medium is used to store a computer program. Among them, when the computer program is executed by a processor, the above-mentioned image super-resolution reconstruction method at the pixel level is implemented.

[0015] The beneficial effects of the present invention are as follows: A pixel-level image super-resolution reconstruction method provided by the present invention. First, the low-resolution image to be processed is subjected to convolution processing and non-linear processing to capture the low-level spatial information of the low-resolution image to be processed, and a shallow feature map is obtained. Then, the shallow feature map is subjected to several global-local feature extraction and fusion processes to fully capture the long-distance dependence and non-local feature information, enhance the multi-scale feature representation ability, and achieve efficient fusion of information at different scales, resulting in a global-local fusion feature map. After the global-local fusion feature map is subjected to convolution processing, it is fused with the shallow feature map to obtain a deep feature map. Finally, the deep feature map is subjected to sub-pixel-based image reconstruction processing to refine the reconstructed image, realize the reconstruction of the low-resolution image, and obtain a high-resolution reconstructed image. This method is used in the field of image super-resolution reconstruction, with good image detail restoration, high image resolution, and good image reconstruction effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] By reading the detailed description of the preferred embodiments below, the solutions and advantages of the present application will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention.

[0017] In the drawings: Figure 1 It is the image reconstruction effect diagram of the method in this embodiment for the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the drawings.

[0019] This embodiment provides a pixel-level image super-resolution reconstruction method, including the following operations: S1. The low-resolution image to be processed is subjected to convolution and non-linear processing to obtain a shallow feature map; S2. The shallow feature map is subjected to several global-local feature extraction and fusion processes to obtain a global-local fusion feature map; after the global-local fusion feature map is subjected to convolution processing, it is fused with the shallow feature map to obtain a deep feature map; S3. The deep feature map is subjected to sub-pixel-based image reconstruction processing to obtain a high-resolution reconstructed image.

[0020] The specific step details are as follows.

[0021] S1. The low-resolution image to be processed is subjected to convolution and non-linear processing to obtain a shallow feature map.

[0022] The low-resolution image to be processed is subjected to 9×9 convolution processing and non-linear processing (which can be implemented by the PReLU activation function) to capture the low-level spatial information of the low-resolution image to be processed, which is used as the initial representation for enhancing the resolution of the input image, and a shallow feature map is obtained.

[0023] The shallow feature map is subjected to several global-local feature extraction and fusion processes to obtain a global-local fusion feature map; after the global-local fusion feature map is subjected to convolutional processing, it is fused with the shallow feature map to obtain a deep feature map.

[0024] The shallow feature map is subjected to several global-local feature extraction and fusion processes to fully capture long-range dependencies and non-local feature information, enhance the multi-scale feature representation ability, and achieve efficient fusion of information at different scales, enabling better reconstruction of local high-frequency information in the image to obtain a global-local fusion feature map; after the global-local fusion feature map is subjected to convolutional processing, it is fused with the shallow feature map to obtain a deep feature map.

[0025] First, the shallow feature map is subjected to several global-local feature extraction and fusion processes to effectively extract and fuse global and local features to obtain a global-local fusion feature map.

[0026] Among them, the operation steps of the global-local feature extraction and fusion process are as follows.

[0027] Step 1: The input feature map (shallow feature map or the output of the previous global-local feature extraction and fusion process) is processed by convolution, batch normalization, and an activation function (which can be implemented by the PReLU activation function) to further enhance the feature representation and obtain a non-linear feature map.

[0028] Step 2: The non-linear feature map is processed by pixel-level non-local self-attention to capture long-range dependency relationships and non-local information in the image to obtain an attention feature map; the attention feature map is fused with the input feature map (which can be implemented by element-wise addition) to obtain an attention residual feature map.

[0029] The operation of pixel-level non-local self-attention processing is specifically as follows: the non-linear feature map is divided into several non-overlapping sub-feature maps, and each non-overlapping sub-feature map of the non-linear feature map is respectively processed by multi-head attention based on pixel similarity and then fused with its respective sub-feature map (which can be implemented by element-wise addition) to obtain several multi-head attention sub-feature maps; each multi-head attention sub-feature map is respectively subjected to feature enhancement, feature information interaction, and non-linear feature extraction in sequence, and then spliced to obtain an attention feature map.

[0030] A single sub-feature map. During the process of performing multi-head attention processing based on pixel similarity, the specific details are as follows: The pixel information within the neighborhood range of each pixel point in the sub-feature map forms their respective neighborhood pixel blocks; Select a pixel point from the sub-feature map as the reference pixel point, and the corresponding neighborhood pixel block as the reference neighborhood pixel block; Taking the current head attention as an example, obtain the similarity between the reference neighborhood pixel block and each neighborhood pixel block (including the reference neighborhood pixel block itself) under the current attention head of the sub-feature map (the similarity between the reference neighborhood pixel block and itself is 1), as the similarity between the current head reference pixel point and each pixel point; The similarity between the current head reference pixel point and each pixel point, after being multiplied by their respective corresponding pixel point embedding vectors, is summed to obtain the current head attention feature; All the head attention features are weighted and summed, and then normalized to obtain the initial attention feature map of the sub-feature map, which is used to perform the operation of fusing with the corresponding sub-feature map.

[0031] The similarity between the reference neighborhood pixel block and the neighborhood pixel blocks is calculated by the following formula: , For the i th attention head, between the reference neighborhood pixel block and the p th neighborhood pixel block , the similarity 、 are respectively the embedding vectors of the reference neighborhood pixel block , the embedding vector of the neighborhood pixel block , 、 are respectively the feature dimension and the total number of attention heads, is the normalization function.

[0032] The embedding vector of the neighborhood pixel block is obtained by performing depth convolution, pointwise convolution, activation function processing (which can be implemented by the LeakyReLU activation function), and flattening processing on the pixel information within the neighborhood range of the pixel point.

[0033] The pixel point embedding vector is obtained by performing pointwise convolution, activation function processing (which can be implemented by the ReLU activation function), and flattening processing on the pixel point information.

[0034] The operation of feature enhancement is specifically as follows: The multi-head attention sub-feature map, after being subjected to layer normalization, fully connected feed-forward network processing, and normalization processing, is added element-wise to the multi-head attention sub-feature map to obtain the multi-head attention enhanced sub-feature map, which is used to perform the operation of feature information interaction.

[0035] The operation of feature information interaction is specifically as follows: After the multi-head attention enhanced feature map is subjected to layer normalization and multi-head self-attention processing based on the inside of the window (preferably implemented through the SW-MSA network), it is added element-wise to the multi-head attention enhanced feature map to obtain a sub-feature interaction map, which is used to perform the operation of non-linear feature extraction.

[0036] The operation of non-linear feature extraction is specifically as follows: After the sub-feature interaction map is subjected to layer normalization and fully connected feed-forward network processing, it is added element-wise to the sub-feature interaction map to obtain an enhanced feature map, which is used to perform the operation of concatenating with other enhanced feature maps.

[0037] Step 3: The attention residual feature map is subjected to layer normalization processing to obtain an attention normalized feature map; The attention normalized feature map is subjected to multi-branch dual-path fusion processing to capture context information at different scales, solve the challenge of scale variation inherent in complex image processing tasks, and achieve information complementarity through the dual path to obtain a cross-branch multi-scale fusion feature; The cross-branch multi-scale fusion feature is fused with the attention normalized feature map (which can be achieved by element-wise addition) to obtain an output feature map, which is used to perform the next global-local feature extraction fusion processing or convolution processing.

[0038] The operation steps of the above cross-branch multi-scale fusion feature are as follows.

[0039] Step a: The attention normalized feature map is subjected to feature extraction with different dilation rates (which can be achieved by dilation convolution processing with different dilation rates and ReLU activation function processing) to obtain several attention branch feature maps.

[0040] Step b: Taking the current attention branch feature map among several attention branch feature maps as an example, the current attention branch feature map is respectively subjected to average pooling and max pooling to obtain the current branch average pooling feature map and the current branch max pooling feature map. The current branch average pooling feature map and the current branch max pooling feature map are subjected to weight generation processing based on parallel convolution to emphasize the most informative features and suppress irrelevant noise to obtain the current first path weight matrix; The current branch average pooling feature map and the current branch max pooling feature map are subjected to weight generation processing based on fusion convolution to control the contribution of each branch during the aggregation process to obtain the current second path weight matrix; The current first path weight matrix and the current second path weight matrix are multiplied element-wise to obtain the current comprehensive weight matrix.

[0041] Each attention branch feature map respectively performs the operations of obtaining the first path weight matrix and the second path weight matrix, and multiplying the first path weight matrix and the second path weight matrix to obtain the comprehensive weight matrix of each attention branch feature map.

[0042] The first-path weight matrix or the weight generation process based on parallel convolution is calculated or implemented by the following formula: , is the first-path weight matrix of the m th attention branch feature map, , are respectively the branch average pooling feature map and the branch max pooling feature map of the m th attention branch feature map, , are the first convolution of the first path and the second convolution of the first path, and the convolution scale is 1×1, is the Sigmoid activation function.

[0043] The second-path weight matrix or the weight generation process based on fusion convolution is calculated or implemented by the following formula: , is the second-path weight matrix of the m th attention branch feature map, , are respectively the branch average pooling feature map and the branch max pooling feature map of the m th attention branch feature map, , are respectively the branch average pooling feature map and the branch max pooling feature map of the n th attention branch feature map, is the concatenation operation, , are the first convolution and the second convolution, and the convolution scale is 1×1, N is the total number of attention branch feature maps.

[0044] Step c: Each attention branch feature map is respectively weighted with its respective comprehensive weight matrix and then subjected to convolution processing to obtain cross-branch multi-scale fusion features.

[0045] Next, the global-local fusion feature map is subjected to convolution processing and then fused with the shallow feature map (which can be achieved by element-wise addition) to obtain a deep feature map.

[0046] S3. The deep feature map is subjected to sub-pixel-based image reconstruction processing to refine the reconstructed image, realizing the reconstruction of the low-resolution image to obtain a high-resolution reconstructed image. See the result graph in Figure 1 . The image has good detail restoration and high image reconstruction quality. The sub-pixel-based image reconstruction processing can be achieved through convolution, sub-pixel convolution, activation function processing (which can be implemented by the PReLU activation function), and convolution processing.

[0047] The image reconstruction method of this embodiment can also establish models corresponding to steps S1, S2, and S3, use a training set formed by several pairs of high-resolution images and low-resolution images to train the models. The training is performed for K rounds of training times, and a loss function is used to iteratively optimize the models. Observe the convergence status of the loss function of the models. When the models converge to the best, an image reconstruction training model is obtained; the low-resolution image to be processed is put into the image reconstruction training model for processing to obtain a high-resolution reconstructed image.

[0048] Among them, in the process of making the training set, the low-resolution images are obtained by performing bicubic interpolation processing, and / or random horizontal flipping, and / or 90° rotation, and / or 270° rotation, and / or random cropping on the corresponding high-resolution images. All high-resolution images and low-resolution images in the training set are normalized.

[0049] This embodiment also provides an image super-resolution reconstruction system at the pixel level for implementing the above-mentioned image super-resolution reconstruction method at the pixel level, including: A shallow feature map generation module, which is used to perform convolution and non-linear processing on the low-resolution image to be processed to obtain a shallow feature map; A deep feature map generation module, which is used to perform several times of global-local feature extraction and fusion processing on the shallow feature map to obtain a global-local fusion feature map; after the global-local fusion feature map is subjected to convolution processing, it is fused with the shallow feature map to obtain a deep feature map; the operation of the global-local feature extraction and fusion processing is: the input feature map is processed by convolution, batch normalization, and an activation function to obtain a non-linear feature map; the non-linear feature map is processed by pixel-level non-local self-attention to obtain an attention feature map; the attention feature map is fused with the input feature map to obtain an attention residual feature map; the attention residual feature map is processed by layer normalization to obtain an attention normalized feature map; the attention normalized feature map is processed by multi-branch dual-path fusion to obtain a cross-branch multi-scale fusion feature; the cross-branch multi-scale fusion feature is fused with the attention normalized feature map to obtain an output feature map for performing global-local feature extraction and fusion processing or convolution processing; the operation of the pixel-level non-local self-attention processing is: each non-overlapping sub-feature map of the non-linear feature map is respectively processed by multi-head attention based on pixel similarity and then fused with its respective sub-feature map to obtain several multi-head attention sub-feature maps; each multi-head attention sub-feature map is respectively subjected to feature enhancement, feature information interaction, and non-linear feature extraction in sequence, and then spliced to obtain an attention feature map; A high-resolution reconstructed image generation module, which is used to perform sub-pixel-based image reconstruction processing on the deep feature map to obtain a high-resolution reconstructed image.

[0050] This embodiment also provides an image super-resolution reconstruction device at the pixel level, including a processor and a memory. When the processor executes the computer program stored in the memory, the above-mentioned image super-resolution reconstruction method at the pixel level is implemented.

[0051] This embodiment also provides a computer-readable storage medium for storing a computer program. When the computer program is executed by a processor, the above-mentioned image super-resolution reconstruction method at the pixel level is implemented.

[0052] A method for image super-resolution reconstruction at the pixel level provided by this embodiment of the invention. First, the low-resolution image to be processed is subjected to convolution processing and nonlinear processing to capture the low-level spatial information of the low-resolution image to be processed, and a shallow feature map is obtained. Then, the shallow feature map is subjected to several global-local feature extraction and fusion processes to fully capture long-distance dependencies and non-local feature information, enhance the multi-scale feature representation ability, and achieve efficient fusion of information at different scales, obtaining a global-local fusion feature map. After the global-local fusion feature map is subjected to convolution processing, it is fused with the shallow feature map to obtain a deep feature map. Finally, the deep feature map is subjected to sub-pixel-based image reconstruction processing to refine the reconstructed image, realize the reconstruction of the low-resolution image, and obtain a high-resolution reconstructed image. This method is used in the field of image super-resolution reconstruction, with good image detail restoration, high image resolution, and good image reconstruction quality.

Claims

1. A pixel-level image super-resolution reconstruction method, characterized in that, Including the following operations: S1. The low-resolution image to be processed undergoes convolution and non-linear processing to obtain a shallow feature map; S2. The shallow feature map undergoes several global-local feature extraction and fusion processes to obtain a global-local fusion feature map; after the global-local fusion feature map undergoes convolution processing, it is fused with the shallow feature map to obtain a deep feature map; The operation of the global-local feature extraction and fusion process is: the input feature map undergoes convolution, batch normalization, and activation function processing to obtain a non-linear feature map; The non-linear feature map undergoes pixel-level non-local self-attention processing to obtain an attention feature map; The attention feature map is fused with the input feature map to obtain an attention residual feature map; The attention residual feature map undergoes layer normalization processing to obtain an attention normalized feature map; The attention normalized feature map undergoes multi-branch and dual-path fusion processing to obtain a cross-branch multi-scale fusion feature; The cross-branch multi-scale fusion feature is fused with the attention normalized feature map to obtain an output feature map, which is used to perform the global-local feature extraction and fusion process or convolution processing; The operation of the pixel-level non-local self-attention processing is: each non-overlapping sub-feature map of the non-linear feature map undergoes multi-head attention processing based on the similarity between pixels, and then is fused with its respective sub-feature map to obtain several multi-head attention sub-feature maps; each multi-head attention sub-feature map undergoes feature enhancement, feature information interaction, and non-linear feature extraction in sequence, and then is concatenated to obtain an attention feature map; S3. The deep feature map undergoes sub-pixel based image reconstruction processing to obtain a high-resolution reconstructed image.

2. The pixel-level image super-resolution reconstruction method according to claim 1, wherein In S2, the operation of the multi-head attention processing based on the similarity between pixels is: The pixel information within the neighborhood range of each pixel point in the sub-feature map forms its respective neighborhood pixel block; Select a pixel point from the sub-feature map as the reference pixel point, and the corresponding neighborhood pixel block as the reference neighborhood pixel block; Obtain the similarity between the reference neighborhood pixel block and each neighborhood pixel block in the sub-feature map under the current attention head, as the similarity between the current head reference pixel point and each pixel point; The similarity between the current head reference pixel point and each pixel point is multiplied by the respective pixel point embedding vector, and then summed to obtain the current head attention feature; After the attention features of all heads undergo weighted sum processing, they are regularized to obtain the initial attention feature map of the sub-feature map, which is used to perform the operation of fusing with the sub-feature map.

3. The pixel-level image super-resolution reconstruction method according to claim 2, characterized in that The similarity between the reference neighborhood pixel block and the neighborhood pixel blocks is calculated through the following formula: , To calculate the similarity between the reference neighborhood pixel block i under the -th attention head and the p -th neighborhood pixel block, where and are the embedding vectors of the reference neighborhood pixel block and the neighborhood pixel block respectively, and are the feature dimension and the total number of attention heads respectively, is the normalization function.

4. The pixel-level image super-resolution reconstruction method according to claim 1, wherein In S2, the operation of the cross-branch multi-scale fusion feature is: The attention normalized feature map undergoes feature extraction with different dilation rates to obtain several attention branch feature maps; The current attention branch feature map undergoes average pooling and max pooling respectively to obtain the current branch average pooling feature map and the current branch max pooling feature map; The current branch average pooling feature map and the current branch max pooling feature map are processed by weight generation based on parallel convolution to obtain the current first path weight matrix; the current branch average pooling feature map and the current branch max pooling feature map are processed by weight generation based on fused convolution to obtain the current second path weight matrix; the current first path weight matrix and the current second path weight matrix are multiplied element-wise to obtain the current comprehensive weight matrix; Each attention branch feature map and its respective comprehensive weight matrix are weighted and then convolved to obtain the cross-branch multi-scale fusion feature.

5. The pixel-level image super-resolution reconstruction method according to claim 4, wherein The second path weight matrix is calculated by the following formula: , is the second path weight matrix of the m th attention branch feature map, , are respectively the branch average pooling feature map and the branch max pooling feature map of the m th attention branch feature map, , are respectively the branch average pooling feature map and the branch max pooling feature map of the n th attention branch feature map, is the concatenation operation, , are the first convolution of the second path and the second convolution of the second path, N is the total number of attention branch feature maps.

6. The pixel-level image super-resolution reconstruction method according to claim 1, wherein In S2, the operation of feature enhancement is specifically: The multi-head attention sub-feature map is processed by layer normalization, fully connected feed-forward network processing, and regularization processing, and then added element-wise to the multi-head attention sub-feature map to obtain the multi-head attention enhanced sub-feature map for performing the operation of feature information interaction.

7. The pixel-level image super-resolution reconstruction method according to claim 1, wherein In S3, the sub-pixel-based image reconstruction processing can be achieved through convolution, sub-pixel convolution, activation function processing, and convolution processing.

8. An image super-resolution reconstruction system at the pixel level, for implementing the image super-resolution reconstruction method at the pixel level according to claim 1, characterized in that, Including: A shallow feature map generation module for obtaining a shallow feature map by subjecting the to-be-processed low-resolution image to convolution and non-linear processing; A deep feature map generation module for obtaining a global-local fusion feature map by performing several global-local feature extraction and fusion processes on the shallow feature map; after the global-local fusion feature map is convolved, it is fused with the shallow feature map to obtain a deep feature map; the operation of global-local feature extraction and fusion processing is: the input feature map is processed by convolution, batch normalization, and activation function to obtain a non-linear feature map; The non-linear feature map is processed by pixel-level non-local self-attention to obtain an attention feature map; The attention feature map is fused with the input feature map to obtain an attention residual feature map; The attention residual feature map is processed by layer normalization to obtain an attention normalized feature map; The attention normalized feature map is processed by multi-branch double-path fusion to obtain a cross-branch multi-scale fusion feature; The cross-branch multi-scale fusion feature is fused with the attention normalized feature map to obtain an output feature map for performing global-local feature extraction and fusion processing or convolution processing; the operation of pixel-level non-local self-attention processing is: each non-overlapping sub-feature map of the non-linear feature map is respectively processed by multi-head attention based on pixel similarity and then fused with its respective sub-feature map to obtain several multi-head attention sub-feature maps; each multi-head attention sub-feature map is respectively processed by feature enhancement, feature information interaction, and non-linear feature extraction in sequence and then concatenated to obtain an attention feature map; A high-resolution reconstructed image generation module for obtaining a high-resolution reconstructed image by subjecting the deep feature map to sub-pixel-based image reconstruction processing.

9. An image super-resolution reconstruction device at the pixel level, characterized in that, Including a processor and a memory, wherein when the processor executes the computer program stored in the memory, it implements the pixel-level image super-resolution reconstruction method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, For storing a computer program, wherein when the computer program is executed by a processor, it implements the pixel-level image super-resolution reconstruction method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Image super-resolution reconstruction model and method based on residual mixed attention network

    CN115222601A

  • Face image super-resolution reconstruction method, system, equipment and medium

    CN115393186A

  • Infrared image super-resolution reconstruction method based on edge enhancement

    CN116071243A

  • Image super-resolution model based on lightweight mixed attention distillation network

    CN118195899A

  • Transform and CNN hybrid network-based sequence image super-resolution reconstruction method

    CN118429188A