Arbitrary scale image super-resolution reconstruction method and system based on dual attention mechanism

Through the image super-resolution reconstruction method based on the dual attention mechanism, the problems of computational resource waste and performance bottlenecks in the existing technology are solved, efficient detail generation of super-resolution reconstruction of images of any scale is achieved, and the image detail performance is enhanced.

CN120410858BActive Publication Date: 2025-09-05GUANGDONG OCEAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510902186.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-09-05
Estimated Expiration
2045-07-01

AI Technical Summary

Technical Problem

Existing technologies suffer from computational resource waste and performance bottlenecks in super-resolution reconstruction of images of arbitrary scales, making it difficult to meet the diverse needs of non-integer or asymmetric scale reconstruction. In addition, traditional methods have limited ability to generate details at unseen continuous scales.

Method used

An arbitrary-scale image super-resolution reconstruction method based on a dual-attention mechanism is adopted. Through shallow and deep feature extraction modules, feature aggregation modules and arbitrary-scale upsampling modules, combined with scale-aware spatial and channel attention modules, feature information of different scales is adaptively processed to generate high-resolution reconstructed images.

Benefits of technology

It improves the feature representation capability in super-resolution tasks of any scale, enhances the image detail expression, and is suitable for a variety of application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120410858B_ABST
    Figure CN120410858B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of machine vision technology, and in particular to a method and system for super-resolution reconstruction of arbitrary-scale images based on a dual-attention mechanism. The method comprises: establishing a super-resolution reconstruction model, wherein the super-resolution reconstruction model comprises a shallow feature extraction module, a deep feature extraction module, a feature aggregation module, and an arbitrary-scale upsampling module; obtaining a data set, and training the super-resolution reconstruction model using the data set to obtain a trained super-resolution reconstruction model; wherein the data set comprises multiple training samples, and the training samples comprise low-resolution images and corresponding ground truth values; obtaining a low-resolution image to be reconstructed, reconstructing the low-resolution image to be reconstructed using the trained super-resolution reconstruction model, and outputting a high-resolution reconstructed image; the present invention can improve the feature representation capability in arbitrary-scale super-resolution tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine vision technology, and in particular to a method and system for super-resolution reconstruction of images of any scale based on a dual-attention mechanism. Background Art

[0002] Super-resolution reconstruction of arbitrary-scale images is a key technology in image processing, aiming to improve image resolution and enhance detail. Traditional methods are typically designed for specific integer magnification ratios, requiring separate training models for different scales. This results in inefficient use of computing resources and storage space, making it difficult to meet the diverse needs of non-integer or asymmetric scale reconstruction in real-world scenarios.

[0003] In addition, the mainstream single-image super-resolution framework relies on a fixed upsampling module, whose discretization characteristics limit the model's ability to generate details at unseen continuous scales, while the strategy of independent training of multi-scale tasks ignores the strong correlation between different scales, which easily leads to performance bottlenecks. Summary of the Invention

[0004] To solve the above problems, the present invention provides a super-resolution reconstruction method and system for arbitrary-scale images based on a dual-attention mechanism, aiming to effectively improve the feature representation capability in arbitrary-scale super-resolution tasks.

[0005] In order to achieve the above object, the present invention provides the following technical solutions:

[0006] In one aspect, an embodiment of the present invention provides a method for super-resolution reconstruction of images of any scale based on a dual attention mechanism, the method comprising the following steps:

[0007] Establishing a super-resolution reconstruction model, the super-resolution reconstruction model includes a shallow feature extraction module, a deep feature extraction module, a feature aggregation module and an arbitrary scale upsampling module; extracting shallow features from the low-resolution image by the shallow feature extraction module, and extracting a channel feature map and multiple intermediate feature maps from the shallow features by the deep feature extraction module; aggregating the channel feature map and the multiple intermediate feature maps by the feature aggregation module to generate an output feature map; processing the output feature map and the upsampling scale factor by the arbitrary scale upsampling module to output a high-resolution reconstructed image;

[0008] Acquire a data set, and train a super-resolution reconstruction model using the data set to obtain a trained super-resolution reconstruction model; wherein the data set includes a plurality of training samples, and the training samples include low-resolution images and corresponding ground truth values;

[0009] A low-resolution image to be reconstructed is obtained, and the low-resolution image to be reconstructed is reconstructed using a trained super-resolution reconstruction model to output a high-resolution reconstructed image.

[0010] Optionally, the training of the super-resolution reconstruction model using the data set to obtain a trained super-resolution reconstruction model includes:

[0011] A loss value of a super-resolution reconstruction model is determined based on the high-resolution reconstructed image and the real image corresponding to the low-resolution image, and a parameter set of the super-resolution reconstruction model is iteratively updated based on the loss value to obtain a trained super-resolution reconstruction model.

[0012] Optionally, the loss value is calculated using the following formula:

[0013] ;

[0014] in, represents the parameter set of the super-resolution reconstruction model, M represents the number of training samples, is the low-resolution image in the j-th sample, is the real image corresponding to the low-resolution image in the jth sample, represents the hyperparameters of the super-resolution reconstruction model, represents the optimal hyperparameters of the super-resolution reconstruction model.

[0015] Optionally, the deep feature extraction module includes a plurality of cascaded RDB-SGDA modules, and the RDB-SGDA module includes a connected RDB module and an SGDA module;

[0016] The upsampling scale factor and the shallow feature map are stacked by the deep feature extraction module to generate a deep feature map and multiple intermediate feature maps, including:

[0017] The shallow feature map is input as the input feature map into the first cascaded RDB-SGDA module for stacking processing. The upsampling scale factor and the input feature map are processed by the RDB-SGDA module. After processing, the intermediate feature map of the first cascaded RDB-SGDA module is obtained. The intermediate feature map is used as the input feature map of the next cascaded RDB-SGDA module. After a set number of RDB-SGDA module stacking processes, the intermediate feature map obtained by the last stacking process is used as the deep feature map.

[0018] Optionally, the SGDA module includes a spatial attention module and a channel attention module, and the stacking process of the intermediate feature maps by the SGDA module to obtain the channel feature maps includes:

[0019] The intermediate feature map is processed by the spatial attention module and the spatial feature map is output;

[0020] The spatial feature map is processed through the channel attention module to output the channel feature map.

[0021] Optionally, processing the intermediate feature map by a spatial attention module to output a spatial feature map includes:

[0022] The intermediate feature map is input into the spatial attention module to generate a spatial attention mask; the intermediate feature map is input into the scale-aware deformable convolution layer to generate a scale-fused feature map with arbitrary scale information;

[0023] After element-wise multiplication of the scale fusion feature map and the spatial attention mask, it is adaptively fused with the intermediate feature map to output the spatial feature map.

[0024] Optionally, performing channel feature processing on the spatial feature map by a channel attention module and outputting the channel feature map includes:

[0025] After the upsampling scale factor is nonlinearly transformed by a multi-layer perceptron and processed by an activation function, two channel attention feature masks are generated;

[0026] After the spatial feature map is transformed by the convolution layer, it is element-wise multiplied with the two channel attention feature masks to obtain two channel weighted feature maps;

[0027] The two channel weighted feature maps and the spatial feature map are added through a residual connection to output a channel feature map.

[0028] On the other hand, an embodiment of the present invention provides an arbitrary-scale image super-resolution reconstruction system based on a dual-attention mechanism, comprising:

[0029] at least one processor;

[0030] at least one memory for storing at least one program;

[0031] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.

[0032] On the other hand, an embodiment of the present invention provides a computer-readable storage medium storing a program executable by a processor. When the program is executed by the processor, it is used to perform the above method.

[0033] The beneficial effects of the present invention are as follows: the present invention discloses a super-resolution reconstruction method and system for arbitrary-scale images based on a dual-attention mechanism, wherein the present invention extracts shallow features from low-resolution images through a shallow feature extraction module, and extracts channel feature maps and multiple intermediate feature maps from the shallow features through a deep feature extraction module; aggregates the channel feature map and multiple intermediate feature maps through a feature aggregation module to generate an output feature map; processes the output feature map and the upsampling scale factor through an arbitrary-scale upsampling module to output a high-resolution reconstructed image; and can effectively improve the feature representation capability in arbitrary-scale super-resolution tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0035] Figure 1 1 is a flow chart of a method for super-resolution reconstruction of images of any scale based on a dual attention mechanism according to an embodiment of the present invention;

[0036] Figure 2 is a diagram of the overall network structure of a super-resolution reconstruction model of any scale in an embodiment of the present invention;

[0037] Figure 3 It is a comparison of feature maps encoded at different scales;

[0038] Figure 4 is a structural diagram of a scale-guided dual attention module in an embodiment of the present invention;

[0039] Figure 5 is a structural diagram of a spatial attention module in an embodiment of the present invention;

[0040] Figure 6 This is a structural diagram of an arbitrary-scale image super-resolution reconstruction system based on a dual-attention mechanism according to an embodiment of the present invention. DETAILED DESCRIPTION

[0041] The following will be combined with the embodiments and drawings to clearly and completely describe the concept, specific structure and technical effects disclosed in the present invention, so as to fully understand the purpose, scheme and effect disclosed in the present invention. It should be noted that the embodiments and features in the embodiments of this application can be combined with each other unless there is a conflict.

[0042] refer to Figure 1 and Figure 2 ,like Figure 1The figure shows a method for super-resolution reconstruction of an arbitrary-scale image based on a dual-attention mechanism provided by an embodiment of the present invention, the method comprising the following steps:

[0043] S100, establishing a super-resolution reconstruction model, the super-resolution reconstruction model comprising a shallow feature extraction module, a deep feature extraction module, a feature aggregation module, and an arbitrary scale upsampling module; extracting shallow features from the low-resolution image by the shallow feature extraction module, extracting a channel feature map and multiple intermediate feature maps from the shallow features by the deep feature extraction module; aggregating the channel feature map and the multiple intermediate feature maps by the feature aggregation module to generate an output feature map; processing the output feature map and the upsampling scale factor by the arbitrary scale upsampling module to output a high-resolution reconstructed image;

[0044] Specifically, in the super-resolution reconstruction model, the shallow feature extraction module is used to extract shallow features from the low-resolution image, and the deep feature extraction module is used to extract channel feature maps and multiple intermediate feature maps from the shallow features; the feature aggregation module is used to aggregate the channel feature map and multiple intermediate feature maps to generate an output feature map; the arbitrary scale upsampling module is used to process the output feature map and the upsampling scale factor to output a high-resolution reconstructed image.

[0045] S200, obtaining a data set, and training a super-resolution reconstruction model using the data set to obtain a trained super-resolution reconstruction model; wherein the data set includes a plurality of training samples, and the training samples include low-resolution images and corresponding ground truth values;

[0046] S300 , obtaining a low-resolution image to be reconstructed, reconstructing the low-resolution image to be reconstructed using a trained super-resolution reconstruction model, and outputting a high-resolution reconstructed image.

[0047] The present invention uses a shallow feature extraction module to extract shallow features from low-resolution images, and a deep feature extraction module to extract channel feature maps and multiple intermediate feature maps from the shallow features. A feature aggregation module aggregates the channel feature maps and multiple intermediate feature maps to generate an output feature map. An arbitrary scale upsampling module processes the output feature map and the upsampling scale factor to output a high-resolution reconstructed image. This method effectively enhances image detail and visual effects, and is suitable for a variety of application scenarios.

[0048] Specifically, to address the limitations of existing arbitrary-scale super-resolution models in feature encoding capabilities, this paper innovatively proposes a plug-and-play dual-attention bottleneck module. This dual-attention bottleneck module consists of two key components: a scale-guided spatial attention module and a scale-guided channel attention module. The SGDA module (Scale-Guidance Dual Attention Block) exhibits strong scale adaptability and can be combined with mainstream single-image super-resolution network architectures (such as the RDN (Residual Dense Network), RCAN (Residual Channel Attention Network), and SwinSR (Image Restoration Using Swin Transformer), thereby significantly enhancing the model's feature representation capabilities for arbitrary-scale super-resolution tasks. This paper will elaborate on the specific implementation details of this module.

[0049] Considering that different scale factors can lead to different degradation characteristics, this paper proposes to use scale information as prior knowledge to guide feature learning, thereby improving image super-resolution performance. Existing research has shown that introducing scale factors into feature extraction networks can effectively enhance the feature encoding capabilities of the backbone network. To verify this, the paper embeds a scale-aware feature adaptation module based on ArbSR after each residual block of the Enhanced Deep Super-Resolution Network (EDSR) baseline network. Figure 3 The module visualizes feature maps at different network layers and upsampling scales (×2, 3, and 4). The brighter the feature map, the greater the feature activation at that location, indicating a higher network focus on that region. When the network depth reaches 16 layers, it is observed that the peak values ​​of the feature maps increase with the upsampling scale, indicating that deeper features contain richer spatial information as the scale increases. Comparing the brightness distribution of the last layer's feature maps at different scales reveals that the highlighted areas of feature maps at smaller scales (e.g., ×2) are smoother, indicating that the network focuses more on low-frequency information (e.g., overall structure) of the image, while the highlighted areas of feature maps at larger scales (e.g., ×4) are concentrated around edges and textures, indicating that the network prefers to extract high-frequency information (e.g., detailed features). This phenomenon reveals the influence of different upsampling scales on network feature learning: smaller scales rely more on low-frequency global information, while larger scales focus more on high-frequency local details. This phenomenon confirms the differential impact of different scale factors on feature encoding and lays the theoretical foundation for the proposed method.

[0050] Specifically, the overall network structure of the arbitrary scale super-resolution (SR) reconstruction model of the present invention is as follows: Figure 2 As shown in Figure 2, the network is divided into four parts: shallow feature extraction module, deep feature extraction module, feature aggregation module and arbitrary scale upsampling module. First, the low-resolution image and the paired ground truth are represented as and . Use convolutional networks to extract shallow features from low-resolution images .

[0051] In some embodiments, the shallow feature map is expressed as:

[0052] ;

[0053] in, represents the shallow feature map, Indicates the use of two layers of 3×3 convolutional layers for shallow feature extraction. Represents a low-resolution image.

[0054] The deep feature extraction module stacks the upsampling scale factor and the shallow feature map to generate multiple intermediate feature maps and channel feature maps, and uses the final output channel feature map as the deep feature map. and the inverse of the upsampling scale factors for height and width As input, deep features are generated through multiple RDB (Residual Dense Block) modules The deep feature extraction module innovatively embeds the SGDA module after each level of the RDB module, gradually enhancing the feature expression capability through this cascade structure.

[0055] After extracting multi-level features through a series of cascaded RDB-SGDA modules, this paper further introduces a dense feature fusion mechanism (DFF) to integrate feature representations at different levels. This mechanism consists of two key parts: global feature fusion and global residual learning. The global feature fusion (GFF) module adaptively aggregates all intermediate feature maps. , enhancing the complementarity of multi-scale features; while the Global Residual Learning Module (GRL) module retains shallow detail information through cross-layer skip connections to ensure effective gradient propagation.

[0056] In some embodiments, the expression of the feature aggregation module is:

[0057] ;

[0058] in, Indicates that through composite function The output feature map generated after aggregating features, After the feature concatenation, 1x1 convolution is performed for channel compression, and then 3×3 convolution is performed for feature fusion.

[0059] After completing the deep feature extraction in the low-resolution space, the present invention further realizes super-resolution reconstruction of any scale in the high-resolution space. In order to enable the network to adaptively process different magnification scales, the upsampling module converts the deep features into , height scale factor and width scale factor After being processed by the Scale-Arbitrary Upsampler (SAB), the high-resolution reconstructed image is finally output. .

[0060] In some embodiments, the expression of the high-resolution reconstructed image is:

[0061] ;

[0062] in, represents the high-resolution reconstructed image, The module dynamically adjusts network parameters to adapt to different magnification ratios, thereby achieving high-quality super-resolution reconstruction at any scale.

[0063] In some embodiments, the training of the super-resolution reconstruction model using the data set to obtain a trained super-resolution reconstruction model includes:

[0064] A loss value of a super-resolution reconstruction model is determined based on the high-resolution reconstructed image and the real image corresponding to the low-resolution image, and a parameter set of the super-resolution reconstruction model is iteratively updated based on the loss value to obtain a trained super-resolution reconstruction model.

[0065] The parameters of the super-resolution reconstruction model are updated based on the loss value until an iterative training condition is met to obtain a trained super-resolution reconstruction model; the iterative training condition includes that the loss value of the super-resolution reconstruction model is lower than a set threshold or the number of iterative training times is reached.

[0066] In some embodiments, the loss function is expressed as:

[0067] ;

[0068] in, represents the parameter set of the super-resolution reconstruction model, M represents the number of training samples, is the low-resolution image in the j-th sample, is the real image corresponding to the low-resolution image in the jth sample, represents the hyperparameters of the super-resolution reconstruction model, represents the optimal hyperparameters of the super-resolution reconstruction model.

[0069] Specifically, in order to optimize the training of the super-resolution reconstruction model, L1 loss is used as the loss function of the super-resolution reconstruction model to obtain the loss value.

[0070] In some embodiments, the deep feature extraction module includes a plurality of cascaded RDB-SGDA modules, wherein the RDB-SGDA module includes a connected RDB module and an SGDA module;

[0071] The upsampling scale factor and the shallow feature map are stacked by the deep feature extraction module to generate a deep feature map and multiple intermediate feature maps, including:

[0072] The shallow feature map F0 is input as the input feature map into the first cascaded RDB-SGDA module for stacking processing. The upsampling scale factor and the input feature map are processed by the RDB-SGDA module. After processing, the intermediate feature map of the first cascaded RDB-SGDA module is obtained. The intermediate feature map is used as the input feature map of the next cascaded RDB-SGDA module. After a set number of RDB-SGDA module stacking processes, the intermediate feature map obtained by the last stacking process is used as the deep feature map.

[0073] The expression of the deep feature map is:

[0074] ;

[0075] in, Represents the deep feature map, n represents the number of stacking of RDB module and SGDA module, , Respectively represent RDB modules and SGDA modules, and Respectively represent RDB modules and SGDA modules, , represents the height scale factor, Indicates the width scale factor.

[0076] It should be noted that after each stacking process of the residual dense block and the SGDA module, the intermediate feature map output by the i-th stacking process is recorded as , the final deep feature map is This cascade structure enables the network to adaptively fuse feature information of different scale factors and enhance the spatial and channel discrimination capabilities of features at different upsampling scales.

[0077] refer to Figure 4 and Figure 5 In some embodiments, the SGDA module includes a spatial attention module and a channel attention module, and the stacking process of the intermediate feature map by the SGDA module to obtain the channel feature map includes:

[0078] The intermediate feature map is processed by the spatial attention module and the spatial feature map is output;

[0079] The spatial feature map is processed through the channel attention module to output the channel feature map.

[0080] Specifically, the specific way to embed the SGDA module is as follows Figure 4 As shown. Based on a given backbone network (such as RDN), the dual attention bottleneck module can be flexibly embedded in the network architecture to achieve super-resolution reconstruction of any scale. Figure 2 Network structure and Figure 5 As shown in Figure 5, the dual-attention bottleneck module deploys two key components sequentially after every K backbone blocks: first, a spatial attention module guided by multi-scale features, followed by a channel attention module guided by the same scale.

[0081] In some embodiments, processing the intermediate feature map by the spatial attention module to output the spatial feature map includes:

[0082] The intermediate feature map is input into the spatial attention module to generate a spatial attention mask; the intermediate feature map is input into the scale-aware deformable convolution layer to generate a scale-fused feature map with arbitrary scale information;

[0083] After element-wise multiplication of the scale fusion feature map and the spatial attention mask, it is adaptively fused with the intermediate feature map to output the spatial feature map.

[0084] Given an intermediate feature map First, a spatial attention mask with a value range of [0, 1] is generated through a spatial attention module consisting of four convolutional layers and a Sigmoid activation function. At the same time, the intermediate feature map It is fed into the scale-aware deformable convolution layer to generate a scale fusion feature map with arbitrary scale information. Finally, the intermediate feature maps are fused through the attention-guided feature fusion mechanism. Scale fusion feature map after scale adjustment Perform adaptive fusion and output the feature map output by the scale-guided spatial attention module, namely the spatial feature map;

[0085] In some embodiments, the expression of the spatial feature map is:

[0086] ;

[0087] in, represents the spatial feature map, represents the i-th intermediate feature map, represents the scale fusion feature map, represents the spatial attention mask.

[0088] In this mechanism, the scale fusion feature map after scale factor fusion Spatial attention mask generated with intermediate feature maps Element-by-element multiplication is performed to achieve spatial difference encoding of features at different scales.

[0089] To ensure the integrity of the original feature information, a residual connection mechanism is also introduced in the spatial attention module.

[0090] For the intermediate feature map Every position on , given The deformable convolution kernel of sampling points (e.g. 3×3 convolution kernel), is the learned offset, is the modulation mask. A predefined regular grid offset satisfy .

[0091] In some embodiments, the fusion scale information can be deformed convolution position The scale fusion feature map output at is expressed as:

[0092] ;

[0093] in, Represents the intermediate feature map The position index on Represents the intermediate feature maps respectively On the positions and corresponding convolution kernel weights, Indicates location The scale fusion feature map output at and Respectively The offset and modulation amount for each position.

[0094] In the deformable convolution module, the convolution kernel weights Together with the dynamically adjusted sampling positions and feature importances, and Respectively The learnable offset scalar and modulation scalar for each position. Specifically, the modulation coefficient Generated by 3×3 convolution and Sigmoid activation function, it is used to adaptively suppress irrelevant areas and enhance key feature responses; it is worth noting that the offset Instead of learning from input features through traditional convolution, it is generated based on scale information prediction, thus achieving scale-aware adjustment of dynamic sampling positions. The specific implementation process is as follows Figure 4 As shown. The number of output channels of the convolutional layer of the entire module is , among which the former Channels corresponding Group 2D offset ,back Channel output modulation coefficient For non-integer sampling positions , whose eigenvalues ​​are calculated by bilinear interpolation , thus supporting feature extraction with sub-pixel accuracy.

[0095] In some embodiments, performing channel feature processing on the spatial feature map by a channel attention module and outputting the channel feature map includes:

[0096] After the upsampling scale factor is nonlinearly transformed by a multi-layer perceptron and processed by an activation function, two channel attention feature masks are generated;

[0097] After the spatial feature map is transformed by the convolution layer, it is element-wise multiplied with the two channel attention feature masks to obtain two channel weighted feature maps;

[0098] The two channel weighted feature maps and the spatial feature map are added through a residual connection to output a channel feature map.

[0099] Specifically, the height and width upsampling scale factors of the given input features are After nonlinear transformation by multi-layer perceptron (MLP), two channel attention feature masks are generated by Sigmoid activation function. ,

[0100] Its expression is as follows:

[0101] ;

[0102] in, represents a multilayer perceptron consisting of two fully connected layers, is the Sigmoid activation function, The operation represents the equal division of the input into two parts along the channel dimension, and the resulting two channel attention feature masks and .

[0103] In some embodiments, the expression of the channel feature map is:

[0104] ;

[0105] in, represents the channel feature map, Represent the two channel attention feature masks respectively, and represents the convolutional layer, Represents a spatial feature map.

[0106] Specifically, the spatial feature map output by the scale-guided spatial attention module After two different 3×3 convolutional layers for feature transformation, the feature maps of the convolution output are combined with the channel attention feature mask and Perform element-wise multiplication to obtain two scale-aware channel weighted feature maps. Finally, the two modulated channel weighted feature maps are combined with the spatial feature map The final output of the scale-guided channel attention module, namely the channel feature map, is formed by adding the residual connections.

[0107] The core function of the channel attention module is to dynamically adjust the feature weights of each channel according to different upsampling scales using a differentiated processing strategy, thereby enhancing the feature representation capability of the backbone network.

[0108] refer to Figure 6 , an embodiment of the present invention further provides an arbitrary-scale image super-resolution reconstruction system based on a dual-attention mechanism, comprising:

[0109] at least one processor;

[0110] at least one memory for storing at least one program;

[0111] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.

[0112] The contents of the above method embodiments are all applicable to this embodiment. The functions specifically implemented by this embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments, which will not be repeated here.

[0113] An embodiment of the present invention further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above method is implemented.

[0114] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiment, the functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0115] Although the description of the present disclosure has been quite detailed and particularly describes several embodiments, it is not intended to be limited to any of these details or embodiments or any particular embodiment, but should be considered to provide a broad possible interpretation of these claims by reference to the appended claims in view of the prior art, thereby effectively covering the intended scope of the present disclosure. In addition, the above description of the present disclosure is based on the embodiments foreseen by the inventors, which is intended to provide a useful description, and those non-substantial changes to the present disclosure that have not yet been foreseen may still represent equivalent changes to the present disclosure.

Claims

1. A super-resolution reconstruction method for arbitrary-scale images based on a dual-attention mechanism, characterized in that: The method comprises the following steps: Establishing a super-resolution reconstruction model, the super-resolution reconstruction model includes a shallow feature extraction module, a deep feature extraction module, a feature aggregation module and an arbitrary scale upsampling module; extracting shallow features from the low-resolution image by the shallow feature extraction module, and extracting a channel feature map and multiple intermediate feature maps from the shallow features by the deep feature extraction module; aggregating the channel feature map and the multiple intermediate feature maps by the feature aggregation module to generate an output feature map; processing the output feature map and the upsampling scale factor by the arbitrary scale upsampling module to output a high-resolution reconstructed image; Acquire a data set, and train a super-resolution reconstruction model using the data set to obtain a trained super-resolution reconstruction model; wherein the data set includes a plurality of training samples, and the training samples include low-resolution images and corresponding real images; Obtaining a low-resolution image to be reconstructed, reconstructing the low-resolution image to be reconstructed using a trained super-resolution reconstruction model, and outputting a high-resolution reconstructed image; The deep feature extraction module includes a plurality of cascaded RDB-SGDA modules, and the RDB-SGDA module includes a connected RDB module and an SGDA module; The method of extracting a channel feature map and a plurality of intermediate feature maps from shallow features by using a deep feature extraction module includes: The shallow feature map is input as the input feature map to the first cascaded RDB-SGDA module for stacking processing, and the upsampling scale factor and the input feature map are processed by the RDB-SGDA module to obtain the intermediate feature map of the first cascaded RDB-SGDA module after processing. The intermediate feature map is used as the input feature map of the next cascaded RDB-SGDA module, and after a set number of RDB-SGDA module stacking processes, the intermediate feature map obtained by the last stacking process is used as the deep feature map; The SGDA module includes a spatial attention module and a channel attention module, and the channel feature map is obtained by: The intermediate feature map is processed by the spatial attention module and the spatial feature map is output; The channel feature processing of the spatial feature map is performed through the channel attention module, and the channel feature map is output.

2. The method according to claim 1, characterized in that The super-resolution reconstruction model is trained using the data set to obtain a trained super-resolution reconstruction model, comprising: A loss value of a super-resolution reconstruction model is determined based on the high-resolution reconstructed image and the real image corresponding to the low-resolution image, and a parameter set of the super-resolution reconstruction model is iteratively updated based on the loss value to obtain a trained super-resolution reconstruction model.

3. The method according to claim 2, characterized in that The loss value is calculated by the following formula: ; in, represents the parameter set of the super-resolution reconstruction model, M represents the number of training samples, is the low-resolution image in the j-th sample, is the real image corresponding to the low-resolution image in the jth sample, represents the hyperparameters of the super-resolution reconstruction model, represents the optimal hyperparameters of the super-resolution reconstruction model.

4. The method according to claim 1, wherein The intermediate feature map is processed by the spatial attention module to output the spatial feature map, including: The intermediate feature map is input into the spatial attention module to generate a spatial attention mask; the intermediate feature map is input into the scale-aware deformable convolution layer to generate a scale-fused feature map with arbitrary scale information; After element-wise multiplication of the scale fusion feature map and the spatial attention mask, it is adaptively fused with the intermediate feature map to output the spatial feature map.

5. The method according to claim 4, characterized in that The channel feature processing is performed on the spatial feature map by the channel attention module, and the channel feature map is output, including: After the upsampling scale factor is nonlinearly transformed by a multi-layer perceptron and processed by an activation function, two channel attention feature masks are generated; After the spatial feature map is transformed by the convolution layer, it is element-wise multiplied with the two channel attention feature masks to obtain two channel weighted feature maps; The two channel weighted feature maps and the spatial feature map are added through a residual connection to output a channel feature map.

6. A super-resolution reconstruction system for arbitrary-scale images based on a dual-attention mechanism, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 5.

7. A computer-readable storage medium storing a program executable by a processor, characterized in that: The processor-executable program is configured to perform the method according to any one of claims 1 to 5 when executed by the processor.

Citation Information

Patent Citations

  • Image super-resolution reconstruction model construction method and device, equipment and storage medium

    CN114926342A

  • AUC optimization-based sharpness perception minimization image classification training method

    CN119785070A