Image super-resolution reconstruction method and device based on omnidirectional feature fusion

Through the image super-resolution reconstruction method of omnidirectional feature fusion, local-global-channel feature interaction and multi-level inter-block feature aggregation, the shortcomings of the self-attention mechanism in the Transformer architecture are solved, and high-quality image reconstruction effect is achieved.

CN120543377APending Publication Date: 2025-08-26Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510514217.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The existing image super-resolution reconstruction method based on Transformer architecture has ineffective interaction and homogeneous aggregation scheme that limits the modeling ability of the self-attention mechanism, and the blockage of network information propagation affects the quality of image reconstruction.

Method used

The local-global-channel feature interaction method is adopted based on omnidirectional feature fusion, and the accuracy and quality of image reconstruction are enhanced through multi-level inter-block feature aggregation and inter-block feature fusion, and the heterogeneous feature aggregation architecture is used.

Benefits of technology

It realizes more accurate and high-quality image reconstruction, solves the unclear details and artifact problems in traditional methods, and improves the super-resolution reconstruction effect of image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120543377A_ABST
    Figure CN120543377A_ABST
Patent Text Reader

Abstract

The invention provides an image super-resolution reconstruction method and device based on omnidirectional feature fusion. The method comprises the following steps: acquiring a to-be-reconstructed low-resolution image, and inputting the to-be-reconstructed low-resolution image into a trained image super-resolution network for image reconstruction to obtain a high-resolution image; the image super-resolution network comprises a shallow feature extraction module, a deep feature extraction module and an image reconstruction module; the deep feature extraction module comprises N1 stacked OEATGs (fully effective attention groups) and a GGFF (global group feature aggregation) module; the image reconstruction process comprises the following steps: extracting a shallow feature F0 of a low-resolution image by using a shallow feature extraction module; inputting the F0 into a deep feature extraction module, extracting N1 intermediate features by using N1 stacked OEATGs, and performing feature aggregation on the N1 intermediate features by using a GGFF module to obtain a deep feature FD; and the image reconstruction module generates a high-resolution image according to the F0 and the FD.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer underlying visual processing technology, and in particular to an image super-resolution reconstruction method and device based on omnidirectional feature fusion. Background Art

[0002] Due to factors such as the hardware performance of image acquisition equipment, shooting environment, and storage and transmission costs, the images actually obtained often have low resolution and cannot meet people's requirements for image quality. Therefore, image super-resolution reconstruction technology has emerged, aiming to convert low-resolution images into high-resolution images through algorithms to improve image clarity and visual effects. Image super-resolution is a classic task in the field of vision and is widely used in medical imaging, security monitoring, remote sensing mapping and other fields. High-resolution images can provide richer and more accurate detail information, which is crucial to improving the effects and performance of these applications. However, image super-resolution is an ill-posed problem because each low-resolution image can be mapped to multiple high-resolution images. The core of image super-resolution research is accurate high-resolution image reconstruction.

[0003] Early image super-resolution reconstruction was mainly based on interpolation algorithms such as nearest neighbor interpolation, bilinear interpolation, and bicubic interpolation. These methods are simple and intuitive, but they do not take into account the texture and structural information of the image. As a result, the reconstructed image is still blurry in details and lacks realism, making it difficult to meet the needs of practical applications. Model-based methods establish a prior model of the image and use the statistical characteristics and local structural information of the image to perform super-resolution reconstruction. However, this method is highly dependent on the prior model. When the characteristics of the image do not match the prior model, the reconstruction effect will be significantly affected. Moreover, these methods are usually computationally complex and require a lot of computing resources and time, making it difficult to meet application scenarios with high real-time requirements.

[0004] Image super-resolution reconstruction methods based on deep learning have achieved significant performance improvements by constructing deep neural network models to automatically learn the mapping relationship between low-resolution and high-resolution images. Compared with traditional methods, deep learning models can better capture the complex features and texture information of images, thereby reconstructing images with higher quality and details. Methods based on convolutional neural networks have achieved good performance in image super-resolution tasks due to their inductive bias and strong local feature modeling capabilities, but it is difficult to break through the performance bottleneck when using large amounts of data. Image super-resolution methods based on the Transformer architecture use long-range feature interactions to establish dependencies between image tokens, achieving performance improvements under large-scale training of large data. Existing methods based on the Transformer architecture have two major shortcomings that hinder further performance improvement: 1) Ineffective interactions and homogeneous aggregation schemes limit the modeling capabilities of the self-attention mechanism; 2) Network information propagation blockages affect image reconstruction quality. These two shortcomings make existing image super-resolution reconstruction results prone to problems such as incorrect structure and unclear details. Summary of the Invention

[0005] In order to overcome the shortcomings of the existing image super-resolution reconstruction method based on the Transformer architecture, the present invention proposes an image super-resolution reconstruction method based on omnidirectional feature fusion. The present invention uses a heterogeneous feature aggregation architecture, designs a "local-global-channel" feature interaction mode, and achieves more accurate and high-quality image reconstruction through multi-level inter-block feature aggregation and inter-block feature fusion.

[0006] In a first aspect, the present invention provides an image super-resolution reconstruction method based on omnidirectional feature fusion, comprising:

[0007] Acquire a low-resolution image to be reconstructed;

[0008] The low-resolution image to be reconstructed is input into a trained image super-resolution network for image reconstruction to obtain a reconstructed high-resolution image; wherein the image super-resolution network includes a shallow feature extraction module, a deep feature extraction module and an image reconstruction module; the deep feature extraction module includes N1 stacked fully effective attention groups (OEATG) and a global group feature aggregation (GGFF) module;

[0009] Correspondingly, the process of the image super-resolution network reconstructing the input low-resolution image includes:

[0010] The shallow feature extraction module is used to extract the shallow feature F0 of the low-resolution image; the shallow feature F0 is input into the deep feature extraction module, and N1 intermediate features are extracted using N1 stacked fully effective attention groups OEATG. The GGFF module is used to aggregate the N1 intermediate features to obtain the deep feature F D ; The image reconstruction module is based on the shallow feature F0 and the deep feature F D Generate high-resolution images.

[0011] Furthermore, the image reconstruction module includes a pixel-shuffle layer and a convolution layer connected in sequence.

[0012] Furthermore, the OEATG includes two stacked fully effective self-attention OEATB modules, a 3×3 convolutional layer and a residual connection; correspondingly, the feature processing process of the OEATG includes: the input feature map is processed in sequence by two OEATB modules and a 3×3 convolutional layer to obtain an intermediate feature map, and the input feature map is element-wise added to the intermediate feature map through a residual connection to obtain an output feature map; wherein, the OEATB module is used to process the input feature map using three different types of Transformer blocks.

[0013] Furthermore, the OEATB module includes a local self-attention L-SA block, a global semantic self-attention GS-SA block and an adaptive channel self-attention AC-SA block connected in sequence; wherein, the L-SA block is used to window the input feature map and aggregate the features of the local window, the GS-SA block is used to adaptively aggregate the same semantic tags according to the input feature map, and perform feature interaction on different semantic tags; the AC-SA block is used to select the channel that is most conducive to accurately reconstructing the high-resolution image.

[0014] Furthermore, the feature processing process of the GS-SA block includes: a meta-token refinement stage and a semantic self-attention stage;

[0015] For the lth GS-SA block, its meta-token refinement stage specifically includes: Let T l Represents the given meta-token input to the lth GS-SA block, using T l Generate query matrix Q, using input feature map X in Generate the key matrix K and the value matrix V; then perform cross-attention calculation on the matrices Q, K, and V to obtain the refined meta-token T l+1 , and record the corresponding attention map A = Softtmax(QK T / τ); where τ represents the learnable parameter, A∈R M×HW ;

[0016] For the lth GS-SA block, its semantic self-attention stage specifically includes: using the attention map A to transform the input feature map X in All of them are aggregated and divided into M feature groups; multi-head attention calculation is performed on each feature group to realize feature interaction and generate new feature groups; all new feature groups are inversely restored to obtain the output feature map.

[0017] Furthermore, the feature processing process of the AC-SA block includes: applying 1×1 convolution and 3×3 depth convolution to the normalized features output by the GS-SA block to generate a query matrix Q, a key matrix K and a value matrix V; applying the reshape function to adjust the size of Q; and using the ReLU operation to calculate the self-attention A=ReLU(QK based on the key matrix and the resized query matrix Q). T / τ); multiply the self-attention A by the value matrix V to obtain the output feature map.

[0018] Furthermore, the OEATB module also includes an adaptive feature fusion AFF module; correspondingly, the fusion process of the AFF module is:

[0019]

[0020] Among them F l-1 and F l Represent the input and output of the lth Transformer block, β l For the lth Transforemr block The learnable parameters within .

[0021] Furthermore, the GGFF module includes a concatenation layer, a 1×1 convolution, and a 3×3 convolution connected in sequence.

[0022] In a second aspect, the present invention provides an image super-resolution reconstruction device based on omnidirectional feature fusion, comprising:

[0023] an acquisition unit, configured to acquire a low-resolution image to be reconstructed;

[0024] A reconstruction unit, configured to input the low-resolution image to be reconstructed into a trained image super-resolution network for image reconstruction to obtain a reconstructed high-resolution image; wherein the image super-resolution network includes a shallow feature extraction module, a deep feature extraction module, and an image reconstruction module; the deep feature extraction module includes N1 stacked fully effective attention groups (OEATGs), a global group feature aggregation (GGFF) module, and a 3×3 convolutional layer;

[0025] Correspondingly, the process of the image super-resolution network reconstructing the input low-resolution image includes:

[0026] The shallow feature extraction module is used to extract the shallow feature F0 of the low-resolution image; the shallow feature F0 is input into the deep feature extraction module, and N1 intermediate features are extracted using N1 stacked fully effective attention groups OEATG. The N1 intermediate features are aggregated using the GGFF module, and the aggregated features are processed by a 3×3 convolutional layer to obtain the deep feature F D ; Add element by element with the shallow feature F0 to obtain the deep feature F D ; The image reconstruction module is based on the shallow feature F0 and the deep feature F D Generate high-resolution images.

[0027] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect when executing the program.

[0028] The beneficial effects of the present invention are:

[0029] The present invention proposes an image super-resolution method based on omnidirectional feature fusion, which overcomes the problems existing in the traditional Transformer architecture-based method driven by the heterogeneous feature aggregation architecture and multi-level feature fusion strategy. By designing a local-global-channel cascade interaction scheme, while promoting information flow, the heterogeneous feature aggregation (i.e., multi-scale and multi-dimensional) is enhanced. At the same time, a global semantic self-attention based on non-local self-similarity is designed in the spatial dimension, and an adaptive sparse global semantic self-attention is designed in the channel dimension. The above two realize the extraction and interaction of effective features, thereby improving the effect of image super-resolution reconstruction. The solution of the present invention can not only perform effective feature modeling on a single image, but also realize the effective fusion of multi-level features, thereby achieving more accurate and high-quality image reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 A schematic diagram of the structure of an image super-resolution network provided by an embodiment of the present invention.

[0031] Figure 2 A schematic diagram of an omnidirectional and efficient feature interaction solution provided by an embodiment of the present invention.

[0032] Figure 3 A schematic diagram of a global semantic self-attention GS-SA block provided in an embodiment of the present invention;

[0033] Figure 4 A schematic diagram of an adaptive channel self-attention (AC-SA) method provided by an embodiment of the present invention;

[0034] Figure 5A schematic structural diagram of an image super-resolution reconstruction device based on omnidirectional feature fusion provided by an embodiment of the present invention;

[0035] Figure 6 This is a structural block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0036] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0037] An embodiment of the present invention provides an image super-resolution reconstruction method based on omnidirectional feature fusion, comprising the following steps:

[0038] S101: Acquire a low-resolution image to be reconstructed;

[0039] S102: Inputting the low-resolution image to be reconstructed into a trained image super-resolution network for image reconstruction to obtain a reconstructed high-resolution image.

[0040] Specifically, if Figure 1 As shown, the image super-resolution network provided by the embodiment of the present invention includes a shallow feature extraction module, a deep feature extraction module and an image reconstruction module; the deep feature extraction module includes N1 stacked fully effective attention groups (OmniEfficient AggregationTransformerGroups, OEATG) and a global group feature aggregation (Global Group Feature Fusion, GGFF) module.

[0041] The training process of the image super-resolution network is as follows: initialize the image super-resolution network parameters, train it on a public low-resolution-high-resolution paired dataset, and obtain a trained image super-resolution network. Correspondingly, the image super-resolution network reconstructs the input low-resolution image through the following steps:

[0042] First, the shallow feature extraction module is used to extract the shallow feature F0 of the low-resolution image; then, the shallow feature F0 is input into the deep feature extraction module, and N1 intermediate features are extracted using N1 stacked fully effective attention groups OEATG. The GGFF module is used to aggregate the N1 intermediate features to obtain the deep feature F D, the process can be expressed by the following formula; the image reconstruction module is based on the shallow feature F0 and the deep feature F D Generate high-resolution images.

[0043]

[0044] in represents the i-th OEATG, represents N1 intermediate features, is the concatenation of the intermediate features generated by the 1st, 2nd, ... N1th OEATG, H GGFF (·) indicates the GGFF module, F D ∈R H×W×C .

[0045] This method effectively addresses the shortcomings of traditional Transformer-based image super-resolution reconstruction methods, leveraging the image modeling capabilities of the self-attention mechanism while also enhancing the ability to utilize network module features. Through effective multi-scale and multi-dimensional feature extraction and aggregation, more accurate and high-quality image reconstruction is achieved. Images processed by this method reveal detailed structures not recovered by other methods, while the generated results are free of artifacts, effectively achieving more accurate and high-quality image reconstruction.

[0046] In one embodiment, the shallow feature extraction module H SL (·) is implemented using a 3×3 convolutional layer. Specifically, given a low-resolution (LR) image I LR ∈H×W×C in , the extracted shallow features F0∈R H×W×C As shown in the following formula:

[0047] F0=H SL (I LR )

[0048] Where H, W represent the height and width of the low-resolution image, C in and C represent the number of input feature and shallow feature channels.

[0049] In one embodiment, the image reconstruction module uses a reconstruction module with a global residual connection to reconstruct a high-resolution result I SR :

[0050] I SR =H Re (F0+F D )

[0051] In this embodiment, the reconstruction module consists of a pixel-shuffle and a convolutional layer.

[0052] In one embodiment, each OEATG includes two OmniEfficientAggregation Transformer Block (OEATB) modules, a 3×3 convolutional layer, and a residual connection. Accordingly, the feature processing process of each OEATG includes: the input feature map is processed by two OEATB modules and a 3×3 convolutional layer in sequence to obtain an intermediate feature map, and the input feature map is element-wise added to the intermediate feature map via a residual connection to obtain an output feature map; wherein the OEATB module is used to process the input feature map using three different types of Transformer blocks.

[0053] As an implementable method, Figure 1 As shown in Figure 1, each OEATB module contains three types of Transformer blocks: Local Self-Attention (L-SA) block, Global Semantic Self-Attention (GS-SA) block, and Adaptive Channel Self-Attention (AC-SA) block. For the i-th OEATG, the feature processing process of the j-th OEATB module can be expressed as:

[0054]

[0055] Where, and denote the input and output features of the jth OEATB module in the i-th OEATG, like Figure 3 Shown is a schematic diagram of the omnidirectional and efficient feature interaction scheme.

[0056] Furthermore, in one embodiment, the L-SA block realizes local feature interaction by window partitioning. Specifically, first, the input feature F l ∈R HW×C Divide the window and set F l Divide into non-overlapping rectangular panes M represents the size of the pane. Then, a standard self-attention operation is performed on each window feature to achieve feature interaction:

[0057] Q=F l′ P Q ,K=F l′ P K ,V=F l′ P V ,

[0058]

[0059] Furthermore, in one embodiment, the GS-SA block is used to adaptively aggregate the same semantic tags according to the input features and perform effective feature interaction on different semantic tags. Its network structure is as follows: Figure 3 Specifically, we first propose a meta-token refinement strategy that can effectively extract representative tokens from spatial information. Then, we perform semantic self-attention to provide beneficial information interaction for each token.

[0060] (1) Meta-token refinement. In the refinement phase, a set of learnable meta-tokens is generated by random initialization. Using meta-token T l ∈R M×C To generate Query, M represents the size of the meta-token; using the input feature X in ∈R HW×C To generate the Key and Value. Subsequently, a cross attention (CA) calculation is performed to facilitate the information exchange between the image token and the meta-token. The above refinement process is expressed as follows:

[0061] Q=T l W Q ,K=X in W K ,V=X in W V ,

[0062] A=Softtmax(QK T / τ),

[0063] T l+1 =CA(T l , X in )=A·V

[0064] Where W Q , W K , W V ∈R C×C is the linear projection matrix, T l Represents the meta-token of the l-th GS-SA block input, T l+1 represents the meta-token output after the l-th GS-SA block refinement; M represents the size of the meta-token, τ is a learnable parameter, A∈R M×HW .

[0065] (2) Semantic self-attention. In the adaptive meta-token refinement stage, the attention map A represents the importance of each meta-token to the input feature map, and its transpose represents the semantic similarity between the input feature and the M meta-features. Therefore, for the input feature x j, divide it into the semantic category η corresponding to the maximum semantic similarity of the row i :

[0066]

[0067] The above formula means: first transpose A to HW*M, then for each row, find the maximum value among the M semantic similarity values; i = 1, 2, 3, ..., M; x j represents the jth row of A, j = 1, 2, 3, ..., HW;

[0068] Then, the input feature map X is transformed into in All pixels in the image are aggregated, i.e., pixels with the same semantic meaning are grouped together into M feature groups. Multi-head attention (MSA) is then performed on each feature group to achieve feature interaction. However, the number of tokens per semantic category is unbalanced. If too many pixels with a particular semantic meaning are aggregated, windowed attention can be used to further perform feature aggregation, improving computational efficiency. Finally, inverse restoration is performed to obtain the features at the correct location.

[0069] The above semantic self-attention process is expressed as follows:

[0070] {ε i}=partition(X in , η)

[0071]

[0072] Among them, ε i represents the i-th feature group, i = 1, 2, 3, ..., M; X out represents the output feature map of the lth GS-SA block.

[0073] Furthermore, in one embodiment, an adaptive channel self-attention (AC-SA) block is used to select the channel that is most conducive to accurate reconstruction, and its network structure is as follows: Figure 4 As shown. Given the normalized feature F generated by the GS-SA block i,j ∈R HW×C , first using 1×1 convolution and then 3×3 depthwise convolution to generate the query Q∈R HW×C , K∈R HW×C , V∈R HW×C . And specifically a reshape function is applied to adjust the size of Q to Q∈R C×HW In order to maintain the most useful attention for feature aggregation, ReLU operation is used instead of softmax to calculate self-attention:

[0074] Q=F i,j WQ ,K=F i,j W K ,V=F i,j W V ,

[0075] A=ReLU(Q T K / τ)

[0076] F i,j+1 =AC-SA(F i,j )=V·A

[0077] where τ is a learnable parameter.

[0078] On the basis of the above embodiments, in order to promote information flow and utilize the complementary advantages of hierarchical features in all stages, this embodiment also introduces an all-block feature fusion (OBFF) in the image super-resolution network, including intra-block feature fusion and inter-block feature aggregation.

[0079] (1) Intra-block feature fusion

[0080] This embodiment establishes a layer-adaptive feature fusion (LAFF) outside each Transformer block. The output features of the Transformer module are fused by adaptively integrating the input features through a trainable adapter β. The formula is

[0081]

[0082] Among them F l-1 and F l Represent the input and output of the lth Transformer block, β l For the lth Transforemr block The learnable parameters within .

[0083] Overall, LAFF solves the fusion of misaligned features across different Transformer blocks and promotes enhanced information flow toward deeper layers of the network, leading to better performance.

[0084] (2) Inter-block feature aggregation

[0085] In the deep feature extraction stage, the GGFF module mines hierarchical features in a global manner:

[0086]

[0087] in It is the concatenation of the intermediate characteristics generated by the 1st, 2nd, ... N1th OEATG. GFA Contains 1×1 and 3×3 convolutions.

[0088] The CT image super-resolution reconstruction method based on omnidirectional feature fusion provided by the present invention includes a global semantic self-attention mechanism based on content self-similarity for aggregating effective features in the spatial dimension; an adaptive sparse channel self-attention mechanism for aggregating effective features in the channel dimension; a "local-global-channel" heterogeneous aggregation method to achieve multi-scale and multi-dimensional feature interaction and promote network feature propagation; and a full-module feature fusion mechanism to enhance intra-module feature propagation and inter-module feature aggregation. The present invention can effectively solve the shortcomings of traditional image super-resolution reconstruction methods based on the Transformer architecture, not only tapping into the image modeling capabilities of the self-attention mechanism, but also enhancing the utilization capabilities of network module features. Detailed structures that have not been restored by other methods are found in the images processed by the method of the present invention, and no artifacts appear in the generated results, which better achieves more accurate and high-quality image reconstruction.

[0089] Based on the same inventive concept, Figure 5 As shown, an embodiment of the present invention provides an image super-resolution reconstruction device based on omnidirectional feature fusion, including an acquisition unit and a reconstruction unit.

[0090] The acquisition unit is used to acquire a low-resolution image to be reconstructed; the reconstruction unit is used to input the low-resolution image to be reconstructed into a trained image super-resolution network for image reconstruction to obtain a reconstructed high-resolution image; wherein the image super-resolution network includes a shallow feature extraction module, a deep feature extraction module and an image reconstruction module; the deep feature extraction module includes N1 stacked fully effective attention groups (OEATG) and a global group feature aggregation (GGFF) module;

[0091] Correspondingly, the image super-resolution network reconstructs the input low-resolution image in the following process: extracting the shallow feature F0 of the low-resolution image using the shallow feature extraction module; inputting the shallow feature F0 into the deep feature extraction module, extracting N1 intermediate features using N1 stacked fully effective attention groups OEATG, and performing feature aggregation on the N1 intermediate features using the GGFF module to obtain the deep feature F D ; The image reconstruction module is based on the shallow feature F0 and the deep feature F D Generate high-resolution images.

[0092] This method effectively addresses the shortcomings of traditional Transformer-based image super-resolution reconstruction methods. It not only exploits the image modeling capabilities of the self-attention mechanism but also enhances the ability to utilize network module features. The processed image reveals detailed structures that other methods fail to recover, while the generated results are free of artifacts, achieving more accurate and high-quality image reconstruction.

[0093] Figure 6 An example of a physical structure diagram of an electronic device is shown below. Figure 6 As shown, the electronic device may include: a processor (processor) 601, a communication interface (Communications Interface) 602, a memory (memory) 603 and a communication bus 604, wherein the processor 601, the communication interface 602, and the memory 603 communicate with each other through the communication bus 604. The processor 601 can call the logic instructions in the memory 603 to execute the image super-resolution reconstruction method based on omnidirectional feature fusion, which includes: obtaining a low-resolution image to be reconstructed; inputting the low-resolution image to be reconstructed into a trained image super-resolution network for image reconstruction to obtain a reconstructed high-resolution image; wherein the image super-resolution network includes a shallow feature extraction module, a deep feature extraction module and an image reconstruction module; the deep feature extraction module includes N1 stacked fully effective attention groups OEATG and a global group feature aggregation GGFF module; correspondingly, the image super-resolution network performs image reconstruction on the input low-resolution image in a process including: extracting shallow features F0 of the low-resolution image using the shallow feature extraction module; inputting the shallow features F0 into the deep feature extraction module, extracting N1 intermediate features using N1 stacked fully effective attention groups OEATG, and performing feature aggregation on the N1 intermediate features using the GGFF module to obtain deep features F D ; The image reconstruction module is based on the shallow feature F0 and the deep feature F D Generate high-resolution images.

[0094] In addition, when the logic instructions in the above-mentioned memory 603 are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0095] An embodiment of the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the image super-resolution reconstruction method based on omnidirectional feature fusion provided by the above-mentioned method embodiments.

[0096] An embodiment of the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the image super-resolution reconstruction method based on omnidirectional feature fusion provided by the above-mentioned method embodiments is implemented.

[0097] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0098] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. An image super-resolution reconstruction method based on omnidirectional feature fusion, characterized in that: include: Acquire a low-resolution image to be reconstructed; The low-resolution image to be reconstructed is input into a trained image super-resolution network for image reconstruction to obtain a reconstructed high-resolution image; wherein the image super-resolution network includes a shallow feature extraction module, a deep feature extraction module and an image reconstruction module; the deep feature extraction module includes N1 stacked fully effective attention groups (OEATG) and a global group feature aggregation (GGFF) module; Correspondingly, the process of the image super-resolution network reconstructing the input low-resolution image includes: The shallow feature extraction module is used to extract the shallow feature F0 of the low-resolution image; the shallow feature F0 is input into the deep feature extraction module, and N1 intermediate features are extracted using N1 stacked fully effective attention groups OEATG. The GGFF module is used to aggregate the N1 intermediate features to obtain the deep feature F D ; The image reconstruction module is based on the shallow feature F0 and the deep feature F D Generate high-resolution images.

2. The image super-resolution reconstruction method based on omnidirectional feature fusion according to claim 1, characterized in that: The image reconstruction module includes a pixel-shuffle layer and a convolution layer connected in sequence.

3. The image super-resolution reconstruction method based on omnidirectional feature fusion according to claim 1, characterized in that: The OEATG includes two stacked fully effective self-attention OEATB modules, a 3×3 convolutional layer and a residual connection; correspondingly, the feature processing process of the OEATG includes: the input feature map is processed in sequence by two OEATB modules and a 3×3 convolutional layer to obtain an intermediate feature map, and the input feature map is element-wise added to the intermediate feature map through a residual connection to obtain an output feature map; wherein, the OEATB module is used to process the input feature map using three different types of Transformer blocks.

4. The image super-resolution reconstruction method based on omnidirectional feature fusion according to claim 3, characterized in that: The OEATB module includes a local self-attention L-SA block, a global semantic self-attention GS-SA block and an adaptive channel self-attention AC-SA block connected in sequence; wherein, the L-SA block is used to window the input feature map and aggregate the features of the local window; the GS-SA block is used to adaptively aggregate the same semantic tags according to the input feature map and perform feature interaction on different semantic tags; the AC-SA block is used to select the channel that is most conducive to accurately reconstructing the high-resolution image.

5. The image super-resolution reconstruction method based on omnidirectional feature fusion according to claim 4, characterized in that: The feature processing process of the GS-SA block includes: meta-token refinement stage and semantic self-attention stage; For the lth GS-SA block, its meta-token refinement stage specifically includes: Let T l Represents the given meta-token input to the lth GS-SA block, using T l Generate query matrix Q, using input feature map X in Generate the key matrix K and the value matrix V; then perform cross-attention calculation on the matrices Q, K, and V to obtain the refined meta-token T l+1 , and record the corresponding attention map A = Softtmax(QK T / τ); where τ represents the learnable parameter, A∈R M×HW ; For the lth GS-SA block, its semantic self-attention stage specifically includes: using the attention map A to transform the input feature map X in All of them are aggregated and divided into M feature groups; multi-head attention calculation is performed on each feature group to realize feature interaction and generate new feature groups; all new feature groups are inversely restored to obtain the output feature map.

6. The image super-resolution reconstruction method based on omnidirectional feature fusion according to claim 4, characterized in that: The feature processing process of the AC-SA block includes: applying 1×1 convolution and 3×3 depthwise convolution to the normalized features output by the GS-SA block to generate a query matrix Q, a key matrix K and a value matrix V; applying the reshape function to adjust the size of Q; and using the ReLU operation to calculate the self-attention A=ReLU(QK based on the key matrix and the resized query matrix Q). T / τ); multiply the self-attention A by the value matrix V to obtain the output feature map.

7. The image super-resolution reconstruction method based on omnidirectional feature fusion according to claim 4, characterized in that: The OEATB module also includes an adaptive feature fusion AFF module; correspondingly, the fusion process of the AFF module is: Among them F l-1 and F l Represent the input and output of the lth Transformer block, β l For the lth Transforemr block The learnable parameters within .

8. The image super-resolution reconstruction method based on omnidirectional feature fusion according to claim 1, characterized in that: The GGFF module includes a concatenation layer, a 1×1 convolution, and a 3×3 convolution connected in sequence.

9. An image super-resolution reconstruction device based on omnidirectional feature fusion, characterized in that: include: an acquisition unit, configured to acquire a low-resolution image to be reconstructed; A reconstruction unit, configured to input the low-resolution image to be reconstructed into a trained image super-resolution network for image reconstruction to obtain a reconstructed high-resolution image; wherein the image super-resolution network includes a shallow feature extraction module, a deep feature extraction module, and an image reconstruction module; the deep feature extraction module includes N1 stacked fully effective attention groups (OEATGs) and a global group feature aggregation (GGFF) module; Correspondingly, the process of the image super-resolution network reconstructing the input low-resolution image includes: The shallow feature extraction module is used to extract the shallow feature F0 of the low-resolution image; the shallow feature F0 is input into the deep feature extraction module, and N1 intermediate features are extracted using N1 stacked fully effective attention groups OEATG. The GGFF module is used to aggregate the N1 intermediate features to obtain the deep feature F D ; The image reconstruction module is based on the shallow feature F0 and the deep feature F D Generate high-resolution images.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.