Efficient visual super-resolution reconstruction method based on multi-dimensional feature enhancement

By combining sliding window attention and adaptive Token sampling technology, the problem of excessive computing and memory consumption in the image super-resolution method is solved, and efficient image reconstruction effect is achieved, especially suitable for large-scale image processing.

CN120339065APending Publication Date: 2025-07-18HENAN UNIV OF SCI & TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510333266.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing image super-resolution methods are over-computing and memory consumption when processing large-scale images, and local attention calculations cannot adequately capture global features.

Method used

Combining the sliding window attention mechanism and adaptive token sampling technology, self-attention calculation is performed by dividing the image into local windows, important tokens are dynamically selected, and Fourier convolution is introduced in the multi-layer perception machine to enhance feature representation.

Benefits of technology

It significantly improves computing efficiency and memory usage efficiency while maintaining high-quality image reconstruction effects, and is suitable for large-scale image processing tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339065A_ABST
    Figure CN120339065A_ABST
Patent Text Reader

Abstract

The invention relates to an efficient visual super-resolution reconstruction method based on multi-dimensional feature enhancement. The method comprises the following steps: firstly, extracting local features of an image through window division and a sliding window attention mechanism, and capturing detail information of the image; local features are integrated to generate global feature representation, adaptive Token sampling is used in the global level, and the most important Token is screened out according to the significance score of the Token so as to reduce the calculation burden and redundant data volume. On the basis of local and global feature extraction, Fourier convolution is further added into a multi-layer perceptron (MLP) structure, frequency features of an image are extracted through frequency domain transformation, and the expression ability of a model on a frequency domain is enhanced. And finally, local, global and frequency features are fused for generating super-resolution image output. Compared with a traditional method, the method has the advantages that the processing efficiency is remarkably improved on the premise that the super-resolution effect is ensured, and the method is suitable for hardware equipment with limited resources and efficient image processing tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing technology, in particular to an efficient visual super-resolution reconstruction method based on multi-dimensional feature enhancement for image super-resolution reconstruction tasks, belonging to the fields of computer vision and deep learning, especially related to image super-resolution technology. Background Art

[0002] With the development of deep learning and neural networks, image super-resolution (SR) has become an important task in the field of computer vision. Image super-resolution aims to recover high-resolution detailed information from low-resolution images, which has important practical value for many application scenarios (such as medical imaging, satellite image analysis, video surveillance, etc.). Existing image super-resolution methods are mostly based on the convolutional neural network (CNN) or vision transformer (Transformer) architecture. Vision Transformers (such as Swin Transformer) have shown significant advantages in image processing tasks. However, due to the global self-attention mechanism of the traditional transformer architecture consuming a large amount of computing resources, there are large memory and computational overheads when processing large-scale images. To address this problem, window attention, as a local attention calculation method, can effectively reduce the consumption of computation and memory. However, its limitation is that it may not be able to fully capture global features. Therefore, how to combine local attention calculation with global feature modeling while further improving computational efficiency has become an important issue in the current research of image super-resolution technology. Summary of the Invention

[0003] The purpose of the present invention is to propose an innovative image super-resolution reconstruction model that combines a sliding window attention mechanism with an adaptive sampling technique, aiming to reduce computational and memory overheads while maintaining good performance;

[0004] To achieve the above object, the present invention provides the following technical solutions:

[0005] S1. Sliding window attention (WindowAttention): Perform self-attention calculation within a local area. By dividing the image into multiple small local windows, the amount of calculation is reduced, and local features can be effectively captured. Through localized calculation, the high computational overhead in the traditional global attention mechanism is avoided;

[0006] S2. Adaptive Token Sampling (ATS): Through an adaptive sampling mechanism, the most important Tokens (feature points) are dynamically selected according to the attention weights, and effective information selection and retention are carried out globally, further reducing the computational amount while ensuring that important features are fully retained;

[0007] S3. Fourier Convolution: Fourier convolution is introduced into the MLP (Multi-Layer Perceptron) structure of the model. By enhancing the feature representation through frequency-domain information, the model's ability to capture details is improved, and feature fusion and generation effects are strengthened.

[0008] Through the combination of these three technologies, the present invention can significantly improve the computational efficiency, reduce the memory consumption, and maintain the high-quality image reconstruction effect in the image super-resolution task;

[0009] Among them, the super-resolution reconstruction model is constructed based on the Swin Transformer network, including:

[0010] S1. First, input the image. The input is a low-resolution or degraded image I low , and the initial feature F0 is extracted through the convolutional layer, F0 = Conv(I LR );

[0011] S2. Use an additional shallow convolutional layer to further extract the shallow feature F SL , F SL = H C3 (F0), where H C3 represents the 3×3 convolutional operation

[0012] S3. Use a deep feature extraction module (based on ASTB) to extract deep features;

[0013] S4. Upsample the features extracted deeply to generate the output feature F up that matches the target resolution;

[0014] S5. Convert the upsampled features into the final high-resolution image I high , I HR = Conv final (F up ).

[0015] Further, in step S3, the deep feature extraction module is stacked by a number of RAS Transform Blocks.

[0016] Among them, the RAS TransformBlock includes several groups of densely connected residual blocks and a convolution connected after the last group of residual blocks. Each group of residual blocks includes a connected convolution and an AS TransformerBlock;

[0017] The RAS TransformBlock module includes Patch Partition, LinearEmbedding, AS TransformerBlock, and Patch Merging connected in sequence;

[0018] Furthermore, each AS TransformBlock specifically includes: a window attention mechanism, a cross-window interaction mechanism, a global adaptive Token sampling module, an MLP module for enhancing features, and a residual connection for fusing features;

[0019] Furthermore, the steps of the window attention mechanism are as follows:

[0020] A1. The window attention mechanism divides the feature map into several non-overlapping small windows, and the size of each window is a fixed value (such as N×N);

[0021] A2. Perform self-attention calculation within each window. First, calculate the query (Q), key (K), and value (V) matrices.

[0022] Q = XW Q , K = XW K , O = AV, V = XW V ; where X is the input feature within the window, and W Q , W K , W V are linear mapping matrices.

[0023] Subsequently, calculate the attention weights based on the dot product of QK T :

[0024] Finally, use the calculated attention weights to perform weighted summation on the V matrix to generate the feature representation within the window: O = AV

[0025] The self-attention based on the window is expressed as:

[0026]

[0027] Capture local feature correlations through the self-attention mechanism in each window:

[0028] F local = WindowAttention(F shallow );

[0029] A3. Re - splice the features of all windows through a shifted window to achieve information interaction between windows and form a local feature map;

[0030] Furthermore, the cross - window interaction mechanism re - defines the boundaries of window division through ShiftedWindow operation, making the window divisions of the front and back layers overlap with each other, so as to achieve cross - window feature interaction;

[0031] Furthermore, the core idea of the global adaptive Token sampling module is to dynamically select the Tokens (i.e., partial features of the image) that need to be calculated according to the saliency of the input image features. By focusing attention on more representative Tokens in an adaptive manner, the computational amount is effectively reduced and the performance of the model is improved.

[0032] Furthermore, the ATS module is mainly implemented through the following steps:

[0033] B1. Calculate the saliency score of each Token in the input image through the self - attention mechanism. The calculation of the score is based on the global information of each Token, including the dependence relationship between this Token and other Tokens and the local feature representation of this Token. By sorting these scores, ATS can identify the most representative Tokens;

[0034] B2. After obtaining the saliency scores of each Token, ATS selects important Tokens for subsequent calculations through an adaptive strategy. For Tokens with low saliency or redundancy, ATS can choose to exclude them, thus reducing the computational amount.

[0035] B3. Adopting a sampling - based method, ATS selects a certain number of high - saliency Tokens from the input Tokens. The global significant features can be expressed as:

[0036] F global = ATS(X)

[0037] The calculation method through the self - attention mechanism in the above B1 is

[0038]

[0039] where S i is the saliency score of the i - th Token, Q i and K j are the query vector and the key vector respectively, and d k is the feature dimension.

[0040] The low - significance or redundant Tokens in B2 above are mainly filled with zeros or ignored through a matrix mask (Mask).

[0041] Further, the local features extracted by the sliding window attention are weighted and summed with the global significant features extracted by the ATS module. The fused features can be expressed as: F fusion = αF local + βF global . Where α and β are learnable fusion weights, ensuring that the fusion process can adaptively adjust the importance of the two types of features.

[0042] The above - fused features contain local details and global semantic information, and have more comprehensive expressive power.

[0043] A further improvement of the technical solution of the present invention is that the global adaptive sampling module and the sliding window attention mechanism are parallel attention calculation modules. Through the calculations of the global adaptive sampling module and the sliding window attention mechanism, both the local features of the picture and the global dependencies of the picture can be obtained. Then, the frequency features are captured through the MLP module that fuses SFB. The finally obtained features have powerful representation capabilities, and the calculation steps are as follows:

[0044] X1 = H ATS (LN(X))

[0045] X2 = H WSA (LN(X))

[0046] X3 = X + X1 + X2

[0047] F = SFBMLP(LN(X3))+X3

[0048] In the above formula, LN is the normalization layer, X1 is the global feature extracted by the adaptive Token sampling module, X2 is the local feature extracted by the sliding window attention WSA, X3 is obtained after the fusion of the global feature, local feature, and initial feature. Subsequently, the fused feature X3 passes through the normalization layer LN and the spatial - frequency fusion SFBMLP and obtains the final network output feature F through a residual connection.

[0049] A further improvement of the technical solution of the present invention is that: the spatial - frequency fusion module SFB is composed of a spatial branch and a frequency branch. After calculating the features of each branch, the features of the two branches are concatenated to obtain the spatial - frequency fusion feature F SFB , and the fast Fourier convolution is used in the calculation process of the frequency branch, enabling this module to mine the high - frequency information of the image, and the spatial branch can learn the global information. Finally, the spatial - frequency fusion module can extract more representative features.

[0050] A further improvement of the technical solution of the present invention lies in: in step 3, after extracting shallow features and deep features, the two features are fused through a long skip connection, and then a high-resolution remote sensing image is calculated through an image reconstruction module: F SFB = H SFB (X). Where F SFB is the enhanced feature, X is the input feature, and H SFB (X) represents the operation of the spatial frequency fusion module SFB, and the main steps are as follows:

[0051] First, the input feature X is first calculated by the spatial branch to obtain the spatial feature F spatial : F spatial = H spatial (X);

[0052] Similarly, the feature obtained by calculating the input feature X by the frequency branch is: F frequency = H frequency (X). Among them, H(X) is the unified operation method.

[0053] Finally, the features calculated by the two branches are spliced and fused to obtain the final spatial frequency fusion feature F SFB : F SFB = H C (|F spatial , F frequency |).

[0054] A further improvement of the technical solution of the present invention lies in: in step S3, after extracting shallow features and deep features, the features respectively extracted from the shallow features and the deep features are fused through a skip connection, and then a high-resolution remote sensing image is calculated through an image reconstruction module.

[0055] F = F SL + F DL

[0056] I HR = H R (F)

[0057] In the formula, F is the fused feature, and I HR represents the high-resolution remote sensing image picture, and H R is the operation of the image reconstruction module.

[0058] Furthermore, the image reconstruction module consists of an upsampling operation and a convolution operation, and the calculation process is as follows:

[0059] First, a 3×3 convolutional layer is used to calculate the feature F, then upsampling is performed through a pixel shuffle operation, and finally a 3×3 convolution is used to refine the feature to obtain the final high-resolution remote sensing image IHR :

[0060] I HR = H C (H PX (H C (F))), where H C represents a 3×3 convolution operation, and H PX represents a pixel shuffle operation.

[0061] Due to the adoption of the above technical solution, the technical effects achieved by the present invention are as follows:

[0062] C1. By combining window attention and adaptive sampling, the amount of calculation is reduced. Especially when processing high-resolution images, the calculation efficiency and memory usage efficiency can be significantly improved.

[0063] C2. The adaptive sampling technology can dynamically select the Tokens to be retained according to the importance of features, thereby reducing unnecessary information loss and ensuring that important details in the image are fully retained.

[0064] C3. By introducing SFB in the MLP module, the ability of the model in image detail, texture, and edge restoration is effectively improved, and the quality of the reconstructed image is enhanced.

[0065] C4. By combining the window attention mechanism and the adaptive sampling technology, the model can reduce the calculation and memory consumption without sacrificing performance, and is applicable to large-scale image processing tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 is a schematic structural diagram of the super-resolution reconstruction model provided by the embodiment of the present invention;

[0067] Figure 2 is a schematic structural diagram of the depth feature extraction module RASTB module provided by the embodiment of the present invention;

[0068] Figure 3 is a schematic structural diagram of the adaptive Token sampling module provided by the embodiment of the present invention;

[0069] Figure 4 is a schematic structural diagram of the SFB MLP module provided by the embodiment of the present invention;

[0070] Figure 5 is a schematic structural diagram of the SFB module provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0071] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0072] The technical solutions of the present application will be further described in detail below in conjunction with specific embodiments.

[0073] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present application and should not be construed as a limitation to the present application. Without conflict, the embodiments of the present application and the technical features in the embodiments can be combined with each other.

[0074] The purpose of the present invention is to improve the performance of image super-resolution technology in the remote sensing image scenario, for reconstructing clearer remote sensing images. While improving the performance, the network structure is further optimized to reduce the computational amount and model parameters.

[0075] As Figure 1 shown, an efficient visual super-resolution reconstruction method based on multi-dimensional feature enhancement includes the following steps:

[0076] Step S1: First, input a low-resolution image I LR , and extract the shallow feature F SL from the input image through the shallow feature extraction module. The specific steps are as follows:

[0077] For the low-resolution image I LR extracted from the remote sensing image, use a 3×3 convolutional layer H C3 to process it and generate the shallow feature F SL . As Figure 1 shown, that is, F SL = H C3 (I LR )

[0078] Step S2: Further process the shallow feature F SL using the deep feature extraction module to obtain the deep feature extraction module F DL . The specific steps are as follows:

[0079] As Figure 1As shown, the deep feature extraction module consists of a sliding window attention mechanism module, an adaptive Token sampling module, and a spatial frequency fusion MLP module. The deep feature extraction process can be described as: F DL = H SFBMLP (H windowattention (F SL ) + H ATS (F SL ))

[0080] where H windowattention and H ATS represent the operation methods of the sliding window attention mechanism module and the adaptive Token sampling respectively, and H SFBMLP is the operation method of the SFBMLP module. By fusing the features extracted by the window attention mechanism module and the adaptive Token sampling module, the features have both local receptive fields and global dependencies; through the SFBMLP module, the fused features further retain high-frequency information; as a whole, the deep features can well retain the texture details of the image.

[0081] As Figure 2 shown, the multi-dimensional feature fusion module RASTB consists of an adaptive Token sampling module, a window attention mechanism module, and a spatial frequency fusion MLP. Among them, WMSA represents window-based multi-head self-attention, and SWMSA represents the attention recalculated after window sliding. The whole process can be expressed as:

[0082] F1 = H WMSA (LN(F SL ))

[0083] F2 = H SWMSA (LN(F SL ))

[0084] F3 = F SL + F1 + F2

[0085] F4 = H ATS (F SL )

[0086] F DL = SFBMLP(LN(F3 + F4)) + F3 + F4

[0087] where LN() is the operation of the normalization layer, F1 is the feature extracted by window-based self-attention, F2 is the feature re-extracted after window sliding, F3 is the local feature extracted by the sliding window attention mechanism, F4 is the significant global feature extracted by the adaptive Token sampling module. The locally-global fused features are input into the SFBMLP, and the final output feature F of the deep feature extraction module is obtained through a skip connection.DL

[0088] Windowed multi-head self-attention MSA divides the input of size H×W×C evenly into non-overlapping windows, each corresponding to an N×n patch. Multi-head self-attention MSA is calculated within each window, so there is no information interaction between windows. In order to overcome the shortcomings of MSA, the shift window method is introduced, and the two methods are used alternately to achieve cross-window connection.

[0089] like Figure 3 As shown in the figure, the ATS module first calculates the significance score of the input Token. The calculation method is: Where A is the attention matrix, and A is calculated as: V represents the attention score. According to the significance score, ATS samples the token through the cumulative distribution function (CDF), and the CDF is calculated as follows: The sampling function is obtained by taking the inverse of the CDF. The sampling function is calculated as: ψ(k) = CDF -1 (k).

[0090] Adaptive sampling technology selects the most representative tokens globally, so that the model not only focuses on important information in the local area, but also performs effective feature selection globally.

[0091] like Figure 4 As shown in Figure 1, based on the traditional multi-layer perceptron, the spatial frequency fusion module SFB is introduced into the MLP. The SFB module consists of a spatial branch and a frequency branch. After calculating each branch feature, the two branch features are concatenated to obtain the spatial frequency fusion feature F. SFB In the process of frequency branch calculation, fast Fourier convolution is used to enable the module to mine the high-frequency information of the image, and the spatial branch can learn global information. Finally, the spatial frequency fusion module can extract more representative features F. SFB : F SFB =H SFB (X). Where F SFB is the enhanced feature, X is the input feature, H SFB (X) represents the operation of the spatial frequency fusion module SFB, and the main steps are as follows:

[0092] First, the input feature X is calculated through two convolutional layers and an activation layer LeakyReLu, and residual connections are used to increase the expressiveness of the model to obtain spatial features. The input feature X is calculated by the spatial branch to obtain the spatial feature F spatial : F spatial =H spatial (X);

[0093] For the frequency branch, the most important operation is the fast Fourier convolution FFC, which transforms the features into the frequency domain to extract global information. First, the two-dimensional fast Fourier transform is used to map the spatial features to the frequency domain to extract global dependencies, and then the inverse fast Fourier transform is performed to map the features back to the spatial domain. The calculation method is: F frequency = H frequency (X).

[0094] Finally, the features calculated by the two branches are concatenated and fused to obtain the final spatial-frequency fusion feature F SFB : F SFB = H C (|F spatial , F frequency |).

[0095] Step 3: Use the image reconstruction module to calculate the reconstructed high-resolution remote sensing image I SL from the shallow features F DL extracted from the input low-resolution remote sensing image and the deep features F HR . The process is as follows:

[0096] First, a 3×3 convolutional layer is used to calculate the features F, then upsampling is performed through the pixel shuffle operation, and finally, a 3×3 convolution is used to refine the features to obtain the final high-resolution remote sensing image I HR : I HR = H C (H PX (H C (F))), where F is the feature after fusing the shallow feature F SL and the deep feature F DL , and H C represents a 3×3 convolutional operation, and H PX represents the pixel shuffle operation.

[0097] Step 4: Evaluate the reconstructed high-resolution image using the peak signal-to-noise ratio.

[0098] In this embodiment, the calculation formula for the peak signal-to-noise ratio is:

[0099] where PSNR is the peak signal-to-noise ratio, n is the number of bits per sampling value, and MSE is the mean square error. For an image with a size of M×N, where is the pixel value of the pixel at the coordinate (i, j) in the reconstructed high-resolution image and the low-resolution image to be reconstructed.

[0100] The super-resolution reconstruction method for remote sensing images provided in this embodiment is applicable to the super-resolution reconstruction of satellite remote sensing images. Currently, during the satellite remote sensing shooting process, due to the excessive attenuation of high-frequency signals caused by the overly long signal transmission distance and the absence of an amplification compensation device in between, the captured images are very blurry, and their quality cannot meet the requirements of ground object target recognition, detailed land detection, etc., and it is even more impossible to achieve good results in application fields such as disaster monitoring, urban economic level assessment, and resource exploration.

[0101] This example provides a super-resolution reconstruction model for remote sensing images based on window attention and adaptive sampling technology, which can maintain excellent image reconstruction quality while reducing computational and memory overhead. By combining local attention calculation and global feature modeling, the model of the present invention shows higher efficiency and better performance in image processing tasks, especially suitable for the super-resolution reconstruction task of large-scale high-resolution images.

[0102] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.

[0103] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An efficient visual super-resolution reconstruction method based on multi-dimensional feature enhancement, characterized in that, Including: Input the low-resolution image to be reconstructed into the trained super-resolution reconstruction model to obtain the reconstructed high-resolution image; Among them, the efficient vision super-resolution reconstruction method based on multi-dimensional feature enhancement combines the image super-resolution reconstruction network of the SwinIR network and the adaptive Token sampling strategy, including Shallow feature extraction module: used to perform shallow feature extraction on the input low-resolution image; Deep feature extraction module: used to perform deep feature extraction on the input low-resolution image; Image reconstruction module: used to upsample the fused shallow features and deep features to obtain the reconstructed high-resolution image; Among them, the shallow feature extraction module uses a convolution to perform shallow feature extraction on the input low-resolution image to be reconstructed; the deep feature extraction module uses several connected RASTBs to perform deep feature extraction on the input low-resolution image to be reconstructed; the image reconstruction module uses the pixel shuffle method combined with convolution to upsample the integrated shallow features and deep features; Among them, the deep feature extraction module includes several groups of densely connected residual blocks and a convolution connected after the last group of residual blocks, and each group of residual blocks includes a connected convolution and an ASTB.

2. The RASTB module according to claim 1, wherein The network structure consists of an input layer, a sliding window attention mechanism, an adaptive Token sampling module (ATS), an MLP layer that fuses the spatial frequency fusion module (SFB), a feature fusion layer, an output layer, several groups of residual connections, and a post-processing layer.

3. The window attention mechanism according to claim 2, wherein Mainly includes the following steps: S1: Divide the input image into multiple local windows; each window contains a certain number of Tokens; S2: Local self-attention calculation: Inside each window, use the self-attention mechanism to calculate the relationship and weight between Tokens. Each Token obtains local context information by interacting with other Tokens within the window, thereby capturing local features; S3: Attention output within the window: The attention calculation result within each window will be used to update the feature representation of each Token within the window; S4: Re-piece together the features of all windows through shifted windows to achieve information interaction between windows and form a local feature map.

4. The adaptive Token sampling module according to claim 2, wherein The module includes the following steps: S1: ATS calculates the saliency score of each Token, and these scores are based on the attention weight of the Token and the norm of the feature. Among them, the calculation method of the attention matrix is as follows: The calculation method of the attention weight is: Ο = AV, where V represents the attention score; The significance score calculation formula is as follows: S2: According to the saliency score, ATS samples the Tokens through the cumulative distribution function (CDF) and selects the Tokens that are most important for the image reconstruction task. In this way, ATS can reduce the computational complexity while retaining the Token information crucial for the task. Among them, the calculation method of the CDF is as follows: With the cumulative distribution function, the sampling function is obtained by taking the reciprocal of the CDF, and the calculation method of the sampling function is: ψ(k) = CDF -1 (k). S3: The adaptive sampling strategy screens out the most representative Tokens globally, enabling the model to not only focus on important information within the local area but also perform effective feature selection globally.

5. The method according to claim 2, wherein Introduce the Spatial Frequency Fusion Module (SFB) into the multi-layer perceptron (MLP) structure. The SFB module consists of a spatial branch and a frequency branch. After calculating the features of each branch, the features of the two branches are concatenated to obtain the spatial frequency fusion feature F SFB During the calculation of the frequency branch, fast Fourier convolution is used, enabling the module to mine the high-frequency information of the image. And the spatial branch can learn global information. Finally, the spatial frequency fusion module can extract more representative features: F SFB = H SFB (X) Among them, F SFB is the enhanced feature, X is the input feature, and H SFB is the method of spatial frequency fusion. The main steps are as follows: First, the input feature X is first calculated by the spatial branch to obtain the spatial feature F spatial : F spatial = H spatial (X); Similarly, the feature obtained by the frequency branch for the input feature X is: F frequency = H frequency (X). Here, H(X) is uniformly the operation method. Finally, the features calculated by the two branches are concatenated and fused to obtain the final spatial frequency fusion feature F SFB .

Citation Information

Cited By

  • Image generation method, electronic device, storage medium and program product

    CN121213708A

  • Image generation methods, electronic devices, storage media and program products

    CN121213708B