Image super-resolution reconstruction method based on double-domain feature fusion and implicit representation
The frequency domain and spatial domain features of remote sensing images are extracted through Haar discrete wavelet transform and Transformer pyramid structure, and the problems of texture blur and structural distortion in remote sensing images are solved through dual-domain cross-attention fusion and implicit representation network, achieving high-quality image reconstruction.
Patent Information
- Application Number
- CN202510710239.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-05
AI Technical Summary
Existing implicit neural representation methods have difficulty in effectively coordinating local and global features in super-resolution reconstruction of remote sensing images, resulting in texture blurring and structural distortion, and lack of generalization capabilities for arbitrary scaling.
Haar discrete wavelet transform and Transformer pyramid structure are used to extract frequency domain and spatial domain features. Through dual-domain cross-attention fusion and implicit representation network, the coordinated expression of frequency domain and spatial domain information is achieved to reconstruct high-quality high-resolution images.
It improves the reconstruction quality of remote sensing images, especially in recovering high-frequency details and maintaining structural consistency, and is suitable for multi-scale super-resolution scenarios.
Smart Images

Figure CN120599489A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the fields of image processing and computers, and in particular to an image super-resolution reconstruction method based on dual-domain feature fusion and implicit representation. Background Art
[0002] In the field of remote sensing, images acquired by satellites or aerial platforms are a crucial data source for geospatial information systems. With the advancement of remote sensing imaging technology, high-resolution remote sensing images are widely used in a variety of practical scenarios, including land cover classification, disaster monitoring and response, and urban infrastructure planning. However, due to the physical limitations of optical imaging systems and the high cost of high-performance hardware, remote sensing sensors struggle to directly acquire image data of sufficiently high spatial resolution, limiting the ability to accurately represent remote sensing images. Therefore, a super-resolution reconstruction algorithm is needed that can restore detailed texture while preserving the integrity of the image structure.
[0003] Most current single image super-resolution (SISR) methods use explicit upsampling structures based on deep neural networks. These trained models are typically only suitable for fixed magnifications and lack generalization capabilities for arbitrary scaling requirements, leading to high computational resource consumption and poor deployment flexibility. To improve the model's adaptability to multiple scales, continuous-scale super-resolution methods that support arbitrary scaling factors have emerged in recent years. Meta-SR first proposed a variable-scale image reconstruction framework, breaking the scale limitations of traditional methods. Subsequent work, such as OverNet, has made further progress in reducing model complexity and improving generalizability. Implicit neural representation (INR), a signal representation method based on coordinate modeling, has recently demonstrated significant potential in image super-resolution. Representative methods, such as LIIF (Local Implicit Image Function), achieve high-quality reconstruction of images of arbitrary scales by constructing a continuous mapping from coordinates to pixel values. However, INR methods have limitations in recovering high-frequency image details (such as texture and edges), manifesting as blurry local regions and unclear edges.
[0004] Frequency domain modeling provides another effective way to represent images. The low-frequency components of an image mainly contain structural information, while the high-frequency components are rich in texture and detail information. Although frequency domain modeling has achieved good results in multiple image reconstruction scenarios, existing methods mostly treat the frequency domain and spatial domain as two independent stages, lacking in-depth information interaction and collaborative modeling, resulting in difficulty in coordinating and unifying the overall structure and detail expression of the image. In remote sensing images, due to the complex structure of the ground objects and frequent texture repetition, if the relationship between local and global features cannot be effectively modeled, it is easy to cause problems such as blurred edges and texture distortion in the reconstructed image. Therefore, it is urgent to design a super-resolution reconstruction method that can achieve efficient information fusion and collaborative expression between the frequency domain and the spatial domain to improve the overall reconstruction quality of remote sensing images. Summary of the Invention
[0005] The purpose of the present invention is to provide an image super-resolution reconstruction method based on dual-domain feature fusion and implicit representation in order to solve the problems that traditional super-resolution methods based on implicit neural representation are mostly limited to spatial domain modeling, which makes it difficult to fully capture the multi-scale fine-grained information in the image, and there are problems of texture blur and structural distortion when facing remote sensing images.
[0006] The above-mentioned purpose of this application is achieved through the following technical solutions: S1: Acquire low-resolution remote sensing images and perform preprocessing; S2: Feature extraction of pre-processed remote sensing images is performed through Haar discrete wavelet transform and Transformer-based pyramid structure to obtain frequency domain features and spatial domain features; S3: Perform dual-domain cross-attention fusion on frequency domain features and spatial domain features to obtain local detail features and global structure features; S4: Through the implicit representation network, the fused local detail features and global structural features are mapped to arbitrary spatial coordinates to obtain the final three-channel high-resolution image, realizing high-quality reconstruction of remote sensing images of any scale.
[0007] Optionally, step S2 includes: S21: Through encoder Preprocessed remote sensing images Extracting shallow features ;
[0008] S22: Using Haar discrete wavelet transform to transform shallow features Processing is performed to extract the frequency domain features of the pre-processed remote sensing image; the frequency domain features include: high frequency information and low frequency information; S23: Transformer-based pyramid structure for shallow features Processing is performed to extract the spatial domain features of the pre-processed remote sensing image; the spatial domain features include: local features and global features.
[0009] Optionally, step S22 includes: Haar discrete wavelet transform is used to transform shallow features Decompose and obtain four sub-band coefficients, including: low-frequency component and three high-frequency detail components 、 、 ; High-frequency detail components 、 、 Represent the vertical, horizontal and diagonal edge information of the image respectively; Use the enhanced residual to enhance the extraction of the low-frequency component and the low-frequency component respectively, as follows: High frequency sub-band After splicing, input the convolution layer to form the initial high-frequency features; low-frequency component Generate features through convolution , i.e. low-frequency information; The initial high-frequency features are fused and the frequency domain enhancement results are output. , that is, high-frequency information.
[0010] Optionally, step S23 includes: Shallow features As input, a four-level pyramid structure is constructed. Each level of feature map is modeled by downsampling and Transformer submodules to achieve different scale context modeling. The feature maps at each level are defined as:
[0011]
[0012] Where H represents the height of the feature map, and W represents the width of the feature map; Indicates the Level feature map; Feature Collection Each level captures local and global dependencies through the Transformer submodule, as follows
[0013] in represents local detail representation, Represents global semantic features; Represents all outputs after P passes through Transformer; Indicates the last output; Represents the first H×W outputs.
[0014] Optionally, step S3 includes: Construct a dual-domain feature fusion module, including: a cross-domain feature alignment unit, a multi-scale orthogonal feature modeling unit, a query-key-value based cross-domain attention fusion unit, and a channel attention enhancement unit; Through the cross-domain feature alignment unit, the feature map of low-frequency information With local features Carry out scale unification and format adjustment; The format-adjusted Query-Key-Value based cross-domain attention fusion unit is fused through the multi-scale orthogonal feature modeling unit and the Query-Key-Value based cross-domain attention fusion unit; Low-frequency information is enhanced by the channel attention unit and global features to integrate; The fusion process includes the following operations:
[0015] in, is low-frequency information; For high-frequency information; It is a local feature; is a global feature; represents the global structural features after fusion, Represents local detail features, Represents the cross-attention operation to achieve complementary fusion of frequency domain and spatial domain information.
[0016] Optionally, the implicit representation network is composed of a set of coordinate-based implicit decoders; The implicit decoder of each coordinate is constructed by a multi-layer perceptron based on a sinusoidal activation function to characterize the texture changes and structural information in the image; the implicit decoder is based on two-dimensional high-resolution coordinates. It is input, combined with the local potential representation in the corresponding feature domain, and outputs the predicted pixel value at that position.
[0017] Optionally, step S4 includes: The fused global structural features are decoded by implicit decoder Local detail features Decode and predict the pixel value of any spatial position in the high-resolution image to obtain the prediction result The prediction results are integrated through the convolution layer to obtain the final three-channel high-resolution image, realizing the mapping from continuous coordinates to pixel values, as follows:
[0018] in Represents a nonlinear function that maps coordinates and fusion features to pixel values, performs pixel-by-pixel implicit function learning, and maps high-resolution coordinates to pixel values; Represents the final three-channel high-resolution image.
[0019] An electronic device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device performs an image super-resolution reconstruction method based on dual-domain feature fusion and implicit representation.
[0020] A computer-readable storage medium stores instructions. When the instructions are executed, an image super-resolution reconstruction method based on dual-domain feature fusion and implicit representation is performed.
[0021] The beneficial effects of the technical solution provided by this application are: The technical solution of the present invention relies on joint modeling in the frequency and spatial domains to enhance the continuous representation capabilities of images. Through the collaborative expression of frequency and spatial domain information, the present invention guides implicit neural representation to achieve higher-precision image reconstruction, improving the quality of recovery of high-frequency details and texture information in complex structures. It outperforms existing methods in mainstream evaluation metrics, particularly in recovering high-frequency details and maintaining structural consistency, and is suitable for refined analysis of remote sensing images and multi-scale super-resolution scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The present application will be further described below with reference to the accompanying drawings and embodiments, in which: Figure 1 is a flow chart in an embodiment of the present application; Figure 2 Schematic diagram of discrete wavelet transform based on Haar wavelet in an embodiment of the present application; Figure 3 is a diagram of the dual-domain implicit super-resolution network structure in an embodiment of the present application; Figure 4 It is a schematic diagram of the structure of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION
[0023] In order to have a clearer understanding of the technical features, purposes and effects of this application, the specific implementation methods of this application are now described in detail with reference to the accompanying drawings.
[0024] The embodiments of the present application provide an image super-resolution reconstruction method based on dual-domain feature fusion and implicit representation.
[0025] Please refer to Figure 1 , Figure 1 : is a flowchart of an image super-resolution reconstruction method based on dual-domain feature fusion and implicit representation in an embodiment of the present application, including: S1: Acquire low-resolution remote sensing images and perform preprocessing; S2: Feature extraction of pre-processed remote sensing images is performed through Haar discrete wavelet transform and Transformer-based pyramid structure to obtain frequency domain features and spatial domain features; S3: Perform dual-domain cross-attention fusion on frequency domain features and spatial domain features to obtain local detail features and global structure features; As an example, dual-domain cross-attention fusion is used to dynamically guide and enhance spatial domain features using frequency domain features. The cross-domain attention mechanism is used to enhance the fusion effect of multi-scale features and achieve collaborative modeling of texture and structural information. Frequency and spatial domain features are fused using dual-domain cross-attention. By performing cross-attention calculations on high-frequency components and local spatial features, a local detail representation of the image is obtained. At the same time, cross-attention calculations are performed on low-frequency components and global spatial features to obtain a global structural representation of the image, achieving the collaborative expression of frequency domain information and spatial structural information.
[0026] S4: Through the implicit representation network, the fused local detail features and global structural features are mapped to arbitrary spatial coordinates to obtain the final three-channel high-resolution image, realizing high-quality reconstruction of remote sensing images of any scale.
[0027] As an embodiment, a dual-path implicit representation network is constructed, which jointly inputs local detail representation, global structure representation and arbitrary high-resolution spatial coordinates, and establishes a nonlinear mapping relationship from coordinates to pixel values to achieve high-quality super-resolution reconstruction of remote sensing images at any scale.
[0028] Step S2 includes: S21: Through encoder Preprocessed remote sensing images Extracting shallow features ;
[0029] S22: Using Haar discrete wavelet transform to transform shallow features Processing is performed to extract the frequency domain features of the pre-processed remote sensing image; the frequency domain features include: high frequency information and low frequency information; S23: Transformer-based pyramid structure for shallow features Processing is performed to extract the spatial domain features of the pre-processed remote sensing image; the spatial domain features include: local features and global features.
[0030] As an embodiment, the method adopts a dual-domain feature extraction strategy to extract features of the input low-resolution remote sensing image in the frequency domain and spatial domain respectively; wherein, the frequency domain features obtain the low-frequency and high-frequency components of the image through wavelet transform, which is used to extract the local detail information of the image; the spatial domain features use the Transformer-based pyramid structure to extract global spatial features and local spatial features respectively to model the structural information of the image.
[0031] Step S22 includes: Haar discrete wavelet transform is used to transform shallow features Decompose and obtain four sub-band coefficients, including: low-frequency component and three high-frequency detail components 、 、 ; High-frequency detail components 、 、 Represent the vertical, horizontal and diagonal edge information of the image respectively; Use the enhanced residual to enhance the extraction of the low-frequency component and the low-frequency component respectively, as follows: High frequency sub-band After splicing, input the convolution layer to form the initial high-frequency features; low-frequency component Generate features through convolution , i.e. low-frequency information; The initial high-frequency features are fused and the frequency domain enhancement results are output. , that is, high-frequency information.
[0032] As an embodiment, the steps of frequency domain feature extraction are as follows: Haar wavelet transform is used to extract the feature map Decompose and obtain four sub-band coefficients: low-frequency component and three high-frequency detail components 、 、 , representing the vertical, horizontal and diagonal edge information of the image respectively. In order to further enhance the feature expression ability of the frequency domain, the enhanced residual is used to enhance the low-frequency and high-frequency information respectively. Specifically, the high-frequency sub-band After splicing, the convolution layer is input to form the initial high-frequency features; the low-frequency components Generate features through convolution The above two features are respectively subjected to residual connection and fusion operations. , output frequency domain enhancement result , effectively improving the ability to restore image details and edges. The internal design includes the maximum pooling operation, The convolution, ReLU activation function and scale cascade fusion structure are designed to further explore the high-frequency detail changes hidden in the low-frequency components and improve the model's ability to express complex texture scenes.
[0033] Step S23 includes: Shallow features As input, a four-level pyramid structure is constructed. Each level of feature map is modeled by downsampling and Transformer submodules to achieve different scale context modeling. The feature maps at each level are defined as:
[0034]
[0035] Where H represents the height of the feature map, and W represents the width of the feature map; Indicates the Level feature map; Feature Collection Each level captures local and global dependencies through the Transformer submodule, as follows
[0036] in represents local detail representation, Represents global semantic features; Represents all outputs after P passes through Transformer; Indicates the last output; Represents the first H×W outputs.
[0037] As an example, the bottom layer focuses on extracting information at the local pixel level, while the top layer integrates semantic information across the entire image through a self-attention mechanism, thereby generating a feature representation that combines local details and overall structure. This is used in the subsequent dual-domain cross-fusion module to enhance the expressiveness of image structure and detail information.
[0038] Step S3 includes: As an embodiment, the frequency domain features and spatial domain features are jointly modeled through a dual-domain cross-attention fusion module to enhance the image detail recovery and structure preservation capabilities. The frequency domain feature extraction uses Haar wavelet transform to decompose the input image to obtain the frequency response information of the image at different scales. This transform divides the original image into multiple sub-bands, including low-frequency sub-bands and high-frequency sub-bands. This method obtains the structure and detail sub-bands of the image through frequency domain decomposition and enhancement, laying the foundation for subsequent dual-domain fusion to obtain low-frequency information and high-frequency information The spatial domain feature extraction adopts the Transformer-based pyramid structure to obtain local features. With global features , used to supplement the image context structure in the frequency domain information.
[0039] Construct a dual-domain feature fusion module, including: a cross-domain feature alignment unit, a multi-scale orthogonal feature modeling unit, a query-key-value based cross-domain attention fusion unit, and a channel attention enhancement unit; Through the cross-domain feature alignment unit, the feature map of low-frequency information With local features Carry out scale unification and format adjustment; The format-adjusted Query-Key-Value based cross-domain attention fusion unit is fused through the multi-scale orthogonal feature modeling unit and the Query-Key-Value based cross-domain attention fusion unit; Low-frequency information is enhanced by the channel attention unit and global features to integrate; The fusion process includes the following operations:
[0040] in, is low-frequency information; For high-frequency information; It is a local feature; is a global feature; represents the global structural features after fusion, Represents local detail features, Represents the cross-attention operation to achieve complementary fusion of frequency domain and spatial domain information.
[0041] As an embodiment, the present invention provides a dual-domain feature fusion module for realizing collaborative modeling between frequency domain and spatial domain features. The module includes: a cross-domain feature alignment unit, a multi-scale orthogonal feature modeling unit, and a cross-domain attention fusion and channel attention enhancement unit. Considering that there are multi-scale objects in remote sensing images, whose structures, textures and edge features have different manifestations at different scales, the present invention introduces multi-scale orthogonal convolution MOR to enhance the model's adaptability to scale changes. The cross-domain feature alignment unit is used to align the feature maps from the frequency domain. Local characteristics of the spatial domain Scale unification and format adjustment are performed to ensure the fusion compatibility of different domain features in terms of dimension and semantics. Specifically, the multi-scale orthogonal feature modeling unit includes two parallel convolution branches, using two sets of asymmetric convolution kernels. and (in ), respectively extracting long-range horizontal and vertical dependency information from the image, effectively modeling long-range spatial relationships while maintaining computational efficiency. Furthermore, the MOR module introduces orthogonal convolution constraints. By encouraging orthogonality between feature channels, it enhances feature independence and diversity, reduces redundant information, and improves model generalization performance.
[0042]
[0043] In order to achieve efficient interaction between high-frequency detail information in the frequency domain and local context information in the spatial domain, this paper introduces a cross-domain attention mechanism based on Query-Key-Value.
[0044] Specifically, the MOR module first feeds high-frequency features from the frequency domain and local features from the spatial domain into its multi-scale representations, generating corresponding attention representations. Subsequently, features from each domain are used as query terms to cross-domain correlate key-value pairs from the other domain, thereby achieving bidirectional attention computation. This mechanism allows features from each domain to be reweighted and selected based on context from the other domain, fully capturing cross-domain semantic dependencies. Finally, through feature concatenation, a local detail-enhanced representation is generated, improving the quality of image reconstruction of edges and detailed structures.
[0045]
[0046]
[0047] In addition to local details, the overall structure and low-frequency information of the image are also crucial for global consistency in reconstruction. The channel attention enhancement unit, built on the Squeeze-and-Excitation architecture, performs channel-wise compression, channel weight learning, and recalibration on the input features to enhance the saliency of key structural regions during the fusion process. This unit acts on both global features in the spatial domain and low-frequency components in the frequency domain, improving the overall structural representation of the fused features.
[0048] The implicit representation network consists of a set of coordinate-based implicit decoders; The implicit decoder of each coordinate is constructed by a multi-layer perceptron based on a sinusoidal activation function to characterize the texture changes and structural information in the image; the implicit decoder is based on two-dimensional high-resolution coordinates. It is input, combined with the local potential representation in the corresponding feature domain, and outputs the predicted pixel value at that position.
[0049] Step S4 includes: The fused global structural features are decoded by implicit decoder Local detail features Decode and predict the pixel value of any spatial position in the high-resolution image to obtain the prediction result The prediction results are integrated through the convolution layer to obtain the final three-channel high-resolution image, realizing the mapping from continuous coordinates to pixel values, as follows:
[0050] in Represents a nonlinear function that maps coordinates and fusion features to pixel values, performs pixel-by-pixel implicit function learning, and maps high-resolution coordinates to pixel values; Represents the final three-channel high-resolution image.
[0051] As an embodiment, the present invention provides an adaptive implicit parser structure for fusing global structural information and local detail information in the frequency domain and the spatial domain to achieve pixel-by-pixel reconstruction of high-resolution images. The structure includes two parallel coordinate parsing submodules, a local neighborhood context encoding unit, and a feature fusion unit. The two coordinate parsing submodules correspond to low-frequency features in the frequency domain and high-frequency features in the spatial domain, respectively, to construct a nonlinear mapping function from coordinates to pixel values. Each submodule adopts a multi-layer perceptron based on sinusoidal activation coding (SIREN-MLP), and its input includes the two-dimensional coordinate points in the target high-resolution image. The global structural characteristics of the point Local detail features To further enhance the perception capability, each parser predicts a high-resolution coordinate point. When , an additional Local neighborhood information. The two coordinate resolvers each output a three-channel intermediate image representation, which is then concatenated in the channel dimension to form a six-channel intermediate feature representation.
[0052]
[0053]
[0054] In order to effectively integrate the complementary information of frequency domain and spatial domain, this paper further introduces a The convolution operation fuses the concatenated six-channel features and finally generates a three-channel high-resolution image result.
[0055]
[0056] This implicit reconstruction structure can integrate information from both frequency and spatial domains at the pixel level, which not only preserves structural continuity but also improves the fidelity of edge and texture details, thereby improving image reconstruction quality.
[0057] As an example, Figure 2 The figure shows a flowchart of the present invention for preprocessing an input image in the frequency domain. The input image first undergoes a discrete wavelet transform (DWT) of the Haar wavelet to extract structural and texture information from the image. The input image passes through a low-pass filter G_L and a high-pass filter G_H to extract low-frequency and high-frequency information, respectively. After downsampling, four subband components are ultimately obtained: A, V, H, and D. Among them, A is an approximate coefficient representing the overall structural information of the image, and V, H, and D are detail coefficients corresponding to the vertical, horizontal, and diagonal edge features in the image, respectively. By operating in the frequency domain, an explicit separation of low-frequency structure and high-frequency texture can be achieved, thereby more specifically enhancing high-frequency details and providing rich information support for image super-resolution reconstruction.
[0058] As an example, Figure 3The figure shows the dual-domain implicit super-resolution network structure (DISR) for continuous image reconstruction proposed by the present invention. The method mainly includes the following steps: shallow feature extraction, which passes the input low-resolution image through multiple convolutional layers to extract basic features; dual-domain feature decomposition, and the extracted features are then divided into frequency domain branches and spatial domain branches: the frequency domain branch obtains low-frequency (structure) and high-frequency (detail) features through Haar wavelet transform; the spatial domain branch uses the Transformer structure to model global and local spatial semantic information; dual-domain cross-fusion, high-frequency features and local spatial information interact through a cross-attention mechanism to generate detail representation; low-frequency features interact with global spatial information to generate structure representation; implicit modeling and reconstruction, the two fused representations are respectively input into two independent implicit function modeling modules (coordinate resolvers), the two-dimensional image coordinate point x and the corresponding contextual features are input, and the pixel value at the coordinate is predicted using an MLP based on sine activation; the two coordinate resolvers respectively output three-channel image prediction results; the final three-channel high-resolution image is obtained through splicing and convolution fusion operations.
[0059] This application also discloses an electronic device. Figure 4 , Figure 4 Schematic diagram of the structure of an electronic device disclosed in an embodiment of the present application. The electronic device 500 may include: at least one processor 501, at least one network interface 504, a user interface 503, a memory 505, and at least one communication bus 502.
[0060] The communication bus 502 is used to implement the connection and communication between these components.
[0061] The user interface 503 may include a display screen, and the optional user interface 503 may also include a standard wired interface or a wireless interface.
[0062] The network interface 504 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).
[0063] The present application also discloses a computer-readable storage medium storing a plurality of instructions suitable for loading by a processor to execute the above-mentioned image super-resolution reconstruction method based on dual-domain feature fusion and implicit representation.
[0064] The above are merely exemplary embodiments of the present disclosure and are not intended to limit the scope of the present disclosure. In other words, any equivalent changes and modifications made according to the teachings of the present disclosure are still within the scope of the present disclosure.
[0065] This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not described herein. The description and examples are to be considered as exemplary only, and the scope and spirit of the present disclosure are to be defined by the claims.
Claims
1. An image super-resolution reconstruction method based on dual-domain feature fusion and implicit representation, characterized in that: The method comprises the following steps: S1: Acquire low-resolution remote sensing images and perform preprocessing; S2: Feature extraction of pre-processed remote sensing images is performed through Haar discrete wavelet transform and Transformer-based pyramid structure to obtain frequency domain features and spatial domain features; S3: Perform dual-domain cross-attention fusion on frequency domain features and spatial domain features to obtain local detail features and global structure features; S4: Through the implicit representation network, the fused local detail features and global structural features are mapped to arbitrary spatial coordinates to obtain the final three-channel high-resolution image, realizing high-quality reconstruction of remote sensing images of any scale.
2. The image super-resolution reconstruction method based on dual-domain feature fusion and implicit representation according to claim 1, characterized in that: Step S2 includes: S21: Through encoder Preprocessed remote sensing images Extracting shallow features ; S22: Using Haar discrete wavelet transform to transform shallow features Processing is performed to extract the frequency domain features of the pre-processed remote sensing image; the frequency domain features include: high frequency information and low frequency information; S23: Transformer-based pyramid structure for shallow features Processing is performed to extract the spatial domain features of the pre-processed remote sensing image; the spatial domain features include: local features and global features.
3. The image super-resolution reconstruction method based on dual-domain feature fusion and implicit representation according to claim 2, characterized in that: Step S22 includes: Haar discrete wavelet transform is used to transform shallow features Decompose and obtain four sub-band coefficients, including: low-frequency component and three high-frequency detail components 、 、 ; High-frequency detail components 、 、 Represent the vertical, horizontal and diagonal edge information of the image respectively; Use the enhanced residual to enhance the extraction of the low-frequency component and the low-frequency component respectively, as follows: High frequency sub-band After splicing, input the convolution layer to form the initial high-frequency features; low-frequency component Generate features through convolution , i.e. low-frequency information; The initial high-frequency features are fused and the frequency domain enhancement results are output. , that is, high-frequency information.
4. The image super-resolution reconstruction method based on dual-domain feature fusion and implicit representation according to claim 2, characterized in that: Step S23 includes: Shallow features As input, a four-level pyramid structure is constructed. Each level of feature map is modeled by downsampling and Transformer submodules to achieve different scale context modeling. The feature maps at each level are defined as: Where H represents the height of the feature map, and W represents the width of the feature map; Indicates the Level feature map; Feature Collection Each level captures local and global dependencies through the Transformer submodule, as follows in represents local detail representation, Represents global semantic features; Represents all outputs after P passes through Transformer; Indicates the last output; Represents the first H×W outputs.
5. The image super-resolution reconstruction method based on dual-domain feature fusion and implicit representation according to claim 2, characterized in that: Step S3 includes: Construct a dual-domain feature fusion module, including: a cross-domain feature alignment unit, a multi-scale orthogonal feature modeling unit, a query-key-value based cross-domain attention fusion unit, and a channel attention enhancement unit; Through the cross-domain feature alignment unit, the feature map of low-frequency information With local features Carry out scale unification and format adjustment; The format-adjusted Query-Key-Value based cross-domain attention fusion unit is fused through the multi-scale orthogonal feature modeling unit and the Query-Key-Value based cross-domain attention fusion unit; Low-frequency information is enhanced by the channel attention unit and global features to integrate; The fusion process includes the following operations: in, is low-frequency information; For high-frequency information; It is a local feature; is a global feature; represents the global structural features after fusion, Represents local detail features, Represents the cross-attention operation to achieve complementary fusion of frequency domain and spatial domain information.
6. The image super-resolution reconstruction method based on dual-domain feature fusion and implicit representation according to claim 5, characterized in that: The implicit representation network consists of a set of coordinate-based implicit decoders; The implicit decoder of each coordinate is constructed by a multi-layer perceptron based on a sinusoidal activation function to characterize the texture changes and structural information in the image; the implicit decoder is based on two-dimensional high-resolution coordinates. It is input, combined with the local potential representation in the corresponding feature domain, and outputs the predicted pixel value at that position.
7. The image super-resolution reconstruction method based on dual-domain feature fusion and implicit representation according to claim 6, characterized in that: Step S4 includes: The fused global structural features are decoded by implicit decoder Local detail features Decode and predict the pixel value of any spatial position in the high-resolution image to obtain the prediction result The prediction results are integrated through the convolution layer to obtain the final three-channel high-resolution image, realizing the mapping from continuous coordinates to pixel values, as follows: in Represents a nonlinear function that maps coordinates and fusion features to pixel values, performs pixel-by-pixel implicit function learning, and maps high-resolution coordinates to pixel values; Represents the final three-channel high-resolution image.
8. An electronic device, characterized in that: The electronic device comprises a processor, a memory, a user interface and a network interface, wherein the memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, and when the instructions are executed by a computer, the method according to any one of claims 1 to 7 is executed.
Citation Information
Cited By
Underwater visual image efficient reconstruction method and system based on implicit entropy learning
CN121074161A
Remote sensing image hiding and recovering method and device based on space-frequency collaborative modeling
CN121767156A
Image processing method, computing device, computer readable storage medium and computer program product
CN121861436A
Remote sensing image super-resolution method based on frequency-space collaborative cross attention network
CN122048660A
A multispectral remote sensing image super-resolution method and system with distortion perception routing and parameter sharing mechanism
CN122312382B