Super-resolution reconstruction method and device for remote sensing image of complex terrain building and medium

By employing a temporal prior bi-branch reconstruction model, and utilizing feature extraction, semantic alignment, and multi-scale convolutional attention processing, the problems of blurring and loss of detail in remote sensing image reconstruction under complex terrain are solved, achieving high-precision image reconstruction.

CN121481852AActive Publication Date: 2026-02-06INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202610025876.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-09
Publication Date
2026-02-06
Estimated Expiration
2046-01-09

AI Technical Summary

Technical Problem

In complex terrain scenarios, existing technologies for super-resolution reconstruction of remote sensing images suffer from performance degradation, resulting in blurred structures and loss of details in the reconstruction results, which affects the accuracy of ground feature identification and change detection.

Method used

A temporal prior dual-branch reconstruction model is adopted. The main branch and the guiding branch are used to extract features and semantically align the image. Feature difference information is calculated and a change weight map is generated. Adaptive weighted fusion is performed and combined with multi-scale convolutional attention processing to finally decode and generate a high-resolution reconstructed image.

Benefits of technology

Without relying on explicit terrain data, the system achieves temporal consistency and high-fidelity reconstruction of remote sensing images of buildings in complex terrain, overcoming the problems of single information utilization, unstable registration and fusion, and lack of change perception, thus improving the accuracy and stability of the reconstruction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121481852A_ABST
    Figure CN121481852A_ABST
Patent Text Reader

Abstract

The invention discloses a super-resolution reconstruction method and device for a remote sensing image of a complex terrain building and a medium. The method comprises the following steps: acquiring a historical low-resolution image and a current high-resolution image; performing feature extraction on the historical low-resolution image by reconstructing a main branch to obtain a main branch feature; performing feature extraction and semantic alignment on the current high-resolution image through a guide branch to obtain a guide branch feature; calculating the difference between the main branch features and the guide branch features to obtain feature difference information, and determining a change weight map based on the feature difference information; performing adaptive weighted fusion on the main branch features and the guide branch features according to the change weight map to obtain fusion features; performing multi-scale convolution attention processing on the fused features to obtain enhanced features; and decoding the enhanced features to obtain a target high-resolution reconstructed image. According to the invention, high-precision and stable image reconstruction can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the fields of computer technology and complex terrain building remote sensing image reconstruction, and more particularly, to a complex terrain building remote sensing image super-resolution reconstruction method, device and medium in the field of computer technology. BACKGROUND

[0002] The super-resolution reconstruction technology of remote sensing images is a method of recovering high-resolution details through algorithms, which is often used to improve the spatial resolution of remote sensing images to meet the application requirements of ground monitoring, urban planning, etc. However, in complex terrain (such as mountainous or hilly) scenes, existing methods often face the challenge of performance degradation, which may be due to the influence of factors such as terrain undulation, light change and shadow, resulting in problems such as structural blurring or detail loss in the reconstruction results. Such problems may affect the accuracy of subsequent analysis tasks such as feature recognition or change detection. Therefore, how to effectively realize the super-resolution reconstruction of complex terrain building remote sensing images has become a technical problem to be solved. SUMMARY

[0003] The embodiments of the present application provide a complex terrain building remote sensing image super-resolution reconstruction method, device and medium, which can utilize the complementary information between multi-temporal images, use the current high-resolution image as a guide, complete the details of the historical low-resolution image and the temporal consistency constraint, thereby realizing high-precision and stable image reconstruction.

[0004] In a first aspect, the embodiments of the present application provide a complex terrain building remote sensing image super-resolution reconstruction method, comprising: obtaining a historical low-resolution image and a current high-resolution image, the historical low-resolution image and the current high-resolution image being remote sensing images of the same geographical region obtained at different time points; extracting features from the historical low-resolution image through a reconstruction main branch to obtain main branch features; extracting features from the current high-resolution image through a guide branch and performing semantic alignment to obtain guide branch features; calculating the difference between the main branch features and the guide branch features to obtain feature difference information, and determining a change weight map based on the feature difference information; adaptively weighting and fusing the main branch features and the guide branch features according to the change weight map to obtain fused features; performing multi-scale convolution attention processing on the fused features to obtain enhanced features; decoding the enhanced features to obtain a target high-resolution reconstructed image.

[0005] In a second aspect, the embodiments of the present application provide a device for super-resolution reconstruction of remote sensing images of complex terrain buildings, comprising: an image acquisition unit configured to acquire a historical low-resolution image and a current high-resolution image, the historical low-resolution image and the current high-resolution image being remote sensing images of the same geographical region acquired at different time points; a first feature acquisition unit configured to perform feature extraction on the historical low-resolution image by a reconstruction main branch to obtain main branch features; a second feature acquisition unit configured to perform feature extraction on the current high-resolution image by a guide branch and perform semantic alignment to obtain guide branch features; an information acquisition unit configured to calculate a difference between the main branch features and the guide branch features to obtain feature difference information, and determine a change weight map based on the feature difference information; a third feature acquisition unit configured to perform adaptive weighted fusion on the main branch features and the guide branch features according to the change weight map to obtain fused features; a fourth feature acquisition unit configured to perform multi-scale convolution attention processing on the fused features to obtain enhanced features; a feature decoding unit configured to decode the enhanced features to obtain a target high-resolution reconstructed image.

[0006] In a third aspect, the embodiments of the present application provide a computer storage medium, which stores a plurality of instructions, the instructions being suitable for being loaded and executed by a processor to perform the method steps described above.

[0007] In a fourth aspect, the embodiments of the present application provide an electronic device, which can include a processor and a memory; wherein the memory stores a computer program, the computer program being suitable for being loaded and executed by the processor to perform the method steps described above.

[0008] In the embodiment of the present application, the historical low-resolution image and the current high-resolution image are first acquired, the main branch feature is extracted from the historical low-resolution image by reconstructing the main branch to construct a structure baseline, and the guided branch feature is extracted from the current high-resolution image by the guided branch to obtain high-frequency priori; then, the feature difference information is obtained by calculating the difference between the main branch feature and the guided branch feature, and the change weight map is determined based on the information, and the main branch feature and the guided branch feature are adaptively weighted and fused according to the change weight map, and the change of the ground object is explicitly modeled through the gating mechanism, so that the temporal consistency is maintained in the unchanged area and the error texture migration is suppressed in the changed area; then, the fused feature is processed by multi-scale convolution attention, which simultaneously captures macro structure and micro detail by using convolution kernels of multiple scales, thereby enhancing the recovery ability of building edges and road textures under complex terrain; finally, the enhanced feature is decoded to generate the target high-resolution reconstruction image, therefore, through the synergistic effect of the above steps, the structure blur and artifact problems caused by single information utilization, unstable registration and fusion and lack of change perception in the prior art are overcome, thereby realizing the temporal consistency and high-fidelity reconstruction of complex terrain building remote sensing images without relying on explicit terrain data. BRIEF DESCRIPTION OF DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0010] Figure 1 A structure diagram of a temporal priori double-branch reconstruction model for complex terrain building remote sensing images provided by the embodiment of the present application; Figure 2 A flowchart of a super-resolution reconstruction method for complex terrain building remote sensing images provided by the embodiment of the present application; Figure 3 A flowchart of a main branch feature acquisition process provided by the embodiment of the present application; Figure 4 A flowchart of a guided branch feature acquisition process provided by the embodiment of the present application; Figure 5 A flowchart of a change weight map acquisition process provided by the embodiment of the present application; Figure 6 A flowchart of a fused feature acquisition process provided by the embodiment of the present application; Figure 7A flowchart of an enhanced feature acquisition process is provided for the embodiments of the present application. Figure 8 A structural diagram of a complex terrain building remote sensing image super-resolution reconstruction device is provided for the embodiments of the present application. Figure 9 A structural diagram of an electronic device is provided for the embodiments of the present application. DETAILED DESCRIPTION

[0011] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0012] Remote sensing image super-resolution reconstruction technology is a key means to improve the spatial resolution of remote sensing images and meet the needs of fine ground monitoring. With the development of deep learning technology, reconstruction methods based on convolutional neural networks and Transformer models have achieved remarkable results under ideal imaging conditions. However, in complex terrain environments such as mountains and hills, the performance of related reconstruction methods is prone to significant degradation due to complex and variable imaging conditions. This degradation often leads to problems such as structure blurring and detail loss in the reconstructed image, seriously affecting the accuracy of subsequent ground object identification and change analysis. Therefore, how to effectively overcome the adverse effects of complex terrain and achieve high-precision and high-stability reconstruction of building remote sensing images has become a technical problem that needs to be solved in this field.

[0013] Specifically, please refer to Figure 1 A structural diagram of a time-series prior double-branch reconstruction model for complex terrain building remote sensing images is provided for the embodiments of the present application. As Figure 1 indicated, the reconstruction model can include a reconstruction main branch 100, a guide branch 200, a change gate fusion module 300, a multi-scale convolution attention module 400, and a decoder 500.

[0014] Among them, the reconstruction main branch 100 includes convolution kernels and residual Swin blocks, the guide branch 200 includes a Swin Transformer encoder, a shift predictor network, and a deformable convolution layer.

[0015] By constructing the reconstruction model, in practical applications, the reconstruction of the main branch 100 extracts features from historical low-resolution images to obtain the main branch features in order to construct the structural baseline. Specifically, the historical low-resolution images are convolved with a preset number of kernels to extract initial features. Then, the initial features are processed hierarchically by residual Swin blocks to output the main branch features of the reconstruction of the main branch 100.

[0016] Simultaneously, the guide branch extracts features from the current high-resolution image and performs semantic alignment to obtain guide branch features and acquire high-frequency priors. Specifically, the Swin Transformer encoder extracts multi-scale features from the current high-resolution image to obtain the first multi-scale features, which provide a basis for subsequent semantic alignment. The Swin Transformer encoder provides the output first multi-scale features to the deformable convolutional layer and the offset prediction sub-network. The offset prediction sub-network can predict the features of the first multi-scale features using a set of learnable parameters with low location cost. Then, the deformable convolutional layer adaptively samples the first multi-scale features according to the set of learnable parameters through deformable convolution to output the guide branch features of guide branch 200.

[0017] Next, the difference between the main branch features and the guiding branch features is calculated by the change gating fusion module 300 to obtain feature difference information. Based on this information, a change weight map is determined. Then, the main branch features and the guiding branch features are adaptively weighted and fused according to this change weight map. This process explicitly models the changes in land cover through a gating mechanism, thereby maintaining temporal consistency in invariant areas and suppressing erroneous texture migration in change areas.

[0018] Then, the fused features are processed by a multi-scale convolutional attention module 400. This process utilizes convolutional kernels of multiple scales to simultaneously capture macroscopic structures and microscopic details, thereby enhancing the ability to restore building edges and road textures in complex terrain. Finally, the enhanced features are decoded by a decoder 500 to generate a high-resolution reconstructed image of the target.

[0019] Through the synergistic effect of the above steps, the embodiments of this application overcome the structural ambiguity and artifact problems caused by the single use of information, unstable registration and fusion, and lack of change perception in the prior art. Thus, without relying on explicit terrain data, the temporal consistency and high-fidelity reconstruction of remote sensing images of complex terrain buildings are achieved.

[0020] based on Figure 1 The model structure shown below will be combined with... Figures 2-7 This application provides a detailed description of the super-resolution reconstruction method for remote sensing images of complex terrain and buildings provided in the embodiments of this application.

[0021] Please seeFigure 2 A flowchart of a complex terrain building remote sensing image super-resolution reconstruction method is provided for the embodiments of the present application. As shown in Figure 2 The method of the embodiments of the present application can include the following steps S101-S107.

[0022] S101, obtaining historical low-resolution images and current high-resolution images; Specifically, the historical low-resolution images involved in the present embodiment can be obtained by satellite sensors or aerial photography; the current high-resolution images can be high-resolution remote sensing images obtained at the latest time point for the same geographical area, wherein low resolution means that the spatial resolution of the image is low and the details are blurred; high resolution means that the spatial resolution of the image is high and the details are clear.

[0023] In order to realize the reconstruction of cross-time images, the historical low-resolution images and the current high-resolution images need to be obtained. As for this step, in some possible implementation manners, the historical low-resolution images and the current high-resolution images can be retrieved from a remote sensing database and preprocessed; in some possible implementation manners, image registration technology can be used to spatially align the historical low-resolution images and the current high-resolution images; in some possible implementation manners, radiation correction can be performed on the historical low-resolution images and the current high-resolution images to eliminate light differences.

[0024] S102, feature extraction is performed on the historical low-resolution images by a reconstruction main branch to obtain main branch features; S103, feature extraction is performed on the current high-resolution images by a guide branch and semantic alignment is performed to obtain guide branch features; Specifically, the present embodiment performs feature extraction on the historical low-resolution images by a reconstruction main branch to obtain main branch features; and performs feature extraction on the current high-resolution images by a guide branch and performs semantic alignment to obtain guide branch features. Wherein, the reconstruction main branch refers to a neural network branch used for processing historical low-resolution images, aiming to extract basic structure features; the main branch features refer to feature maps representing the structure of the historical low-resolution images extracted from the reconstruction main branch; the guide branch refers to a neural network branch used for processing current high-resolution images, aiming to provide high-frequency detail guidance; the guide branch features refer to feature maps representing the details of the current high-resolution images extracted from the guide branch; semantic alignment refers to the process of aligning the guide branch features with the main branch features at the semantic level. In some possible implementation manners, a convolutional neural network can be used as the basic architecture of the reconstruction main branch and the guide branch; in some possible implementation manners, the feature extraction capability of the reconstruction main branch and the guide branch can be enhanced through an attention mechanism.

[0025] S104, calculate the difference between the main branch feature and the guide branch feature to obtain feature difference information, and determine a change weight map based on the feature difference information; Specifically, the embodiment calculates the difference between the main branch feature and the guide branch feature to obtain feature difference information, and determines a change weight map based on the feature difference information. The feature difference information refers to a measure representing the difference between the main branch feature and the guide branch feature. The change weight map refers to a weight map generated based on the feature difference information, which is used to indicate the change degree of each pixel or region. In some possible implementation ways, the absolute difference between the main branch feature and the guide branch feature can be calculated as the feature difference information, and then a change weight map is generated through a gating network. In some possible implementation ways, the feature difference information can be processed by a pooling operation to generate a change weight map.

[0026] S105, adaptively weighted fusion of the main branch feature and the guide branch feature according to the change weight map to obtain a fusion feature; Specifically, the embodiment adaptively weighted fusion of the main branch feature and the guide branch feature according to the change weight map to obtain a fusion feature. The fusion feature refers to a weighted feature map combining the main branch feature and the guide branch feature. In some possible implementation ways, the change weight map can be applied to the guide branch feature and the main branch feature to obtain the fusion feature through element multiplication and addition operations. In some possible implementation ways, the main branch feature and the guide branch feature can be fused using a weighted average strategy to generate the fusion feature.

[0027] S106, multi-scale convolution attention processing of the fusion feature to obtain an enhanced feature; Specifically, the embodiment multi-scale convolution attention processing of the fusion feature to obtain an enhanced feature. The enhanced feature refers to a feature map after multi-scale convolution attention processing, which has enhanced details and context information. In some possible implementation ways, multiple convolution kernel scales can be used to extract the fusion feature, and the results are fused through an attention mechanism. In some possible implementation ways, a channel attention module can be introduced after multi-scale convolution to generate the enhanced feature.

[0028] S107, decoding of the enhanced feature to obtain a target high-resolution reconstructed image; Specifically, the enhanced feature is decoded to obtain a target high-resolution reconstruction image. The target high-resolution reconstruction image refers to a final output high-resolution image, representing the reconstruction result at a historical time point. In some possible implementation manners, an up-sampling and convolution layer can be used as a decoder to convert the enhanced feature into a pixel-level image; in some possible implementation manners, an inverse convolution operation can be used to enhance the edge details of the target high-resolution reconstruction image.

[0029] Regarding this step, in some possible implementation manners, post-processing can be performed on the target high-resolution reconstruction image, including denoising and sharpening, to improve the visual quality.

[0030] In the embodiments of the present application, a historical low-resolution image and a current high-resolution image are first obtained, a main branch feature is extracted from the historical low-resolution image to construct a structure baseline, and a guide branch feature is extracted from the current high-resolution image to obtain high-frequency priori information. Then, the difference between the main branch feature and the guide branch feature is calculated to obtain feature difference information, and a change weight map is determined based on the information. The main branch feature and the guide branch feature are adaptively weighted and fused according to the change weight map. The process explicitly models the feature change through a gating mechanism, thereby maintaining temporal consistency in unchanged areas and suppressing error texture migration in changed areas. Then, the fused feature is subjected to multi-scale convolution attention processing. This process uses convolution kernels of multiple scales to capture macro structures and micro details at the same time, thereby enhancing the recovery ability of building edges and road textures under complex terrain. Finally, the enhanced feature is decoded to generate a target high-resolution reconstruction image. Therefore, through the synergistic effect of the above steps, the present application overcomes the structure blurring and artifact problems caused by the single information utilization, unstable registration and fusion, and lack of change perception in the prior art, thereby achieving temporal consistency and high-fidelity reconstruction of complex terrain building remote sensing images without relying on explicit terrain data.

[0031] Please refer to Figure 3 A flowchart of a complex terrain building remote sensing image super-resolution reconstruction method is provided for the embodiments of the present application. As shown in Figure 3 The method of the embodiments of the present application can include the following steps S201-S202.

[0032] S201, performing convolution mapping on the historical low-resolution image to obtain an initial feature; S202, performing hierarchical processing on the initial feature based on a residual Swin block to obtain the main branch feature; In detail, the historical low-resolution image referred to in the embodiment refers to low-resolution remote sensing image data representing a geographical area acquired from a historical time point, used to reflect the spatial information of the ground surface at a past time; the initial feature refers to initial feature representation generated after convolution operation on the historical low-resolution image, used as the basis input for subsequent feature extraction; in order to extract low-level visual information from the historical low-resolution image, the historical low-resolution image needs to be subjected to convolution mapping to obtain the initial feature. Wherein, the initial feature refers to feature data containing image basic texture and edge information, serving as the input source of hierarchical processing. As to this step, in some possible implementation manners, a convolution kernel of a preset size can be used to perform convolution operation on the historical low-resolution image to generate the initial feature; in some possible implementation manners, the convolution result can be processed in combination with an activation function and a normalization step to obtain the initial feature.

[0033] The calculation formula of the initial feature is: F0=Conv 3×3 (L t ); Wherein, F0 represents the initial feature; Conv 3×3 represents a 3x3 convolution operation, used for low-level feature extraction on the input image; L t represents the historical low-resolution image.

[0034] Based on the obtained initial feature, the residual Swin block referred to in the embodiment refers to a neural network module integrating residual connection and SwinTransformer structure, used to enhance the feature extraction capability; the main branch feature refers to deep semantic feature representation extracted after hierarchical processing of the initial feature, used to reconstruct the structural information of the image; in order to capture multi-level context dependency from the initial feature, the initial feature needs to be subjected to hierarchical processing based on the residual Swin block to obtain the main branch feature. Wherein, the main branch feature refers to a feature map representing the geometric contour and spatial relationship of the historical low-resolution image. As to this step, in some possible implementation manners, the initial feature can be input into multiple residual Swin blocks connected in series for processing to extract the main branch feature; in some possible implementation manners, the depth of feature extraction can be controlled by adjusting the number of stacked layers of the residual Swin block to obtain the main branch feature.

[0035] The calculation formula of the main branch feature is: F r =E r (F0); Wherein, F r represents the main branch feature; E rAn encoder structure composed of residual Swin blocks, used for hierarchical processing of initial features; F0 represents the initial features.

[0036] In the embodiments of the present application, through the processing of the convolutional mapping and the residual Swin block, the main and branch features with local texture details and macro spatial structure can be effectively extracted from historical low-resolution images. The initial features generated by the convolutional mapping operation retain the basic edges and texture information of the image, providing a stable input for subsequent deep feature extraction. The hierarchical processing based on the residual Swin block uses its residual connection and windowed self-attention mechanism to enhance the efficiency of feature transmission while capturing long-range context dependencies, thereby more accurately constructing the geometric shapes of key features such as building outlines and road directions. Therefore, the main and branch features formed by this process not only represent the static structure baseline at the historical moment, but also provide a reliable spatial reference for the subsequent dynamic alignment and fusion of guiding features, which is the basis for realizing cross-time high-fidelity reconstruction.

[0037] Please refer to Figure 4 A flowchart of a complex terrain building remote sensing image super-resolution reconstruction method is provided in the embodiments of the present application. As shown in Figure 4 The method of the embodiments of the present application can include the following steps S301-S303.

[0038] S301, using a lightweight Swin Transformer encoder to process the current high-resolution image to obtain first multi-scale features; Specifically, the lightweight Swin Transformer encoder involved in the present embodiment refers to a lightweight feature extraction network based on the Transformer architecture, which aims to efficiently extract multi-level semantic features; the current high-resolution image refers to high-resolution remote sensing image data of the same geographical area obtained at the latest time point.

[0039] In order to extract rich multi-scale features from the current high-resolution image and provide a basis for subsequent semantic alignment, a lightweight Swin Transformer encoder needs to be used to process the current high-resolution image to obtain first multi-scale features. The first multi-scale features refer to feature representations containing different spatial resolutions and semantic levels, which are used to capture global context and local texture information.

[0040] In some possible implementation manners, an input of the lightweight Swin Transformer encoder can be set as the current high-resolution image, and a first multi-scale feature can be extracted by stacking multiple Swin Transformer blocks. In some possible implementation manners, the output of the lightweight Swin Transformer encoder can be down-sampled to enhance the scale diversity of the first multi-scale feature.

[0041] A calculation formula of the first multi-scale feature is as follows: F g =E g (H T ); wherein, F g represents the first multi-scale feature, E g represents the lightweight Swin Transformer encoder, and H T represents the current high-resolution image.

[0042] S302, predicting a learnable offset field based on the first multi-scale feature; Specifically, the learnable offset field referred to in this embodiment refers to a set of learnable parameters used to indicate the offset of the feature sampling position.

[0043] In order to realize the spatial alignment of cross-time features and reduce the geometric misalignment caused by terrain undulations and illumination changes, it is necessary to predict a learnable offset field based on the first multi-scale feature. Regarding this step, in some possible implementation manners, the first multi-scale feature can be input into an offset prediction subnetwork to generate the learnable offset field. In some possible implementation manners, the first multi-scale feature can be processed through a convolution layer to obtain the learnable offset field.

[0044] A calculation formula of the learnable offset field is as follows: Δ=P(F g ); wherein, Δ represents the learnable offset field, and P represents the offset prediction subnetwork.

[0045] S303, adaptively sampling the first multi-scale feature through a deformable convolution based on the learnable offset field to obtain the guide branch feature. Specifically, the guide branch feature referred to in this embodiment refers to a feature representation after spatial alignment and detail enhancement, and is used to guide the historical image reconstruction.

[0046] In order to obtain features consistent with the main branch feature semantics and spatially aligned, and to promote cross-time information fusion, the first multi-scale features are adaptively sampled by deformable convolution according to the learnable offset field to obtain the guided branch features. Regarding this step, in some possible implementation manners, the learnable offset field can be taken as a sampling control parameter, and the first multi-scale features are subjected to deformable convolution operation to generate the guided branch features. In some possible implementation manners, the first multi-scale features are resampled by bilinear interpolation combined with the learnable offset field to obtain the guided branch features.

[0047] The calculation formula of the guided branch features is as follows: F g ′=DeformConv(F g ,Δ); Wherein, F g ′ represents the guided branch features, and DeformConv represents the deformable convolution operation.

[0048] In the embodiments of the present application, the current high-resolution image is processed by the lightweight Swin Transformer encoder, which can effectively extract the first multi-scale features containing rich texture and structure information. Subsequently, the learnable offset field predicted based on the first multi-scale features provides accurate geometric correction information for subsequent feature spatial alignment. Finally, the deformable convolution adaptively samples the first multi-scale features according to the learnable offset field, so that the guided branch features obtained by sampling retain high-frequency details while achieving semantic alignment with the feature space of the historical low-resolution image. This process effectively overcomes the cross-time feature misalignment problem caused by terrain undulations, view angle differences and illumination changes, ensuring the accuracy and reliability of the guided branch features as reconstruction priors, and laying a high-quality input foundation for subsequent change gating fusion.

[0049] Please refer to Figure 5 , a flowchart of a complex terrain building remote sensing image super-resolution reconstruction method is provided. As Figure 5 shown, the method of the embodiments of the present application can include the following steps S401-S402.

[0050] S401, input the feature difference information into a preset gating network; Specifically, the feature difference information related to the present embodiment refers to the measurement information representing the difference degree between the main branch features and the guided branch features. In order to generate a change weight map to control the feature fusion ratio, the feature difference information needs to be input into a preset gating network. The preset gating network refers to a lightweight neural network structure used to predict weight values based on input features.

[0051] The calculation formula of the feature difference information can be: ΔF = |F r – F g ′ |; Wherein, ΔF represents the feature difference information, F r represents the main branch feature, and F g ′ represents the guide branch feature.

[0052] S402, output the change weight map through the gating network; Specifically, the change weight map is output through the gating network. The change weight map refers to a weight map generated based on the feature difference information, and is used to indicate the change degree of each pixel or region.

[0053] In some possible implementation manners, the feature difference information can be processed through the gating network, a Sigmoid function is applied to constrain the output to the interval [0, 1], and further, the output of the gating network can be normalized to obtain the change weight map.

[0054] The prediction formula of the change weight map can be: M = σ(G(ΔF)); Wherein, M represents the change weight map, σ represents an activation function, G represents the gating network, and ΔF represents the feature difference information.

[0055] In the embodiment of the present application, the feature difference information between the main branch feature and the guide branch feature is calculated, and the information is input into the gating network to generate the change weight map, so that the explicit modeling of the cross-time feature difference is realized. The change weight map serves as a gating signal, and can adaptively control the proportion of the main branch feature and the guide branch feature during fusion. For a region with small feature difference, the change weight map gives the guide branch feature a high weight, so as to maintain the time sequence consistency; for a region with significant feature difference, the change weight map reduces the influence of the guide branch feature, and relies more on the main branch feature, so as to suppress the error texture migration caused by real object change or light difference. This adaptive fusion mechanism based on feature difference enables the model to distinguish between “real change” and “light or angle difference”, while maintaining the structure updating capability in a dynamic scene, and improves the authenticity and time sequence consistency of the reconstruction result.

[0056] Please refer to Figure 6 , a flowchart of a complex terrain building remote sensing image super-resolution reconstruction method is provided. As Figure 6 shown, the method of the embodiment of the present application can include the following steps S501-S503.

[0057] S501, multiply the change weight map and the guide branch feature to obtain a first weighted feature; S502, point-multiplying the change weight map with the main branch feature after taking the inverse processing to obtain a second weighted feature; S503, merging the first weighted feature and the second weighted feature to obtain the fusion feature; Specifically, in order to realize the adaptive weighting of the guide branch feature in feature fusion, the change weight map needs to be point-multiplied with the guide branch feature to obtain a first weighted feature; wherein the first weighted feature refers to the guide branch feature weighted by the change weight map, and is used to reflect the influence degree of the current time phase information in the fusion process. In some possible implementation manners, the change weight map can be applied to the guide branch feature through element-by-element multiplication operation to generate the first weighted feature.

[0058] The change weight map is point-multiplied with the main branch feature after taking the inverse processing to obtain a second weighted feature; wherein the second weighted feature refers to the main branch feature weighted by the complement of the change weight map, and is used to reflect the influence degree of the historical time phase information in the fusion process. In some possible implementation manners, the complement of the change weight map (i.e. 1 minus the change weight map) can be calculated, and the element-by-element multiplication operation is performed with the main branch feature to generate the second weighted feature.

[0059] The first weighted feature and the second weighted feature are merged to obtain the fusion feature; wherein the fusion feature refers to the final feature representation combining the main branch feature and the guide branch feature, and is used for subsequent multi-scale processing and decoding reconstruction. In some possible implementation manners, the first weighted feature and the second weighted feature can be merged through addition operation to generate the fusion feature.

[0060] Optionally, the calculation formula of the fusion feature can be: F f =M⊙F g ′+(1-M)⊙F r ; Wherein, F f represents the fusion feature; M represents the change weight map; F g ′ represents the guide branch feature; F r represents the main branch feature.

[0061] In the embodiment of the present application, by calculating the feature difference information between the main branch feature and the guide branch feature, and inputting the information into the gating network to generate the change weight map, the explicit modeling of the cross-time feature difference is realized. The change weight map as a gating signal can adaptively control the proportion of the main branch feature and the guide branch feature in the fusion. For the area with small feature difference, the change weight map gives the guide branch feature a higher weight to maintain the time consistency; for the area with significant feature difference, the change weight map reduces the influence of the guide branch feature and relies more on the main branch feature, thereby suppressing the error texture migration caused by the real object change or light difference. This adaptive fusion mechanism based on feature difference enables the model to distinguish between "real change" and "light or angle difference", while maintaining the structure updating capability in the dynamic scene, and improves the authenticity and time consistency of the reconstruction result.

[0062] Please refer to Figure 7 A flowchart of a complex terrain building remote sensing image super-resolution reconstruction method is provided in the embodiment of the present application. As shown in Figure 7 The method of the embodiment of the present application can include the following steps S601-S602.

[0063] S601, respectively using a plurality of scale convolution kernels to perform convolution processing on the fusion feature to obtain a plurality of groups of convolution features of different scales; S602, adaptively fusing the plurality of groups of convolution features of different scales through a channel attention mechanism to obtain the enhanced feature; Specifically, the plurality of groups of convolution kernels of different scales referred to in the embodiment refer to a set of convolution operations with different receptive field sizes, which are used to capture multi-level context information; the fusion feature refers to a feature representation obtained after processing by the change gating fusion module, which combines the information of the historical image and the current image; In order to enhance the detail recovery capability of the feature, a plurality of scale convolution kernels are respectively used to perform convolution processing on the fusion feature to obtain a plurality of groups of convolution features of different scales. Among them, the plurality of groups of convolution features of different scales refer to a plurality of groups of feature maps extracted from the fusion feature under different receptive fields, which respectively capture local edge, regional texture and global structure information.

[0064] In some possible implementation manners, 3x3 convolution kernels, 5x5 convolution kernels and 7x7 convolution kernels, etc. can be used to perform parallel convolution operations on the fusion feature to generate a plurality of groups of convolution features of different scales, ensuring that the fusion feature is completely utilized in processing.

[0065] On this basis, the channel attention mechanism refers to a mechanism for automatically learning the importance weight of different feature channels; To improve the pertinence and effectiveness of feature expression, the multiple sets of convolution features of different scales are adaptively fused through a channel attention mechanism to obtain the enhanced features. The enhanced features refer to feature representations processed by the multi-scale convolution attention mechanism, and have enhanced details and context information. Further, the adaptive fusion process can calculate the statistics of each scale feature through global average pooling, and then generate weight values through a fully connected layer. The importance weight can specifically represent the normalized coefficient output by the attention module.

[0066] In the embodiments of the present application, three scale convolutions are provided, which are 3x3 convolution kernel, 5x5 convolution kernel and 7x7 convolution kernel. The output convolution features can be F1, F2 and F3, which correspond to feature maps of different scales respectively. Wherein, F1=Conv 3×3 (F f ), F2=Conv 5×5 (F f ), F3=Conv 7×7 (F f ), F f represents the fusion feature.

[0067] The importance weight can be calculated, and the specific calculation formula can be: α i =Softmax(W i ×GAP(F i )); Wherein, αi represents the importance weight of the i-th scale feature, i can take values 1, 2 and 3, W i represents a learnable weight matrix, and GAP(F i ) represents a global average pooling operation on the i-th convolution feature.

[0068] After obtaining the importance weight of the convolution of different scales, the enhanced feature can be calculated by the following formula: ; Wherein , F h represents the enhanced feature , α i represents the importance weight , F i represents the convolution feature of the i-th scale.

[0069] After obtaining the enhanced feature, the enhanced feature can be further decoded to obtain the target high-resolution reconstructed image.

[0070] In the embodiment of the present application, the fusion features are processed by multi-scale convolution kernels respectively, which can extract structural information at different levels from local to global, so that the model can perceive both the edge details and the overall layout of the building. Subsequently, the channel attention mechanism generates weights representing the importance of features at different scales by globally average pooling the multiple groups of convolution features at different scales. The weights can adaptively adjust the contribution ratio of features at different scales in the fusion. Finally, by multiplying the weights with the corresponding convolution features and summing them up, the enhanced features not only retain the fine texture extracted by 3x3 convolution, but also fuse the macro profile captured by 7x7 convolution. Therefore, the multi-scale convolution attention module in the embodiment dynamically optimizes the fusion features in function, so that the reconstructed image can maintain edge sharpness and structural continuity in complex structures such as mountainous building-dense areas and road intersections, thereby effectively solving the problem of detail loss caused by fixed receptive field in traditional methods under complex terrain.

[0071] In the embodiment of the present application, the reconstructed model formed by the above method can use a joint loss function during training. The specific implementation process can be: calculating the reconstruction loss between the target high-resolution reconstructed image and the real image, the perception loss based on the feature space, the edge loss based on the gradient, and the gating constraint loss based on the change weight map, and weighting and summing the loss terms through weight coefficients.

[0072] The formula of the joint loss function is: L = λ1 × L rec + λ2 × L perc + λ3 × L edge + λ4 × L gate ; Wherein, L represents the joint loss function, λ1, λ2, λ3, λ4 represent weight coefficients for balancing the optimization intensity of each loss term, L rec represents the reconstruction loss, which is a loss function for constraining pixel-level accuracy, L perc represents the perception loss, which is a loss function for maintaining texture consistency in the feature space, L edge represents the edge loss, which is a loss function for strengthening the clarity of the boundaries of buildings and roads, L gate represents the gating constraint loss, which is a loss function for guiding the learning of the change weight map, which is determined based on the difference between the change weight map and the real change mask. The real change mask can be data indicating the real change area of ground objects between the historical low-resolution image and the current high-resolution image, such as data generated by manual annotation or high-precision change detection algorithm.

[0073] The following will be combined with Figure 8This application provides a detailed description of the super-resolution reconstruction apparatus for remote sensing images of complex terrain and buildings, as provided in the embodiments of this application. It should be noted that... Figure 8 The super-resolution reconstruction apparatus for remote sensing images of complex terrain and buildings is specifically used to perform the functions described in this application. Figures 1-7 The methods shown in the embodiments are for illustrative purposes only, illustrating the parts relevant to the embodiments of this application. For specific technical details not disclosed, please refer to this application. Figures 1-7 The example shown.

[0074] Please see Figure 8 This is a schematic diagram of the structure of a super-resolution reconstruction device for remote sensing images of complex terrain and buildings, provided in an embodiment of this application. Figure 8 As shown, the device 1 may include an image acquisition unit 11, a first feature acquisition unit 12, a second feature acquisition unit 13, an information acquisition unit 14, a third feature acquisition unit 15, a fourth feature acquisition unit 16, and a feature decoding unit 17.

[0075] The image acquisition unit 11 is used to acquire historical low-resolution images and current high-resolution images, wherein the historical low-resolution images and the current high-resolution images are remote sensing images of the same geographical area acquired at different time points. The first feature acquisition unit 12 is used to extract features from the historical low-resolution image by reconstructing the main branch, and obtain the main branch features. The second feature acquisition unit 13 is used to extract features from the current high-resolution image and perform semantic alignment through the guide branch to obtain guide branch features; The information acquisition unit 14 is used to calculate the difference between the main branch feature and the guiding branch feature, obtain feature difference information, and determine the change weight map based on the feature difference information; The third feature acquisition unit 15 is used to adaptively weight and fuse the main branch features and the guiding branch features according to the change weight map to obtain the fused features; The fourth feature acquisition unit 16 is used to perform multi-scale convolutional attention processing on the fused features to obtain enhanced features; The feature decoding unit 17 is used to decode the enhanced features to obtain a high-resolution reconstructed image of the target.

[0076] In one embodiment, the first feature acquisition unit 12 is specifically used for: The historical low-resolution images are convolutionally mapped to obtain initial features; The initial features are hierarchically processed based on the residual Swin block to obtain the main branch features.

[0077] In one embodiment, the second feature acquisition unit 13 is specifically used for: The current high-resolution image is processed by using a lightweight Swin Transformer encoder to obtain a first multi-scale feature; A learnable offset field is predicted based on the first multi-scale feature; According to the learnable offset field, the first multi-scale feature is adaptively sampled by a deformable convolution to obtain the guide branch feature.

[0078] In one embodiment, the information acquisition unit 14 is specifically configured to: Calculate the difference between the main branch feature and the guide branch feature to obtain feature difference information; Input the feature difference information into a preset gating network; Output the change weight map through the gating network.

[0079] In one embodiment, the third feature acquisition unit 15 is specifically configured to: Point multiply the change weight map and the guide branch feature to obtain a first weighted feature; Point multiply the change weight map after the inversion processing and the main branch feature to obtain a second weighted feature; Merge the first weighted feature and the second weighted feature to obtain the fusion feature.

[0080] In one embodiment, the fourth feature acquisition unit 16 is specifically configured to: Convolve the fusion feature using a plurality of scale convolution kernels to obtain a plurality of groups of convolution features of different scales; Adaptively fuse the plurality of groups of convolution features of different scales through a channel attention mechanism to obtain the enhanced feature.

[0081] In one embodiment, the plurality of scale convolution kernels include 3x3 convolution kernels, 5x5 convolution kernels, and 7x7 convolution kernels.

[0082] In one embodiment, when performing the step of adaptively fusing the plurality of groups of convolution features of different scales through a channel attention mechanism to obtain the enhanced feature, the fourth feature acquisition unit 16 specifically performs the following operations: Perform global average pooling processing on each group of convolution features to determine the importance weight of each group of convolution features; According to the importance weight, the plurality of groups of convolution features of different scales are weighted summed to obtain the enhanced feature.

[0083] In one embodiment, the device further comprises: an optimization unit configured to train and optimize the method based on a joint loss function.

[0084] In one embodiment, the joint loss function comprises a reconstruction loss, a perception loss, an edge loss, and a gating constraint loss.

[0085] In one embodiment, the gating constraint loss is determined based on a difference between the change weight map and a ground truth change mask.

[0086] In one embodiment, the target high-resolution reconstructed image is used to restore a historical scene of a complex terrain area with high precision, and the change weight map is used to indicate a building change area between the historical low-resolution image and the current high-resolution image.

[0087] In the embodiments of the present application, a historical low-resolution image and a current high-resolution image are first acquired, a main branch is used to extract features from the historical low-resolution image to construct a structure baseline, and a guide branch is used to extract features from the current high-resolution image and perform semantic alignment to obtain high-frequency priors; then, a difference between the features of the main branch and the guide branch is calculated to obtain feature difference information, and a change weight map is determined based on the information; the features of the main branch and the guide branch are adaptively weighted and fused according to the change weight map, and a gating mechanism is used to explicitly model building changes, so that temporal consistency is maintained in unchanged areas and error texture migration is suppressed in changed areas; then, multi-scale convolution attention processing is performed on the fused features, which uses convolution kernels of multiple scales to capture macro structures and micro details at the same time, thereby enhancing the recovery capability of building edges and road textures under complex terrain; finally, the enhanced features are decoded to generate a target high-resolution reconstructed image. Therefore, through the synergistic effect of the above steps, the present application overcomes the problems of structure blurring and artifacts caused by single information utilization, unstable registration and fusion, and missing change perception in the prior art, thereby achieving temporal consistency and high-fidelity reconstruction of complex terrain building remote sensing images without relying on explicit terrain data.

[0088] It should be noted that the complex terrain building remote sensing image super-resolution reconstruction device provided in the above embodiments is used to execute the complex terrain building remote sensing image super-resolution reconstruction method, and only the division of the above functional modules is used as an example for illustration. In actual applications, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the complex terrain building remote sensing image super-resolution reconstruction device and the complex terrain building remote sensing image super-resolution reconstruction method provided in the above embodiments belong to the same concept, and the implementation process is described in detail in the method embodiments. Therefore, it is not repeated here.

[0089] The above-mentioned sequence numbers of the embodiments of the present application are only for description, and do not represent advantages or disadvantages of the embodiments. In some cases, the actions or steps recited in the claims can be executed in a different order than the order in the embodiments and still achieve the desired result. In addition, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing can be utilized or can be advantageous.

[0090] The embodiment of the present application further provides a storage medium, wherein the storage medium stores a computer program, and the computer program is executed by a processor to implement the method for super-resolution reconstruction of remote sensing images of complex terrain buildings as shown in the above Figures 2-7 The specific implementation process of the method for super-resolution reconstruction of remote sensing images of complex terrain buildings of the embodiment shown in the above can be referred to the specific description of the embodiment shown in the above, and will not be described here. Figures 2-7 The specific implementation process of the method for super-resolution reconstruction of remote sensing images of complex terrain buildings of the embodiment shown in the above can be referred to the specific description of the embodiment shown in the above, and will not be described here.

[0091] Please refer to Figure 9 which shows a structural schematic diagram of an electronic device provided by an example embodiment of the present application. The electronic device in the present application can include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140 and a bus 150. The processor 110, the memory 120, the input device 130 and the output device 140 can be connected through the bus 150.

[0092] The processor 110 can include one or more processing cores. The processor 110 connects various parts in the entire electronic device by using various interfaces and lines, executes various functions of the electronic device and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 120, and calling data stored in the memory 120. Optionally, the processor 110 can be implemented in at least one of a hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), a programmable logic array (PLA). The processor 110 can integrate a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU) and a modem. Among them, the CPU is mainly used to process operating systems, user pages and application programs; the GPU is used to be responsible for rendering and drawing display content; and the modem is used to process wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor 110, but be realized by a separate communication chip.

[0093] The memory 120 can include a Random Access Memory (RAM) and can also include a Read-Only Memory (ROM). Optionally, the memory 120 includes a Non-Transitory Computer-Readable Storage Medium. The memory 120 can be used to store instructions, programs, codes, code sets, or instruction sets. The memory 120 can include a program storage area and a data storage area, where the program storage area can store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playing function, an image playing function, etc.), instructions for implementing the various method embodiments described above, and the like, and the operating system can be an Android system, an IOS system developed by Apple Inc., a system developed based on the Android system or the IOS system, or other systems.

[0094] The memory 120 can be divided into an operating system space and a user space, where the operating system runs in the operating system space, and native and third-party applications run in the user space. In order to ensure that different third-party applications can achieve good running effects, the operating system allocates corresponding system resources to different third-party applications. However, there are also differences in the demand for system resources in different application scenarios in the same third-party application, for example, in the local resource loading scenario, the third-party application has a higher requirement for the disk reading speed; in the animation rendering scenario, the third-party application has a higher requirement for the GPU performance. However, the operating system and the third-party application are independent of each other, and the operating system often cannot timely perceive the current application scenario of the third-party application, resulting in that the operating system cannot perform targeted system resource adaptation according to the specific application scenario of the third-party application.

[0095] In order to enable the operating system to distinguish the specific application scenario of the third-party application, it is necessary to open up the data communication between the third-party application and the operating system, so that the operating system can obtain the current scenario information of the third-party application at any time, and then perform targeted system resource adaptation based on the current scenario.

[0096] The input device 130 is configured to receive input instructions or data, and the input device 130 includes but is not limited to a keyboard, a mouse, a camera, a microphone, or a touch device. The output device 140 is configured to output instructions or data, and the output device 140 includes but is not limited to a display device and a speaker. In one example, the input device 130 and the output device 140 can be combined, and the input device 130 and the output device 140 are a touch display screen.

[0097] The touch display screen can be designed as a full screen, a curved screen, or a special-shaped screen. The touch display screen can also be designed as a combination of a full screen and a curved screen, a combination of a special-shaped screen and a curved screen, and the present application does not limit this.

[0098] In addition, those skilled in the art can understand that the structure of the electronic device shown in the above figure does not constitute a limitation on the electronic device, and the electronic device can include more or fewer components than the figure, or combine certain components, or different component arrangements. For example, the electronic device also includes radio frequency circuit, input unit, sensor, audio circuit, WiFi module, power supply, Bluetooth module and other components, which will not be described here.

[0099] In Figure 9 In the electronic device shown, the processor 110 can be used to call the computer application stored in the memory 120, and specifically perform the following operations: obtain historical low-resolution images and current high-resolution images, the historical low-resolution images and the current high-resolution images being remote sensing images obtained at different time points for the same geographical area; extract features from the historical low-resolution images through a main branch, to obtain main branch features; extract features from the current high-resolution images through a guide branch and perform semantic alignment, to obtain guide branch features; calculate the difference between the main branch features and the guide branch features, to obtain feature difference information, and determine a change weight map based on the feature difference information; adaptively weight and fuse the main branch features and the guide branch features according to the change weight map, to obtain fused features; perform multi-scale convolution attention processing on the fused features, to obtain enhanced features; decode the enhanced features, to obtain a target high-resolution reconstructed image.

[0100] In one embodiment, when the processor 110 performs feature extraction from the historical low-resolution images through a main branch to obtain main branch features, it specifically performs the following operations: perform convolution mapping on the historical low-resolution images, to obtain initial features; perform hierarchical processing on the initial features based on a residual Swin block, to obtain the main branch features.

[0101] In one embodiment, when the processor 110 performs feature extraction from the current high-resolution images through a guide branch and performs semantic alignment to obtain guide branch features, it specifically performs the following operations: The current high-resolution image is processed by using a lightweight Swin Transformer encoder to obtain first multi-scale features. A learnable offset field is predicted based on the first multi-scale features. According to the learnable offset field, the first multi-scale features are adaptively sampled by a deformable convolution to obtain the guided branch features.

[0102] In one embodiment, when the processor 110 performs determining a change weight map based on the feature difference information, the processor 110 specifically performs the following operations: The feature difference information is input into a preset gating network. The change weight map is output by the gating network.

[0103] In one embodiment, when the processor 110 performs adaptively weighting and fusing the main branch features and the guided branch features according to the change weight map to obtain fused features, the processor 110 specifically performs the following operations: The change weight map is point-multiplied with the guided branch features to obtain first weighted features. The change weight map is point-multiplied with the main branch features after being inverted to obtain second weighted features. The first weighted features and the second weighted features are merged to obtain the fused features.

[0104] In one embodiment, when the processor 110 performs multi-scale convolution attention processing on the fused features to obtain enhanced features, the processor 110 specifically performs the following operations: The fused features are respectively convoluted by convolution kernels of multiple scales to obtain multiple groups of convolution features of different scales. The multiple groups of convolution features of different scales are adaptively fused by a channel attention mechanism to obtain the enhanced features.

[0105] In one embodiment, the convolution kernels of multiple scales include 3x3 convolution kernels, 5x5 convolution kernels, and 7x7 convolution kernels.

[0106] In one embodiment, when the processor 110 performs adaptively fusing the multiple groups of convolution features of different scales by a channel attention mechanism to obtain the enhanced features, the processor 110 specifically performs the following operations: Each group of convolution features is globally average-pooled to determine importance weights of the groups of convolution features. The multiple groups of convolution features of different scales are weighted-summed according to the importance weights to obtain the enhanced features.

[0107] In one embodiment, the processor 110 further performs the following operations: The method is trained and optimized based on a joint loss function.

[0108] In one embodiment, the joint loss function includes a reconstruction loss, a perception loss, an edge loss, and a gating constraint loss.

[0109] In one embodiment, the gating constraint loss is determined based on a difference between the change weight map and a true change mask.

[0110] In one embodiment, the target high-resolution reconstructed image is used for high-precision restoration of a historical scene of a complex terrain area, and the change weight map is used to indicate a feature change area between the historical low-resolution image and the current high-resolution image.

[0111] In the embodiments of the present application, a historical low-resolution image and a current high-resolution image are first acquired, feature extraction is performed on the historical low-resolution image by a reconstruction main branch to obtain main branch features to construct a structure baseline, and feature extraction is performed on the current high-resolution image by a guide branch to obtain guide branch features to obtain high-frequency priors; then, difference information between the main branch features and the guide branch features is obtained by calculating the difference between the main branch features and the guide branch features, and a change weight map is determined based on the information; then, the main branch features and the guide branch features are adaptively weighted and fused according to the change weight map, and the process explicitly models feature changes through a gating mechanism, thereby maintaining temporal consistency in unchanged areas and suppressing false texture migration in changed areas; then, multi-scale convolution attention processing is performed on the fused features, and the process simultaneously captures macro structures and micro details using convolution kernels of multiple scales, thereby enhancing the recovery capability of building edges and road textures under complex terrain; finally, the enhanced features are decoded to generate a target high-resolution reconstructed image. Therefore, through the synergistic effect of the above steps, the present application overcomes the problems of structure blurring and artifacts caused by single information utilization, unstable registration and fusion, and missing change perception in the prior art, thereby achieving temporal consistency and high-fidelity reconstruction of complex terrain building remote sensing images without relying on explicit terrain data.

[0112] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, can include the processes of the above-mentioned embodiments. The storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM), a random access memory (RAM), etc.

[0113] The above descriptions are only the preferred embodiment of the application, of course, cannot be used to limit the scope of the application, thus the equivalent variations made by the claims of the application, still belongs to the scope of the application covered.

Claims

1. A method for super-resolution reconstruction of remote sensing images of buildings in complex terrain, characterized in that, include: Acquire historical low-resolution images and current high-resolution images, wherein the historical low-resolution images and the current high-resolution images are remote sensing images of the same geographic area acquired at different times. The main branch features are obtained by reconstructing the main branch and extracting features from the historical low-resolution image. The guide branch features are obtained by extracting features from the current high-resolution image and performing semantic alignment. Calculate the difference between the main branch feature and the guiding branch feature to obtain feature difference information, and determine the change weight map based on the feature difference information; Based on the change weight map, the main branch features and the guiding branch features are adaptively weighted and fused to obtain the fused features; The fused features are subjected to multi-scale convolutional attention processing to obtain enhanced features; The enhanced features are decoded to obtain a high-resolution reconstructed image of the target.

2. The method according to claim 1, characterized in that, The step of extracting features from the historical low-resolution image by reconstructing the main branch to obtain the main branch features includes: The historical low-resolution images are convolutionally mapped to obtain initial features; The initial features are hierarchically processed based on the residual Swin block to obtain the main branch features.

3. The method according to claim 1, characterized in that, The step of extracting features from the current high-resolution image using a guide branch and performing semantic alignment to obtain guide branch features includes: The current high-resolution image is processed using a lightweight Swin Transformer encoder to obtain the first multi-scale features; Predict a learnable offset field based on the first multi-scale feature; Based on the learnable offset field, the first multi-scale features are adaptively sampled using deformable convolution to obtain the guiding branch features.

4. The method according to claim 1, characterized in that, The step of determining the change weight map based on the feature difference information includes: The feature difference information is input into a preset gating network; The changed weight map is output through the gating network.

5. The method according to claim 1, characterized in that, The adaptive weighted fusion of the main branch features and the guiding branch features based on the change weight map to obtain fused features includes: The first weighted feature is obtained by multiplying the change weight map with the guiding support feature; The second weighted feature is obtained by inverting the change weight map and multiplying it by the main branch feature. The first weighted feature and the second weighted feature are combined to obtain the fused feature.

6. The method according to claim 1, characterized in that, The process of performing multi-scale convolutional attention processing on the fused features to obtain enhanced features includes: The fused features are convolved using convolution kernels of various scales to obtain multiple sets of convolutional features of different scales; The enhanced features are obtained by adaptively fusing multiple sets of convolutional features of different scales through a channel attention mechanism.

7. The method according to claim 6, characterized in that, The various convolutional kernels include 3×3, 5×5, and 7×7 kernels.

8. The method according to claim 6, characterized in that, The enhanced features are obtained by adaptively fusing multiple sets of convolutional features of different scales through a channel attention mechanism, including: Global average pooling is performed on each group of convolutional features to determine the importance weights of each group of convolutional features; The enhanced features are obtained by weighted summation of the multiple sets of convolutional features at different scales according to the importance weights.

9. The method according to claim 1, characterized in that, The method further includes: The method is trained and optimized based on the joint loss function.

10. The method according to claim 9, characterized in that, The joint loss function includes reconstruction loss, perception loss, edge loss, and gating constraint loss.

11. The method according to claim 10, characterized in that, The gated constraint loss is determined based on the difference between the change weight map and the actual change mask.

12. The method according to claim 1, characterized in that, The target high-resolution reconstructed image is used to accurately restore historical scenes in complex terrain areas, and the change weight map is used to indicate the area of ​​ground feature change between the historical low-resolution image and the current high-resolution image.

13. A super-resolution reconstruction device for remote sensing images of buildings in complex terrain, characterized in that, include: The image acquisition unit is used to acquire historical low-resolution images and current high-resolution images, wherein the historical low-resolution images and the current high-resolution images are remote sensing images of the same geographical area acquired at different time points. The first feature acquisition unit is used to extract features from the historical low-resolution image by reconstructing the main branch, and obtain the main branch features. The second feature acquisition unit is used to extract features from the current high-resolution image and perform semantic alignment through the guide branch to obtain guide branch features; An information acquisition unit is used to calculate the difference between the main branch feature and the guiding branch feature, obtain feature difference information, and determine a change weight map based on the feature difference information. The third feature acquisition unit is used to adaptively weight and fuse the main branch features and the guiding branch features according to the change weight map to obtain the fused features; The fourth feature acquisition unit is used to perform multi-scale convolutional attention processing on the fused features to obtain enhanced features; The feature decoding unit is used to decode the enhanced features to obtain a high-resolution reconstructed image of the target.

14. A computer storage medium, characterized in that, The computer storage medium stores a plurality of instructions adapted for loading by a processor and executing the steps of the method as described in any one of claims 1 to 12.

Citation Information

Patent Citations

  • Remote sensing image space-time fusion method and device based on mixed attention mechanism

    CN120411699A

  • Lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel cavity coordinate attention

    CN120876231A

  • DEM intelligent super-resolution method based on high spatial resolution remote sensing data

    CN120932067A

  • Generative adversarial network super-resolution image reconstruction method and system based on Swin Transform generator and double discriminators

    CN121190308A

  • Generation method, system and apparatus capable of visual resolution enhancement, and storage medium

    WO2022242029A1