Super-resolution reconstruction method and device based on multi-frame remote sensing image and medium

CN122434737BActive Publication Date: 2026-08-21HANGZHOU INST FOR ADVANCED STUDY UCAS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610883591.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-08-21
Estimated Expiration
2046-06-18

AI Technical Summary

Technical Problem

然而,当前影像帧原始分辨率本就较低,当前影像帧所包含的有效信息极为有限,只能通过猜测来确定当前影像帧中缺失的信息,这种方式导致当前影像帧的分辨率增强效果不佳,影像的增强质量难以达到要求

Benefits of technology

[0015]根据本发明提供的一种基于多帧遥感影像的超分辨率重建方法、装置及介质,与目前仅依赖当前影像帧自身的低分辨率像素信息进行上采样重建为高分辨率影像帧的方式相比,本发明通过参考同一连续遥感影像帧序列中已经增强分辨率后的前序影像帧的历史传播特征来辅助当前待处理影像帧的超分辨率重建,能够充分挖掘遥感影像帧间的时空冗余信息,利用时空冗余信息能够充分确定当前待处理影像帧中的纹理信息和具体结构,能够为当前待处理影像帧中的弱纹理提供参考信息,从而能够提高影像的分辨率增强效果;通过利用遥感影像训练的编码器来提取空间特征,能够使编码器直接学习遥感影像专用的时空先验表征,从而提高特征提取精度;通过高频增强模块对空间特征进行增强处理,使得边缘、纹理和细小结构在跨帧传播前就被补足,避免后续影像对齐聚合过程中高频信息逐步衰减的问题;通过全局分支捕获长距离运动趋势,局部分支建模细粒度形变,二者融合后得到更精确的偏移预测,克服了现有单一卷积堆叠结构对遥感影像复杂运动建模不足的缺陷,从而能够提高特征的捕捉全面性;对全局运动趋势特征和局部细粒度形变特征进行二阶空间变换,使历史传播特征与增强特征实现像素级精确配准,避免错位融合导致的结构模糊和伪影,从而提高了遥感影像的分辨率增强效果,保障增强后的遥感影像的影像质量。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122434737B_ABST
    Figure CN122434737B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multi-frame remote sensing image super-resolution method, device and medium, it is related to remote sensing image processing technical field, can improve remote sensing image resolution enhancement effect.It includes: obtaining the historical propagation characteristics of image frame to be processed and previous image frame;Obtain preset image super-resolution model;The image frame to be processed and historical propagation characteristics are input into model, spatial feature extraction is carried out to image frame to be processed by encoder, spatial feature is enhanced into enhanced feature by high-frequency enhancement module, global trend extraction and local fine-grained extraction are carried out to enhanced feature and historical propagation characteristics by global branch and local branch, historical propagation characteristics are processed into the alignment historical propagation characteristics of alignment with enhanced feature by alignment module according to global trend feature and local fine-grained feature, high-resolution image frame is reconstructed to enhanced feature according to alignment historical propagation characteristics by image reconstruction module.The application is suitable for remote sensing image processing scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of remote sensing image processing technology, and in particular to a method, apparatus and medium for super-resolution reconstruction based on multi-frame remote sensing images. Background Technology

[0002] Remote sensing imagery is acquired through sensors mounted on platforms such as satellites and drones. Limited by factors such as sensor manufacturing processes, payload capacity, data transmission bandwidth, and storage capacity, the raw remote sensing images often suffer from low spatial resolution. Low-resolution remote sensing images exhibit blurred target details, making it difficult to effectively identify and accurately locate ground objects (such as vehicles, people, and buildings), severely restricting the application effectiveness of remote sensing imagery in downstream tasks such as target detection, change monitoring, and situational awareness. Therefore, resolution enhancement of remote sensing imagery is a key technical means to improve its usability.

[0003] Currently, high-resolution image frames are typically reconstructed by upsampling based solely on the low-resolution pixel information of the current image frame itself. However, the original resolution of the current image frame is already low, and the effective information contained in the current image frame is extremely limited. The missing information in the current image frame can only be determined by guesswork, resulting in poor resolution enhancement and failing to meet the required image enhancement quality. Summary of the Invention

[0004] This invention provides a method, apparatus, and medium for super-resolution reconstruction based on multi-frame remote sensing images, which mainly improves the resolution enhancement effect of remote sensing images and ensures the image quality of the enhanced remote sensing images.

[0005] According to a first aspect of the present invention, a super-resolution reconstruction method based on multi-frame remote sensing images is provided, comprising: The historical propagation features of the current image frame to be processed and the previous image frames in the remote sensing image frame sequence are obtained, wherein the historical propagation features are obtained by extracting features from at least two previous image frames and fusing the extracted features. A preset image super-resolution model is obtained, wherein the preset image super-resolution model includes an encoder for feature extraction, a high-frequency enhancement module for feature enhancement, a global-local offset estimator for cross-frame alignment, and an image reconstruction module for image reconstruction. The global-local offset estimator includes a global branch, a local branch, and a second-order deformable alignment module. The preset image super-resolution model is pre-trained based on a sample dataset. The current image frame to be processed and the historical propagation features are input into the preset image super-resolution model. The encoder extracts spatial features from the current image frame to be processed. The high-frequency enhancement module enhances the spatial features into enhanced features with image frame edge details and image frame texture details. The global-local offset estimator extracts global motion trends and local fine-grained deformations between frames from the concatenated results of the enhanced features with image frame edge details and image frame texture details and the historical propagation features. The global motion trend features corresponding to the global branch and the local fine-grained deformation features corresponding to the local branch are obtained. The second-order deformable alignment module performs spatial transformation on the historical propagation features based on the global motion trend features and the local fine-grained deformation features to obtain aligned historical propagation features that are aligned across frames with the enhanced features. The image reconstruction module reconstructs the high-resolution remote sensing image frame corresponding to the current image frame to be processed based on the aligned historical propagation features.

[0006] Optionally, before obtaining the preset image super-resolution model, the method further includes: Obtain the shared encoder; Acquire sample low-resolution remote sensing images, wherein the sample low-resolution remote sensing images include a sequence of low-resolution remote sensing image frames with standard high-resolution remote sensing image frame annotation information, wherein the standard high-resolution remote sensing image frame annotation information is obtained by annotating the standard high-resolution remote sensing image frame corresponding to each low-resolution remote sensing image frame in the sequence of low-resolution remote sensing image frames. The shared encoder is used to extract the frame-level features of each low-resolution remote sensing image frame in the sample low-resolution remote sensing image. The central low-resolution remote sensing image frame and the neighboring low-resolution remote sensing image frames adjacent to the central low-resolution remote sensing image frame are determined in the sample low-resolution remote sensing image. The center frame-level features of the central low-resolution remote sensing image frame are subjected to block-level random masking, and the neighboring frame-level features of the neighboring low-resolution remote sensing image frames are subjected to auxiliary masking. The block-level random masked center frame-level features and the auxiliary masked neighboring frame-level features are fused to obtain fused features. The fused features are decoded into the predicted center high-resolution remote sensing image frame corresponding to the central low-resolution remote sensing image frame using a preset decoder. A loss function is determined based on the difference between the standard high-resolution remote sensing image frame annotation information corresponding to the central low-resolution remote sensing image frame and the predicted central high-resolution remote sensing image frame. The shared encoder is iteratively trained based on the loss function, and the encoder is determined based on the iterative training results.

[0007] Optionally, Block-level random masking is performed on the center frame-level features of the central low-resolution remote sensing image frame, including: The central frame-level feature is divided into multiple non-overlapping central feature blocks. Multiple central feature blocks to be masked are randomly selected from the multiple central feature blocks at a first preset ratio. Each of the central feature blocks to be masked in the central frame-level feature is masked to obtain a binary central mask map. The central frame-level feature after block-level random masking is determined based on the binary central mask map and the central frame-level feature. Auxiliary masking is performed on the neighboring frame-level features of the adjacent low-resolution remote sensing image frames, including: The neighboring frame-level features are divided into multiple non-overlapping neighboring feature blocks. Multiple neighboring feature blocks to be masked are randomly selected from the multiple neighboring feature blocks at a second preset ratio. Each of the neighboring feature blocks to be masked in the neighboring frame-level features is masked to obtain a binary neighbor mask map. The neighboring frame-level features after auxiliary masking are determined based on the binary neighbor mask map and the neighboring frame-level features. The second preset ratio is less than the first preset ratio.

[0008] Optionally, The high-frequency enhancement module includes, in sequence, a Gaussian filter layer, a high-frequency residual signal construction layer, a high-frequency feature mapping layer, and a gating layer; The spatial features are enhanced into enhanced features with image frame edge details and image frame texture details through the high-frequency enhancement module, including: The low-frequency component is separated from the current image frame to be processed by the Gaussian filtering layer, and the low-frequency component is constructed into a high-frequency residual signal by the high-frequency residual signal construction layer. The high-frequency residual signal is converted into high-frequency features with image frame edge details and image frame texture details that are consistent with the spatial feature dimension through the high-frequency feature mapping layer; The high-frequency features containing image frame edge details and image frame texture details are injected into the spatial features through the gated layer to obtain enhanced features containing image frame edge details and image frame texture details.

[0009] Optionally, the high-frequency features containing image frame edge details and image frame texture details are injected into the spatial features through the gate layer to obtain enhanced features containing image frame edge details and image frame texture details, including: The spatial features are compressed by encoding channel dimension to obtain a single-channel spatial attention map. The single-channel spatial attention map is then subjected to adaptive histogram equalization to obtain an enhanced single-channel spatial attention map. A gated weight map is generated based on the enhanced single-channel spatial attention map, wherein each pixel value in the gated weight map corresponds to the high-frequency information injection intensity coefficient at the corresponding position of the spatial feature. Based on the high-frequency information injection intensity coefficient, the high-frequency features with image frame edge details and image frame texture details are injected into the spatial features to obtain enhanced features with image frame edge details and image frame texture details.

[0010] Optionally, The global-local offset estimator also includes a spatial pyramid network and an offset prediction branch; The second-order deformable alignment module performs spatial transformation on the historical propagation features based on the global motion trend features and the local fine-grained deformation features to obtain aligned historical propagation features that are aligned across frames with the enhanced features, including: The spatial pyramid network determines the first-order optical flow information between the enhancement feature with image frame edge details and image frame texture details and the historical propagation feature, and performs bilinear interpolation on the historical propagation feature based on the first-order optical flow information to obtain an initial aligned historical propagation feature that is aligned across frames with the enhancement feature with image frame edge details and image frame texture details. The offset prediction branch predicts the cross-frame alignment offset based on the global motion trend features and the local fine-grained deformation features; The second-order deformable alignment module performs a secondary correction on the initial alignment history propagation feature based on the cross-frame alignment offset to obtain the alignment history propagation feature that is aligned across frames with the enhanced feature containing image frame edge details and image frame texture details.

[0011] Optionally, the global-local offset estimator performs inter-frame global motion trend extraction and local fine-grained deformation extraction on the concatenated results of the enhanced features and the historical propagation features, which contain image frame edge details and image frame texture details, including: The concatenated results are subjected to layer normalization through the global branch, and channel attention map and spatial attention map are determined based on the concatenated results after layer normalization. The global motion trend features between the current image frame to be processed and the preceding image frame are determined based on the channel attention map and spatial attention map. Different deformation feature maps are obtained by performing dilated convolution operations with different dilation rates on the cascaded results through the local branches. Each deformation feature map is then concatenated along the channel dimension to obtain a stitched feature map. Channel attention weighting is applied to the stitched feature map to obtain a channel weight vector, where each element in the channel weight vector corresponds to the importance coefficient of the corresponding channel in the stitched feature map. The channel elements of the stitched feature map are weighted and summed using the importance coefficients. The stitched feature map after weighted summation of the channel elements is then convolved and dimensionality reduced to obtain the local fine-grained deformation features between the current image frame to be processed and the previous image frame. The local fine-grained deformation features include at least one of local displacement features, edge deformation features, and small target motion features.

[0012] According to a second aspect of the present invention, a super-resolution reconstruction apparatus based on multi-frame remote sensing images is provided, comprising: The image acquisition unit is used to acquire the historical propagation features of the current image frame to be processed and the previous image frames in the remote sensing image frame sequence, wherein the historical propagation features are obtained by extracting features from at least two previous image frames and fusing the extracted features. The model acquisition unit is used to acquire a preset image super-resolution model, wherein the preset image super-resolution model includes an encoder for feature extraction, a high-frequency enhancement module for feature enhancement, a global-local offset estimator for cross-frame alignment, and an image reconstruction module for image reconstruction. The global-local offset estimator includes a global branch, a local branch, and a second-order deformable alignment module. The preset image super-resolution model is pre-trained based on a sample dataset. The resolution enhancement unit is used to input the current image frame to be processed and the historical propagation features into the preset image super-resolution model. The encoder extracts spatial features from the current image frame to be processed. The high-frequency enhancement module enhances the spatial features into enhanced features with image frame edge details and image frame texture details. The global-local offset estimator extracts global motion trends and local fine-grained deformations between frames from the concatenated results of the enhanced features with image frame edge details and image frame texture details and the historical propagation features. This yields global motion trend features corresponding to the global branch and local fine-grained deformation features corresponding to the local branch. The second-order deformable alignment module performs spatial transformation on the historical propagation features based on the global motion trend features and local fine-grained deformation features to obtain aligned historical propagation features that are aligned across frames with the enhanced features. The image reconstruction module reconstructs the enhanced features with image frame edge details and image frame texture details based on the aligned historical propagation features to obtain the high-resolution remote sensing image frame corresponding to the current image frame to be processed.

[0013] According to a third aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the above-described super-resolution reconstruction method based on multi-frame remote sensing images.

[0014] According to a fourth aspect of the present invention, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described super-resolution reconstruction method based on multi-frame remote sensing images.

[0015] The present invention provides a method, apparatus, and medium for super-resolution reconstruction based on multi-frame remote sensing images. Compared with current methods that rely solely on the low-resolution pixel information of the current image frame for upsampling and reconstruction into a high-resolution image frame, the present invention assists in the super-resolution reconstruction of the current image frame by referencing the historical propagation characteristics of the preceding image frames in the same continuous remote sensing image frame sequence that have already been enhanced in resolution. This fully exploits the spatiotemporal redundancy information between remote sensing image frames. Utilizing this spatiotemporal redundancy information, the texture information and specific structure in the current image frame can be fully determined, providing reference information for weak textures in the current image frame, thereby improving the resolution enhancement effect of the image. By using an encoder trained on remote sensing images to extract spatial features, the encoder can directly learn the spatiotemporal features specific to remote sensing images. The system employs several techniques to enhance features, including: First, it improves feature extraction accuracy by employing a high-frequency enhancement module to augment spatial features, ensuring edges, textures, and fine structures are filled in before cross-frame propagation, thus preventing the gradual attenuation of high-frequency information during subsequent image alignment and aggregation. Second, it captures long-distance motion trends through global branches and models fine-grained deformations through local branches, fusing the two to obtain more accurate offset predictions. This overcomes the shortcomings of existing single convolutional stacking structures in modeling complex motions in remote sensing images, thereby improving the comprehensiveness of feature capture. Third, it performs second-order spatial transformations on global motion trend features and local fine-grained deformation features, enabling pixel-level precise registration between historical propagation features and enhanced features. This avoids structural blurring and artifacts caused by misaligned fusion, thereby improving the resolution enhancement effect of remote sensing images and ensuring the image quality of the enhanced remote sensing images. Attached Figure Description

[0016] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings: Figure 1 A flowchart of a super-resolution reconstruction method based on multi-frame remote sensing images provided by an embodiment of the present invention is shown. Figure 2A flowchart of another super-resolution reconstruction method based on multi-frame remote sensing images provided by an embodiment of the present invention is shown; Figure 3 A schematic diagram of the structure of a super-resolution reconstruction device based on multi-frame remote sensing images provided in an embodiment of the present invention is shown. Figure 4 A schematic diagram of another super-resolution reconstruction device based on multi-frame remote sensing images provided by an embodiment of the present invention is shown. Figure 5 A schematic diagram of the physical structure of a computer device provided in an embodiment of the present invention is shown. Detailed Implementation

[0017] The present invention will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the present application can be combined with each other.

[0018] Currently, the method relies solely on the low-resolution pixel information of the current image frame itself for upsampling and reconstruction into a high-resolution image frame. Since the effective information contained in the current image frame is extremely limited, the missing information can only be determined through guesswork. This method results in poor resolution enhancement of the current image frame, and the image enhancement quality fails to meet requirements.

[0019] To address the aforementioned problems, embodiments of the present invention provide a super-resolution reconstruction method based on multi-frame remote sensing images, such as... Figure 1 As shown, the method includes: 101. Obtain the historical propagation features of the current image frame to be processed and the preceding image frames in the remote sensing image frame sequence. The historical propagation features are obtained by extracting features from at least two preceding image frames and fusing the extracted features.

[0020] Among them, remote sensing image frame sequences can be image frame sequences in any scene, such as satellite remote sensing image frame sequences and aerial remote sensing image frame sequences; satellite remote sensing image frame sequences include Earth observation image sequences taken by optical satellites, covering continuous low-resolution image frame sequences of scenes such as urban area monitoring, agricultural planting area change detection, ocean surface dynamic observation, forest cover change tracking, and glacier movement monitoring; aerial remote sensing image frame sequences include aerial reconnaissance images and disaster emergency monitoring images acquired by UAVs or manned aircraft equipped with optical / infrared cameras.

[0021] In this embodiment of the invention, low-resolution image frames at the current time t are extracted sequentially from the remote sensing image frame sequence in chronological order as the current image frames to be processed. For the current time t, super-resolution reconstruction and feature aggregation have been completed for image frames from time 1 to t-1, resulting in historical propagation features of the preceding image frames. These historical propagation features can be aggregated features of image frames from time 1 to t-1 that have undergone super-resolution reconstruction, or aggregated features of two or more image frames from time 1 to t-1 that have undergone super-resolution reconstruction. The super-resolution method of each image frame from time 1 to t-1 can be the same as or different from the super-resolution method of the current image frame to be processed in this embodiment of the invention. For example, taking any target image frame from time 1 to time t-1 as an example, the pre-propagation features of the target image frame and the preceding image frame that has undergone super-resolution processing are jointly input into a preset image super-resolution model. The encoder extracts the target spatial features of the target image frame. The high-frequency enhancement module enhances the target spatial features into target enhancement features with image frame edge details and image frame texture details. The global-local offset estimator extracts the global motion trend and local fine-grained deformation between frames from the cascaded results of the target enhancement features with image frame edge details and image frame texture details and the pre-propagation features, obtaining global motion trend features and local fine-grained deformation features. The second-order deformable alignment module performs spatial transformation on the pre-propagation features based on the global motion trend features and the local fine-grained deformation features to obtain aligned pre-propagation features that are aligned across frames with the target enhancement features. The image reconstruction module reconstructs the target enhancement features with image frame edge details and image frame texture details based on the aligned pre-propagation features to obtain the high-resolution image frame corresponding to the target image frame. Then, features from the high-resolution image frame are extracted using a CNN feature extraction model as the features of the target image frame. The features of multiple consecutive super-resolution processed preceding image frames are then fused to obtain historical propagation features. This process is repeated to determine the historical propagation features between image frames from time t-1 to previous times. It should be noted that if the current image frame to be processed is the second frame in the remote sensing image frame sequence (i.e., the image frame at time 2), its corresponding historical propagation features are only the features of the first frame that has undergone super-resolution processing (the image frame at time 1). If the current image frame to be processed is the first frame in the remote sensing image frame sequence (i.e., the image frame at time 1), its corresponding historical propagation features are set to 0.

[0022] This invention, through obtaining the historical propagation features of preceding image frames that have undergone super-resolution processing, enables the reconstruction of the current frame to leverage the high-resolution structure and texture information recovered from the preceding frames. This fully exploits the spatiotemporal redundancy information between remote sensing image frames, breaks through the bottleneck of single-frame information, and thus improves the resolution enhancement effect of the image.

[0023] 102. Obtain a preset image super-resolution model, wherein the preset image super-resolution model includes an encoder for feature extraction, a high-frequency enhancement module for feature enhancement, a global-local offset estimator for cross-frame alignment, and an image reconstruction module for image reconstruction. The global-local offset estimator includes a global branch, a local branch, and a second-order deformable alignment module. The preset image super-resolution model is pre-trained based on a sample dataset.

[0024] In this embodiment of the invention, an encoder is pre-trained and constructed based on a remote sensing image dataset. By using the encoder trained on the remote sensing image dataset to extract spatial features, the encoder can directly learn the spatiotemporal prior representations specific to remote sensing images, thereby improving feature extraction accuracy. Then, the encoder parameters are fixed, and a preset image super-resolution model with fixed encoder parameters is trained and constructed using a sample dataset. Specifically, the process of training and constructing a preset image super-resolution model with fixed encoder parameters using a sample dataset includes: constructing a preset initial image super-resolution model; obtaining a sample dataset, wherein the sample dataset includes low-resolution remote sensing images with high-resolution remote sensing image frame annotation information; dividing the sample dataset into a training set and a test set; training the preset initial image super-resolution model using the training set; and testing the trained preset initial image super-resolution model using the test set. Finally, the trained preset initial image super-resolution model that meets the test conditions is taken as the preset image super-resolution model.

[0025] The model structure of the preset initial image super-resolution model is the same as that of the preset image super-resolution model. Specifically, during model training, the preset initial image super-resolution model is first constructed, and then a sample dataset is obtained. Ensure the dataset contains all necessary files. Convert the data to a format that the preset initial image super-resolution model can understand, and finally train and test the model. Specifically, the dataset can be divided first: using randomness or a specific strategy (such as stratified sampling), the sample dataset is divided into a training set and a test set. The model is then trained using the training set, and tested using the test set to evaluate its performance on unseen data. Precision and recall metrics on the test set are calculated and recorded. If the model performance does not meet the requirements, it can return to the training phase for further iterations or adjustments. This process yields a preset image super-resolution model that meets the requirements. Furthermore, after training and constructing the preset image super-resolution model, it is stored. When super-resolution processing of the current image frame needs to be performed, the preset image super-resolution model can be directly retrieved from its storage address for application.

[0026] 103. Input the current image frame to be processed and the historical propagation features into the preset image super-resolution model. Extract spatial features from the current image frame through the encoder. Enhance the spatial features into enhanced features with image frame edge details and image frame texture details through the high-frequency enhancement module. Extract global motion trend and local fine-grained deformation between frames through the cascaded results of the enhanced features with image frame edge details and image frame texture details and the historical propagation features through the global-local offset estimator. Obtain global motion trend features corresponding to the global branch and local fine-grained deformation features corresponding to the local branch. Perform spatial transformation on the historical propagation features based on the global motion trend features and local fine-grained deformation features through the second-order deformable alignment module to obtain aligned historical propagation features that are aligned with the enhanced features across frames. Reconstruct the high-resolution remote sensing image frame corresponding to the current image frame to be processed based on the aligned historical propagation features through the image reconstruction module.

[0027] In this embodiment of the invention, the current image frame to be processed is input into the encoder of a preset image super-resolution model to extract the spatial features of the current frame. The spatial features are then input into a high-frequency enhancement module, where low-frequency components are first separated by Gaussian filtering and a high-frequency residual signal is constructed. Edge and texture details are then injected into the spatial features through convolutional mapping and gated modulation to obtain enhanced features that carry the edge and texture details of the image frame. The enhanced features are concatenated with historical propagation features and input into a global-local offset estimator. The global branch of the global-local offset estimator extracts global motion trend features through channel attention and spatial attention, while the local branch extracts local fine-grained deformation features through multi-scale dilated convolution. After the global motion trend features and local fine-grained deformation features are fused, a second-order deformable alignment module performs spatial transformation on the historical propagation features based on the fused features to obtain aligned historical propagation features that are aligned across frames with the enhanced features. The image reconstruction module performs fusion reconstruction on the enhanced features based on the aligned historical propagation features, using an upsampling and convolutional reconstruction structure to achieve ×n super-resolution restoration, and outputs the high-resolution remote sensing image frame corresponding to the current image frame to be processed. The embodiments of this invention employ a sequential processing approach of encoder extraction, high-frequency enhancement, global-local joint alignment, and image reconstruction. This approach enables the current frame to simultaneously possess remote sensing-specific spatiotemporal priors, explicitly compensated edge texture details, and accurate cross-frame alignment information after joint correction of global motion trends and local fine-grained deformations during reconstruction. This overcomes the shortcomings of general image super-resolution models in remote sensing scenarios, such as inaccurate cross-frame alignment, insufficient detail recovery, and overly smooth results, thereby improving the super-resolution effect of the image.

[0028] The super-resolution reconstruction method based on multi-frame remote sensing imagery provided by this invention, compared with the current method that relies solely on the low-resolution pixel information of the current image frame for upsampling and reconstruction into a high-resolution image frame, assists in the super-resolution reconstruction of the current image frame by referencing the historical propagation characteristics of the preceding image frames in the same continuous remote sensing image frame sequence that have already been enhanced in resolution. This fully exploits the spatiotemporal redundancy information between remote sensing image frames, and utilizes this redundancy to fully determine the texture information and specific structure in the current image frame, providing reference information for weak textures in the current image frame, thereby improving the resolution enhancement effect of the image. By using an encoder trained on remote sensing images to extract spatial features, the encoder can directly learn the spatiotemporal prior representations specific to remote sensing images. This improves feature extraction accuracy. A high-frequency enhancement module enhances spatial features, ensuring edges, textures, and fine structures are supplemented before cross-frame propagation, preventing the gradual attenuation of high-frequency information during subsequent image alignment and aggregation. Global branches capture long-distance motion trends, while local branches model fine-grained deformations; their fusion yields more accurate offset predictions, overcoming the shortcomings of existing single convolutional stacking structures in modeling complex motions in remote sensing images, thus improving the comprehensiveness of feature capture. A second-order spatial transformation is performed on global motion trend features and local fine-grained deformation features, enabling pixel-level precise registration of historical propagation features and enhanced features, avoiding structural blurring and artifacts caused by misaligned fusion, thereby improving the resolution enhancement effect of remote sensing images and ensuring the image quality of the enhanced remote sensing images.

[0029] Furthermore, to better illustrate the above-described process of super-resolution reconstruction of remote sensing images, as a refinement and extension of the above embodiments, this invention provides another super-resolution reconstruction method based on multi-frame remote sensing images, such as... Figure 2 As shown, the method includes: 201. Obtain the historical propagation features of the current image frame to be processed and the preceding image frames in the remote sensing image frame sequence. The historical propagation features are obtained by extracting features from at least two preceding image frames and fusing the extracted features.

[0030] Specifically, the preceding image frames that have undergone super-resolution processing are obtained from the image database that has already undergone super-resolution processing, and the historical propagation features of the preceding image frames are extracted using feature extraction models such as CNN.

[0031] 202. Obtain a preset image super-resolution model, wherein the preset image super-resolution model includes an encoder for feature extraction, a high-frequency enhancement module for feature enhancement, a global-local offset estimator for cross-frame alignment, and an image reconstruction module for image reconstruction. The high-frequency enhancement module includes a Gaussian filter layer, a high-frequency residual signal construction layer, a high-frequency feature mapping layer, and a gating layer. The global-local offset estimator includes a global branch, a local branch, and a second-order deformable alignment module. The preset image super-resolution model is pre-trained based on a sample dataset.

[0032] In this embodiment of the invention, to improve the feature extraction accuracy of the encoder for remote sensing images, it is first necessary to train and construct the encoder using a remote sensing image dataset. Based on this, the method includes: acquiring a shared encoder; acquiring sample low-resolution remote sensing images, wherein the sample low-resolution remote sensing images include a sequence of low-resolution remote sensing image frames with standard high-resolution remote sensing image frame annotation information, wherein the standard high-resolution remote sensing image frame annotation information is obtained by annotating the standard high-resolution remote sensing image frames corresponding to each low-resolution remote sensing image frame in the low-resolution remote sensing image frame sequence; using the shared encoder to extract frame-level features of each low-resolution remote sensing image frame in the sample low-resolution remote sensing images, and determining the central low-resolution remote sensing image frame and the frame-level features corresponding to the central low-resolution remote sensing image frame in the sample low-resolution remote sensing images. For a central low-resolution remote sensing image frame, the center frame-level features of the central low-resolution remote sensing image frame are subjected to block-level random masking, and the neighboring frame-level features of the neighboring low-resolution remote sensing image frames are subjected to auxiliary masking. The center frame-level features after block-level random masking and the neighboring frame-level features after auxiliary masking are fused to obtain fused features. The fused features are decoded into the predicted central high-resolution remote sensing image frame corresponding to the central low-resolution remote sensing image frame using a preset decoder. A loss function is determined based on the difference between the standard high-resolution remote sensing image frame annotation information corresponding to the central low-resolution remote sensing image frame and the predicted central high-resolution remote sensing image frame. The shared encoder is iteratively trained based on the loss function, and the encoder is determined based on the iterative training results. The method for performing block-level random masking on the center frame-level features and auxiliary masking on the neighboring frame-level features includes: dividing the center frame-level features into multiple non-overlapping center feature blocks; randomly selecting multiple center feature blocks to be masked from the multiple center feature blocks at a first preset ratio; performing masking processing on each of the center feature blocks to be masked in the center frame-level features to obtain a binary center mask image; and determining the center frame-level features after block-level random masking based on the binary center mask image and the center frame-level features; dividing the neighboring frame-level features into multiple non-overlapping neighboring feature blocks; randomly selecting multiple neighboring feature blocks to be masked from the multiple neighboring feature blocks at a second preset ratio; performing masking processing on each of the neighboring feature blocks to be masked in the neighboring frame-level features to obtain a binary neighboring mask image; and determining the neighboring frame-level features after auxiliary masking based on the binary neighboring mask image and the neighboring frame-level features, wherein the second preset ratio is less than the first preset ratio.

[0033] Specifically, the shared encoder, i.e., the universal encoder, is first obtained. During the training of the shared encoder using sample low-resolution remote sensing images, these sample low-resolution images consist of multiple consecutive frames. For example, given a sequence of five low-resolution images centered at time t, the sequence of five low-resolution images... As shown below: , in, , , , , These are consecutive image frames at times t-2, t-1, t, t+1, and t+2. Then, a shared encoder is used. Extract features from each image frame separately, i.e., frame-level features. ,in, For a moment From the image frames, a frame-level feature set is obtained. As shown below:

[0034] in, , , , , These are the frame-level features corresponding to the image frames at times t-2, t-1, t, t+1, and t+2, respectively. Further, the center frame-level features... Block-level random masking is used: center frame-level features are... Divide into non-overlapping blocks of size P×P (central feature blocks), according to a first preset ratio. Randomly select blocks to construct a binary center mask diagram Center frame level features after masking Represented as: , This indicates element-wise multiplication. A lower-ratio (second preset ratio) auxiliary mask is used for features at the neighboring frame level: for each neighboring frame... Neighboring frame-level features According to the second preset ratio Construct a mask diagram (binary nearest neighbor mask diagram) Neighboring frame-level features after masking for: This is to avoid the model relying solely on complete neighboring frame information for copy-based recovery. Then, a center-guided gated temporal fusion unit, such as a center-guided gated fusion unit, is used to fuse the masked center frame-level features with neighboring frame-level features. Specifically, if... and These represent the center frame-level features after masking. Features of neighboring frames after masking The projection function, This represents the gated prediction function. It considers the neighboring frame-level features after each mask. First, calculate the gating graph. :

[0035] Where [·,·] represents channel cascading, and σ(·) is the Sigmoid function. Fusion Features As shown below:

[0036] Furthermore, the fusion features Input a lightweight decoder D(·) to reconstruct high-resolution remote sensing image frames with predicted centers. ,in, During encoder training, high-resolution remote sensing image frames with reconstructed prediction centers are determined. Annotation information of standard high-resolution remote sensing images The difference between them, and the loss function is determined based on that difference. As shown below:

[0037] Where ε is the numerical stability constant (e.g., ε = 10). - ³). This pre-training process enables the encoder to learn a spatiotemporal representation more suitable for remote sensing imagery. After pre-training, the encoder is retained as the feature extraction backbone of the downstream pre-defined imagery super-resolution model, while the decoder is discarded. During the training of the downstream pre-defined imagery super-resolution model, the encoder remains fixed and does not participate in parameter updates, ensuring that the pre-trained data is retained for the final task.

[0038] Furthermore, a preset initial image super-resolution model is obtained. This preset initial image super-resolution model includes the pre-trained encoder with fixed parameters, as well as a high-frequency enhancement module, a global-local offset estimator, and an image reconstruction module. During the training of this preset initial image super-resolution model, the parameters of the encoder remain unchanged, and only the parameters of the high-frequency enhancement module, the global-local offset estimator, and the image reconstruction module are iteratively updated, thereby obtaining a preset image super-resolution model with super-resolution accuracy that meets the requirements.

[0039] 203. The current image frame to be processed and historical propagation features are input into a preset image super-resolution model. An encoder extracts spatial features from the current image frame. A Gaussian filtering layer separates low-frequency components from the current image frame. A high-frequency residual signal construction layer constructs high-frequency residual signals from the low-frequency components. A high-frequency feature mapping layer converts the high-frequency residual signals into high-frequency features with image frame edge and texture details consistent with the spatial feature dimension. A gating layer injects these high-frequency features with image frame edge and texture details into the spatial features to obtain enhanced features with image frame edge and texture details. This is then processed through a global-local mapping layer. The offset estimator extracts global motion trends and local fine-grained deformations between frames from the concatenated results of enhancement features and historical propagation features containing image frame edge details and image frame texture details. This yields global motion trend features corresponding to global branches and local fine-grained deformation features corresponding to local branches. A second-order deformable alignment module performs spatial transformation on the historical propagation features based on the global motion trend features and local fine-grained deformation features to obtain aligned historical propagation features that are aligned with the enhancement features across frames. Finally, an image reconstruction module reconstructs the enhancement features containing image frame edge details and image frame texture details based on the aligned historical propagation features to obtain the high-resolution remote sensing image frame corresponding to the current image frame to be processed.

[0040] In this embodiment of the invention, the current image frame to be processed and historical propagation features are jointly input into a preset image super-resolution model. An encoder is used to extract the backbone features, i.e., spatial features, from the current image frame to be processed. Simultaneously, the current image frame to be processed... and its corresponding spatial features The image is input to the high-frequency enhancement module. In the high-frequency enhancement module, the current image frame to be processed is first... The input is fed into a Gaussian filter layer, where the low-frequency components are separated by Gaussian filtering. Then the low-frequency components The input is fed into the high-frequency residual signal construction layer to construct a high-frequency residual signal. Among them, high-frequency residual signals The construction process is as follows: Then, the high-frequency residual signal The input is fed into a high-frequency feature mapping layer. In this layer, a 1×1 convolution is first used to linearly project the high-frequency residual signal along its channel dimension to obtain channel-aligned intermediate features. This ensures that the number of channels in the high-frequency residual signal matches the number of channels in the spatial features, thus linearly transforming the high-frequency residual signal from its original feature space to the same feature space. Then, nonlinear activation processing is applied to the intermediate features, such as using the ReLU (Rectified Linear Unit) activation function to retain positive value regions and suppress negative value regions. Positive value regions correspond to brightness jump regions (i.e., edges) in the image frame, while negative value regions correspond to brightness inverse jump regions. This embodiment of the invention uses the ReLU activation function to selectively enhance edge and texture details in the positive gradient direction of the image frame, while avoiding interference from negative high-frequency responses in subsequent feature injection processes. Furthermore, the features activated by ReLU are locally refined in the spatial dimension, for example, through a 3×3 convolutional layer, to ensure that the spatial resolution of the output features remains consistent with the input. For edge details: the 3×3 convolutional layer can capture the directional information and gradient intensity of edges in the local neighborhood, aggregating scattered edge pixels into continuous edge contours; for texture details: the 3×3 convolutional kernel can capture the repetition patterns and directions of textures in the local neighborhood, aggregating scattered texture responses into complete texture patterns. Then, the features output by the 3×3 convolutional kernel are sequentially processed through residual connections to obtain the final high-frequency features containing image frame edge details and image frame texture details. By introducing residual connections, the high-frequency feature mapping layer can fuse the refined high-frequency features obtained after 3×3 convolution with the original high-frequency residual signal, ensuring that edge and texture details in the original high-frequency residual signal are not lost due to multi-layer transformations, while simultaneously enhancing local high-frequency patterns using the 3×3 convolutional layer. Thus, this embodiment of the invention can transform the high-frequency residual signal through the convolutional mapping of the convolutional layers in the high-frequency feature mapping layer. Transformed into high-frequency features with image frame edge details and image frame texture details consistent with spatial feature dimensions. Furthermore, to obtain enhanced features, high-frequency features need to be injected into spatial features in the gating layer. Based on this, the method includes: compressing the encoded channel dimension of the spatial features to obtain a single-channel spatial attention map; performing adaptive histogram equalization on the single-channel spatial attention map to obtain an enhanced single-channel spatial attention map; generating a gating weight map based on the enhanced single-channel spatial attention map, wherein each pixel value in the gating weight map corresponds to the high-frequency information injection intensity coefficient at the corresponding position of the spatial feature; and injecting the high-frequency features containing image frame edge details and image frame texture details into the spatial features based on the high-frequency information injection intensity coefficient to obtain enhanced features containing image frame edge details and image frame texture details.

[0041] Specifically, high-frequency features containing image frame edge details and image frame texture details will be included. The spatial features of the current image frame to be processed are input together to the gating layer. The spatial features output by the encoder are multi-channel spatial features. In order to focus on the detailed distribution of the spatial dimension, a 1×1 convolutional layer or global average pooling can be used in the gating layer to compress the channel dimension of the spatial features output by the encoder, reducing it to a single-channel spatial attention map. This attention map reflects the saliency distribution of the current frame in the spatial dimension. Secondly, in order to enhance the response of edge, texture and other detailed areas in the image frame and suppress noise in smooth areas, this embodiment of the invention performs adaptive histogram equalization on the single-channel spatial attention map. Specifically, the histogram distribution of the single-channel spatial attention map is calculated, and the pixel gray value is non-linearly mapped by the cumulative distribution function, thereby increasing the contrast between the detailed area and the background area, and obtaining the enhanced single-channel spatial attention map. Then, the enhanced single-channel spatial attention map is input to an activation function (such as the sigmoid function) and normalized to the [0.1] interval to generate the final gating weight map. Each pixel value in this weight map is the high-frequency information injection intensity coefficient of the corresponding position. The larger the value, the more high-frequency details need to be added at that position. Finally, based on the generated gating weight graph High-frequency features with edge and texture details Multiply element-wise by the weight coefficients and then by the learnable scaling parameter. (Used to control the overall injection volume), it is adaptively superimposed onto the original spatial features. Enhanced features with image frame edge details and image frame texture details. The determination process is as follows:

[0042] This invention, through the introduction of an adaptive histogram equalization mechanism to generate a gated weight map, significantly enhances the model's sensitivity to detailed regions (such as edges and textures) in image frames. This allows high-frequency features to be injected more accurately into the areas requiring enhancement, while effectively suppressing noise amplification in smooth regions. This adaptive injection strategy solves the problem of pseudo-texture or detail loss caused by uneven injection of high-frequency information in traditional methods, thereby improving the detail fidelity of remote sensing image super-resolution reconstruction while maintaining temporal consistency.

[0043] In another embodiment of the present invention, in order to determine the enhancement features through the gating layer, the method further includes: performing gating mapping processing on the spatial features to obtain a gating weight map, wherein the gating mapping processing includes convolutional mapping and Sigmoid (activation function) activation; controlling the intensity of high-frequency features injected into the spatial features based on the gating weight map to obtain gating modulated high-frequency features; adding the gating modulated high-frequency features to the spatial features to obtain enhancement features with image frame edge details and image frame texture details.

[0044] In another embodiment of the present invention, during the construction of the high-frequency residual signal in the high-frequency residual signal construction layer, Gaussian fuzzy residuals can be used to construct the high-frequency residual signal, or Laplace pyramids, wavelet decomposition, or frequency domain filtering can be used to extract high-frequency information, which is then injected into the backbone features through weighted fusion or gated fusion. In the high-frequency enhancement module, the gated generation method can also be replaced with channel attention, spatial attention, or a combination of both, to adaptively adjust the amount of high-frequency information injected according to the texture intensity and noise level of different ground features.

[0045] Furthermore, after determining the enhanced features with image frame edge details and image frame texture details, it is also necessary to extract global motion trend features and local fine-grained deformation features in the image frame through a global-local offset estimator. Based on this, step 203 specifically includes: performing layer normalization processing on the concatenated result through the global branch, and determining the channel attention map and spatial attention map based on the concatenated result after layer normalization processing; determining the global motion trend features between the current image frame to be processed and the previous image frame based on the channel attention map and spatial attention map; and performing dilated convolution operations with different dilation rates on the concatenated result through the local branch. Different deformation feature maps are obtained accordingly. Each deformation feature map is stitched together along the channel dimension to obtain a stitched feature map. Channel attention weighting is applied to the stitched feature map to obtain a channel weight vector. Each element in the channel weight vector corresponds to the importance coefficient of the corresponding channel in the stitched feature map. The channel elements of the stitched feature map are weighted and summed using the importance coefficients. The stitched feature map after channel element weighting and summing is convolved and dimensionality reduced to obtain the local fine-grained deformation features between the current image frame to be processed and the previous image frame. The local fine-grained deformation features include at least one of local displacement features, edge deformation features, and small target motion features.

[0046] Specifically, the enhanced features containing image frame edge details and image frame texture details are fused with the historical propagation features of the preceding image frame in advance to form a cascaded result. Then the cascaded results will be... The input is fed into the global branch of the global-local offset estimator, where the cascaded results are processed. After performing layer normalization, channel attention maps can be generated by connecting a 1×1 fully connected layer or a convolution after global average pooling. It is used to capture the dependencies between feature channels. At the same time, it can be used to capture the dependencies between feature channels through concatenated results, such as 7×7 large kernel convolution. Generate spatial attention map This is used to capture the saliency distribution in the spatial dimension of the feature map. Then, the global contextual relationship between the channel dimension and the spatial dimension is jointly modeled to obtain global motion trend features. As shown below:

[0047] in, , These are learnable parameters set according to actual needs.

[0048] In this embodiment of the invention, it is also necessary to utilize the local branch in the global-local offset estimator to extract local fine-grained deformation features. Specifically, the local branch can employ a multi-scale dilated convolutional pyramid, where the cascaded results are... Simultaneously, three parallel dilated convolution branches are input, and the dilation rate d of each branch is set to be different, such as d=1, 2, 3, to extract local deformation features under different receptive fields. , and The deformation feature maps under these three different expansion rates are concatenated along the channel dimension to obtain a concatenated feature map. Next, channel attention weighting is applied to this concatenated feature map to generate a channel weight vector. The importance coefficients in this vector are then used to perform a weighted summation of the channel elements of the concatenated feature map to adaptively fuse multi-scale local information. Finally, the weighted summated feature map is subjected to convolutional dimensionality reduction (e.g., 1×1 convolution) to output fine-grained local deformation features containing information on local displacement, edge deformation, and small target motion. As shown below:

[0049] in, It is a convolution function. This is a concatenation function.

[0050] This invention, through the decoupling of the extraction mechanism for global motion trends and local fine-grained deformations, enables the model to simultaneously perceive large-scale overall translational motion and minute local deformations (such as edge jitter and small target displacement) in remote sensing images. This complementarity of global and local features significantly improves the accuracy of offset estimation, thus laying a solid foundation for subsequent high-precision inter-frame alignment.

[0051] Furthermore, after extracting global motion trend features and local fine-grained deformation features, it is necessary to use the second-order deformable alignment module in the global-local offset estimator to perform cross-frame alignment of historical propagation features to enhanced features. Based on this, the method includes: the global-local offset estimator further includes a spatial pyramid network and an offset prediction branch; the spatial pyramid network determines the first-order optical flow information between the enhanced features with image frame edge details and image frame texture details and the historical propagation features, and performs bilinear interpolation on the historical propagation features based on the first-order optical flow information to obtain an initial aligned historical propagation feature that is aligned across frames with the enhanced features with image frame edge details and image frame texture details; the offset prediction branch predicts the cross-frame alignment offset based on the global motion trend features and the local fine-grained deformation features; the second-order deformable alignment module performs secondary correction on the initial aligned historical propagation features based on the cross-frame alignment offset to obtain an aligned historical propagation feature that is aligned across frames with the enhanced features with image frame edge details and image frame texture details.

[0052] Specifically, in the global-local offset estimator, the first-order optical flow information between the current enhanced feature with edge texture details and the historical propagation features is first calculated using a spatial pyramid network. This first-order optical flow information is determined by constructing a feature-related cost volume and using a convolutional neural network to regress the optical flow field. Specifically, the feature similarity between each pixel in the current enhanced feature and all displacement-position pixels in the historical propagation features is calculated. For a pixel position in the current enhanced feature, the matching score under different displacement amounts is calculated within a certain search range centered on that position in the previous historical propagation features. These matching scores are stacked to form a multi-channel cost volume tensor, which explicitly encodes the correspondence between features in two frames. Then, optical flow regression is performed: the constructed cost volume is input into an optical flow regression network, which typically employs an encoder-decoder convolutional neural network. The encoder compresses the spatial dimension of the cost volume through multiple convolutional and pooling operations to extract high-level motion features; the decoder gradually restores the resolution of the feature map through upsampling operations. Finally, the feature map is mapped to an optical flow field containing two components: horizontal and vertical displacement, through the final convolutional layer of the network; this is the final first-order optical flow information. Further, based on this first-order optical flow information, bilinear interpolation or similar resampling techniques are used to spatially transform the historical propagation features, thereby obtaining initial aligned historical propagation features that are initially aligned with the enhanced features in macroscopic motion. To further correct the alignment error, this embodiment introduces an offset prediction branch in the global-local offset estimator to predict fine-grained cross-frame alignment offsets. Specifically, the current conditional features (obtained by cascading or fusing the enhanced features and historical propagation features), first-order optical flow information, and second-order optical flow information are cascaded to obtain the input features for the offset prediction branch. This input feature is then fed into the offset prediction branch to predict the cross-frame alignment offset, where the offset integrates global motion trends and local fine-grained deformation information, representing the pixel-level displacement difference between two frames. Finally, the initial alignment history propagation features obtained above are corrected by a second-order deformable alignment module based on the cross-frame alignment offset predicted by the offset prediction branch. Specifically, the initial alignment history propagation features can be corrected into alignment history propagation features based on the cross-frame alignment offset through spatial transformation operations (such as bilinear interpolation or similar resampling techniques). By applying the offset to the initial alignment history propagation features, residual local deformation and fine-grained deviations are eliminated, ultimately resulting in alignment history propagation features that are perfectly aligned across frames with the enhanced features at the pixel level. This embodiment of the invention effectively solves the problem of difficult accurate modeling of complex motions in remote sensing images by constructing a cascaded alignment mechanism of first-order optical flow coarse alignment + offset prediction + second-order offset fine correction.The global-local offset estimator's dual-branch structure (global and local) enables the model to capture both the overall motion trend of objects in the image frame and to sensitively perceive edge jitter and small target displacement, thereby generating high-precision offsets and improving the accuracy of cross-frame alignment.

[0053] Finally, the aligned historical propagation features, along with the enhanced features containing image frame edge details and image frame texture details, are input into the image reconstruction module. The specific implementation process of the image reconstruction module mainly includes two stages: feature fusion and upsampling reconstruction. The aim is to enhance the detail recovery capability of the current frame by utilizing historical information that has been precisely spatiotemporally aligned. Specifically, firstly, feature fusion is performed: the aligned historical propagation features obtained after second-order correction are concatenated with the enhanced features at the current moment (i.e., features containing explicit edge and texture details) in the channel dimension. Then, one or more convolutional layers are used to process the concatenated features to promote deep interaction and fusion between historical temporal information and current spatial detail information. After that, the fused features are input into the upsampling module. In the upsampling module, subpixel convolution or transposed convolutional layers can be used to gradually increase the resolution of the current image frame to be processed to the target high-resolution scale, thereby generating a high-resolution remote sensing image frame.

[0054] In another embodiment of the present invention, in terms of cross-frame alignment, the global-local offset estimator can also be replaced by an attention-based cross-frame aligner, a correlation matching module, an explicit multi-scale optical flow fusion unit, or other offset prediction structures that can simultaneously take into account large-scale motion and local micro-deformation.

[0055] The present invention provides another super-resolution reconstruction method based on multi-frame remote sensing images. Compared with the current method that relies solely on the low-resolution pixel information of the current image frame itself for upsampling and reconstruction into a high-resolution image frame, the present invention assists in the super-resolution reconstruction of the current image frame by referencing the historical propagation characteristics of the preceding image frames in the same continuous remote sensing image frame sequence that have already been enhanced in resolution. This fully exploits the spatiotemporal redundancy information between remote sensing image frames. Utilizing this spatiotemporal redundancy information, the texture information and specific structure in the current image frame can be fully determined, providing reference information for weak textures in the current image frame, thereby improving the resolution enhancement effect of the image. By using an encoder trained on remote sensing images to extract spatial features, the encoder can directly learn the spatiotemporal prior representations specific to remote sensing images. This improves feature extraction accuracy. A high-frequency enhancement module enhances spatial features, ensuring edges, textures, and fine structures are supplemented before cross-frame propagation, preventing the gradual attenuation of high-frequency information during subsequent image alignment and aggregation. Global branches capture long-distance motion trends, while local branches model fine-grained deformations; their fusion yields more accurate offset predictions, overcoming the shortcomings of existing single convolutional stacking structures in modeling complex motions in remote sensing images, thus improving the comprehensiveness of feature capture. A second-order spatial transformation is performed on global motion trend features and local fine-grained deformation features, enabling pixel-level precise registration of historical propagation features and enhanced features, avoiding structural blurring and artifacts caused by misaligned fusion, thereby improving the resolution enhancement effect of remote sensing images and ensuring the image quality of the enhanced remote sensing images.

[0056] Furthermore, as Figure 1 In specific implementation, embodiments of the present invention provide a super-resolution reconstruction device based on multi-frame remote sensing images, such as... Figure 3 As shown, the device includes: an image acquisition unit 31, a model acquisition unit 32, and a resolution enhancement unit 33.

[0057] The image acquisition unit 31 can be used to acquire the historical propagation features of the current image frame to be processed and the previous image frames in the remote sensing image frame sequence. The historical propagation features are obtained by extracting features from at least two previous image frames and fusing the extracted features.

[0058] The model acquisition unit 32 can be used to acquire a preset image super-resolution model, wherein the preset image super-resolution model includes an encoder for feature extraction, a high-frequency enhancement module for feature enhancement, a global-local offset estimator for cross-frame alignment, and an image reconstruction module for image reconstruction. The global-local offset estimator includes a global branch, a local branch, and a second-order deformable alignment module. The preset image super-resolution model is pre-trained based on a sample dataset.

[0059] The resolution enhancement unit 33 can be used to input the current image frame to be processed and the historical propagation features into the preset image super-resolution model. The encoder extracts spatial features from the current image frame to be processed. The high-frequency enhancement module enhances the spatial features into enhanced features with image frame edge details and image frame texture details. The global-local offset estimator extracts global motion trends and local fine-grained deformations between frames from the concatenated results of the enhanced features with image frame edge details and image frame texture details and the historical propagation features. The global motion trend features corresponding to the global branch and the local fine-grained deformation features corresponding to the local branch are obtained. The second-order deformable alignment module performs spatial transformation on the historical propagation features based on the global motion trend features and the local fine-grained deformation features to obtain aligned historical propagation features that are aligned across frames with the enhanced features. The image reconstruction module reconstructs the enhanced features with image frame edge details and image frame texture details based on the aligned historical propagation features to obtain the high-resolution remote sensing image frame corresponding to the current image frame to be processed.

[0060] In specific application scenarios, in order to train and build an encoder, such as Figure 4 As shown, the device also includes a construction unit 34.

[0061] The construction unit 34 can be used to acquire a shared encoder; acquire sample low-resolution remote sensing images, wherein the sample low-resolution remote sensing images include a sequence of low-resolution remote sensing image frames with standard high-resolution remote sensing image frame annotation information, the standard high-resolution remote sensing image frame annotation information being obtained by annotating the standard high-resolution remote sensing image frames corresponding to each low-resolution remote sensing image frame in the low-resolution remote sensing image frame sequence; extract frame-level features of each low-resolution remote sensing image frame in the sample low-resolution remote sensing images using the shared encoder; determine the central low-resolution remote sensing image frame and the neighboring low-resolution remote sensing image frames adjacent to the central low-resolution remote sensing image frame in the sample low-resolution remote sensing images; and perform [further processing / processing]. The center frame-level features of the central low-resolution remote sensing image frame are subjected to block-level random masking, and the neighboring frame-level features of the adjacent low-resolution remote sensing image frames are subjected to auxiliary masking. The center frame-level features after block-level random masking and the neighboring frame-level features after auxiliary masking are fused to obtain fused features. The fused features are decoded into the predicted center high-resolution remote sensing image frame corresponding to the central low-resolution remote sensing image frame using a preset decoder. A loss function is determined based on the difference between the standard high-resolution remote sensing image frame annotation information corresponding to the central low-resolution remote sensing image frame and the predicted center high-resolution remote sensing image frame. The shared encoder is iteratively trained based on the loss function, and the encoder is determined based on the iterative training results.

[0062] In specific application scenarios, in order to perform masking processing on the center frame-level features and neighboring frame-level features, the construction unit 34 can be specifically used to divide the center frame-level features into multiple non-overlapping center feature blocks, randomly select multiple center feature blocks to be masked from the multiple center feature blocks at a first preset ratio, and perform masking processing on each of the center feature blocks to be masked in the center frame-level features to obtain a binary center mask map. Based on the binary center mask map and the center frame-level features, the center frame-level features after block-level random masking are determined. Similarly, the neighboring frame-level features are divided into multiple non-overlapping neighboring feature blocks, and multiple neighboring feature blocks to be masked are randomly selected from the multiple neighboring feature blocks at a second preset ratio. Each of the neighboring feature blocks to be masked in the neighboring frame-level features is then masked to obtain a binary neighboring mask map. Based on the binary neighboring mask map and the neighboring frame-level features, the neighboring frame-level features after auxiliary masking are determined. The second preset ratio is less than the first preset ratio.

[0063] In specific application scenarios, the high-frequency enhancement module sequentially includes a Gaussian filter layer, a high-frequency residual signal construction layer, a high-frequency feature mapping layer, and a gating layer; in order to enhance the details of spatial features, the resolution enhancement unit 33 includes a construction module 331, a conversion module 332, and an injection module 333.

[0064] The construction module 331 can be used to separate low-frequency components in the current image frame to be processed through the Gaussian filtering layer, and construct the low-frequency components into high-frequency residual signals through the high-frequency residual signal construction layer.

[0065] The conversion module 332 can be used to convert the high-frequency residual signal into high-frequency features with image frame edge details and image frame texture details that are consistent with the spatial feature dimension through the high-frequency feature mapping layer.

[0066] The injection module 333 can be used to inject the high-frequency features with image frame edge details and image frame texture details into the spatial features through the gate layer to obtain enhanced features with image frame edge details and image frame texture details.

[0067] In specific application scenarios, in order to process spatial features into enhanced features with image frame edge details and image frame texture details, the injection module 333 can be specifically used to compress the encoding channel dimension of the spatial features to obtain a single-channel spatial attention map, perform adaptive histogram equalization on the single-channel spatial attention map to obtain the enhanced single-channel spatial attention map, generate a gated weight map based on the enhanced single-channel spatial attention map, wherein each pixel value in the gated weight map corresponds to the high-frequency information injection intensity coefficient of the corresponding position of the spatial feature; and inject the high-frequency features with image frame edge details and image frame texture details into the spatial features based on the high-frequency information injection intensity coefficient to obtain the enhanced features with image frame edge details and image frame texture details.

[0068] In specific application scenarios, the global-local offset estimator also includes a spatial pyramid network and an offset prediction branch; in order to spatially transform the historical propagation features into aligned historical propagation features that are aligned with the enhanced features across frames, the resolution enhancement unit 33 also includes an interpolation module 334, a prediction module 335 and a correction module 336.

[0069] The interpolation module 334 can be used to determine the first-order optical flow information between the enhancement feature with image frame edge details and image frame texture details and the historical propagation feature through the spatial pyramid network, and perform bilinear interpolation on the historical propagation feature based on the first-order optical flow information to obtain an initial aligned historical propagation feature that is aligned across frames with the enhancement feature with image frame edge details and image frame texture details.

[0070] The prediction module 335 can be used to predict cross-frame alignment offset based on the global motion trend features and the local fine-grained deformation features through the offset prediction branch.

[0071] The correction module 336 can be used to perform secondary correction on the initial alignment history propagation feature based on the cross-frame alignment offset by the second-order deformable alignment module to obtain the alignment history propagation feature that is aligned across frames with the enhanced feature containing image frame edge details and image frame texture details.

[0072] In specific application scenarios, in order to extract global motion trends and local fine-grained deformations between frames, the resolution enhancement unit 33 also includes a feature extraction module 337.

[0073] The feature extraction module 337 can be used to perform layer normalization processing on the cascaded results through the global branch, and determine the channel attention map and spatial attention map based on the cascaded results after layer normalization processing. Based on the channel attention map and the spatial attention map, the global motion trend features between the current image frame to be processed and the preceding image frame are determined. Different deformation feature maps are obtained by performing dilated convolution operations with different dilation rates on the cascaded results through the local branch. Each deformation feature map is stitched along the channel dimension to obtain a stitched feature map. The stitched feature map is then weighted by channel attention to obtain a channel weight vector, where each element in the channel weight vector corresponds to the importance coefficient of the corresponding channel in the stitched feature map. The importance coefficient is used to perform channel element weighted summation on the stitched feature map. The stitched feature map after channel element weighted summation is then subjected to convolution dimensionality reduction to obtain the local fine-grained deformation features between the current image frame to be processed and the preceding image frame. The local fine-grained deformation features include at least one of local displacement features, edge deformation features, and small target motion features.

[0074] It should be noted that other corresponding descriptions of the functional modules involved in the super-resolution reconstruction device based on multi-frame remote sensing images provided in this embodiment of the invention can be found in [reference]. Figure 1 The corresponding description of the method shown will not be repeated here.

[0075] Based on the above, Figure 1The method shown, correspondingly, also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the following steps: acquiring the historical propagation features of the current image frame to be processed and the preceding image frames in a remote sensing image frame sequence, wherein the historical propagation features are obtained by extracting features from at least two preceding image frames and fusing the extracted features; acquiring a preset image super-resolution model, wherein the preset image super-resolution model includes an encoder for feature extraction, a high-frequency enhancement module for feature enhancement, a global-local offset estimator for cross-frame alignment, and an image reconstruction module for image reconstruction, wherein the global-local offset estimator includes a global branch, a local branch, and a second-order deformable alignment module, and the preset image super-resolution model is pre-trained based on a sample dataset; inputting the current image frame to be processed and the historical propagation features into the preset image super-resolution model, and through the encoder... Spatial features are extracted from the current image frame to be processed. The high-frequency enhancement module enhances the spatial features into enhanced features with image frame edge details and image frame texture details. The global-local offset estimator extracts the global motion trend and local fine-grained deformation between frames from the concatenated results of the enhanced features with image frame edge details and image frame texture details and the historical propagation features. The global motion trend features corresponding to the global branch and the local fine-grained deformation features corresponding to the local branch are obtained. The second-order deformable alignment module performs spatial transformation on the historical propagation features based on the global motion trend features and the local fine-grained deformation features to obtain aligned historical propagation features that are aligned with the enhanced features across frames. The image reconstruction module reconstructs the enhanced features with image frame edge details and image frame texture details based on the aligned historical propagation features to obtain the high-resolution remote sensing image frame corresponding to the current image frame to be processed.

[0076] Based on the above, Figure 1 The method shown and as Figure 3 The embodiment of the device shown in the invention also provides a physical structure diagram of a computer device, such as... Figure 5As shown, the computer device includes: a processor 41, a memory 42, and a computer program stored in the memory 42 and executable on the processor. Both the memory 42 and the processor 41 are mounted on a bus 43. When the processor 41 executes the program, it performs the following steps: acquiring historical propagation features of the current image frame to be processed and previous image frames in a remote sensing image frame sequence, wherein the historical propagation features are obtained by feature extraction from at least two previous image frames and fusion of the extracted features; acquiring a preset image super-resolution model, wherein the preset image super-resolution model includes an encoder for feature extraction, a high-frequency enhancement module for feature enhancement, a global-local offset estimator for cross-frame alignment, and an image reconstruction module for image reconstruction. The global-local offset estimator includes a global branch, a local branch, and a second-order deformable alignment module. The preset image super-resolution model is pre-trained based on a sample dataset; and inputting the current image frame to be processed and the historical propagation features into the preset image super-resolution model. In the image super-resolution model, the encoder extracts spatial features from the current image frame to be processed. The high-frequency enhancement module enhances the spatial features into enhanced features with image frame edge details and image frame texture details. The global-local offset estimator extracts global motion trends and local fine-grained deformations between frames from the concatenated results of the enhanced features with image frame edge details and image frame texture details and the historical propagation features. This yields global motion trend features corresponding to the global branch and local fine-grained deformation features corresponding to the local branch. The second-order deformable alignment module performs spatial transformation on the historical propagation features based on the global motion trend features and local fine-grained deformation features to obtain aligned historical propagation features that are aligned across frames with the enhanced features. The image reconstruction module reconstructs the high-resolution remote sensing image frame corresponding to the current image frame to be processed based on the aligned historical propagation features and the enhanced features with image frame edge details and image frame texture details.

[0077] Through the technical solution of this invention, the present invention assists in the super-resolution reconstruction of the current image frame by referencing the historical propagation characteristics of the preceding image frames in the same continuous remote sensing image frame sequence that have already been enhanced in resolution. This fully exploits the spatiotemporal redundancy information between remote sensing image frames. Utilizing this spatiotemporal redundancy information, the texture information and specific structure in the current image frame can be fully determined, providing reference information for weak textures in the current image frame, thereby improving the resolution enhancement effect of the image. By using an encoder trained on remote sensing images to extract spatial features, the encoder can directly learn the spatiotemporal prior representations specific to remote sensing images, thereby improving the feature extraction accuracy. The spatial features are enhanced through a high-frequency enhancement module. This approach ensures that edges, textures, and fine structures are filled in before cross-frame propagation, avoiding the problem of high-frequency information gradually attenuating during subsequent image alignment and aggregation. By capturing long-distance motion trends through global branches and modeling fine-grained deformations through local branches, the fusion of the two yields more accurate offset predictions, overcoming the shortcomings of existing single convolution stacked structures in modeling complex motions in remote sensing images, thereby improving the comprehensiveness of feature capture. A second-order spatial transformation is performed on global motion trend features and local fine-grained deformation features, enabling pixel-level accurate registration of historical propagation features and enhancement features, avoiding structural blurring and artifacts caused by misaligned fusion, thus improving the resolution enhancement effect of remote sensing images and ensuring the image quality of the enhanced remote sensing images.

[0078] It is obvious to those skilled in the art that the modules or steps of the present invention described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.

[0079] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A super-resolution reconstruction method based on multi-frame remote sensing images, characterized in that, include: The historical propagation features of the current image frame to be processed and the previous image frames in the remote sensing image frame sequence are obtained, wherein the historical propagation features are obtained by extracting features from at least two previous image frames and fusing the extracted features. A preset image super-resolution model is obtained, wherein the preset image super-resolution model includes an encoder for feature extraction, a high-frequency enhancement module for feature enhancement, a global-local offset estimator for cross-frame alignment, and an image reconstruction module for image reconstruction. The global-local offset estimator includes a global branch, a local branch, and a second-order deformable alignment module. The preset image super-resolution model is pre-trained based on a sample dataset. The current image frame to be processed and the historical propagation features are input into the preset image super-resolution model. The encoder extracts spatial features from the current image frame to be processed. The high-frequency enhancement module enhances the spatial features into enhanced features with image frame edge details and image frame texture details. The global-local offset estimator extracts global motion trends and local fine-grained deformations between frames from the concatenated results of the enhanced features with image frame edge details and image frame texture details and the historical propagation features. The global motion trend features corresponding to the global branch and the local fine-grained deformation features corresponding to the local branch are obtained. The second-order deformable alignment module performs spatial transformation on the historical propagation features based on the global motion trend features and the local fine-grained deformation features to obtain aligned historical propagation features that are aligned across frames with the enhanced features. The image reconstruction module reconstructs the high-resolution remote sensing image frame corresponding to the current image frame to be processed based on the aligned historical propagation features.

2. The super-resolution reconstruction method based on multi-frame remote sensing images according to claim 1, characterized in that, Before obtaining the preset image super-resolution model, the method further includes: Obtain the shared encoder; Acquire sample low-resolution remote sensing images, wherein the sample low-resolution remote sensing images include a sequence of low-resolution remote sensing image frames with standard high-resolution remote sensing image frame annotation information, wherein the standard high-resolution remote sensing image frame annotation information is obtained by annotating the standard high-resolution remote sensing image frame corresponding to each low-resolution remote sensing image frame in the sequence of low-resolution remote sensing image frames. The shared encoder is used to extract the frame-level features of each low-resolution remote sensing image frame in the sample low-resolution remote sensing image. The central low-resolution remote sensing image frame and the neighboring low-resolution remote sensing image frames adjacent to the central low-resolution remote sensing image frame are determined in the sample low-resolution remote sensing image. The center frame-level features of the central low-resolution remote sensing image frame are subjected to block-level random masking, and the neighboring frame-level features of the neighboring low-resolution remote sensing image frames are subjected to auxiliary masking. The block-level random masked center frame-level features and the auxiliary masked neighboring frame-level features are fused to obtain fused features. The fused features are decoded into the predicted center high-resolution remote sensing image frame corresponding to the central low-resolution remote sensing image frame using a preset decoder. A loss function is determined based on the difference between the standard high-resolution remote sensing image frame annotation information corresponding to the central low-resolution remote sensing image frame and the predicted central high-resolution remote sensing image frame. The shared encoder is iteratively trained based on the loss function, and the encoder is determined based on the iterative training results.

3. The super-resolution reconstruction method based on multi-frame remote sensing images according to claim 2, characterized in that, Block-level random masking is performed on the center frame-level features of the central low-resolution remote sensing image frame, including: The central frame-level feature is divided into multiple non-overlapping central feature blocks. Multiple central feature blocks to be masked are randomly selected from the multiple central feature blocks at a first preset ratio. Each of the central feature blocks to be masked in the central frame-level feature is masked to obtain a binary central mask map. The central frame-level feature after block-level random masking is determined based on the binary central mask map and the central frame-level feature. Auxiliary masking is performed on the neighboring frame-level features of the adjacent low-resolution remote sensing image frames, including: The neighboring frame-level features are divided into multiple non-overlapping neighboring feature blocks. Multiple neighboring feature blocks to be masked are randomly selected from the multiple neighboring feature blocks at a second preset ratio. Each of the neighboring feature blocks to be masked in the neighboring frame-level features is masked to obtain a binary neighbor mask map. The neighboring frame-level features after auxiliary masking are determined based on the binary neighbor mask map and the neighboring frame-level features. The second preset ratio is less than the first preset ratio.

4. The super-resolution reconstruction method based on multi-frame remote sensing images according to claim 1, characterized in that, The high-frequency enhancement module includes, in sequence, a Gaussian filter layer, a high-frequency residual signal construction layer, a high-frequency feature mapping layer, and a gating layer; The spatial features are enhanced into enhanced features with image frame edge details and image frame texture details through the high-frequency enhancement module, including: The low-frequency component is separated from the current image frame to be processed by the Gaussian filtering layer, and the low-frequency component is constructed into a high-frequency residual signal by the high-frequency residual signal construction layer. The high-frequency residual signal is converted into high-frequency features with image frame edge details and image frame texture details that are consistent with the spatial feature dimension through the high-frequency feature mapping layer; The high-frequency features containing image frame edge details and image frame texture details are injected into the spatial features through the gated layer to obtain enhanced features containing image frame edge details and image frame texture details.

5. The super-resolution reconstruction method based on multi-frame remote sensing images according to claim 4, characterized in that, The high-frequency features containing image frame edge details and image frame texture details are injected into the spatial features through the gated layer to obtain enhanced features containing image frame edge details and image frame texture details, including: The spatial features are compressed by encoding channel dimension to obtain a single-channel spatial attention map. The single-channel spatial attention map is then subjected to adaptive histogram equalization to obtain an enhanced single-channel spatial attention map. A gated weight map is generated based on the enhanced single-channel spatial attention map, wherein each pixel value in the gated weight map corresponds to the high-frequency information injection intensity coefficient at the corresponding position of the spatial feature. Based on the high-frequency information injection intensity coefficient, the high-frequency features with image frame edge details and image frame texture details are injected into the spatial features to obtain enhanced features with image frame edge details and image frame texture details.

6. The super-resolution reconstruction method based on multi-frame remote sensing images according to claim 1, characterized in that, The global-local offset estimator also includes a spatial pyramid network and an offset prediction branch; The second-order deformable alignment module performs spatial transformation on the historical propagation features based on the global motion trend features and the local fine-grained deformation features to obtain aligned historical propagation features that are aligned across frames with the enhanced features, including: The spatial pyramid network determines the first-order optical flow information between the enhancement feature with image frame edge details and image frame texture details and the historical propagation feature, and performs bilinear interpolation on the historical propagation feature based on the first-order optical flow information to obtain an initial aligned historical propagation feature that is aligned across frames with the enhancement feature with image frame edge details and image frame texture details. The offset prediction branch predicts the cross-frame alignment offset based on the global motion trend features and the local fine-grained deformation features; The second-order deformable alignment module performs a secondary correction on the initial alignment history propagation feature based on the cross-frame alignment offset to obtain the alignment history propagation feature that is aligned across frames with the enhanced feature containing image frame edge details and image frame texture details.

7. The super-resolution reconstruction method based on multi-frame remote sensing images according to claim 1, characterized in that, The global-local offset estimator performs inter-frame global motion trend extraction and local fine-grained deformation extraction on the concatenated results of the enhanced features and the historical propagation features, which contain image frame edge details and image frame texture details, including: The concatenated results are subjected to layer normalization through the global branch, and channel attention map and spatial attention map are determined based on the concatenated results after layer normalization. The global motion trend features between the current image frame to be processed and the preceding image frame are determined based on the channel attention map and spatial attention map. Different deformation feature maps are obtained by performing dilated convolution operations with different dilation rates on the cascaded results through the local branches. Each deformation feature map is then concatenated along the channel dimension to obtain a stitched feature map. Channel attention weighting is applied to the stitched feature map to obtain a channel weight vector, where each element in the channel weight vector corresponds to the importance coefficient of the corresponding channel in the stitched feature map. The channel elements of the stitched feature map are weighted and summed using the importance coefficients. The stitched feature map after weighted summation of the channel elements is then convolved and dimensionality reduced to obtain the local fine-grained deformation features between the current image frame to be processed and the previous image frame. The local fine-grained deformation features include at least one of local displacement features, edge deformation features, and small target motion features.

8. A super-resolution reconstruction device based on multi-frame remote sensing images, characterized in that, include: The image acquisition unit is used to acquire the historical propagation features of the current image frame to be processed and the previous image frames in the remote sensing image frame sequence, wherein the historical propagation features are obtained by extracting features from at least two previous image frames and fusing the extracted features. The model acquisition unit is used to acquire a preset image super-resolution model, wherein the preset image super-resolution model includes an encoder for feature extraction, a high-frequency enhancement module for feature enhancement, a global-local offset estimator for cross-frame alignment, and an image reconstruction module for image reconstruction. The global-local offset estimator includes a global branch, a local branch, and a second-order deformable alignment module. The preset image super-resolution model is pre-trained based on a sample dataset. The resolution enhancement unit is used to input the current image frame to be processed and the historical propagation features into the preset image super-resolution model. The encoder extracts spatial features from the current image frame to be processed. The high-frequency enhancement module enhances the spatial features into enhanced features with image frame edge details and image frame texture details. The global-local offset estimator extracts global motion trends and local fine-grained deformations between frames from the concatenated results of the enhanced features with image frame edge details and image frame texture details and the historical propagation features. This yields global motion trend features corresponding to the global branch and local fine-grained deformation features corresponding to the local branch. The second-order deformable alignment module performs spatial transformation on the historical propagation features based on the global motion trend features and local fine-grained deformation features to obtain aligned historical propagation features that are aligned across frames with the enhanced features. The image reconstruction module reconstructs the enhanced features with image frame edge details and image frame texture details based on the aligned historical propagation features to obtain the high-resolution remote sensing image frame corresponding to the current image frame to be processed.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Video super-resolution method based on local feature enhancement

    CN119313566A

  • Remote sensing video super-resolution reconstruction method based on inter-frame information compensation

    CN121746187A