A spatio-spectral joint super-resolution reconstruction method based on a giant remote sensing satellite cluster

By employing a spatiotemporal-spectral joint super-resolution reconstruction method based on giant remote sensing constellations, the problem of integrated spatiotemporal-spectral modeling of multi-source heterogeneous remote sensing data was solved. This method enables the reconstruction of images with high spatial resolution, temporal resolution, and high spectral fidelity, and is suitable for tasks such as urban fine mapping, agricultural monitoring, ecological assessment, disaster emergency response, and battlefield situational awareness.

CN121073780BActive Publication Date: 2026-02-10CHINA UNIV OF MINING & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511604132.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2026-02-10
Estimated Expiration
2045-11-05

AI Technical Summary

Technical Problem

Existing remote sensing super-resolution reconstruction methods struggle to achieve integrated spatiotemporal spectral modeling in multi-source heterogeneous scenarios, making it difficult to align cross-load radiation responses, stably achieve sub-pixel accuracy in joint registration, and suffer from temporal consistency issues due to cloud cover and imaging condition fluctuations. Furthermore, improving spatial clarity often comes at the expense of spectral fidelity.

Method used

A spatiotemporal spectral joint super-resolution reconstruction method based on giant remote sensing constellations is adopted. This method generates reconstructed images with high spatial and temporal resolution while maintaining high spectral fidelity by standardizing radiometric and atmospheric corrections of multi-source satellite data, joint registration of unified projection and reference time, extracting temporal variation features by constructing 3D convolution with time windows, fusing multi-view information by 2D convolution, refining spectral features by 1D convolution, and super-resolution reconstruction by attention cross-modal weighting and conditional diffusion.

Benefits of technology

Maintaining stable reconstruction quality under complex weather and extreme perspective conditions significantly improves the engineering usability and physical interpretability of multi-source fusion reconstruction, ensuring the accuracy and stability of subsequent classification and inversion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121073780B_ABST
    Figure CN121073780B_ABST
Patent Text Reader

Abstract

The application discloses a kind of spatiotemporal spectrum joint super-resolution reconstruction methods based on giant remote sensing star group, including from multiple isomerism satellites acquisition same area time series, multi-angle and multispectral data and carry out radiation calibration and atmospheric correction;Subpixel level alignment of multi-source data is realized using joint registration model;Three-dimensional convolution is used to extract time-varying characteristics, two-dimensional convolution is used to extract spatial structure and texture characteristics, and one-dimensional convolution is used to extract and reduce dimension spectral characteristics;Through attention mechanism, time, space and spectral characteristics are adaptively weighted and fused to generate joint feature tensor;Super-resolution reconstruction is carried out, and high spatial resolution, high time resolution and high spectral fidelity target image are obtained.The method considers resolution improvement and spectral authenticity, and is suitable for fine city mapping, agricultural monitoring, ecological environment assessment, disaster emergency and battlefield situation awareness, long-range reconnaissance, target change detection and damage assessment and other high-precision remote sensing application scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of remote sensing information acquisition and intelligent reconstruction, and particularly relates to a space-time spectrum joint super-resolution reconstruction method based on a giant remote sensing star cluster. BACKGROUND

[0002] With the advancement of commercial spaceflight and mass-produced satellite manufacturing, the giant remote sensing star cluster with high revisit and high coverage capacity realizes rapid deployment and on-orbit cooperation, so that multi-source data acquisition in a shorter time scale, more observation angles and a wider spectral range for the same area becomes the norm. This trend provides more timely and multi-dimensional observation basis for applications such as urban fine mapping, agricultural growth monitoring, ecological environment assessment, disaster emergency response and battlefield situation awareness. However, the differences in multi-source heterogeneous data in imaging mechanism, radiation response, time sampling, geometric imaging and noise statistics are significant: the differences in spectral response function and calibration link of different loads lead to the fact that the intensity and spectral shape cannot be directly compared across sources; the projection deviation and parallax effect caused by orbit and attitude disturbance and viewing angle difference make it difficult for multi-source images to naturally align under the same geographic reference; the time instability introduced by clouds, aerosols and changes in sunlight conditions, and the coupling of real changes in ground objects and changes in imaging conditions increase the difficulty of time series discrimination.

[0003] In the aspect of super-resolution reconstruction, existing single-source or weakly fused methods mostly focus on improving spatial clarity, often at the expense of spectral consistency, which in turn affects the reliability of subsequent classification, parameter inversion and physical simulation. Some studies attempt to improve alignment accuracy through more precise radiation calibration and atmospheric correction, RPC-based rigorous geometric models or deep registration networks, and introduce adversarial generation, variational coding or a hybrid structure of convolution and Transformer in the reconstruction stage to enhance texture details. However, these approaches still lack a unified space-time spectrum integrated modeling and fusion framework, making it difficult to simultaneously consider spatial resolution improvement, time consistency maintenance and spectral authenticity preservation; they lack systematic processing of multi-source heterogeneity, especially in spectral response alignment, cross-source statistical normalization and quality weight modeling; and the reconstruction process rarely incorporates physical degradation models such as point spread, spectral response and down-sampling mapping into end-to-end constraints, which can easily result in over-sharpening, false colors and geometric drift.

[0004] However, existing remote sensing super-resolution reconstruction methods in multi-source heterogeneous scenarios generally lack space-time spectrum integrated modeling, cross-load radiation and spectral response alignment, joint registration to sub-pixel level, time consistency easily disturbed by clouds, aerosols and imaging conditions, and spectral fidelity sacrificed when improving spatial clarity, making it difficult to meet the dual requirements of accuracy and stability for subsequent classification and inversion. In view of the above challenges, the present application proposes a space-time spectrum joint super-resolution reconstruction method based on a giant remote sensing star cluster. Summary of the Invention

[0005] The purpose of this invention is to provide a spatiotemporal spectral joint super-resolution reconstruction method based on giant remote sensing constellations. This method addresses key challenges such as inconsistent radiometric responses from multi-source heterogeneous data, difficulties in achieving stable sub-pixel accuracy through joint registration, complex temporal dynamics leading to temporal mismatch, and the often-at-the-cost trade-off between improving spatial clarity and spectral fidelity. This method maintains stable reconstruction quality under complex weather conditions, extreme perspectives, and signal-to-noise imbalances, and provides reliable input for subsequent classification, parameter inversion, and mechanism simulation, significantly improving the engineering usability and physical interpretability of multi-source fusion reconstruction.

[0006] To achieve the above objectives, this invention provides a spatiotemporal spectral joint super-resolution reconstruction method based on giant remote sensing constellations, comprising the following steps:

[0007] S1. Data collected from multiple satellite sources in the same region, with radiation and atmospheric correction standardization;

[0008] S2. Unified projection and reference time, joint registration to subpixel;

[0009] S3. Construct a time window and use 3D convolution to extract temporal variation features;

[0010] S4 and 2D convolution fuse multi-view information to enhance structure and texture;

[0011] S5 and 1D convolution are used to extract spectral features and appropriately reduce dimensionality;

[0012] S6. Attention is weighted across modalities to generate a joint feature tensor;

[0013] S7. Conditional diffusion super-resolution reconstruction with multi-scale optimization to preserve spectral fidelity.

[0014] Preferably, step S1 is as follows:

[0015] S1.1-1 Obtaining the same region at different times from a giant remote sensing constellation Different sources With different spectral bands Original image digital values ​​( ), that is, the first Time, Number Source, No. Band pixels are ;

[0016] S1.1-2. For each source and its band, use the spaceborne calibration coefficients to... Linear mapping to top atmospheric radiance To eliminate the difference between the readout link and the electronic gain:

[0017] ;

[0018] in, For gain, As a bias, this unifies the outputs of different loads to physically interpretable radiation dimensions;

[0019] S1.1-3. For reflectivity bands such as visible light and near-infrared, considering the Earth-Sun distance and solar geometry, calculate the reflectivity of the upper atmosphere to eliminate the diurnal variation of incident irradiance:

[0020] ;

[0021] in, For the first The distance between the Earth and the Sun, For band Solar irradiance, The solar zenith angle;

[0022] S1.1-4. Based on Planck's law, the thermal infrared radiation brightness is inverted into brightness temperature for subsequent thermodynamic quantity inversion and quality control:

[0023] ;

[0024] in, , For this band The Planck constant term;

[0025] S1.1-5, Using a parameterized radiative transfer framework for... Correction yields the surface reflectance. Explicitly separated path radiation and atmospheric transmittance terms:

[0026] ;

[0027] in, For path radiation, and These represent the upward and downward atmospheric transmittance, respectively. This refers to the downward atmospheric irradiance.

[0028] S1.1-6. To eliminate the differences in the spectral response functions (SRF) of different sensors, the source data are mapped to a unified target spectral reference for any target band. The equivalent radiance is expressed as:

[0029] ;

[0030] in Let be the spectral response function of the target band. Obtained by source band interpolation, mean-variance normalization was then performed using the reference source "ref" as a benchmark to unify the cross-source statistical scale, and surface reflectance was selected as the normalization object:

[0031] ;

[0032] in, , For source band The mean and standard deviation of surface reflectance, , For reference, the corresponding statistics from the same source in the same band are all calculated on the effective pixels of the quality mask;

[0033] S1.1-7. Utilizing orbit and attitude parameters, combined with RPC and a rigorous geometric model, orthorectify each image and project it uniformly onto a common map coordinate system and pixel grid to obtain a geometrically consistent data cube. Simultaneously, generate a quality mask based on cloud detection, shadow determination, water body identification, and saturation marking. This ultimately forms a standardized input dataset for subsequent joint registration and feature learning.

[0034] .

[0035] Preferably, step S2 is as follows:

[0036] S2.1-1. Select a unified map projection and pixel grid as the geographic reference system, and set the alignment reference time as... The source index is pixel coordinates are The three-dimensional points on the Earth's surface are and , No. Source in time The image is The imaging geometry uses a rational polynomial camera model:

[0037] ;

[0038] in, The exterior and interior orientation parameters are obtained through time interpolation. Used for time unification;

[0039] S2.1-2, Utilizing and Orthorectify the images from each source onto the reference frame to obtain Select reference source Reference image The remaining sources are based on the initial displacement. Coarse alignment:

[0040] ;

[0041] S2.1-3, Using a shared feature network Calculate feature map Using normalized cross-correlation as the matching criterion, the most similar location is searched in the local neighborhood of the candidate location in the reference image to construct a set of candidate corresponding points:

[0042] ;

[0043] And on Mesh homogenization is applied, retaining only the pair with the highest cross-correlation peaks within each mesh; matching is performed only within the effective pixels of the quality mask to eliminate the effects of clouds, shadows, water bodies, and saturation; a random sampling consensus method is used to perform minimum sampling, symmetric propagation error assessment, and interior set update; linear equations are established using four pairs of matching to obtain the homography parameters and homography matrix. The error metric is the symmetric propagation error, and the interior point determination threshold is... To obtain the optimal set of interior points Then, using weighted least squares in The initial homography matrix is ​​obtained by precise estimation. Define the initial geometric mapping: ;

[0044] S2.2-4. Parameterization of non-rigid deformation fields using B-splines The control point parameters are The composite mapping is represented as:

[0045] ;

[0046] in, For the reference pixel coordinates, This is the initial geometric mapping obtained in the previous stage. For B-spline parameterized non-rigid incremental deformation fields, This is the control parameter vector for the corresponding viewpoint. It is pass Resampled image;

[0047] Construct the joint energy function:

[0048] ;

[0049] in, For reference only. for Robust loss, As a projection operator, it projects three-dimensional surface points. Projected to view At any moment pixel coordinates, For the corresponding projection model parameters, Yes The first-order spatial gradient, and ;

[0050] S2.2-5, Establish a scale pyramid index , in scale To linearize the data terms, the Gauss-Newton method is used to calculate the parameter increments:

[0051] ;

[0052] in for about Jacobi, For approximate Hessian matrix, The gradient vector is then updated. ;

[0053] S2.2-6, in the block domain Residual translation estimated using phase correlation :

[0054] ;

[0055] in, For reference, at the base time Block-domain image, Sources to be aligned at time The block-domain image has been resampled through a preceding mapping. It is the two-dimensional discrete Fourier transform and inverse transform;

[0056] And perform translation compensation on the mapping:

[0057] ;

[0058] S2.2-7. When the acquisition time is inconsistent with the reference time, i.e. Introducing a linear slowly varying motion model ,in Define the time correction mapping:

[0059] ;

[0060] In combined energy replace Minimize to improve spatiotemporal consistency;

[0061] S2.2-8. Resample the optimized geometric map using cubic convolution interpolation to generate a registered image consistent with the reference raster:

[0062] ;

[0063] Acceptance is based on two indicators: root mean square error of projected coordinate reprojection and normalized cross-correlation. The criterion is... If the target is not met, the process returns to multi-scale iteration for further refinement.

[0064] in, This represents the root mean square error of the projected coordinate reprojection. It is the geometric error threshold. Represents the normalized cross-correlation coefficient. This is the correlation threshold;

[0065] S2.2-9. Summarize the registration results and quality masks from all sources as unified input for subsequent spatiotemporal spectral feature extraction and super-resolution reconstruction:

[0066] .

[0067] Preferably, the step in S3 of extracting the target's change pattern over time from the time series using three-dimensional convolution is as follows:

[0068] S3.1-1, Let the reconstruction time be... The time window length is Time index set The registered image and mask of S2 are as follows:

[0069] ;

[0070] Quality weighting and fallback factor:

[0071] ;

[0072] in This is a fallback factor for missing measurements, used to prevent the entire time chain from breaking, thus resulting in a weighted time stack:

[0073] ;

[0074] S3.1-2. To alleviate the amplitude imbalance across sources and time periods, at the source... With band Perform zero-mean unit variance normalization:

[0075] ;

[0076] in, For the time stack with quality weights, at the target time... time window Data from internal tissues, It is the normalized temporal tensor. and exist Regarding the source With band Calculate the mean and standard deviation. As a stability constant, the first-order time difference is calculated and used together with the normalized intensity as the input channel:

[0077] ;

[0078] in, yes The forward first-order time difference, and ;

[0079] Will and The channels are stitched together to form a three-dimensional convolutional input. :

[0080] ;

[0081] in, This indicates a splicing operation along the channel dimension. Indicates at time The input tensor of the 3D convolution, For each source With band In the window By normalizing the intensity tensor, the amplitude differences between the source and time step are suppressed. That is The first-order temporal difference explicitly characterizes short-term dynamics, highlighting clues of motion and abrupt changes;

[0082] S3.1-3, Let the first... Layer input is Using core size 3D convolution with bias and activation:

[0083] ;

[0084] in , It is a non-linear activation function. ;

[0085] S3.1-4, Stacked residual 3D convolutional blocks, and time dilation rate in the time dimension. :

[0086] ;

[0087] Increasing It covers short-term details and long-term gradual changes without downsampling, and the residual path ensures gradient stability.

[0088] S3.1-5. Within the same time window, the effectiveness of frames varies significantly. Frames obscured by clouds or affected by strong aerosols should be weakened, while frames after the clouds have cleared are usually more informative. Therefore, attention aggregation is introduced in the time dimension and multiplied by the soft weights derived from the quality mask:

[0089] ;

[0090] in, and Depend on Linear mapping yields, The temporal aggregation feature is obtained by temporal smoothing of the quality mask to mitigate mask noise.

[0091] ;

[0092] S3.1-6. Applying one-dimensional convolution along the time dimension to obtain a low-frequency trend representation:

[0093] ;

[0094] The significance of the changes was represented by time-variance type aggregation:

[0095]

[0096] The overall time representation is achieved by concatenating the aggregation features, trend branches, and significant change branches along the channel dimension:

[0097] ;

[0098] S3.1-7, To Lightweight normalization and optional channel recalibration are performed to bring its statistics to the same order of magnitude as the spatial and spectral branches. This tensor is then registered in the joint feature cache.

[0099] .

[0100] Preferably, step S4 is as follows:

[0101] S4.1-1, The joint registration image and quality mask of S2 are denoted as:

[0102] ;

[0103] The incident angle and azimuth angle are calculated from the imaging geometry and encoded into a viewpoint embedding:

[0104] ;

[0105] Credibility of space generated by mask:

[0106] ;

[0107] S4.1-2, at the source band Zero-mean unit variance normalization:

[0108] ;

[0109] in, and Statistics on quality effective pixels, As a numerical stability constant, to explicitly introduce structural priors regarding edges and textures, the first-order spatial gradient is calculated:

[0110] ;

[0111] And based on this, a structure tensor is formed:

[0112] ;

[0113] in, Reflects the strength of the dominant edge. The texture fineness is represented, and the initial spatial tensor is obtained by splicing them together:

[0114] ;

[0115] S4.1-3, Define the perspective response and normalize the source:

[0116] ;

[0117] ;

[0118] Based on this, the initial image for multi-view linear fusion is obtained:

[0119] ;

[0120] S4.1-4, Constructing a two-dimensional convolution input:

[0121] ;

[0122] Definition of the first Two-dimensional convolutional unit:

[0123] ;

[0124] in It is a non-linear activation function. ;

[0125] S4.1-5, Residual blocks using void ratio rotation:

[0126] ;

[0127] in By alternating blocks, a hollow spatial pyramid is formed to expand the receptive field without downsampling;

[0128] S4.1-6, Multiple Predicting deformable migration fields With weight Bilinear interpolation is performed at the convolution sampling positions:

[0129] ;

[0130] S4.1-7, Assume the low-frequency structural branch is a pair The smoothed residual convolution stack is denoted as The high-frequency texture branch is a convolutional stack under high-pass prior constraints, denoted as... The gating coefficients are generated by shallow convolution and then passed through... Compression, represented as:

[0131] ;

[0132] The final spatial fusion is represented as:

[0133] ;

[0134] S4.1-8. Apply consistency regularization to the spatial response of source statistics:

[0135] ;

[0136] in, For the first The spatial response obtained from the source's separate pre-transmission and Standardize the channels and register them as spatial branch outputs:

[0137] .

[0138] Preferably, step S5 is as follows:

[0139] S5.1-1, in pixels At this point, the registration outputs of S2 are stacked by band to form a source-sensing spectral vector:

[0140] ;

[0141] Define spectral weights in conjunction with the quality mask:

[0142] ;

[0143] Robust weighted aggregation was performed on the source dimension to obtain the pixel-level observation spectrum:

[0144] ;

[0145] S5.1-2, Perform zero-mean unit variance normalization in the band dimension:

[0146] ;

[0147] in and Obtained from quality effective pixel statistics, The numerical stability constant is used as the spectral baseline by performing a smooth baseline fitting along the wavelength axis. The detrended spectrum is obtained:

[0148] ;

[0149] S5.1-3, with As input, let the first... The length of the layer spectral convolution kernel is Calculation in band dimension:

[0150] ;

[0151] in, For band axis neighborhood index, Nonlinear activation;

[0152] S5.1-4. Introduce band-dimensional dilated convolution without downsampling and stack it in the form of residuals:

[0153] ;

[0154] in, The band expansion rate is determined by the increasing sequence of residual blocks.

[0155] S5.1-5. Several branches are set in parallel on the one-dimensional convolution backbone, with the branch kernel length and dilation rate being respectively... Output of each branch First, channel compression is performed using pointwise one-dimensional convolution, and then weighted fusion is performed using learnable weights:

[0156] ;

[0157] S5.1-6, To Channel statistics are obtained by performing a global average across the band dimensions. Then, a weighted vector is obtained through two layers of perceptrons. And superimposed band weights obtained by smoothing with quality and band-level signal-to-noise ratio. The weighted representation is obtained as follows:

[0158] ;

[0159] S5.1-7. Apply L2 smoothing regularization to the band dimension:

[0160] ;

[0161] The spectral characterization is then mapped to a low-dimensional bottleneck to reduce the number of parameters and alignment costs.

[0162] ;

[0163] To preserve local geometric relationships, Apply approximately orthogonal constraints:

[0164] ;

[0165] S5.1-8, To Perform lightweight standardization and optional channel recalibration to bring its statistics to the same order of magnitude as the temporal and spatial branches, and finally record it as the spectral branch output:

[0166] .

[0167] Preferably, the step in S6 of using an attention mechanism to adaptively weight and fuse temporal, spatial, and spectral features to generate a joint feature tensor is as follows:

[0168] S6.1-1 Mapping the three branches to a unified channel dimension To avoid imbalances in the magnitude of integration:

[0169] ;

[0170] in, Layer normalization is adopted;

[0171] S6.1-2 Constructing a two-dimensional location code With relative time encoding Injecting three features additively:

[0172] ;

[0173] S6.1-3. Using the time branch as the query and the spatial and spectral branches as the key values, construct mutual attention:

[0174] ;

[0175] ;

[0176] Calculate attention weights and perform weighted convergence:

[0177] ;

[0178] Similarly, construct symmetric units where space is the query and spectrum is the query. ;

[0179] S6.1-4, Expanding mutual attention to The individual components are then stitched together in the first dimension, followed by residual fusion using gating coefficients:

[0180] ;

[0181] in, Indicates the first Under the attention of the head, time branches are used as Mutual attention output derived from spatial and spectral branches This means to focus all attention on the head. The outputs are concatenated in the first dimension to form a convergent feature. For element-wise multiplication, The features are that the temporal, spatial, and spectral paths are aligned to the same number of channels. It is a pixel-level gating system, obtained by concatenating three features and then passing them through a small multilayer perceptron. This system is used to balance the introduced information with the original mode. The spatial and spectral sides are obtained in the same form. and ;

[0182] S6.1-5. Construct a confidence scalar to suppress anomalous contribution signals, and set spatial weights. Time weight , graded weights As estimated in the preceding steps, three confidence levels are formed using aggregation functions and then pruned to a stable range:

[0183] ;

[0184] And confidence modulation is applied to the three-way fusion result:

[0185] ;

[0186] S6.1-6, Pixel-by-pixel summation and normalization, followed by channel recalibration using a feedforward network and establishment of a second residual path:

[0187] ;

[0188] S6.1-7. To improve generalization ability and constrain the spread of attention in irrelevant regions, consistency and entropy regularization are designed:

[0189] ;

[0190] Training two regularization methods together with the main task loss can significantly improve robustness in complex environments;

[0191] S6.1-8. Perform lightweight normalization and optional channel compression on the fusion result to obtain the joint feature tensor that can be directly used by the subsequent reconstruction network:

[0192] .

[0193] Preferably, step S7 is as follows:

[0194] S7.1-1, During Reconstruction On the unified reference grid, by joint feature tensor The generation space magnification is High-resolution reconstructed images Let the observation degradation operator be... This includes the point spread function, modulation transfer function, and downsampling raster mapping for the source and band; let the spectral response mapping be... Used to project from the target spectral reference to the sensor band, in order to Represents high-resolution pixel coordinates. Indicates the coordinates of the original reference raster cell;

[0195] S7.1-2. Using a conditional denoising diffusion probability model, define a forward Markov process:

[0196] ;

[0197] Noise schedule Monotonically increasing, making , During training, real high-resolution targets sampling:

[0198] ;

[0199] S7.1-3, with As a condition, the upstream network output co-modulates the denoising network at both the channel dimension and the diffusion time. Let the conditional encoder be:

[0200] ;

[0201] Time embedding as Employing characteristic linear modulation injection:

[0202] ;

[0203] in This represents the intermediate features of the denoising network. , Adaptive conditional control for different diffusion steps is achieved by predicting with a small perceptron.

[0204] S7.1-4. The denoising network adopts an encoder-decoder structure with cross-layer skip connections, and performs the denoising ratio increase step by step in the spatial domain. Upsampling, making This represents the pixel rearrangement upsampling operator, with the scale-increasing sequence being... , No. Level decoding output:

[0205] ;

[0206] in For cascading, Features from the same layer of the encoder are used to insert conditional modulation and residual convolutional blocks at each scale to ensure that details are recovered step by step while spectral statistics remain stable.

[0207] S7.1-5, Learning the inverse process using noise prediction:

[0208] ;

[0209] The training loss is calculated using weighted noise regression.

[0210] ;

[0211] in, By adjusting the relative importance of each diffusion step, a prediction correction strategy can be used in the inference phase to reduce the number of steps;

[0212] S7.1-6. To ensure that the reconstruction results are consistent with the observations from each source under the physical degradation model, a re-projection is performed for any source and band:

[0213] ;

[0214] in, This represents the high-resolution spectral image at the target time for reconstruction. For spectral response mapping, It is an observation degeneracy operator. This is a simulated observation by re-projection;

[0215] and by quality weight Constraints and Real Observations Consistency:

[0216] ;

[0217] in, These are real, registered observations. Indicates the weight of quality and credibility;

[0218] S7.1-7. To maintain the authenticity of the spectral shape and the reliability of subsequent classification and inversion, spectral angles and relative information divergence are used:

[0219] ;

[0220] in, In the band The normalized spectral values, The target spectrum is used as a reference or pseudo-reference.

[0221] S7.1-8. To improve texture perception quality and suppress oversmoothing, structure and perception terms are introduced:

[0222] ;

[0223] in, For a downsampled reconstruction map that matches the reference scale, It is the initial image obtained by spatial branching linear pre-fusion. This represents a fixed perceptual feature extraction layer;

[0224] Further, a local contrast consistency term is introduced to stabilize the edges:

[0225] ;

[0226] in, Represents the spatial gradient operator. express Paradigm;

[0227] S7.1-9. Based on the above, the training objective is:

[0228] ;

[0229] Weights By using a validation set search, we can ensure physical consistency and spectral fidelity while improving perception quality and edge sharpness.

[0230] S7.1-10, The reasoning stage uses piecewise cosine reduction to shorten the schedule and number of steps. To enhance conditional dependencies, a classifier free guidance is introduced:

[0231] ;

[0232] in, To guide the intensity, The condition is discarded and substituted into the inverse mean expression for iterative generation. ;

[0233] S7.1-11. After generation, perform multi-scale physical consistency acceptance:

[0234] ;

[0235] in, Set a threshold for the reference or pseudo-reference spectrum. and If the condition is not met, the diffusion steps are regressed by several steps and the value is adaptively increased. and A correction resampling step is performed, ultimately outputting a reconstructed image with high spatial and temporal resolution and high spectral fidelity. It is also registered for use by downstream classification and inversion tasks.

[0236] Beneficial effects:

[0237] 1. This invention, through the synergistic optimization of the spatiotemporal spectrum integrated reconstruction framework and the physical-data dual constraint mechanism, jointly models the time, space and spectral domains under a unified spatiotemporal benchmark. Combined with gain and bias calibration, atmospheric correction and spectral response function alignment and mean-variance normalization, it significantly improves the radiometric and spectral shape comparability of multi-load data, laying a stable benchmark for subsequent fusion.

[0238] 2. Joint registration integrating rigorous imaging geometry and depth feature matching, and introducing non-rigid deformation and phase-related refinement, stably achieves sub-pixel-level alignment and suppresses geometric errors caused by off-axis, parallax, and pose perturbations. In representation learning, 3D convolution combines temporally dilated residual blocks with quality-modulated temporal attention aggregation to effectively acquire long-short-term dependencies and resist cloud and fog anomalies. On the 2D convolution side, viewpoint-weighted pre-fusion, hollow spatial pyramids, and deformable convolution balance receptive field and high-frequency sensitivity, and structure / texture dual-branch gating reduces over-smoothing and pseudo-textures. On the 1D convolution side, a multi-scale spectral pyramid and channel attention are constructed along the band axis, and smoothing and approximately orthogonal constraints are applied to maintain key absorption bands and interband correlations while appropriately reducing dimensionality, ensuring spectral authenticity.

[0239] 3. The three-domain features are adaptively fused within the same semantic scale through position / relative time encoding and bidirectional mutual attention. The source, temporal sequence and band-level confidence are used to suppress the interference of low-quality samples on attention weights, which significantly enhances the robustness of the fusion. Finally, the conditional diffusion super-resolution network combines observation feedback consistency, spectral angle and information divergence, structure-perception and other joint losses to achieve both high spatial resolution and high temporal resolution while maintaining high spectral fidelity and reducing oversharpening, false colors and geometric drift.

[0240] 4. This method has good cross-platform portability and engineering scalability, and can significantly improve the accuracy and robustness of classification, inversion and change detection in downstream tasks such as urban fine mapping, agricultural monitoring, ecological assessment, disaster emergency response and battlefield situational awareness. Attached Figure Description

[0241] Figure 1 This is a schematic diagram of the overall process of the present invention;

[0242] Figure 2 This is a flowchart of multi-source data acquisition and radiation and atmospheric preprocessing in an embodiment of the present invention;

[0243] Figure 3 This is a schematic diagram illustrating the working principle of unified benchmark and joint registration in this embodiment of the invention;

[0244] Figure 4 This is a schematic diagram of temporal branching 3D convolution and temporal attention structure in an embodiment of the present invention;

[0245] Figure 5 This is a schematic diagram of spatial branching multi-view coding and hollow pyramid structure in an embodiment of the present invention;

[0246] Figure 6 This is a schematic diagram of the spectral branch one-dimensional convolution pyramid and channel attention structure in an embodiment of the present invention;

[0247] Figure 7 This is a flowchart of cross-modal attention fusion and gated modulation collaborative optimization in an embodiment of the present invention;

[0248] Figure 8 This is a flowchart illustrating the joint constraint process of conditional diffusion super-resolution reconstruction and observation re-projection in an embodiment of the present invention. Detailed Implementation

[0249] The invention will now be further described with reference to the accompanying drawings.

[0250] like Figure 1 As shown, a spatiotemporal spectral joint super-resolution reconstruction method based on giant remote sensing constellations includes the following steps:

[0251] S1. Obtain time-series, different perspectives and multispectral observation data of the same area from multiple heterogeneous satellites, and perform gain-bias calibration, solar geometric correction and atmospheric correction on the images from each source in sequence to obtain comparable standardized input;

[0252] S2. Construct a joint registration model under a unified spatial projection and reference time, and combine rigorous imaging geometry and feature matching to accurately align multi-source images to the sub-pixel level;

[0253] S3. Construct a time window around the target reconstruction time, use a three-dimensional convolution and dilated residual structure to model the registered time series, and introduce time attention and quality weights to extract the continuity and abrupt change features of the target's evolution over time.

[0254] S4. Spatial modeling of multi-view information on a unified grid: First, perform radiation normalization and structural prior extraction, then use two-dimensional convolution, hollow spatial pyramid and deformable convolution to fuse multi-view differences and enhance contour structure and fine-grained texture.

[0255] S5. Along the band dimension, a one-dimensional convolution and multi-scale dilation strategy is adopted to mine interband correlations and key absorption bands from multimodal spectra. The channel attention and bottleneck mapping are used to achieve appropriate dimensionality reduction, maintaining the authenticity of the spectrum and the compactness of the representation.

[0256] S6. Introducing a cross-modal attention mechanism, adaptive weighting and mutual attention interaction are performed on the temporal, spatial and spectral features. Under the constraints of confidence modulation and gating fusion, a single joint feature tensor is generated, which gathers complementary information from multiple sources and domains.

[0257] S7. Input the joint features into the diffusion super-resolution network, and optimize it by combining the U-Net structure with stepwise upsampling and joint loss. The output is a target reconstruction image with high spatial resolution, high temporal resolution and high spectral fidelity, which meets the accuracy requirements of subsequent classification and inversion.

[0258] Furthermore, such as Figure 2 As shown, step S1 is as follows:

[0259] S1.1-1 Obtaining the same region at different times from a giant remote sensing constellation Different sources Different spectral bands The original image digital value (DN), i.e., the first Time, Number Source, No. Band pixels are: ;

[0260] S1.1-2. For each source and its respective band, the DN is converted to top atmospheric radiance using calibrated gain and bias. :

[0261] ;

[0262] in, For gain, For bias;

[0263] S1.1-3. For the reflectivity bands of visible light and near-infrared light, considering the Earth-Sun distance and the angle of solar incidence, calculate:

[0264] ;

[0265] in, For the first The distance between the Earth and the Sun, For band Solar irradiance, The solar zenith angle;

[0266] S1.1-4. For the thermal infrared band, the brightness temperature is inverted according to Planck's law:

[0267] ;

[0268] in, , This is the Planck constant term for this band;

[0269] S1.1-5. Adopting a parametric radiation transmission architecture, for Correction is performed to obtain the surface reflectance. :

[0270] ;

[0271] in, For path radiation, and These represent the upward and downward atmospheric transmittance, respectively. This refers to the downward atmospheric irradiance.

[0272] S1.1-6. To eliminate the differences in the spectral response functions (SRF) of different sensors, all source data are mapped to a unified target spectral reference, and any target band... The equivalent radiance is expressed as:

[0273] ;

[0274] in Let be the spectral response function of the target band. Obtained by source band interpolation;

[0275] At the statistical level, mean-variance normalization is performed using the reference sensor 'ref' as a benchmark, letting... :

[0276] ;

[0277] in , For the sensor in band The mean and standard deviation of surface reflectance, , To reference the corresponding statistics of the sensor in the same band, all statistics are on the quality mask. Calculated on the specified valid pixels;

[0278] S1.1-7. Using orbital and attitude parameters, combined with a rational polynomial camera model and a rigorous geometric model, orthorectify each image and project it onto a unified map projection and pixel grid; simultaneously generate quality masks for clouds, shadows, water bodies, and saturation. This forms a standardized input dataset for subsequent joint registration and feature extraction:

[0279] .

[0280] Furthermore, such as Figure 3 As shown, step S2 is as follows:

[0281] S2.1-1 First, determine a unified map projection and pixel grid as the reference system. And select the alignment reference time. The source is pixel coordinates 3D points on the Earth's surface and , No. Source in time The observation is recorded as Imaging geometry is expressed using RPC:

[0282] ;

[0283] in, The interior and exterior orientation parameters are obtained through time-series interpolation. Used for time normalization;

[0284] S2.1-2, based on and Orthophotos from various sources get ,select For reference, then:

[0285] ;

[0286] Set initial translation for non-reference sources Perform coarse alignment:

[0287] ;

[0288] S2.1-3, Employing a shared feature extraction network Calculate the feature map:

[0289] ;

[0290] Construct a set of candidate corresponding points between the reference and the images to be registered:

[0291] ;

[0292] in Obtained by normalized cross-correlation search:

[0293] ;

[0294] Using the random sampling consensus method from Estimate initial single pressure ;

[0295] S2.2-4, Using B-splines to represent the deformation field Modeling, For the control point parameters, the composite geometric mapping is:

[0296] ;

[0297] in, For the reference pixel coordinates, This is the initial geometric mapping obtained in the previous stage. For B-spline parameterized non-rigid incremental deformation fields, This is the control parameter vector for the corresponding viewpoint. It is pass Resampled image;

[0298] Construct the joint energy function:

[0299] ;

[0300] in, For reference only. for Robust loss, As a projection operator, it projects three-dimensional surface points. Projected to view At any moment pixel coordinates, For the corresponding projection model parameters, Yes The first-order spatial gradient, and ;

[0301] S2.2-5, Constructing a Scale Pyramid At every scale Linearize the data items and iterate using the Gauss-Newton method:

[0302] ;

[0303] in for about Jacobi, For approximate Hessian matrix, The gradient vector is used to update the gradient: ;

[0304] S2.2-6, in the block domain Residual translation estimated using phase correlation :

[0305] ;

[0306] in, For reference, at the base time Block-domain image, Sources to be aligned at time The block-domain image has been resampled through a preceding mapping. It is the two-dimensional discrete Fourier transform and inverse transform;

[0307] And perform translation compensation on the mapping:

[0308] ;

[0309] S2.2-7, when Introduce a linear slowly varying motion model:

[0310] ;

[0311] Based on this, the time correction mapping is defined as follows:

[0312] ;

[0313] S2.2-8. Resample the optimized geometric mapping using cubic convolution interpolation to generate a registration result consistent with the reference raster:

[0314] ;

[0315] Then, using the root mean square error of the reprojection of projected coordinates and the normalized cross-correlation as dual indicators, and under the condition that... ;

[0316] in, This represents the root mean square error of the projected coordinate reprojection. It is the geometric error threshold. Represents the normalized cross-correlation coefficient. This is the correlation threshold;

[0317] S2.2-9. The registration results from various sources are collected together with the quality mask and used as a unified input for subsequent spatiotemporal spectral feature extraction and super-resolution reconstruction.

[0318] .

[0319] Furthermore, such as Figure 4 As shown, step S3 is as follows:

[0320] S3.1-1, Around the Reconstruction Moment Set length as Symmetric time window Using the registered image and quality mask of S2 as unified input:

[0321] ;

[0322] To enhance robustness to low-quality frames, soft quality weights are introduced:

[0323] ;

[0324] And based on this, a weighted time stack is formed:

[0325] ;

[0326] S3.1-2. To eliminate amplitude bias across sources and time periods, zero-mean unit variance standardization is performed within each source and band:

[0327] ;

[0328] in, For the time stack with quality weights, at the target time... time window Data from internal tissues, It is the normalized temporal tensor. and exist Regarding the source With band Calculate the mean and standard deviation. To determine the stability constant, calculate the first-order time difference:

[0329] ;

[0330] in, yes The forward first-order time difference, and ;

[0331] The two are concatenated along the channel dimension and used as input for a 3D convolution to obtain a composite tensor. :

[0332] ;

[0333] in, This indicates a splicing operation along the channel dimension. Indicates at time The input tensor of the 3D convolution, For each source With band In the window By normalizing the intensity tensor, the amplitude differences between the source and time step are suppressed. That is The first-order temporal difference explicitly characterizes short-term dynamics, highlighting clues of motion and abrupt changes;

[0334] S3.1-3, Let the first... Layer input is Using core size 3D Convolution and Nonlinear Activation:

[0335] ;

[0336] in , It is a non-linear activation function. ;

[0337] S3.1-4. To expand the temporal receptive field and maintain gradient stability, stack residual blocks containing temporally dilated convolutions:

[0338] ;

[0339] Among them, the time dilation rate As the layers increase, to cover The long and short contexts;

[0340] S3.1-5. Aggregation is performed in the temporal dimension using content-based attention, and smooth quality weights are combined to suppress the influence of abnormal frames. The attention weights are calculated as follows:

[0341] ;

[0342] in, and Depend on Linear mapping yields, Obtained by temporal smoothing from a quality mask, the temporal aggregation feature is as follows:

[0343] ;

[0344] S3.1-6, Along the time dimension Apply one-dimensional convolution to extract low-frequency trends:

[0345] ;

[0346] And the significance of the change was measured using variance-based aggregation:

[0347] ;

[0348] Will , , By splicing along the channel dimensions, we obtain:

[0349] ;

[0350] S3.1-7. To align numerically with the spatial and spectral branches, for Perform lightweight normalization and optional channel recalibration, and register them to the joint feature cache:

[0351] .

[0352] Furthermore, such as Figure 5 As shown, step S4 is as follows:

[0353] S4.1-1. Explicitly characterize source differences within a uniform raster, using the joint registration output of S2 as the basic input for spatial branching:

[0354] ;

[0355] Pixel-level incident angle and azimuth angle are extracted from imaging geometry and denoted as follows: and Constructing perspective embedding:

[0356] ;

[0357] Credibility of space generated by mask:

[0358] ;

[0359] S4.1-2. To avoid texture contrast imbalance caused by inconsistent radiation scales, standardization should be performed first:

[0360] ;

[0361] in, and Statistics on quality effective pixels, To find the numerical stability constant, calculate the first-order gradient:

[0362] ;

[0363] Forming a structure tensor:

[0364] ;

[0365] Will , , and perspective embedding By splicing along the channel dimension, the initial tensor of the spatial branch is formed. ;

[0366] S4.1-3. Before nonlinear convolution, perform a weighted average based on the degree of off-axis rotation and quality:

[0367] ;

[0368] ;

[0369] Based on this, the initial image for multi-view linear fusion is obtained:

[0370] ;

[0371] S4.1-4, will , , , With coordinate channels The combination is:

[0372] ;

[0373] Definition of the first Two-dimensional convolutional unit:

[0374] ;

[0375] in, , It is a nonlinear function;

[0376] S4.1-5. Without downsampling, construct a void space pyramid using multiple void ratios:

[0377] ;

[0378] in By alternating blocks, a hollow space pyramid is formed;

[0379] S4.1-6, based on Predicting deformable migration fields With weight Deformable convolution is achieved using bilinear interpolation:

[0380] ;

[0381] S4.1-7, with Divide and conquer modeling low-frequency structures With high frequency texture Generated by shallow convolution and then processed compression to obtain the gating coefficient Through adaptive fusion:

[0382] ;

[0383] S4.1-8. Suppress off-axis pseudo-textures by constraining consistency through source statistics:

[0384] ;

[0385] in, For the first The spatial response obtained from the separate pre-transmission source will Standardize the channels and register them as spatial branch outputs:

[0386] .

[0387] Furthermore, such as Figure 6 As shown, step S5 is as follows:

[0388] S5.1-1. Stack the registration outputs of S2 at each pixel according to the spectral band to form the source-sensing spectral vector:

[0389] ;

[0390] Spectral weights are defined based on the quality mask:

[0391] ;

[0392] And the source dimensional weights converge into a pixel-level observation spectrum:

[0393] ;

[0394] S5.1-2, Perform zero-mean unit variance normalization in the band dimension:

[0395] ;

[0396] in and Obtained from quality effective pixel statistics, This is the numerical stability constant, and then a smooth baseline is fitted along the wavelength axis. And remove trends:

[0397] ;

[0398] S5.1-3, with As input, let the first... The length of the layer spectral convolution kernel is Calculation in band dimension:

[0399] ;

[0400] in, For band axis neighborhood index, Nonlinear activation;

[0401] S5.1-4, Introduce band-dimensional dilated convolution and stack it in residual form:

[0402] ;

[0403] in, The band expansion rate is determined by the increasing sequence of residual blocks.

[0404] S5.1-5. Construct several branches in parallel, with the branch kernel length and expansion rate being respectively... Output of each branch First, channel compression is performed using pointwise one-dimensional convolution, and then weighted fusion is performed using learnable weights:

[0405] ;

[0406] S5.1-6, To Channel statistics are obtained by performing a global average across the band dimensions. Then, a weighted vector is obtained through two layers of perceptrons. And superimposed band weights obtained by smoothing with quality and band-level signal-to-noise ratio. The weighted representation is obtained as follows:

[0407] ;

[0408] S5.1-7, First apply spectral smoothing regularization:

[0409] ;

[0410] Then the spectral characterization is mapped to a low-dimensional bottleneck:

[0411] ;

[0412] And on Apply approximately orthogonal constraints:

[0413] ;

[0414] S5.1-8, To Perform lightweight normalization and optional channel recalibration to register it as a spectral branch output:

[0415] .

[0416] Furthermore, such as Figure 7 As shown, step S6 is as follows:

[0417] S6.1-1, Put Projecting to the same semantic and channel scale avoids amplitude imbalance dominating fusion:

[0418] ;

[0419] in, Layer normalization is used to suppress inter-batch drift;

[0420] S6.1-2. To prevent spatial mismatch and phase loss over time, grid position encoding and relative window phase are added to each branch:

[0421] ;

[0422] S6.1-3. Using the time branch as the query and the spatial and spectral branches as the key values, construct mutual attention:

[0423] ;

[0424] ;

[0425] Calculate attention weights and perform weighted convergence:

[0426] ;

[0427] Similarly, we define the space as the query and the spectrum as the symmetric unit of the query. This enables bidirectional supplementation of information from three channels.

[0428] S6.1-4, in Each subspace captures complementary relationships in parallel and performs adaptive fusion with cell-level gating:

[0429] ;

[0430] in, Indicates the first Under the attention of the head, time branches are used as Mutual attention output derived from spatial and spectral branches This means to focus all attention on the head. The outputs are concatenated in the first dimension to form a convergent feature. For element-wise multiplication, The features are that the temporal, spatial, and spectral paths are aligned to the same number of channels. It is a pixel-level gating system, obtained by concatenating three features and then passing them through a small multilayer perceptron. This system is used to balance the introduced information with the original mode. The spatial and spectral sides are obtained in the same form. and ;

[0431] S6.1-5, Summarize the quality indicators from S1 to S5 to form branch confidence levels:

[0432] ;

[0433] And confidence modulation is applied to the three-way fusion result:

[0434] ;

[0435] S6.1-6. Sum the features after three-way confidence modulation, normalize them, and express them using feedforward residual boosting:

[0436] ;

[0437] S6.1-7. To improve generalization and constrain attention diffusion, branch consistency and attention entropy regularization are introduced:

[0438] ;

[0439] S6.1-8. Finally, after lightweight standardization, a unified interface for reconstructing networks is obtained:

[0440] .

[0441] Furthermore, such as Figure 8 As shown, step S7 is as follows:

[0442] S7.1-1, During Reconstruction On the unified reference grid, by joint feature tensor The generation space magnification is High-resolution reconstructed images Let the observation degradation operator be... Mapped to the spectrum ,by Represents high-resolution pixel coordinates. Indicates the coordinates of the original reference raster cell;

[0443] S7.1-2. Using a conditional denoising diffusion probability model, define a forward Markov process:

[0444] ;

[0445] set up During training, real high-resolution targets sampling:

[0446] ;

[0447] S7.1-3, with As a condition, the upstream network output co-modulates the denoising network at both the channel dimension and the diffusion time. Let the conditional encoder be:

[0448] ;

[0449] Time embedding as Injection using characteristic linear modulation:

[0450] ;

[0451] S7.1-4. Employs a cross-layer jump-connected encoder-decoder structure, enabling... For pixel rearrangement, the upsampling sequence is , No. Level decoding output:

[0452] ;

[0453] in For cascading, Features from the same layer of the encoder;

[0454] S7.1-5, Learning the inverse process using noise prediction:

[0455] ;

[0456] The training loss is calculated using weighted noise regression.

[0457] ;

[0458] S7.1-6. Introduce an observation consistency term to perform backcasting for any source and band:

[0459] ;

[0460] in, This represents the high-resolution spectral image at the target time for reconstruction. For spectral response mapping, It is an observation degeneracy operator. This is a simulated observation by re-projection;

[0461] and by quality weight Constraints and Real Observations Consistency:

[0462] ;

[0463] in, These are real, registered observations. Indicates the weight of quality and credibility;

[0464] S7.1-7, Using a combined constraint of spectral angle and relative information divergence:

[0465] ;

[0466] in, In the band The normalized spectral values, The target spectrum is used as a reference or pseudo-reference.

[0467] S7.1-8. To improve texture perception quality and suppress oversmoothing, structure and perception terms are introduced:

[0468] ;

[0469] in, For a downsampled reconstruction map that matches the reference scale, It is the initial image obtained by spatial branching linear pre-fusion. This represents a fixed perceptual feature extraction layer;

[0470] Further, a local contrast consistency term is introduced to stabilize the edges:

[0471] ;

[0472] in, Represents the spatial gradient operator. express Paradigm;

[0473] S7.1-9. Based on the above, the training objective is:

[0474] ;

[0475] Weights Determined through validation set search;

[0476] S7.1-10, Adopt To shorten the schedule and enhance conditional dependencies, a classifier-driven approach is introduced:

[0477] ;

[0478] in, To guide the intensity, This indicates that the condition has been discarded.

[0479] S7.1-11. After generation, perform multi-scale physical consistency acceptance:

[0480] ;

[0481] in, Set a threshold for the reference or pseudo-reference spectrum. and If the condition is not met, the diffusion steps are regressed by several steps and the value is adaptively increased. and A correction resampling step is performed, ultimately outputting a reconstructed image with high spatial and temporal resolution and high spectral fidelity. It is also registered for use by downstream classification and inversion tasks.

Claims

1. A spatiotemporal spectral joint super-resolution reconstruction method based on giant remote sensing constellations, characterized in that, Includes the following steps: S1. Collect time-series, multi-view, and multispectral data of the same region from different satellites, and perform radiometric calibration and atmospheric correction; S2. Use a joint registration model to align multi-source data to the sub-pixel level under a unified spatiotemporal coordinate system; S3. Use 3D convolution to extract the pattern of target changes over time from the time series; S4. Use 2D convolution to fuse multi-view information and highlight structural and texture details; S5. Use 1D convolution to extract effective features from multimodal spectra and reduce dimensionality appropriately; S6. Use an attention mechanism to weight and fuse temporal, spatial, and spectral features into a single feature tensor. S6.1-1. Map the representations of the three branches to the same semantic scale, making subsequent attention weights comparable and weighted within the same metric space, including the temporal branch. Emphasizing dynamic consistency, spatial branching Emphasis on geometric and textural details, spectral branching Emphasizing interband absorption, spectral stability, and the fact that these three factors naturally differ in distribution and variance scale, linear projection and normalization are used to align them to the channel dimension. This avoids the magnitude of a particular branch dominating subsequent fusion: ; in, Layer normalization is employed to reduce inter-batch drift and improve training stability; S6.1-2 Constructing a two-dimensional location code To express grid coordinate relationships and construct relative time encoding To express the phase of the current target moment within a local time window, it is then injected into each branch in an additive manner: ; S6.1-3. In actual observations, time clues often need to draw supplementary information from spatial and spectral clues; Using the time branch as the query and the spatial and spectral branches as the key-value pairs, we construct mutual attention: ; ; Calculate attention weights and perform weighted convergence: ; Similarly, we define the space as the query and the spectrum as the symmetric unit of the query. , , , ; S6.1-4, Expanding mutual attention to The data is then stitched together in the head dimension, followed by residual fusion using gating coefficients. The gating is determined by the shallow perceptron based on the three inputs at the current pixel, in order to balance external information with prior knowledge of the current modality and prevent overfitting or information overload. , ; in, Indicates the first Under the attention of the head, time branches are used as Mutual attention output derived from spatial and spectral branches This means to focus all attention on the head. The outputs are concatenated in the first dimension to form a convergent feature. For element-wise multiplication, The features are that the temporal, spatial, and spectral paths are aligned to the same number of channels. It is a pixel-level gating system, obtained by concatenating three features and then passing them through a small multilayer perceptron. This system is used to balance the introduced information with the original mode. The spatial and spectral sides are obtained in the same form. and ; S6.1-5, the quality metrics from S1 to S5 reflect imaging noise, cloud shadow occlusion, off-axis geometry, and band-level signal-to-noise ratio differences. A confidence scalar is constructed to suppress anomalous contribution signals, and spatial weights are set. Time weight , graded weights As estimated in the preceding steps, three confidence levels are formed using aggregation functions and then pruned to a stable range: ; And confidence modulation is applied to the three-way fusion result: ; S6.1-6. After completing mutual attention, gating, and confidence modulation, the three information streams need to be converged into a unified representation while maintaining numerical stability and smooth gradients. The features of the three streams are summed pixel by pixel and normalized. Then, the feedforward network is used to recalibrate the channels and establish a second residual path to improve the nonlinear expressive power without destroying the learned scale relationship. ; S6.1-7, Design Consistency and Entropy Regularization: The former encourages different sources to form similar attention distributions for the same location, while the latter encourages a sharper allocation of attention to highlight key matches. ; Two regularization methods are trained together with the main task loss, which significantly improves robustness in complex environments. S6.1-8 Finally, lightweight normalization and optional channel compression are applied to the fusion result to obtain the joint feature tensor that can be directly used by the subsequent reconstruction network: ; S7. The fused features are input into the conditional diffusion super-resolution network, and the reconstruction is completed through multi-scale upsampling and joint loss optimization to generate high-resolution images while maintaining spectral fidelity to meet the accuracy requirements of subsequent classification and inversion.

2. The spatiotemporal spectral joint super-resolution reconstruction method based on giant remote sensing constellations according to claim 1, characterized in that, The steps in S1 are as follows: S1.1-1 Obtaining the same region at different times from a giant remote sensing constellation Different sources Different spectral bands Original image digital values That is, the first Time, Number Source, No. Band pixels are ; S1.1-2. For each sensor and each band, use its calibration coefficients to... Converted to top atmospheric radiance : ; in, For gain, For bias; S1.1-3. For the visible and near-infrared reflectivity bands, considering the Earth-Sun distance and the angle of solar incidence, calculate... Reflectivity: ; in, For the first The distance between the Earth and the Sun, For band Solar irradiance, The solar zenith angle; S1.1-4. For the thermal infrared band, the radiance is inverted into brightness temperature according to Planck's law: ; in, , This is the Planck constant term for this band; S1.1-5, Using a parametric radiative transfer architecture for... Correction is performed to obtain the surface reflectance. : ; in, For path radiation, and These represent the upward and downward atmospheric transmittance, respectively. This refers to the downward atmospheric irradiance. S1.1-6. Map all source data to a unified target spectral reference, for any target band. The equivalent radiance is expressed as: ; in Let be the spectral response function of the target band. Obtained by source band interpolation; At the statistical level, mean-variance normalization is performed using the reference sensor "ref" as a benchmark, letting... : ; in , For the sensor in band The mean and standard deviation of surface reflectance, , To reference the corresponding statistics of the sensor in the same band, all statistics are on the quality mask. Calculated on the specified valid pixels; S1.1-7. Using orbit and attitude parameters, combined with a rational polynomial camera model and a rigorous geometric model, orthorectify each image and project it onto a unified map projection and pixel grid; simultaneously generate quality masks for clouds, shadows, water bodies, and saturation. This forms a standardized input dataset for subsequent joint registration and feature extraction: 。 3. The spatiotemporal spectral joint super-resolution reconstruction method based on giant remote sensing constellations according to claim 1, characterized in that, The steps in S2 are as follows: S2.1-1. Select a unified map projection and pixel grid as the geographic reference system. Alignment reference time is The source is pixel coordinates are The three-dimensional points on the Earth's surface are and , No. Source in The image of the moment is The imaging geometry uses a rational polynomial camera model: ; in, The exterior and interior orientation parameters are obtained through element interpolation. Used for time unification; S2.1-2, Utilizing and Orthorectify the images from each source onto the reference frame. get Select reference source The reference image is: ; The remaining sources are based on the initial displacement vector. Coarse alignment: ; S2.1-3, Employing a shared feature extraction network Calculate the feature map: ; Using normalized cross-correlation as the matching criterion, the most similar location is searched within the neighborhood of candidate points on the reference image to construct a set of candidate corresponding points. Implement grid homogenization constraints, retaining only the pair with the highest cross-correlation peaks in each grid; perform matching on the effective pixels of the quality mask to remove cloud shadows, water bodies, and saturated regions. ; The random sampling consensus method is used for minimum sampling, error assessment, and interior point set update. A linear equation is constructed using four pairs of matching to obtain the homography parameters and homography matrix. ; Error measurement uses symmetrical propagation error, and the interior point determination threshold is... : ; When the error of a certain pair of matches Less than When this happens, we classify it as an interior point and obtain the optimal set of interior points. Then, using weighted least squares in The initial homography matrix is ​​obtained by precise estimation. And define the initial geometric mapping: ; S2.2-4. Parameterization of non-rigid deformation fields using B-splines The control point parameters are The composite mapping is represented as: ; in, For the reference pixel coordinates, This is the initial geometric mapping obtained in the previous stage. For B-spline parameterized non-rigid incremental deformation fields, This is the control parameter vector for the corresponding viewpoint. It is pass Resampled image; Construct the joint energy function: ; in, For reference only. for Robust loss, As a projection operator, it projects three-dimensional surface points. Projected to view At any moment pixel coordinates, For the corresponding projection model parameters, Yes The first-order spatial gradient, and ; S2.2-5, Establish a scale pyramid index , in scale The data items are linearized using the Gauss-Newton method, and the parameters are updated accordingly. ; ; in for about Jacobi, For approximate Hessian matrix, The gradient vector; S2.2-6, in the block domain Residual translation estimated using phase correlation And update the mapping: ; ; in, For reference, at the base time Block-domain image, Sources to be aligned at time The block-domain image has been resampled through a preceding mapping. It is the two-dimensional discrete Fourier transform and inverse transform; S2.2-7. When the acquisition time is inconsistent with the reference time, i.e. A linear slowly varying motion model is introduced, followed by a time-corrected mapping, and finally, the joint energy is expressed as... replace Minimize: ; ; S2.2-8. The optimized geometric mapping is resampled using cubic convolution interpolation to generate a registered image consistent with the reference raster. Then, the root mean square error of the projected coordinate reprojection and the normalized cross-correlation are used as dual indicators, based on the following conditions: : ; in, This represents the root mean square error of the projected coordinate reprojection. It is the geometric error threshold. Represents the normalized cross-correlation coefficient. This is the correlation threshold; S2.2-9. Summarize the registration results and quality masks from all sources as unified input for subsequent spatiotemporal spectral feature extraction and super-resolution reconstruction: 。 4. The spatiotemporal spectral joint super-resolution reconstruction method based on giant remote sensing constellations according to claim 1, characterized in that, The steps in S3 are as follows: S3.1-1, First, reconstruct the timeline around the target. Build length is Symmetric time window The registration result output by S2 is used as the basic input and represented using unified notation as follows: ; Considering that cloud and shadow quality issues can disrupt temporal consistency, a controllable quality weight is introduced. ,in This is a fallback factor for missing measurements, used to prevent the entire time chain from breaking, thus resulting in a weighted time stack: ; S3.1-2, Perform zero-mean unit variance normalization within the source and band: ; in, For the time stack with quality weights, at the target time... time window Data from internal tissues, It is the normalized temporal tensor. and exist Regarding the source With band Calculate the mean and standard deviation. As a stability constant, the first-order time difference is calculated and used as the input channel along with the normalized intensity: ; in, yes The forward first-order time difference, and ; Ultimately and The channels are stitched together to form a three-dimensional convolutional input. : ; in, This indicates a splicing operation along the channel dimension. Indicates at time The input tensor of the 3D convolution, For each source With band In the window By normalizing the intensity tensor, the amplitude differences between the source and time step are suppressed. That is The first-order temporal difference explicitly characterizes short-term dynamics, highlighting clues of motion and abrupt changes; S3.1-3. Spatial texture and temporal evolution of remote sensing time series are tightly coupled. To couple them within the same operator, a three-dimensional convolution with a unified kernel is used. Let the first... Layer input is Core size The output will be: ; in , It is a non-linear activation function. ; S3.1-4, Stacking residual 3D convolutional blocks and introducing dilated convolution in the time dimension to exponentially expand the context coverage: ; Among them, the time dilation rate Incrementing by layer; S3.1-5. Introduce attention aggregation in the temporal dimension and multiply it with the soft weights derived from the quality mask, so that the contribution of the frame is simultaneously constrained by both "relevance to the current semantics" and "imaging quality reliability": ; in, and Depend on Linear mapping yields, The final temporal aggregation feature is obtained by temporal smoothing from a quality mask: ; S3.1-6. To structure the "long-term gradual change" and "short-term abrupt change" for subsequent module processing before fusion, two complementary branches are constructed, with the trend branch corresponding to... Perform one-dimensional convolution along the time dimension to emphasize low-frequency and slowly varying patterns: ; The significance of change branch calculation uses time-dimensional variance-based convergence to measure the intensity of fluctuations and instability. ; By concatenating the time-attention aggregation features, trend branches, and change significance branches along the channel dimension, we obtain the overall time representation: ; S3.1-7, To Lightweight normalization and optional channel recalibration are performed, and then the tensor is registered to the joint feature cache: 。 5. The spatiotemporal spectral joint super-resolution reconstruction method based on giant remote sensing constellations according to claim 1, characterized in that, The steps in S4 are as follows: S4.1-1, using the joint registration output of S2 as the basic input for the spatial branch, denoted as: ; Based on imaging geometry, the incident angle and azimuth angle are calculated at each pixel location and denoted as follows: and Encode geometric angles into learnable viewpoint embeddings: ; Simultaneously, spatial credibility weights are constructed using the mass mask: ; S4.1-2. Perform zero-mean unit variance normalization within the source and band: ; in, and Statistics on quality effective pixels, As a numerical stability constant, to explicitly introduce structural priors regarding edges and textures, the first-order spatial gradient is calculated: ; And based on this, a structure tensor is formed: ; in, Reflects the strength of the dominant edge. Characterize texture fineness, and then... , , and perspective embedding By splicing along the channel dimension, an initial tensor for spatial branching is formed, denoted as . This tensor carries radiation consistency and geometric structure information before entering the convolution; S4.1-3. Before nonlinear convolution, a linear weighting is performed based on the degree of geometric off-axis and quality reliability to reduce interference from extreme off-axis or low-quality sources. Define the viewpoint response function: ; And normalize the source dimension: ; Based on this, the initial image for multi-view linear fusion is obtained: ; S4.1-4, Merging Images Structural features and The perspective statistics and coordinate channels are combined to construct a two-dimensional convolutional input: ; in, This represents statistical aggregation over the source dimension; the first definition is... Two-dimensional convolutional unit: ; in It is a non-linear activation function. This unit simultaneously captures structural contours and fine-grained texture patterns within a local neighborhood; With fine-grained texture mode; S4.1-5, Introduce residual two-dimensional convolution and use multiple sets of dilation rates in the spatial dimension to construct a dilated spatial pyramid: ; in By alternating blocks, a hollow space pyramid is formed; S4.1-6, by Predicting deformable migration fields With weight Bilinear interpolation is performed at the convolution sampling positions: ; S4.1-7. Spatial representation includes both low-frequency contours and high-frequency textures. To achieve divide-and-conquer and complementarity, low-frequency structural branches and high-frequency texture branches are constructed, and then adaptive fusion is performed at the feature layer; let the low-frequency structural branch be a pair of... The smoothed residual convolution stack is denoted as The high-frequency texture branch is a convolutional stack under high-pass prior constraints, denoted as... The gating coefficients are generated by shallow convolution and then passed through... Compression, represented as: ; The final spatial fusion is represented as: ; S4.1-8. Apply consistency regularization to the spatial response of source statistics: ; in, For the first The spatial response obtained from a single source forward is used to measure the consistency error between sources; Standardize the channels and register them as spatial branch outputs: 。 6. The spatiotemporal spectral joint super-resolution reconstruction method based on giant remote sensing constellations according to claim 1, characterized in that, The steps in S5 are as follows: S5.1-1. Assemble the source-sensing spectral vectors at the pixel level and perform quality-weighted robust convergence to form a unified observation spectrum. Stack the registration outputs of S2 at each pixel level by band to form the source-sensing spectral vector. ; To suppress the effects of cloud and sensor anomalies, spectral weights are defined in conjunction with a quality mask: ; Then, the multi-source spectra are robustly converged in the source dimension using a weighted approach to obtain the pixel-level observation spectra: ; S5.1-2. Perform radiation normalization and baseline detrending in the band dimension to amplify weak absorption structures. First, perform zero-mean unit variance normalization in the band dimension: ; in and Obtained from quality effective pixel statistics, The numerical stability constant is then used, followed by a smooth baseline fitting along the wavelength axis to remove trends, with the convex hull or a low-order polynomial serving as the spectral baseline. The detrended spectrum is obtained: ; S5.1-3. One-dimensional convolution is used in the band dimension to capture local absorption shapes and inter-band correlations, in order to As input, let the first... The length of the layer spectral convolution kernel is Calculation in band dimension: ; in, For band axis neighborhood index, As a nonlinear activation, this formula encodes narrowband absorption, shoulder peaks, and interband coupling into discriminative channel responses; S5.1-4. A residual one-dimensional convolution and multi-scale time-invariant hole strategy are adopted to expand the spectral receptive field while maintaining resolution, and to cover both narrowband and broadband structures. Band-dimensional dilated convolution is introduced and stacked in residual form without downsampling. ; in, The band expansion rate is determined by the increasing sequence of residual blocks. S5.1-5. Constructing a spectral pyramid and enabling separable fusion to improve the discriminability of structures with different bandwidths, several branches are set in parallel on the one-dimensional convolutional backbone, with branch kernel length and dilation rate respectively. Output of each branch First, channel compression is performed using pointwise one-dimensional convolution, and then weighted fusion is performed using learnable weights: ; S5.1-6. Implement channel attention and quality modulation in the band dimension, highlighting the mission-relevant spectral bands, and introduce Squeeze–Excitation type channel attention, first for Channel statistics are obtained by performing a global average across the band dimensions. Then, a weighted vector is obtained through two layers of perceptrons. And superimposed band weights obtained by smoothing with quality and band-level signal-to-noise ratio. The weighted representation is obtained as follows: ; S5.1-7, Apply spectral smoothing regularization and perform low-dimensional bottleneck mapping, then apply L2 smoothing regularization in the band dimension: ; The spectral characterization is then mapped to a low-dimensional bottleneck to reduce the number of parameters and alignment costs. ; To preserve local geometric relationships, Apply approximately orthogonal constraints: ; S5.1-8. Standardize and interface with spectral branch outputs to ensure a consistent interface with both temporal and spatial branches. Perform lightweight standardization and optional channel recalibration to bring its statistics to the same order of magnitude as the temporal and spatial branches, and finally record it as the spectral branch output: 。 7. The spatiotemporal spectral joint super-resolution reconstruction method based on giant remote sensing constellations according to claim 1, characterized in that, The steps in S7 are as follows: S7.1-1, During Reconstruction On the unified reference grid, by joint feature tensor The generation space magnification is High-resolution reconstructed images Let the observation degradation operator be... This includes the point spread function, modulation transfer function, and downsampling raster mapping for the source and band; let the spectral response mapping be... Used to project from the target spectral reference to the sensor band, in order to Represents high-resolution pixel coordinates. Indicates the coordinates of the original reference raster cell; S7.1-2. Using a conditional denoising diffusion probability model, define a forward Markov process: ; Noise schedule Monotonically increasing, making , During training, real high-resolution targets sampling: ; S7.1-3, with As a condition, the upstream network output co-modulates the denoising network at both the channel dimension and the diffusion time. Let the conditional encoder be: ; Time embedding as Injection using characteristic linear modulation: ; in This represents the intermediate features of the denoising network. , Adaptive conditional control for different diffusion steps is achieved by predicting with a small perceptron. S7.1-4. The denoising network adopts an encoder-decoder structure with cross-layer skip connections, and performs the denoising ratio increase step by step in the spatial domain. Upsampling, making This represents the pixel rearrangement upsampling operator, with the scale-increasing sequence being... , No. Level decoding output: ; in For cascading, In-layer features from the encoder, conditional modulation is inserted at each scale. With residual convolution blocks, details are restored step by step while spectral statistics remain stable; S7.1-5, Learning the inverse process using noise prediction: ; The training loss is calculated using weighted noise regression. ; in, Adjust the relative importance of each diffusion step, and use a prediction correction strategy in the inference phase to reduce the number of steps; S7.1-6. Introduce an observation consistency term to perform backcasting for any source and band: ; in, This represents the high-resolution spectral image at the target time for reconstruction. For spectral response mapping, It is an observation degeneracy operator. This is a simulated observation by re-projection; and by quality weight Constraints and Real Observations Consistency: ; in, These are real, registered observations. Indicates the weight of quality and credibility; ; in, In the band The normalized spectral values, The target spectrum is used as a reference or pseudo-reference. S7.1-8, Introduction of Structure and Perception Terms: ; in, For a downsampled reconstruction map that matches the reference scale, It is the initial image obtained by spatial branching linear pre-fusion. This represents a fixed perceptual feature extraction layer; Further, a local contrast consistency term is introduced to stabilize the edges: ; in, Represents the spatial gradient operator. express Paradigm; S7.1-9. Based on the above, the training objective is: ; Weights Determined through validation set search; S7.1-10, The reasoning stage uses piecewise cosine reduction to shorten the schedule and number of steps. To enhance conditional dependencies, a classifier free guidance is introduced: ; in, To guide the intensity, This indicates that the condition is discarded. Substitute the mean expression from S7.1-5 and iterate to generate ; S7.1-11. After generation, perform multi-scale physical consistency acceptance: ; in, Set a threshold for the reference or pseudo-reference spectrum. and If the condition is not met, the diffusion steps are regressed by several steps and the value is adaptively increased. and A correction resampling step is performed, ultimately outputting a reconstructed image with high spatial and temporal resolution and high spectral fidelity. It is also registered for use by downstream classification and inversion tasks.

Citation Information

Patent Citations

  • Deep learning method for remote sensing parameter space-time spectrum fusion

    CN118351456A

  • Deep and shallow double-branch super-resolution method for forest hyperspectral satellite image

    CN120450964A