Remote sensing multi-parameter integrated inversion normal form method, system and equipment based on AI-Agent

By using the AI-Agent-driven DL-C-PSK paradigm, combined with the RM-Transformer-MoE nested model and multi-source database, the coupling problem in remote sensing multi-parameter inversion is solved, achieving high-precision and adaptive parameter inversion, and improving LST inversion accuracy and model interpretability.

CN121960179APending Publication Date: 2026-05-01INST OF AGRI RESOURCES & REGIONAL PLANNING CHINESE ACADEMY OF AGRI SCI

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INST OF AGRI RESOURCES & REGIONAL PLANNING CHINESE ACADEMY OF AGRI SCI
Filing Date
2026-01-22
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing remote sensing multi-parameter inversion techniques suffer from problems such as complex parameter coupling, strong prior dependence, and insufficient adaptability. Traditional methods have low computational efficiency, while AI methods have poor interpretability. Furthermore, existing AI integrated frameworks lack innovative integration of dynamic agents and nested models.

Method used

We construct the DL-C-PSK paradigm by combining AI-Agent-driven deep learning neural networks with RM-Transformer-MoE nested models, physical methods, statistical methods, and knowledge. We perform multi-parameter inversion using a high-precision multi-source database and radiative transfer equations, and introduce a refinement mechanism and MoE routing to achieve adaptive expert modeling.

Benefits of technology

It significantly improves the accuracy of multi-parameter synchronous inversion, with LST inversion accuracy improved by about 20%, reaching below 1 K. The RMSE and MAE values ​​are significantly better than traditional methods, improving the interpretability and adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121960179A_ABST
    Figure CN121960179A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing multi-parameter integrated inversion normal form method, system and equipment based on AI-Agent. According to the method, a deep learning neural network is dynamically driven through AI-Agent, a refining mechanism (RM)-Transformer-MoE size nested model, a physical method, a statistical method and expert knowledge are coupled, a DL-C-PSK normal form is constructed, a high-precision multi-source database is established based on the normal form, an appropriate radiation transfer equation is constructed through geophysical logical reasoning, and a high-precision multi-source database is established. And inversion of parameters such as surface temperature, surface emissivity, atmospheric water vapor content and near-surface air temperature is realized. According to a causal relationship between an input wave band and an output parameter, a direct synchronous inversion or iterative inversion mode is adopted to ensure multi-parameter high-precision synchronous inversion. Wherein the core of the deep learning neural network comprises RM logic derivation, SHAP model interpretation, Transform model architecture and a Transform-MoE size nested model, so that the interpretability, the adaptability and the precision of the model are improved. Through an AI-Agent driven RM-Transform-MoE nested model, deep coupling of physics-statistics-knowledge is realized, compared with a traditional SW method, the inversion precision is greatly improved, and verification shows that the technology is suitable for the fields of global climate observation, environment monitoring and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of remote sensing parameter inversion technology, and particularly relates to a method, system and device for integrated remote sensing multi-parameter inversion paradigm based on AI-Agent. Background Technology

[0002] According to the Global Climate Observing System (GCOS), land surface temperature (LST) is a wavelength-independent dynamic quantity representing the thermodynamic temperature of a specific surface layer. As a key physical quantity, it plays a crucial role in thermal radiation, surface-atmosphere interactions, and energy exchange. Strongly coupled with LST is land surface emissivity (LSE), which represents the inherent ability of the surface to emit radiation compared to a blackbody at the same LST; its variation affects LST retrieval. For example, studies using a single-window algorithm to retrieve LST based on thermal infrared (TIR) ​​data show that a 1% LSE error leads to approximately a 0.6 K LST error. During radiative transfer, the upward radiation energy from the surface is attenuated by various gases in the atmosphere; atmospheric water vapor is a major factor in atmospheric effects, and its absorption effect cannot be ignored. Simultaneously, the atmospheric effective mean atmospheric temperature (MAT) is influenced by the atmospheric profile temperature distribution and atmospheric water vapor. Near-surface air temperature (NSAT), which is closely related to the air temperature of each layer of the atmospheric profile, is generally defined as the atmospheric temperature observed near the Earth's surface (usually at a height of 2 m above the ground). It is a key parameter linking the energy exchange process between the Earth and the atmosphere, and is also influenced by LST, LSE, and atmospheric water vapor content (WVC). The coupling of these geophysical parameters means that retrieving one parameter requires accurately obtaining other related parameters, making the decoupling study of surface and atmospheric information of significant theoretical and practical importance. Accurate estimation of these parameters helps different fields understand surface and climate change processes and provides important algorithmic and data support for research on global warming, environmental change, surface evapotranspiration, and agricultural meteorological disasters.

[0003] Acquiring LST, LSE, WVC, and NSAT data using limited ground monitoring stations and radiosonde techniques ensures data accuracy, but is costly, inefficient, and cannot achieve temporal and spatial data continuity in large-scale studies. In contrast, remote sensing technology has the ability to accurately monitor various elements of the Earth system globally. Among them, TIR remote sensing is often used to invert key parameters in the surface-atmosphere system. Traditional inversion methods can be divided into single-channel, SW, and multi-channel methods. Single-channel methods (such as the general single-channel algorithm proposed by Jiménez-Muñoz et al., 2003) use a single TIR channel to establish the Radiative Transfer Equation (RTE) to estimate LST, requiring prior knowledge such as LSE, atmospheric transmittance, and MAT. SW methods (such as the algorithm of Wan and Dozier, 1996) add a TIR channel, using two RTEs to invert LST, and correcting for atmospheric effects through two adjacent TIR bands. Multi-channel methods further add TIR or MIR bands to invert LST, improving accuracy. However, traditional methods still have shortcomings, such as reliance on prior knowledge, day-night image matching, uncertainty in LSE variation, limited dynamic range of bands, and the influence of solar radiation. Existing techniques, such as Gillespie et al.'s (1998) temperature-emissivity separation algorithm, rely on multi-channel data but are insufficient in handling LSE variation uncertainty; Mushkin et al. (2005) extended surface temperature and emissivity retrieval to the mid-infrared (3-5 μm) range using a multispectral thermal imager (MTI); Schmugge et al. (2002) discussed separating temperature and emissivity from multispectral thermal infrared observations; Peres and DaCamara (2005) used emissivity maps to retrieve surface temperature from MSG / SEVIRI; and Sobrino et al. (2012) used a classification method to map the emissivity of urban areas using the DESIREX experiment. This invention addresses these problems through the DL-C-PSK paradigm.

[0004] Existing AI methods, such as deep learning-based LST inversion (referencing "Artificial Intelligence-Based Joint Retrieval Algorithm for Land Surface Temperature, Emissivity, and Atmospheric Water Vapor" published on EarthArXiv 2023), utilize convolutional neural networks (CNNs) combined with radiative transfer simulation data to invert LST and WVC. However, they neglect the deep integration of physical constraints, resulting in insufficient model generalization ability under extreme weather conditions. Furthermore, they do not involve NSAT inversion, and their handling of LSE relies on prior spectral libraries rather than real-time refinement mechanisms. Another example is "A Mechanism-Learning Deeply Coupled Deep Model for Remote Sensing Retrieval of Global Land Surface Temperature" published on arXiv 2025. This model learns to couple physical models through a mechanism, but lacks oscillatory refinement using statistical methods and expert guidance from knowledge sources, making it prone to overfitting to noisy data. In addition, the related research published by Tan et al. (2021) in the journal Remote Sensing uses LSTM networks to process time series data, which improves the spatiotemporal continuity of LST inversion, but does not integrate WVC and NSAT, resulting in incomplete parameter decoupling.

[0005] To further objectively assess the state of existing technologies, it is necessary to compare relevant patent technologies, such as Chinese invention patent CN118797218A (A Remote Sensing Multi-Parameter AI Integrated Inversion Paradigm Method, System, and Device). This patent discloses an AI-based remote sensing multi-parameter integrated inversion method, mainly focusing on using deep learning models combined with physical and statistical methods to achieve the inversion of parameters such as LST, LSE, WVC, and NSAT. The core of this patent lies in constructing a deep learning framework to improve the accuracy of parameter estimation through multi-source data fusion and the inversion process of the radiative transfer equation. Specifically, CN118797218A describes using neural networks to process remote sensing data, establish a high-precision database, and perform direct or iterative inversion based on band combinations. This patent emphasizes the synchronous inversion of multiple parameters to solve the problem of difficult decoupling of parameter coupling in traditional methods, and ensures the reliability of the results through verification steps. However, this patent has limitations in its AI-driven mechanism, mainly relying on static deep learning models without introducing dynamic agents (such as AI-Agents) to optimize the coupling process. This results in poor adaptability in complex remote sensing scenarios. For example, in areas with strong surface heterogeneity or variable atmospheric conditions, the inversion accuracy may be limited by the model's fixed parameter sharing paradigm. Furthermore, CN118797218A not only fails to use the Transformer architecture but also fails to implement nested integration of MoE (Mixture of Experts), thus failing to effectively handle the differentiated responses of different surface types. While the system implementation of this patent includes an inversion module and database construction, its integrated inversion module does not clearly distinguish the selection logic between direct synchronization and iterative modes, relying solely on band combination judgments. Moreover, the verification part relies on simulation and cross-validation but does not quantify accuracy indicators such as specific thresholds for RMSE and MAE. In contrast, this invention, through an AI-Agent-driven DL-C-PSK paradigm, achieves the organic coupling of physical methods, statistical methods, and knowledge. The RM-Transformer-MoE nested model introduces a refinement mechanism (RM) and SHAP model interpretation, improving the model's interpretability and adaptability. This dynamic optimization mechanism improves the accuracy of the invention by approximately 20% in multi-parameter synchronous inversion, with the LST error controlled below 1K, significantly outperforming the accuracy level reported in CN118797218A. This invention overcomes these shortcomings by introducing an AI-Agent as a coordinator to dynamically adjust the parameters of the deep learning model, achieving comprehensive coupling of physics, statistics, and knowledge. Furthermore, the real-time feedback from the AI-Agent optimizes the mode selection, enhancing the system's robustness.

[0006] Overall, existing technologies have made significant progress in the field of remote sensing multi-parameter inversion, but still face challenges such as complex parameter coupling, strong prior dependence, and insufficient adaptability. Traditional methods, while physically sound, suffer from low computational efficiency; AI methods, while flexible, lack interpretability. CN118797218A, as a typical example, provides an integrated AI framework, but lacks dynamic intervention by AI-Agents and innovative integration of nested models. This invention addresses these issues by proposing an AI-Agent-based DL-C-PSK paradigm. Through a refinement mechanism and MoE routing, it achieves adaptive expert modeling, ensuring high-precision synchronous inversion of multiple parameters. The introduction of this paradigm not only overcomes the limitations of existing technologies but also provides more reliable tool support for global climate monitoring. Summary of the Invention

[0007] To address the aforementioned technical problems, this invention proposes an AI-Agent-based integrated multi-parameter remote sensing inversion paradigm, method, system, and device. This addresses the issues of low inversion accuracy, strong reliance on prior knowledge, and difficulty in decoupling parameter coupling in existing technologies. This invention achieves high-precision synchronous inversion of multiple parameters by coupling AI-Agent-driven deep learning with a nested RM-Transformer-MoE model. Compared to the traditional SW method, the LST inversion accuracy is improved by approximately 20%, reaching below 1K.

[0008] To achieve the above objectives, the present invention provides the following technical solution:

[0009] A remote sensing multi-parameter integrated inversion paradigm based on AI-Agent includes the following steps:

[0010] The DL-C-PSK paradigm is constructed by coupling RM-Transformer-MoE nested models, physical methods, statistical methods, and knowledge with AI-Agent-driven deep learning neural networks.

[0011] A high-precision multi-source database is constructed based on the DL-C-PSK paradigm;

[0012] A well-posed radiative transfer equation is constructed based on geophysical logical reasoning;

[0013] Based on the high-precision multi-source database and radiative transfer equation, atmospheric parameters are obtained by inverting remote sensing multi-parameters.

[0014] A three-theme parameter inversion multi-source database is constructed based on the atmospheric parameters;

[0015] When the three subject parameters of the multi-source database satisfy the dominant band combination, direct synchronous inversion is performed through the radiative transfer equation.

[0016] If the conditions are not met, add prior knowledge for iterative inversion.

[0017] Experimental results show (e.g.) Figure 2-7 As shown), in the inversion verification during day and night, this invention verifies the combination of LST and LSE (e.g.) Figure 2 ), RMSE value less than 1K; for WVC inversion (such as Figure 3 MAE value less than 0.5 g / cm²; for NSAT inversion (e.g. Figure 4 The accuracy is about 20% better than the traditional SW method, and the cross-validation MAE is less than 1.9K, which significantly improves the effect of multi-parameter synchronous inversion.

[0018] Preferably, the main architecture of the deep learning neural network is as follows:

[0019]

[0020] in This is the input value of the k-th neuron in the first hidden layer. The value of the j-th neuron in the input layer. Here, J represents the synaptic weights from the j-th neuron in the input layer to the k-th neuron in the first hidden layer, where J is the number of neurons in the input layer and k = 1 to K is the number of neurons in the first hidden layer.

[0021]

[0022] in This represents the output value of the z-th neuron in the last hidden layer. f is the input value of the z-th neuron in the last hidden layer, f is the activation function, and z = 1 to Z (Z is the number of neurons in the last hidden layer).

[0023]

[0024] in Let m be the input value of the m-th neuron in the output layer. This represents the output value of the z-th neuron in the last hidden layer. Let M be the synaptic weight from the z-th neuron in the last hidden layer to the m-th neuron in the output layer, where M is the number of neurons in the output layer.

[0025]

[0026] in Let f be the final output value of the m-th neuron in the output layer, f be the activation function, and adaptive expert modeling is achieved through RM-Transformer-MoE nesting.

[0027] Preferably, the method for constructing a three-theme parameter inversion multi-source database includes: combining surface temperature and surface emissivity as the first theme, constructing a database with atmospheric water vapor content, near-surface air temperature and auxiliary data sources, and applying the RM mechanism to correct biases.

[0028] Preferably, the radiance at the top of the atmosphere The expression is:

[0029]

[0030] in At the top of the atmosphere at the observed zenith angle Observation azimuth Radiance under temperature Ti in remote sensing data of band i This represents the integral from the lower band limit λ1 to the upper band limit λ2. Let i be the normalized spectral response function for band i. Let λ be the atmospheric transmittance at wavelength λ. Ground radiance ( (where ) represents the temperature of the ground in band i (remote sensing data). This is upward thermal radiation from the atmosphere. This is thermal radiation scattered upwards by the atmosphere.

[0031] Preferably, the ground radiance The expression is:

[0032]

[0033]

[0034]

[0035] in Temperature of ground in band i remote sensing data Observing the zenith angle Observation azimuth and radiance at wavelength λ, This represents the integral from the lower band limit λ1 to the upper band limit λ2. Let i be the normalized spectral response function for band i. Let λ be the surface emissivity at wavelength λ. The values ​​of the Planck function are given by the surface temperature Ts and wavelength λ. This is downward atmospheric thermal radiation. This refers to the downward-scattering of solar radiation by the atmosphere. Bidirectional reflectance of the earth's surface ( The zenith angle of the sun. (The azimuth angle of the sun) The solar irradiance at the top of the atmosphere, cos( () is the cosine of the solar zenith angle. Atmospheric transmittance in the direction of the sun.

[0036] Preferably, the calculation of the effective surface emissivity is as described in claim 10.

[0037] Preferably, it further includes a verification step, as described in claims 11 and 12.

[0038] The system and device are as described in claims 13 and 14. Attached Figure Description

[0039] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0040] Figure 1 This is a flowchart of the AI-Agent-based remote sensing multi-parameter integrated inversion method according to an embodiment of the present invention. The diagram illustrates the complete process from data acquisition to the final inversion result output, including data acquisition & preprocessing, multi-source database construction, radiative transfer equation derivation, direct inversion mode, iterative inversion mode, and parameter estimation & verification.

[0041] Figure 2 The LST and LSE inversion results for different band combinations in this embodiment of the invention are shown below. C1 represents the 5BT_to_LST_5LSE (27, 28, 29, 31, 32) combination, C2 represents the 4BT_to_LST_5LSE (27,29,31,32), C3 represents the 4BT_to_LST_4LSE (27,28,31,32), C4 represents the 4BT_to_LST_5LSE (28,29,31,32), C5 represents the 3BT_to_LST_3LSE (29,31,32) combination, C6 represents the 2BT_to_LST_2LSE (31,32) combination, and C7 represents the 1BT_to_LST_1LSE (32). The average precision of the C1, C2, C3, C4, C5, C6, and C7 combinations are 0.41K, 0.51K, 0.51K, 0.54K, 1.09K, 1.44K, and 3.62K, respectively; the average emissivity error is below 0.01, and the C1 and C2 combinations are below 0.005.

[0042] Figure 3The images show the verification results of the multi-source database for atmospheric water vapor (WVC) inversion in this embodiment of the invention, where (a) is the daytime verification result with an average accuracy of 0.377 g / cm²; and (b) is the nighttime verification result with an average accuracy of 0.268 g / cm². All verifications are based on MODIS data from 2020 to 2022.

[0043] Figure 4 The near-surface air temperature (NSAT) verification results for this embodiment of the invention are shown in (a) daytime verification results with an average accuracy of 2.14 K and (b) nighttime verification results with an average accuracy of 1.64 K. All verifications are based on MODIS data from 2020 to 2022.

[0044] Figure 5 This is a scatter plot of the LST inversion multi-source database cross-validation results from an embodiment of the present invention, where (a) is daytime and (b) is nighttime. All validations are based on MODIS data from 2020 to 2022.

[0045] Figure 6 This is a scatter plot of the WVC inversion multi-source database cross-validation results from an embodiment of the present invention, where (a) is daytime and (b) is nighttime. All validations are based on MODIS data from 2020 to 2022.

[0046] Figure 7 This is a scatter plot of the NSAT inversion multi-source database cross-validation results from an embodiment of the present invention, where (a) is daytime and (b) is nighttime. All validations are based on MODIS data from 2020 to 2022. Detailed Implementation

[0047] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0048] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0049] Example 1

[0050] like Figure 1 As shown, this embodiment provides a remote sensing multi-parameter integrated inversion paradigm method based on AI-Agent, including the following steps:

[0051] The DL-C-PSK (deep learning coupled with physical methods + statistical methods + expert knowledge) paradigm is constructed by using an AI-Agent-driven deep learning neural network coupled with a nested RM-Transformer-MoE model, physical methods, statistical methods, and knowledge. A high-precision multi-source database is built based on the DL-C-PSK paradigm, and a well-defined radiative transfer equation is constructed based on geophysical logical reasoning. Based on the high-precision multi-source database and the radiative transfer equation, high-precision atmospheric parameters are obtained by inverting multiple remote sensing parameters.

[0052] The atmospheric parameters include surface temperature, surface emissivity, atmospheric water vapor content, and near-surface air temperature. A three-theme parameter inversion multi-source database is constructed based on the atmospheric parameters.

[0053] When the three-theme parameter inversion multi-source database satisfies the dominant band combination, the atmospheric parameters are directly inverted through the radiative transfer equation to obtain the final atmospheric parameters.

[0054] When the multi-source database for the inversion of the three thematic parameters does not satisfy the dominant band combination, prior knowledge from the previous inversion is added to assist in iterative inversion to obtain the final atmospheric parameters.

[0055] The DL-C-PSK paradigm and integrated inversion are explained in detail below:

[0056] (1) DL-C-PSK paradigm

[0057] The DL-C-PSK paradigm, driven by AI-Agent, couples RM-Transformer-MoE nested models, physical methods, statistical methods, and knowledge to improve the accuracy and efficiency of TIR remote sensing parameter inversion. This paradigm goes beyond simply combining elements; it organically integrates them based on a deep analysis of their respective advantages, thus pioneering a new research approach. In terms of physical methods, the DL-C-PSK paradigm focuses on establishing the radiative transfer (RTE) equation for the thermal radiation transfer process. The RTE equation is similar to mathematically modeling the energy exchange relationships between the Earth's surface and atmosphere for various parameters in the Earth-atmosphere system, serving as a bridge connecting satellite remote sensing data (such as BT) with surface and atmospheric parameters (such as LST, LSE, WVC, and NSAT). By constructing an accurate and appropriate RTE, high-precision LST, LSE, WVC, and NSAT parameters can be obtained through inversion. The DL-C-PSK paradigm optimizes the RTE solution process by introducing the RM-Transformer-MoE nested model, making it more efficient and accurate. In terms of statistical methods, the DL-C-PSK paradigm establishes a high-precision multi-source database. This database provides richer information and strong support for parameter inversion. In other words, it represents the space of theoretically representative "solutions" (true values) for the parameter. Based on the concept of geospatial big data, this data can include remote sensing data from different satellites, reanalysis datasets, measured data from ground monitoring stations and radiosondes, and simulation data from physical models. The RM-Transformer-MoE nested model can learn the inherent laws of radiative transfer processes from massive amounts of data, thereby achieving more accurate parameter inversion. Knowledge utilization is another important aspect of the DL-C-PSK paradigm. By integrating theoretical knowledge from the field of remote sensing, previous research results, and practical experience, the DL-C-PSK paradigm can make more rational decisions in data processing and parameter inversion, similar to a regulator incorporating human wisdom. Specifically, expert knowledge plays a crucial role in analyzing the feasibility of global methods and verifying the obtained results, also possessing a "reverse translation" function. The "reverse translation" introduced in this embodiment from the field of AI research can also be interpreted as "abductive learning," which refers to the process of selectively inferring certain facts based on background knowledge to explain phenomena and observations. Therefore, it avoids blindly relying on theoretical reasoning or actual situations. This knowledge-based method not only ensures data quality but also ensures that the inversion process conforms to actual laws, improving the reliability of the inversion results.

[0058] The RM-Transformer-MoE nested model, as a core component of the DL-C-PSK paradigm, includes the following detailed mechanisms:

[0059] Refinement Mechanism (RM): During radiative transfer, noise factors such as atmospheric absorption / scattering effects, sensor radiometric calibration errors, and surface heterogeneity cause the observed LST to deviate from the true LST, resulting in a systematic bias. To suppress the impact of this systematic bias on the learning objective, a deep learning refinement mechanism is proposed. This mechanism uses the physical retrieval results as the initial LST / LSE / WVC / NSAT, and employs a neural network to learn the mapping between the initial LST / LSE / WVC / NSAT and the true LST / LSE / WVC / NSAT. This corrects the bias in the initial LST, gradually approximating the true LST, as shown in the formula. ,in For target parameters (e.g., surface temperature, surface emissivity, atmospheric water vapor content, and near-surface air temperature). For deep learning models ( (where x is the model parameter) and x is the input brightness temperature (BT). Considering that the network may still be limited by sample distribution and model capacity, an oscillation refinement mechanism based on mathematical statistics is further proposed. Perform stepwise convergence correction:

[0060]

[0061] in and LST / LSE / WVC / NSAT for k and k+1 iterations respectively, where k is the number of oscillations (k>0). (0, 1) represents the oscillation weight coefficients; when the convergence condition is met, it is considered the relative true value, where K is the final iteration number. If a value converges to a certain stable value in the mean-square sense, then we consider it as... Relative truth value: ,in Finally, refine the LST / LSE / WVC / NSAT values. Let LST / LSE / WVC / NSAT be the LST / LSE / WVC / NSAT values ​​after the k-th iteration, where K is the final iteration number. By coupling a deep learning refinement mechanism and an oscillatory refinement mechanism, the influence of atmospheric and surface uncertainties and other factors on LST / LSE / WVC / NSAT inversion is gradually reduced, resulting in a final LST / LSE / WVC / NSAT value. Theoretically, it converges to a refined solution that is closer to the physical truth.

[0062] SHAP Model: Brightness Temperature (BT) is a core observation characterizing Earth's surface thermal radiation and has a clear physical relationship with LST / LSE / WVC / NSAT. However, the sensitivity differences of BT across different thermal infrared channels are often difficult to quantify in deep learning inversion, lacking interpretability. The SHAP model, proposed by Lundberg et al. in 2017, can be used for global and local interpretability analysis of models. The core idea of ​​the SHAP model is to explain the importance of each input feature by evaluating the marginal contribution of features across different feature subsets and then weighting the average, as shown in the formula:

[0063]

[0064] in Let S be the Shapley value of feature i, S be the subset of features that do not include feature i, and N be the set of all features. Let S be the size of the subset S, |N| be the number of all features, v(S∪{i}) be the model output value after using the subset S and adding feature i, and v(S) be the model output using the subset S; the prediction decomposition is as follows: , where g(x′) is the output value of the explanatory model, and x′ is the simplified feature vector (binary representation of whether the feature exists). Let M be a basic constant and M be the number of simplified features. Let xi be the Shapley value of the i-th feature, and xi′ be the value of the i-th simplified feature (1 indicates presence, 0 indicates absence). The global importance of each feature is obtained by calculating its Shapley value, and the features are then ranked accordingly.

[0065] The Transformer model, proposed by Vaswani et al. in 2017, is a representative deep learning architecture. Due to its powerful feature representation and learning capabilities, it is suitable for geophysical parameter inversion. Essentially, the Transformer is an encoder-decoder structure, internally composed of multiple stacked encoders and decoders. The core of the Transformer model is the multi-head attention mechanism, which consists of multiple scaled dot-product attention mechanisms. The scaled dot-product attention function assigns weights to different value vectors based on the similarity between the query vector and the key vector, as shown in the formula:

[0066]

[0067] Where Attention(Q,K,V) is the output of the attention mechanism, Q is the query matrix, K is the key matrix, and V is the value matrix. This is the transpose of the key matrix. Let be the dimension of the key matrix, and be the softmax activation function. The multi-head self-attention mechanism obtains multi-view feature relationships by projecting Q, K, and V onto multiple subspaces to compute attention in parallel. Finally, the outputs of each head are concatenated and fused together using a linear transformation. The Transformer, with its self-attention mechanism at its core, can model the nonlinear relationship between multi-channel BT and LST in parallel, but it does not contain positional information itself. Positional encoding is needed to supplement the relative or absolute positional information of the elements. Positional encoding uses sine and cosine functions, as shown in the formula:

[0068]

[0069]

[0070] in and These are the position encoding values ​​of position pos in even-numbered dimensions 2i and odd-numbered dimensions 2i+1, respectively, where pos is the position index in the sequence and i is the dimension index. The model dimensions are defined as follows: positional encoding uses a sine function in even-numbered dimensions and a cosine function in odd-numbered dimensions, thereby generating a unique vector for each sequence position.

[0071] Transformer-MoE Nested Model: Remote sensing scenes are complex and surface types are diverse. While a single large Transformer model can capture global dependencies through self-attention, it still struggles to simultaneously account for the differentiated responses of different surfaces under a unified parameter-sharing learning paradigm. To improve the model's adaptability to the characteristics of remote sensing data, a Mixture of Experts (MoE) model is embedded within the Transformer model to construct a nested model (large-small model). Figure 2 The specific operating steps are as follows:

[0072] Step 1: Encode the location of the input samples and extract features from the sample data using the Transformer model.

[0073] Step 2: Perform linear mapping on the input features of each layer to obtain the query matrix, key matrix and value matrix. Then, use a scaled dot product attention mechanism to calculate the weights from Q and K, and sum them with weights over V to obtain the single-head output.

[0074] Step 3: After multi-head attention is computed in parallel in multiple subspaces, the outputs of each head are concatenated and linearly mapped to obtain the multi-head attention result.

[0075] Step 4: Input the output of multi-head attention into the feedforward neural network (FFN) for feature transformation.

[0076] Step 5: Embed the MoE model after the FFN output to learn surface difference features. The gated network calculates expert weights for each and uses sparse routing to activate a minority of experts, as shown in the formula.

[0077]

[0078] Where G(x) is the output weight vector of the gating network, x is the input token, Wg is the weight matrix of the gating network, and softmax is the softmax activation function; sparse routing uses noisy top k selection, as shown in the formula.

[0079] , ϵ∼

[0080] in This is the input after adding Gaussian noise. Let ϵ be the original input token, and ϵ be Gaussian random noise. The mean is μ and the variance is μ. The normal distribution smooth activation function.

[0081] Step 6: Use the MoE output as a compensation term for FFN, and perform residual connection and normalization processing, as shown in the formula:

[0082]

[0083]

[0084] in The result of layer normalization (LayerNorm) is the residual connection between the multi-head attention output z and the MoE output MoE(z), where z′ is... With FFN output FFN( The final result after residual connection and layer normalization is as follows: LayerNorm is the layer normalization function, z is the output of the multi-head attention mechanism, MoE(z) is the output of the MoE layer, and FFN( ) represents the output of the feedforward network.

[0085] Step 7: Output the output through the standard output layer to obtain the final result.

[0086] (2) Integrated inversion

[0087] This embodiment, based on the DL-C-PSK paradigm, establishes three multi-source databases ("LST&E", "WVC", and "NSAT") for thematic parameter inversion. These three databases are used for inversion with LST&E, WVC, and NSAT as themes (objects), respectively. The merging of LST and LSE is due to their strong coupling relationship, making them suitable for synchronous inversion. Furthermore, two parameter inversion modes for integrated inversion are determined: one is "direct synchronous inversion" satisfying the dominant band combination, where the four parameters can be directly inverted by establishing the RTE of the TIR bands; the other is to improve the inversion accuracy of non-dominant band combinations by adding prior knowledge from previous inversions ("iterative inversion"). In the integrated method, a dominant band combination is defined as a band combination where the input band information and output parameters are highly correlated (strong causal relationship). This combination needs to meet the GLR's requirement for the number of equations (directly establishing a well-posed equation set) or constructing a well-posed equation set by adding prior knowledge. A non-dominant band combination, on the other hand, means a band combination where the input band information and output parameters have a weak correlation (weak causality). At this point, regardless of whether the RTE is well-posed or under-posed, adding prior knowledge can not only assist in constructing a well-posed RTE, but also amplify the relevant information to help improve the accuracy of parameter inversion. Therefore, the three databases can be inverted simultaneously, achieving synchronous inversion of all four parameters, i.e., "direct synchronous inversion." Both direct inversion and iterative inversion by adding prior knowledge in a loop can achieve integrated multi-parameter inversion from the same data source. Theoretically, this method can maximize accuracy and ensure the uniformity, convenience, and efficiency of the inversion method, distinguishing it from traditional methods and other AI-based methods for single-parameter inversion. The RM-Transformer-MoE nested model enhances its ability to handle surface heterogeneity through adaptive expert modeling in this process.

[0088] To verify the above-mentioned technology, this embodiment also conducted the following experiments for comparison and cross-validation, please see details below. Figures 2 to 6 .

[0089] 1. Data Selection

[0090] The application utilizes MODIS data with multiple bands, including LST, LSE, WVC, and NSAT data that have not yet been released. The MODIS sensor, a hyperspectral radiometer on the Terra and Aqua satellites, has 36 bands covering the visible, near-infrared, short-wave infrared, and mid-wave infrared spectral ranges. It offers advantages such as global coverage, high radiometric resolution, and dynamic measurement, and exhibits high calibration accuracy in the thermal infrared (TIR) ​​band. These TIR bands can be used to retrieve sea surface temperature, LST, and atmospheric characteristics, acquiring various meteorological and environmental parameters such as surface reflectance, surface temperature, atmospheric temperature, water vapor content, clouds, and aerosols. The analysis period is three years, with data from January to December 2020-2021 used for training and validation, and data from January to December 2022 used for independent testing. The LST retrieval MAE using this method is 0.78K, superior to the 1.5K of the MYD21 product.

[0091] 1.1 MODIS Data

[0092] Several key principles must be considered when selecting spectral bands to optimize the extraction of target object information. These principles include maximizing the signal response of the target object, minimizing interference from the transmission medium, and selecting specific spectral regions to distinguish the target object. These choices collectively improve the visibility of target object information, minimize adverse effects, and ultimately achieve more accurate measurements. Therefore, this embodiment focuses specifically on MODIS bands 27, 28, 29, 31, and 32. These bands are widely recognized in studies across multiple fields and are highly correlated with the inversion of radiation-related parameters of the Earth-atmosphere system, such as WVC, LST, LSE, and NSAT. The important role of these bands in radiation observations further highlights their importance for accurate parameter inversion. Bands 27 and 28 are absorption bands specifically designed to capture atmospheric water vapor information; these bands are highly sensitive to humidity in the upper troposphere. Bands 29, 31, and 32 are primarily suitable for TIR inversion studies of LST, LSE, WVC, and NSAT. A combination of these five bands can better invert multiple parameters.

[0093] 1.2 Auxiliary Data

[0094] Six auxiliary data sources were used, including MODTRAN simulation data, satellite products, reanalysis datasets, site monitoring data, emissivity spectral library data, and spectral response function data. Integrating these data sources ensured the accuracy and reliability of the multi-source database.

[0095] MODTRAN data. MODTRAN has been widely used to construct RTE and validate models for various remote sensing analyses and applications worldwide. In addition to the usual settings of environmental parameters and the settings of spectral regions for each band, the following four parameters were modified to simulate the radiometric information of the five TIR bands (bands 27, 28, 29, 31 and 32) of MODIS captured by the sensor: (1) LST: range of 270-330 K, step size of 2 K; (2) observation angle: range of 0~65°, step size of 3°; (3) LSE: 17 land cover types including vegetation, soil, rock, sand and artificial materials were selected; (4) WVC: range of 0.1-4.5 g / cm2, interval of 0.2 g / cm2.

[0096] MODIS includes LST&E products based on the TES algorithm (MYD21), WVC products based on the TIR SW algorithm (MYD05), cloud mask products (MYD35), and geolocation products (MYD03). Except for MYD05, which has a spatial resolution of 5 km, all other products have a spatial resolution of 1 km. MODIS is renowned for its global coverage, high radiometric resolution, and precise calibration in the visible, NIR, and TIR bands. The MxD21 product algorithm further reduces atmospheric effects, offering advantages over the MxD11 product developed using the SW algorithm. High-quality data (“good quality”) under clear-sky conditions was extracted using the MYD21 LST&E quality control file. For the MYD05 WVC data, high-quality data under the “Best Quality” condition was selected. MYD35 was used for cloud pixel removal during both day and night, using only “confident clear” data. MYD03 was used for geometric correction of MODIS imagery. The acquisition time for MODIS data was consistent with the imaging time for AHI.

[0097] The ERA5 reanalysis dataset is used. ERA5 is one of the most advanced reanalysis products, validated by research in many fields. Since NSAT currently lacks a widely validated, officially verified satellite remote sensing product, using a high-precision reanalysis dataset is a reasonable choice to obtain NSAT data that matches the pixel BT of AHI. ERA5 provides global products hourly with a spatial resolution of 0.1°, recorded in UTC time. Therefore, the first step is to interpolate the ERA5 imagery temporally to match the imaging time of the AHI data. Furthermore, data provided by ERA5, such as LST, WVC, and NSAT, can also be used to collaboratively verify the reliability of site-measured data and satellite product data values. If significant errors exist, their reliability needs to be comprehensively considered.

Claims

1. A remote sensing multi-parameter integrated inversion paradigm method based on AI-Agent, characterized in that, Includes the following steps: This paradigm constructs a deep learning coupled physics-statistics-knowledge (DL-C-PSK) model by using AI-Agent to drive a deep learning neural network coupled with a nested RM-Transformer-MoE model, physical methods, statistical methods, and knowledge. The nested RM-Transformer-MoE model includes the nested integration of refinement mechanism logic derivation, SHAP model interpretation, Transformer architecture, and MoE routing to achieve adaptive expert modeling. Unlike existing deep learning methods, this paradigm improves the accuracy of multi-parameter synchronous inversion by dynamically optimizing the coupling process through AI-Agent. A high-precision multi-source database is constructed based on the DL-C-PSK paradigm, wherein the high-precision multi-source database contains a representative solution space of unknown parameters; A well-posed radiative transfer equation is constructed based on geophysical logical reasoning; Based on the high-precision multi-source database and radiative transfer equation, remote sensing multi-parameters are inverted to obtain atmospheric parameters, including surface temperature, surface emissivity, atmospheric water vapor content, and near-surface air temperature. Based on the atmospheric parameters, a multi-source database for inversion of three themes is constructed. The three themes include the combination of surface temperature and surface emissivity, atmospheric water vapor content, near-surface air temperature and its auxiliary data sources. When the three-theme parameter inversion multi-source database satisfies the dominant band combination, the atmospheric parameters are directly and synchronously inverted through the radiative transfer equation to obtain the final atmospheric parameters. When the multi-source database for the inversion of the three thematic parameters does not satisfy the dominant band combination, prior knowledge from the previous inversion is added to assist in iterative inversion to obtain the final atmospheric parameters.

2. The method according to claim 1, characterized in that, The main architecture of the deep learning neural network includes an input layer, a hidden layer, and an output layer. The hidden layer uses an RM-Transformer-MoE nested model to implement a multi-head attention mechanism and hybrid expert routing. The RM-Transformer-MoE nested model includes a refinement mechanism logic derivation, a SHAP model, a Transformer model, and a Transformer-MoE nested architecture.

3. The method according to claim 2, characterized in that, The logical derivation of the refinement mechanism includes: during radiative transfer, using a deep learning refinement mechanism with the physical retrieval result as the original LST, using a neural network to learn the mapping between the original LST and the true LST, and correcting the initial LST for deviation, as shown in the formula: ,in For target parameters (e.g., surface temperature, surface emissivity, atmospheric water vapor content, and near-surface air temperature). For deep learning models ( (where x is the model parameter) and x is the input brightness temperature (BT); further, an oscillatory refinement mechanism based on mathematical statistics is used for gradual convergence correction, as shown in the formula. ,in and LST / LSE / WVC / NSAT for k and k+1 iterations respectively, where k is the number of oscillations (k>0). (0, 1) is the oscillation weight coefficient; when the convergence condition is met, it is considered as the relative true value, where K is the final iteration number.

4. The method according to claim 2, characterized in that, The SHAP model is used to quantify the sensitivity differences of different thermal infrared channels (BT), providing global and local interpretability analysis; the SHAP value calculation formula is as follows: ,in Let S be the Shapley value of feature i, S be the subset of features that do not include feature i, and N be the set of all features. Let S be the size of the subset S, |N| be the number of all features, v(S∪{i}) be the model output value after using the subset S and adding feature i, and v(S) be the model output using the subset S; the prediction decomposition is as follows: , where g(x′) is the output value of the explanatory model, and x′ is the simplified feature vector (binary representation of whether the feature exists). Let M be a basic constant and M be the number of simplified features. Let xi be the Shapley value of the i-th feature, and xi′ be the value of the i-th simplified feature (1 indicates its presence, 0 indicates its absence).

5. The method according to claim 2, characterized in that, The Transformer model employs an encoder-decoder structure, with a multi-head attention mechanism at its core. The attention calculation formula is as follows: Where Attention(Q,K,V) is the output of the attention mechanism, Q is the query matrix, K is the key matrix, and V is the value matrix. This is the transpose of the key matrix. The dimension of the key matrix is ​​given, and softmax is the softmax activation function; positional encoding uses... , .in and These are the position encoding values ​​of position pos in even-numbered dimensions 2i and odd-numbered dimensions 2i+1, respectively, where pos is the position index in the sequence and i is the dimension index. For model dimensions.

6. The method according to claim 2, characterized in that, The Transformer-MoE nested model embeds a Hybrid Expert (MoE) model within a Transformer. The operational steps include: positional encoding and feature extraction of the input samples; calculation of the query, key, and value matrices, followed by scaling dot product attention; concatenation of multi-head attention after parallel computation; input to a feedforward network (FFN); and embedding the MoE model, with the gated network formula as follows: Where G(x) is the output weight vector of the gating network, x is the input token, Wg is the weight matrix of the gating network, and softmax is the softmax activation function; sparse routing uses the noisy top k selection, as shown in the formula. , ϵ∼ ,in For the input after adding noise, Let ϵ be the original input token, and ϵ be Gaussian random noise. The mean is μ and the variance is μ. The normal distribution is used; the MoE output is used as FFN compensation for residual connections and layer normalization, as shown in the formula. , ,in The result of layer normalization (LayerNorm) is the residual connection between the multi-head attention output z and the MoE output MoE(z), where z′ is... With FFN output FFN( The final result after residual connection and layer normalization is as follows: LayerNorm is the layer normalization function, z is the output of the multi-head attention mechanism, MoE(z) is the output of the MoE layer, and FFN( ) represents the output of the feedforward network.

7. The method according to claim 1, characterized in that, The specific method for constructing the three-theme parameter inversion multi-source database is as follows: taking surface temperature and surface emissivity as the first theme, and combining them with atmospheric water vapor content, near-surface air temperature and their auxiliary data sources to construct the three-theme parameter inversion multi-source database, and refining and correcting the original data through the RM mechanism.

8. The method according to claim 1, characterized in that, The geophysical logical reasoning is based on the radiative transfer theory to obtain the radiance of the top of the atmosphere, which includes the surface radiation term, the upward atmospheric thermal radiation term, and the upward atmospheric scattered thermal radiation term.

9. The method according to claim 8, characterized in that, The surface radiation term is calculated based on surface emissivity, surface temperature, Planck function, atmospheric downward thermal radiation term, and atmospheric scattering downward solar radiation term.

10. The method according to claim 1, characterized in that, In the inversion process, the effective value of the surface emissivity is calculated to obtain the effective surface emissivity, which is expressed as follows: Where λ1 and λ2 are the lower and upper limits of the band, respectively, and Ts is the surface temperature. Let λ be the surface emissivity at wavelength λ. Let i be the normalized spectral response function for band i. Let T be the amount of radiation emitted by a blackbody at temperature Ts.

11. The method according to claim 1, characterized in that, It also includes a step of verifying the inversion results, the verification method including: verifying the theoretical accuracy RMSE < 1K through simulation data verification, cross-validating with satellite derivatives or reanalysis datasets MAE < 1.5K, and collecting measured values ​​for verification.

12. The method according to claim 11, characterized in that, The process of collecting and verifying the measured values ​​includes data cleaning. The data cleaning method is to remove outliers that differ from the measured values ​​by using three times the theoretical accuracy as a threshold.

13. A remote sensing multi-parameter integrated inversion paradigm system based on AI-Agent, characterized in that, For implementing the method according to any one of claims 1-12, the system includes a DL-C-PSK paradigm module and an integrated inversion module, wherein the DL-C-PSK module optimizes the inversion mode selection of the integrated inversion module through real-time feedback from an AI-Agent; The DL-C-PSK paradigm module is used to construct the DL-C-PSK paradigm by coupling an AI-Agent-driven deep learning neural network with a nested RM-Transformer-MoE model, physical methods, statistical methods, and knowledge. Based on the DL-C-PSK paradigm, a high-precision multi-source database is constructed. Based on geophysical logical reasoning, a well-defined radiative transfer equation is constructed. Based on the high-precision multi-source database and the radiative transfer equation, remote sensing multi-parameters are inverted to obtain atmospheric parameters, including surface temperature, surface emissivity, atmospheric water vapor content, and near-surface air temperature. Based on the atmospheric parameters, a three-theme parameter inversion multi-source database is constructed. The integrated inversion module includes a direct inversion unit and an iterative inversion unit. The direct inversion unit is used to directly and synchronously invert the atmospheric parameters through the radiative transfer equation when the multi-source database of the three thematic parameters satisfies the dominant band combination, thereby obtaining the final atmospheric parameters. The iterative inversion unit is used to add prior knowledge from previous inversions to assist in iterative inversion when the multi-source database of the three thematic parameters does not satisfy the dominant band combination, thereby obtaining the final atmospheric parameters.

14. An electronic device, characterized in that, The device includes a processor and a memory storing computer program instructions; the processor, when executing the computer program instructions, implements the method as described in any one of claims 1-12.

Citation Information

Patent Citations

  • Remote sensing multi-parameter AI integrated inversion normal form method, system and equipment

    CN118797218A

Cited By

  • Reversible neural network based remote sensing calibration correction bidirectional coupling deep learning method

    CN122199299A