Interpretation method and device for star-ground collaborative remote sensing image based on causal reasoning

By constructing a causal variable graph on the ground and updating the satellite-borne model through knowledge distillation, the problems of interpretation accuracy and environmental cause inference in the interpretation of satellite-ground collaborative remote sensing images were solved, achieving dynamic coordination and high-precision interpretation.

CN121121532BActive Publication Date: 2026-02-24XINGHAN SPACE TIME (SHENZHEN) AEROSPACE INTELLIGENT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511667610.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-14
Publication Date
2026-02-24
Estimated Expiration
2045-11-14

AI Technical Summary

Technical Problem

In existing satellite-ground collaborative remote sensing image interpretation, the ground end cannot coordinate dynamically with the satellite end in a timely manner, which affects the accuracy of interpretation and cannot meet the needs of inferring environmental causes.

Method used

Initial interpretation features of multimodal remote sensing data from the ground-based terminal are acquired. A causal variable graph is constructed through causal reasoning to determine the causal labels and interpretation results of the labeled regions. The interpretation model of the spaceborne terminal is updated through knowledge distillation to achieve dynamic coordinated training.

Benefits of technology

The accuracy of interpretation of spaceborne remote sensing data has been optimized, meeting the need for inferring environmental causes in remote sensing images and improving the interpretability and system adaptability of the interpretation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121121532B_ABST
    Figure CN121121532B_ABST
Patent Text Reader

Abstract

The application discloses a star-ground collaborative remote sensing image interpretation method and device based on causal reasoning, relates to the technical field of remote sensing image processing, and mainly aims to solve the problem that the existing method cannot meet the inference demand of the environment cause in the remote sensing image. The method comprises the following steps: a ground terminal obtains initial interpretation features obtained by performing initial interpretation on multi-modal remote sensing data by a spaceborne terminal, and performs interpretation processing on the initial interpretation features based on a first interpretation model that has completed model training, to obtain core interpretation features; a causal variable is determined based on the core interpretation features, and a causal variable graph is constructed based on the causal variable; a cause label of a label region and an interpretation result are determined according to the causal variable graph, and a verification result corresponding to the cause label and the interpretation result is obtained; learning core features of the first interpretation model are determined based on knowledge distillation and the verification result to update the training, so that a second interpretation model in the spaceborne terminal is coordinated based on the learning core features fed back by the ground terminal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of remote sensing image processing technology, and in particular to a method and apparatus for interpreting satellite-ground collaborative remote sensing images based on causal reasoning. Background Technology

[0002] With the development of aerospace technology, space-ground collaborative remote sensing has become an important means of remote sensing monitoring. Through space-ground collaborative interpretation, real-time acquisition, transmission, analysis, and application of remote sensing images can be achieved. At this point, by linking the spaceborne and ground-based ends, and combining the high coverage of space-based data with the high precision of ground processing, a multi-dimensional interpretation methodology can be formed.

[0003] Currently, when interpreting satellite-ground collaborative remote sensing images, both the satellite-based and ground-based terminals can utilize artificial intelligence and other technologies to process and interpret the images. However, the ground-based terminals cannot dynamically coordinate with the satellite-based terminals in a timely manner during the interpretation process. When the monitored scene changes, the satellite-based terminals cannot quickly retrieve effective learning samples, affecting the accuracy of the interpretation. Furthermore, when generating interpretation content for remote sensing images, only conclusive information is obtained, which cannot meet the need for inferring the environmental causes in the remote sensing images. Therefore, there is an urgent need for a satellite-ground collaborative remote sensing image interpretation method based on causal reasoning to solve the above problems. Summary of the Invention

[0004] In view of this, this application provides a method and apparatus for interpreting satellite-ground collaborative remote sensing images based on causal reasoning. The main purpose is to solve the problem that the interpretation accuracy of existing remote sensing images is poor and cannot meet the inference needs of environmental causes in remote sensing images.

[0005] According to one aspect of this application, a method for interpreting satellite-ground collaborative remote sensing images based on causal reasoning is provided, applied at the ground end, including:

[0006] The initial interpretation features obtained by the satellite-borne terminal in the initial interpretation of multimodal remote sensing data are obtained, and the initial interpretation features are processed by the first interpretation model that has completed model training to obtain the core interpretation features.

[0007] Causal variables are determined based on the core interpretation features, and a causal variable graph is constructed based on the causal variables. The causal variables are used to characterize the variables that affect the monitoring results in multimodal remote sensing data. The causal variable graph contains the proportion of the influence of different causal variables on different causal associations.

[0008] The causal labels and interpretation results of the labeled regions are determined based on the causal variable diagram, and the verification results corresponding to the causal labels and interpretation results are obtained.

[0009] If the verification result is successful, the learning core features of the first interpretation model are determined based on knowledge distillation and the core interpretation features, and the learning core features are sent to the satellite terminal so that the second interpretation model in the satellite terminal can perform coordinated training and interpretation.

[0010] Furthermore, the step of determining causal variables based on the core interpretation features and constructing a causal variable graph based on the causal variables includes:

[0011] Causal variables are selected according to a preset variable extraction mapping relationship, which includes the matching relationship between different core interpretation features and at least one of different physical variables, different quantitative variables and different correlation variables.

[0012] Causal relationships are determined based on historical statistical data and geographical knowledge, and an initial causal variable graph is constructed based on the causal variables and the causal relationships.

[0013] The influence ratio of each causal variable on different causal associations is determined by causal intervention verification, and the influence ratio is configured in the initial causal variable diagram to obtain the causal variable diagram.

[0014] Furthermore, the determination of the proportion of influence of each causal variable in different causal associations through causal intervention verification includes:

[0015] The backdoor path is used to filter the causal variables of each causal variable, and the causal variables obtained after the filtering are determined to be causally related, so as to determine the proportion of the influence.

[0016] Furthermore, after determining the proportion of influence of each causal variable on different causal associations through causal intervention verification and configuring the proportion of influence in the initial causal variable graph to obtain the causal variable graph, the method further includes:

[0017] The influence percentages of each causal variable are ranked, and the interpretation results of the influencing factors are determined based on the ranking results.

[0018] Furthermore, the multimodal remote sensing data includes optical image data, infrared image data, and SAR image data. The step of determining the causal labels and interpretation results of the labeled regions based on the causal variable map, and obtaining the verification results corresponding to the causal labels and interpretation results, includes:

[0019] The label region corresponding to the multimodal remote sensing data is determined, and the causal labels and interpretation results of the label region are generated based on the interpretation results of the influencing factors. The label region is a feature region corresponding to the core interpretation feature. The causal labels are used to mark the causal variables that affect the core interpretation feature, and the interpretation results are used to characterize the influence of the causal variables.

[0020] Furthermore, obtaining the verification result corresponding to the cause label and the interpretation result includes:

[0021] Output the cause label and the interpretation result, and receive the verification result from the user verifying the cause label and the interpretation result.

[0022] According to one aspect of this application, another method for interpreting satellite-ground collaborative remote sensing images based on causal reasoning is provided, applied to a satellite-borne terminal, including:

[0023] Multimodal remote sensing data is acquired, and the multimodal features of the multimodal remote sensing data are fused to obtain a multimodal feature vector;

[0024] The multimodal feature vector is interpreted based on the first interpretation model to generate interpretation features, which are then sent to the ground terminal. The first interpretation model is a lightweight deep learning model.

[0025] The first interpretation model is obtained through coordinated training based on the core features learned from ground-based feedback. The core features are determined by the second interpretation model in the ground-based system after interpretation, and verified by a causal variable graph. The causal variable graph is constructed based on the core interpretation features obtained after interpretation. The causal variable graph contains the influence ratio of different causal variables on different causal associations. The causal variables are used to characterize the variables that affect the monitoring results in multimodal remote sensing data.

[0026] Furthermore, the multimodal remote sensing data includes optical image data, infrared image data, and SAR image data. Before performing feature fusion on the multimodal features of the multimodal remote sensing data to obtain a multimodal feature vector, the method further includes:

[0027] Texture spatial variation information is extracted from optical image data that has undergone radiometric and geometric corrections to obtain a gradient intensity map. The gradient mean, gradient variance, and gradient entropy of the gradient intensity map are then combined to obtain texture features.

[0028] SAR image data that has been radiometrically calibrated and noise-removed is filtered and dimensionality reduced to obtain a backscatter map. The backscatter map is then linearly projected to obtain the backscatter coefficients.

[0029] Temperature inversion is performed on infrared image data calibrated by thermal infrared band to obtain a surface temperature map, and the surface temperature map is linearly mapped to obtain surface temperature characteristics.

[0030] The texture features, the backscattering coefficient, and the surface temperature features have the same dimension.

[0031] Furthermore, the feature fusion of the multimodal features of the multimodal remote sensing data to obtain the multimodal feature vector includes:

[0032] Multimodal similarity matrices are generated for the texture features, backscattering coefficients, and surface temperature features, respectively. Feature fusion is then performed based on the feature associations and modal weights of the multimodal similarity matrices to obtain a multimodal feature vector. The modal weights are dynamically configured based on different scenarios.

[0033] Furthermore, the feature fusion based on the feature association and modality weights of the multimodal similarity matrix to obtain the multimodal feature vector includes:

[0034] Feature associations are determined based on the multimodal similarity matrix, and these feature associations are used to characterize the association strength between features of different modalities.

[0035] Modal weights are assigned according to different scenarios, including normal scenarios, abnormal high temperature scenarios, and cloud and fog coverage scenarios.

[0036] The feature associations and modal weights are calculated using a weighted summation method to obtain a multimodal feature vector.

[0037] Furthermore, before allocating modal weights according to different scenarios, the method further includes:

[0038] The feature association consistency index is calculated based on the multimodal similarity matrix, and the mean temperature and standard deviation of the temperature in the whole region are calculated based on the surface temperature feature vector. The surface temperature feature vector is extracted after inversion of the infrared image data.

[0039] When the feature association consistency index, the average temperature of the entire region, and the standard deviation of the temperature of the entire region match the preset normal scenario conditions, the scenario is determined to be a normal scenario.

[0040] Furthermore, before allocating modal weights according to different scenarios, the method further includes:

[0041] Based on the surface temperature characteristics, calculate the temperature value and standard deviation of the sub-region, and mark the candidate abnormal region when the temperature value and standard deviation of the sub-region match the preset temperature conditions.

[0042] Once the texture features and the backscattering coefficient are verified, the candidate anomaly region is determined to be a high-temperature anomaly scene.

[0043] Furthermore, before allocating modal weights according to different scenarios, the method further includes:

[0044] The suspected cloud and fog coverage areas were determined based on the average gray value and contrast of the surface temperature characteristics.

[0045] Once it is determined that there is cloud and fog interference in the suspected cloud and fog coverage area based on the gradient entropy of the texture features, the consistency index of the backscattering coefficient and the land cover of the texture features is calculated.

[0046] When the land cover consistency index matches the preset cloud and fog scene conditions, the scene is determined to be a cloud and fog coverage scene.

[0047] According to one aspect of this application, an interpretation apparatus for satellite-ground collaborative remote sensing images based on causal reasoning is provided, comprising:

[0048] The acquisition module is used to acquire the initial interpretation features obtained by the satellite-borne terminal in the initial interpretation of multimodal remote sensing data, and to perform interpretation processing on the initial interpretation features based on the first interpretation model that has completed model training to obtain the core interpretation features;

[0049] A construction module is used to determine causal variables based on the core interpretation features and construct a causal variable graph based on the causal variables. The causal variables are used to characterize the variables that affect the monitoring results in multimodal remote sensing data. The causal variable graph contains the proportion of the influence of different causal variables on different causal associations.

[0050] The determination module is used to determine the causal labels and interpretation results of the label regions based on the causal variable diagram, and to obtain the verification results corresponding to the causal labels and interpretation results;

[0051] The sending module is configured to, when the verification result is passed, determine the learning core features of the first interpretation model based on knowledge distillation and the core interpretation features, and send the learning core features to the satellite terminal so that the second interpretation model in the satellite terminal can perform coordinated training and interpretation.

[0052] Furthermore,

[0053] The construction module is specifically used to select causal variables according to a preset variable extraction mapping relationship, which includes the matching relationship between different core interpretation features and at least one of different physical variables, different quantitative variables, and different correlation variables; determine causal associations based on historical statistical data and geographical knowledge information, and construct an initial causal variable graph based on the causal variables and the causal associations; determine the influence ratio of each causal variable on different causal associations through causal intervention verification, and configure the influence ratio in the initial causal variable graph to obtain the causal variable graph.

[0054] Furthermore, the device also includes:

[0055] The filtering module is used to filter the causal variables through a backdoor path, and determine the causal relationships based on the filtered causal variables to determine the proportion of influence.

[0056] Furthermore, the determining module is also used to sort the influence proportions corresponding to each of the causal variables, and determine the interpretation result of the influencing factors based on the sorting result.

[0057] Furthermore, the multimodal remote sensing data includes optical image data, infrared image data, and SAR image data.

[0058] The determining module is further configured to determine the label region corresponding to the multimodal remote sensing data, and generate the causal label of the label region and the interpretation result based on the interpretation result of the influencing factors. The label region is a feature region corresponding to the core interpretation feature. The causal label is used to mark the causal variables that affect the core interpretation feature. The interpretation result is used to characterize the influence result of the causal variables.

[0059] Furthermore, the device also includes:

[0060] The output module is used to output the cause label and the interpretation result, and to receive the verification result of the user verifying the cause label and the interpretation result.

[0061] According to one aspect of this application, another interpretation apparatus for satellite-ground collaborative remote sensing images based on causal reasoning is provided, comprising:

[0062] The acquisition module is used to acquire multimodal remote sensing data and perform feature fusion on the multimodal features of the multimodal remote sensing data to obtain a multimodal feature vector.

[0063] The sending module is used to interpret the multimodal feature vector based on the first interpretation model, generate interpretation features, and send them to the ground terminal. The first interpretation model is a lightweight deep learning model.

[0064] The first interpretation model is obtained through coordinated training based on the core features learned from ground-based feedback. The core features are determined by the second interpretation model in the ground-based system after interpretation, and verified by a causal variable graph. The causal variable graph is constructed based on the core interpretation features obtained after interpretation. The causal variable graph contains the influence ratio of different causal variables on different causal associations. The causal variables are used to characterize the variables that affect the monitoring results in multimodal remote sensing data.

[0065] Furthermore, the multimodal remote sensing data includes optical image data, infrared image data, and SAR image data, and the device further includes:

[0066] The extraction module is used to extract texture space variation information from optical image data that has undergone radiometric and geometric corrections, obtain a gradient intensity map, and merge the gradient mean, gradient variance, and gradient entropy of the gradient intensity map to obtain texture features.

[0067] The projection module is used to filter and reduce the dimensionality of SAR image data after radiometric calibration and noise removal to obtain a backscatter map, and to perform linear projection on the backscatter map to obtain the backscatter coefficients.

[0068] The mapping module is used to perform temperature inversion on infrared image data calibrated by thermal infrared band radiometry to obtain a surface temperature map, and to perform linear mapping on the surface temperature map to obtain surface temperature characteristics.

[0069] The texture features, the backscattering coefficient, and the surface temperature features have the same feature dimension.

[0070] Furthermore, the acquisition module is specifically used to generate multimodal similarity matrices corresponding to the texture features, the backscattering coefficients, and the surface temperature features, respectively, and to perform feature fusion based on the feature association and modal weights of the multimodal similarity matrices to obtain multimodal feature vectors, wherein the modal weights are dynamically configured based on different scenarios.

[0071] Furthermore, the acquisition module is specifically used to determine feature associations based on the multimodal similarity matrix, wherein the feature associations are used to characterize the association strength between different modal features; to assign modal weights according to different scenarios, wherein the scenarios include normal scenarios, high temperature abnormal scenarios, and cloud and fog coverage scenarios; and to calculate the feature associations and the modal weights based on a weighted summation method to obtain a multimodal feature vector.

[0072] Furthermore, the device also includes:

[0073] The calculation module is used to calculate the feature association consistency index based on the multimodal similarity matrix, and to calculate the mean temperature and standard deviation of the temperature in the whole region based on the surface temperature feature vector, wherein the surface temperature feature vector is extracted by inversion of the infrared image data;

[0074] The determination module is used to determine the scenario as a normal scenario when the feature association consistency index, the average temperature of the whole region, and the standard deviation of the temperature of the whole region match the preset normal scenario conditions.

[0075] Furthermore,

[0076] The calculation module is also used to calculate the temperature value of the sub-region and the standard deviation of the sub-region based on the surface temperature characteristics, and to mark the candidate abnormal region when the temperature value of the sub-region and the standard deviation of the sub-region match the preset temperature conditions.

[0077] The determining module is further configured to determine the candidate abnormal region as a high-temperature abnormal scene after the texture features and the backscattering coefficient have been verified.

[0078] Furthermore,

[0079] The determining module is also used to determine suspected cloud and fog coverage areas based on the average gray value and contrast of the surface temperature characteristics.

[0080] The calculation module is also used to calculate the consistency index between the backscattering coefficient and the land cover of the texture feature after determining that there is cloud and fog interference in the suspected cloud and fog coverage area based on the gradient entropy of the texture feature.

[0081] The determining module is further configured to determine the scene as a cloud and fog coverage scene when the land cover consistency index matches the preset cloud and fog scene conditions.

[0082] According to one aspect of this application, a storage medium is provided, wherein at least one executable instruction is stored therein, the executable instruction causing a processor to perform operations corresponding to the above-described method for interpreting satellite-ground collaborative remote sensing images based on causal reasoning.

[0083] According to one aspect of this application, a terminal is provided, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other through the communication bus;

[0084] The memory is used to store at least one executable instruction, which causes the processor to perform the operation corresponding to the above-described interpretation method of satellite-ground collaborative remote sensing images based on causal reasoning.

[0085] According to one aspect of this application, another storage medium is provided, wherein at least one executable instruction is stored therein, the executable instruction causing a processor to perform operations corresponding to the above-described method for interpreting satellite-ground collaborative remote sensing images based on causal reasoning.

[0086] According to one aspect of this application, another terminal is provided, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus;

[0087] The memory is used to store at least one executable instruction, which causes the processor to perform the operation corresponding to the above-described interpretation method of satellite-ground collaborative remote sensing images based on causal reasoning.

[0088] By employing the above technical solutions, the technical solutions provided in the embodiments of this application have at least the following advantages:

[0089] This application provides a method and apparatus for interpreting remote sensing images based on causal reasoning and space-ground collaboration. Compared with the prior art, the embodiments of this application obtain initial interpretation features from the initial interpretation of multimodal remote sensing data by the spaceborne end at the ground end, and perform interpretation processing on the initial interpretation features based on the first interpretation model that has completed model training to obtain core interpretation features; determine causal variables based on the core interpretation features, and construct a causal variable graph based on the causal variables; determine the causal labels and interpretation results of the labeled regions according to the causal variable graph, and obtain the verification results corresponding to the causal labels and interpretation results; determine the learning core features for updating and training the first interpretation model based on knowledge distillation and verification results, and send the learning core features to the spaceborne end so that the second interpretation model in the spaceborne end is obtained by coordinated training based on the learning core features fed back from the ground end. This achieves the purpose of timely dynamic coordinated training between the ground end and the spaceborne end, optimizes the interpretation accuracy of remote sensing data at the spaceborne end, can meet the inference needs of environmental causes in remote sensing images in a diversified and precise manner, significantly improves the interpretability of the interpretation results and the adaptability of the spaceborne-ground system, and greatly increases the diversity of applicable scenarios for remote sensing monitoring.

[0090] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0091] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0092] Figure 1 A flowchart of a satellite-ground collaborative remote sensing image interpretation method based on causal reasoning, provided in an embodiment of this application, is shown.

[0093] Figure 2 This illustration shows a schematic diagram of the interaction process between a large cloud model on the ground and a satellite-based terminal, according to an embodiment of this application.

[0094] Figure 3 This illustration shows a schematic diagram of a ground-based depth interpretation process provided in an embodiment of this application;

[0095] Figure 4 A flowchart of another method for interpreting satellite-ground collaborative remote sensing images based on causal reasoning, provided in an embodiment of this application, is shown.

[0096] Figure 5 This illustration shows a timing diagram of the interaction between a large cloud model on the ground and a satellite-based terminal, according to an embodiment of this application.

[0097] Figure 6 This illustration shows a schematic diagram of data interaction between a large cloud model on the ground and a satellite-based terminal, provided in an embodiment of this application.

[0098] Figure 7 This paper illustrates a block diagram of a satellite-ground collaborative remote sensing image interpretation device based on causal reasoning, according to an embodiment of this application.

[0099] Figure 8 This paper illustrates a block diagram of another satellite-ground collaborative remote sensing image interpretation device based on causal reasoning, provided in an embodiment of this application.

[0100] Figure 9 This illustration shows a structural diagram of a terminal provided in an embodiment of this application;

[0101] Figure 10 A schematic diagram of the structure of a terminal provided in an embodiment of this application is shown. Detailed Implementation

[0102] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0103] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0104] This application provides a method for interpreting satellite-ground collaborative remote sensing images based on causal reasoning, applied to the ground end, such as... Figure 1 As shown, the method includes:

[0105] 101. Obtain the initial interpretation features obtained by the satellite-borne terminal in the initial interpretation of multimodal remote sensing data, and perform interpretation processing on the initial interpretation features based on the first interpretation model that has completed model training to obtain the core interpretation features.

[0106] In this embodiment, the current execution entity is the ground end, which is the device on the Earth side corresponding to the spaceborne end. It can be a cloud service device, etc., and this embodiment does not specify any particular limitation. The multimodal remote sensing data includes optical image data, infrared image data, and SAR image data. Optical image data is an image formed by passively receiving sunlight (visible light, near-infrared, etc.) reflected from the Earth's surface. Optical image data depends on lighting conditions, is greatly affected by weather, cannot work at night, and is blocked by clouds. It has high resolution and is suitable for high-precision map production, land classification, and disaster assessment, such as visual interpretation of flood inundation areas. Infrared image data is an image formed by detecting surface thermal radiation (such as mid-infrared and far-infrared). In this case, no lighting is required, and it can directly reflect temperature anomalies. It has lower resolution, is sensitive to ambient temperature, and is mainly used for fire monitoring, volcano early warning, and urban heat island effect research. SAR image data is images acquired by Synthetic Aperture Radar (SAR). It is an image formed by actively emitting microwave pulses and receiving echo signals. It can work in all weather conditions and penetrate clouds, meaning it can work around the clock and in all weather. It has strong penetration and is not limited by weather. However, it has low resolution and low signal-to-noise ratio. It is often used for geological disaster monitoring (earthquake deformation), military reconnaissance, and marine monitoring.

[0107] It should be noted that, due to limited computing resources, the interpretation model on the spaceborne end is small and has few parameters. Therefore, in this embodiment, a lightweight interpretation model is used on the spaceborne end. After acquiring multimodal remote sensing data, the spaceborne end performs initial interpretation to obtain preliminary interpretation features. The ground-based execution unit receives the initial interpretation features from the spaceborne end and processes them based on a first interpretation model that has completed model training to obtain core interpretation features. At this point, the ground-based end has abundant hardware resources, high computing power, a large model, and many parameters. Therefore, the first interpretation model can be an interpretation model that has been accurately trained with a large number of samples to perform secondary interpretation and obtain core interpretation features, such as features related to floods and fires.

[0108] In some embodiments, the first interpretation model on the ground serves as a large model in the cloud, providing deep and accurate interpretation services. Relying on the massive computing power and storage resources of the ground server, the first interpretation model's core objective is to perform high-precision, multi-dimensional, deep interpretation of the core interpretation features of the key data transmitted from the spaceborne terminal, thereby providing final decision support. It can support complex tasks, including fine-grained land cover classification (distinguishing between trees / shrubs / grasslands, industrial buildings / civilian buildings), dynamic change analysis (such as fire spread trend prediction and crop growth status assessment), and multi-source data fusion inference (combining historical remote sensing data, meteorological data, and geographic information data). Simultaneously, it supports high-precision results, such as land cover classification accuracy exceeding 90% and anomaly detection recall rate ≥95%, thus providing directly applicable and accurate results for industry applications (such as emergency disaster relief, agricultural monitoring, and ecological assessment). In addition, the ground-based interpreter, serving as the primary interpreter for the large cloud-based model, prioritizes performance limits. Since there are no resource constraints on the ground, the computing power of the cloud server's GPU cluster (such as multiple A100 GPUs) and petabyte-level storage can be fully utilized. This is characterized by the following three points:

[0109] 1. Large number of parameters: The number of parameters is usually in the billions (10B-100B). For example, the "RemoteSensingGPT" model, which is designed for remote sensing, has about 50B of parameters to learn the complex feature associations in massive remote sensing data, such as the dynamic relationship between vegetation texture and temperature in different seasons.

[0110] 2. High computational complexity: It employs complex structures such as full-size Transformer self-attention, multi-scale feature fusion, and cross-modal causal reasoning. The processing time for a single anomalous region image may reach several minutes, but it only processes the core interpretation features selected as key data by the satellite end, resulting in higher overall efficiency.

[0111] 3. High dependence requires the integration of multi-source data, including key features transmitted from the spacecraft, historical remote sensing image databases stored on the ground (data from the same area over the past 10 years), meteorological data (temperature / precipitation), and geographic information data (administrative divisions, topography), to achieve data-driven deep interpretation combined with knowledge reasoning.

[0112] In some embodiments, to achieve deep interpretation, the first interpretation model employs a complex architecture combined with continuous learning, possessing adaptive optimization capabilities. The complex architecture typically comprises four layers: a multimodal feature encoding layer (processing optical / SAR / infrared / meteorological data), a cross-modal attention fusion layer (establishing dynamic correlations between different data), a causal reasoning layer (analyzing causal relationships of ground cover changes, such as the correlation between fire and vegetation type), and a fine-grained decoding layer (outputting pixel-level interpretation results). Dynamic learning means that the model can be continuously optimized through incremental training (e.g., fine-tuning the model with each additional month of remote sensing data). Scene determination no longer relies on fixed thresholds but instead autonomously learns the distribution of normal / abnormal features in different regions and seasons through the optimized first interpretation model. For example, the high-temperature threshold in tropical regions is autonomously adjusted to 45℃, and in frigid regions to 35℃. This embodiment does not impose specific limitations.

[0113] 102. Determine causal variables based on the core interpretation features, and construct a causal variable graph based on the causal variables.

[0114] In this embodiment, after obtaining the core interpretation features, the ground terminal, as the current execution terminal, determines the causal variables. These causal variables characterize the variables affecting the monitoring results in the multimodal remote sensing data, including but not limited to rainfall, temperature, slope, soil permeability, and vegetation cover. This embodiment does not impose specific limitations on these variables. Furthermore, a causal variable graph is constructed based on the determined causal variables. The causal variable graph contains the proportion of influence of different causal variables on different causal associations. Causal associations characterize the correlation between different causal variables. The causal variable graph characterizes the influence between causal variables. This embodiment does not impose specific limitations on these variables.

[0115] 103. Determine the causal labels and interpretation results of the label regions based on the causal variable diagram, and obtain the verification results corresponding to the causal labels and interpretation results.

[0116] In this embodiment, to optimize the interpretation of causes, the ground terminal determines the causal labels and interpretation results of the labeled regions based on causal variables. The labeled regions are the image regions whose causes need to be interpreted. These labels can be manually assigned or based on feature extraction; this embodiment does not impose specific limitations. The causal labels can represent the interpretation causes that lead to the visual results in the image. Correspondingly, the interpretation results are the variable results derived from the causal variable graph that cause the aforementioned causes, forming a visual result output to the front-end interface of the current ground terminal, allowing the ground terminal user to verify and obtain the verification results.

[0117] 104. When the verification result is that the verification is passed, the learning core features of the first interpretation model are determined based on knowledge distillation and the core interpretation features, and the learning core features are sent to the satellite terminal.

[0118] In this embodiment, after receiving the verification result from the user, the current execution terminal determines the core learning features of the first interpretation model based on knowledge distillation, such as the water scattering features and terrain gradient features obtained from the distillation. This embodiment does not specifically limit the features. Knowledge distillation is a model compression technique that transfers knowledge from a large, complex model (the first interpretation model) to a small, lightweight model (the second interpretation model) to reduce computational resource requirements while maintaining performance. Figure 2 As shown, this enables the second interpretation model in the spaceborne terminal to perform coordinated training and interpretation.

[0119] It should be noted that the first interpretation model in the ground end (such as the Transformer model) obtains core interpretation features (which are already verified) through transfer and causal reasoning during the distillation process based on knowledge distillation technology. These features include water scattering features of floods and topographic gradient features of landslides. This application does not specifically limit these features in the embodiments.

[0120] In some implementation examples, the first interpretation model used for flood monitoring, as a large cloud-based model, can preserve the SAR backscattering characteristics of water bodies and the Normalized Difference Water Index (NDWI) of optical images. The first interpretation model used for landslide monitoring, as a large cloud-based model, can preserve the gradient characteristics of surface deformation and the abrupt changes in vegetation cover. The loss function in the first interpretation model in this application embodiment is expressed as:

[0121] ;

[0122] in, For classifying losses, Characteristic distillation loss, This is a loss of causal characteristics. This indicates that the cosine similarity between the feature vector output by the second interpretation model on the satellite-borne end and the first interpretation model on the ground end is ≥0.9. This indicates that key features for causal reasoning, such as the correlation between slope and rainfall, are assigned higher weights. =0.4, =0.3, =0.3. At this point, the performance comparison of the models before and after distillation is as follows: The cloud-based large model used on the ground side is preferably the SegViT model, with 80M parameters and 93% accuracy. The lightweight second interpreter model on the spaceborne side, after distillation, has 5.2M parameters and 91% accuracy, a loss of 2% in accuracy. When the scene changes from plain flooding to mountain landslides, the accuracy of the second interpreter model on the spaceborne side before distillation decreases from 89% to 62%, a decrease of 27%, and the accuracy after distillation decreases from 89% to 82%, a decrease of 7%.

[0123] In another embodiment of this application, for further definition and explanation, the steps of determining causal variables based on the core interpretation features and constructing a causal variable graph based on the causal variables include:

[0124] Select causal variables according to the pre-defined variable extraction mapping relationship;

[0125] Causal relationships are determined based on historical statistical data and geographical knowledge, and an initial causal variable graph is constructed based on the causal variables and the causal relationships.

[0126] The influence ratio of each causal variable on different causal associations is determined by causal intervention verification, and the influence ratio is configured in the initial causal variable diagram to obtain the causal variable diagram.

[0127] To achieve the goal of interpreting the causes of remote sensing data, the ground-based execution unit first selects causal variables according to a preset variable extraction mapping relationship when determining causal variables. This preset variable extraction mapping relationship includes matching relationships between different core interpretation features and at least one of different physical variables, different quantified variables, and different correlation variables. Physical variables are those directly physically related to changes in the remote sensing target, such as temperature and wind speed. Quantified variables are those that affect the remote sensing target after quantification, extracted through multimodal data quantification, such as the intensity of human activity, quantified through nighttime light and building features in SAR image data. Correlation variables are those with a significant correlation to the core interpretation features in historical data, such as variables with a correlation coefficient greater than 0.6. This application embodiment does not impose specific limitations on these variables. Furthermore, the preset variable extraction mapping relationship can be pre-configured based on the precise and diverse scenario requirements for interpreting the causes of remote sensing targets, allowing for direct selection and matching. For example, observable and quantifiable geographical factors directly related to changes in the remote sensing target can be selected as preset variable extraction mapping relationships. This relationship can be stored in list form. This application embodiment does not impose specific limitations on these relationships. In flood monitoring scenarios, variables that match the core interpretation features of floods in the preset variable extraction mapping relationship include rainfall, terrain slope, soil permeability, vegetation coverage, and intensity of human activities, etc., and this application embodiment does not make specific limitations.

[0128] In some embodiments, when configuring the preset variable extraction mapping relationship, restrictions such as required variables, optional variables, and excluded variables can also be set. For example, required variables include rainfall (continuous value, unit mm), terrain slope (continuous value, unit °), and soil permeability (discrete value, divided into high / medium / low); optional variables include vegetation coverage (continuous value, unit %) and human activity intensity (discrete value, divided into high / medium / low based on nighttime light data); excluded variables include "regional GDP" (which has no direct causal relationship with flood changes) and "distance from the administrative center" (a non-geographical driving factor), etc. The embodiments of this application do not impose specific limitations.

[0129] It should be noted that after obtaining the causal variables, an initial causal variable graph is constructed based on the causal variables and causal relationships. At this time, the causal relationships are determined based on historical statistical data and geographical knowledge information. Geographical knowledge information is used to characterize the geographical conditions affecting remote sensing targets, and historical statistical data is data from historical remote sensing targets for causal interpretation. Correspondingly, when determining causal relationships, one can search for causal relationships related to the causal variables from historical statistical data and geographical knowledge information. At this time, the causal relationships can characterize the causal direction. For example, rainfall → flooding is a positive correlation. At this time, historical data shows that the probability of flooding increases by 80% when rainfall > 100mm; terrain slope → flooding is a negative correlation. The probability of flooding decreases by 60% when the slope > 15°; human activities (such as land reclamation from lakes) → flooding is a positive correlation. Areas with high activity intensity have a 30% increase in flood frequency. This application embodiment does not make specific limitations. Furthermore, an initial causal variable graph is constructed based on the causal variables and causal relationships, with the structure graph having variable nodes and causal relationships between nodes as edges.

[0130] For example, the core interpretation feature, as the remote sensing interpretation target, can be target Y-fire. Causal variable matching includes intervention variables X1-X4, where X1 represents infrared temperature features (a core causal variable; fire directly leads to increased temperature), X2 represents SAR backscattering coefficient (fire causes vegetation burning, reducing the scattering coefficient), X3 represents optical texture features (fire causes irregular abrupt changes in surface texture), and X4 represents meteorological data, such as wind speed, which indirectly affects fire spread but has a weaker direct impact on determining whether it is a fire. In this case, causal relationships are constructed based on domain knowledge, with arrows indicating direct causal effects, resulting in X1→Y (high temperature → more likely to be determined as a fire), X2→Y (low scattering coefficient → more likely to be determined as a fire), X3→Y (texture anomalies → more likely to be determined as a fire), and X4→X1 (wind speed affects temperature diffusion, indirectly affecting the effect of X1 on Y), to construct a causal variable graph.

[0131] In some embodiments, to accurately construct a causal variable graph for use as a basis for querying or filtering causal relationships, the ground-based system can determine the influence ratio of each causal variable on different causal associations through causal intervention verification, and configure the influence ratio in the initial causal variable graph to obtain the causal variable graph. Here, the influence ratio is used to characterize the proportion of influence of each causal relationship on the core interpretation feature, and can be configured in the initial causal variable graph using a labeling method to obtain the causal variable graph; this embodiment does not impose specific limitations. Specifically, when determining the influence ratio of each causal variable on different causal associations through causal intervention verification, the causal intervention verification "Do-Calculus" uses intervention operations (do (X=x)) to remove spurious associations between variables, calculates the average treatment effect (ATE) of each variable X on the target Y, and then normalizes the ATE into a comprehensive weight to obtain the influence ratio.

[0132] In a specific implementation scenario, rainfall exceeding 150 mm accounts for 60% of the impact on landslides, while a terrain slope of less than 5 degrees accounts for 30%. Determining the percentage of impact involves the following steps:

[0133] Step 1: Identify the factors that cause the landslide and construct a relationship diagram of the influencing variables;

[0134] The outcome variable (which can be understood as the core interpretable feature) is whether a landslide will occur. There are four factors influencing landslides: X1 = rainfall (the more rain, the greater the likelihood of a landslide), X2 = terrain slope (the steeper the slope, the greater the likelihood of a landslide), X3 = vegetation cover (the more grass and trees, the more soil stabilization, and the lower the likelihood of a landslide), and X4 = human excavation (such as digging into mountains to build houses, which increases the likelihood of a landslide). Therefore, the relationship between "factors → landslide" is plotted as an "influence relationship diagram," i.e., landslide ← rainfall (X1), terrain slope (X2), vegetation cover (X3), human excavation (X4), to visually determine which factors affect the outcome.

[0135] Step 2: Identify and remove "interference factors";

[0136] Some factors can interact and influence each other, reducing the accuracy of the individual factor's true impact on landslides. These become "distractors" and need to be eliminated first. For example, rainfall (X1) and terrain slope (X2) may influence each other through "regional climate zones." Rainy areas often have gentler terrain (plains have more rainfall and gentler slopes; mountainous areas have steeper slopes but may receive less rainfall). This can lead to the assumption that "X1 and X2 influence each other," resulting in incorrect causal inferences such as "more rainfall leads to gentler slopes" or "gentle slopes lead to more rainfall," thus miscalculating the landslide's impact percentage. Therefore, it's necessary to screen for such interactions. For instance, when studying the impact of rainfall on landslides, the terrain slope should be kept constant; when studying terrain slope, the rainfall should be kept constant, thereby improving the accuracy of calculating the true impact percentage of individual factors.

[0137] Step 3: Calculate the impact of each factor one by one, that is, calculate the degree of influence, as the percentage of influence.

[0138] In calculating the impact of rainfall (X1), the values ​​of the other three factors were fixed: terrain slope X2 = 25 degrees, vegetation coverage X3 = 30%, and human excavation X4 = 0 (no excavation). When the rainfall was changed from "50 mm (sparse rain)" to "150 mm (heavy rain)," the probability of a landslide increased from 10% to 60%, a rise of 50 percentage points. Therefore, rainfall (X1) accounts for 50% of the landslide impact. Similarly, in calculating terrain slope (X2), with the other three factors fixed, rainfall X1 = 100 mm, vegetation coverage X3 = 30%, and human excavation X4 = 0, the probability of a landslide increased from 40% to 70% when the terrain slope was changed from "10 degrees (gentle slope)" to "30 degrees (steep slope)," a rise of 30 percentage points. Therefore, terrain slope (X2) accounts for 30% of the landslide impact. Similarly, using the same method to calculate vegetation coverage X3 and human excavation X4, the final calculated impact percentages are 15% and 5%, respectively.

[0139] In this embodiment of the application, when the step calculates the influence magnitude through causal intervention verification as the influence ratio, the causal intervention verification Do-Calculus is to remove spurious associations between variables by intervention operation (do (X=x)), calculate the average intervention effect (ATE) of each X on Y, and finally normalize the ATE into a comprehensive weight.

[0140] In a scenario where fire detection is used as an example, such as Figure 3 As shown, the steps include:

[0141] Step 1: Control profane variables and eliminate the influence of backdoor paths;

[0142] Before calculating the intervention effect of each factor variable X, it is necessary to filter and control the identified confounding variables, such as cloud cover and season, through the backdoor path corresponding to Do-Calculus to ensure that the association between X and Y is a valid pure causal association. For example, after controlling for the cloud cover variable, the association between X1 (infrared temperature) and Y (fire) is only caused by the direct causal relationship of fire → temperature increase, without any spurious association caused by cloud cover interference; after controlling for the season variable, the association between X3 (optical texture) and Y (fire) is only caused by fire → texture anomaly, without any interference from winter vegetation leaf drop.

[0143] Step 2: Calculate the average intervention effect (ATE) for each variable;

[0144] At this point, the average intervention effect ATE is defined as the probability that the target variable Y changes when the intervention variable X changes from the baseline value to an outlier value, expressed by the formula: ;

[0145] Where do (X=x) means forcibly setting X to x, cutting off the influence of other variables on X, and ensuring that the calculation is of the direct causal effect of X on Y. Based on the characteristics of remote sensing data, the ATE calculations for each X are shown in Table 1 below:

[0146] Table 1

[0147]

[0148] Among them, ATE calculation needs to be based on historical interpretation data of large models in the cloud and on the ground, such as 100,000+ fire / non-fire annotations, which are obtained by statistically analyzing the probability changes of Y before and after intervention.

[0149] Step 3: ATE normalization to obtain the overall weight;

[0150] In this study, the absolute value of ATE for each variable is taken as the causal influence strength, and converted into a proportion weight through a normalization formula, which is expressed as:

[0151] .

[0152] The ATE results calculated using the normalization formula include:

[0153] The total ATE is 0.8 + 0.48 + 0.24 + 0.08 = 1.6;

[0154] X1 weight = (0.8 / 1.6) × 100% = 50%;

[0155] X2 weight = (0.48 / 1.6) × 100% = 30%;

[0156] X3 weight = (0.24 / 1.6) × 100% = 15%;

[0157] X4 weight = (0.08 / 1.6) × 100% = 5%;

[0158] The final overall weights are: X1 (50%) > X2 (30%) > X3 (15%) > X4 (5%).

[0159] It should be noted that the weights obtained above represent the strength of causal influence, that is, the proportion of influence in the embodiments of this application, representing the causal influence. For example, a 50% weight for X1 indicates that the direct causal influence of infrared temperature features on fire determination is the highest, indicating that the core interpretation feature of fire is high temperature, and infrared temperature is the most direct basis for determination; while the weight for X4 (wind speed) is only 5%, because wind speed does not directly determine whether it is a fire, but only indirectly affects temperature diffusion, and the causal relationship is the weakest. At this time, the weights will also change dynamically with the interpretation target. For example, if the interpretation target changes from fire determination to vegetation drought assessment, the variable definition and ATE calculation will be adjusted synchronously, and the weights will also change. At this time, X1 may become NDVI value, that is, drought leads to a decrease in NDVI; X2 may become soil moisture, that is, drought directly leads to low moisture; X3 may become optical vegetation texture; and X4 may become precipitation. If the ATE of X2 (soil moisture) is 0.7, the ATE of X1 (NDVI) is 0.56, the ATE of X4 (precipitation) is 0.28, and the ATE of X3 is 0.14, the normalized weights may become X2 (40%) > X1 (32%) > X4 (16%) > X3 (12%). This achieves the purpose of determining the causal weights by interpreting the target. Do-Calculus ensures that the weights only reflect pure causal effects through intervention operations, avoiding interference from spurious associations.

[0160] In another embodiment of this application, for further definition and explanation, the step of determining the proportion of influence of each causal variable in different causal associations through causal intervention verification includes:

[0161] The backdoor path is used to filter the causal variables of each causal variable, and the causal variables obtained after the filtering are determined to be causally related, so as to determine the proportion of the influence.

[0162] To avoid cross-influence between factors and improve the accuracy of the true impact of individual factors on landslides, the ground-based system filters associated causal variables during the calculation of impact percentages. Specifically, in step 102 of the aforementioned embodiment, a backdoor path can be used to filter associated causal variables, and the resulting associated causal variables are used to determine the causal relationship and thus the impact percentage. In causal interpretation, it is usually necessary to clarify the direct causal relationship between the core interpretable feature of the observed variable and the causal variable of the target variable, such as "fire" → "infrared high-temperature feature" or "SAR low-scattering feature." In this case, the backdoor path is an interference path that, without passing through the target variable (causal variable), can simultaneously affect both the target variable and the observed variable, leading to a false association between them. Its core characteristic is the presence of "uncontrolled confounders." That is, when a causal variable is identified that affects both the target variable and the observed variable, the corresponding path is identified as a backdoor path, and this causal variable is deleted as a confounder to break the interference path of a direct "cause-effect" relationship.

[0163] It should be noted that the backdoor path in this embodiment is essentially a causal interference channel caused by confounding variables. The process of identifying the backdoor path is the process of separating false associations from data associations and extracting true causal relationships, in order to achieve the purpose of connecting the initial screening at the spaceborne end with the in-depth interpretation at the ground end. At this time, it can help the spaceborne end reduce invalid data transmission and provide accurate causal relationships for the cloud-based large model at the ground end, ensuring that the interpretation results are upgraded from inferences based on statistical associations to reliable conclusions based on causal logic, meeting the high-precision and error-free requirements of industry scenarios such as emergency rescue, disaster relief, and agricultural monitoring.

[0164] In a specific scenario where cloud cover is used as a confounding variable in a backdoor path implementation, the target variable (cause) is the actual surface temperature, such as "normal surface" vs. "high-temperature fire zone"; the observed variable (effect) is the temperature value retrieved from infrared imagery; and the confounding variable is cloud cover. In this case, the backdoor path is: Cloud cover → Actual surface temperature, as cloud cover blocks sunlight and lowers the surface temperature; Cloud cover → Infrared temperature retrieval value, as the thermal radiation interference from cloud cover leads to an overestimation of the retrieved temperature. Here, "cloud cover," as a confounding variable, affects both the cause and effect, interfering with the direct causal relationship between the "infrared temperature retrieval value" and the "actual surface temperature." If this backdoor path is not identified, "abnormal retrieval temperature caused by cloud cover" may be misjudged as "actual high surface temperature (fire)," resulting in interpretation errors. Therefore, the confounding variable cloud cover will be removed according to this backdoor path.

[0165] In a scenario where a backdoor path is implemented using a specific seasonal factor as a confounding variable, the target variable (cause) is vegetation growth status, such as "healthy vegetation" vs. "drought-stricken and withered vegetation"; the observed variable (effect) is the NDVI of the optical image, which reflects vegetation cover through the normalized vegetation index; the confounding variable is the season, such as "winter". In this case, the backdoor path is: winter → vegetation growth status. In winter, vegetation naturally sheds its leaves, resulting in a decline in growth status; winter → NDVI value. Low vegetation cover in winter naturally leads to a lower NDVI value. Here, "season" acts as a confounding variable, causing "low NDVI," which creates a spurious association with "drought and withering." If this backdoor path is not identified, "normally low NDVI in winter" may be misjudged as "drought-stricken and withered vegetation," leading to interpretation errors and affecting the interpretation accuracy in scenarios such as agricultural monitoring. In this case, the essence of the backdoor path refers to the spurious association channel mediated by the confounding variable. Identifying the deviation between the observed variable and the target variable from the true causal relationship becomes the core source of "misjudgment" in remote sensing interpretation.

[0166] In this application embodiment, by identifying backdoor paths, interference from useless variables can be eliminated to ensure the causal validity of the interpretation results. That is, the core function of identifying backdoor paths is to control confounding variables, remove spurious associations, and restore the true causal relationship between the target variable and the observed variable, so as to directly serve the deep interpretation target of the cloud-based large model on the ground.

[0167] In one specific embodiment, due to the limitations of the lightweight model on the spaceborne end, which is constrained by fixed rules and low complexity, preliminary interpretation errors may occur when backdoor paths cannot be identified, leading to misclassification of temperature anomalies caused by cloud and fog interference as fires. Simultaneously, ground-based causal inference, by identifying backdoor paths, can quickly filter out false anomaly data affected by confounding variables. For example, after identifying the "cloud and fog → infrared temperature" backdoor path, the ability of SAR imagery to penetrate clouds and fog can be combined. SAR imagery is not affected by clouds and fog and can reflect the true surface structure; therefore, it is possible to verify whether there are SAR scattering anomalies in the infrared temperature anomaly area, such as vegetation structure damage caused by fire. If there are no SAR anomalies, the infrared anomaly is determined to be caused by the backdoor path (cloud and fog), eliminating the need for further in-depth interpretation and reducing unnecessary computational power consumption on the ground. If there are SAR anomalies, it is confirmed to be a real fire, and computational power is prioritized for analysis to correct the errors in the preliminary interpretation on the spaceborne end and reduce unnecessary data transmission.

[0168] In one specific embodiment, backdoor paths can lead to "causal inversion" in complex tasks such as fine-grained land cover classification and dynamic change analysis. For example, the causal relationship of vegetation type → soil moisture might be misjudged as "soil moisture → vegetation type." Therefore, identifying backdoor paths can be achieved by controlling confounding variables to restore the true causal relationship. For instance, in "crop yield assessment," the target variable is yield level, the observed variable is optical NDVI value, and the confounding variable is precipitation, which affects both crop yield and NDVI value. After identifying the backdoor path from precipitation to yield and from precipitation to NDVI, stratified analysis (such as grouping data into "high precipitation," "medium precipitation," and "low precipitation") can be used to analyze the correlation between NDVI and yield under the same precipitation conditions. At this point, the interference of precipitation can be removed, and yield can be predicted more accurately through NDVI. This avoids the false correlation of high NDVI and high yield in high precipitation areas, which mistakenly equates high NDVI directly with high yield, while ignoring the real situation where there may be high NDVI but low yield in low precipitation areas. This improves the accuracy of surface depth interpretation and avoids causal reversal.

[0169] In one specific embodiment, due to the fixed threshold of the lightweight model at the spaceborne end, such as the high temperature threshold... While backdoor paths can easily fail in different regions or seasons, identifying them can help large cloud-based models on the ground dynamically adapt to different scenarios. For example, in detecting high-temperature anomalies in winter in cold regions, if the threshold rules for temperate regions are used, the backdoor path might be identified as "low winter temperature → low actual surface temperature" and "higher infrared inversion temperature due to winter atmospheric back radiation," leading to a misjudgment of slightly higher inversion temperatures as an anomaly. After identifying this backdoor path, the judgment logic can be adjusted autonomously by introducing region / season as a control variable, such as adjusting the high-temperature threshold in cold winter regions to "...". This ensures consistency of interpretation results across different scenarios, avoids a "one-size-fits-all" approach, and thus possesses the generalization ability to support cross-scenario interpretation, adapting to scenarios in different regions or seasons.

[0170] In another embodiment of this application, for further definition and explanation, the step of determining the influence ratio of each causal variable on different causal associations through causal intervention verification, and configuring the influence ratio in the initial causal variable graph to obtain the causal variable graph, the method further includes:

[0171] The influence percentages of each causal variable are ranked, and the interpretation results of the influencing factors are determined based on the ranking results.

[0172] To accurately identify the interpretation results affecting remote sensing targets, after calculating the influence percentage, the influence percentages of all causal variables are ranked to obtain the final influence mapping interpretation result. In a specific implementation scenario, the current step of ranking and determining the interpretation results of influencing factors can be taken as step 4 after step 3 above. The influence percentages of the four factors obtained above are ranked, resulting in rainfall (X1) 50% > terrain slope (X2) 30% > vegetation coverage (X3) 15% > human excavation (X4) 5%. Thus, it is concluded that the main cause of the landslide is "heavy rain (strong rainfall)" plus "steep terrain slope". These two factors together account for 80% of the influence on the landslide, which is the influence mapping interpretation result. This application embodiment does not make specific limitations.

[0173] In another embodiment of this application, for further definition and explanation, the step of determining the causal labels and interpretation results of the label regions based on the causal variable diagram, and obtaining the verification results corresponding to the causal labels and interpretation results includes:

[0174] The label regions corresponding to the multimodal remote sensing data are determined, and the causal labels of the label regions and the interpretation results are generated based on the interpretation results of the influencing factors.

[0175] To improve the protocol interpretation effect between the ground and spaceborne terminals, the ground terminal of the current execution entity can determine the label region corresponding to the multimodal remote sensing data and generate causal labels and interpretation results for the label region according to the obtained influence mapping interpretation results. At this time, the label region is the feature region corresponding to the core interpretation feature. It can be labeled through image change recognition or manually labeled to be output as a visual map to the front-end user. This application embodiment does not make specific limitations. At the same time, since the interpretation results of influencing factors contain different causal variables and their corresponding influence ratios, after statistically analyzing the causal variables and their corresponding influence ratios, causal labels are determined. That is, causal labels are used to mark the causal variables that affect the core interpretation feature. Correspondingly, the interpretation results are used to characterize the influence results of the causal variables and can be output to the front-end user in the form of a natural language report.

[0176] In another embodiment of this application, for further definition and explanation, the step of obtaining the verification result corresponding to the cause label and the interpretation result includes:

[0177] Output the cause label and the interpretation result, and receive the verification result from the user verifying the cause label and the interpretation result.

[0178] To ensure the validity of the causal labels and interpretation results, thereby achieving effective collaborative interpretation between the ground and spaceborne ends, the system outputs causal labels and interpretation results, and receives verification results from users to validate the causal labels and interpretation results, which serve as the basis for knowledge transfer of the aforementioned first interpretation model.

[0179] In some implementation scenarios, the selection of causal variables can also consider both human and natural factors. Specifically, the characteristics for determining natural factors are that the changes in factors conform to natural laws (such as seasonal fluctuations in rainfall and long-term stable slopes), and the judgment is based on evidence that multimodal data shows no non-natural traces (such as the linear characteristics of mechanical excavation in SAR images and the absence of artificial building materials in optical images). The characteristics for determining human factors are abnormal changes occurring in a short period of time (such as the addition of a large area of ​​exposed land within 24 hours); the judgment is based on evidence such as corner reflector characteristics shown in SAR images (building materials) and abnormally enhanced nighttime lighting shown in infrared images (construction activities). This application does not impose specific limitations on these aspects.

[0180] This application provides a method for interpreting remote sensing images based on causal reasoning and space-ground collaboration. Compared with existing technologies, this application obtains initial interpretation features from the initial interpretation of multimodal remote sensing data by the spaceborne end from the ground end, and processes these initial interpretation features based on a first interpretation model that has completed model training to obtain core interpretation features. Causal variables are determined based on these core interpretation features, and a causal variable graph is constructed based on these causal variables. The causal labels and interpretation results for labeled regions are determined according to the causal variable graph, and verification results corresponding to the causal labels and interpretation results are obtained. Based on knowledge distillation and the verification results, the learning core features for updating and training the first interpretation model are determined, and these learning core features are sent to the spaceborne end. This allows the second interpretation model in the spaceborne end to be trained in coordination with the learning core features based on feedback from the ground end, achieving the goal of timely dynamic coordination training between the ground end and the spaceborne end. This optimizes the interpretation accuracy of remote sensing data from the spaceborne end, can meet the inference needs of environmental causes in remote sensing images in a diversified and precise manner, significantly improves the interpretability of the interpretation results and the adaptability of the spaceborne-ground system, and greatly increases the diversity of applicable scenarios for remote sensing monitoring.

[0181] This application provides another method for interpreting satellite-ground collaborative remote sensing images based on causal reasoning, applied to satellite-borne terminals, such as... Figure 2 As shown, the method includes:

[0182] 201. Obtain multimodal remote sensing data, and perform feature fusion on the multimodal features of the multimodal remote sensing data to obtain a multimodal feature vector.

[0183] In this embodiment, the spaceborne terminal refers to an artificial celestial device in a satellite system capable of capturing images of the Earth's surface in space, or a spacecraft with imaging capabilities within aerospace equipment. In this case, the spaceborne terminal system mainly includes a spaceborne base station, a telemetry and control data transmission terminal, and a multi-functional electronic terminal, among other devices. Its functions cover key tasks such as communication, data processing, and attitude control, enabling real-time acquisition of multimodal remote sensing data of the Earth. This multimodal remote sensing data includes optical image data, infrared image data, and SAR image data. The spaceborne terminal can acquire data through optical sensors, infrared sensors, and radar sensors; however, this embodiment does not impose specific limitations on these methods.

[0184] It should be noted that, in order to realize the lightweight interpretation model for interpreting multimodal remote sensing data, the current execution end extracts multimodal features from the multimodal remote sensing data, so as to perform feature fusion based on the extracted texture features, backscattering coefficients and surface temperature features to obtain multimodal feature vectors.

[0185] In some embodiments, for the multimodal features to be fused, since the multimodal remote sensing data includes optical image data, infrared image data, and SAR image data, the spaceborne end first performs feature extraction after acquiring the multimodal remote sensing data to obtain multimodal features, such as texture features, backscattering coefficients, and surface temperature features. For example, texture features of three RGB channels are extracted from optical image data, expressed as gradient values, with a dimension of 64; backscattering coefficients are extracted from SAR image data, expressed as dB values, with a dimension of 32, and mapped to a 64-dimensional space by projecting a 32×64 matrix W, consistent with the dimension of optical features; surface temperature features, such as thermal infrared band inversion values, expressed as °C, with a dimension of 32, are extracted from infrared image data, and mapped to a 64-dimensional space by projecting a 32×64 matrix W', which can supplement temperature dimension information, such as high temperature anomalies in fire monitoring, temperature differences between water and land, etc., and this application embodiment does not make specific limitations.

[0186] 202. Based on the first interpretation model, the multimodal feature vector is interpreted to generate interpretation features, which are then sent to the ground terminal.

[0187] In this embodiment of the application, the first interpretation model of the spaceborne terminal is a lightweight deep learning model. At this time, the first interpretation model is obtained by coordinated training based on the learning core features fed back from the ground terminal. The learning core features are determined by the verification results of the causal variable graph after the second interpretation model in the ground terminal interprets the data. The causal variable graph is constructed based on the core interpretation features obtained after interpretation. The causal variable graph contains the influence ratio of different causal variables on different causal associations. The causal variables are used to characterize the variables that affect the monitoring results in multimodal remote sensing data.

[0188] It should be noted that, as Figure 5 As shown, the satellite-based terminal can interpret the acquired multimodal remote sensing data using the first interpretation model before co-training to obtain initial interpretation features, which are then sent to the ground terminal. After the ground terminal executes the aforementioned steps 101-104, it feeds back the learned core features to the satellite-based terminal of the current executing entity for retraining. After training is completed, the first interpretation model can be reused to interpret the newly acquired multimodal remote sensing data, thus achieving the purpose of co-training the interpretation models between the ground and satellite terminals. This application embodiment does not impose specific limitations.

[0189] In some embodiments, the first interpretation model on the satellite-borne end, due to its core objective of adapting to the limited computing and storage resources of the satellite, needs to complete the preliminary processing of a single 10km×10km multimodal image within 5 seconds or less. Therefore, the core objective is data filtering and preliminary analysis. At this time, only basic tasks need to be completed, such as scene determination (normal / high temperature anomaly / cloud and fog coverage), simple object classification (vegetation / water body / bare land), and coarse location of anomaly areas (such as the approximate range of high temperature points), etc., to output initial interpretation features. At the same time, more efficient data transmission can be completed at low precision. For example, the ground object classification accuracy only needs to be 70%-80%, which can quickly output key data on whether there are anomalies or which needs to be transmitted first, reducing the amount of invalid data transmission for subsequent ground processing. At this time, only anomaly areas or key features are transmitted, rather than the full amount of raw data. In addition, the satellite-based computing power (usually embedded chips, such as FPGAs and low-power CPUs, with computing power less than 1% of that of ground servers), storage (only able to store a small amount of model parameters and temporary data), and power consumption (requiring strict control of energy consumption to extend satellite lifespan) are limited. Therefore, the characteristics of the first interpretation model must meet the "three lows" requirement:

[0190] 1. Low parameter count, typically in the millions (1M-10M). For example, the first interpretation model used above can include MobileNetV2 with approximately 3.5M parameters or ShuffleNetV2 with approximately 1.4M parameters, in order to avoid consuming too much storage.

[0191] 2. Low computational complexity: It adopts lightweight structures such as depthwise separable convolution, grouped convolution, and sliding window attention. The computational cost of a single image is only 1 / 100 to 1 / 10 of that of the cloud model used on the ground, ensuring fast operation.

[0192] 3. Low dependency: No external data support is required. It only relies on the multi-modal feature vectors obtained in real time by the satellite, such as 64-dimensional optical mode vector, 64-dimensional SAR mode vector, and 64-dimensional infrared mode vector. There is no need to call historical data or external databases.

[0193] In this embodiment, the first interpreter model for the lightweight model can include, but is not limited to, lightweight convolutional neural network (CNN) models such as MobileNet and ShuffleNet, and can also include lightweight models with attention mechanisms such as Swin Transformer Tiny and TinyViT. This embodiment does not impose specific limitations. To ensure fast operation, the first interpreter model adopts a combination of simplified structure and fixed decision rules to achieve structural simplification. For example, the MobileNetV2 model retains only the inverted residual structure plus a 1×1 convolution, removing the complex feature pyramid network; Swin Transformer Tiny uses only 2 layers of sliding window attention instead of 12 layers of full attention. Furthermore, rule solidification can be implemented, meaning that scene determination depends on a fixed threshold, such as C≥0.7 for a normal scene. For high temperature anomalies, the weight allocation is based on preset values, such as 0.4 for optical, 0.4 for SAR, and 0.2 for infrared in normal scenarios. These values ​​can be dynamically adjusted based on new data.

[0194] like Figure 5 The illustrated collaborative interpretation process between the satellite-borne and ground-based terminals illustrates a scenario where, in a satellite-borne interpretation interaction, the interpretation model on the satellite-borne terminal is a lightweight model, supporting real-time onboard response. Applications focusing on local processing on the satellite-borne terminal include, for example, fire monitoring, which can quickly locate areas of abnormal high temperatures by transmitting only the characteristic data of these areas (rather than full-frame images), reducing the amount of satellite downlink data (from GB to KB). In cloud and fog obstruction monitoring, cloud and fog scenes can be determined in real time, automatically increasing SAR weights to ensure that the initial results output by the satellite-borne terminal are not affected by optical obstruction. The interpretation model on the ground-based terminal, serving as a large model in the cloud, supports ground-based industry applications, focusing on in-depth decision-making on the ground. For instance, in emergency rescue scenarios, combining the fire anomaly area characteristics transmitted by the satellite-borne terminal with historical terrain and wind direction data to predict fire spread paths, the interpretation results provide precise deployment plans for rescue teams. In agricultural monitoring scenarios, the ground-based terminal will select areas of abnormal crop growth with core interpretation features from the satellite-based terminal, and perform in-depth interpretation by combining historical yield data and soil moisture data to analyze the causes of the abnormalities, such as drought or pests and diseases, and thus output the final interpretation results for precise irrigation or fertilization.

[0195] In another embodiment of this application, for further definition and explanation, before the step of feature fusion of the multimodal features of the multimodal remote sensing data to obtain a multimodal feature vector, the method further includes:

[0196] Texture spatial variation information is extracted from optical image data that has undergone radiometric and geometric corrections to obtain a gradient intensity map. The gradient mean, gradient variance, and gradient entropy of the gradient intensity map are then combined to obtain texture features.

[0197] SAR image data that has been radiometrically calibrated and noise-removed is filtered and dimensionality reduced to obtain a backscatter map. The backscatter map is then linearly projected to obtain the backscatter coefficients.

[0198] Temperature inversion is performed on infrared image data calibrated by thermal infrared band to obtain a surface temperature map, and the surface temperature map is linearly mapped to obtain surface temperature characteristics.

[0199] In order to achieve the goal of fusion features with uniform dimensions and effective features, the current execution subject extracts features from the multimodal remote sensing data to obtain multimodal features including texture features, backscattering coefficients, and surface temperature features.

[0200] For feature extraction from optical image data, the spatial variation information of texture is first extracted from the optical image data after radiometric and geometric correction to obtain a gradient intensity map. At this point, radiometric correction is performed on the RGB three-channel optical image data to eliminate atmospheric scattering and sensor errors. Simultaneously, geometric correction is used to match geographic coordinates, ensuring that texture features correspond to the actual surface location. Furthermore, in the process of generating the gradient intensity map, multi-scale gradient operators can be used to extract the spatial variation information of the texture. The specific steps include: first, selecting either the Sobel operator or the Scharr operator, and calculating the horizontal gradient (Gx) and vertical gradient (Gy) for each pixel in the RGB three-channel optical image data; second, L2 norm normalizing is performed on the gradient values ​​of each channel to obtain a single-channel gradient intensity map, which is then combined to form a three-channel gradient intensity map. After obtaining the gradient intensity map, the current satellite-based end merges the gradient mean, gradient variance, and gradient entropy of the gradient intensity map to obtain texture features. This involves performing block statistics on the three-channel gradient intensity map, including dividing the 10km×10km image data into 64 equally sized sub-regions (e.g., each sub-region is 1.25km×1.25km), calculating the gradient mean, gradient variance, and gradient entropy for each sub-region, and the statistics for the three categories. Finally, the statistical results of the three channels are merged. The result is: three channels × 64 sub-regions × one category of core statistics (if multiple statistics are not explicitly specified, the gradient mean can be used by default) = a 64-dimensional feature vector. The final output dimension is consistent with the backscattering coefficient and infrared features.

[0201] For feature extraction from SAR image data, the SAR image data is first transformed into its physical meaning. That is, the SAR image data acquired by the spaceborne terminal is the original gray value, i.e., DN value, which is converted into the backscattering coefficient in physical meaning, denoted as σ.0 (Unit: dB) The formula is expressed as:

[0202] ;

[0203] Where K is the sensor calibration constant, provided by the satellite manufacturer; for example, the K value for Sentinel-1 can be -15.6 dB; PRF is the pulse repetition frequency; and R is the slant range of the radar reaching the Earth's surface. The incident angle can be read from the SAR image metadata, and this application embodiment does not impose specific limitations. Meanwhile, to eliminate noise, Lee filtering or Gamma filtering can be used during radiometric calibration and noise removal to eliminate speckle noise in the SAR image data, obtaining a backscatter map to avoid noise interference with scattering coefficient statistics. Simultaneously, the denoised backscattering coefficient map is divided into 32 equally sized sub-regions, such as each sub-region being 3.125km × 3.125km, and the mean backscattering coefficient of each sub-region is calculated to form a 32-dimensional dB feature vector. Furthermore, to ensure the consistency of feature dimensions, the spaceborne end can map the 32-dimensional features to 64 dimensions using a projection matrix W (32 × 64 dimensions), specifically employing a linear projection algorithm: first, the original feature vector in the SAR image data is set as... The projected feature vector is Then Y = W×X, where the projection matrix W is pre-learned through feature alignment of the training set to ensure that the backscattering coefficients after mapping are consistent with the numerical range and distribution trend of the optical texture features, which facilitates subsequent cross-modal fusion.

[0204] For feature extraction from infrared image data, the spaceborne end performs temperature inversion on the radiatively converted infrared image data to obtain a surface temperature map. Specifically, the grayscale values ​​(DN) of the infrared image data are first converted into the apex radiance. The unit is W / (m 2 sr μm), the formula is:

[0205] ;

[0206] in, , The maximum or minimum radiance of that infrared band can be read from image metadata, such as the TIRS band of Landsat-8. Approximately 15.3 W / (m 2 sr The embodiments in this application are not specifically limited to μm. Furthermore, infrared image data can be calibrated using thermal infrared band radiometry, and temperature inversion can be performed on the calibrated infrared image data, i.e., using a single-window algorithm suitable for spaceborne infrared data, such as Landsat-8, expressed by the formula:

[0207] ;

[0208] in, Surface temperature (°C). Surface emissivity can be obtained from land cover type retrieved from optical imagery, and water bodies... ≈0.98, vegetation ≈0.96, bare land ≈0.92; The average atmospheric temperature can be obtained from real-time meteorological data or atmospheric profile products. Atmospheric transmittance can be calculated based on the observation angle and altitude, such as in low-altitude areas. (≈0.9); a, b, and c are algorithm coefficients determined by band characteristics. For the TIRS10 band of Landsat-8, a≈-1.101, b≈1.522, and c≈0.004, respectively. This application does not impose specific limitations on the embodiments. In addition, to ensure the consistency of feature dimensions, the satellite-borne end performs linear mapping on the surface temperature map to obtain surface temperature features. That is, the inverted surface temperature map is first divided into 32 equally sized sub-regions, and the temperature mean of each sub-region is calculated to form a 32-dimensional feature vector of temperature value ℃. Then, linear mapping is performed through the projection matrix W' (32×64 dimensions). At this time, the mapping logic is the same as that of the backscattering coefficient, expanding the 32-dimensional feature to 64 dimensions to ensure the consistency with the optical and backscattering coefficient dimensions and support subsequent cross-modal attention fusion.

[0209] In another embodiment of this application, for further definition and explanation, the step of feature fusion of the multimodal features of the multimodal remote sensing data to obtain a multimodal feature vector includes:

[0210] Multimodal similarity matrices are generated for the texture features, backscattering coefficients, and surface temperature features, respectively. Feature fusion is then performed based on the feature associations and modal weights of the multimodal similarity matrices to obtain multimodal feature vectors.

[0211] To achieve the goal of interpreting large models using multimodal features, the current spaceborne end performs feature fusion on the extracted multimodal features. Specifically, the cosine similarity matrix quantization method is first used to calculate the multimodal similarity matrix M1 between texture features and backscattering coefficients, the multimodal similarity matrix M2 between surface temperature features and texture features, and the multimodal similarity matrix M3 between surface temperature features and backscattering coefficients. Among them, M1, M2, and M3 are 64×64 dimensions. M1[i][j]=cos(texture feature i, backscattering coefficient j) is used to establish the correspondence between surface texture details and radar reflection intensity, such as matching vegetation areas distinguished by texture with SAR low scattering features. M2[i][j]=cos(surface temperature feature i, texture feature j) is used to strengthen the association between temperature anomaly areas and surface cover types, such as matching high temperature points with bare land. M3[i][j]=cos(surface temperature feature i, backscattering coefficient j) is used for temperature feature calibration under complex weather conditions, such as verifying whether infrared high temperature points are real fire points in clouds and fog through SAR structural features.

[0212] In another embodiment of this application, for further definition and explanation, the step of performing feature fusion based on the feature association of the multimodal similarity matrix and the modality weights to obtain a multimodal feature vector includes:

[0213] Feature associations are determined based on the multimodal similarity matrix.

[0214] Assign modal weights according to different scenarios;

[0215] The feature associations and modal weights are calculated using a weighted summation method to obtain a multimodal feature vector.

[0216] Since the modal weights are dynamically configured based on different scenarios, and feature association is used to characterize the association strength between different modal features during fusion, they can be determined based on the multimodal similarity matrix. Meanwhile, the scenarios include normal scenarios, high temperature abnormal scenarios, and cloud and fog coverage scenarios. The corresponding modal weights can be assigned according to different scenarios. The feature associations and modal weights are calculated based on the weighted summation method to obtain the fused multimodal feature vector. The modal weights include optical weights, SAR weights, and infrared weights.

[0217] In some specific embodiments, the current spaceborne terminal fuses the aforementioned multimodal similarity matrices M1, M2, and M3 based on feature associations and modal weights of the multimodal similarity matrices. Specific steps include:

[0218] Step 1: Establish feature associations based on multimodal similarity matrices. Three multimodal similarity matrices (M1, M2, M3) are used to quantify the association strength of different modal features. Specifically, the cosine similarity values ​​(range [-1, 1]) of the elements in the matrices determine the semantically related feature dimensions in texture features, backscattering coefficients, and surface temperature features. For example, the texture dimension and radar scattering dimension corresponding to high similarity values ​​in M1 represent the feature mapping of the same surface target, providing a basis for subsequent feature fusion rather than directly participating in numerical addition.

[0219] The second step is to dynamically allocate modal weights according to the scene to complete feature fusion. The core of this fusion is to weight and fuse the original feature vectors of three feature modes: texture features, backscattering coefficient, and surface temperature features, rather than weighting the multimodal similarity matrix. The modal weight allocation methods for different scenes include: normal scenes, high-temperature anomaly scenes, and cloud / fog-covered scenes. Specifically, in normal scenes (without high-temperature anomalies), the optical weight is configured as 0.4, the SAR weight as 0.4, and the infrared weight as 0.2; in high-temperature anomaly scenes (such as fires or geothermal areas), the optical weight is configured as 0.3, the SAR weight as 0.3, and the infrared weight as 0.4; and in cloud / fog-covered scenes, the infrared weight is configured as 0.2, the SAR weight as 0.5, and the infrared weight as 0.3. Finally, the formula for calculating the fused feature vector is: Fusion Feature = Texture Feature × Optical Weight + Backscattering Coefficient × SAR Weight + Surface Temperature Feature × Infrared Weight.

[0220] In another embodiment of this application, for further definition and explanation, before the step of allocating modal weights according to different scenarios, the method further includes:

[0221] The feature association consistency index is calculated based on the multimodal similarity matrix, and the mean temperature and standard deviation of the temperature in the whole region are calculated based on the surface temperature feature vector.

[0222] When the feature association consistency index, the average temperature of the entire region, and the standard deviation of the temperature of the entire region match the preset normal scenario conditions, the scenario is determined to be a normal scenario.

[0223] To achieve accurate configuration of modal weights and thus improve the depth interpretation of remote sensing targets in different scenarios, the current satellite-based terminal, when determining the scenario to be a conventional scenario, first calculates the feature association consistency index based on the multimodal similarity matrix. Specifically, the feature association consistency index can be calculated based on the three generated multimodal similarity matrices (M1, M2, M3). First, the proportion of high-similarity elements in multimodal similarity matrix M1 is counted, and the number of elements with a similarity value ≥ 0.6 is divided by the total number of elements, and labeled as R1. The proportion of high-similarity elements in multimodal similarity matrix M2 is counted and labeled as R2. Finally, the consistency index C = (R1 + R2) / 2 is calculated. If C ≥ 0.7, it indicates that the descriptions of the Earth's surface by each feature are highly consistent, with no obvious abnormal interference. Furthermore, the mean temperature and standard deviation of the entire region are calculated based on the surface temperature feature vector. This involves extracting the surface temperature feature vector retrieved from the infrared image data and calculating the mean temperature (T_avg) and standard deviation (T_std) for the entire region. T_avg falls within the historical normal temperature range for the same period in the region (e.g., T_avg 25-35℃ in temperate regions during summer), and T_std ≤ 5℃, indicating no localized extreme high temperatures. Finally, if the feature correlation consistency index and the mean and standard deviation of the entire region's temperature are set to pre-defined conventional scenario conditions, then the scenario is considered conventional. For example, if the feature correlation consistency index C ≥ 0.7 and there are no extreme temperature anomalies, this is considered a pre-defined conventional scenario condition.

[0224] In another embodiment of this application, for further definition and explanation, before the step of allocating modal weights according to different scenarios, the method further includes:

[0225] Based on the surface temperature characteristics, calculate the temperature value and standard deviation of the sub-region, and mark the candidate abnormal region when the temperature value and standard deviation of the sub-region match the preset temperature conditions.

[0226] Once the texture features and the backscattering coefficient are verified, the candidate anomaly region is determined to be a high-temperature anomaly scene.

[0227] To achieve accurate configuration of modal weights and thus improve the depth interpretation of remote sensing targets in different scenarios, the current satellite-based system, when determining a high-temperature anomaly scenario, first locates the high-temperature region based on a statistical outlier detection algorithm. The specific steps include: Step 1: Calculating the mean temperature T_avg of all sub-regions in the infrared temperature feature vector based on surface temperature characteristics, and three times the sub-region standard deviation (3×T_std); Step 2: Setting a high-temperature threshold. If a sub-region temperature exists (Preset temperature conditions) marked as candidate anomalous areas. For example, if a region has T_avg=30℃ and T_std=4℃, then T_t h=42℃. Sub-regions with temperatures ≥42℃ are included as candidates. Furthermore, to avoid misclassifying sensor noise and cloud top high temperatures as surface high temperatures, cross-modal verification is required. This includes optical feature verification, such as verifying that if the optical texture features of the candidate high-temperature area show irregular gray-scale abrupt changes, such as the smoke texture of a fire or the columnar texture of an industrial chimney, rather than a uniform cloud texture, then an anomaly is confirmed. Another example is backscattering coefficient verification, where the SAR backscattering coefficient shows high scattering (buildings / bare ground) or low scattering (water bodies excluded), rather than moderate scattering (cloud areas), further eliminating cloud interference and completing the verification. Finally, when there are... When a candidate region is selected and its texture features and backscattering coefficients pass verification, it is determined to be a high-temperature abnormal scene.

[0228] In another embodiment of this application, for further definition and explanation, before the step of allocating modal weights according to different scenarios, the method further includes:

[0229] The suspected cloud and fog coverage areas were determined based on the average gray value and contrast of the surface temperature characteristics.

[0230] Once it is determined that there is cloud and fog interference in the suspected cloud and fog coverage area based on the gradient entropy of the texture features, the consistency index of the backscattering coefficient and the land cover of the texture features is calculated.

[0231] When the land cover consistency index matches the preset cloud and fog scene conditions, the scene is determined to be a cloud and fog coverage scene.

[0232] To achieve accurate configuration of modal weights and thus improve the depth interpretation of remote sensing targets in different scenarios, the current satellite-based system, when determining a scene to be covered by clouds and fog, first identifies suspected cloud and fog coverage areas based on the average grayscale value and contrast of surface temperature features. Specifically, a cloud and fog recognition algorithm based on surface temperature and texture features is first used to determine the quality of optical image data. Surface temperature feature analysis includes calculating the average grayscale value (G_avg) and contrast (Contrast) of the RGB three channels in the texture features. Due to scattering, cloud and fog areas typically have higher G_avg (high brightness) and lower Contrast (blurred details). If G_avg ≥ 200 (8-bit grayscale value, range 0-255) and Contrast ≤ 30, it is marked as highly suspected of being covered by clouds and fog. At the same time, texture feature analysis includes the gradient entropy of texture features to reflect texture complexity. If it is significantly lower than the historical normal level of the area, such as gradient entropy = 1.8 for normal vegetation areas and 0.6 for cloud and fog areas, it indicates loss of surface texture, further confirming cloud and fog interference. Furthermore, when calculating the consistency index of land cover between backscattering coefficient and texture features, the physical characteristics of SAR radar waves penetrating clouds and fog are used to compare the consistency of land cover between texture features and backscattering coefficient. At the same time, the proportion of high similarity elements R1 of the multimodal similarity matrix M1 is calculated. If R1 ≤ 0.4 as a preset cloud and fog scene condition (far lower than 0.7 in the normal scene), it indicates that the land surface described by the optical image is very different from the land surface described by SAR after penetrating clouds and fog, confirming cloud and fog occlusion.

[0233] In an interpretation scenario involving monitoring landslides triggered by torrential rain in mountainous areas using both satellite-based and ground-based systems, such as... Figure 6As shown, the satellite-borne data includes: Sentinel-1SAR: C-band, VV / VH polarization, 10m resolution, imaging time 10:00 on August 5, 2024; Landsat-8 infrared: thermal infrared (TIR) ​​band, 30m resolution, temperature inversion error ±1℃; geographic knowledge base: DEM data: 30m resolution, slope calculation error ±1°; historical data: landslide records in the area from 2010 to 2023 (a total of 12, of which 8 were related to rainfall >100mm); real-time weather: cumulative rainfall of 120mm from August 4-5, 2024 (weather station data, time resolution 1 hour). Preliminary interpretation output from the spaceborne terminal: Preliminary interpretation features of the suspected landslide area, with boundary coordinates of 118.5°-118.6°E and 30.2°-30.3°N. Feature data includes SAR backscattering abrupt change values ​​(-15dB→-5dB) and infrared temperature anomalies (2°C higher than the surrounding area). The coarse-screened regional heat map generated by the spaceborne terminal is sent to the ground terminal. The darker the red, the higher the probability of landslide. After constructing a causal variable map on the ground terminal, the variable weight table includes rainfall (55%), slope (35%), and vegetation coverage (10%). Nodes (circles) represent variables, and edges (arrows) represent causal directions. The thickness of the edges corresponds to the weights; for example, the edges for landslides are thickest, indicating rainfall. Users can use a web annotation tool to select missed areas, such as 118.55°E, 30.25°N, and note that the area shows signs of recent road construction and excavation. The cloud-based large model adds a human excavation variable (X4), extracts linear excavation trace features (length > 50m, width > 5m) from SAR imagery, and retrains the large model. The influence weight of X4, the core feature to be learned, is adjusted to 15% (originally 5%). Through knowledge distillation, the feature weights associated with X4 are synchronized to the spaceborne model. After the update, the spaceborne model's recognition rate of landslides with excavation traces increases from 60% to 83%.

[0234] In an interpretation scenario of forest fire monitoring between a satellite-based terminal and a ground-based terminal, such as Figure 7As shown, the satellite data includes: BAOK: Gaofen-6 optical image, resolution 2m, red / green / blue / NIR bands, used to identify fire point smoke textures; Sentinel-1 SAR: C-band, VV polarization, resolution 10m, used to penetrate smoke to identify the range of heat sources; Landsat-8 infrared: thermal infrared band, resolution 30m, temperature inversion range -30℃~150℃, error ±1℃; Geographic knowledge base: historical fire records (2015-2023, a total of 32 fires, of which 28 were related to wind speed >3m / s); real-time weather data is as follows: September 10, 2024, wind speed 5m / s (southwest direction), relative humidity 40%. The initial satellite-based interpretation outputs preliminary features of a suspected fire area, located at 119.2°–119.3°E and 29.8°–29.9°N. Feature data includes infrared temperature anomalies (85–120°C), SAR heat source range (0.8 km²), and optical smoke area (1.2 km²). A fused heat map is generated and sent to the ground station, with red indicating high-temperature areas and yellow indicating smoke-covered areas. The ground station's variable weighting table includes peak infrared temperature (45%), wind speed (30%), and vegetation cover (25%). The causal labels and interpretation results indicate that when temperature > 100°C and wind speed > 5 m / s, the fire spreads southwestward at a 20% faster rate, primarily due to high wind speeds and exacerbated by low vegetation humidity (< 30%). After users verified the missed forest fires (infrared temperature 50-70℃, no obvious smoke), the cloud model added forest fire feature templates (infrared temperature 50-70℃ + SAR low scattering value). That is, after retraining, the cloud-based large model's causal inference variables were supplemented with vegetation canopy closure (weight 10%). Finally, the fire feature weights were synchronized to the spaceborne model through knowledge distillation. After the update, the fire recognition rate increased from 65% to 81%.

[0235] This application provides another method for interpreting remote sensing images based on causal reasoning using a satellite-ground collaborative system. The method involves acquiring initial interpretation features from multimodal remote sensing data at the ground-based end, processing these initial features using a first interpretation model that has completed training, and obtaining core interpretation features. Causal variables are then determined based on these core interpretation features, and a causal variable graph is constructed. The causal labels and interpretation results for labeled regions are determined according to the causal variable graph, and verification results corresponding to these labels and results are obtained. Based on knowledge distillation and the verification results, the learning core features for updating and training the first interpretation model are determined, and these learning core features are sent to the satellite-based end. This allows the second interpretation model at the satellite-based end to be trained in coordination with the learning core features based on feedback from the ground-based end. This achieves timely dynamic coordination training between the ground-based and satellite-based ends, optimizes the interpretation accuracy of remote sensing data at the satellite-based end, and can more diversely and accurately meet the inference needs of environmental causes in remote sensing images. It significantly improves the interpretability of the interpretation results and the adaptability of the satellite-ground system, greatly increasing the diversity of applicable scenarios for remote sensing monitoring.

[0236] Furthermore, as a response to the above Figure 1 The implementation of the method shown in this application provides an interpretation device for satellite-ground collaborative remote sensing images based on causal reasoning, such as... Figure 7 As shown, the device includes:

[0237] The acquisition module 31 is used to acquire the initial interpretation features obtained by the satellite-borne terminal in the initial interpretation of multimodal remote sensing data, and to perform interpretation processing on the initial interpretation features based on the first interpretation model that has completed model training to obtain the core interpretation features;

[0238] The construction module 32 is used to determine causal variables based on the core interpretation features and construct a causal variable graph based on the causal variables. The causal variables are used to characterize the variables that affect the monitoring results in multimodal remote sensing data. The causal variable graph contains the proportion of the influence of different causal variables on different causal associations.

[0239] The determination module 33 is used to determine the causal labels and interpretation results of the label region based on the causal variable diagram, and to obtain the verification results corresponding to the causal labels and interpretation results;

[0240] The sending module 34 is used to determine the learning core features of the first interpretation model based on knowledge distillation and the core interpretation features when the verification result is successful, and send the learning core features to the satellite terminal so that the second interpretation model in the satellite terminal can perform coordinated training and interpretation.

[0241] Furthermore,

[0242] The construction module is specifically used to select causal variables according to a preset variable extraction mapping relationship, which includes the matching relationship between different core interpretation features and at least one of different physical variables, different quantitative variables, and different correlation variables; determine causal associations based on historical statistical data and geographical knowledge information, and construct an initial causal variable graph based on the causal variables and the causal associations; determine the influence ratio of each causal variable on different causal associations through causal intervention verification, and configure the influence ratio in the initial causal variable graph to obtain the causal variable graph.

[0243] Furthermore, the device also includes:

[0244] The filtering module is used to filter the causal variables through a backdoor path, and determine the causal relationships based on the filtered causal variables to determine the proportion of influence.

[0245] Furthermore, the determining module is also used to sort the influence proportions corresponding to each of the causal variables, and determine the interpretation result of the influencing factors based on the sorting result.

[0246] Furthermore, the multimodal remote sensing data includes optical image data, infrared image data, and SAR image data.

[0247] The determining module is further configured to determine the label region corresponding to the multimodal remote sensing data, and generate the causal label of the label region and the interpretation result based on the interpretation result of the influencing factors. The label region is a feature region corresponding to the core interpretation feature. The causal label is used to mark the causal variables that affect the core interpretation feature. The interpretation result is used to characterize the influence result of the causal variables.

[0248] Furthermore, the device also includes:

[0249] The output module is used to output the cause label and the interpretation result, and to receive the verification result of the user verifying the cause label and the interpretation result.

[0250] This application provides a satellite-ground collaborative remote sensing image interpretation device based on causal reasoning. Compared with the prior art, this application acquires initial interpretation features obtained by the satellite-borne end from the initial interpretation of multimodal remote sensing data at the ground end, and processes the initial interpretation features based on the first interpretation model that has completed model training to obtain core interpretation features; causal variables are determined based on the core interpretation features, and a causal variable graph is constructed based on the causal variables; the causal labels of the labeled regions and the interpretation results are determined according to the causal variable graph, and the verification results corresponding to the causal labels and interpretation results are obtained; the learning core features for updating and training the first interpretation model are determined based on knowledge distillation and the verification results, and the learning core features are sent to the satellite-borne end so that the second interpretation model in the satellite-borne end is obtained by coordinated training based on the learning core features fed back from the ground end. This achieves the purpose of timely dynamic coordinated training between the ground end and the satellite-borne end, optimizes the interpretation accuracy of remote sensing data at the satellite-borne end, can meet the inference needs of environmental causes in remote sensing images in a diversified and accurate manner, significantly improves the interpretability of the interpretation results and the adaptability of the satellite-ground system, and greatly increases the diversity of applicable scenarios for remote sensing monitoring.

[0251] Furthermore, as a response to the above Figure 4 To implement the method shown, this application provides another interpretation device for satellite-ground collaborative remote sensing images based on causal reasoning, such as... Figure 8 As shown, the device includes:

[0252] The acquisition module 41 is used to acquire multimodal remote sensing data and perform feature fusion on the multimodal features of the multimodal remote sensing data to obtain a multimodal feature vector.

[0253] The sending module 42 is used to interpret the multimodal feature vector based on the first interpretation model, generate interpretation features, and send them to the ground terminal. The first interpretation model is a lightweight deep learning model.

[0254] The first interpretation model is obtained through coordinated training based on the core features learned from ground-based feedback. The core features are determined by the second interpretation model in the ground-based system after interpretation, and verified by a causal variable graph. The causal variable graph is constructed based on the core interpretation features obtained after interpretation. The causal variable graph contains the influence ratio of different causal variables on different causal associations. The causal variables are used to characterize the variables that affect the monitoring results in multimodal remote sensing data.

[0255] Furthermore, the multimodal remote sensing data includes optical image data, infrared image data, and SAR image data, and the device further includes:

[0256] The extraction module is used to extract texture space variation information from optical image data that has undergone radiometric and geometric corrections, obtain a gradient intensity map, and merge the gradient mean, gradient variance, and gradient entropy of the gradient intensity map to obtain texture features.

[0257] The projection module is used to filter and reduce the dimensionality of SAR image data after radiometric calibration and noise removal to obtain a backscatter map, and to perform linear projection on the backscatter map to obtain the backscatter coefficients.

[0258] The mapping module is used to perform temperature inversion on infrared image data calibrated by thermal infrared band radiometry to obtain a surface temperature map, and to perform linear mapping on the surface temperature map to obtain surface temperature characteristics.

[0259] The texture features, the backscattering coefficient, and the surface temperature features have the same feature dimension.

[0260] Furthermore, the acquisition module is specifically used to generate multimodal similarity matrices corresponding to the texture features, the backscattering coefficients, and the surface temperature features, respectively, and to perform feature fusion based on the feature association and modal weights of the multimodal similarity matrices to obtain multimodal feature vectors, wherein the modal weights are dynamically configured based on different scenarios.

[0261] Furthermore, the acquisition module is specifically used to determine feature associations based on the multimodal similarity matrix, wherein the feature associations are used to characterize the association strength between different modal features; to assign modal weights according to different scenarios, wherein the scenarios include normal scenarios, high temperature abnormal scenarios, and cloud and fog coverage scenarios; and to calculate the feature associations and the modal weights based on a weighted summation method to obtain a multimodal feature vector.

[0262] Furthermore, the device also includes:

[0263] The calculation module is used to calculate the feature association consistency index based on the multimodal similarity matrix, and to calculate the mean temperature and standard deviation of the temperature in the whole region based on the surface temperature feature vector, wherein the surface temperature feature vector is extracted by inversion of the infrared image data;

[0264] The determination module is used to determine the scenario as a normal scenario when the feature association consistency index, the average temperature of the whole region, and the standard deviation of the temperature of the whole region match the preset normal scenario conditions.

[0265] Furthermore,

[0266] The calculation module is also used to calculate the temperature value of the sub-region and the standard deviation of the sub-region based on the surface temperature characteristics, and to mark the candidate abnormal region when the temperature value of the sub-region and the standard deviation of the sub-region match the preset temperature conditions.

[0267] The determining module is further configured to determine the candidate abnormal region as a high-temperature abnormal scene after the texture features and the backscattering coefficient have been verified.

[0268] Furthermore,

[0269] The determining module is also used to determine suspected cloud and fog coverage areas based on the average gray value and contrast of the surface temperature characteristics.

[0270] The calculation module is also used to calculate the consistency index between the backscattering coefficient and the land cover of the texture feature after determining that there is cloud and fog interference in the suspected cloud and fog coverage area based on the gradient entropy of the texture feature.

[0271] The determining module is further configured to determine the scene as a cloud and fog coverage scene when the land cover consistency index matches the preset cloud and fog scene conditions.

[0272] This application provides another interpretation device for satellite-ground collaborative remote sensing images based on causal reasoning. Compared with the prior art, this application obtains initial interpretation features from the initial interpretation of multimodal remote sensing data by the satellite-borne end at the ground end, and processes the initial interpretation features based on the first interpretation model that has completed model training to obtain core interpretation features; causal variables are determined based on the core interpretation features, and a causal variable graph is constructed based on the causal variables; the causal labels of the labeled regions and the interpretation results are determined according to the causal variable graph, and the verification results corresponding to the causal labels and interpretation results are obtained; the learning core features for updating and training the first interpretation model are determined based on knowledge distillation and the verification results, and the learning core features are sent to the satellite-borne end so that the second interpretation model in the satellite-borne end is obtained by coordinated training based on the learning core features fed back from the ground end. This achieves the purpose of timely dynamic coordinated training between the ground end and the satellite-borne end, optimizes the interpretation accuracy of remote sensing data at the satellite-borne end, can meet the inference needs of environmental causes in remote sensing images in a diversified and precise manner, significantly improves the interpretability of the interpretation results and the adaptability of the satellite-ground system, and greatly increases the diversity of applicable scenarios for remote sensing monitoring.

[0273] According to one embodiment of this application, a storage medium is provided, the storage medium storing at least one executable instruction that can execute the interpretation method of satellite-ground collaborative remote sensing images based on causal reasoning in any of the above method embodiments.

[0274] Figure 9The diagram shows a structural schematic of a terminal according to one embodiment of the present application. The specific embodiments of the present application do not limit the specific implementation of the terminal.

[0275] like Figure 9 As shown, the terminal may include: a processor 502, a communications interface 504, a memory 506, and a communications bus 508.

[0276] The processor 502, communication interface 504, and memory 506 communicate with each other via communication bus 508.

[0277] Communication interface 504 is used to communicate with other network elements such as clients or other servers.

[0278] The processor 502 is used to execute program 510, which can specifically execute the relevant steps in the above-described embodiment of the satellite-ground collaborative remote sensing image interpretation method based on causal reasoning.

[0279] Specifically, program 510 may include program code that includes computer operation instructions.

[0280] Processor 502 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The terminal includes one or more processors, which may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs.

[0281] Memory 506 is used to store program 510. Memory 506 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0282] Specifically, program 510 can be used to cause processor 502 to perform the following operations:

[0283] The initial interpretation features obtained by the satellite-borne terminal in the initial interpretation of multimodal remote sensing data are obtained, and the initial interpretation features are processed by the first interpretation model that has completed model training to obtain the core interpretation features.

[0284] Causal variables are determined based on the core interpretation features, and a causal variable graph is constructed based on the causal variables. The causal variables are used to characterize the variables that affect the monitoring results in multimodal remote sensing data. The causal variable graph contains the proportion of the influence of different causal variables on different causal associations.

[0285] The causal labels and interpretation results of the labeled regions are determined based on the causal variable diagram, and the verification results corresponding to the causal labels and interpretation results are obtained.

[0286] If the verification result is successful, the learning core features of the first interpretation model are determined based on knowledge distillation and the core interpretation features, and the learning core features are sent to the satellite terminal so that the second interpretation model in the satellite terminal can perform coordinated training and interpretation.

[0287] According to one embodiment of this application, another storage medium is provided, the storage medium storing at least one executable instruction that can execute the interpretation method of satellite-ground collaborative remote sensing images based on causal reasoning in any of the above method embodiments.

[0288] Figure 10 The diagram shows a structural schematic of another terminal provided according to one embodiment of the present application. The specific embodiments of the present application do not limit the specific implementation of the terminal.

[0289] like Figure 10 As shown, the terminal may include: a processor 602, a communications interface 604, a memory 606, and a communications bus 608.

[0290] The processor 602, communication interface 604, and memory 606 communicate with each other via communication bus 608.

[0291] Communication interface 604 is used to communicate with other network elements such as clients or other servers.

[0292] The processor 602 is used to execute program 610, specifically to execute the relevant steps in the above-described embodiment of the satellite-ground collaborative remote sensing image interpretation method based on causal reasoning.

[0293] Specifically, program 610 may include program code that includes computer operation instructions.

[0294] The processor 602 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The terminal includes one or more processors, which may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs.

[0295] Memory 606 is used to store program 610. Memory 606 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0296] Specifically, program 610 can be used to cause processor 602 to perform the following operations:

[0297] Multimodal remote sensing data is acquired, and the multimodal features of the multimodal remote sensing data are fused to obtain a multimodal feature vector;

[0298] The multimodal feature vector is interpreted based on the first interpretation model to generate interpretation features, which are then sent to the ground terminal. The first interpretation model is a lightweight deep learning model.

[0299] The first interpretation model is obtained through coordinated training based on the core features learned from ground-based feedback. The core features are determined by the second interpretation model in the ground-based system after interpretation, and verified by a causal variable graph. The causal variable graph is constructed based on the core interpretation features obtained after interpretation. The causal variable graph contains the influence ratio of different causal variables on different causal associations. The causal variables are used to characterize the variables that affect the monitoring results in multimodal remote sensing data.

[0300] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.

[0301] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for interpreting satellite-ground collaborative remote sensing images based on causal reasoning, applied to the ground end, characterized in that, include: The initial interpretation features obtained by the satellite-borne terminal in the initial interpretation of multimodal remote sensing data are obtained, and the initial interpretation features are processed by the first interpretation model that has completed model training to obtain the core interpretation features. Causal variables are determined based on the core interpretation features, and a causal variable graph is constructed based on the causal variables. The causal variables are used to characterize the variables that affect the monitoring results in multimodal remote sensing data. The causal variable graph contains the proportion of the influence of different causal variables on different causal associations. The causal labels and interpretation results of the labeled regions are determined based on the causal variable diagram, and the verification results corresponding to the causal labels and interpretation results are obtained. When the verification result is successful, the learning core features of the first interpretation model are determined based on knowledge distillation and the core interpretation features, and the learning core features are sent to the satellite terminal so that the second interpretation model in the satellite terminal can perform coordinated training and interpretation. The process of determining causal variables based on the core interpretation features and constructing a causal variable graph based on the causal variables includes: Causal variables are selected according to a preset variable extraction mapping relationship, which includes the matching relationship between different core interpretation features and at least one of different physical variables, different quantitative variables and different correlation variables. Causal relationships are determined based on historical statistical data and geographical knowledge, and an initial causal variable graph is constructed based on the causal variables and the causal relationships. The influence ratio of each causal variable on different causal associations is determined by causal intervention verification, and the influence ratio is configured in the initial causal variable diagram to obtain the causal variable diagram.

2. The method according to claim 1, characterized in that, The determination of the proportion of influence of each causal variable in different causal associations through causal intervention verification includes: The backdoor path is used to filter the causal variables of each causal variable, and the causal variables obtained after the filtering are determined to be causally related, so as to determine the proportion of the influence.

3. The method according to claim 1, characterized in that, After determining the influence ratio of each causal variable in different causal associations through causal intervention verification and configuring the influence ratio in the initial causal variable graph to obtain the causal variable graph, the method further includes: The influence percentages of each causal variable are ranked, and the interpretation results of the influencing factors are determined based on the ranking results.

4. The method according to claim 3, characterized in that, The multimodal remote sensing data includes optical image data, infrared image data, and SAR image data. The step of determining the causal labels and interpretation results of the labeled regions based on the causal variable map, and obtaining the verification results corresponding to the causal labels and interpretation results, includes: The label region corresponding to the multimodal remote sensing data is determined, and the causal labels and interpretation results of the label region are generated based on the interpretation results of the influencing factors. The label region is a feature region corresponding to the core interpretation feature. The causal labels are used to mark the causal variables that affect the core interpretation feature, and the interpretation results are used to characterize the influence of the causal variables.

5. The method according to claim 1, characterized in that, The step of obtaining the verification result corresponding to the cause label and the interpretation result includes: Output the cause label and the interpretation result, and receive the verification result from the user verifying the cause label and the interpretation result.

6. A method for interpreting satellite-ground collaborative remote sensing images based on causal reasoning, applied to a satellite-borne terminal, characterized in that, include: Multimodal remote sensing data is acquired, and the multimodal features of the multimodal remote sensing data are fused to obtain a multimodal feature vector; The multimodal feature vector is interpreted based on the first interpretation model to generate interpreted features, which are then sent to the ground terminal. The first interpretation model is a lightweight deep learning model. The first interpretation model is obtained through coordinated training based on the core features learned from ground-based feedback. The core features are determined by the second interpretation model in the ground-based system after interpretation, and verified by the causal variable graph. The causal variable graph is constructed based on the core interpretation features obtained after interpretation. The causal variable graph contains the influence ratio of different causal variables on different causal associations. The causal variables are used to characterize the variables that affect the monitoring results in multimodal remote sensing data. The causal variable diagram is obtained by selecting causal variables according to the preset variable extraction mapping relationship at the ground end, determining causal relationships based on historical statistical data and geographical knowledge information, constructing an initial causal variable diagram based on the causal variables and the causal relationships, determining the influence ratio of each causal variable on different causal relationships through causal intervention verification, and configuring the influence ratio in the initial causal variable diagram. The preset variable extraction mapping relationship includes the matching relationship between different core interpretation features and at least one of different physical variables, different quantitative variables, and different correlation variables.

7. The method according to claim 6, characterized in that, The multimodal remote sensing data includes optical image data, infrared image data, and SAR image data. Before performing feature fusion on the multimodal features of the multimodal remote sensing data to obtain a multimodal feature vector, the method further includes: Texture spatial variation information is extracted from optical image data that has undergone radiometric and geometric corrections to obtain a gradient intensity map. The gradient mean, gradient variance, and gradient entropy of the gradient intensity map are then combined to obtain texture features. SAR image data that has been radiometrically calibrated and noise-removed is filtered and dimensionality reduced to obtain a backscatter map. The backscatter map is then linearly projected to obtain the backscatter coefficients. Temperature inversion is performed on infrared image data calibrated by thermal infrared band to obtain a surface temperature map, and the surface temperature map is linearly mapped to obtain surface temperature characteristics. The texture features, the backscattering coefficient, and the surface temperature features have the same dimension.

8. The method according to claim 7, characterized in that, The feature fusion of the multimodal features of the multimodal remote sensing data to obtain the multimodal feature vector includes: Multimodal similarity matrices are generated for the texture features, backscattering coefficients, and surface temperature features, respectively. Feature fusion is then performed based on the feature associations and modal weights of the multimodal similarity matrices to obtain a multimodal feature vector. The modal weights are dynamically configured based on different scenarios.

9. The method according to claim 8, characterized in that, The feature fusion based on the feature association and modality weights of the multimodal similarity matrix to obtain the multimodal feature vector includes: Feature associations are determined based on the multimodal similarity matrix, and these feature associations are used to characterize the association strength between features of different modalities. Modal weights are assigned according to different scenarios, including normal scenarios, abnormal high temperature scenarios, and cloud and fog coverage scenarios. The feature associations and modal weights are calculated using a weighted summation method to obtain a multimodal feature vector.

10. The method according to claim 9, characterized in that, Before allocating modal weights according to different scenarios, the method further includes: The feature association consistency index is calculated based on the multimodal similarity matrix, and the mean temperature and standard deviation of the temperature in the whole region are calculated based on the surface temperature feature vector. The surface temperature feature vector is extracted after inversion of the infrared image data. When the feature association consistency index, the average temperature of the entire region, and the standard deviation of the temperature of the entire region match the preset normal scenario conditions, the scenario is determined to be a normal scenario.

11. The method according to claim 9, characterized in that, Before allocating modal weights according to different scenarios, the method further includes: Based on the surface temperature characteristics, calculate the temperature value and standard deviation of the sub-region, and mark the candidate abnormal region when the temperature value and standard deviation of the sub-region match the preset temperature conditions. Once the texture features and the backscattering coefficient are verified, the candidate anomaly region is determined to be a high-temperature anomaly scene.

12. The method according to claim 9, characterized in that, Before allocating modal weights according to different scenarios, the method further includes: The suspected cloud and fog coverage areas were determined based on the average gray value and contrast of the surface temperature characteristics. Once it is determined that there is cloud and fog interference in the suspected cloud and fog coverage area based on the gradient entropy of the texture features, the consistency index of the backscattering coefficient and the land cover of the texture features is calculated. When the land cover consistency index matches the preset cloud and fog scene conditions, the scene is determined to be a cloud and fog coverage scene.

13. An interpretation device for satellite-ground collaborative remote sensing images based on causal reasoning, characterized in that, include: The acquisition module is used to acquire the initial interpretation features obtained by the satellite-borne terminal in the initial interpretation of multimodal remote sensing data, and to perform interpretation processing on the initial interpretation features based on the first interpretation model that has completed model training to obtain the core interpretation features; A construction module is used to determine causal variables based on the core interpretation features and construct a causal variable graph based on the causal variables. The causal variables are used to characterize the variables that affect the monitoring results in multimodal remote sensing data. The causal variable graph contains the proportion of the influence of different causal variables on different causal associations. The determination module is used to determine the causal labels and interpretation results of the label regions based on the causal variable diagram, and to obtain the verification results corresponding to the causal labels and interpretation results; The sending module is used to determine the learning core features of the first interpretation model based on knowledge distillation and the core interpretation features when the verification result is successful, and send the learning core features to the spaceborne terminal so that the second interpretation model in the spaceborne terminal can perform coordinated training and interpretation. The construction module is specifically used to select causal variables according to the preset variable extraction mapping relationship. The preset variable extraction mapping relationship includes the matching relationship between different core interpretation features and at least one of different physical variables, different quantitative variables and different correlation variables. Causal relationships are determined based on historical statistical data and geographical knowledge, and an initial causal variable graph is constructed based on the causal variables and the causal relationships. The influence ratio of each causal variable on different causal relationships is determined through causal intervention verification, and the influence ratio is configured in the initial causal variable graph to obtain the causal variable graph.

14. An interpretation device for satellite-ground collaborative remote sensing images based on causal reasoning, characterized in that, include: The acquisition module is used to acquire multimodal remote sensing data and perform feature fusion on the multimodal features of the multimodal remote sensing data to obtain a multimodal feature vector. The sending module is used to interpret the multimodal feature vector based on the first interpretation model, generate interpretation features, and send them to the ground terminal. The first interpretation model is a lightweight deep learning model. The first interpretation model is obtained through coordinated training based on the core features learned from ground-based feedback. The core features are determined by the second interpretation model in the ground-based system after interpretation, and verified by the causal variable graph. The causal variable graph is constructed based on the core interpretation features obtained after interpretation. The causal variable graph contains the influence ratio of different causal variables on different causal associations. The causal variables are used to characterize the variables that affect the monitoring results in multimodal remote sensing data. The causal variable diagram is obtained by selecting causal variables according to the preset variable extraction mapping relationship at the ground end, determining causal relationships based on historical statistical data and geographical knowledge information, constructing an initial causal variable diagram based on the causal variables and the causal relationships, determining the influence ratio of each causal variable on different causal relationships through causal intervention verification, and configuring the influence ratio in the initial causal variable diagram. The preset variable extraction mapping relationship includes the matching relationship between different core interpretation features and at least one of different physical variables, different quantitative variables, and different correlation variables.

15. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method of claim 1.

16. A computer device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method of claim 1.

17. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method of claim 6.

18. A computer device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method of claim 6.

Citation Information

Patent Citations

  • Satellite-ground collaborative black and odorous water body recognition model automatic optimization system

    CN115631408A

  • Satellite-ground cooperation-based satellite-borne ground target recognition algorithm automatic updating method

    CN117372883A