A method and system for screening a soil health evaluation index system based on multi-modal data

By constructing a causal graph and making causal intervention inferences, the problem of insufficient data integration in the existing soil health assessment index screening scheme is solved. This improves the accuracy and reliability of the soil health assessment index system, which can truly reflect the physical driving mechanism of soil environmental evolution and provide accurate data support for actual intervention decisions.

CN122492032APending Publication Date: 2026-07-31INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS
Filing Date
2026-06-29
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing soil health assessment index screening schemes rely on single-dimensional quantitative detection data, lacking effective feature extraction and quantitative fusion of unstructured qualitative survey semantic data. This results in limited data representation space, making it difficult to comprehensively depict the complex soil habitat state. Furthermore, traditional statistical algorithms struggle to isolate data noise caused by multicollinearity and network coupling, leading to low accuracy and robustness of the evaluation index system.

Method used

By acquiring quantitative monitoring data and qualitative survey semantic data, and combining them with soil evolution causal constraint rules to construct a causal graph, causal intervention inferences are made to filter out spurious correlation interference and improve the accuracy and reliability of indicator selection.

Benefits of technology

It has achieved accurate identification and improved reliability of the soil health assessment index system, which can truly reflect the physical driving mechanism of soil environmental evolution and provide accurate data support for actual intervention decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492032A_ABST
    Figure CN122492032A_ABST
Patent Text Reader

Abstract

This application discloses a method and system for screening soil health assessment indexes based on multimodal data. This application can acquire initial multimodal soil state characteristic data corresponding to each sub-region to be evaluated within a target area, including quantitative monitoring data and qualitative survey semantic data; acquire preset soil evolution causal constraint rules and construct an initial soil characteristic causal graph accordingly; perform feature mapping processing on the qualitative survey semantic data and quantitative monitoring data to obtain aligned multimodal characteristic data, and fuse it into the corresponding nodes of the initial soil characteristic causal graph to obtain the target soil characteristic causal graph; perform causal intervention processing on each node in the graph to obtain the average causal effect value corresponding to various factors affecting soil state; and screen the target factors affecting soil state from these factors. Therefore, this application improves the accuracy and reliability of soil health assessment index screening.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, specifically to a method and system for screening soil health evaluation index systems based on multimodal data. Background Technology

[0002] Accurate quantitative assessment of soil health status is a core data processing task in the field of intelligent ecological and environmental monitoring. The reliability of the evaluation results directly depends on the accuracy of multi-source heterogeneous data fusion during the feature index screening process, as well as the technical implementation level of the degree of restoration of the real driving mechanism. Accurately identifying and screening key feature indicators of soil health is the data foundation for building subsequent environmental dynamic monitoring and intervention decision-making models.

[0003] Currently, when screening soil health assessment indicators, a data processing scheme based on statistical correlation is typically used. Specifically, the system usually first acquires structured quantitative detection data such as soil physicochemical properties within the target area, and then uses conventional statistical algorithms such as correlation coefficient matrix or principal component analysis to calculate the correlation weight between each quantitative detection data and the system state. Finally, the evaluation indicators are ranked and screened based on the magnitude of the correlation weight.

[0004] However, the aforementioned screening schemes have significant technical limitations in actual data processing: on the one hand, existing schemes mainly rely on single-dimensional quantitative detection data, lacking effective feature extraction and quantitative fusion mechanisms for unstructured qualitative survey semantic data such as farmland appearance, resulting in limited data representation space on the system input side and difficulty in comprehensively depicting complex soil habitat conditions; on the other hand, existing feature weight allocation mechanisms are inherently highly dependent on statistical covariance among data. In highly nonlinear dynamic environmental systems, complex coupling feedback effects often exist between multidimensional variables, and traditional statistical algorithms struggle to effectively remove data noise caused by multicollinearity and network coupling, easily extracting deceptive redundant correlation features. This screening mechanism based on appearance correlation cannot truly reflect the intrinsic driving force of various indicators on the evolution of system state, resulting in low accuracy and robustness of the final output evaluation index system, failing to provide accurate data support for actual intervention decisions. Summary of the Invention

[0005] This application provides a method and system for screening soil health assessment indicators based on multimodal data. By aligning and integrating quantitative monitoring data with qualitative survey semantic data into a causal graph constrained by physical evolution, and by performing causal intervention inference based on graph networks to effectively filter out spurious correlation interference between data, the accuracy and reliability of soil health assessment indicator screening are improved, which is conducive to providing accurate data support for actual intervention decisions.

[0006] This application provides a method for screening a soil health assessment index system based on multimodal data. The method includes:

[0007] Acquire initial multimodal soil state characteristic data corresponding to multiple sub-regions to be evaluated within the target area. The initial multimodal soil state characteristic data includes quantitative monitoring data and qualitative survey semantic data. The quantitative monitoring data contains information characterizing the objective physical and chemical properties of the soil, while the qualitative survey semantic data contains information characterizing the subjective apparent state of the soil.

[0008] Obtain the preset causal constraint rules for soil evolution. These rules characterize the inherent driving relationships of various factors influencing soil state during the physical and chemical evolution of soil.

[0009] An initial soil characteristic causal graph is constructed based on the soil evolution causal constraint rules. Nodes in the initial soil characteristic causal graph represent factors that affect soil state, and edges represent directed causal transmission paths restricted by the soil evolution causal constraint rules.

[0010] The qualitative survey semantic data and quantitative monitoring data are processed by feature mapping to obtain aligned multimodal feature data. The aligned multimodal feature data is then fused into the corresponding nodes of the initial soil feature causal graph to obtain the target soil feature causal graph. The target soil feature causal graph is a causal directed graph model in which each node is assigned aligned multimodal feature data.

[0011] Causal intervention was performed on each node in the causal diagram of the target soil characteristics to obtain the average causal effect value corresponding to various factors affecting soil state. The average causal effect value reflects the true impact of various factors affecting soil state on soil health under the condition of cutting off the transmission path of confounding factors.

[0012] Based on the average causal effect value, indicators were screened from various factors affecting soil state to obtain target factors affecting soil state. These target factors are the core indicators for soil health assessment after filtering out statistical pseudo-correlation.

[0013] This application also provides a screening system for a soil health assessment index system based on multimodal data, the system comprising:

[0014] The multimodal feature acquisition module is used to acquire initial multimodal soil state feature data corresponding to multiple sub-regions to be evaluated within the target area. The initial multimodal soil state feature data includes quantitative monitoring data and qualitative survey semantic data. The quantitative monitoring data contains information characterizing the objective physical and chemical properties of the soil, and the qualitative survey semantic data contains information characterizing the subjective appearance of the soil.

[0015] The evolution constraint rule acquisition module is used to acquire preset soil evolution causal constraint rules. The soil evolution causal constraint rules characterize the inherent driving relationship of various factors affecting soil state in the process of soil physical and chemical evolution.

[0016] The initial causal graph construction module is used to construct an initial soil characteristic causal graph based on the soil evolution causal constraint rules. The nodes in the initial soil characteristic causal graph represent factors that affect soil state, and the edges represent directed causal transmission paths restricted by the soil evolution causal constraint rules.

[0017] The multimodal causal instantiation module is used to perform feature mapping processing on qualitative survey semantic data and quantitative monitoring data to obtain aligned multimodal feature data. The aligned multimodal feature data is then fused into the corresponding nodes of the initial soil feature causal graph to obtain the target soil feature causal graph. The target soil feature causal graph is a causal directed graph model in which each node is assigned aligned multimodal feature data.

[0018] The causal intervention assessment module is used to perform causal intervention processing on each node in the causal diagram of the target soil characteristics to obtain the average causal effect value corresponding to various factors affecting soil state. The average causal effect value reflects the true impact of various factors affecting soil state on soil health under the condition of cutting off the transmission path of confounding factors.

[0019] The core indicator screening module is used to screen indicators from various factors affecting soil state based on the average causal effect value, and obtain the target factors affecting soil state. The target factors affecting soil state are the core indicators for soil health assessment that have been filtered out of statistical pseudo-correlation.

[0020] This application also provides an electronic device, including a processor and a memory, the memory storing multiple instructions; the processor loads instructions from the memory to execute the steps in any of the screening methods for a soil health evaluation index system based on multimodal data provided in this application.

[0021] This application also provides a computer-readable storage medium storing multiple instructions adapted for loading by a processor to execute steps in any of the screening methods for a soil health evaluation index system based on multimodal data provided in this application.

[0022] In this application, initial multimodal soil state characteristic data corresponding to multiple sub-regions to be evaluated within the target area are first obtained. This data includes quantitative monitoring data and qualitative survey semantic data. The quantitative monitoring data contains information characterizing the objective physicochemical properties of the soil, while the qualitative survey semantic data contains information characterizing the subjective apparent state of the soil. Since single-dimensional characteristic data cannot comprehensively depict the system state, and there is a problem of cross-modal incompatibility between qualitative semantic data and quantitative numerical data, an initial soil characteristic causal graph is constructed by obtaining soil evolution causal constraint rules characterizing inherent driving relationships. Feature mapping processing is then performed on the qualitative survey semantic data and quantitative monitoring data, and the aligned multimodal characteristic data is fused into the corresponding nodes of the initial soil characteristic causal graph, thereby obtaining the target soil characteristic causal graph. Then, based on this target soil characteristic causal graph, causal intervention processing is performed on each node in the graph to obtain the average causal effect value corresponding to various factors affecting soil state. Furthermore, this average causal effect value can truly reflect the degree of influence of various factors affecting soil state on soil health under the condition of cutting off the transmission path of confounding factors. Based on the average causal effect value, indicators are screened from various factors affecting soil state to obtain target factors affecting soil state. This effectively filters out the interference of statistical spurious correlations, enabling dynamic and accurate identification of core evaluation indicators. The constructed indicator system can objectively reflect the true physical driving mechanism of soil environmental evolution. Therefore, this application improves the accuracy and reliability of soil health evaluation indicator selection by aligning and integrating quantitative monitoring data with qualitative survey semantic data into a causal graph constrained by physical evolution, and by performing causal intervention inference based on graph networks to effectively filter out spurious correlation interference between data. This provides accurate data support for actual intervention decisions. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a flowchart illustrating a method for screening a soil health evaluation index system based on multimodal data, as provided in an embodiment of this application.

[0025] Figure 2 This is a schematic diagram of the structure of a screening system for a soil health evaluation index system based on multimodal data, provided in an embodiment of this application. Detailed Implementation

[0026] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0027] This application provides a method and system for screening soil health evaluation index systems based on multimodal data.

[0028] The following sections provide detailed descriptions of each example. It should be noted that the sequence numbers of the following embodiments are not intended to limit the preferred order of the embodiments.

[0029] In this embodiment, a method for screening a soil health evaluation index system based on multimodal data is provided, such as... Figure 1 The specific process of the screening method for a soil health assessment index system based on multimodal data is as follows:

[0030] S101. Obtain initial multimodal soil state characteristic data corresponding to multiple sub-regions to be evaluated within the target area. The initial multimodal soil state characteristic data includes quantitative monitoring data and qualitative survey semantic data. The quantitative monitoring data contains information characterizing the objective physical and chemical properties of the soil, and the qualitative survey semantic data contains information characterizing the subjective appearance of the soil.

[0031] The target area is the macro-geographical spatial range within which an overall soil health assessment and monitoring are required. The target area can be defined based on administrative divisions, natural topography, or agricultural planning. For example, the target area could be all basic farmland within the jurisdiction of a city, a specific watershed ecological restoration demonstration area, or a large-scale modern agricultural planting base covering thousands of acres.

[0032] The sub-region to be evaluated refers to an independent micro-evaluation unit obtained after spatially dividing the target area into grids or dividing it according to natural plot boundaries. To achieve accurate evaluation, the target area is usually composed of multiple sub-regions to be evaluated. For example, the sub-region to be evaluated can be a standard sampling area with a specific grid number within the aforementioned large planting base (such as a 10m×10m spatial grid), a single experimental field with independent micro-topographic features, or a single cultivated plot independently contracted by a specific farmer.

[0033] The objective physicochemical properties of soil refer to the inherent material composition, structural state, and chemical reaction characteristics of soil, which are not affected by human subjective perception. They encompass absolute physicochemical indicators of soil at both the microscopic and macroscopic levels, mainly manifested in the soil's nutrient state, pH, pore structure characteristics, heavy metal enrichment level, and water and heat conduction properties, among other physical and chemical attributes.

[0034] Quantitative monitoring data refers to structured, objective data with clear numerical values ​​and physical dimensions obtained by measuring the objective physicochemical properties of soil using specialized testing instruments, sensor equipment, or laboratory physicochemical analysis methods. Quantitative monitoring data can include real-time data such as soil moisture content (e.g., 25%), soil temperature (e.g., 15°C), and electrical conductivity collected in real time by in-situ IoT sensors in the field; it can also include laboratory test results for soil organic matter content (e.g., 20 g / kg), soil pH value (e.g., 6.5), total nitrogen content, available phosphorus content, and concentrations of specific heavy metals (e.g., cadmium, lead), obtained through standard laboratory procedures.

[0035] The subjective appearance of soil refers to the perception and empirical evaluation of the macroscopic habitat characteristics and suitability for cultivation of soil based on human vision, touch and long-term field operation experience. It reflects the comprehensive characteristics of soil in actual agricultural production, such as ease of cultivation, degree of compaction, color variation, and intuitive manifestation of plant root development, which are difficult to be directly quantified by conventional instruments.

[0036] Qualitative survey semantic data refers to unstructured or semi-structured text data formed by textual and semantic descriptions of the subjective appearance of soil through manual or semi-manual information collection methods such as farmer interviews, expert questionnaires, field inspections, or extraction of agricultural records. Qualitative survey semantic data usually exists in the form of natural language text labels or fuzzy rating terms. For example, it may be semantic information records such as farmers' feedback such as "severe soil compaction", "obvious cracking of topsoil after rain", "extremely high tillage resistance", "dark gray soil color", "extremely slow water infiltration", or "heavy texture".

[0037] Initial multimodal soil state characteristic data refers to the original multi-source datasets acquired synchronously for the same sub-region to be evaluated, which have not yet undergone cross-modal alignment and feature mapping processing. This data is a collection of the aforementioned "quantitative monitoring data" and "qualitative survey semantic data". As the underlying input of the evaluation system of this application, it breaks through the limitations of a single data dimension, including not only absolute values ​​representing the underlying physicochemical habitat (quantitative modality), but also natural language text representing the macro-cultivation state (qualitative modality), thus constituting a heterogeneous data source capable of jointly depicting the original evolutionary state of the soil in this sub-region from both subjective and objective perspectives.

[0038] In some embodiments, the problem of spatiotemporal misalignment of features caused by inconsistent sampling frequencies of various field IoT sensors and network transmission delays is effectively solved, and dirty data noise caused by equipment failure or extreme environmental changes is intelligently filtered out, thereby completely eliminating underlying data gaps and distribution distortions, ensuring the accuracy of subsequent causal network node feature mapping and the robustness of causal intervention inference from the source. Specifically, before acquiring the initial multimodal soil state feature data corresponding to multiple sub-regions to be evaluated within the target area, the method further includes:

[0039] Obtain multiple types of initial observation value sequences corresponding to each sub-region to be evaluated. The observation values ​​in the initial observation value sequences carry timestamps, and different types of initial observation value sequences correspond to different soil state observation targets.

[0040] Based on the regional identifier information and timestamps corresponding to each sub-region to be evaluated, the initial observation value sequences of multiple types are subjected to joint spatiotemporal alignment processing to obtain spatiotemporally normalized observation data. The spatiotemporally normalized observation data is a data set that merges different types of observation values ​​belonging to the same sub-region to be evaluated within a preset time window to a unified time reference point.

[0041] By utilizing the characteristics of temporal continuity and spatial proximity, missing value detection and imputation are performed on spatiotemporally normalized observation data to obtain spatiotemporally continuous observation data. The spatiotemporally continuous observation data is a complete set of data with inferred and imputed missing data nodes.

[0042] Based on the physical reasonable range and statistical distribution characteristics of each soil condition observation target, outlier identification and removal are performed on the spatiotemporal continuous observation data to obtain quantitative monitoring data. The quantitative monitoring data is observation enhancement characterization data that has been filtered out of sampling errors and environmental abrupt changes.

[0043] The "type" refers to the classification of the specific physical, chemical, or biological detection mechanism upon which the data acquisition terminal or underlying sensor relies. For example, types can include electrochemical detection based on potential analysis mechanisms (such as in-situ soil pH sensors), optical remote sensing based on electromagnetic spectrum reflection mechanisms (such as multispectral or hyperspectral cameras), temperature and humidity sensing based on thermodynamic and dielectric constant mechanisms (such as soil temperature and humidity probes), and trace element analysis based on X-ray fluorescence mechanisms (such as portable heavy metal XRF spectrometers), etc.

[0044] Observations are structured digital quantities output by target data acquisition equipment after performing a single detection action at a specific time node and spatial coordinates, used to accurately characterize the objective physical or chemical state of the underlying soil. Observations are usually expressed as numerical values ​​with clear physical dimensions. For example, soil volumetric water content output by temperature and humidity sensors (e.g., 15.2%), effective hydrogen ion concentration index output by electrochemical sensors (e.g., pH 6.5), and specific heavy metal enrichment concentration values ​​detected and output by trace element analysis equipment (e.g., fluorescence spectrometer), such as the mass fraction of cadmium (Cd) in the soil (e.g., 0.25 mg / kg) or the absolute content of lead (Pb) (e.g., 35.2 mg / kg).

[0045] A timestamp is a precise system time identifier that records the generation or uploading of a corresponding observation. It can be a standard Unix timestamp or time series data formatted as "2024-05-2014:30:00".

[0046] Soil condition monitoring targets refer to specific physical or chemical parameters that are continuously tracked and measured using specific types of sensors. Examples include topsoil moisture content, available phosphorus concentration, and nitrate nitrogen content.

[0047] An initial observation sequence refers to a one-dimensional set of original observations arranged in chronological order for the same soil condition observation target. Due to the different sampling frequencies of different types of sensors (e.g., temperature sensors sample once every 10 minutes, while nutrient sensors sample once every 4 hours), multiple initial observation sequences with varying temporal densities will be formed.

[0048] Regional identification information refers to spatial coding or coordinate system data used to uniquely distinguish and locate different sub-regions to be evaluated within a macro target area. It can be a polygonal boundary coordinate set composed of GPS latitude and longitude, or a unique grid hash code assigned in a GIS system.

[0049] Joint spatiotemporal alignment processing refers to the algorithmic operation of merging multi-source, heterogeneous, and asynchronously sampled data streams into a unified physical space and temporal dimension. The specific processing steps include: parsing the geographic coordinates inherent in each initial observation sequence and mapping them to the sub-regions to be evaluated corresponding to the regional identifiers using a hash matching algorithm; dividing the global time axis into continuous discrete time slots with a preset frequency (e.g., 1 hour) as the step size; extracting the timestamps carried by the observations and determining which time slot they fall into; if multiple observations from the same type of sensor fall into the same time window, generating a unique representative value using a mean aggregation or median aggregation algorithm; if no data falls into the same time window, leaving the corresponding matrix node empty (denoted as NaN).

[0050] A preset time window refers to a fixed span tolerance range used to divide time slots when performing joint spatiotemporal alignment processing. Considering the relatively slow evolution of agricultural habitats, this preset time window can typically be set to 1 hour or 24 hours.

[0051] Spatiotemporally normalized observation data refers to the structured data output after the above-mentioned joint spatiotemporal alignment processing. It is a multi-dimensional matrix set after strictly aligning different types of observations belonging to the same sub-region to be evaluated within a preset time window to a unified time reference point.

[0052] Temporal continuity and spatial proximity characteristics refer to the smooth autocorrelation of soil environmental variables over time and the distance-decreasing correlation in geographic spatial distribution. Temporal continuity indicates that soil moisture at the current moment is highly correlated with that of the previous hour; spatial proximity indicates that two adjacent grids 5 meters apart should have highly similar soil pH values ​​without human intervention.

[0053] Missing value detection and imputation refers to the algorithmic operation of scanning null nodes in the data matrix and estimating the true value using surrounding valid data. The specific process includes: using a sliding window algorithm to traverse the spatiotemporally normalized observation data matrix to locate data nodes with NaN values. If the missing window of this node in the temporal dimension is short (e.g., consecutive missing values ​​less than 3 steps), then using the temporal continuity feature, cubic spline interpolation or Lagrange interpolation algorithms are used to imput the temporal dimension based on its preceding and following historical data. If the node has severe missing values ​​in the temporal dimension, but its adjacent sub-regions to be evaluated have valid observations at the same time, then using the spatial proximity feature, inverse distance weighted interpolation (IDW) or Kriging spatial interpolation algorithms are used to cross-implement the spatial dimension.

[0054] Spatiotemporal continuous observation data refers to a complete set of multidimensional tensors that, after the above-mentioned filling process, eliminates data gaps and has valid values ​​for all time reference points and spatial grids.

[0055] Physically reasonable range and statistical distribution characteristics are objective criteria for judging the authenticity of data. Physically reasonable range refers to the absolute extreme value boundary that conforms to natural laws (for example, soil moisture content cannot be negative, and pH value must be between 0 and 14); statistical distribution characteristics refer to the probability density distribution law that the data follows at the historical population level (for example, the monthly average temperature of a certain region follows a normal distribution, and its extreme values ​​should not exceed the range of the mean plus or minus 3 standard deviations).

[0056] Outlier identification and removal refers to the algorithmic operation of identifying and filtering out abrupt, contaminated data that violates natural laws or statistical distributions from continuous sequences. The specific process includes: retrieving the threshold library of physically reasonable ranges corresponding to each soil condition observation target; directly identifying and removing observations exceeding the absolute physical upper and lower limits in the spatiotemporal continuous observation data as extreme values ​​of equipment failure; based on statistical distribution characteristics, introducing the Isolation Forest algorithm or the 3Sigma statistical rule algorithm to calculate the distribution deviation score of each observation; identifying and removing observations with scores exceeding the tolerance threshold as outliers caused by accidental environmental mutations; and then using the aforementioned interpolation algorithm to perform secondary smoothing and filling of the vacant values, ultimately outputting the quantitative monitoring data with a high signal-to-noise ratio.

[0057] S102. Obtain the preset causal constraint rules for soil evolution. The causal constraint rules for soil evolution characterize the inherent driving relationship of various factors affecting soil state in the process of soil physical and chemical evolution.

[0058] Among them, soil state factors refer to various endogenous or exogenous independent variable characteristics in the ecological, physical, and chemical evolution network of the soil system that can cause changes in soil quality and fluctuations in physicochemical properties. Soil state factors include, but are not limited to: macro-environmental and topographical factors (such as annual precipitation, surface slope, and groundwater dynamics), factors related to human agricultural management activities (such as the amount of chemical fertilizer applied, the proportion of straw returned to the field, and the frequency of human irrigation), and endogenous characteristic factors reflecting the heterogeneity of the soil's own physicochemical habitat (such as the clay content in the soil's mechanical composition, the type of natural parent material, and the enrichment abundance of exogenous or inherent heavy metal elements).

[0059] Soil physical and chemical evolution refers to the material cycling, energy flow, and structural dynamics of multiphase soil media under the interaction of natural habitats and human cultivation. For example, the long-term excessive application of ammonium nitrogen fertilizer leads to the release of hydrogen ions through nitrification by soil microorganisms, resulting in the "soil acidification evolution process"; or the "heavy metal migration evolution process" occurs when heavy metals (such as cadmium) enter the soil with irrigation water and undergo adsorption, desorption, precipitation, and dissolution under specific pH conditions, thereby altering their form and bioavailability.

[0060] The inherent driving relationship refers to the objectively existing asymmetric, irreversible, unidirectional causal transmission path among various factors affecting soil state, based on underlying physical, chemical, and agro-meteorological mechanisms. This driving relationship is not a numerical correlation derived from statistical calculations, but rather an inherent law of the objective world. For example, applying large amounts of nitrogen fertilizer will directly drive a decrease in soil pH, which is a unidirectional driving force with a "physicochemical mechanism." Conversely, adjusting soil pH will not automatically and spontaneously increase the total nitrogen content in the soil.

[0061] The pre-defined causal constraint rules for soil evolution refer to transforming the inherent driving relationships among the aforementioned factors influencing soil state into hard-coded control logic or conditional graph rules that can be directly identified by the graph network model and used as the basis for topology pruning. This manifests as a priori directional constraint logic, such as: fertilizer application rate → organic matter content = true, and organic matter content → fertilizer application rate = false. This restricts the causal network to only propagate in one direction when generating edges, thereby eliminating reverse spurious associations.

[0062] In some embodiments, the creation process of the pre-defined causal constraint rules for soil evolution involves the system pre-collecting standard agricultural scientific literature databases, soil quality standards (such as the "Soil Environmental Quality Standard"), farmland improvement technical specifications, and experience texts from senior experts. Using Named Entity Recognition (NER) technology in Natural Language Processing (NLP), technical entities corresponding to "factors affecting soil state" are automatically extracted, along with verb phrases representing physical evolution drivers (such as "promote," "lead to," "exacerbate," and "passivate"). Then, the extracted entities are used as graph nodes, and the verb phrases are transformed into directed edges labeled with physical mechanisms. For example, based on the classic chemical mechanism of "long-term application of physiologically acidic fertilizers leading to soil acidification," a unidirectional directed edge is created in the graph from [physiologically acidic fertilizer application rate] to [soil pH value], and it is assigned the attribute label [hydrogen ion replacement mechanism], thereby constructing a closed-loop soil ecological mechanism knowledge graph at the semantic layer. Finally, the directed edge logic in the soil ecological mechanism knowledge graph is transformed into a constraint condition matrix (i.e., adjacency matrix constraint) executable by a computer graph network. If there is a unidirectional physical drive from node A to node B in the graph, then in the preset constraint topology matrix, element M[A][B] is marked as "causal edges are allowed (1)", and the reverse element M[B][A] is forcibly hard-coded as "causal edges are strictly prohibited (0)", and some fixed edges restricted by absolute natural laws are set (such as terrain cannot be driven in reverse by soil nutrients). Through this matrix-based fixed writing, the preset soil evolution causal constraint rules are finally generated, which are used to perform strong constraint pruning on the structural search space of the graph model in subsequent steps.

[0063] S103. Construct an initial soil characteristic causal graph based on the soil evolution causal constraint rules. The nodes in the initial soil characteristic causal graph represent factors that affect soil state, and the edges represent directed causal transmission paths restricted by the soil evolution causal constraint rules.

[0064] In this context, a node refers to a core data storage unit in a graph network data structure, used to carry a specific feature space and represent an independent physical or chemical entity. In this embodiment, each node is uniquely mapped to a "soil state influencing factor" as defined above, and each node is assigned a unique index number in computer memory (e.g., V1, V2, ..., V...). n It also includes a data slot for dynamically attaching multimodal features. For example, node V1 is mapped to [chemical fertilizer application rate], node V2 to [soil pH value], and node V3 to [heavy metal bioavailability].

[0065] An edge is a data link used to connect any two different nodes in a graph network, representing the relationship between the two nodes. In the embodiments of this application, the edge is not a statistical connection generated freely by pure data, but a topological control line whose "existence" and "direction" must be strictly restricted by the aforementioned "soil evolution causal constraint rule" matrix.

[0066] A directed causal transmission path refers to a chain of directed edges, each with a clearly defined start and end node, connected end-to-end, representing a unidirectional physical driving force of a physicochemical mechanism. For example, consider a directed causal transmission path starting from node V1, passing through node V2, and finally reaching node V3: [Chemical fertilizer application rate V1 → Soil pH value V2 → Effective heavy metal activity V3]. This path indicates that V1 is the physical cause of V2, and V2 is the physical cause of V3, thus forming a unidirectional transmission channel for energy and matter evolution in a graph network.

[0067] The initial soil feature causal graph refers to the initial graph network architecture generated purely based on prior physical evolution mechanism constraints before the system is injected with measured multimodal feature data. It is stored in the form of a directed acyclic graph (DAG) or a topological adjacency matrix.

[0068] In the specific execution process, the system constructs the initial soil characteristic causal graph as follows: Based on the pre-analyzed n factors affecting soil state, the computer system allocates a corresponding graph node object array V={V1,V2,…,V...} in memory. n At this point, the feature vector slots of each node are initialized to empty (e.g., null or zero vector). The system retrieves the preset soil evolution causal constraint rule matrix M created in step S102 above. n×n Traverse any pair of nodes (V) i V j The system reads the corresponding element value from the constraint matrix. When matrix element M[i][j]=1, the system creates a slave node V in memory. i Pointing to node V j Directed edge pointer Eij =〈V i V j > Construct the initial topology; when matrix element M[i][j]=0, the system forcibly locks and prohibits the establishment of a new topology from V. i To V j Any connections. After traversal, a graph data structure object G, consisting of a set of nodes V and a set of directed edges E, is generated in the computer. initial =〈V i V j The initial soil characteristic causal graph is perfectly constrained in topology by the underlying physicochemical evolution mechanism, thus achieving active isolation of "false transmission paths" at the algorithm framework level.

[0069] S104. Perform feature mapping processing on the qualitative survey semantic data and the quantitative monitoring data to obtain aligned multimodal feature data, and then fuse the aligned multimodal feature data into the corresponding nodes of the initial soil feature causal graph to obtain the target soil feature causal graph. The target soil feature causal graph is a causal directed graph model in which each node is assigned aligned multimodal feature data.

[0070] Feature mapping refers to the mathematical computation process of projecting and transforming heterogeneous data that originally belong to different data modalities and have different dimensions and scales to the same low-dimensional continuous dense vector space (i.e., a unified high-dimensional feature space) using feature extraction algorithms or transformation matrices.

[0071] The specific processing steps include: extracting the qualitative survey semantic data from step S101 (such as the text paragraph "severe soil compaction" input by farmers), calling a pre-trained lightweight text encoder (such as a semantic extraction network model composed of a multi-layer Transformer architecture), and mapping it into a dense continuous feature vector of fixed dimensions. .

[0072] Then, quantitative monitoring data (such as pH value 6.5 and heavy metal mass fraction 0.25 mg / kg) from the same sub-region and at the same time point are obtained. These data are then transformed into numerical feature vectors with the same dimensions using a fully connected neural network (MLP) or a one-dimensional mapping matrix. The two-stream mapping mechanism described above eliminates the heterogeneity of data format and dimensions at the algorithm level.

[0073] Aligned multimodal feature data refers to a set of multi-source features that, after the aforementioned feature mapping process, achieve complete matching in spatial, temporal, and semantic dimensions and have the same mathematical representation (i.e., same-dimensional vector format). For example, for a specific sub-region to be evaluated. The corresponding quantitative physicochemical indicators and qualitative apparent texts are mapped and transformed into joint correlation feature tensors at the same time. (in (This represents a vector concatenation or fusion operator), where the qualitative and quantitative components in the tensor have been double-aligned in both spatial coordinates and time steps.

[0074] The target soil feature causal graph refers to the feature-rich network model generated by dynamically injecting and mounting the aligned multimodal feature data as attribute vectors into the corresponding nodes of the initial soil feature causal graph.

[0075] A feature-rich graph network model refers to a data structure based on an initial soil feature causal graph, where the data slots of each node are fully written with aligned multimodal feature data. Specifically, the "rich features" in this model are manifested in the fact that the nodes in the graph are no longer empty nodes without injected data, but are substantially loaded with joint correlation feature tensors (i.e., attribute vectors) composed of the aforementioned quantitative monitoring data and qualitative survey semantic data. The "graph network" in this model is manifested in the fact that these nodes loaded with joint correlation feature tensors are connected strictly according to the directed causal transmission paths constrained by the aforementioned soil evolution causal rules. Within a single structure, this feature-rich graph network model contains both dense, continuous feature vectors representing factors influencing soil state and retains topological connections representing unidirectional physical drives of physicochemical mechanisms.

[0076] The specific process by which the system fuses the aligned multimodal feature data into the corresponding nodes of the initial soil feature causal graph is as follows: First, the system traverses the initial soil feature causal graph. Each empty node object in the graph network is used. Based on the attribute labels carried by the feature data, its corresponding memory address in the graph network is located. Then, the aligned multimodal feature data is written into the feature vector slot (Data Slot) of the corresponding node. For example, this involves fusing features containing the semantic vector of "soil compaction" and the numerical vector of "volume limiting water content". Assign a value to the node (Soil texture / moisture node). Repeat the above steps until all nodes in the initial graph are updated from "empty nodes" to "feature-rich nodes", and finally encapsulate and output the target soil feature causal graph object in memory.

[0077] The causal directed graph model is a high-level composite data structure that combines "structural causal relationship constraints" with "multimodal high-dimensional feature representation." This model is no longer a purely symbolic graph, nor a black-box neural network lacking mechanistic constraints. Its topological structure (edges) represents the causal transmission direction restricted by absolute physical and chemical evolutionary laws, while its entities (nodes) carry dense feature vectors reflecting the multimodal distribution of the real world. As the core underlying computational foundation of this application, it enables the system to possess both the ability to perceive and represent complex multimodal data using machine learning and the reasoning ability to intervene and control through causal inference algorithms.

[0078] In some embodiments, by introducing fuzzy semantic set theory and scale-space transformation mechanisms, this method achieves for the first time the conversion of farmers' colloquial and highly subjective qualitative descriptive text into feature parameters with the same dimensions and numerical scale as objective test reports. Dynamic weighting based on the reliability of the multi-source data itself completely eliminates the measurement barriers between cross-modal data at the algorithm's source, effectively removing fuzzy noise and measurement errors in human subjective perception, and significantly improving the precision and reliability of the input multimodal data. Specifically, feature mapping processing is performed on qualitative survey semantic data and quantitative monitoring data to obtain aligned multimodal feature data, including:

[0079] Based on a pre-defined fuzzy semantic set of soil conditions, parameters are extracted from the qualitative survey semantic data to obtain semantic distribution feature parameters, which reflect the fuzzy distribution state of the qualitative survey semantic data.

[0080] Based on the scale space of the quantitative monitoring data, the semantic distribution feature parameters are scale normalized and numerically transformed to obtain qualitative survey numerical data, which is data with the same dimensions as the quantitative monitoring data.

[0081] Based on the reliability characterization value of multimodal data, the qualitative survey numerical data and quantitative monitoring data are weighted and fused to obtain aligned multimodal feature data. The aligned multimodal feature data is feature data that has eliminated dimensional differences and subjective cognitive noise.

[0082] The pre-defined soil state fuzzy semantic set refers to a pre-constructed standard benchmark library used to map fuzzy terms describing the apparent state of soil in human natural language to mathematical fuzzy sets or membership functions. For example, for the observation target of "soil compaction degree", the fuzzy semantic set may include standard semantic labels such as "slight compaction", "moderate compaction", and "severe hardening by heavy metals", with each label corresponding to a specific triangular fuzzy number or Gaussian membership function at the underlying level.

[0083] Parameter extraction processing refers to the text mining and mathematical quantification process of identifying the original qualitative text and calculating its membership parameters in the fuzzy semantic set. The specific process includes: the system uses a text similarity matching algorithm (such as cosine similarity or keyword matching) to compare the input qualitative survey semantic data (e.g., "This plot of land is severely cracked and hard after the rain") with a preset fuzzy semantic set of soil conditions, identifying the corresponding standard fuzzy semantic label as "severe compaction". The system then retrieves the preset fuzzy function corresponding to this label, using the text's tone intensity as the operator input, to calculate the expected value, hyperentropy, and left and right boundary values ​​of the text in the fuzzy matrix, generating a one-dimensional or multi-dimensional feature vector.

[0084] Semantic distribution feature parameters refer to the feature vectors output after the aforementioned parameter extraction and processing, which can express the fuzzy uncertainty in qualitative text using precise mathematical language. For example, the triplet parameters extracted using the CloudModel algorithm. , where the expected value It expresses the core point of the probability distribution of hardening reflected in the text: entropy. and hyperentropy This quantitatively expresses the ambiguity and randomness of the farmer's subjective experience (i.e., subjective cognitive noise).

[0085] The scale space of quantitative monitoring data is a mathematical space consisting of the numerical variation domain, discrete resolution, and probability distribution boundary of quantitative monitoring data collected by objective instruments under specific physical dimensions. For example, if the current quantitative data is the soil penetration resistance value (used to measure compaction), its scale space is [0MPa, 5MPa], and its resolution is 0.01MPa.

[0086] Scale normalization and numerical transformation refers to the mathematical transformation path that proportionally maps semantic distribution feature parameters, which have no physical dimensions, to the physical domain of quantitative data. The specific process includes extracting the expected value from the semantic distribution feature parameters. (It is usually in the [0,1] interval). Read the scale space upper and lower limits of the quantitative monitoring data (e.g., 0 MPa to 5 MPa), and use linear interpolation or scale translation transformation formulas to... Mapped to this space. The calculation formula is as follows: Transformation Value .when At that time, the calculated qualitative survey numerical data was 4.25. This value was then forcibly assigned a physical dimension label (e.g., MPa) consistent with the quantitative data.

[0087] Qualitative survey numerical data refers to standardized numerical values ​​with clear physical dimensions that can be directly used in matrix operations with objective measured data after the aforementioned scaling and transformation processes. This includes characteristic numerical values ​​with pressure units (such as 4.25 MPa) derived from the subjective text.

[0088] The reliability characterization value of multimodal data refers to the weighting coefficient used to quantitatively assess the signal-to-noise ratio and reliability of qualitative and quantitative data sources. This value is dynamically assessed based on the data source. For example, if the quantitative monitoring data comes from a high-precision, newly calibrated IoT sensor, its reliability characterization value will be given a high weight (e.g., ...). If the qualitative semantic data comes from oral accounts by farmers who have not received professional training, its reliability characterization value is given a low weight due to the high degree of subjective noise (e.g., ).

[0089] Aligned multimodal feature data refers to the dense feature matrix that is finally mounted on the graph network nodes after heterogeneous data has been weighted and merged using the aforementioned weights. The specific processing (weighted fusion processing) includes: the system multiplies qualitative survey numerical data (e.g., 4.25 MPa) and actual quantitative monitoring data (e.g., measured resistance value 4.10 MPa) at the same time reference point by their corresponding reliability characterization values. and Perform superposition operations (such as...) This process generates the final aligned multimodal feature data. This data, serving as the attribute vectors of feature-rich nodes, not only inherits complete information from both subjective and objective sources but also mathematically eliminates subjective cognitive noise from qualitative texts through weighted smoothing.

[0090] In some embodiments, a reverse and forward cloud generator architecture is introduced to map qualitative text into cloud feature parameters composed of expectation, entropy, and hyperentropy. The cloud model is then scaled and spatially simulated within the physical range of quantitative data. This mechanism not only quantitatively characterizes fuzzy boundaries and random discrete noise in subjective experience but also achieves adaptive weighted fusion of subjective and objective data by introducing a dual dynamic weighting system composed of "feature certainty (representing qualitative reliability)" and "measurement confidence (representing quantitative reliability)." This completely avoids the loss of nonlinear feature information caused by traditional hard boundary partitioning or fixed weight allocation, thus improving the robustness of the aligned multimodal feature data. Specifically, the semantic distribution feature parameters are soil state cloud feature parameters, which include the expected value of soil state, soil state entropy, and soil state hyperentropy.

[0091] Based on the scale space of the quantitative monitoring data, the semantic distribution feature parameters are subjected to scale normalization and numerical transformation to obtain qualitative survey numerical data, including:

[0092] Obtain the effective numerical range corresponding to the quantitative monitoring data, and perform scale normalization correction on the soil state entropy based on the ratio of the range of the effective numerical range to the expected value of the soil state, to obtain the corrected soil state entropy.

[0093] Based on the expected value of soil state, the corrected soil state entropy, and the soil state hyperentropy, multiple soil state space simulation points are generated.

[0094] The feature certainty is determined from multiple soil state space simulation points to determine the corresponding soil state fuzzy semantic set. The feature certainty reflects the probability that a soil state space simulation point belongs to the corresponding fuzzy semantic set.

[0095] Based on the degree of characteristic certainty, the soil state space simulation points are numerically transformed to obtain qualitative survey numerical data.

[0096] Based on the reliability characterization values ​​of multimodal data, a weighted fusion process is performed on the qualitative survey numerical data and the quantitative monitoring data to obtain aligned multimodal feature data, including:

[0097] The feature certainty and the measurement confidence of the quantitative monitoring data are determined as the reliability characterization values;

[0098] By using reliability characterization values, the qualitative survey numerical data and quantitative monitoring data are weighted and fused to obtain aligned multimodal feature data.

[0099] Among them, the expected value of soil condition ( ) refers to the central location point of the cloud model in the data domain space, representing the core concept that best reflects the fuzzy semantics of the soil state.

[0100] Soil state entropy ( ) refers to the measurable granularity of fuzzy concepts, representing the fuzzy span (i.e., the uncertainty of the boundary) of the corresponding fuzzy semantics of soil state.

[0101] Soil state hyperentropy ( Entropy refers to the entropy of soil state, representing an uncertainty measure of soil state entropy. It is used to quantitatively characterize the randomness and discrete noise in human experience texts when expressing that state.

[0102] Soil state cloud characteristic parameters refer to the composite cloud digital characteristics composed of the three mathematical indicators mentioned above, which are usually represented as a set of triples in the algorithm. .

[0103] The effective value range corresponding to quantitative monitoring data is the reasonable physical variation range of a certain indicator under a specific habitat in objective testing or instrument measurement. For example, the effective value range for soil compaction (penetration resistance) can be expressed as follows: .

[0104] The ratio of the effective numerical range to the expected value of the soil condition refers to the dynamic correction coefficient used to project and scale the abstract cloud parameters, which are located in the [0,1] relative domain, to the space of real physical dimensions. The calculation formula is: .

[0105] Corrected soil entropy refers to the soil entropy multiplied by the correction factor. Then, the soil state entropy is aligned with the scale of the actual physical region and has actual dimensions. The calculation process is as follows: .

[0106] Soil state space simulation points ( A virtual data sample point that satisfies a specific probability distribution is generated in physical space by the forward cloud generator algorithm.

[0107] The specific processing steps (generating multiple soil state space simulation points) include: generating... As the mean, with (or the corrected hyperentropy) is a normally distributed random number with variance. ; generate with As the mean, with The variance is a normally distributed random number, which is the single soil state space simulation point generated. Generated through a loop These are spatial simulation points for soil state.

[0108] Feature Determinism () refers to a specific soil state spatial simulation point. The mathematical probability (or membership value) of belonging to the corresponding soil state fuzzy semantic set.

[0109] The specific processing steps include: according to the formula Calculate the corresponding value for each simulation point The value is strictly within the range of [0,1].

[0110] Numerical transformation and processing and qualitative survey of numerical data This refers to using characteristic certainty as probability weights to perform expected integration over all simulated points, thereby eliminating randomness and refining unique physical quantity measurements. The specific processing includes using the centroid method or a weighted average algorithm to complete the transformation, as shown in the formula: This value is the final numerical data obtained from the qualitative survey.

[0111] Measurement confidence level of quantitative monitoring data ( The confidence score refers to a quantitative data reliability score dynamically calculated by the system based on the current operating status of the data acquisition hardware, environmental noise, or signal transmission quality. For example, when the IoT probe has sufficient battery power and no signal loss, its measurement confidence score is 0.95; when the device is nearing the end of its life or is operating in an extreme permafrost environment, the confidence score decreases to 0.40.

[0112] The reliability characterization value refers to the composite weight scalar pair assigned to the qualitative and quantitative data channels respectively in the multimodal fusion matrix. Among them, qualitative channel reliability Directly equal to the average feature determination in the aforementioned simulation process Quantitative channel reliability Directly equal to the measurement confidence level .

[0113] In some embodiments, millisecond-level precise addressing and mounting of multimodal attributes of arbitrary nodes is achieved. Furthermore, a joint update mechanism deeply couples the static features of nodes with the states of the causal evolution propagation path, enabling the generated graph network model to reflect the dynamic response of the measured state to the causal network structure in real time. This provides a highly complete topological foundation for subsequent high-precision causal intervention and counterfactual decision calculations. Specifically, the aligned multimodal feature data is fused into the corresponding nodes of the initial soil feature causal graph to obtain the target soil feature causal graph, including:

[0114] Obtain a preset node feature mapping dictionary, which records the binding relationship between node semantics and multimodal data feature dimensions;

[0115] Based on the node identifiers of the initial soil feature causal graph, candidate causal nodes are obtained by performing association matching on the aligned multimodal feature data through a node feature mapping dictionary.

[0116] By using aligned multimodal feature data, the states of candidate causal nodes and the causal evolution paths of the initial soil feature causal graph are jointly updated to obtain the target soil feature causal graph.

[0117] The preset node feature mapping dictionary refers to a key-value pair data structure that is pre-built and stored in the system memory. It is used to establish accurate mapping rules between the abstract semantics of nodes in the graph network and the dimensions of the multimodal feature vectors on the input side.

[0118] For example, in the dictionary, the key is the unique textual semantics or identifier of the node (such as [soil compaction degree]), and the value is the index interval or feature channel number of the corresponding slice dimension in the multimodal feature tensor (such as [channel 0: text cloud parameter, channel 1: sensor resistance value]), thereby guiding the system to route the corresponding data stream to the correct memory block.

[0119] A node identifier is a unique numerical code (ID) or hash string assigned to each graph node object when the initial soil feature causal graph is stored at the underlying level. This ID is used for unique addressing of the node in the memory stack. For example, it can be a string like NODE_001_pH, NODE_002_FERTILIZER, or a specific 32-bit UUID.

[0120] Candidate causal nodes refer to graph node instances that, after the system retrieves the mapping dictionary based on node identifiers, have successfully opened a data buffer and locked the feature write pointer, and are in a state of waiting for feature injection.

[0121] The core task of the association matching stage is to accurately "place" the processed multimodal feature data onto the corresponding nodes in the causal graph. The specific processing includes the following steps: Traversal retrieval: The system, like a patrol, uses pointers to scan all nodes in the graph one by one to find the node that needs to be processed (e.g., the compaction node with ID NODE_003_COMPACTION). Dictionary routing: After obtaining the node ID, the system looks up the "dictionary" (preset mapping rules) to determine which region in memory this node corresponds to, or in other words, which dimension in the feature vector it corresponds to (the first dimension). (Dimensional). Pointer locking: After finding the location, the system adds a "write lock" to that location (to prevent data conflicts) and then copies the feature data into it. After this step is completed, an "empty shell" node waiting for data becomes a candidate causal node with actual data.

[0122] The core task of the joint update processing (synchronous evolution of feature and path states) stage is to activate nodes using newly injected data and allow this change to propagate along the causal chain, ultimately generating the latest soil feature causal graph. The specific processing includes the following steps: Node state activation: The specific values ​​after fusion (such as the calculated compaction degree of 4.04 MPa) are forcibly written into the previously locked node buffer. This action will trigger a switch, changing the node state from "no signal (0)" to "data available (1)", indicating that the node has been activated. Structural bias calculation: With new data, the system will use a specific algorithm (such as a linear activation function or probability matrix) to calculate what impact this new state will have on other nodes it connects to, that is, to calculate the "physical state bias". Path dynamic evolution: This is the key to causal reasoning. The system will automatically update the transmission strength on the causal edges according to the physicochemical mechanism. For example: If the [fertilization amount] node is injected with high concentration data, the system will automatically calculate and update its impact on the downstream [soil pH value] node. Generate target topology: Repeat the above process until all nodes and edges in the graph have completed state updates and weight adjustments based on the latest data. Finally, the system encapsulates and outputs a causal graph of the target soil features that reflects the current real soil conditions in memory.

[0123] In some embodiments, the precise instantiation evolution of node states from abstract mechanisms to real observation boundaries is achieved; simultaneously, the final generated instantiated causal network is not only structurally controlled by macroscopic evolutionary laws, but also highly consistent with the actual driving intensity specific to the current sub-region to be evaluated at the quantitative level, fundamentally ensuring the local adaptability and computational accuracy of subsequent causal intervention assessments. Specifically, using aligned multimodal feature data, the states of candidate causal nodes and the causal evolution path of the initial soil feature causal map are jointly updated to obtain the target soil feature causal map, including:

[0124] Determine the numerical features and variation distribution features corresponding to candidate causal nodes from the aligned multimodal feature data;

[0125] By utilizing numerical features and variation distribution features, the initial prior distribution information contained in the candidate causal nodes is updated to obtain instantiated causal nodes. The instantiated causal nodes reflect the actual observed state of the sub-region to be evaluated under specific soil state factors.

[0126] Based on the actual observation state of each instantiated causal node, the causal evolution path in the initial soil feature causal graph is weighted and processed to obtain the updated causal directed edge. The updated causal directed edge is used to characterize the actual driving intensity of different factors affecting soil state in the current sub-region to be evaluated.

[0127] Based on the instantiated causal nodes and the updated causal directed edges, the initial soil feature causal graph is reconstructed into a graph network to obtain the target soil feature causal graph. The target soil feature causal graph is a multimodal instantiated causal network graph that fully characterizes the local soil environment evolution mechanism of the target area.

[0128] Numerical features refer to scalar values ​​extracted from aligned multimodal feature data that reflect the central trend or expected level of various soil indicators. For example, the average measured value of soil pH in the current sub-region (e.g., 5.2) or the expected value of total nitrogen content (e.g., 1.5 g / kg) obtained from the aforementioned steps.

[0129] Variation distribution characteristics are statistical measures (such as standard deviation, variance, or covariance matrix) used to characterize the degree of dispersion and fluctuation of numerical features in spatiotemporal sampling or subjective perception space. Within a one-hour sampling period, the standard deviation of pH value (e.g., σ=0.15) due to local micro-topography quantifies the spatial non-uniformity or measurement uncertainty of this feature index in the current sub-region.

[0130] Initial prior distribution information refers to the probability density function (such as preset generalized Gaussian distribution or Beta distribution parameters) assigned to each map node object before the current sub-region's measured data is injected, based on historical data or general soil science knowledge.

[0131] Specific factors affecting soil state are the soil physicochemical properties corresponding to a particular causal node that is currently being updated. For example, "the node currently being operated on: [Soil Organic Matter Content]".

[0132] Instantiated causal nodes refer to graph node instances whose probability distribution parameters completely converge to the current measured state after performing conjugate probability updates (such as Bayesian estimation updates) on the initial prior distribution information by using numerical features and variation distribution features as observational evidence.

[0133] The true observed state refers to the deterministic state vector or high-confidence core value extracted from the latest posterior probability distribution of the aforementioned instantiated causal nodes, which best represents the actual status of the current sub-region to be evaluated.

[0134] The causal evolution path in the initial soil characteristic causal graph is a topological connection structure in the initial graph network that connects different factors affecting soil state and represents the prior unidirectional causal relationship.

[0135] Weight assignment is a mathematical process that dynamically calculates the conduction coefficients on causal edges based on the actual observed states of upstream and downstream nodes, using conditional probability tables (CPT), structural equation modeling (SEM) coefficient solutions, or nonlinear mapping functions. The specific process includes: Parent-child node state extraction: The system locates a causal evolution path (e.g., [soil pH] → [heavy metal bioavailability]) and reads the actual observed state values ​​of the parent node (pH) and child node (heavy metal). Driven response bias calculation: A preset physicochemical response operator is invoked, and the current actual pH value (e.g., 5.2, indicating a slightly acidic habitat) is input. Based on the electrochemical evolution mechanism that heavy metals are more easily released under strongly acidic conditions, the positive driven amplification bias of this path in the current state is calculated. Edge weighting: The calculated amplification coefficient is assigned as a weight component to the causal directed edge, completing the weighting process.

[0136] The updated causal directed edge refers to a graph network edge object whose data structure has been injected with explicit conditional transmission coefficients or influence weight scalars after the aforementioned weight assignment process.

[0137] The actual driving strength is the absolute value of the weight carried on the updated causal directed edge. It quantitatively characterizes the extent to which the child node will produce a chain evolution response when the parent node experiences a one-unit physical fluctuation under the specific habitat state of the current sub-region.

[0138] The specific processing steps for graph network reconstruction include the following: Retrieving the original topology of the initial soil feature causal graph, locking its directed acyclic graph (DAG) connectivity skeleton, and allocating a new instantiated dynamic memory area for the graph network within the system stack. Using multi-threaded parallel processing, all instantiated causal nodes with the latest posterior probability distributions are sequentially written into this memory area according to their corresponding node numbers, overwriting any empty prior nodes lacking actual numerical support. The actual driving strength coefficients on all updated causal directed edges are synchronously compiled and written into the corresponding topological adjacency matrix of this memory area, ensuring that while the graph network possesses multimodal characteristics, the conduction impedance or amplification effect on the edges is adaptively aligned based entirely on the actual measured state of the current sub-region. After reconstruction, the system performs DAG checks and numerical convergence verification on the feature-rich and weighted graph network. Upon successful verification, the target soil feature causal graph is encapsulated and output in the system's video memory, thus completely and computably reconstructing the real environmental evolution mechanism of this local sub-region under multimodal data constraints.

[0139] Rich features refer to the fact that the nodes in the graph network are no longer empty prior nodes without actual numerical support, but instantiated causal nodes that integrate numerical features and variation distribution features.

[0140] Rich weights refer to the fact that the causal evolution paths in the graph network are no longer just connections representing prior relations, but updated causal directed edges that are given actual driving strength.

[0141] S105. Perform causal intervention processing on each node in the causal diagram of the target soil characteristics to obtain the average causal effect value corresponding to various factors affecting soil state. The average causal effect value reflects the true impact of various factors affecting soil state on soil health under the condition of cutting off the transmission path of mixed factors.

[0142] Among them, the confounding factor transmission path refers to a false, non-directly physically driven data flow association channel formed between the independent variable and the dependent variable in a graph network structure, where one or more common exogenous variables (i.e., confounding factors) simultaneously drive the independent variable (a factor affecting soil state) and the dependent variable (soil health state).

[0143] For example, in soil habitat networks, [regional meteorological precipitation] (a confounding factor) This will directly drive the soil moisture content (independent variable). Fluctuations in soil nutrient levels can simultaneously drive the loss of surface nutrients through leaching, thereby affecting soil health (dependent variable). At this point, a backdoor path will be created at the data level: [soil moisture content] ←Precipitation →Soil health status If traditional statistical methods are used, the existence of this confounding factor transmission path will incorrectly add the indirect effects of precipitation to the water content, resulting in serious statistical bias.

[0144] Causal intervention refers to the use of causal inference graph network theory. - Operator ( -calculus), at the algorithm level, forcibly blocks all backdoor unidirectional transmission paths of the target node (i.e., cuts off all directed edges pointing to the independent variable node), thereby completely decoupling the independent variable node from its original parent node constraints, and artificially assigning it a specific value and observing the mathematical simulation process of the system state response.

[0145] The specific processing steps include: when the system is ready to process a specific soil state factor (such as soil moisture content)... When intervening, the algorithm control module retrieves the target soil feature causal graph generated in the preceding steps, locates the node in computer memory, and forcibly deletes all directed edge pointers pointing to that node. This is equivalent to cutting off the [precipitation] node in the algorithm. →Soil moisture content This physical dependence makes the confounding factor... Unable to access via the backdoor path Traditional noise is measured. The characteristic probability distribution model of this node is changed to... The operator forcibly sets its state value to a fixed intervention constant. Based on the structural equations of graph networks, this intervention signal is directed along the pruned network topology (only...). A positive edge pointing to a subsequent node propagates unidirectionally downstream. Using the probability marginalization integral formula, the system's performance after removing clutter noise is calculated. The ultimate [soil health state] caused by purely physical changes Backdoor adjustment probability distribution .

[0146] The average causal effect value refers to the expected difference between the predicted soil health response values ​​when a specific soil state factor is artificially intervened to different baseline levels after completely severing the transmission paths of all confounding factors through the aforementioned causal intervention treatment. Its mathematical expression is usually defined as: ACE(X)=E[Y|do(X=x1)]-E[Y|do(X=x0)].

[0147] For example, regarding the influencing factor [chemical fertilizer application rate X], x1 represents the intervention of "scientific reduction standard application rate," and x0 represents the intervention of "traditional excessive application rate." The system calculates the expected value of the comprehensive soil health score after network-wide transmission under a pure topology that blocks interference from all confounding factors such as soil fertility and crop type. If the calculated E[Y|do(X=x1)] = 85 points and E[Y|do(X=x0)] = 65 points, then the average causal effect value corresponding to this factor is: ACE (fertilizer application rate) = 85 - 65 = 20. This value quantifies and reflects, extremely purely and objectively, the extent to which the single physical action of "optimizing fertilizer application rate" can bring about a substantial improvement (i.e., the true driving force) to the soil health status of the sub-region being evaluated, under an ideal physical intervention state that completely eliminates external environmental and management confounding interference.

[0148] In some embodiments, to isolate and eliminate statistical spurious correlation interference from uneven natural distribution, sampling bias, or environmental common factors in observational data, specifically, causal intervention processing is performed on each node in the causal graph of the target soil characteristics to obtain the average causal effect values ​​corresponding to various factors affecting soil state, including:

[0149] From the causal diagram of the target soil characteristics, determine the set of confounding factor nodes corresponding to various factors affecting soil state. The nodes in the set of confounding factor nodes are common cause nodes that simultaneously drive the factors affecting soil state and the evolution of soil health.

[0150] Using a pre-defined causal intervention quantifier, the causal transmission paths pointing to factors affecting soil state in the causal graph of the target soil characteristics are truncated to obtain a decontamination intervention graph, which is a directed graph model that eliminates the interference of contamination factor node set paths.

[0151] Based on the joint probability distribution of the confounding factor node set in the causal map of the target soil features, the deconfounding intervention map is marginalized under different variable intervention states to obtain the causal intervention prediction values ​​corresponding to different variable intervention states.

[0152] Based on the expected difference between the causal intervention prediction values ​​of soil state factors under different variable intervention states, the effects of various soil state factors are evaluated to obtain the average causal effect values ​​corresponding to each type of soil state factor.

[0153] Among them, the set of mixed factor nodes refers to the set of nodes that simultaneously serve as a common cause (common cause) of a certain "factor affecting soil state" and "soil health evolution state" in the causal diagram of the target soil characteristics.

[0154] For example, "topography (such as slope and aspect)". Topography directly determines "soil moisture content" (waterlogging in depressions, drought on slopes) and also directly affects "soil health evolution state" (nutrient loss is more likely on eroded slopes). If "topography" is not isolated as a confounding factor, algorithms will incorrectly attribute nutrient loss caused by topography entirely to the level of moisture content based solely on observational data.

[0155] Preset causal interference operators refer to mathematical operators in graph network models that are based on causal inference theories (such as Pearl's do-calculus theory) and are used to simulate the artificial application of external control to change the natural generation mechanism of variables.

[0156] For example, in natural observation, farmers' "fertilizer application rate" is randomly determined based on their own experience (affected by confounding factors such as income and habits). In contrast, causal interference quantifiers issue mandatory instructions within the algorithm (such as the mandatory command do(fertilizer application rate = standard value)) to simulate scientific experiments that forcibly change the state of this factor under identical external conditions.

[0157] A causal transmission path refers to a directed path of information and energy transfer in a causal graph of target soil characteristics, starting from an influencing factor node, following directed edges, and ultimately pointing to another node or a soil health status node. It characterizes the inherent driving chain in the physical, chemical, or biological evolution of soil (e.g., fertilizer application rate → total nitrogen content in soil → microbial activity → soil health).

[0158] Network truncation refers to the operation of using causal interference econometrics to analyze the topological structure of the causal graph of the target soil characteristics and forcibly clearing the connection weights of all directed edges pointing to the currently examined factor to zero. Graphically, this is represented by "cutting off all arrows pointing to the node," turning it into an independent exogenous variable without a parent node.

[0159] For example, when assessing the real impact of "soil moisture content" on health, we can use network truncation to separate the two edges of "rainfall → soil moisture content" and "irrigation habits → soil moisture content." In this case, the distribution of soil moisture content is no longer affected by weather and farmer behavior, achieving "controlled variables" at the digital level.

[0160] A decontamination intervention graph refers to a new topological directed graph model obtained after the target soil characteristic causal graph has been truncated. In this graph, the soil state factors under investigation have been completely isolated from the bias and interference of the observation data in the natural distribution, and their relationship with other nodes is a purely independent and controlled state.

[0161] Joint probability distribution refers to the statistical probability distribution pattern of various environmental and ecological factors (such as temperature, humidity, and organic matter baseline values) in the set of mixed factor nodes occurring simultaneously and jointly in multiple sub-regions to be evaluated within the target area. It preserves the complex background environmental and ecological characteristics of the real world.

[0162] Intervention states refer to specific controlled values ​​or state ranges that are artificially assigned to the soil state factors under investigation. They typically include intervention constants for the control group (e.g., no fertilization, no irrigation, denoted as state 0) and intervention constants for the treatment group (e.g., standard fertilization, standard irrigation, denoted as state 1) as comparisons.

[0163] Marginalization inference processing refers to the process of integrating the background by using the joint probability distribution of each confounding factor in the real world as weights, based on the deconfounded intervention map, using the law of total probability or integral methods. Its mathematical essence is to perform a weighted ergonomic calculation of a specific intervention behavior across all possible historical and environmental contexts.

[0164] The causal intervention prediction value refers to the expected value of the soil health status (e.g., the expected health prediction of the control group and the expected health prediction of the treatment group) that the system finally outputs after marginalization inference under a specific variable intervention state (e.g., fertilization or no fertilization) in a decongested intervention map.

[0165] The expected difference refers to the mathematical difference (subtraction result) between the expected health prediction of the treatment group (health after applying a certain intervention) and the expected health prediction of the control group (health without applying the intervention).

[0166] Effect assessment refers to a systematic calculation process that evaluates the influence of a soil state factor on the final direction (positive promotion or negative inhibition) and the strength of its driving force by quantifying the expected difference mentioned above.

[0167] The average causal effect value refers to the final output value of the effect assessment treatment. It eliminates all spurious correlation bubbles caused by uneven data sampling and environmental common factors, and represents the true and essential physical / chemical driving force of the change in soil health status in the target area when a controlled change of this soil state factor occurs by one unit.

[0168] Understandably, traditional soil assessment methods based on statistical correlation often only capture superficial, symbiotic relationships in the data. They are easily affected by common environmental factors such as topography, climate, and irrigation history, leading to the misjudgment of "statistical correlation" as "causal driving force" (for example, areas with high rainfall typically have high soil health; without intervention, the algorithm might mistakenly consider high water content as the sole determinant of health, thus ignoring the risk of salinization caused by excessive irrigation). This application's embodiments introduce a causal intervention quantifier to truncate the graph network, simulating a physical science experiment of "controlled variables" in the digital space. This method actively cuts off all interference paths pointing to the investigated factor, thereby completely isolating and eliminating statistical spurious correlation interference from uneven natural distribution, sampling bias, or environmental common factors in the observed data. By calculating the expected difference under different intervention states, the true causal driving strength of each soil state factor on the final soil health evolution state can be accurately and purely quantified. This not only greatly improves the accuracy of core indicator selection and avoids the omission and misjudgment of key indicators in traditional methods, but also provides a causal evidence chain with essential scientific logic for subsequent practical agricultural decisions such as precision fertilization and soil improvement.

[0169] In some embodiments, the natural distribution bias of observational data caused by environmental selection preferences, uneven sampling areas, etc., is isolated. Specifically, using a preset causal intervention quantifier, network truncation processing is performed on the causal transmission paths pointing to factors affecting soil state in the causal graph of the target soil characteristics to obtain a decontamination intervention graph, including:

[0170] The graph structure topology sequence of the target soil feature cause-effect graph is parsed to obtain the node adjacency matrix;

[0171] Using the causal intervention quantifier, the connection weights of all input directed edges pointing to soil state factors in the node adjacency matrix are cleared to zero, resulting in the updated adjacency matrix after intervention.

[0172] Based on the updated adjacency matrix after intervention, the target soil feature causal graph is restructured to obtain a decontamination intervention graph, which is a directed graph model that isolates the natural distribution bias of the observation data.

[0173] Based on the joint probability distribution of the confounding factor node set in the causal map of the target soil features, the deconfounded intervention map is marginalized under different variable intervention states to obtain the causal intervention prediction values ​​corresponding to different variable intervention states, including:

[0174] Obtain the intervention constants for the control group and the treatment group corresponding to the factors affecting soil state;

[0175] Using the intervention constants of the control group and the treatment group, the node states of the soil state factors influencing the unmixed intervention diagram were forcibly assigned values, resulting in intervention sub-plots of the control group and the treatment group, respectively.

[0176] Based on the joint probability distribution of the mixed factor node set, the intervention subgraphs of the control group and the treatment group are subjected to full probability integral inference processing to obtain the expected health prediction of the control group and the expected health prediction of the treatment group. The expected health prediction of the control group and the expected health prediction of the treatment group are used as the causal intervention prediction values.

[0177] Among them, the graph structure topological sequence refers to the topological sequence obtained by linearly arranging the nodes (factors affecting soil state) and directed edges (driving paths) in the causal graph of the target soil characteristics according to their causal order. It reflects the underlying logical chain sequence of various ecological factors from "cause" to "effect" in the process of soil evolution.

[0178] A node adjacency matrix is ​​a mathematical matrix that quantifies, in matrix form, whether there are direct pointing relationships between nodes in a causal graph of soil characteristics that influence soil state factors, and the strength of these driving paths. The rows and columns of the matrix correspond to different factors influencing soil state, and the non-zero elements in the matrix represent the connection weights of directed edges (i.e., the actual driving strength).

[0179] For example, suppose the matrix is... The row represents the "amount of fertilizer applied", the first row... The column represents "soil total nitrogen content". If a strong causal relationship exists between the two, then the (th column) in the adjacency matrix... , The connection weight corresponding to an element may be 0.85; if there is no causal link, the value of the element is 00.

[0180] Zeroing out refers to a matrix mathematical operation that uses causal interference quantifiers to forcibly change the connection weights of all in-degree edges in the node adjacency matrix that point to the currently examined "factors affecting soil state" to a value of 0.

[0181] The updated adjacency matrix refers to the new matrix obtained after being zeroed out. In this matrix, all weights that originally represented external environmental disturbances (pointed to by the parent node) in the rows or columns corresponding to the soil state factors under investigation have disappeared, becoming a pure state without input.

[0182] Graph restructuring refers to the process of remapping the updated adjacency matrix into a graph network topology, thereby constructing a completely new causal directed graph model in computer memory.

[0183] A decontamination intervention graph refers to a directed graph model generated by graph structure restructuring, in which all arrows pointing to the examined factors have been removed from the topology.

[0184] For example, in natural observation data, farmers often increase irrigation during dry seasons when rainfall is low. This natural association can lead to a significant bias in the observed irrigation data. By clearing and reorganizing the data, the directed edge between rainfall and irrigation is severed. The resulting decontamination intervention graph ensures that irrigation is no longer controlled by rainfall, becoming an independent, controlled node.

[0185] Natural distribution bias in observational data refers to the non-randomness and selective bias in the quantitative and qualitative data collected during actual farmland environmental sampling due to uneven geographical distribution, differences in farmers' personal management habits, or objective climate phenomena.

[0186] The intervention constant of the control group refers to a deterministic value that is artificially set and used as a benchmark to represent "no artificial intervention" or "maintaining the minimum baseline state".

[0187] For example, when assessing the impact of a new type of organic fertilizer on soil health, the application rate of the organic fertilizer is set to 0.0 kg / hm², and this value of 0.0 is the intervention constant for the control group.

[0188] The intervention constant for the treatment group refers to a deterministic value that is artificially set, used for comparative verification, and represents "the application of a certain standard behavior" or "the state of improvement of the target".

[0189] For example, continuing from the previous example, the application rate of the organic fertilizer is set to the standard promotion rate of 500 kg / hm², and this value of 500 is the intervention constant of the treatment group.

[0190] Forced assignment refers to directly erasing the original natural observation distribution of the node being examined and forcibly injecting the intervention constant of the control group or the intervention constant of the treatment group into the digital causal graph network.

[0191] The control group intervention subgraph is a locally branched causal directed subgraph model derived from the decontamination intervention graph by forcibly assigning the node states of the target factors affecting soil state to the "control group intervention constant".

[0192] The treatment group intervention subgraph is a locally branched causal directed subgraph model derived from the decontamination intervention graph by forcibly assigning the node states of the target factors affecting soil state to the "treatment group intervention constant".

[0193] Joint probability distribution refers to the joint statistical probability law of other mixed environmental factors (such as soil base pH, groundwater level, historical tillage times, etc.) in multiple regions within the target area, in addition to the factor under investigation.

[0194] Full probability integral inference refers to a statistical inference process that uses the "joint probability distribution" of the set of confounding factor nodes as a weighted statistical background in the intervention subgraph of the control / treatment group, and performs traversal integration (or summation) of the conditional probabilities in the graph network. Its physical meaning is to eliminate the randomness of a single background environment and examine the intervention behavior within all possible overall environmental contexts.

[0195] The expected health prediction of the control group refers to the expected value of the soil health status score of the target area when the intervention is not implemented at all, which is the output of the full probability integral inference treatment under the control group intervention subgraph.

[0196] The expected health prediction of the treatment group refers to the expected value of the soil health status score of the target area when the standard intervention behavior is forcibly implemented under the intervention subgraph of the treatment group, which is the output of the full probability integral inference treatment.

[0197] Understandably, traditional soil index analysis based on big data or machine learning typically only performs statistical correlation calculations at the matrix level, unable to intervene in the data generation mechanism. This application addresses this by resolving the graph topology into a node adjacency matrix and using causal interference quantifiers to zero out the weights of specific input directed edges. This matrix-level severing operation completely isolates the natural distribution bias of observational data caused by environmental selection preferences and uneven sampling areas in the underlying digital space of the computer. Furthermore, this application forcibly assigns values ​​to the intervention constants of the control and treatment groups, reconstructing mutually isolated intervention subgraphs. Combined with full probability integral inference, it achieves accurate deduction and prediction of expected soil health values ​​under different extreme or standard intervention strategies in a digital twin network, even when physical temperature, water, and fertilizer control experiments cannot be conducted in the target area. This significantly improves the robustness and generalization ability of graph network models in complex, variable, and high-noise real-world farmland ecological environments, laying a solid computational architecture foundation for truly achieving "physical experiment-free" reverse screening of core soil health indicators.

[0198] For example, by forcibly assigning the values ​​of "irrigation amount" to 00 (no irrigation) and 100 (standard irrigation) respectively, and after performing a full probability integral to eliminate complex rainfall and topographic interference, the expected prediction of soil health under no irrigation is calculated to be 65 points, and the expected prediction under standard irrigation is 85 points. These two are used as causal intervention prediction values, and the expected difference between the two (85-65=20 points) is the pure causal effect of irrigation amount on soil health.

[0199] S106. Based on the average causal effect value, indicators are screened from various factors affecting soil state to obtain target factors affecting soil state. These target factors are the core indicators for soil health evaluation after filtering out statistical pseudo-correlation.

[0200] Among them, statistical pseudo-correlation refers to the phenomenon in natural observation data where two soil variables show a strong trend of coordinated change or a high degree of correlation in numerical values, but there is no direct or indirect causal driving relationship between them in terms of the actual physical, chemical or biological evolutionary mechanism. It is usually caused by a third hidden confounding factor, resulting in a "data illusion".

[0201] For example, in many farmland samples, "soil surface moisture content" and "final crop yield (a component of soil health)" show a very high positive correlation. However, both are actually driven by the hidden confounding factor "rainfall." If there is prolonged and continuous rain, simply having high moisture content will not only fail to increase yield but may also lead to root rot in crops. If this statistical spurious correlation is not eliminated, the assessment system will mistakenly treat "persistently high moisture content" as a favorable core indicator of soil health.

[0202] The target soil state influencing factors refer to the set of minimal soil characteristic indicators that are ultimately retained after causal intervention and screening in this application, which have an essential causal driving force on soil health and can filter out statistical pseudo-correlation interference.

[0203] For example, when processing a large amount of initial multimodal feature data, it initially contained hundreds of potential factors affecting soil state (such as soil pH, total nitrogen, available potassium, soil surface color, looseness semantics recorded by inspectors, and surrounding vegetation cover). After causal intervention calculations to eliminate complex background interference from topography, climate, and other factors, the final calculated average causal effect value showed that only three factors—pH, total nitrogen, and available potassium—had a truly powerful driving force across noise on the evolution of soil health. These three factors were identified as the target factors affecting soil state (i.e., the core indicators for soil health assessment), while pseudo-correlated indicators such as vegetation cover (which is greatly affected by external human planting) were successfully filtered out.

[0204] In some embodiments, this approach overcomes the limitation of traditional correlation analysis where "positive and negative effects cannot be directly compared," enabling a pure reflection of the magnitude and strength of the intervention of various factors in driving soil health evolution. It also precisely extracts a simplified core indicator system composed of target soil state influencing factors that truly possess causal driving force. This scheme not only significantly reduces the sampling and calculation costs of subsequent soil health monitoring but also provides the most core and reliable data asset support for establishing an intelligent and refined soil evaluation system with inherent physical and chemical evolutionary laws. Specifically, based on the average causal effect value, indicators are screened from various soil state influencing factors to obtain target soil state influencing factors, including:

[0205] The average causal effect values ​​corresponding to various soil state factors are mapped by absolute value processing to obtain the causal effect intensity values ​​corresponding to various soil state factors. The causal effect intensity values ​​represent the pure intervention magnitude of the corresponding soil state factors on the change of soil health status.

[0206] Based on the causal effect intensity values ​​of various factors affecting soil state, the factors affecting soil state are sorted in descending order to obtain the causal driving contribution sequence. The causal driving contribution sequence reflects the order of strength of various factors affecting soil state in driving soil health evolution.

[0207] Obtain a preset causal effect cutoff threshold, which is the minimum causal effect intensity required to determine whether a factor affecting soil state has substantial intervention significance;

[0208] By using the causal effect truncation threshold, the causal driving contribution sequence is truncated and screened to obtain the target soil state factors. The target soil state factors are the set of soil state factors whose causal effect intensity value is greater than the causal effect truncation threshold.

[0209] Absolute value mapping refers to the calculation process of converting the average causal effect value with positive and negative directions into a purely scalar value through absolute value mathematical operations (such as |ACE|). Its physical significance lies in eliminating the difference in driving direction, placing "positive promoting effect" and "negative inhibiting effect" on the same dimension for equal comparison of driving power. Suppose that the average causal effect of "soil total nitrogen content" on health is +0.45 (numerical positive promoting), while the average causal effect of "soil heavy metal cadmium content" on health is -0.60 (numerical negative inhibiting). A direct comparison shows that -0.60 < +0.45, failing to objectively reflect its influence. After absolute value mapping, both are transformed into +0.45 and -0.60, indicating that in terms of the absolute destructive or constructive magnitude of the driving force, the driving power of cadmium content is even greater than that of total nitrogen content.

[0210] The causal effect strength value refers to the deterministic positive numerical value output after absolute value mapping processing. It intuitively quantifies the magnitude of pure intervention of a certain factor affecting soil state on the change of soil health.

[0211] Pure intervention amplitude refers to the purest, unadulterated span of physicochemical change in soil health when a soil factor is artificially altered by a single unit of state after excluding background noise and confounding factors from the entire environmental context in a digital causal network.

[0212] Understandably, after eliminating complex confounding factors such as "rainfall" and "irrigation history", forcibly changing the "soil total nitrogen content" by one unit in the algorithm model resulted in a pure increase of +0.45 units in the soil health score. This +0.45 is the causal effect strength value of the total nitrogen content, which represents the pure intervention magnitude of this factor.

[0213] Descending order sorting refers to a sequence organization operation that reorders various factors influencing soil state according to their corresponding causal influence strength values ​​from largest to smallest. The resulting queue is the causal driving contribution sequence, which reflects the order of strength of various factors influencing soil state in driving soil health evolution.

[0214] For example, the system calculated the causal effect strength values ​​of five factors: total nitrogen (+0.45), pH (0.72), cadmium (0.60), soil temperature (0.12), and a certain trace element (0.05). After sorting in descending order, the causal driving contribution sequence was obtained: ["pH", "cadmium", "total nitrogen", "soil temperature", "a certain trace element"]. This sequence clearly shows the actual contribution of each factor to changes in soil health and the order of their destructive power.

[0215] The preset causal effect cutoff threshold refers to the boundary value set by humans or obtained through training by historical data statistics to determine whether a certain soil state factor has actual regulatory value, that is, the minimum causal effect intensity lower limit (i.e. the pass line).

[0216] For example, in agricultural engineering, if a factor has a very weak driving effect on soil health, it is often uneconomical to artificially improve it. Therefore, a value of 0.30 is set as a preset causal effect cutoff threshold. All factors with a causal effect strength below 0.30 are considered to be weak factors lacking practical return on investment in engineering, as their influence does not reach the minimum causal effect strength limit.

[0217] The truncation screening process refers to a streamlined convergence procedure that compares the causal driver contribution sequence one by one with a preset causal effect truncation threshold, forcibly eliminating all factors in the sequence below the threshold and retaining only those above it. The retained factors are considered to have substantial intervention significance, meaning that in real-world agricultural management, artificial intervention (such as fertilization, tillage, and irrigation) can effectively and significantly alter the health evolution of the soil, possessing practical management value that aligns with economic and ecological returns.

[0218] For example, the system compares the causal driver contribution sequence ["pH value": 0.72, "Cadmium heavy metal": 0.60, "Total nitrogen": 0.45, "Soil temperature": 0.12, "A certain trace element": 0.05] with a cutoff threshold of 0.30. Through truncation screening, the system automatically removes "soil temperature" and "a certain trace element" from the list, ultimately selecting the set of target factors influencing soil state consisting of ["pH value", "Cadmium heavy metal", "Total nitrogen"]. These three indicators have substantial intervention significance for soil health in reality and become the core indicators for the final soil health assessment.

[0219] To fully verify the technical effectiveness of the method in this application, the following four comparison schemes are set up:

[0220] Option A (Traditional Correlation Screening Method): The correlation between each quantitative monitoring indicator and the comprehensive soil health score is calculated using the Pearson correlation coefficient matrix. The indicators are then sorted in descending order of the absolute value of the correlation coefficient and a threshold is set for index screening. Only quantitative monitoring data is used, and qualitative survey semantic data is not integrated.

[0221] Option B (Principal Component Analysis Screening Method): Principal component analysis (PCA) is performed on the full quantitative monitoring data to extract the principal components with a cumulative contribution rate of over 85% as the core evaluation dimensions. Based on this, key indicators are determined, without integrating qualitative survey semantic data.

[0222] Option C (Multimodal fusion + correlation screening, no causal intervention): Introducing the cloud model quantification mechanism and multimodal data alignment fusion method of this application, after fusing qualitative survey semantic data and quantitative monitoring data, correlation analysis is still used for indicator screening, without using causal intervention.

[0223] Solution D (Complete Method of This Application): The complete technical solution proposed in this application is adopted, including cloud model quantization, multimodal alignment and fusion, physical constraint causal graph construction and do-operator causal intervention inference, and index screening is performed based on the average causal effect value.

[0224] The experimental results of each scheme on 480 sub-regions to be evaluated are shown in Table 1.

[0225] Option A (Traditional Correlation) 52.0% 31.4% 9.83 5 Option B (PCA screening) 60.0% 38.7% 8.61 5 Option C (Multimodal + Correlation) 72.0% 54.2% 6.47 5 Option D (Method of this application) 92.0% 89.6% 2.31 5

[0226] Table 1 Comparison of main evaluation indicators for each scheme

[0227] In summary, the embodiments of this application improve the accuracy and reliability of soil health assessment index screening by aligning and fusing quantitative monitoring data with qualitative survey semantic data into a causal graph constrained by physical evolution, and by performing causal intervention inference based on graph networks to effectively filter out spurious correlation interference between data, which is conducive to providing accurate data support for actual intervention decisions.

[0228] To better implement the above methods, this application also provides a screening system for a soil health evaluation index system based on multimodal data, specifically integrated into an electronic device, to illustrate the method of this application embodiment in detail.

[0229] For example, such as Figure 2 As shown, the screening system for the soil health evaluation index system based on multimodal data may include a multimodal feature acquisition module 201, an evolutionary constraint rule acquisition module 202, an initial causal graph construction module 203, a multimodal causal instantiation module 204, a causal intervention evaluation module 205, and a core index screening module 206, as follows:

[0230] (I) Multimodal feature acquisition module 201.

[0231] The multimodal feature acquisition module 201 is used to acquire initial multimodal soil state feature data corresponding to multiple sub-regions to be evaluated within the target area. The initial multimodal soil state feature data includes quantitative monitoring data and qualitative survey semantic data. The quantitative monitoring data contains information characterizing the objective physical and chemical properties of the soil, and the qualitative survey semantic data contains information characterizing the subjective appearance of the soil.

[0232] (II) Evolutionary constraint rule acquisition module 202.

[0233] The evolution constraint rule acquisition module 202 is used to acquire preset soil evolution causal constraint rules. The soil evolution causal constraint rules characterize the inherent driving relationship of various factors affecting soil state in the process of soil physical and chemical evolution.

[0234] (III) Initial Cause-and-Effect Graph Construction Module 203.

[0235] The initial causal graph construction module 203 is used to construct an initial soil characteristic causal graph based on the soil evolution causal constraint rules. The nodes in the initial soil characteristic causal graph represent factors that affect soil state, and the edges represent directed causal transmission paths restricted by the soil evolution causal constraint rules.

[0236] (iv) Multimodal causal instantiation module 204.

[0237] The multimodal causal instantiation module 202 is used to perform feature mapping processing on qualitative survey semantic data and quantitative monitoring data to obtain aligned multimodal feature data. The aligned multimodal feature data is then fused into the corresponding nodes of the initial soil feature causal graph to obtain the target soil feature causal graph. The target soil feature causal graph is a causal directed graph model in which each node is assigned aligned multimodal feature data.

[0238] (v) Causal intervention assessment module 205.

[0239] The causal intervention assessment module 205 is used to perform causal intervention processing on each node in the causal diagram of the target soil characteristics to obtain the average causal effect value corresponding to various factors affecting soil state. The average causal effect value reflects the true impact of various factors affecting soil state on soil health under the condition of cutting off the transmission path of confounding factors.

[0240] (vi) Core indicator screening module 206.

[0241] The core indicator screening module 206 is used to screen indicators from various factors affecting soil state based on the average causal effect value, and obtain the target factors affecting soil state. The target factors affecting soil state are the core indicators for soil health evaluation that have been filtered out of statistical pseudo-correlation.

[0242] In practice, each of the above units can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units, please refer to the previous method embodiments, which will not be repeated here.

[0243] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0244] Therefore, embodiments of this application provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the screening methods for a soil health evaluation index system based on multimodal data provided in embodiments of this application.

[0245] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0246] Since the instructions stored in the storage medium can execute the steps in any of the methods for screening soil health evaluation index systems based on multimodal data provided in the embodiments of this application, the beneficial effects that any of the methods for screening soil health evaluation index systems based on multimodal data provided in the embodiments of this application can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.

[0247] The above provides a detailed description of the method and system for screening soil health evaluation index system based on multimodal data provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for screening an index system for soil health evaluation based on multi-modal data, characterized in that, The method includes: Acquire initial multimodal soil state characteristic data corresponding to multiple sub-regions to be evaluated within the target area. The initial multimodal soil state characteristic data includes quantitative monitoring data and qualitative survey semantic data. The quantitative monitoring data contains information characterizing the objective physical and chemical properties of the soil, and the qualitative survey semantic data contains information characterizing the subjective apparent state of the soil. Obtain preset causal constraint rules for soil evolution, which characterize the inherent driving relationships of various factors affecting soil state in the process of soil physical and chemical evolution; An initial soil characteristic causal graph is constructed based on the soil evolution causal constraint rules. The nodes in the initial soil characteristic causal graph represent the factors that affect soil state, and the edges represent the directed causal transmission paths restricted by the soil evolution causal constraint rules. The qualitative survey semantic data and the quantitative monitoring data are subjected to feature mapping processing to obtain aligned multimodal feature data. The aligned multimodal feature data is then fused into the corresponding nodes of the initial soil feature causal graph to obtain the target soil feature causal graph. The target soil feature causal graph is a causal directed graph model in which each node is assigned the aligned multimodal feature data. Causal intervention processing is performed on each node in the causal graph of the target soil characteristics to obtain the average causal effect value corresponding to various factors affecting soil state. The average causal effect value reflects the true degree of influence of various factors affecting soil state on soil health state under the condition of cutting off the transmission path of confounding factors. Based on the average causal effect value, indicators are screened from the various factors affecting soil state to obtain target factors affecting soil state. These target factors affecting soil state are core indicators for soil health evaluation that have been filtered out of statistical pseudo-correlation.

2. The method of claim 1, wherein, The step of performing feature mapping processing on the qualitative survey semantic data and the quantitative monitoring data to obtain aligned multimodal feature data includes: Based on a preset fuzzy semantic set of soil conditions, parameter extraction processing is performed on the qualitative survey semantic data to obtain semantic distribution feature parameters, which reflect the fuzzy distribution state of the qualitative survey semantic data. Based on the scale space of the quantitative monitoring data, the semantic distribution feature parameters are subjected to scale normalization and numerical conversion to obtain qualitative survey numerical data, which is data with the same dimensions as the quantitative monitoring data. Based on the reliability characterization value of multimodal data, the qualitative survey numerical data and the quantitative monitoring data are weighted and fused to obtain aligned multimodal feature data. The aligned multimodal feature data is feature data that has eliminated dimensional differences and subjective cognitive noise.

3. The method of claim 2, wherein, The semantic distribution feature parameters are soil state cloud feature parameters, which include soil state expectation, soil state entropy, and soil state hyperentropy. The step involves performing scale normalization and numerical transformation on the semantic distribution feature parameters based on the scale space of the quantitative monitoring data to obtain qualitative survey numerical data, including: Obtain the effective numerical range corresponding to the quantitative monitoring data, and perform scale normalization correction on the soil state entropy based on the ratio of the range of the effective numerical range to the expected value of the soil state, to obtain the corrected soil state entropy. Based on the expected value of soil state, the corrected soil state entropy, and the soil state hyperentropy, multiple soil state space simulation points are generated. The feature certainty degree corresponding to the soil state fuzzy semantic set is determined from the plurality of soil state space simulation points, and the feature certainty degree reflects the probability that the soil state space simulation point belongs to the corresponding fuzzy semantic set; Based on the aforementioned feature certainty, the soil state space simulation points are subjected to numerical transformation to obtain qualitative survey numerical data. The reliability characterization value based on multimodal data involves weighted fusion processing of the qualitative survey numerical data and the quantitative monitoring data to obtain the aligned multimodal feature data, including: The feature certainty and the measurement confidence of the quantitative monitoring data are determined as the reliability characterization value; Using the reliability characterization value, the qualitative survey numerical data and the quantitative monitoring data are weighted and fused to obtain the aligned multimodal feature data.

4. The method as described in claim 1, characterized in that, The step of fusing the aligned multimodal feature data into the corresponding nodes of the initial soil feature causal map to obtain the target soil feature causal map includes: Obtain a preset node feature mapping dictionary, which records the binding relationship between node semantics and multimodal data feature dimensions; Based on the node identifiers of the initial soil feature causal graph, the aligned multimodal feature data is subjected to association matching processing through the node feature mapping dictionary to obtain candidate causal nodes; Using the aligned multimodal feature data, the states of the candidate causal nodes and the causal evolution path of the initial soil feature causal map are jointly updated to obtain the target soil feature causal map.

5. The method as described in claim 4, characterized in that, The step of jointly updating the state of the candidate causal nodes and the causal evolution path of the initial soil feature causal map using the aligned multimodal feature data to obtain the target soil feature causal map includes: Determine the numerical features and variation distribution features corresponding to the candidate causal nodes from the aligned multimodal feature data; Using the numerical features and the variation distribution features, the initial prior distribution information contained in the candidate causal nodes is updated to obtain instantiated causal nodes. The instantiated causal nodes reflect the actual observed state of the sub-region to be evaluated under specific soil state factors. Based on the actual observation state of each instantiated causal node, the causal evolution path in the initial soil feature causal graph is weighted to obtain the updated causal directed edge. The updated causal directed edge is used to characterize the actual driving intensity of different factors affecting soil state in the current sub-region to be evaluated. Based on the instantiated causal nodes and the updated causal directed edges, the initial soil feature causal graph is reconstructed into a graph network to obtain the target soil feature causal graph. The target soil feature causal graph is a multimodal instantiated causal network graph that fully characterizes the local soil environment evolution mechanism of the target area.

6. The method as described in claim 1, characterized in that, The causal intervention process is applied to each node in the causal graph of the target soil characteristics to obtain the average causal effect values ​​corresponding to various factors affecting soil state, including: From the causal graph of the target soil characteristics, determine the set of mixed factor nodes corresponding to various factors affecting soil state. The nodes in the mixed factor node set are common cause nodes that simultaneously drive the factors affecting soil state and the soil health evolution state. Using a preset causal intervention metric, the causal transmission paths pointing to the soil state factors in the target soil feature causal graph are truncated to obtain a decontamination intervention graph, which is a directed graph model that eliminates the path interference of the contamination factor node set. Based on the joint probability distribution of the set of confounding factor nodes in the causal graph of the target soil features, the deconfounding intervention graph is marginalized under different variable intervention states to obtain the causal intervention prediction values ​​corresponding to the different variable intervention states. Based on the expected difference between the causal intervention prediction values ​​of the soil state factors under different variable intervention states, the effects of the various soil state factors are evaluated to obtain the average causal effect value corresponding to each type of soil state factor.

7. The method as described in claim 6, characterized in that, The process involves using a pre-defined causal intervention algorithm to perform network truncation on the causal transmission paths pointing to the soil state factors in the target soil characteristic causal graph, resulting in a decontamination intervention graph. This includes: The graph structure topology sequence of the target soil feature causal graph is parsed to obtain the node adjacency matrix; Using the causal intervention quantifier, the connection weights of all directed edges pointing to the soil state factors in the node adjacency matrix are cleared to zero, resulting in the updated adjacency matrix after intervention. Based on the updated adjacency matrix after the intervention, the target soil feature causal graph is restructured to obtain a decontamination intervention graph, which is a directed graph model that isolates the natural distribution bias of the observation data. The method involves performing marginalization inference processing on the decontamination intervention map under different variable intervention states based on the joint probability distribution of the confounding factor node set in the causal map of the target soil features, to obtain the causal intervention prediction values ​​corresponding to the different variable intervention states, including: Obtain the intervention constants of the control group and the treatment group corresponding to the factors affecting soil state; Using the intervention constants of the control group and the treatment group, the node states of the soil state factors in the decontamination intervention diagram are forcibly assigned values ​​to obtain the control group intervention sub-diagram and the treatment group intervention sub-diagram, respectively. Based on the joint probability distribution of the set of mixed factor nodes, the intervention subgraphs of the control group and the treatment group are subjected to full probability integral inference processing to obtain the expected health prediction of the control group and the expected health prediction of the treatment group. The expected health prediction of the control group and the expected health prediction of the treatment group are used as causal intervention prediction values.

8. The method as described in claim 1, characterized in that, The step of screening indicators from various soil state factors based on the average causal effect value to obtain target soil state factors includes: The average causal effect values ​​corresponding to the various soil state factors are subjected to absolute value mapping to obtain the causal effect intensity values ​​corresponding to the various soil state factors. The causal effect intensity values ​​represent the pure intervention magnitude of the corresponding soil state factors on the change of soil health status. Based on the causal effect intensity values ​​corresponding to the various factors affecting soil state, the various factors affecting soil state are sorted in descending order to obtain a causal driving contribution sequence. The causal driving contribution sequence reflects the order of strength of the various factors affecting soil state in driving the evolution of soil health. Obtain a preset causal effect cutoff threshold, which is the minimum causal effect intensity required to determine whether a factor affecting soil state has substantial intervention significance; Using the causal effect truncation threshold, the causal driving contribution sequence is truncated and screened to obtain the target soil state factors. The target soil state factors are the set of soil state factors whose causal effect intensity value is greater than the causal effect truncation threshold.

9. The method as described in claim 1, characterized in that, Before acquiring the initial multimodal soil state characteristic data corresponding to multiple sub-regions to be evaluated within the target area, the method further includes: Obtain multiple types of initial observation value sequences corresponding to each sub-region to be evaluated. The observation values ​​in the initial observation value sequences carry timestamps, and different types of initial observation value sequences correspond to different soil state observation targets. Based on the region identification information corresponding to each of the sub-regions to be evaluated and the timestamp, the initial observation value sequences of the multiple types are subjected to joint spatiotemporal alignment processing to obtain spatiotemporally normalized observation data. The spatiotemporally normalized observation data is a data set that merges different types of observation values ​​belonging to the same sub-region to be evaluated within a preset time window to a unified time reference point. By utilizing the characteristics of temporal continuity and spatial proximity, missing value detection and imputation processing are performed on the spatiotemporal normalized observation data to obtain spatiotemporal continuous observation data, which is a complete data set with inferred and imputed missing data nodes. Based on the physical reasonable range and statistical distribution characteristics of each soil condition observation target, outlier identification and removal are performed on the spatiotemporal continuous observation data to obtain quantitative monitoring data. The quantitative monitoring data is observation enhancement characterization data that has been filtered out of sampling errors and environmental abrupt changes.

10. A screening system for soil health evaluation indexes based on multimodal data, characterized in that, The system includes: The multimodal feature acquisition module is used to acquire initial multimodal soil state feature data corresponding to multiple sub-regions to be evaluated within the target area. The initial multimodal soil state feature data includes quantitative monitoring data and qualitative survey semantic data. The quantitative monitoring data contains information characterizing the objective physical and chemical properties of the soil, and the qualitative survey semantic data contains information characterizing the subjective apparent state of the soil. The evolution constraint rule acquisition module is used to acquire preset soil evolution causal constraint rules, which characterize the inherent driving relationship of various factors affecting soil state in the process of soil physical and chemical evolution. The initial causal graph construction module is used to construct an initial soil characteristic causal graph according to the soil evolution causal constraint rules. The nodes in the initial soil characteristic causal graph represent the factors affecting soil state, and the edges represent the directed causal transmission paths restricted by the soil evolution causal constraint rules. The multimodal causal instantiation module is used to perform feature mapping processing on the qualitative survey semantic data and the quantitative monitoring data to obtain aligned multimodal feature data, and to fuse the aligned multimodal feature data into the corresponding nodes of the initial soil feature causal graph to obtain the target soil feature causal graph. The target soil feature causal graph is a causal directed graph model in which each node is assigned the aligned multimodal feature data. The causal intervention assessment module is used to perform causal intervention processing on each node in the causal diagram of the target soil characteristics to obtain the average causal effect value corresponding to various factors affecting soil state. The average causal effect value reflects the true impact of various factors affecting soil state on soil health state under the condition of cutting off the transmission path of confounding factors. The core indicator screening module is used to screen indicators from the various factors affecting soil state based on the average causal effect value, and obtain target factors affecting soil state. The target factors affecting soil state are core indicators for soil health evaluation that have been filtered out of statistical pseudo-correlation.