A method for ecological data fusion and causal mining of mine development based on a dual-ancestry knowledge graph
By constructing a dual-spectrum knowledge graph and introducing a time-series graph neural network with physical constraints, we have achieved deep integration and causal mining of ecological data in mine development. This has solved the problems of difficult integration of multi-source heterogeneous data and difficulty in confirming causal relationships in mine ecological environment management, and enabled full-process tracing and precise control of issues such as surface deformation and water and soil pollution.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CENT SOUTH UNIV
- Filing Date
- 2026-03-16
- Publication Date
- 2026-07-24
AI Technical Summary
Mining enterprises face problems such as spatiotemporal mismatch of multimodal data, lack of physical mechanisms in prediction models, and lag in tracing the source of environmental problems in ecological and environmental management. Existing technologies are unable to achieve deep integration and causal mining of multi-source ecological data, resulting in low efficiency of ecological monitoring.
A method for ecological data fusion and causal mining in mining development based on a dual-spectrum knowledge graph is constructed. Through cross-spectrum mapping, spatiotemporal alignment of multimodal data, and the introduction of a time-series graph neural network with physical constraints, the dynamic correlation and prediction of production activities and ecological responses are realized, the source of disasters is identified, and control instructions are generated.
It has enabled full-process traceability and precise control of mining environmental problems, solved the problems of difficulty in integrating multi-source heterogeneous data and difficulty in confirming causal relationships, and achieved precise prediction and control of problems such as surface deformation and water and soil pollution.
Smart Images

Figure CN121859265B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the interdisciplinary field of smart mining and environmental protection, specifically a method for ecological data fusion and causal mining in mining development based on a dual-spectrum knowledge graph. Background Technology
[0002] Mineral resource development is a complex industrial process encompassing geological exploration, mining, mineral processing, and smelting. Its impact on the ecological environment is characterized by multiple sources, cumulative effects, cross-media transmission, and delayed effects. Currently, mining enterprises face the following severe challenges in ecological environment management: Multimodal data suffers from spatiotemporal mismatch, making fusion analysis difficult: The frequency of data sources in mines varies greatly. There is a time gap between production control data at the second / minute level, ground sensor data at the hour level, and satellite remote sensing data at the day / month level. This makes it impossible to correlate and analyze production processes and ecological responses on the same time axis, forming dynamic data silos.
[0003] Predictive models lack physical mechanisms, resulting in insufficient credibility and interpretability: Existing data-driven environmental prediction models are mostly black-box models, failing to incorporate the physicochemical laws of fields such as hydrogeology and rock strata movement. This makes the model prediction results susceptible to noise interference, potentially leading to conclusions that violate the law of conservation of mass or basic scientific common sense, and they cannot provide explainable causal paths leading to environmental changes, making it difficult to support precise regulatory decisions.
[0004] Environmental problems are often poorly traced, hindering early warning and accurate attribution: existing monitoring systems are mostly "post-event response" models, typically only issuing alerts after ecological damage has occurred. The lack of a dynamic correlation model spanning the entire chain of "production disturbance - environmental stress - ecological response" makes it difficult to detect early signs and quickly pinpoint specific upstream production processes causing macro-ecological anomalies, resulting in delayed and inefficient control measures.
[0005] While existing technologies have attempted to integrate information using knowledge graphs, these efforts primarily focus on static knowledge management or equipment fault diagnosis, failing to provide a systematic solution tailored to the unique needs of mine ecological management. Therefore, an intelligent method is needed that can deeply integrate multi-source ecological data from mines, embed domain-specific physical laws, and support interpretable causal mining across different stages, in order to achieve a shift from end-point monitoring to intelligent management throughout the entire process. Summary of the Invention
[0006] The purpose of this application is to provide a method for ecological data fusion and causal mining in mining development based on a dual-spectrum knowledge graph, so as to solve the technical problems faced by mining enterprises in ecological environment management in the prior art.
[0007] To achieve the above objectives, this application provides a method for mine development ecological data fusion and causal mining based on a dual-spectrum knowledge graph, the method comprising: S1: Construct a dual-spectrum domain ontology model. This dual-spectrum domain ontology model is based on preset physical coupling rules and performs cross-spectrum mapping between the production activity spectrum describing the underground mining process and the ecological response spectrum describing the evolution of surface environmental elements. S2: Obtain the preset time granularity as the baseline value, use high-frequency sensor data as a proxy variable, interpolate and generate low-frequency remote sensing data, and perform spatiotemporal alignment and standardization of multimodal data; S3: Based on the aligned data, the entity is instantiated as a graph node, and the real-time monitoring values are mapped to the dynamic attribute features of the graph node to construct a dynamic graph reflecting the real-time status of the mine. S4: Use a time-series graph neural network with physical constraints to learn the evolution of dynamic graphs and predict future ecological indicators; S5: When the predicted future ecological indicators are abnormal, construct an intention subgraph based on the abnormal nodes, and use counterfactual reasoning to calculate the contribution of each production link to the abnormality, and pinpoint the source of the disaster. S6: Based on the causal analysis results, generate production control instructions and feed them back for execution.
[0008] Preferably, S1 includes: S11: Construct the production activity genealogy ontology. The construction process includes defining a set of production entities, a set of production attributes, and a set of relationships. The set of production entities is used to characterize the key units of the process flow topology, the equipment that performs specific production actions, and the materials that flow, are consumed, and are transformed during the production process. The set of production attributes is used to characterize a set of dynamic attributes that evolve over time associated with each entity. The set of relationships is used to describe the process logic and spatial membership between entities and to form a directed graph. S12: Construct an ecological response spectrum ontology. The construction process includes defining a set of ecological entities and a set of ecological attributes. The set of ecological entities is used to characterize natural units that carry pollution or responses, monitoring indicators, and concrete data collection locations. The set of ecological attributes is used to characterize the ecological state observed at the data collection locations. S13: Define cross-spectral physical coupling rules. These rules predefine the potential force channels and intensity of production activities on the ecological environment through preset physical mechanism equations, and serve as the initial weights for cross-layer relationship edges in the knowledge graph.
[0009] Preferably, S2 includes: S21: Set the time granularity, which is a uniform time step, and the time step is determined based on the standard production shift; S22: For data with a sampling frequency greater than the reciprocal of the time granularity, the original reading sequence within a shift is aggregated into a feature vector characterizing the working conditions during that period, so that the high-frequency stream data is reduced to shift status features. S23: For data with a sampling frequency less than the reciprocal of the time granularity, an anchor-surrogate generation mechanism is adopted. Each pre-processed low-frequency ecological environment monitoring data is used as an anchor point. High-frequency surrogate variables that are physically related to the target ecological variables and have been aligned with production data are used to learn a continuous dynamic model of the ecological state, thereby generating a continuous and reasonable virtual high-frequency sequence between anchor points.
[0010] Preferably, S3 includes: S31: Define the basic storage unit of a knowledge graph as a time-weighted quadruple. This time-weighted quadruple is set with a timestamp, which should be consistent with the time granularity. The time-weighted quadruple includes a head entity, a relation type, a tail entity, and a dynamic weight. The head entity and relation type are determined based on the production activity spectrum, the tail entity is determined based on the ecological response spectrum, and the dynamic weight represents the influence strength of the head entity on the tail entity. This influence strength is determined based on the physical coupling rule. S32: For each node in the knowledge graph, at each time granularity, traverse all entities in the production activity spectrum and the ecological response spectrum, and use the data stream obtained from S2 to instantiate the node and calculate the comprehensive confidence. S33: Dynamically construct and quantify relational edges based on physical coupling rules and real-time data at each time granularity; S34: At the end of each time granularity, generate a snapshot of the graph at the current moment to perform graph evolution and lifecycle management.
[0011] Preferably, S4 includes: S41: Construct training samples and representation inputs, and determine the weighted adjacency matrix based on dynamic attribute features during node feature extraction; S42: A spatiotemporal graph neural network model with an embedded dual-spectral domain ontology model is constructed using an encoder-decoder architecture; the encoder consists of spatial aggregation and temporal evolution modules. S43: Define the physical constraint loss function of the spatiotemporal graph neural network model, which is used to characterize at least the water balance constraint term and the subsidence mechanism constraint term; train the spatiotemporal graph neural network model through the physical constraint loss function to predict future ecological indicators; S44: Connect the trained spatiotemporal graph neural network model to the real-time data stream of S2 for rolling prediction. When the predicted future ecological indicators exceed the limit and / or change abruptly, an anomaly is determined and S5 is executed.
[0012] Preferably, S5 includes: S51: Centered on the abnormal nodes of the warning, based on the physical transmission path and spatiotemporal range, automatically crop out the intention sub-graph from the global dynamic graph for deep causal analysis. S52: Based on the spatiotemporal graph neural network model and intention subgraph, the impact of virtual intervention on historical production activities on future ecological prediction results is observed, thereby quantifying the causal responsibility of each production link, performing causal contribution quantification based on counterfactual reasoning, and outputting a structured source tracing report. The source tracing report includes the target abnormal event, a list of key source nodes sorted by contribution, the contribution value of each source, and the main impact path.
[0013] Preferably, S6 includes: S61: Generate a candidate strategy set by searching a predefined strategy library based on the source of the disaster, and generate preliminary control suggestions by matching them; wherein, the rule library for matching consists of condition-action rules, and the conditions are matched with the entity type and disaster mode in the graph; S62: Based on the spatiotemporal graph neural network model, candidate strategies are pre-tested and the optimal strategy is obtained through screening. S63: Execute the optimal strategy and drive the optimization updates of S4 and S5 through result feedback.
[0014] Preferably, the physical coupling rule includes at least a stress transmission rule and a material migration rule; wherein, the stress transmission rule is used to quantify the deformation impact of underground mining on a specific location on the surface, and the material migration rule is used to quantify the migration lag and attenuation characteristics of pollutants in groundwater and determine the correlation strength.
[0015] Preferably, the comprehensive confidence calculation specifically involves assigning a comprehensive confidence score to the state vector of each node. This is used for weighting the loss function during S4 model training; where the overall confidence level is... The formula for calculation is: in: Assess the confidence level of the data source; if the data comes from a high-precision sensor, then... If the data comes from interpolation and generation of low-frequency remote sensing data, the corresponding [data] is determined based on the generation process. If data is missing and the mean is used for filling, then ; This is the time-related decay factor. For static attributes that are not updated in real time, their accuracy in describing the current state decreases over time. The corresponding calculation formula is: , This is the last update time of this attribute. The attenuation coefficient is... To represent the first The end time of each time granularity.
[0016] Preferably, the causal contribution quantification based on counterfactual reasoning specifically involves: Intention subgraph sequence The input is fed into a trained spatiotemporal graph neural network model, and a forward computation is performed to obtain the target node. The actual predicted value at the time of the warning ; For each candidate production source node in the intent subgraph, perform the following operations on the nodes belonging to the production activity spectrum: Constructing an intervention scenario, specifically: assuming nodes The node has been in a safe / downtime state for some time. Throughout the entire sequence time window Dynamic attribute values within Replace with preset counterfactual benchmark value While keeping the attributes of other nodes in the subgraph unchanged, generate a sequence of counterfactual subgraphs. Among them, the counterfactual benchmark value Use the historical statistical average of the node under long-term normal operating conditions, or set it to zero to simulate its complete shutdown; Performing counterfactual predictions specifically involves: interpreting the subgraph sequence after intervention. Input the spatiotemporal graph neural network model to obtain the counterfactual prediction value under this intervention. ; Calculating causal contribution involves: for each node that was intervened upon. Calculate its attribution score The attribution score measures the marginal contribution of the node's activity to the prediction of ecological anomalies. The formula for calculation is: Among them, molecules The denominator represents the decrease in predicted outliers after eliminating the influence of nodes. The predicted value exceeds the safety threshold. The total amount; The determination and output of the source of the disaster are as follows: Calculate the source of the disaster for all candidate production nodes. Sort in descending order and set a threshold for contribution significance. ,Will The node was identified as a key source of the disaster.
[0017] Beneficial Effects: The mining development ecological data fusion and causal mining method based on dual-spectrum knowledge graph proposed in this application achieves a complete closed loop from physical fusion of multi-source heterogeneous data to accurate prediction based on mechanism and data dual-drive, and then to scientific regulation based on causal reasoning. This is achieved by constructing a dual-spectrum ontology that includes production activities and ecological responses, building a dynamic knowledge graph that is consistent across the entire mining area in time and space, and introducing a spatiotemporal graph neural network model with physical conservation constraints. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A flowchart illustrating the method for mine development ecological data fusion and causal mining based on a dual-spectrum knowledge graph, provided in this application embodiment; Figure 2 A flowchart of S1 provided in an embodiment of this application; Figure 3 A flowchart of S2 provided in an embodiment of this application; Figure 4 A flowchart of S3 provided in an embodiment of this application; Figure 5 A flowchart of S4 provided in an embodiment of this application; Figure 6 A flowchart of S5 provided in an embodiment of this application; Figure 7 This is a flowchart of S6 provided in an embodiment of this application.
[0020] The implementation, functional features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0021] The technical solutions in the embodiments of this application will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0022] In this document, the term "comprising" is intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0023] This embodiment discloses a method for fusion and causal mining of ecological data in mine development based on a dual-spectrum knowledge graph, aiming to solve the problems of difficulty in fusion of multi-source heterogeneous data and difficulty in confirming causal relationships in mine ecological environment monitoring. By constructing a dual-spectrum ontology that includes production activities and ecological responses, a dynamic knowledge graph with spatiotemporal consistency is built for the entire mine area. A spatiotemporal graph neural network model with physical conservation constraints is introduced to achieve full-process tracing and precise control of mine environmental problems such as surface deformation and water and soil pollution.
[0024] This example uses a typical underground copper mine. The mine includes an underground mining system (stopes and haulage tunnels), a surface beneficiation plant, a tailings dam, and surrounding affected groundwater systems, soil, and vegetation areas. The main ecological problems it faces include surface subsidence caused by mining, aquifer damage, and the migration of heavy metal pollution.
[0025] Reference Figure 1 , Figure 1 A flowchart illustrating the method for ecological data fusion and causal mining of mining development based on a dual-spectrum knowledge graph, provided in this application embodiment.
[0026] like Figure 1 As shown, this embodiment discloses a method for mine development ecological data fusion and causal mining based on a dual-spectrum knowledge graph, including: S1: Construct a dual-spectrum domain ontology model. This dual-spectrum domain ontology model is based on preset physical coupling rules and performs cross-spectrum mapping between the production activity spectrum describing the underground mining process and the ecological response spectrum describing the evolution of surface environmental elements.
[0027] In summary, a dual-spectrum domain ontology model of production activities and ecological responses was constructed based on S1. The aim is to establish a computable semantic framework connecting the mining industry system and the natural ecosystem. By constructing ontology in two dimensions, production and ecology, and defining physical coupling rules, a structured knowledge skeleton is provided for subsequent fusion analysis and causal reasoning.
[0028] Reference Figure 2 , Figure 2 This is a flowchart of S1 provided in an embodiment of this application.
[0029] like Figure 2 As shown, specifically, S1 includes: S11: Construct the ontology of production activities. The construction process includes defining the set of production entities, the set of production attributes, and the set of relationships. The set of production entities is used to represent the key units of the process flow topology, the equipment that performs specific production actions, and the materials that flow, are consumed, and are transformed during the production process. The set of production attributes is used to represent a set of dynamic attributes that each entity is associated with and that evolve over time. The set of relationships is used to describe the process logic and spatial affiliation between entities and to form a directed graph.
[0030] In summary, a production activity spectrum ontology was constructed based on S11. The production activity spectrum describes the industrial processes of mineral resource development, i.e., the sources of environmental disturbance.
[0031] In a specific application of this embodiment, S11 includes: Define the set of production entities At least including: Spatial unit entity Key units that characterize the topology of the process flow include mining areas, roadways, ore dressing plants, and tailings ponds.
[0032] Equipment Entity This includes equipment that performs specific production actions, such as rock drilling rigs, loaders, drainage pumps, fans, ball mills, and flotation machines.
[0033] Material entity This includes substances that are transferred, consumed, and transformed during the production process, such as raw ore, concentrate, tailings, and mineral processing reagents.
[0034] Define production attribute set At least including: Each entity is associated with a set of dynamic attributes that evolve over time. For example, for any mining entity... Its attribute vector is defined as: in, The propulsion speed (m / d) For the mining height (m), The ore output (t / h) Recovery rate (%) This corresponds to the time.
[0035] Define relation set This includes relationships such as location, transport to, consumption, and emission, describing the technological logic and spatial affiliation between entities, forming a directed graph.
[0036] S12: Construct an ecological response spectrum ontology. The construction process includes defining a set of ecological entities and a set of ecological attributes. The set of ecological entities is used to characterize the natural units that carry pollution or response, monitoring indicators, and concrete data collection locations. The set of ecological attributes is used to characterize the ecological state observed at the data collection locations.
[0037] In summary, an ecological response spectrum ontology was constructed based on S12. The ecological response spectrum describes the ecological and environmental elements affected by mining activities and their state evolution.
[0038] In a specific application of this embodiment, S12 includes: Define the set of ecological entities At least including: Environmental media entity Natural units that bear or respond to pollution include groundwater, surface water, soil, atmosphere, and vegetation.
[0039] Pollutant / Indicator Entities This includes water level, heavy metal concentration, vegetation index (NDVI), and surface deformation. Monitoring indicators, etc.
[0040] Monitoring point entities This includes concrete data collection locations such as water level observation wells, water quality monitoring points, settlement observation stations, soil sampling points, sewage outlets, and air quality stations.
[0041] Define the set of ecological attributes At least including: For monitoring point entities Its attribute vector records the observed ecological state: ,in, These are observed values (e.g., water level elevation in m, concentration in mg / L). The rate of change; This is the national standard safety threshold.
[0042] S13: Define cross-spectral physical coupling rules. These rules predefine the potential force channels and intensity of production activities on the ecological environment through preset physical mechanism equations, and serve as the initial weights for cross-layer relationship edges in the knowledge graph.
[0043] In summary, cross-spectral physical coupling rules are defined based on S13. The potential force channels and strengths of production activities on the ecological environment are predefined through physical mechanism equations, serving as the initial weights for cross-layer relationship edges in the knowledge graph.
[0044] As a preferred embodiment of this example, the physical coupling rule includes at least a stress transmission rule and a material migration rule; wherein, the stress transmission rule is used to quantify the deformation impact of underground mining on a specific location on the surface, and the material migration rule is used to quantify the migration lag and attenuation characteristics of pollutants in groundwater and determine the correlation strength.
[0045] In the specific application of this embodiment, a probability integral method model is used to calculate the stress transmission rules in the mining area. For surface points The impact weight of sinking The calculation formula is as follows: in: The subsidence coefficient reflects the weakness of the overlying strata and is obtained by inversion from historical surface movement observation data of the mining area. For mining area Extracted volume; To deepen the burial depth of the mining area; For mining area Central projection point and surface monitoring point The horizontal distance between them; The main influence radius of the mining area, ( (The main influencing angle).
[0046] In the specific application of this embodiment, regarding the rules of material migration, for pollution sources to monitoring point The potential impact is estimated using the Green's function form of the one-dimensional convection-diffusion equation, with the following calculation formula: Source emission intensity is The weighting of its impact on downstream monitoring points for: in, As the source to monitoring point streamline distance; The average velocity of groundwater; The dispersion coefficient was obtained through on-site pumping tests and tracer tests. It represents the first-order decay coefficient of pollutants (such as adsorption / degradation). This refers to the migration time. This indicates the dispersion of pollutants over time during their migration process. ) and transport distance ( ) continuously expands and dilutes; This reflects the attenuation effect; the longer the migration time, the more pollutants are lost due to various processes, and the lower the impact weight.
[0047] S2: Obtain the preset time granularity as the baseline value, use high-frequency sensor data as a proxy variable, interpolate and generate low-frequency remote sensing data, and perform spatiotemporal alignment and standardization of multimodal data.
[0048] In summary, S2 enables spatiotemporal alignment and standardization of multimodal data. Addressing the inconsistency in data frequency between mineral development and ecological environment data (e.g., production data at the second level, water quality sampling data at the hour level, and remote sensing data at the month level), a unified dataset with spatiotemporal alignment of all elements is generated using production shifts as a unified benchmark.
[0049] Reference Figure 3 , Figure 3 This is a flowchart of S2 provided in an embodiment of this application.
[0050] like Figure 3 As shown, specifically, S2 includes: S21: Set the time granularity, which is a uniform time step, and this time step is determined based on the standard production shift.
[0051] In a specific application of this embodiment, the reference time granularity is set based on S21. In one specific application, a uniform time step is set for the entire system. For a standard production shift (e.g.) Define the global time series set as follows: ,in Representing the The end time of each shift.
[0052] S22: For data with a sampling frequency greater than the reciprocal of the time granularity, the original reading sequence within a shift is aggregated into a feature vector representing the working conditions during that period, thus reducing the dimensionality of the high-frequency stream data to shift status features.
[0053] In a specific application of this embodiment, high-frequency data dimensionality reduction and aggregation are implemented based on S22. In one specific application, for the sampling frequency... Data (such as equipment sensors, online water quality meters), and its distribution in the first... Original reading sequence within each class Aggregates into a feature vector representing the operating conditions during that period. This allows high-frequency streaming data to be reduced in dimensionality to shift status features.
[0054] in, This is the average value per shift. The standard deviation of the fluctuation within each shift. Reflects the stability of operating conditions. The slope represents the linear regression, reflecting the trend of this indicator within each shift. These are the minimum and maximum values, used to capture possible instantaneous shocks or extreme events; For the proportion exceeding the limit, Quantify the proportion of time this indicator remains in an abnormally high-risk state. These are limits preset based on equipment safety thresholds, process control limits, or environmental standards.
[0055] S23: For data with a sampling frequency less than the reciprocal of the time granularity, an anchor-surrogate generation mechanism is adopted. Each pre-processed low-frequency ecological environment monitoring data is used as an anchor point. High-frequency surrogate variables that are physically related to the target ecological variables and have been aligned with production data are used to learn a continuous dynamic model of the ecological state, thereby generating a continuous and reasonable virtual high-frequency sequence between anchor points.
[0056] In this specific application, low-frequency data virtual generation is implemented based on S23. In this specific application, low-frequency data refers to mine ecological environment monitoring data with a sampling frequency significantly lower than the production shift cycle, i.e. Specifically, this includes monthly satellite remote sensing data such as the National Distance Vibration Index (NDVI) and surface deformation, and quarterly manual sampling data such as soil heavy metal content and water sediment pollutant concentration. To address the time scale inconsistency between low-frequency and high-frequency production data, an anchor-surrogate generation mechanism is adopted. Each pre-processed low-frequency ecological environment monitoring data point serves as an anchor point. High-frequency surrogate variables, physically correlated with the target ecological variables and aligned with the production data, are used to learn a continuous dynamic model of the ecological state, thereby generating a continuous and reasonable virtual high-frequency sequence between anchor points. This mechanism is based on Neural ODEs, and the specific steps are as follows: Define low-frequency anchor point data: Low-frequency ecological environment monitoring data (denoted as...) Each valid observation is taken as an anchor point, and the time between two adjacent anchor points is denoted as . and Their observed values are respectively and ,in .
[0057] Define high-frequency proxy variables: Select a set of variables related to the target ecosystem. High-frequency data sequences related to physical mechanisms are used as proxy variables. The proxy variables for the vegetation index (NDVI) include daily average temperature, daily rainfall, and soil moisture content (IoT sensor, minute-level), while the proxy variables for surface deformation include the advance speed of the mining face, the volume of extracted ore, and the strength of the backfill.
[0058] Constructing a Neural ODE model: Assuming the target ecological state The change is a continuous dynamic process, and its rate of change is determined by the current state and high-frequency proxy variables. A neural network is used to illustrate this. The dynamics are modeled with the following parameters: The dynamic equation is expressed as: in, This is an estimate of the ecological state at a given time. This is the proxy variable vector corresponding to that moment. It is a function parameterized by a neural network, and its output is the rate of change of state.
[0059] Model Solving and Training: For each target ecological variable, a training sample set is constructed from historical data from the past 3-5 years. Anchor point of time Initially, by solving the defined ordinary differential equation through numerical integration, a path is generated from... arrive Continuous state trajectory The integration operation is implemented using the fourth-order Runge-Kutta method (dopri5) with an adaptive step size, denoted as: The model is trained by minimizing the following loss function: The first term is the anchor point matching loss, which forces the generated trajectory to precisely match the next real anchor point at its end; the second term is the smoothing regularization loss, which penalizes excessively high rates of state change to ensure that the generated trajectory is smooth and avoids non-physical, violent oscillations. To balance hyperparameters, backpropagation and optimization algorithms are used to optimize the neural network parameters by minimizing the loss function. .
[0060] Virtual value generation and output: For all shift times between any two anchor points Virtual observations at the shift level are generated through numerical integration. and for each Assign confidence to virtual values This confidence level is calculated based on the sigmoid function of the standardized error.
[0061] in, To generate the root mean square error of the model on the independent validation set, Let the standard deviation of the target variable be the value on the validation set. This is a scaling factor (usually 1-3), which controls the rate at which the confidence level decreases as the error increases.
[0062] Ultimately, the output is a continuous sequence of ecological states perfectly aligned with the production data timeline. And its virtual value confidence.
[0063] S3: Based on the aligned data, entities are instantiated as graph nodes, and real-time monitoring values are mapped to the dynamic attribute features of graph nodes to construct a dynamic graph reflecting the real-time status of the mine.
[0064] In summary, dynamic knowledge graph instantiation and evolution management were implemented based on S3. This step constructed a dynamic knowledge graph sequence that evolves over time. This provides a directly computable structured input for subsequent spatiotemporal prediction models.
[0065] Reference Figure 4 , Figure 4 This is a flowchart of S3 provided in an embodiment of this application.
[0066] like Figure 4 As shown, specifically, S3 includes: S31: Define the basic storage unit of the knowledge graph as a time-weighted quadruple. The time-weighted quadruple is set with a timestamp, which should be time-granular. The time-weighted quadruple includes a head entity, relation type, tail entity, and dynamic weight. The head entity and relation type are determined based on the production activity spectrum, the tail entity is determined based on the ecological response spectrum, and the dynamic weight represents the influence strength of the head entity on the tail entity. The influence strength is determined based on the physical coupling rule.
[0067] In summary, a graph structure definition based on time-weighted quadruples was implemented based on S31. To characterize the dynamic interaction strength between mining entities, this embodiment defines the basic storage unit of the knowledge graph as a time-weighted quadruple. ,in For head entities (such as "Mining Area No. 3"), For relation types (such as "discharge to"), For example, "surface subsidence point P5". For dynamic weights, it means right The intensity of the influence is calculated based on physical coupling rules. This is a timestamp, corresponding to the unified baseline time granularity set by S2, i.e., the production shift.
[0068] S32: For each node in the knowledge graph, at each time granularity, traverse all entities in the production activity spectrum and the ecological response spectrum, and use the data stream obtained from S2 to instantiate the node and calculate the comprehensive confidence.
[0069] In summary, node instantiation and confidence initialization are implemented based on S32. In practical applications, at each shift time... Iterate through the entity set defined in S1 and instantiate it using the data stream in S2.
[0070] For node instantiation, at each time Based on the S1 ontology, instantiate all and Nodes, data vectors aligned to S2 and The attributes assigned to the corresponding node constitute the state vector of that node. Each node's state vector is composed of static and dynamic attributes, representing the entity's condition within the overall mine dynamic system at a specific moment, rather than simply a list of its attributes. Define a node. At any moment The status is: in, Static properties (such as lithology, design mining height, and location coordinates) that represent the S1 ontology definition; This represents the dynamic attributes calculated by S2 (such as the amount of ore produced during the shift, the interpolated NDVI, and the water level value).
[0071] Specifically, the calculation of the overall confidence score for a node involves assigning an overall confidence score to the state vector of each node to differentiate the reliability of data from different sources. This is used for weighting the loss function during subsequent S4 model training; among which, the comprehensive confidence score... The formula for calculation is: in: Assess the confidence level of the data source; if the data comes from a high-precision sensor, then... If the data comes from interpolation and generation of low-frequency remote sensing data, the corresponding [data] is determined based on the generation process. If the data comes from the virtual generation of S23, then it can be used directly. As the corresponding attribute If data is missing and the mean is used for filling, then ; This is the time-related decay factor. For static attributes that are not updated in real time, their accuracy in describing the current state decreases over time. The corresponding calculation formula is: , This is the last update time of this attribute. The attenuation coefficient is... To represent the first The end time of each time granularity.
[0072] S33: Dynamically construct and quantify relational edges based on physical coupling rules and real-time data at each time granularity.
[0073] In the specific application of this embodiment, S33 is based on the physical rules of S13 and Real-time data at any given moment dynamically constructs and quantifies relational edges.
[0074] For spatial topology updates, for production entities whose spatial locations change dynamically, such as stopes and tunnel faces, updates are based on their current shift's production data (such as advance rate). Update its spatial coordinates Subsequently, its horizontal distances to all relevant ecological entities (such as surface monitoring points) were recalculated. Spatial screening is implemented only when the mining area is open. With surface monitoring points horizontal distance Less than its current mining depth The radius of influence of the decision Only then can a relationship be established or maintained between the two.
[0075] For edge weight instantiation, the physical coupling rules of S13 are utilized and Real-time parameters at each moment are used to calculate dynamic weights for each relation edge. That is, in the quadruple .
[0076] The stress transmission edge weight is calculated using the formula in S13, and is then substituted into... Time parameters: in, This represents the volume extracted during this shift. If no extraction is performed during this shift, then... The weight is set to zero.
[0077] The weights of the material migration edges are also substituted into the real-time parameters: in, As a source of pollution exist The emission intensity at any given moment is derived from production or monitoring data. If If there are no emissions upstream at a given time, the flow rate is 0, and the weight is zero.
[0078] S34: At the end of each time granularity, generate a snapshot of the graph at the current moment to perform graph evolution and lifecycle management.
[0079] In this specific application, graph evolution and lifecycle management are implemented based on S34. In this specific application, in each shift... At the end, generate a snapshot of the map at the current moment. .
[0080] Regarding node activation and dormancy: When a new node is created, if a new entity identifier (such as a newly excavated face or a newly installed sensor) appears in the S2 data stream, a new node is automatically created in the map, inheriting the class attributes of S1. When a node is dormant, for entities that have ended their operations, their status is marked as Inactive. Their historical data and relationships are available for querying for a certain retention period, after which they are moved to the archive and no longer participate in real-time calculations to improve efficiency. For example, a closed stope retains its relationship edges with the surface during the strata movement period of 3-6 months; after the stabilization period, all out-degree edges are automatically disconnected, and it becomes a historical archive node, no longer participating in real-time calculations.
[0081] This involves generating an evolutionary snapshot sequence. At the end of each shift, a package is created containing all currently active entities, valid relationships, and their node states. and the combined confidence vector Generate snapshot Snapshot The mathematical expression is: in, The set of active nodes, i.e. The set of all entity nodes that are in an active state and whose data can be used for current analysis and prediction; For the set of active relation edges, i.e. All connections at all times A time-weighted quadruple with a middle node and a weight that is not zero or exceeds the activation threshold. A set; The state vector matrix of all active nodes; This is the combined confidence vector for all active nodes. The final output is a dynamic graph sequence. , which serves as the direct input to the S4 spatiotemporal graph neural network.
[0082] S4: Utilize a time-series graph neural network with introduced physical constraints to learn the evolution of dynamic graphs and predict future ecological indicators.
[0083] In summary, spatiotemporal graph neural network prediction and anomaly detection with physical constraints were realized based on S4. This was achieved using a constructed dual-spectral dynamic graph sequence. We train a spatiotemporal graph neural network (Physics-Informed ST-GNN) that incorporates the mechanisms of the mining field to learn the complex spatiotemporal correlation between production activities and ecological responses, thereby achieving accurate prediction of future ecological conditions and early anomaly identification.
[0084] Reference Figure 5 , Figure 5 This is a flowchart of S4 provided in an embodiment of this application.
[0085] like Figure 5 As shown, specifically, S4 includes: S41: Construct training samples and representation inputs, and determine the weighted adjacency matrix based on dynamic attribute features during node feature extraction.
[0086] In the specific application of this embodiment, S41 is specifically as follows: Sliding window sampling is performed by setting the historical observation window length to... Each production shift (e.g.) (i.e., 10 days), prediction step size is (For example, the next 3 train services).
[0087] In any current shift Extracting a sequence of consecutive snapshots of the map from previous data as input features: .
[0088] In the future Within each class, key state observations of ecological phylogenetic nodes are used as training labels: in These are the key attribute values for ecological nodes in the graph.
[0089] Node feature extraction is performed, specifically: for each map snapshot in the input sequence... Extracting the node feature matrix and weighted adjacency matrix ,in The total number of nodes. The node state vector Dimensions The dynamic weight matrix calculated by S33 based on physical rules directly reflects the intensity of physical influence between entities in the current shift.
[0090] S42: A spatiotemporal graph neural network model with an embedded dual-spectral domain ontology model is constructed using an encoder-decoder architecture; the encoder consists of spatial aggregation and temporal evolution modules.
[0091] For physically guided spatial aggregation modules. At every historical moment. This module utilizes graph convolutional layers to capture the spatial impact of production nodes on ecological nodes. Unlike traditional graph convolution, this module uses physical weights calculated by S3. To transmit messages. At any given moment. ,node Aggregate its neighbor information to generate spatial features : in: For nodes The set of neighbors in the graph; for and The physical dynamic weight of the relation edge; if there is no edge, then... ; The learnable spatial transformation weight matrix, For nodes At any moment The hidden state feature vector, For bias vectors, It is a non-linear activation function.
[0092] For the time-series evolution module. Nodes. It updates its own state using its previous state and current spatial feature information. This step is implemented through a gated recurrent unit (GRU) to capture the cumulative and hysteresis effects of ecological indicators (such as the dispersion process of pollutants in the aquifer). The corresponding mathematical expression is: This module will consider the spatial impact at the current moment. Compared to the previous hidden state The simulation integrates and simulates the cumulative and delayed processes of environmental impacts (such as the migration and decay of pollutants and the continuous development of deposition).
[0093] For the prediction output layer. After processing a length of... After the input sequence, take each ecological node. In the final moments Hidden state It can directly predict the future through a multilayer perceptron (MLP). The status value of each shift.
[0094] S43: Define the physical constraint loss function of the spatiotemporal graph neural network model, which is used to characterize at least the water balance constraint term and the subsidence mechanism constraint term; train the spatiotemporal graph neural network model using the physical constraint loss function to predict future ecological indicators.
[0095] To prevent the model from outputting predictions that violate physical laws, this embodiment designs a composite loss function. Guided model training: in, The data fitting term is represented by the weighted mean square error between the predicted and observed values: in, For nodes The overall confidence level, It is a collection of ecological nodes. For nodes exist The overall confidence level of the observations at each time point For the model to nodes exist The predicted value of the state at time step. These are actual observed values.
[0096] For water balance constraints (for hydrological nodes), based on the application of Kirchhoff's laws in hydraulics, for any groundwater node... The difference between its inflow and outflow should be equal to the change in water volume caused by the change in water level: in , respectively via edge and The predicted flow rate can be estimated based on the flow velocity and cross-sectional area predicted by the model. The aquifer yield; For nodes in Predict the change in water level height over a given time period.
[0097] As a subsidence mechanism constraint (for surface subsidence nodes), the predicted subsidence increment must conform to the laws of rock mechanics, and the subsidence increment should be proportional to the output of the associated mining area.
[0098] in, For surface nodes The predicted settlement increment For rectified linear unit function, when When the value is less than 0 (predicted to be rising), this item is positive, resulting in a loss. This is the subsidence coefficient. To determine the extracted volume of the associated mining site u within the corresponding time period, Let u be the average mining depth of the mining area. Let v be the horizontal distance between the mining area and the surface point v. A function that characterizes how the effect decays with distance.
[0099] S44: Connect the trained spatiotemporal graph neural network model to the real-time data stream of S2 for rolling prediction. When the predicted future ecological indicators exceed the limit and / or change abruptly, an anomaly is determined and S5 is executed.
[0100] In this specific application, the trained model is connected to a real-time data stream for rolling prediction. When a certain ecological node... Predicted value A red alert is triggered and step S5 is initiated when any of the following conditions are met: Exceeding limits: , This is a preset safety threshold.
[0101] mutation: , This is the upper limit of the normal rate of change determined based on historical data statistics.
[0102] Output: Warning signal and corresponding target anomaly node .
[0103] S5: When the predicted future ecological indicators are abnormal, construct an intention subgraph based on the abnormal nodes, and use counterfactual reasoning to calculate the contribution of each production link to the abnormality, thereby identifying the source of the disaster.
[0104] When S4 detects an anomaly, it signifies a high probability of an ecological exceedance event in the near future, thus triggering step S5 for causal attribution. First, highly relevant local structures are extracted from the dynamic knowledge graph of mineral development to construct an intent subgraph. Then, based on counterfactual reasoning, the causal contribution of each production stage to the anomaly is quantified within this subgraph, thereby accurately pinpointing the source of the disaster.
[0105] Reference Figure 6 , Figure 6 This is a flowchart of S5 provided in an embodiment of this application.
[0106] like Figure 5 As shown, specifically, S5 includes: S51: Centered on the abnormal nodes of the warning, based on the physical transmission path and spatiotemporal range, automatically crop out the intention subgraph for deep causal analysis from the global dynamic graph.
[0107] In the specific application of this embodiment, S51 is specifically as follows: Identify anomaly anchor points: Receive early warning signals from S44, and set the ecological node it points to as the target anomaly node for this source tracing analysis, denoted as... (For example: Groundwater monitoring well_W03).
[0108] Perform reverse spacetime backtracking: Starting from the first point, a directed search is performed in the reverse direction of all relation edges in the dynamic knowledge graph. The search is based on the maximum causal time span of each relation edge. Compared with the system reference time granularity Determine the backtracking depth The calculation formula is: in, The maximum causal time span (such as the maximum transport time of pollutants or the period of stress transmission influence) can be set according to domain knowledge (such as setting it to 30 days based on the historical maximum lag event). The production shift duration is defined in S2.1.
[0109] Generate intent subgraph sequence: Extract all in The nodes that appear on the reverse path constitute the node set of the intention subgraph. This set typically contains: the target node. This includes upstream production activity nodes (such as mining sites and sewage outlets) and intermediate environmental media nodes (soil and aquifers). Simultaneously, all relational edges connecting these nodes are extracted to form the edge set of the intended subgraph. Cut out these nodes and edges in the past. A sequence of dynamic states at consecutive time steps constitutes an intentional subgraph sequence for causal analysis. .
[0110] S52: Based on the spatiotemporal graph neural network model and intention subgraph, the impact of virtual intervention on historical production activities on future ecological prediction results is observed, thereby quantifying the causal responsibility of each production link, performing causal contribution quantification based on counterfactual reasoning, and outputting a structured source tracing report. The source tracing report includes the target abnormal event, a list of key source nodes sorted by contribution, the contribution value of each source, and the main impact path.
[0111] Specifically, causal contribution quantification based on counterfactual reasoning is performed, specifically as follows: Define the fact prediction baseline: the sequence of intention subgraphs obtained in step S51 that reflects the true history. The input is fed into the S4-trained model, and a forward computation is performed to obtain the target node. The actual predicted value at the time of the warning .
[0112] Constructing counterfactual intervention: sequentially process each candidate production source node in the intent subgraph. (Right now For nodes belonging to the production activity spectrum, perform the following operations: Constructing an intervention scenario: Hypothetical nodes The node has been in a "safe / downtime" state for some time. Throughout the entire sequence time window Dynamic attribute values within Replace with preset counterfactual benchmark value While keeping the attributes of other nodes in the subgraph unchanged, generate a sequence of counterfactual subgraphs. Among them, the benchmark value Typically, the historical statistical average of the node under long-term normal operating conditions is used, or it is set to zero to simulate its complete shutdown.
[0113] Perform counterfactual prediction: Transform the subgraph sequence after intervention. Inputting the ST-GNN model yields the counterfactual predictions under this intervention. .
[0114] Calculate causal contribution: for each node that was intervened. Calculate its attribution score This score measures the marginal contribution of the node's activity to the prediction of ecological anomalies. The corresponding formula is: Among them, molecules The denominator represents the decrease in predicted outliers after eliminating the influence of nodes. This refers to the total amount of predicted values that exceed the safety threshold. This intuitively explains "if the nodes can be completely controlled". "How much percentage of the risk of exceeding the standard can be eliminated by this activity?"
[0115] Disaster source identification and output: Calculated for all candidate production nodes Sort in descending order. Set a threshold for contribution significance. (like =0.15), will The node was identified as a key source of the disaster.
[0116] The system outputs a structured source tracing report, including: the target anomaly event, a list of key source nodes sorted by contribution, the contribution value of each source, and the main impact path. For example: "Regarding the 'predicted copper ion exceedance in monitoring well W03' event, the main cause is 'leachate discharge from tailings dam No. 2' (contribution 68%), and the secondary cause is 'acidic water inrush in stope 3102' (contribution 22%)." S6: Based on the causal analysis results, generate production control instructions and feed them back for execution.
[0117] In summary, S6 enables the generation of control strategies and closed-loop feedback.
[0118] Reference Figure 7 , Figure 7 This is a flowchart of S6 provided in an embodiment of this application.
[0119] like Figure 6 As shown, specifically, S6 includes: S61: Generate a candidate strategy set by retrieving a predefined strategy library based on the source of the disaster, and generate preliminary control suggestions by matching them; wherein, the rule library for matching consists of condition-action rules, and the conditions are matched with the entity type and disaster mode in the graph.
[0120] In summary, S61 implements the generation of control strategies based on source tracing results. In the specific application of this embodiment, the source node locked in S5 is used... Based on its attribute type, a predefined strategy library is retrieved to generate a candidate strategy set. This data is used to generate preliminary control suggestions. The rule base consists of "condition-action" rules, where conditions are matched with entity types and disaster-causing modes in the graph. For example: like If the mining area is affected by subsidence, then... Reduce propulsion speed Up to 50%; Increase the strength of the filling material Up to 5 MPa.
[0121] like If the ore dressing plant is affected by heavy metal emissions, then... Adjust the dosage ratio of the medicine.
[0122] S62: Based on the spatiotemporal graph neural network model, candidate strategies are pre-tested and the optimal strategy is obtained through screening.
[0123] In the specific application of this embodiment, in order to avoid blind regulation, the pre-trained S4 spatiotemporal graph neural network model is used as a digital sandbox to quickly preview the candidate strategies.
[0124] Policy simulation: Each candidate policy... This is transformed into a counterfactual intervention on the state of the corresponding node in the graph (such as modifying its "progress speed" attribute value), input into the S4 model, and predicting the evolution trajectory of key ecological indicators after the implementation of this strategy. .
[0125] Effect Evaluation and Screening: The core criterion for evaluating the simulation effect of each strategy is whether the predicted trajectory enables the ecological indicators to return to the safe threshold in the fastest and most stable way. Within this range. The system prioritizes strategies that meet this core security objective and have minimal impact on production. As a final recommendation.
[0126] S63: Execute the optimal strategy and drive the optimization and update of S4 and S5 through result feedback.
[0127] In the specific application of this embodiment, the optimized strategy will be used. The process is pushed to the production management system for execution, and the results feedback drives the system's self-evolution.
[0128] Execution tracking and data collection: Record the actual execution parameters of the strategy and multi-source monitoring data over a subsequent period to form a real-world strategy-response case study. Record the actual ecosystem data after strategy execution. .
[0129] Experience accumulation and model fine-tuning: The entire data chain from the anomaly map to the regulation strategy to the actual ecological response is used as an enhanced sample and stored in the historical case library. This library can be used to enrich the causal rules of S5 or for manual analysis. The accumulated new samples are periodically (e.g., monthly) mixed with the old data to fine-tune the S4 model. This allows the model to gradually adapt to the slow changes in mine geological conditions and production patterns, achieving a gradual improvement in prediction and source tracing accuracy without the need for complex online parameter inversion.
[0130] Through the detailed steps of S1 to S6 described above, this embodiment constructs a dual-spectrum ontology encompassing production activities and ecological responses, builds a dynamic knowledge graph that is consistent across the entire mine in time and space, and introduces a spatiotemporal graph neural network model with physical conservation constraints. This achieves a complete closed loop from the physical fusion of multi-source heterogeneous data to accurate prediction based on both mechanism and data-driven approaches, and then to scientific regulation based on causal reasoning. It enables full-process tracing and precise regulation of mine environmental issues such as surface deformation and water and soil pollution, thereby solving the problems of difficulty in fusion of multi-source heterogeneous data and difficulty in confirming causal relationships in mine ecological environment monitoring.
[0131] In the embodiments provided in this application, it should be understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, code, or any suitable combination thereof. For hardware implementation, the processor may be implemented in one or more of the following: application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, other electronic units designed to implement the functions described herein, or combinations thereof. For software implementation, some or all of the processes of the embodiments may be performed by a computer program instructing the associated hardware. During implementation, the program may be stored in a computer-readable storage medium or transmitted as one or more instructions or code on a computer-readable storage medium. Computer-readable storage media include computer storage media and communication media, wherein communication media include any medium that facilitates the transmission of a computer program from one place to another. Storage media may be any available medium accessible to a computer. Computer-readable storage media may include, but are not limited to, RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code having the form of instructions or data structures and accessible to a computer.
[0132] Finally, it should be noted that the above description is only a preferred embodiment of this application and is not intended to limit this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for ecological data fusion and causal mining in mining development based on a dual-spectrum knowledge graph, characterized in that, The method includes: S1: Construct a dual-spectrum domain ontology model. This dual-spectrum domain ontology model is based on preset physical coupling rules to perform cross-spectrum mapping between the production activity spectrum describing the underground mining process and the ecological response spectrum describing the evolution of surface environmental elements. The physical coupling rules predefine the potential force channels and intensity of the production activities' impact on the ecological environment through preset physical mechanism equations. The physical coupling rules include at least stress transmission rules and material migration rules. The stress transmission rules are used to quantify the deformation impact of underground mining on specific locations on the surface, and the material migration rules are used to quantify the migration lag and attenuation characteristics of pollutants in groundwater and determine the correlation strength. S2: Obtain the preset time granularity as the baseline value, use high-frequency sensor data as a proxy variable, interpolate and generate low-frequency remote sensing data, and perform spatiotemporal alignment and standardization of multimodal data; S2 includes: S21: Set the time granularity, which is a uniform time step, and the time step is determined based on the standard production shift; S22: For data with a sampling frequency greater than the reciprocal of the time granularity, aggregate the original reading sequence within a shift into a feature vector characterizing the working conditions of that shift, so that the high-frequency stream data is reduced to shift status features. S23: For data with a sampling frequency less than the reciprocal of the time granularity, an anchor-surrogate generation mechanism is adopted. Each preprocessed low-frequency ecological environment monitoring data point is used as an anchor point. High-frequency surrogate variables that are physically related to the target ecological variable and aligned with the time granularity are used to learn a continuous dynamic model of the ecological state, thereby generating a continuous and reasonable virtual high-frequency sequence between anchor points; the time intervals between two adjacent anchor points are... and Their observed values are respectively and Utilizing high-frequency proxy variables that are physically related to the target ecological variables A continuous dynamic model of ecological state is learned through the constant differential equations of the gods: in: This is an estimate of the ecological state at a given time. For neural networks; continuous virtual high-frequency sequences are generated by numerical integration using the fourth-order Runge-Kutta method; S3: Based on the aligned data, the entity is instantiated as a graph node, and the real-time monitoring values are mapped to the dynamic attribute features of the graph node to construct a dynamic graph reflecting the real-time status of the mine. S4: Utilizing a time-series graphical neural network with introduced physical constraints to learn the evolutionary patterns of dynamic graphs and predict future ecological indicators; S4 includes: S41: Construct training samples and representation inputs, and determine the weighted adjacency matrix based on dynamic attribute features during node feature extraction; S42: A spatiotemporal graph neural network model with an embedded dual-spectrum domain ontology model is constructed using an encoder-decoder architecture. The encoder consists of spatial aggregation and temporal evolution modules. The spatial aggregation in the encoder includes graph convolutional layers, and the temporal evolution includes gated recurrent units. The prediction output layer of the decoder is a multilayer perceptron. S43: Define the physical constraint loss function of the spatiotemporal graph neural network model, which is used to characterize at least the water balance constraint term and the subsidence mechanism constraint term; train the spatiotemporal graph neural network model through the physical constraint loss function to predict future ecological indicators; S44: Connect the trained spatiotemporal graph neural network model to the real-time data stream of S2 for rolling prediction. When the predicted future ecological indicators exceed the limit and / or change abruptly, an anomaly is determined and S5 is executed. S5: When the predicted future ecological indicators are abnormal, construct an intention subgraph based on the abnormal nodes, and use counterfactual reasoning to calculate the contribution of each production link to the abnormality, and pinpoint the source of the disaster. S6: Based on the causal analysis results, generate production control instructions and feed them back for execution.
2. The method for mine development ecological data fusion and causal mining based on dual-spectral knowledge graphs according to claim 1, characterized in that, S1 includes: S11: Construct the production activity genealogy ontology. The construction process includes defining a set of production entities, a set of production attributes, and a set of relationships. The set of production entities is used to characterize the key units of the process flow topology, the equipment that performs specific production actions, and the materials that flow, are consumed, and are transformed during the production process. The set of production attributes is used to characterize a set of dynamic attributes that evolve over time associated with each entity. The set of relationships is used to describe the process logic and spatial membership between entities and to form a directed graph. S12: Construct an ecological response spectrum ontology. The construction process includes defining a set of ecological entities and a set of ecological attributes. The set of ecological entities is used to characterize natural units that carry pollution or responses, monitoring indicators, and concrete data collection locations. The set of ecological attributes is used to characterize the ecological state observed at the data collection locations. S13: Define the physical coupling rules across spectrums and use them as the initial weights for cross-layer relation edges in the knowledge graph.
3. The method for mine development ecological data fusion and causal mining based on dual-spectral knowledge graphs according to claim 1, characterized in that, S3 includes: S31: Define the basic storage unit of the knowledge graph as a time-weighted quadruple. The time-weighted quadruple is set with a timestamp, which corresponds to the time granularity. The time-weighted quadruple includes a head entity, a relation type, a tail entity, and a dynamic weight. The head entity and relation type are determined based on the production activity spectrum, the tail entity is determined based on the ecological response spectrum, and the dynamic weight represents the influence strength of the head entity on the tail entity. The influence strength is determined based on the physical coupling rule. S32: For each node in the knowledge graph, at each time granularity, traverse all entities in the production activity spectrum and the ecological response spectrum, and use the data stream obtained from S2 to instantiate the node and calculate the comprehensive confidence. S33: Dynamically construct and quantify relational edges based on physical coupling rules and real-time data at each time granularity; S34: At the end of each time granularity, generate a snapshot of the graph at the current moment to perform graph evolution and lifecycle management.
4. The method for mine development ecological data fusion and causal mining based on dual-spectrum knowledge graphs according to claim 1, characterized in that, S5 includes: S51: Centered on the abnormal nodes of the warning, based on the physical transmission path and spatiotemporal range, automatically crop out the intention sub-graph from the global dynamic graph for deep causal analysis. S52: Based on the spatiotemporal graph neural network model and intention subgraph, the impact of virtual intervention on historical production activities on future ecological prediction results is observed, thereby quantifying the causal responsibility of each production link, performing causal contribution quantification based on counterfactual reasoning, and outputting a structured source tracing report. The source tracing report includes the target abnormal event, a list of key source nodes sorted by contribution, the contribution value of each source, and the main impact path.
5. The method for mine development ecological data fusion and causal mining based on dual-spectral knowledge graphs according to claim 1, characterized in that, S6 includes: S61: Generate a candidate strategy set by searching a predefined strategy library based on the source of the disaster, and generate preliminary control suggestions by matching them; wherein, the rule library for matching consists of condition-action rules, and the conditions are matched with the entity type and disaster mode in the graph; S62: Based on the spatiotemporal graph neural network model, candidate strategies are pre-tested and the optimal strategy is obtained through screening. S63: Execute the optimal strategy and drive the optimization updates of S4 and S5 through result feedback.
6. The method for mine development ecological data fusion and causal mining based on dual-spectral knowledge graphs according to claim 3, characterized in that, The comprehensive confidence score calculation specifically involves assigning a comprehensive confidence score to the state vector of each node. This is used for weighting the loss function during S4 model training; where the overall confidence level is... The formula for calculation is: in: Assess the confidence level of the data source; if the data comes from a high-precision sensor, then... If the data comes from interpolation and generation of low-frequency remote sensing data, the corresponding [data] is determined based on the generation process. If data is missing and the mean is used for filling, then ; This is the time-related decay factor. For static attributes that are not updated in real time, their accuracy in describing the current state decreases over time. The corresponding calculation formula is: , This is the last update time of this attribute. The attenuation coefficient is... To represent the first The end time of each time granularity.
7. The method for mine development ecological data fusion and causal mining based on dual-spectral knowledge graphs according to claim 4, characterized in that, The aforementioned causal contribution quantification based on counterfactual reasoning specifically includes: Intention subgraph sequence The input is fed into a trained spatiotemporal graph neural network model, and a forward computation is performed to obtain the target node. The actual predicted value at the time of the warning ; For each candidate production source node in the intent subgraph, perform the following operations on the nodes belonging to the production activity spectrum: Constructing an intervention scenario, specifically: assuming nodes The node has been in a safe / downtime state for some time. Throughout the entire sequence time window Dynamic attribute values within Replace with preset counterfactual benchmark value While keeping the attributes of other nodes in the subgraph unchanged, generate a sequence of counterfactual subgraphs. Among them, the counterfactual benchmark value Use the historical statistical average of the node under long-term normal operating conditions, or set it to zero to simulate its complete shutdown; Performing counterfactual predictions specifically involves: interpreting the subgraph sequence after intervention. Input the spatiotemporal graph neural network model to obtain the counterfactual prediction value under this intervention. ; Calculating causal contribution involves: for each node that was intervened upon. Calculate its attribution score The attribution score measures the marginal contribution of the node's activity to the prediction of ecological anomalies. The formula for calculation is: Among them, molecules The denominator represents the decrease in predicted outliers after eliminating the influence of nodes. The predicted value exceeds the safety threshold. The total amount; The determination and output of the source of the disaster are as follows: Calculate the source of the disaster for all candidate production nodes. Sort in descending order and set a threshold for contribution significance. ,Will The node was identified as a key source of the disaster.