A method and system for identifying soil pollution sources and constructing a database based on a spatiotemporal knowledge graph and dynamic updates

By constructing a four-dimensional model based on spatiotemporal knowledge graphs and a dynamic update mechanism, combined with random forest and graph neural network algorithms, the problems of accuracy and data consistency in soil pollution source identification were solved, achieving efficient pollution source identification and regulatory support.

CN121434189BActive Publication Date: 2026-03-17SOUTH CHINA INST OF ENVIRONMENTAL SCI MEP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610000391.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-04
Publication Date
2026-03-17
Estimated Expiration
2046-01-04

AI Technical Summary

Technical Problem

Existing soil pollution source identification technologies are inaccurate in complex scenarios, have limited data dimensions, are outdated, and lack consistency. Existing databases cannot effectively serve the dynamic identification and precise monitoring of pollution sources.

Method used

Based on spatiotemporal knowledge graphs and dynamic updates, a four-dimensional spatiotemporal association model is constructed. Combining random forest models and graph neural network algorithms, pollution sources are identified. A dynamic update and conflict resolution mechanism is designed to build a time-series database of soil pollution sources.

Benefits of technology

It enables accurate identification and efficient updating of pollution sources, improves the scientific rigor and reliability of identification results, ensures the timeliness and consistency of data, and supports precise soil environmental supervision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121434189B_ABST
    Figure CN121434189B_ABST
Patent Text Reader

Abstract

This invention provides a method and system for soil pollution source identification and database construction based on spatiotemporal knowledge graphs and dynamic updates. Relating to the field of environmental monitoring technology, the method includes: constructing a four-dimensional spatiotemporal knowledge graph containing spatiotemporal information of pollution sources, pollutants, environmental media, and environmental factors; and combining a fusion model of random forest and graph neural networks to mine the complex relationships between pollution sources and pollutants, achieving accurate identification of potential pollution sources and quantification of pollution contributions. Furthermore, a soil pollution source time-series database is designed, introducing multi-dimensional attribute tags to structure the identification results and associated data. Dynamic data update rules are constructed using timestamps and triggers, and the database data is continuously corrected and optimized by combining priority rule chains and conflict resolution mechanisms. This invention improves the accuracy and timeliness of soil pollution source identification, providing reliable data support for precise soil environmental monitoring and pollution prevention and control decisions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of environmental monitoring technology, and in particular to a method and system for identifying and constructing a database of soil pollution sources based on spatiotemporal knowledge graphs and dynamic updates. Background Technology

[0002] In terms of soil pollution source identification technology, existing research mainly falls into two categories: one is the traditional method based on field monitoring and simulation, which obtains pollutant concentration data by setting up monitoring points and combining numerical models to invert the location and intensity of pollution sources, such as source tracing technology based on Gaussian diffusion models and HYDRUS models. This type of method relies on dense monitoring data support. In complex terrain or multi-source pollution scenarios, the deviation between model assumptions and actual working conditions can easily lead to insufficient identification accuracy and make it difficult to uncover the implicit correlation between pollution sources and pollutants. The other category is intelligent identification methods based on machine learning, such as using algorithms like support vector machines and single neural networks to process monitoring data. Although this improves data processing efficiency, it is limited by the single data dimension and does not fully integrate the temporal evolution patterns and spatial distribution characteristics. It has a weak comprehensive identification ability for historical pollution legacy and current new pollution, and there is also the problem of redundant data interference leading to biased identification results. Moreover, the formulas used are mostly traditional general formulas that have not been optimized in combination with the spatiotemporal characteristics of soil pollution data.

[0003] In terms of soil pollution data management, existing databases are mostly statically stored, with data sources primarily including interim results from pollution source surveys and routine monitoring. While these databases can archive and organize basic data, they suffer from three major shortcomings: First, the data update mechanism is lagging, relying heavily on manual, periodic data entry, making it difficult to capture dynamic pollution information such as sudden industrial emissions and cyclical changes in agricultural sources in real time, resulting in insufficient data timeliness. Second, the data correlation is low, with data often stored according to a single dimension, failing to establish spatiotemporal relationships between pollution sources, pollutants, and environmental media, thus failing to provide multi-dimensional data support for source tracing analysis. Third, a conflict resolution mechanism is lacking; when data from different sources differs, there is a lack of scientific priority determination and conflict reconciliation strategies, making it difficult to guarantee database data consistency. These problems prevent existing databases from effectively serving the dynamic identification and precise monitoring of pollution sources, hindering the implementation and application of soil pollution prevention and control technologies.

[0004] Against this backdrop, addressing the pain points of existing technologies such as inaccurate identification of potential pollution sources, limited data dimensions, delayed updates, and insufficient consistency, the development of an identification method and database product that integrates spatiotemporal features and dynamic updating capabilities has become an inevitable requirement. This invention, based on the fields of environmental monitoring and big data processing technology, deeply integrates spatiotemporal knowledge graphs with machine learning algorithms to construct a four-dimensional spatiotemporal correlation model to uncover complex pollution relationships. Simultaneously, it designs dynamic updating and conflict resolution mechanisms, forming a complete technical system of "intelligent identification - dynamic storage - precise service." This aims to solve the core shortcomings of traditional technologies and provide technical support and data assurance for precise soil environmental monitoring. Summary of the Invention

[0005] To overcome the shortcomings of existing technologies, the purpose of this invention is to provide a method and system for soil pollution source identification and database construction based on spatiotemporal knowledge graphs and dynamic updates. This invention solves the problems that existing soil pollution source identification and data management technologies generally suffer from inaccurate pollution source identification, single data dimension, insufficient spatiotemporal correlation, lagging database updates, and lack of data consistency in complex scenarios.

[0006] To achieve the above objectives, the present invention provides the following solution:

[0007] A method for identifying and constructing a database of soil pollution sources based on spatiotemporal knowledge graphs and dynamic updates includes:

[0008] Based on the needs of soil pollution prevention and control, multi-source raw data is acquired, and the multi-source raw data is preprocessed to obtain preprocessed multi-source data.

[0009] Based on the preprocessed multi-source data, with pollution sources, pollutants, environmental media and spatiotemporal as core entities, the entities and their attributes as well as the relationships between entities are organized in the form of triples to construct a four-dimensional spatiotemporal knowledge graph, and the four-dimensional spatiotemporal knowledge graph is stored in a graph database.

[0010] The entity features and relation features in the four-dimensional spatiotemporal knowledge graph are input into a fusion model composed of a random forest model and a graph neural network model to obtain the first identification probability and the second identification probability of each candidate pollution source.

[0011] Based on the dual-algorithm fusion confidence calculation formula, the final identification confidence is obtained according to the first identification probability and the second identification probability, and the target pollution source is determined according to the preset threshold. The redundant records are eliminated by combining the spatiotemporal correlation redundant data elimination formula, and the pollution contribution of each target pollution source is quantified by the pollution source contribution quantification formula.

[0012] Based on the four-dimensional spatiotemporal knowledge graph and the target pollution source, a soil pollution source time series database is designed, and multi-dimensional attribute labels are defined for pollution sources, pollutants and environmental media in the soil pollution source time series database, and the corresponding data is entered into the database according to the multi-dimensional attribute labels.

[0013] In the soil pollution source time series database, dynamic update rules and data update priority rule chains based on timestamps and triggers are set for different types of pollution sources;

[0014] Based on the dynamic update rules and data update priority rule chain of the timestamp and trigger, a conflict resolution result dataset is obtained;

[0015] Based on the conflict resolution result dataset, current status detection and consistency verification are performed on the data in the soil pollution source time series database, and corresponding update and correction operations are triggered to obtain a dynamically updated and optimized soil pollution source time series database dataset.

[0016] A system for identifying and constructing a database of soil pollution sources based on spatiotemporal knowledge graphs and dynamic updates, comprising:

[0017] The basic information module is used to acquire multi-source raw data based on the needs of soil pollution prevention and control, and to perform data preprocessing on the multi-source raw data to obtain preprocessed multi-source data.

[0018] The knowledge graph construction module is used to construct a four-dimensional spatiotemporal knowledge graph based on the preprocessed multi-source data, with pollution sources, pollutants, environmental media and spatiotemporal as core entities, and to organize the entities and their attributes and the relationships between entities in the form of triples, and to store the four-dimensional spatiotemporal knowledge graph in a graph database.

[0019] The association module is used to input the entity features and relation features in the four-dimensional spatiotemporal knowledge graph into a fusion model composed of a random forest model and a graph neural network model to obtain the first identification probability and the second identification probability of each candidate pollution source, respectively.

[0020] The identification module is used to calculate the confidence level based on the dual-algorithm fusion formula, obtain the final identification confidence level according to the first identification probability and the second identification probability, determine the target pollution source according to the preset threshold, remove redundant records by combining the spatiotemporal correlation redundant data removal formula, and quantify the pollution contribution of each target pollution source by using the pollution source contribution quantification formula.

[0021] The dynamic update module is used to design a soil pollution source time series database based on the four-dimensional spatiotemporal knowledge graph and the target pollution source, define multi-dimensional attribute labels for pollution sources, pollutants and environmental media in the soil pollution source time series database, and input the corresponding data into the database according to the multi-dimensional attribute labels.

[0022] The rule-making module is used to set dynamic update rules and data update priority rule chains based on timestamps and triggers for different types of pollution sources in the soil pollution source time-series database.

[0023] The conflict module is used to obtain a conflict resolution result dataset based on the dynamic update rules and data update priority rule chain of the timestamp and trigger.

[0024] The resolution module is used to perform currentity detection and consistency verification on the data in the soil pollution source time series database based on the conflict resolution result dataset, and trigger corresponding update and correction operations to obtain a dynamically updated and optimized soil pollution source time series database dataset.

[0025] The present invention discloses the following technical effects:

[0026] This invention provides a method and system for soil pollution source identification and database construction based on spatiotemporal knowledge graphs and dynamic updates. Regarding the accuracy of pollution source identification, this invention constructs a four-dimensional spatiotemporal knowledge graph containing entities such as pollution sources, pollutants, and environmental media based on historical pollution source tracing results. Combining random forest and graph neural network algorithms, it deeply mines the complex nonlinear mapping relationships between pollution sources and pollutants in the graph. This not only comprehensively integrates multi-dimensional spatiotemporal information, making the association between pollution sources and pollutants more contextualized and logical, but also improves the scientific validity and reliability of the identification results. Furthermore, the algorithm model's filtering capabilities effectively eliminate redundant data, achieving accurate identification of historical and current potential pollution sources. Secondly, regarding data update efficiency and timeliness, this invention achieves automatic identification, conversion, and updating of data changes by designing data dynamic update rules based on timestamps and triggers, and an intelligent change tracking mechanism. This breaks away from the traditional, lagging mode that relies on manual input and periodic maintenance. At the same time, by combining priority rule chains and rule-based conflict resolution strategies, differentiated update management is implemented for the different characteristics of industrial and agricultural sources. This enables industrial source data to respond to changes in real time and agricultural source data to be updated at reasonable intervals. Thus, while ensuring the timeliness of high-risk pollution source data, the consistency and stability of the overall data are also taken into account, significantly improving the timeliness and consistency of the database. Furthermore, regarding its practicality in supporting precise soil environmental monitoring, this invention upgrades traditional, primarily statically stored soil pollution data products into a technological system that provides continuous "intelligent services" through high-precision pollution source identification results, a soil pollution source time-series database with multi-dimensional attribute tags, and a data management mechanism with dynamic updates and conflict resolution capabilities. This enables regulatory authorities to quickly and accurately identify key pollution sources and promptly grasp the spatiotemporal evolution of pollution, providing scientific and efficient data support and decision-making basis for the formulation and adjustment of pollution prevention and control measures. Consequently, it promotes the transformation of soil environmental monitoring from extensive management to precise and refined management, significantly improving the pertinence and effectiveness of soil pollution prevention and control work. Attached Figure Description

[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0028] Figure 1 A flowchart illustrating a method for identifying and constructing a database of soil pollution sources based on spatiotemporal knowledge graphs and dynamic updates, provided for an embodiment of the present invention;

[0029] Figure 2This is a schematic diagram illustrating the performance comparison and analysis of the soil pollution source identification method in chemical industrial parks provided in this embodiment of the invention. Detailed Implementation

[0030] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0031] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0032] like Figure 1 As shown, this invention provides a method for soil pollution source identification and database construction based on spatiotemporal knowledge graphs and dynamic updates, including:

[0033] Step 100: Based on the needs of soil pollution prevention and control, acquire multi-source raw data, and perform data preprocessing on the multi-source raw data to obtain preprocessed multi-source data;

[0034] Step 200: Based on the preprocessed multi-source data, with pollution sources, pollutants, environmental media and spatiotemporal as core entities, organize the entities and their attributes and the relationships between entities in the form of triples to construct a four-dimensional spatiotemporal knowledge graph, and store the four-dimensional spatiotemporal knowledge graph in a graph database;

[0035] Step 300: Input the entity features and relation features in the four-dimensional spatiotemporal knowledge graph into the fusion model composed of the random forest model and the graph neural network model to obtain the first identification probability and the second identification probability of each candidate pollution source.

[0036] Step 400: Based on the dual-algorithm fusion confidence calculation formula, the final identification confidence is obtained according to the first identification probability and the second identification probability, and the target pollution source is determined according to the preset threshold. Redundant records are removed by combining the spatiotemporal correlation redundant data removal formula, and the pollution contribution of each target pollution source is quantified by the pollution source contribution quantification formula.

[0037] Step 500: Based on the four-dimensional spatiotemporal knowledge graph and the target pollution source, design a soil pollution source time series database, define multi-dimensional attribute labels for pollution sources, pollutants and environmental media in the soil pollution source time series database, and input the corresponding data into the database according to the multi-dimensional attribute labels;

[0038] Step 600: In the soil pollution source time series database, set dynamic update rules and data update priority rule chains based on timestamps and triggers for different types of pollution sources;

[0039] Step 700: Based on the timestamp and trigger dynamic update rules and data update priority rule chain, obtain the conflict resolution result dataset;

[0040] Step 800: Based on the conflict resolution result dataset, perform currentity detection and consistency verification on the data in the soil pollution source time series database, and trigger corresponding update and correction operations to obtain a dynamically updated and optimized soil pollution source time series database dataset.

[0041] Specifically, the recognition accuracy is high (95% accuracy): it stems from the "construction of a four-dimensional spatiotemporal knowledge graph" (comprehensive integration of spatiotemporal and entity information) and the "fusion of random forest and graph neural network algorithms" (accurately mining complex relationships). By using optimal parameters (10-year time range, 100m×100m spatial precision, 200 decision trees, and 256-dimensional hidden layers), redundant data is effectively removed, improving the reliability of recognition.

[0042] Highly timely updates (0.5-minute delay for industrial sources, 3-day cycle for agricultural sources): This is due to the "intelligent change tracking mechanism of timestamp + trigger" and "priority rule chain". It achieves differentiated and efficient updates through optimal update parameters, solving the problem of lagging updates in traditional technologies.

[0043] Fast query response (0.1-second response for a single condition): This is due to the database product's "B+ tree association index", "distributed storage", and "layered architecture" design. Optimal structural parameters (5-layer B+ tree, 5 storage nodes, and 500 records / second concurrent processing) ensure query and application efficiency.

[0044] Furthermore, the multi-source raw data includes:

[0045] Historical pollution source tracing results data, basic information data on pollution sources, pollutant data, and environmental media data.

[0046] Specifically, historical pollution source tracing data:

[0047] Sources include soil pollution survey reports from environmental protection departments over the years, pollution source tracing research results from scientific research institutions, and pollution control acceptance reports from enterprises; the usage range is 1,000 to 50,000, with an optimal usage of 10,000, which must cover source tracing cases of different regions and different pollution types over at least the past 10 years (optimal time span).

[0048] Basic information data on pollution sources:

[0049] The data sources include the National Pollution Source Census Database, enterprise discharge permit filing information from local environmental protection departments, and agricultural pollution source survey data from the Ministry of Agriculture and Rural Affairs; the usage range is 5,000 to 100,000 data points, with an optimal usage of 30,000 data points, covering mainstream pollution sources such as industrial sources (12 optimal types including chemical and metallurgical), agricultural sources (fertilizer application, livestock and poultry breeding, etc.), and domestic sources.

[0050] Pollutant data:

[0051] The sources are the priority control pollutant list in existing documents and soil pollutant detection data from environmental monitoring agencies; the usage range is 30 to 80 pollutants, and the usage of attribute data such as concentration and migration coefficient for each pollutant is no less than 500 records.

[0052] Environmental media data:

[0053] The data sources are rainfall and wind speed data from meteorological departments, groundwater runoff data from water conservancy departments, and soil type distribution data from land and resources departments; the data used covers four optimal media: soil, groundwater, atmospheric deposition, and surface runoff, with no less than 800 data points on the physicochemical properties and spatiotemporal distribution of each media.

[0054] Algorithm model parameter data:

[0055] The data sources are open-source machine learning datasets and research results on algorithm optimization in the field of soil pollution; the amount of data used must meet the training requirements of random forest and graph neural network models, with the training set accounting for 70% to 80% of the total data and the test set accounting for 20% to 30%.

[0056] Furthermore, the step of preprocessing the multi-source raw data to obtain preprocessed multi-source data includes:

[0057] The missing value identification and completion processing is performed on the multi-source raw data. A spatiotemporal weighting method based on the spatial distance of sampling points, time difference, and distribution of neighborhood data is used to estimate and fill the missing data, resulting in a multi-source dataset with missing value filling.

[0058] Based on the multi-source dataset after missing value imputation, outlier determination is performed on various types of data, and data records that deviate from the normal range are removed using statistical criteria based on mean and standard deviation, resulting in a multi-source dataset after outlier removal.

[0059] Based on the multi-source dataset after removing outliers, the pollutant-related data are standardized. Taking into account the pollutant toxicity level, data mean and fluctuation range, the numerical range of different pollutants is scaled and feature-enhanced, resulting in a multi-source dataset that has undergone pollution feature enhancement and standardization.

[0060] Based on the multi-source dataset that has undergone pollution feature enhancement and standardization, the data is normalized and, combined with the dispersion of pollutant data, the amount of effective data, and the spatial distribution characteristics of sampling points and the geometric center of the region, various types of data are adaptively mapped to a unified numerical range, resulting in preprocessed multi-source data.

[0061] Furthermore, the expression for the multi-source dataset after missing value imputation is:

[0062] ;

[0063] in, As a spatiotemporal comprehensive weight, This is the spatiotemporal weighting balance coefficient; Sampling points With neighboring sampling points Spatial Euclidean distance, The maximum effective neighborhood distance; For the sampling time difference, The maximum effective time window; Sampling points The set of K nearest neighbors. This represents the raw observation value of the j-th type of data (such as the concentration of a certain type of pollutant, the physicochemical properties of the environmental medium, etc.) corresponding to sampling point k (neighboring sampling points). tᵢ represents the sampling time of sampling point i. k This represents the sampling time of the neighboring sampling point k.

[0064] Furthermore, the formula for standardization is:

[0065] ;

[0066] in, This is the toxicity correction factor. pollutants Toxicity level, TL max This is the maximum value of the pollutant's toxicity level; , Pollutants The mean and standard deviation of the original data. This represents the original observation data of the j-th type of pollutant corresponding to sampling point i. This represents the standardized result value of the pollutant of type j corresponding to sampling point i after the pollution feature enhancement and standardization process.

[0067] Furthermore, the formula for the normalization process is:

[0068] ;

[0069] in, For spacetime correction terms; pollutants standard deviation pollutants The effective amount of data; Sampling points Distance to the geometric center of the region The average sampling interval for the region. This represents the minimum value of pollutant of type j across all sampling points. This represents the maximum value of pollutant of type j across all sampling points. This represents the standard deviation of all sampling point data for pollutant class j, used to reflect the dispersion of this type of data. This represents the final result value of the j-th type of pollutant corresponding to sampling point i after spatiotemporal adaptive normalization.

[0070] Furthermore, based on the preprocessed data, we define "pollution source", "pollutant", "environmental medium" and "spatiotemporal" as the four core entities, and construct a four-dimensional spatiotemporal knowledge graph that includes entities, attributes and relationships.

[0071] The time dimension granularity ranges from 1 day to 1 month (optimal 7 days), and the spatial dimension granularity ranges from 10m×10m to 1000m×1000m; entity relationships are defined using RDF triples; the graph storage uses the Neo4j graph database, with a node storage capacity of no less than 100,000 and a relationship storage capacity of no less than 500,000.

[0072] Furthermore, the constructed knowledge graph data is input into a random forest and graph neural network fusion model to mine complex mapping relationships between entities, identify potential sources of pollution, and remove redundant data.

[0073] Random forest parameters: number of decision trees ranges from 50 to 500 (optimal 200), feature sampling ratio ranges from 0.5 to 0.9 (optimal 0.7), minimum number of samples for node splitting ranges from 2 to 10, and the model output is the preliminary probability of pollution source identification.

[0074] Graph Neural Network Parameters: The GCN model is used, with hidden layer dimensions ranging from 64 to 512, a learning rate ranging from 0.001 to 0.05, and the number of iterations ranging from 100 to 1000. The ReLU activation function is used, with the following formula:

[0075] ;

[0076] Redundant data removal: Cosine similarity is used for determination, with a similarity threshold ranging from 0.1 to 0.3. When the similarity between two data points is higher than the threshold, the one with more complete information is retained.

[0077] The confidence calculation formula is fused using a dual-algorithm approach, replacing the output of a single algorithm.

[0078] ;

[0079] in, As a source of pollution The final identification confidence level (value from 0 to 1); , The pollution sources are respectively the outputs of random forest and graph neural network. First recognition probability and second recognition probability; As a source of pollution With Top k Characteristic similarity of high-confidence pollution sources Pollution source output by graph neural network Feature vector , , The weighting coefficients are determined through grid search optimization, and the confidence threshold is set to 0.75 (optimal value). (Optimal value 8).

[0080] Furthermore, the expression for the spatiotemporal correlation redundant data removal formula is as follows:

[0081] ;

[0082] in, Let p be the redundancy between data p and data q. The feature similarity weight coefficient; This is the spatial redundancy threshold. This is the time redundancy threshold. The feature vector corresponding to data p (a vector representation extracted by models such as graph neural networks, containing the spatiotemporal attributes and pollution characteristics of data p). This represents the feature vector corresponding to data q (and) (They have the same meaning, representing the feature vector of another set of data).

[0083] A new formula for quantifying the contribution of pollution sources has been added to quantify the contribution of each pollution source to the polluted area:

[0084] ;

[0085] in: As a source of pollution With pollutants The association weights (learned by the graph neural network); pollutants With environmental media Migration weights (calibrated based on migration coefficients); as an environmental medium Pollutants The number of times the standard was exceeded; As a source of pollution With monitoring points distance, This is the time decay coefficient.

[0086] Recognition result verification: The accuracy rate should reach 90%~95% (optimal 95%) when evaluated using a confusion matrix.

[0087] Furthermore, referring to the pollution source census classification system and combining the characteristics of pollutant migration, a time-series database architecture was designed, which includes four major modules: basic information, spatiotemporal correlation, dynamic updates, and conflict resolution.

[0088] The system adopts a distributed storage architecture with a storage node count ranging from 3 to 10, and the data synchronization delay between nodes does not exceed 1 second. The database management system combines MySQL and Redis, with MySQL used for structured data storage and Redis used for caching frequently accessed data. The cache expiration time is set to 1 to 12 hours.

[0089] Define multi-dimensional attribute labels for entities such as pollution sources, pollutants, and environmental media, and classify and store the valid data after the first stage of identification according to the labels.

[0090] Pollution source attribute labels include 20 optimal fields such as name, type, location, and emission intensity; pollutant attribute labels include 15 optimal fields such as chemical name, CAS number, toxicity level, and migration coefficient; environmental media attribute labels include 10 optimal fields such as type, physicochemical properties, and monitoring method; data is entered into the database in batches, with each batch containing 1,000 to 10,000 data entries, and the success rate of data entry must reach over 99.9%.

[0091] Furthermore, design dynamic update rules based on timestamps and triggers to establish a data intelligent change tracking mechanism;

[0092] The timestamp precision is set to milliseconds; the trigger conditions include three types: data addition, modification, and deletion, with a trigger delay of no more than 0.1 seconds; the tracking mechanism adopts a log recording method, and the log retention time range is 1 to 5 years;

[0093] Establish a data update priority rule chain for different pollution sources such as industrial and agricultural sources, and develop a rule-based conflict resolution strategy;

[0094] The priority rule chain sets 5 optimal rules, with the priority from high to low as follows: industrial source heavy metal emission data > industrial source organic pollutant emission data > agricultural source fertilizer application data > agricultural source livestock and poultry breeding data > domestic source pollution data.

[0095] Conflict resolution employs a multi-dimensional data-driven conflict resolution scoring formula:

[0096] ;

[0097] in: For data Authority For data source The weights are (environmental protection department = 0.9, research institution = 0.7, enterprise self-testing = 0.4). Historical reliability of the data source (value ranges from 0.8 to 1.0). For data Timeliness For data validity periods, industrial sources = 7 days, and agricultural sources = 90 days. For data Consistency with data from other sources; weighting coefficients , , (Determined by the Analytic Hierarchy Process), the data with the highest score is retained.

[0098] Regularly perform timeliness and consistency checks on database data to ensure data quality;

[0099] Industrial source data is monitored in real time, with a monitoring frequency of once per second; agricultural source data is monitored periodically, with a monitoring cycle of 1 to 7 days; consistency testing uses hash value verification, and the data hash value matching rate must reach more than 95%.

[0100] This embodiment also provides a soil pollution source identification and database construction system based on spatiotemporal knowledge graphs and dynamic updates, including:

[0101] The basic information module is used to acquire multi-source raw data based on the needs of soil pollution prevention and control, and to perform data preprocessing on the multi-source raw data to obtain preprocessed multi-source data.

[0102] The knowledge graph construction module is used to construct a four-dimensional spatiotemporal knowledge graph based on the preprocessed multi-source data, with pollution sources, pollutants, environmental media and spatiotemporal as core entities, and to organize the entities and their attributes and the relationships between entities in the form of triples, and to store the four-dimensional spatiotemporal knowledge graph in a graph database.

[0103] The association module is used to input the entity features and relation features in the four-dimensional spatiotemporal knowledge graph into a fusion model composed of a random forest model and a graph neural network model to obtain the first identification probability and the second identification probability of each candidate pollution source, respectively.

[0104] The identification module is used to calculate the confidence level based on the dual-algorithm fusion formula, obtain the final identification confidence level according to the first identification probability and the second identification probability, determine the target pollution source according to the preset threshold, remove redundant records by combining the spatiotemporal correlation redundant data removal formula, and quantify the pollution contribution of each target pollution source by using the pollution source contribution quantification formula.

[0105] The dynamic update module is used to design a soil pollution source time series database based on the four-dimensional spatiotemporal knowledge graph and the target pollution source, define multi-dimensional attribute labels for pollution sources, pollutants and environmental media in the soil pollution source time series database, and input the corresponding data into the database according to the multi-dimensional attribute labels.

[0106] The rule-making module is used to set dynamic update rules and data update priority rule chains based on timestamps and triggers for different types of pollution sources in the soil pollution source time-series database.

[0107] The conflict module is used to obtain a conflict resolution result dataset based on the dynamic update rules and data update priority rule chain of the timestamp and trigger.

[0108] The resolution module is used to perform currentity detection and consistency verification on the data in the soil pollution source time series database based on the conflict resolution result dataset, and trigger corresponding update and correction operations to obtain a dynamically updated and optimized soil pollution source time series database dataset.

[0109] Specifically, the technical parameters for constructing a four-dimensional spatiotemporal knowledge graph are as follows:

[0110] Build scope:

[0111] The time dimension covers the past 5-30 years; the spatial dimension is based on administrative divisions, with an accuracy range of 10m×10m~1000m×1000m; the physical dimension includes pollution sources, pollutants, and environmental media (3-6 types such as soil, groundwater, and atmospheric deposition, focusing on media directly related to soil pollution).

[0112] Data source:

[0113] Historical pollution source tracing reports, environmental monitoring data, pollution source census data, etc., with a data volume ranging from 100,000 to 10 million records (the minimum effective data volume to meet the requirements of algorithm training and graph construction).

[0114] Parameters of the pollution source-pollutant relationship mining algorithm:

[0115] Random Forest Algorithm:

[0116] The number of decision trees ranges from 50 to 500; the feature sampling ratio ranges from 0.5 to 0.9; and the minimum number of samples for node splitting ranges from 2 to 10.

[0117] Graph Neural Network Algorithm:

[0118] The hidden layer dimension ranges from 64 to 512; the learning rate ranges from 0.001 to 0.05; the number of iterations ranges from 100 to 1000; and the redundancy data removal threshold ranges from 0.1 to 0.3.

[0119] This product is an intelligent database that integrates spatiotemporal knowledge graphs and dynamic update mechanisms. Its composition, structure, physical properties, and implementation scope are as follows:

[0120] Product composition and usage (data level):

[0121] Core components:

[0122] The system comprises a basic information module, a spatiotemporal association module, a dynamic update module, and a conflict resolution module. The data proportions of each module range from 20%-30%, 30%-40%, 25%-35%, and 5%-10%, respectively. The spatiotemporal association module includes a knowledge graph construction module, an association module, and an identification module. The dynamic update module includes a dynamic update module and a rule formulation module. The conflict resolution module includes a conflict module and a resolution module.

[0123] Basic Information Module:

[0124] It includes basic information on pollution sources (name, location, type, etc., 15-25 fields), basic information on pollutants (chemical name, CAS number, toxicity level, etc., 12-18 fields), and basic information on environmental media (type, properties, monitoring methods, etc., 8-12 fields).

[0125] Spatiotemporal correlation module:

[0126] The system stores four-dimensional spatiotemporal knowledge graph data, with the proportion of relational data to 60%-80% of the module's total data volume, ensuring the integrity of spatiotemporal relationships.

[0127] Dynamic update module:

[0128] It includes an updated rule base (20-50 rules) and intelligent tracking scripts (5-12 scripts) to support automatic data updates.

[0129] Conflict resolution module:

[0130] It includes priority rule chains (3-8 rule chains) and conflict determination model.

[0131] Logical structure:

[0132] It adopts a three-tier architecture: a data storage layer (using a distributed storage architecture with 3-10 nodes to achieve load balancing), a data processing layer (including data transformation, cleaning, and association engines with a concurrent processing capacity of 100-1000 records / second), and an application service layer (providing data query, statistics, and alert interfaces with an interface response time of 0.1-1 second).

[0133] Entity association structure:

[0134] Using the pollution source ID as the core index, a many-to-many association structure is established for "pollution source-pollutant", "pollution source-environmental medium", and "pollutant-environmental medium". The association index adopts a B+ tree structure (tree depth 3-8 levels to improve query efficiency).

[0135] Product physical properties (performance indicators):

[0136] Data storage performance:

[0137] Single-node storage capacity: 100GB-2TB.

[0138] Data update performance:

[0139] Industrial data updates are delayed by 0.1-1 minute; agricultural data updates are 1-7 days in cycle.

[0140] Data query performance:

[0141] Single-condition query response time: 0.05-0.5 seconds; multi-condition combined query response time: 0.1-1 seconds.

[0142] Data accuracy:

[0143] The accuracy rate of pollution source identification is 90%-99%; the accuracy rate of data update is 90%-95%.

[0144] Furthermore, this invention provides an implementation case of soil pollution source identification and database construction in a chemical industrial park;

[0145] The chemical industrial park covers an area of ​​approximately 10 km². 2 The survey covers 86 industrial enterprises across 12 categories, including chemical, pharmaceutical, and dyeing industries, which suffer from combined pollution from heavy metals and organic pollutants. Traditional identification methods can only locate about 70% of the pollution sources, and data updates rely on manual intervention, resulting in a lag of over 15 days. This implementation adopts the method of this invention, aiming to achieve a pollution source identification accuracy rate of ≥90% and an industrial source data update delay of ≤1 minute.

[0146] Historical pollution source tracing data: 10,000 pollution investigation reports and enterprise governance acceptance reports collected from the park from 2014 to 2023 (optimal usage), covering typical pollution types such as heavy metals and VOCs.

[0147] Basic information data on pollution sources: 30,000 data points were obtained from local pollution source census, discharge permit and environmental statistics databases, including information on 12 types of industrial sources such as enterprise discharge permits and production processes.

[0148] Pollutant data: 45 basic items and 15 common pollutants were selected from existing documents, with 1,000 data points for each pollutant's concentration, migration coefficient and other attributes.

[0149] Environmental media data: 1,500 data points each of the physical and chemical properties and spatiotemporal distribution of soil and groundwater were collected, sourced from monitoring data from meteorological and water conservancy departments.

[0150] Algorithm model parameter data: divided into 75% training set and 25% test set to meet the training requirements of random forest and graph neural network models.

[0151] Missing value handling: A spatiotemporally weighted missing value imputation formula is used, with α=0.5 and D... max =3000m, T max =730 days, k=10 (optimal value), fill in missing fields such as Cr concentration and benzene concentration.

[0152] Outlier removal: Using the 3σ criterion, outlier data such as groundwater pH and COD concentration that deviated from the mean by 3 times the standard deviation were removed, resulting in the removal of 7 invalid data entries.

[0153] Dimension definition: The time dimension granularity is set to 7 days, the spatial dimension precision is 100m×100m (optimal value), and the entity dimension includes two types of environmental media: pollution source, pollutant, and soil / groundwater.

[0154] Graph storage: The Neo4j graph database is used to construct 300,000 entity nodes and 1 million relations, and the relationships such as "pollution source-emission-pollutant" and "pollutant-migration-environmental medium" are defined in the form of RDF triples.

[0155] Random forest parameter settings: number of decision trees 200 (optimal value), feature sampling ratio 0.7, minimum number of samples for node splitting 5, output preliminary probability of pollution source identification.

[0156] Graph Neural Network Parameter Settings: GCN model is used, hidden layer dimension is 256 (optimal value), learning rate is 0.01, iterations are 500, and activation function is ReLU.

[0157] Dual-algorithm fusion confidence calculation: Set ω1=0.4, ω2=0.4, ω3=0.2, k=8, calculate the final identification confidence of each pollution source, and set the confidence threshold to 0.75.

[0158] Redundant data removal: A spatiotemporal correlation redundant data removal formula was adopted, with γ=0.6, D0=300m, T0=60 days, and θ=0.25. Redundant data with similarity higher than the threshold were removed, and 98 valid data were retained in the end.

[0159] Results verification: Through confusion matrix evaluation, the accuracy of pollution source identification reached 95% (optimal value), which is 23 percentage points higher than the traditional method.

[0160] Storage architecture: A distributed storage architecture with 5 nodes (optimal value) is adopted, with MySQL storing structured data and Redis caching frequently accessed data (cache expiration time of 6 hours), and data synchronization latency between nodes is 0.5 seconds.

[0161] Module division: Basic information module (25%), spatiotemporal correlation module (35%), dynamic update module (30%), conflict resolution module (10%).

[0162] Attribute tag definition: Set 20 attribute fields such as name, location, and emission intensity for pollution sources, 15 attribute fields such as CAS number and toxicity level for pollutants, and 10 attribute fields such as monitoring method for environmental media.

[0163] Batch insertion: Insert 5,000 data entries at a time, for a total of 860,000 data entries, with a success rate of 99.95%.

[0164] Update mechanism: Set millisecond-level timestamps, trigger delay of 0.05 seconds, and logs are saved for 3 years; industrial sources are updated in real time, while agricultural sources (supporting planting areas within the park) are updated every 3 days.

[0165] Conflict resolution: The priority rule chain is set as "industrial heavy metals > industrial organic matter > agricultural fertilizers > agricultural livestock and poultry breeding > domestic sources", with an authority weight of 0.7 and a time weight of 0.3, and the conflict resolution accuracy rate is 99.5%.

[0166] Pollution source identification: A total of 23 potential pollution sources were identified, including 17 chemical enterprises, 4 solid waste storage sites, and 2 sewage pipeline leakage points, with a 95% match rate with the actual pollution investigation results in the park.

[0167] Database performance: Single-condition query response time is 0.1 seconds, and multi-condition combined query response time is 0.3 seconds; the average update latency of industrial source data is 0.3 minutes, and the data timeliness is improved by 82% compared with traditional static databases.

[0168] Application results: Through real-time database updates, a single illegal chemical waste discharge point can be quickly located, providing precise data support for pollution control in the industrial park and improving pollution control response efficiency by 60%.

[0169] Specifically, such as Figure 2 As shown, the experiment identified 23 potential pollution sources, including 17 chemical enterprises, 4 solid waste storage sites, and 2 sewage pipeline leaks. The accuracy rate of pollution source identification reached 95% (optimal value), which is 23 percentage points higher than the traditional identification method (accuracy rate of 72%). The redundant data removal rate reached 38%, effectively reducing the data processing pressure. The identification results matched the actual pollution investigation results of the park with 95%. The core pollution sources under control were clarified through the pollution source contribution quantification formula, providing quantitative support for precise supervision.

[0170] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0171] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for soil pollution source identification and database construction based on a spatiotemporal knowledge graph and dynamic update, characterized in that, The method comprises the following steps: Based on the demand for soil pollution prevention and control, multi-source original data is obtained, and the multi-source original data is preprocessed to obtain preprocessed multi-source data; Based on the preprocessed multi-source data, taking pollution sources, pollutants, environmental media and space-time as core entities, the entities and their attributes and the association relationship between the entities are organized in the form of triples, a four-dimensional space-time knowledge graph is constructed, and the four-dimensional space-time knowledge graph is stored in a graph database; The entity features and relationship features in the four-dimensional space-time knowledge graph are input into a fusion model composed of a random forest model and a graph neural network model to obtain a first identification probability and a second identification probability of each candidate pollution source; Based on a double-algorithm fusion confidence calculation formula, the final identification confidence is obtained according to the first identification probability and the second identification probability, and the target pollution source is determined according to a preset threshold, redundant records are removed by using a space-time association redundancy data removal formula, and the pollution contribution of each target pollution source is quantified by using a pollution source contribution quantification formula; Based on the four-dimensional space-time knowledge graph and the target pollution source, a soil pollution source time series database is designed, multi-dimensional attribute labels are defined for pollution sources, pollutants and environmental media in the soil pollution source time series database, and corresponding data are stored according to the multi-dimensional attribute labels; In the soil pollution source time series database, dynamic update rules based on timestamps and triggers and data update priority rule chains are set for different types of pollution sources; Based on the dynamic update rules based on timestamps and triggers and the data update priority rule chains, a conflict processing result data set is obtained; Based on the conflict processing result data set, the data in the soil pollution source time series database is subjected to present situation detection and consistency verification, and corresponding update and correction operations are triggered to obtain a dynamically updated and optimized soil pollution source time series database data set. 2.The method of claim 1, wherein, The multi-source original data comprises: historical pollution traceability result data, pollution source basic information data, pollutant data and environmental medium data. 3.The method of claim 1, wherein, The data preprocessing of the multi-source original data to obtain preprocessed multi-source data comprises: missing value identification and completion processing of the multi-source original data, estimation and filling of missing data by using a space-time weighted method based on sampling point spatial distance, time difference and neighborhood data distribution to obtain a multi-source data set after missing value filling; determination of abnormal values of each type of data based on the multi-source data set after missing value filling, removal of data records deviating from the normal range by using a statistical criterion based on mean and standard deviation to obtain a multi-source data set after removal of abnormal values; standardization processing of pollutant-related data based on the multi-source data set after removal of abnormal values, scale unification and feature enhancement of the numerical range of different pollutants by comprehensively considering the toxicity grade, data mean and fluctuation amplitude of the pollutants to obtain a multi-source data set after pollution feature enhancement standardization processing; Based on the multi-source data set subjected to the contaminated feature-enhanced standardization processing, data is normalized, and various types of data are adaptively mapped to a unified numerical interval in combination with the dispersion degree of the pollutant data, the effective data amount, and the spatial distribution characteristics of the sampling points and the regional geometric center, so as to obtain the preprocessed multi-source data. 4.The method of claim 3, wherein, An expression of the multi-source data set after the missing value filling is as follows: ; wherein, is a spatio-temporal integrated weight, is a spatio-temporal weight balance coefficient; is a sampling point a spatial Euclidean distance between the sampling point and the neighborhood sampling point is a maximum effective neighborhood distance; is a sampling time difference, is a maximum effective time window; is a K-nearest neighbor set of the sampling point , wherein denotes the original observation value of the j-th type of data corresponding to the sampling point k, denotes the sampling time of the sampling point i, denotes the sampling time of the neighborhood sampling point k. 5.The method of claim 3, wherein, An expression of a formula of the standardization processing is as follows: ; wherein, is a toxicity correction factor, is a toxicity level of the pollutant TL max is a maximum value of the toxicity level of the pollutant; , is a toxicity level of the pollutant is a mean and a standard deviation of the original data of the pollutant represents the original observation data of the jth pollutant corresponding to the sampling point i, represents the standardized result value of the jth pollutant corresponding to the sampling point i after the pollution feature enhancement standardization processing. 6.The method of claim 3, wherein, An expression of a formula of the normalization processing is as follows: ; wherein, is a space-time correction term; is the effective data amount of the pollutant ; is the distance of the sampling point to the geometric center of the area, is the average sampling distance of the area, denotes the minimum value of the jth pollutant in all sampling point data, denotes the maximum value of the jth pollutant in all sampling point data, denotes the standard deviation of all sampling point data of the jth pollutant, denotes the final result value of the jth pollutant after space-time adaptive normalization processing corresponding to the sampling point i, denotes the original observation data of the jth pollutant corresponding to the sampling point i. 7.The method of claim 1, wherein, An expression of a formula of the double-algorithm fusion confidence calculation is as follows: ; in, As a source of pollution The final identification confidence level; , The pollution sources are respectively the outputs of random forest and graph neural network. First recognition probability and second recognition probability; As a source of pollution With Top k Characteristic similarity of high-confidence pollution sources Pollution source output by graph neural network Eigenvectors. 8.The method of claim 1, wherein, An expression of a formula of the spatiotemporal correlation redundant data elimination is as follows: ; wherein, is a redundancy of data p and data q, is a feature similarity weight coefficient; is a spatial redundancy threshold, is a temporal redundancy threshold, represents a feature vector corresponding to data p, represents a feature vector corresponding to data q.

9. A system for soil pollution source identification and database construction based on a spatiotemporal knowledge graph and dynamic updating, characterized in that, The method comprises the following steps: An information base module is configured to acquire multi-source original data based on the demand for soil pollution prevention and control, and to perform data preprocessing on the multi-source original data to obtain preprocessed multi-source data. A knowledge graph construction module is configured to organize entities and their attributes and the correlation between entities in the form of triples based on the preprocessed multi-source data, with pollution sources, pollutants, environmental media, and space-time as core entities, to construct a four-dimensional spatiotemporal knowledge graph, and to store the four-dimensional spatiotemporal knowledge graph in a graph database. An association module is configured to input entity features and relationship features in the four-dimensional spatiotemporal knowledge graph into a fusion model composed of a random forest model and a graph neural network model to obtain first identification probabilities and second identification probabilities of each candidate pollution source. An identification module is configured to obtain a final identification confidence based on a double-algorithm fusion confidence calculation formula according to the first identification probabilities and the second identification probabilities, to determine target pollution sources according to a preset threshold, to eliminate redundant records in combination with a spatiotemporal correlation redundant data elimination formula, and to quantify the pollution contribution of each target pollution source using a pollution source contribution quantification formula. A dynamic update module is configured to design a soil pollution source time series database based on the four-dimensional spatiotemporal knowledge graph and the target pollution sources, to define multi-dimensional attribute labels for pollution sources, pollutants, and environmental media in the soil pollution source time series database, and to perform corresponding data storage according to the multi-dimensional attribute labels. A rule setting module is configured to set dynamic update rules based on timestamps and triggers and data update priority rule chains for different types of pollution sources in the soil pollution source time series database. A conflict module is configured to obtain a conflict processing result data set based on the dynamic update rules based on timestamps and triggers and the data update priority rule chains. A solution module is configured to perform present situation detection and consistency verification on data in the soil pollution source time series database based on the conflict processing result data set, to trigger corresponding update and correction operations, and to obtain a soil pollution source time series database data set subjected to dynamic update and optimization.

Citation Information

Patent Citations

  • Intelligent tracing method and system for agricultural non-point source pollution based on knowledge graph

    CN120670485A

  • Intelligent water quality monitoring and pollution source identification method based on full spectrum analysis

    CN121141544A