A method and system for the convergence and fusion of multi-source heterogeneous data in the environmental governance industry
By employing quality assessment, semantic mapping, blockchain notarization, and dynamic knowledge graph anomaly detection of multi-source heterogeneous data, the problems of low accuracy and insufficient credibility of data fusion in environmental governance have been solved, achieving high-quality data fusion and source tracing analysis, and improving the scientific nature and efficiency of environmental governance.
Patent Information
- Application Number
- CN202511406321.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-09-29
AI Technical Summary
Existing technologies for environmental governance suffer from problems such as low accuracy of multi-source heterogeneous data fusion, insufficient data reliability, poor adaptability to anomaly detection, lack of dynamic knowledge evolution capabilities, and insufficient reliability of source tracing and inference.
Through quality assessment and repair of multi-source heterogeneous data, semantic mapping and spatiotemporal alignment, blockchain notarization, multimodal feature extraction and knowledge graph construction, combined with anomaly detection and tracing models of environmental dynamic knowledge graphs, high-quality data fusion is achieved.
Outputting high-quality data fusion results improves the accuracy and scientific nature of environmental governance data, provides reliable data support, and ensures the interpretability and usability of the data.
Smart Images

Figure CN120873998B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of environmental data processing, and in particular to a method and system for the convergence and fusion of multi-source heterogeneous data in the environmental governance industry. Background Technology
[0002] Existing technologies typically employ traditional data integration and preprocessing methods to handle multi-source heterogeneous data. Environmental data from sources such as data extraction, sensor networks, and manual reporting undergoes initial cleaning and format standardization, followed by centralized storage in relational or database systems. At the data analysis level, most systems utilize rule-based quality verification methods and statistical anomaly detection algorithms for data filtering, and employ basic spatiotemporal interpolation techniques and semantic mapping rules to align data from different sources. Some advanced systems incorporate traditional machine learning methods for data fusion and simple source tracing analysis, with an overall architecture primarily based on hierarchical processing and batch computation.
[0003] However, existing technologies have several shortcomings: First, traditional data preprocessing methods struggle to effectively address semantic ambiguity and spatiotemporal scale inconsistencies unique to the environmental domain, resulting in limited data fusion accuracy. Second, centralized storage architectures suffer from insufficient data credibility, high tampering risks, and a lack of effective credible evidence storage mechanisms. Third, rule-based quality control and anomaly detection methods have poor adaptability and cannot effectively identify complex anomaly patterns under complex environmental conditions. Furthermore, existing methods lack the ability to dynamically evolve knowledge throughout the entire process, leading to insufficient reliability of source tracing and inference results. Finally, traditional data fusion methods fail to adequately consider data credibility assessment and uncertainty quantification, severely limiting the interpretability and practicality of data fusion results. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a method and system for the convergence and fusion of multi-source heterogeneous data in the environmental governance industry.
[0005] To achieve the above objectives, in a first aspect, this invention provides a method for the convergence and fusion of multi-source heterogeneous data in the environmental governance industry. The method includes the following steps: acquiring multi-source heterogeneous environmental data, and performing quality assessment and repair processing on the multi-source heterogeneous environmental data to obtain a standardized environmental dataset; performing semantic mapping and spatiotemporal alignment on the standardized environmental dataset to obtain a unified semantic-spatiotemporal dataset, and storing the unified semantic-spatiotemporal dataset on blockchain to obtain a multi-source heterogeneous data convergence result; extracting multimodal features from the multi-source heterogeneous data convergence result to establish an environmental dynamic knowledge graph, and performing anomaly detection on the environmental dynamic knowledge graph to obtain a knowledge-enhanced environmental dataset; constructing an environmental data tracing model to perform tracing analysis on the knowledge-enhanced environmental dataset to obtain tracing inference results, and performing data fusion based on the tracing inference results to obtain a multi-source heterogeneous data fusion result. This invention, through quality repair, semantic-spatiotemporal alignment, blockchain storage, knowledge graph construction, and tracing fusion, solves problems such as poor data quality, inconsistent semantic-spatiotemporal structure, and difficulty in tracing anomalies in environmental governance, outputting high-quality data fusion results, providing accurate data support for environmental governance decision-making, and improving the efficiency and scientific nature of environmental governance.
[0006] Optionally, the step of acquiring multi-source heterogeneous environmental data and performing quality assessment and remediation processing on the multi-source heterogeneous environmental data to obtain a standardized environmental dataset includes: acquiring the multi-source heterogeneous environmental data of the environmental governance industry through a multi-dimensional environmental data monitoring network; establishing data quality access rules, and performing quality assessment on the multi-source heterogeneous environmental data according to the data quality access rules to obtain qualified data; constructing an environmental data remediation mechanism, and performing remediation processing on the qualified data according to the environmental data remediation mechanism to obtain environmental remediation data; and performing cross-validation and conflict resolution on the environmental remediation data to obtain the standardized environmental dataset. This invention ensures comprehensive data sources through a multi-dimensional monitoring network, filters qualified data through data quality access rules, compensates for data defects through an environmental data remediation mechanism, and ensures data consistency through cross-validation and conflict resolution, ultimately obtaining a standardized dataset. This lays a high-quality foundation for subsequent data processing and improves the accuracy of data analysis results.
[0007] Optionally, the step of semantically mapping and spatiotemporally aligning the standardized environment dataset to obtain a unified semantic-spatiotemporal dataset includes: constructing a semantic recognition mechanism; performing semantic mapping on the standardized environment dataset based on the semantic recognition mechanism to obtain an environment semantic association dataset; establishing a spatiotemporal calibration mechanism; performing spatiotemporal alignment on the standardized environment dataset based on the spatiotemporal calibration mechanism to obtain an environment spatiotemporal calibration dataset; and generating a unified semantic-spatiotemporal identifier by combining the environment semantic association dataset and the environment spatiotemporal calibration dataset to establish the unified semantic-spatiotemporal dataset. This invention combines a semantic recognition mechanism to achieve deep semantic association of data and a spatiotemporal calibration mechanism to unify the spatiotemporal dimensions of data, thereby generating a unified semantic-spatiotemporal identifier and constructing a unified semantic-spatiotemporal dataset. This solves the problems of semantic ambiguity and spatiotemporal misalignment in standardized data, providing a unified basis for data interpretation.
[0008] Optionally, obtaining the multi-source heterogeneous data convergence result by blockchain-based notarization of the semantic-spatiotemporal unified dataset includes: obtaining the cryptographic hash value of the semantic-spatiotemporal unified dataset; encrypting and storing the semantic-spatiotemporal unified dataset and recording the data storage address; constructing notarization metadata, which includes the semantic-spatiotemporal unified identifier, the cryptographic hash value, and the data storage address; automatically verifying the notarization metadata based on a smart contract on the blockchain network, and initiating an on-chain request after successful verification; performing consensus verification on the on-chain request through consensus nodes of the blockchain network, storing the notarization transaction information on the blockchain network after successful verification, and generating a digital notarization certificate bound to the notarization transaction information; and using the semantic-spatiotemporal unified dataset with the digital notarization certificate as the multi-source heterogeneous data convergence result. This invention ensures data security through cryptographic hash values, records key information through notarization metadata, ensures the authenticity and credibility of notarization through automatic verification by smart contracts and verification by consensus nodes, and makes the generated digital notarization certificate traceable. The final multi-source heterogeneous data convergence result enhances the credibility of the data.
[0009] Optionally, the step of extracting multimodal features from the multi-source heterogeneous data aggregation results to establish an environmental dynamic knowledge graph includes: extracting multimodal environmental features from the multi-source heterogeneous data aggregation results to obtain multimodal environmental features, including the time-series dynamic features of numerical time-series data, the textual semantic features of text-based report data, and the visual spatial features of image-based monitoring data; constructing an environmental governance domain ontology to define core entity types and core relationship types, and building a knowledge graph pattern by combining the core entity types and core relationship types; mapping the multimodal environmental features to entity relationship attribute values, and performing relationship reasoning based on environmental spatiotemporal constraint rules and domain business logic rules to construct an initial environmental knowledge graph; and establishing an incremental update and version management mechanism for the knowledge graph to iteratively update the initial environmental knowledge graph to establish the environmental dynamic knowledge graph. This invention comprehensively mines data features through multimodal feature extraction, builds a knowledge graph pattern based on an environmental governance domain ontology, improves graph associations through relationship reasoning, and ensures the dynamic nature of the graph through incremental updates and version management mechanisms. The constructed environmental dynamic knowledge graph can intuitively present environmental data associations, providing structured support for subsequent anomaly detection.
[0010] Optionally, the step of detecting anomalies in the dynamic environmental knowledge graph to obtain a knowledge-enhanced environmental dataset includes: using a graph neural network algorithm to detect anomalies in the node attributes and network structure of the dynamic environmental knowledge graph to obtain graph anomaly information; performing collaborative analysis and correlation parsing on the graph anomaly information based on the multimodal environmental features to obtain composite associated anomaly events; and using the composite associated anomaly events as new knowledge information, assigning confidence labels to the new knowledge information, and injecting it into the dynamic environmental knowledge graph to form the knowledge-enhanced environmental dataset. This invention utilizes graph neural network algorithms to accurately detect graph anomalies, performs collaborative analysis of multimodal features to parse composite associated anomaly events, and injects new knowledge information into the graph to form a knowledge-enhanced environmental dataset. This not only accurately identifies anomalies but also enriches the data knowledge dimensions, enhances data value, and provides a comprehensive data foundation for source tracing analysis.
[0011] Optionally, the step of constructing an environmental data tracing model to perform tracing analysis on the knowledge-enhanced environmental dataset to obtain tracing inference results includes: constructing an environmental process mechanism model and a probabilistic graphical model to obtain the environmental data tracing model; using the environmental process mechanism model to perform reverse simulation on the knowledge-enhanced environmental dataset to obtain anomaly tracing pending information; performing probabilistic reasoning on the anomaly tracing pending information based on the probabilistic graphical model to obtain an anomaly tracing list, and obtaining the tracing evidence chain of the anomaly tracing list; and using the anomaly tracing list and the tracing evidence chain as the tracing inference result. This invention constructs an environmental data tracing model through an environmental process mechanism model and a probabilistic graphical model, accurately locates anomaly sources through reverse simulation and probabilistic reasoning, ensures the reliability of the obtained tracing evidence chain, and provides anomaly data references for data fusion, thus helping to improve the reliability of data fusion results.
[0012] Optionally, the construction of the environmental process mechanism model and the probabilistic graphical model to obtain the environmental data tracing model includes: establishing the environmental process mechanism model based on the migration and evolution laws of environmental elements; constructing the probabilistic graphical model based on the abnormal propagation path of the knowledge-enhanced environmental dataset; and coupling the environmental process mechanism model and the probabilistic graphical model to obtain the environmental data tracing model. The environmental process mechanism model constructed in this invention closely matches actual environmental processes, and the probabilistic graphical model has specific tracing characteristics. The environmental data tracing model formed by the coupling of the two combines mechanistic accuracy with scientific reasoning, enabling more precise and efficient tracing analysis and providing strong technical support for tracing inference results.
[0013] Optionally, the step of obtaining a multi-source heterogeneous data fusion result based on the source tracing and inference results includes: evaluating the knowledge-enhanced environment dataset based on the source tracing and inference results to obtain data point credibility; obtaining dynamic fusion weights for data points based on the data point credibility; fusing the knowledge-enhanced environment dataset according to the dynamic fusion weights of the data points and a weighted fusion algorithm to obtain a preliminary fused data field; and performing uncertainty quantification on the preliminary fused data field to obtain the multi-source heterogeneous data fusion result. This invention ensures high-quality data fusion through data point credibility evaluation and dynamic weight setting, efficiently integrates data to form a preliminary fused data field using a weighted fusion algorithm, and makes the fusion result more transparent through uncertainty quantification. The final multi-source heterogeneous data fusion result has high quality and strong credibility, better meeting the environmental governance industry's demand for accurate data.
[0014] Secondly, this invention provides a multi-source heterogeneous data aggregation and fusion system for the environmental governance industry. The system executes the multi-source heterogeneous data aggregation and fusion method for the environmental governance industry provided by this invention. The system includes input devices, output devices, a processor, and a memory, which are interconnected. The memory stores a computer program, which includes program instructions, and the processor is configured to invoke the program instructions. This invention, through high-performance hardware collaboration, achieves automated and systematic processing of multi-source heterogeneous environmental data aggregation and fusion, avoiding human error, improving data processing efficiency, and providing stable hardware support for data processing in the environmental governance industry. Attached Figure Description
[0015] Figure 1 This is a flowchart of a multi-source heterogeneous data aggregation and fusion method for the environmental governance industry according to an embodiment of the present invention;
[0016] Figure 2 This is a schematic diagram of an unstructured data processing algorithm according to an embodiment of the present invention;
[0017] Figure 3 This is a schematic diagram of the data cleaning and preprocessing process according to an embodiment of the present invention;
[0018] Figure 4 This is a schematic diagram of the data quality assessment system according to an embodiment of the present invention;
[0019] Figure 5 This is a schematic diagram of the intelligent quality control algorithm development process according to an embodiment of the present invention;
[0020] Figure 6 This is a schematic diagram of cross-modal data fusion according to an embodiment of the present invention;
[0021] Figure 7 This is a framework diagram of a multi-source heterogeneous data aggregation and fusion system for the environmental governance industry, according to an embodiment of the present invention. Detailed Implementation
[0022] Specific embodiments of the present invention will now be described in detail. It should be noted that the embodiments described herein are for illustrative purposes only and are not intended to limit the invention. In the following description, numerous specific details are set forth in order to provide a thorough understanding of the invention. However, it will be apparent to those skilled in the art that these specific details are not necessary to practice the invention. In other instances, well-known circuits, software, or methods have not been specifically described to avoid obscuring the invention.
[0023] Throughout this specification, references to "an embodiment," "an embodiment," "an example," or "an example" mean that a particular feature, structure, or characteristic described in connection with that embodiment or example is included in at least one embodiment of the invention. Therefore, the phrases "in an embodiment," "in an embodiment," "an example," or "an example" appearing in various places throughout the specification do not necessarily refer to the same embodiment or example. Furthermore, specific features, structures, or characteristics can be combined in one or more embodiments or examples in any suitable combination and / or sub-combination. Moreover, those skilled in the art will understand that the illustrations provided herein are for illustrative purposes and are not necessarily drawn to scale.
[0024] Please see Figure 1 One embodiment of the present invention provides a method for the convergence and fusion of multi-source heterogeneous data in the environmental governance industry, the method comprising the following steps:
[0025] S1. Obtain multi-source heterogeneous environmental data, and perform quality assessment and repair processing on the multi-source heterogeneous environmental data to obtain a standardized environmental dataset.
[0026] In this embodiment, the acquisition of multi-source heterogeneous environmental data is achieved by constructing a multi-dimensional environmental data monitoring network. This network consists of spatially distributed monitoring equipment, including ground-based fixed monitoring stations, mobile inspection devices, remote sensing satellite observation platforms, and manually reported data terminals. The monitoring data covers multiple environmental elements such as air pollutant concentrations, water quality parameters, and soil pollution indicators. Data acquisition employs various communication protocols, including IoT-specific protocols and traditional industrial bus protocols, ensuring seamless access for various monitoring devices. The acquired raw data includes structured time-series data, semi-structured report data, and unstructured data, thus forming multi-source heterogeneous environmental data. Metadata information is recorded simultaneously during data acquisition, including key attributes such as data source device information, acquisition timestamps, geographical coordinates, and measurement units, laying the foundation for subsequent data processing.
[0027] Please see Figure 2In one optional embodiment, intelligent processing technology for multimodal data is designed for unstructured data processing. The text processing module, based on natural language processing technology, has developed core algorithms such as Chinese word segmentation, named entity recognition, keyword extraction, and text classification. The Chinese word segmentation employs a hybrid strategy combining dictionary and statistical methods, and a dedicated dictionary database has been established for environmental protection terminology, achieving a segmentation accuracy of over 80%. The named entity recognition algorithm is based on a BERT-BiLSTM-CRF model constructed using Bidirectional Encoder Representations from Transformers (BERT), Bidirectional Long Short-Term Memory (BiLSTM), and Conditional Random Field (CRF). This model accurately identifies entity types such as company names, personal names, place names, and technology names, providing a reliable foundation for subsequent data fusion. In terms of semantic understanding, a semantic analysis engine has been developed, achieving deep semantic understanding of unstructured text by constructing an environmental protection industry graph. Knowledge extraction algorithms can automatically extract entity, relationship, and event information from documents such as technical reports, project introductions, and company profiles, constructing structured knowledge representations. Entity extraction, based on deep learning models and combined with contextual information and domain knowledge, can accurately identify specialized entities such as technology names, product models, and performance parameters. Relationship extraction employs methods based on dependency parsing and semantic role labeling, enabling the identification of various semantic relationships between entities, such as technology application relationships, cooperative relationships, and competitive relationships.
[0028] Please see Figure 3 In one optional embodiment, a complete data cleaning and preprocessing workflow is established, employing a standardized "collection-cleaning-transformation-verification-database entry" process to ensure a balance between data quality and processing efficiency. The data collection phase supports access from multiple data sources, including file import, web crawling, and manual entry, and a unified data collection adapter is established to shield the technical differences between different data sources. The data cleaning phase implements functions such as deduplication, missing value handling, outlier detection, and format standardization. The deduplication algorithm, based on multi-field matching and similarity calculation, accurately identifies and merges duplicate records. Missing value handling employs various strategies, including mean imputation, regression prediction, and expert rules, selecting the optimal strategy based on field characteristics and business requirements. The data transformation phase implements functions such as data format conversion, encoding conversion, and unit standardization, establishing a unified data format specification to ensure data consistency and comparability.
[0029] Furthermore, to improve data processing efficiency, an automated preprocessing system based on a rules engine was developed. This system can automatically select appropriate processing strategies based on data characteristics, reducing the need for manual intervention. The rules engine component, based on expert experience and business logic, has established a rich data processing rule base, covering common data quality issues and processing methods. The system also establishes a processing log and auditing mechanism to record detailed information for each processing step, supporting process backtracking and problem diagnosis.
[0030] After obtaining multi-source heterogeneous environmental data, a systematic quality assessment is conducted based on established data quality admission rules. These rules encompass four dimensions: integrity rules, scope reasonableness rules, logical consistency rules, and timeliness rules. Integrity rules require data records to include essential fields, and the missing rate must not exceed a preset threshold (e.g., 5%). Scope reasonableness rules stipulate that parameter values must be within physically possible ranges (e.g., water temperature between 0℃ and 100℃) and historical statistical ranges (e.g., based on a three-standard-deviation criterion). Logical consistency rules check the logical relationships between related parameters (e.g., the negative correlation between dissolved oxygen and water temperature). Timeliness rules ensure that the delay between data acquisition and reception is within acceptable limits. The quality assessment employs an automated processing workflow based on a rule engine to obtain data quality scores. Multiple rules are applied in parallel to each data record, and data that passes the quality assessment is considered quality-qualified.
[0031] Please see Figure 4 In one optional embodiment, a data quality assessment system is constructed, establishing a structured data quality assessment system covering six dimensions: completeness, accuracy, timeliness, consistency, validity, and reliability. Differentiated quality assessment standards and testing rules are formulated for different types of structured data, such as information from alliance members, project data, and expert information. Completeness testing assesses the completeness of data through indicators such as field fill rate and record completeness, establishing a field weighting system based on business importance and imposing stricter completeness requirements on key fields. Accuracy testing employs various technical means, including format validation, range checks, and logical verification, ensuring data accuracy through regular expressions, data dictionaries, and business rules. Consistency testing focuses on the data consistency of the same entity across different systems, achieving unified data management and consistency maintenance through the establishment of a master data management mechanism.
[0032] Please see Figure 5In one optional embodiment, regarding intelligent quality control algorithms, a machine learning-based intelligent data quality control module has been developed. This module can automatically learn and identify data anomaly patterns, improving the intelligence level of quality control. The anomaly detection algorithm, based on unsupervised learning methods such as Isolation Forest and Local Outlier Factor (LOF), can identify statistically abnormal data points. The data repair algorithm, based on historical data patterns and business rules, automatically repairs or marks detected abnormal data for manual processing. Furthermore, the quality control module integrates data lineage tracking, enabling the tracing of data sources and change history, providing support for root cause analysis of data quality issues. The quality control system also establishes a real-time monitoring mechanism. By setting quality thresholds and early warning rules, it can promptly detect and address data quality problems, ensuring the continuous stability of data quality.
[0033] Furthermore, for repairable defects in qualified data, an environmental data repair mechanism is used to obtain environmentally repaired data. A spatiotemporal graph attention network model is employed as the environmental data repair mechanism. For missing values caused by transient sensor malfunctions and outliers significantly deviating from the normal range, corrections are made using the spatiotemporal graph attention network model. The spatiotemporal correlation characteristics of environmental elements are fully considered during the repair process; for example, upstream monitoring point data is used to assist downstream data repair, and historical data from the same period is used to assist current data repair. All repair operations on the environmentally repaired data are logged, including the original value, the repaired value, the repair method, and the repair confidence level, ensuring the traceability of the repair process.
[0034] The above spatiotemporal graph attention network model satisfies the following relationship:
[0035]
[0036] in, For in position time The repair value, For adjacent positions, For the set of adjacent positions, Spatial attention weights, Adjacent positions In time The observed values, The time series contribution coefficient. This represents the temporal characteristics of historical location data.
[0037] It should be noted that the temporal contribution coefficient is obtained by learning from historical data using a spatiotemporal graph attention network.
[0038] In this embodiment, cross-validation and conflict resolution are performed on the restored environmental remediation data to generate a standardized environmental dataset. Cross-validation employs multi-source data comparison methods, such as cross-validating measurements from different monitoring devices in the same area, and cross-validating measurement results of the same parameter from different monitoring methods. Conflict resolution is based on evidence theory and a reliability weighting method, assigning reliability weights to data from different sources and resolving data conflicts through weighted fusion. The final result is a standardized environmental dataset using a unified data model, containing consistent field naming conventions, data encoding formats, and a measurement unit system. It also includes complete data quality metadata, including quality scores, remediation records, and validation results, providing a high-quality, standardized data foundation for subsequent processing.
[0039] S2. Perform semantic mapping and spatiotemporal alignment on the standardized environment dataset to obtain a unified semantic spatiotemporal dataset, and perform blockchain notarization on the unified semantic spatiotemporal dataset to obtain the multi-source heterogeneous data aggregation result.
[0040] In this embodiment, a semantic recognition mechanism is established based on the ontology of environmental governance. This mechanism employs multimodal semantic parsing technology: for structured data fields, a semantic annotation rule base is established, and the original field names are converted into standardized semantic tags through regular expression matching and dictionary mapping; for semi-structured report data, entity recognition technology is used, combined with conditional random field algorithms to extract environmental entities and their relationships; for unstructured text data, knowledge graph embedding technology is used for semantic recognition. Finally, an environmental semantic association dataset is generated and stored in triplet format, with each data point associated with a clear semantic description, forming a machine-understandable semantic network.
[0041] Furthermore, a unified framework is established for multi-source spatiotemporal reference systems to obtain a spatiotemporal calibration mechanism. The time calibration mechanism adopts the unified time representation method based on the International Organization for Standardization (ISO), converting all time information to Coordinated Universal Time (UTC) with millisecond accuracy. For data collected at different frequencies, downsampling is used for high-frequency sensor data, and time series interpolation is used for low-frequency manual monitoring data. Spatial calibration employs a multi-level coordinate system transformation system, uniformly transforming various coordinate systems to the standard geographic coordinate system. For mobile monitoring equipment, spatial location correction is performed by combining positioning trajectory data and map matching algorithms. During the spatiotemporal calibration process, the spatiotemporal characteristics of environmental data are considered, using spatiotemporal kriging interpolation to fill data gaps, and spatiotemporal correlation analysis is used to verify the rationality of the calibration results. Finally, an environmental spatiotemporal calibration dataset with a unified spatiotemporal reference framework is generated, providing a consistent spatiotemporal benchmark for subsequent analysis.
[0042] Furthermore, semantic feature vectors, including entity semantic hashes, spatiotemporal hashes, and data quality hashes, are extracted from the environmental semantic association dataset. Spatiotemporal feature vectors, including time-series geo-hash encoding and spatial grid encoding, are extracted from the environmental spatiotemporal calibration dataset. A Merkle Hash Tree structure is used to fuse the semantic and spatiotemporal feature vectors at multiple levels. A unique semantic spatiotemporal unified identifier is generated through a secure hash algorithm, comprising a version number, semantic digest, spatiotemporal digest, and checksum, ensuring the identifier's uniqueness, stability, and verifiability. This semantic spatiotemporal unified identifier is then combined to establish a semantic spatiotemporal unified dataset with complete semantic annotations, supporting efficient spatiotemporal range queries and semantic retrieval, laying the data foundation for blockchain-based evidence storage.
[0043] The above-mentioned semantic-spatial unified identifiers satisfy the following relationship:
[0044]
[0045] in, For semantic and spatiotemporal unified identifiers, For cryptographic hash functions, For entity semantic hashing, Represents an entity, For spatiotemporal hashing, Indicates time, Represents coordinates, For data quality hashing, Indicates the data quality score. This indicates a splicing operation.
[0046] In this embodiment, the cryptographic hash value of the semantic spatiotemporal unified dataset is first calculated. A batch hashing method based on Merkle trees is used to divide the dataset into blocks, calculating the hash value of each block separately, and then constructing the complete Merkle tree level by level. Encrypted storage employs a hybrid encryption scheme: the dataset is symmetrically encrypted using an encryption algorithm to obtain a symmetric key, which generates an encrypted data file; the symmetric key is further encrypted to generate a key envelope. The encrypted data is stored in a distributed file system, where the system records the data storage address, encryption algorithm parameters, and access control information to ensure data confidentiality and accessibility.
[0047] The above cryptographic hash values satisfy the following relationship:
[0048]
[0049] in, The root hash value of the Merkle tree. For cryptographic hash functions, Divide the dataset into chunks. This indicates a splicing operation.
[0050] Furthermore, evidence storage metadata is constructed, including core metadata segments and extended metadata segments. Core metadata includes a semantically and spatiotemporally unified identifier, cryptographic hash value, data storage address, encryption algorithm parameters, and digital signature information. Extended metadata includes data source information, environmental monitoring standard compliance statements, data quality indicators, and access control policies. All metadata fields conform to the Data Catalog Vocabulary (DCAT) standard to ensure machine readability and interoperability.
[0051] Furthermore, the hash value of the evidence storage metadata itself is calculated, and together with the data hash, it forms a complete evidence storage information package.
[0052] In this embodiment, a smart contract deployed on the blockchain network executes an automated verification process to initiate an on-chain request. The smart contract first verifies the compliance of the notarization request format, checking whether required fields are complete and whether the data format conforms to specifications. Then, it performs business logic verification: verifying the identity and permissions of the data provider by querying the on-chain registry; checking the legality and uniqueness of the semantic spatiotemporal unified identifier; and verifying the validity of the digital signature. The smart contract also performs a consistency check, ensuring data integrity and consistency by comparing the hash value in the metadata with the actual calculated data hash value.
[0053] Specifically, during the consensus verification phase, the consensus nodes of the blockchain network perform distributed consensus verification on on-chain requests that have passed smart contract verification. This includes five stages: request, pre-preparation, preparation, submission, and response. The consensus verification mainly includes: verification of the legality of the evidence storage request, verification of the correctness of the smart contract execution result, and verification of the synchronization of the timestamp. After successful verification, each consensus node packages the evidence storage transaction information (including metadata hash, timestamp, block height, etc.) into a new block and stores it in the blockchain network via cryptographic links.
[0054] Furthermore, a digital evidence certificate is generated that is bound to the evidence transaction information, containing key information of the evidence transaction: blockchain network identifier, block height, transaction hash, timestamp information, and evidence digital information.
[0055] In this embodiment, the digital evidence is associated with the original semantic spatiotemporal unified dataset to generate a multi-source heterogeneous data aggregation result with complete and credible proof, providing a reliable foundation for subsequent data analysis and applications. All operations in the entire evidence storage process are recorded in the audit log to ensure the traceability and auditability of the entire process.
[0056] S3. The results of the aggregation of the multi-source heterogeneous data are subjected to multimodal feature extraction to establish an environmental dynamic knowledge graph, and anomaly detection is performed on the environmental dynamic knowledge graph to obtain a knowledge-enhanced environmental dataset.
[0057] In this embodiment, a hierarchical feature extraction architecture is employed to process the convergence results of multi-source heterogeneous data to obtain multimodal environmental features. For numerical time-series data, data resampling and smoothing are first performed. A time-series decomposition algorithm is used to separate trend terms, periodic terms, and residual terms. Then, a long short-term memory network autoencoder is used to extract time-series dynamic features, capturing the long-range dependencies and evolution patterns of environmental parameters. For text-based report data, a multi-task learning model is used, combined with a pre-trained language model in the environmental domain, to simultaneously perform named entity recognition, relation extraction, and semantic analysis tasks, extracting textual semantic features containing environmental entities, numerical indicators, and state descriptions. For image-based monitoring data, a deep convolutional neural network combined with an attention mechanism is used to extract visual spatial features, focusing on the visual patterns of key areas such as pollution discharge outlets and the operational status of treatment facilities. Multi-scale visual spatial features are then fused through a feature pyramid network. After normalizing the time-series dynamic features, textual semantic features, and visual spatial features, a cross-modal attention mechanism is used for feature alignment and fusion to form unified multimodal environmental features, providing a rich feature foundation for knowledge graph construction.
[0058] In this embodiment, the construction of the environmental governance ontology adopts a standardized hierarchical modeling method: First, by analyzing environmental monitoring standards, pollution source coding specifications, and environmental business procedures, the core concept system is extracted, and entity types and their hierarchical relationships are defined; then, semantic associations between entities are established through object attributes and data attributes, and spatiotemporal constraints and business rules are added as axiom constraints; finally, formal modeling is performed to obtain the environmental governance ontology, which supports machine-readable semantic reasoning and consistency checks.
[0059] Furthermore, based on the ontology definition in the field of environmental governance, core entity types and core relationship types are defined hierarchically. First, core entity types are defined using domain expert knowledge and standards, including pollution sources, monitoring sites, environmental media, treatment facilities, and regulatory units. Each entity type includes attribute patterns and data type constraints. Then, relationship types between entities are defined, including spatial relationships (e.g., location, adjacency), physical relationships (e.g., emission, impact), process relationships (e.g., treatment, transformation), and management relationships (e.g., regulation, affiliation). Finally, combining the core entity types and core relationship types, a knowledge graph model is established, including a hierarchical structure of entity types, attribute inheritance relationships, and relationship constraints, providing a structured framework for graph construction.
[0060] In this embodiment, multimodal environmental features are mapped to knowledge graph patterns. Entity linking technology is used to match multimodal environmental features with entity types, and a similarity-based entity parsing algorithm is used to resolve ambiguity issues related to entities with the same name. Multimodal environmental features are mapped to entity relationship attribute values, and timestamps and confidence level metadata are added. Relationship reasoning is performed based on environmental spatiotemporal constraints (such as pollutant diffusion models and hydrological correlation models) and domain business logic rules (such as emission standards and treatment process procedures) to identify implicit environmental relationships, thereby establishing an initial environmental knowledge graph, including data sources, time validity, and confidence level scores, forming a complete knowledge network containing entities, attributes, and relationships, and stored in a triplet format.
[0061] Specifically, the multimodal environment features are mapped to entity relation attribute values that satisfy the following relationship:
[0062]
[0063] in, For joint mapping functions, For the first A multimodal environment feature, The first in the environmental dynamics knowledge graph Entity type, This indicates that the optimal entity relationship attribute value has been obtained. For entity relationship attribute values, These are learnable weight coefficients. For semantic similarity, The confidence score is... It serves as a measure of spatiotemporal consistency.
[0064] The above relational reasoning satisfies the following relation:
[0065]
[0066] in, For nodes In the Layer feature representation, For activation function, Adjacent nodes For nodes The set of adjacent nodes, The normalization coefficient is... For the first The weight matrix of the layer, For nodes In the Layer feature representation.
[0067] In this embodiment, an incremental update and version management mechanism for the knowledge graph is designed to establish a dynamic environmental knowledge graph. A time-window-based incremental update strategy is established to receive new multi-source data in real time, extract features, and update the knowledge graph. A graph difference algorithm is used to identify knowledge changes, updating only the subgraphs that have changed, thus improving update efficiency. Version management employs a graph version control system to store historical versions and change records of the knowledge graph, supporting version rollback and change traceability. Simultaneously, a knowledge quality assessment mechanism is established to periodically check the consistency of the knowledge graph and resolve conflicts, ensuring the accuracy and timeliness of the knowledge. The resulting dynamic environmental knowledge graph possesses real-time update, historical traceability, and quality assurance capabilities, providing a reliable knowledge foundation for anomaly detection.
[0068] In this embodiment, a heterogeneous graph neural network model based on an attention mechanism is used to process the dynamic knowledge graph of the environment. First, feature engineering is performed on the knowledge graph, embedding node attribute features into a low-dimensional space using a multilayer perceptron. A relational graph convolutional network is then used to capture structural features under different types of relationships. For node attribute anomaly detection, a multi-scale temporal convolutional module is designed to analyze the temporal change patterns of node attributes and detect outliers by combining autoencoder reconstruction errors. For network structure anomaly detection, normal graph structure patterns are learned, and abnormal connections and subgraphs are identified through reconstruction errors. Model training uses normal samples to train a baseline model, and a dynamic threshold adjustment mechanism distinguishes between normal and abnormal patterns. The detected graph anomaly information includes a list of abnormal nodes, a list of abnormal edges, and abnormal subgraph patterns. Each anomaly is accompanied by an anomaly score and anomaly type classification, providing detailed anomaly clues for subsequent analysis.
[0069] Furthermore, in the collaborative analysis and correlation resolution stage, a multimodal feature-enhanced anomaly correlation analysis framework is established. First, the graph anomaly information detected by the graph neural network is correlated across modalities with multimodal environmental features. An attention mechanism is used to calculate the correlation weights between anomaly nodes and each modal feature, identifying the dominant anomaly feature modality. Then, anomaly propagation analysis based on causal reasoning is employed to construct an anomaly propagation graph model, analyzing the propagation path and impact range of anomalies in the spatiotemporal and logical dimensions. For composite correlated anomaly events, a hierarchical clustering algorithm is used to cluster related anomalies into anomaly events. Sequence pattern mining is used to analyze the temporal evolution patterns of anomaly events, and spatial analysis techniques are used to determine the spatial distribution characteristics of anomaly events. Finally, composite correlated anomaly events are generated, including event type, impact range, temporal evolution process, spatial distribution pattern, and multimodal feature manifestations, achieving a comprehensive understanding of anomaly events.
[0070] Furthermore, a knowledge fusion mechanism based on confidence assessment is designed to enhance and inject new knowledge-based information into a knowledge-enhanced environment dataset. First, composite and related anomaly events are represented as knowledge-based data, converted into triples acceptable to the knowledge graph, including anomaly event entities and anomaly type attributes. Then, a multi-source confidence assessment model is established, comprehensively considering factors such as anomaly detection confidence, data source reliability, time validity, and expert verification results to calculate the comprehensive confidence score for each anomaly knowledge. During the knowledge injection phase, high-confidence knowledge is directly integrated into the knowledge graph, medium-confidence knowledge is stored as hypotheses to be verified, and low-confidence knowledge undergoes a manual review process. Consistency checks are performed during injection to prevent knowledge conflicts, and the confidence of relevant nodes is updated through a knowledge propagation algorithm. The resulting knowledge-enhanced environment dataset not only contains the original environmental data but also integrates new knowledge discovered through anomaly detection, significantly improving the dataset's knowledge density and value, and providing stronger data support for environmental governance decisions.
[0071] S4. Construct an environmental data tracing model to perform tracing analysis on the knowledge-enhanced environmental dataset to obtain tracing inference results, and perform data fusion based on the tracing inference results to obtain multi-source heterogeneous data fusion results.
[0072] In this embodiment, an environmental process mechanism model is established based on the migration and evolution laws of environmental elements. For the migration process of air pollutants, a partial differential equation system including advection-diffusion-reaction terms is constructed using computational fluid dynamics. The advection term uses the Navier-Stokes equations to describe the wind field effect, the diffusion term is modeled based on Fick's law, and the reaction term describes the chemical transformation process using the Arrhenius equation. For the migration of water pollutants, considering the spatial variability of hydrological parameters (flow velocity, water depth, riverbed roughness) and water quality parameters (degradation coefficient, adsorption coefficient), a hydrodynamic coupling model is established. The model parameters are optimized using Bayesian calibration based on historical monitoring data, and the posterior distribution of parameters is sampled using the Markov chain Monte Carlo algorithm to ensure the model's consistency with actual conditions. Furthermore, the finite volume method is employed, with adaptive adjustment of the time step, and boundary conditions are dynamically set based on topographic and meteorological data, ultimately forming an environmental process mechanism model that can simulate the spatiotemporal evolution of pollutants.
[0073] In this embodiment, a Bayesian network structure is constructed as a probabilistic graphical model based on the anomaly propagation paths of the knowledge-enhanced environmental dataset. First, a causal discovery algorithm is used to learn the causal relationships between variables from historical anomaly data, constructing the skeleton structure of a directed acyclic graph. Nodes are categorized into three types: observed variables (sensor data), latent variables (unobservable environmental conditions), and intervention variables (pollution source intensity). Edge directions are constrained based on environmental mechanism knowledge (e.g., downstream diffusion of pollutants) and time priority relationships (cause precedes effect). The conditional probability distribution is modeled parametrically: Gaussian process regression is used for continuous variables, and multinomial distribution is used for discrete variables. Conditional probabilities considering spatiotemporal dependence are estimated using the spatiotemporal kriging method. Model parameters are learned using the expectation-maximization algorithm, and historical data is used for model training and validation, ultimately constructing a probabilistic graphical model capable of representing the probabilistic relationships of anomaly propagation.
[0074] Furthermore, in the model coupling stage, a co-simulation framework is employed to achieve deep integration of the environmental process mechanism model and the probabilistic graphical model. The environmental process mechanism model undergoes parameter dimensionality reduction, and key parameter features are extracted through principal component analysis. These key parameter features are connected to the probabilistic graphical model via shared variables: the output of the environmental process mechanism model (e.g., concentration field) serves as observational evidence for the probabilistic graphical model, and the inference results of the probabilistic graphical model are used to adjust the input parameters of the environmental process mechanism model. The coupling interface employs a variational inference-based coupling mechanism to achieve bidirectional probabilistic information transfer. Ultimately, this coupling forms an environmental data tracing model, supporting reverse probabilistic inference from observational data to pollution sources.
[0075] The above coupling interface satisfies the following relationship:
[0076]
[0077] in, It is the natural logarithm function. For the posterior probability distribution, This represents the hidden state of the environmental process mechanism model. For observation data, For expectation operator, For variational posterior distribution, These are the latent variables in the probabilistic graphical model. Let be the likelihood function. for divergence, Let be the posterior probability distribution.
[0078] In this embodiment, an inversion method combining the adjoint equation method and regularized optimization is used to perform inverse simulation on a knowledge-enhanced environmental dataset using an environmental process mechanism model to obtain undetermined information on anomaly source tracing. First, the source tracing problem is transformed into an optimization problem, with the objective function including a fitting term between the observed data and the model output, as well as a regularization term. Using the constructed environmental process mechanism model, the simulated concentration field is obtained by numerically solving partial differential equations. During the inverse simulation, the adjoint method is used to efficiently calculate the gradient of the objective function with respect to the source strength parameter: an adjoint equation corresponding to the original partial differential equation is constructed, and the sensitivity field is calculated by inverse time integration. For chemical reaction processes with strong nonlinearity, iterative regularized optimization is used, gradually approximating the optimal solution through sequence linearization. The uncertainty of the observed data is considered during the inversion process, and the contribution of reliable data is increased through covariance matrix weighting. Finally, undetermined information on anomaly source tracing is output, including the probability distribution of possible pollution source locations, the time variation curve of the source strength parameter, and its uncertainty range, providing preliminary source tracing hypotheses for subsequent probabilistic inference.
[0079] Furthermore, based on a probabilistic graphical model, probabilistic inference is performed on the pending information for anomaly tracing to obtain an anomaly source list and a source tracing evidence chain. First, the pending source tracing information is transformed into a prior distribution of the probabilistic graphical model, and then the posterior probability is calculated using Bayes' theorem. For discrete variable inference, a connection tree algorithm is used to accurately calculate marginal probabilities; for continuous variables, stochastic variational inference is used to approximate the posterior distribution by optimizing the variational lower bound. The inference process considers spatiotemporal correlation, defining the dependence strength between variables by introducing a spatiotemporal kernel function. The anomaly source list is generated using maximum a posteriori probability estimation, listing the top K potential pollution sources with the highest probabilities. Each source includes location coordinates, source strength estimate, probability confidence, and contribution rate index. The source tracing evidence chain is constructed through reverse path tracing: starting from the anomaly node, the model is traversed backwards along the edges of the probabilistic graphical model, extracting node state changes, conditional probability values, and information gain indices on key inference paths to form a complete evidence chain containing data evidence, model evidence, and inference logic as the source tracing evidence chain. The final output of the source tracing inference results includes a list of anomalies sorted by probability and the corresponding source tracing evidence chain, providing credible decision support for environmental regulation.
[0080] In this embodiment, a multi-dimensional credibility assessment is performed on the knowledge-enhanced environment dataset based on the source tracing and inference results. A credibility evaluation index system is established, including four dimensions: data source reliability, temporal consistency, spatial coordination, and mechanistic conformity. Each dimension is assigned an adaptive weight, which is dynamically adjusted according to the source tracing results: the weight of data from high-risk areas confirmed by source tracing is reduced; the weight of data from confirmed reliable benchmark sites is increased. Finally, each data point obtains a credibility score between 0 and 1, along with a credibility decomposition index, forming a data quality profile.
[0081] Furthermore, dynamic fusion weights for data points are obtained based on their credibility. Discrete data point credibility is converted into a continuous probability distribution through kernel density estimation. Then, a weight allocation strategy based on information entropy is used to obtain dynamic fusion weights for data points. Data points with higher credibility receive higher weights, while considering data diversity to avoid over-reliance on a single source.
[0082] The above dynamic fusion weights satisfy the following relationship:
[0083]
[0084] in, For data points Dynamic fusion weights, The observed variance of the data points. It is an exponential function. To adjust hyperparameters, for divergence, For the posterior probability distribution, For data points, To trace the source and infer the results, Let be the prior probability distribution.
[0085] In this embodiment, the weighted fusion algorithm employs a robust statistical method to fuse the knowledge-enhanced environment dataset according to the dynamic fusion weights of the data points, obtaining a preliminary fused data field. Weighted least squares fitting is used for numerical data, and weighted voting integration is used for categorical data. Spatiotemporal heterogeneity is specially handled during the fusion process: in the spatial dimension, adaptive kernel regression is used, with the kernel bandwidth dynamically adjusted according to the data density; in the temporal dimension, variational mode decomposition is used to decompose the data into different frequency components, which are then fused separately. This yields a preliminary fused data field, whose spatiotemporal points contain fused values and contributor source lineage information.
[0086] Please see Figure 6In one optional embodiment, a cross-modal data fusion technique is designed to achieve deep fusion of different modalities such as text and structured data. The fusion framework adopts a three-stage processing mode of "feature alignment - semantic mapping - fusion decision". First, through feature extraction and alignment techniques, data from different modalities are mapped to a unified feature space; then, through semantic mapping techniques, semantic relationships between data from different modalities are established; finally, through a fusion decision algorithm, comprehensive fusion and consistent representation of multimodal information are achieved. Regarding feature alignment, BERT-based sentence vector representation is used for text data, deep feature extraction based on Residual Network (ResNet) is used for image data, and vector representation based on graph embedding is used for structured data. Cross-modal contrastive learning techniques are used to achieve the alignment of features from different modalities.
[0087] It is worth noting that at the fusion decision-making level, an attention-based fusion decision-making algorithm was developed. This algorithm can perform weighted fusion based on the importance and credibility of data from different modalities. The attention mechanism can automatically learn the contribution of different modalities to the final fusion result, avoiding the limitations of traditional fixed-weight methods. The fusion algorithm also integrates a conflict detection and resolution mechanism. When conflicting information exists between different modalities, it can intelligently arbitrate based on factors such as the authority, timeliness, and completeness of the data source. In addition, a fusion effect evaluation system was designed to assess the quality of the fusion result through indicators such as accuracy, completeness, and consistency, and a feedback mechanism was established to continuously optimize the performance of the fusion algorithm. The successful implementation of cross-modal fusion technology has laid a solid technical foundation for the alliance to build unified, complete, and accurate data assets, significantly improving the value and application effect of the data.
[0088] In this embodiment, uncertainty quantification is performed on the initial fused data field, and the Monte Carlo method is used to simulate uncertainty propagation: sampling is performed from the data confidence distribution, model parameter distribution, and fusion algorithm error distribution to generate multiple fusion results. The probability distribution of the fused value at each spatiotemporal point is statistically calculated using the multiple fusion results, and the mean field, variance field, and confidence interval field are extracted. Uncertainty quantification specifically distinguishes between cognitive uncertainty (originating from insufficient knowledge) and stochastic uncertainty (originating from inherent randomness), and the probability box method is used to represent mixed uncertainty. Finally, a multi-source heterogeneous data fusion result is generated, containing three levels of information: the best estimate field provides directly usable multi-source heterogeneous fused data; the uncertainty field provides data reliability indicators; and the source lineage field provides data source and historical processing information, meeting the comprehensive needs for data accuracy, reliability, and traceability in environmental governance industry applications.
[0089] Please see Figure 7In an optional embodiment, the present invention provides a multi-source heterogeneous data aggregation and fusion system for the environmental governance industry. The system includes input devices, output devices, a processor, and a memory, all interconnected. The memory stores a computer program comprising program instructions, and the processor is configured to invoke the program instructions to execute specific steps as described in the relevant embodiments of the multi-source heterogeneous data aggregation and fusion method for the environmental governance industry provided by the present invention. The multi-source heterogeneous data aggregation and fusion system for the environmental governance industry provided by the present invention has a complete structure, is objective and stable, and enhances the overall applicability and practical application capabilities of the present invention.
[0090] In summary, the present invention provides a method and system for the convergence and fusion of multi-source heterogeneous data in the environmental governance industry. It obtains a standardized environmental dataset through quality assessment and remediation processing, forms a unified semantic-spatiotemporal dataset through semantic mapping and spatiotemporal alignment, and uses blockchain for notarization to ensure data credibility. Furthermore, it constructs a dynamic environmental knowledge graph through multimodal feature extraction, and uses graph neural networks for anomaly detection to obtain a knowledge-enhanced environmental dataset. Finally, it combines environmental process mechanism models and probabilistic graphical models for source tracing analysis, and performs credibility-weighted data fusion and uncertainty quantification based on the source tracing inference results to generate high-quality multi-source heterogeneous data fusion results. This achieves credible and intelligent processing of environmental data throughout the entire process from collection, processing, notarization to knowledge-based analysis and fusion applications. The method of this invention is easy to understand, computationally simple, requires minimal workload, and is convenient for engineering applications, providing a theoretical foundation and technical support for the further development of the environmental data processing field.
[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the claims and specification of the present invention.
Claims
1. A method for merging and integrating multi-source heterogeneous data in the environmental governance industry, characterized in that, Includes the following steps: Acquire multi-source heterogeneous environmental data, and perform quality assessment and repair processing on the multi-source heterogeneous environmental data to obtain a standardized environmental dataset; Semantic mapping and spatiotemporal alignment are performed on the standardized environment dataset to obtain a unified semantic spatiotemporal dataset, and blockchain notarization is performed on the unified semantic spatiotemporal dataset to obtain the multi-source heterogeneous data aggregation result. The results of the aggregation of the multi-source heterogeneous data are subjected to multimodal feature extraction to establish an environmental dynamic knowledge graph, and anomaly detection is performed on the environmental dynamic knowledge graph to obtain a knowledge-enhanced environmental dataset. An environmental data tracing model is constructed to perform tracing analysis on the knowledge-enhanced environmental dataset to obtain tracing inference results. Based on the tracing inference results, data fusion is performed to obtain multi-source heterogeneous data fusion results. The step of extracting multimodal features from the aggregated results of the multi-source heterogeneous data to establish a dynamic environmental knowledge graph includes: The results of the aggregation of multi-source heterogeneous data are subjected to multimodal feature extraction to obtain multimodal environmental features, including the temporal dynamic features of numerical time series data, the textual semantic features of textual report data, and the visual spatial features of image-based monitoring data. An ontology for the field of environmental governance is constructed to define core entity types and core relationship types, and a knowledge graph model is built by combining the core entity types and the core relationship types. The multimodal environment features are mapped to entity relationship attribute values, and relationship reasoning is performed based on environmental spatiotemporal constraint rules and domain business logic rules to construct an initial environmental knowledge graph; Establish an incremental update and version management mechanism for the knowledge graph, and iteratively update the initial environment knowledge graph to establish the dynamic environment knowledge graph; The knowledge-enhanced environment dataset obtained by performing anomaly detection on the dynamic knowledge graph of the environment includes: A graph neural network algorithm is used to detect anomalies in the node attributes and network structure of the environmental dynamic knowledge graph to obtain graph anomaly information; Based on the multimodal environmental characteristics, the map anomaly information is analyzed collaboratively and correlated to obtain composite correlated anomaly events; Using the composite associated anomaly events as new knowledge information, the new knowledge information is assigned a confidence label and then injected into the environmental dynamic knowledge graph to form the knowledge-enhanced environment dataset; The constructed environmental data tracing model performs tracing analysis on the knowledge-enhanced environmental dataset to obtain tracing inference results, including: Construct an environmental process mechanism model and a probabilistic graphical model to obtain the environmental data tracing model; The environmental process mechanism model is used to perform reverse simulation on the knowledge-enhanced environment dataset to obtain anomaly tracing information; Based on the probabilistic graphical model, probabilistic reasoning is performed on the anomaly tracing pending information to obtain an anomaly tracing list, and the tracing evidence chain of the anomaly tracing list is obtained. The anomaly tracing list and the tracing evidence chain are used as the tracing inference results; The process of fusing data based on the source tracing and inference results to obtain multi-source heterogeneous data fusion results includes: The credibility of data points is obtained by evaluating the knowledge enhancement environment dataset based on the source inference results, and the dynamic fusion weight of data points is obtained based on the credibility of data points. Based on the dynamic fusion weights of the data points, a weighted fusion algorithm is used to fuse the knowledge-enhanced environment dataset to obtain a preliminary fused data field; Uncertainty quantification is performed on the preliminary fused data field to obtain the multi-source heterogeneous data fusion result.
2. The method for multi-source heterogeneous data aggregation and fusion in the environmental governance industry according to claim 1, characterized in that, The process of acquiring multi-source heterogeneous environmental data and performing quality assessment and restoration on the multi-source heterogeneous environmental data to obtain a standardized environmental dataset includes: The multi-source heterogeneous environmental data of the environmental governance industry is acquired through a multi-dimensional environmental data monitoring network; Establish data quality access rules, and conduct quality assessments on the multi-source heterogeneous environment data according to the data quality access rules to obtain qualified data. An environmental data restoration mechanism is constructed, and the qualified data is restored according to the environmental data restoration mechanism to obtain environmental restoration data; The environmental remediation data is cross-validated and conflict resolution is performed to obtain the standardized environmental dataset.
3. The method for merging and integrating multi-source heterogeneous data in the environmental governance industry according to claim 1, characterized in that, The process of semantically mapping and spatiotemporally aligning the standardized environment dataset to obtain a unified semantic-spatiotemporal dataset includes: A semantic recognition mechanism is constructed, and the standardized environment dataset is semantically mapped based on the semantic recognition mechanism to obtain an environment semantic association dataset. A spatiotemporal calibration mechanism is established, and the standardized environmental dataset is spatiotemporally aligned according to the spatiotemporal calibration mechanism to obtain an environmental spatiotemporal calibration dataset. A semantic spatiotemporal unified identifier is generated by combining the environmental semantic association dataset and the environmental spatiotemporal calibration dataset to establish the semantic spatiotemporal unified dataset.
4. The method for merging and integrating multi-source heterogeneous data in the environmental governance industry according to claim 3, characterized in that, The step of obtaining multi-source heterogeneous data aggregation results by performing blockchain notarization on the semantic spatiotemporal unified dataset includes: Obtain the cryptographic hash value of the semantic spatiotemporal unified dataset, encrypt and store the semantic spatiotemporal unified dataset, and record the data storage address; Construct evidence storage metadata, which includes the semantic spatiotemporal unified identifier, the cryptographic hash value, and the data storage address; The smart contract based on the blockchain network automatically verifies the stored metadata and initiates an on-chain request after the verification is successful. The blockchain network's consensus nodes verify the on-chain request. Once the verification is successful, the evidence storage transaction information is stored in the blockchain network, and a digital evidence storage certificate bound to the evidence storage transaction information is generated. The semantic spatiotemporal unified dataset with the digital evidence attached is used as the result of the multi-source heterogeneous data aggregation.
5. The method for merging and integrating multi-source heterogeneous data in the environmental governance industry according to claim 1, characterized in that, The construction of the environmental process mechanism model and probabilistic graphical model to obtain the environmental data tracing model includes: A mechanism model of the environmental process is established based on the migration and evolution laws of environmental elements; The probabilistic graphical model is constructed based on the anomaly propagation paths of the knowledge-enhanced environment dataset. The environmental process mechanism model and the probabilistic graphical model are coupled to obtain the environmental data tracing model.
6. A multi-source heterogeneous data aggregation and fusion system for the environmental governance industry, characterized in that, The system includes an input device, an output device, a processor, and a memory, which are interconnected. The memory stores a computer program, which includes program instructions. The processor is configured to invoke the program instructions to execute the multi-source heterogeneous data aggregation and fusion method for the environmental governance industry as described in any one of claims 1-5.
Citation Information
Patent Citations
Multi-source heterogeneous data knowledge graph construction method for railway disaster prevention monitoring
CN120492447A
Knowledge graph construction method and system based on large language model
CN120633803A