Deep Learning-Based Intelligent Analysis and Standardization System, Method, Electronic Equipment, and Storage Medium for Multi-Source Heterogeneous Data of Contaminated Sites
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2026-08-14
AI Technical Summary
[0011]本发明要解决的技术问题是现有污染场地数据处理过程中存在的数据处理效率低、格式标准不统一、人工处理错误率高、不同场地间风险评估结果难以标准化比较、技术导则更新不及时以及历史案例知识难以有效利用等缺陷,提出了一种基于深度学习的污染场地多源异构数据智能解析与标准化系统及方法,该系统通过集成深度学习文档理解模型、多模态信息提取技术、标准化风险表征算法和图数据库存储架构,实现了从非结构化原始资料到标准化数据的高效自动转换,为污染场地修复决策提供高质量、统一标准的数据基础
1、本发明通过深度学习文档理解模型实现了对多种格式污染场地报告的自动解析,数据提取效率提升5.8倍,关键信息提取F1值达到0.85,显著降低了人工处理成本和错误率;
Smart Images

Figure CN121095038B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of interdisciplinary technology of environmental engineering and artificial intelligence, and in particular to a system, method, electronic device and storage medium for intelligent analysis and standardization of multi-source heterogeneous data of contaminated sites based on deep learning. Background Technology
[0002] Rapid industrialization and urbanization have exacerbated the urgency of remediating contaminated sites, while the investigation, assessment, and remediation of contaminated sites generate a large amount of heterogeneous data from multiple sources. This data comes from a wide range of sources and takes various forms, including site investigation reports (PDF, Word, etc.), testing data sheets (Excel, CSV, etc.), risk assessment reports, technical guidance documents, and historical case data.
[0003] Currently, data processing for contaminated sites mainly relies on manual methods, which presents the following prominent problems: First, data processing efficiency is low, and manually extracting key information is time-consuming and labor-intensive; second, data from different sources has inconsistent formats, making it difficult to integrate and analyze; third, manual processing is prone to errors, and accuracy is difficult to guarantee; fourth, due to differences in exposure scenario construction, model application, and parameter selection, risk assessment results for different sites are difficult to standardize and compare; fifth, data standards need to be manually updated after technical guidelines are updated; and sixth, historical case knowledge is difficult to effectively utilize and transfer.
[0004] The following is a comparative analysis of existing technologies: 1. Commercial OCR systems (such as ABBYY FineReader, Kofax Transformation) While such systems can perform basic document recognition, they have significant limitations when processing specialized reports on contaminated sites: (1) The accuracy rate of specialized terminology recognition is low (usually only 50-65%), and they cannot recognize specialized pollutant names and their chemical formulas such as "Triphenylphosphine oxide" and "O,O,O-Triethylphosphorothioate". (2) The system has limited ability to handle complex table structures, with a recognition rate of less than 70% for merged cells and cross-page tables. (3) The system lacks the ability to understand the specific structure of contaminated site reports and cannot extract structured pollution information. (4) The system does not have the ability to correlate data across documents and cannot connect related sampling and analysis results.
[0005] 2. General data integration platforms (such as Informatica, Talend)
[0006] While these platforms offer data integration capabilities, their processing capabilities for contaminated site data are insufficient. (1) Lack of a dedicated knowledge model for contaminated sites, making it unable to understand industry-specific contexts. (2) Reliance on predefined rules for processing unstructured documents, resulting in poor adaptability. (3) Inability to handle cross-validation of multimodal data (text, charts, maps). Lack of a standardized processing mechanism for risk assessment models.
[0007] 3. Environmental data management system (such as EnviroData, EQuIS)
[0008] While these systems focus on environmental data management, they still have shortcomings in intelligent analysis and standardization: (1) They mainly rely on manual input or simple template matching, resulting in low automation. (2) They lack deep learning technology support and cannot handle historical reports with varying formats. (3) They do not support dynamic updates of guidelines, and standard parameters need to be maintained manually. (4) Their ability to standardize risk assessment results is limited, making it difficult to cross-site requirements.
[0009] While existing research has proposed several data processing methods, such as rule-based information extraction and simple OCR recognition, these methods have significant limitations when dealing with complex and diverse contaminated site data. Traditional OCR technology suffers from insufficient accuracy when processing professional reports containing tables and images; rule-based methods struggle to adapt to different report formats; existing data standardization methods have limited ability to recognize technical terms and parameters; and there is a lack of a unified integration framework for multi-source data. With the rapid development of deep learning and natural language processing technologies, especially the significant progress made by document understanding models based on the Transformer architecture in text analysis, table recognition, and image processing, new technical pathways have been provided to address the problem of processing multi-source heterogeneous data from contaminated sites. However, these advanced technologies have not yet been fully applied to data processing systems in the field of environmental engineering.
[0010] Therefore, there is an urgent need to develop a deep learning-based intelligent analysis and standardization system for multi-source heterogeneous data of contaminated sites, which can automatically and efficiently process raw data in various formats, extract key information, perform standardization processing, and construct a unified data structure to provide high-quality data support for subsequent contaminated site analysis and decision-making. Summary of the Invention
[0011] The technical problem this invention aims to solve is the shortcomings of existing contaminated site data processing, such as low data processing efficiency, inconsistent format standards, high error rate of manual processing, difficulty in standardizing and comparing risk assessment results between different sites, untimely updates of technical guidelines, and difficulty in effectively utilizing historical case knowledge. This invention proposes a deep learning-based intelligent analysis and standardization system and method for multi-source heterogeneous data of contaminated sites. This system integrates a deep learning document understanding model, multimodal information extraction technology, standardized risk characterization algorithms, and a graph database storage architecture to achieve efficient and automatic conversion from unstructured raw data to standardized data, providing a high-quality, unified standard data foundation for contaminated site remediation decisions.
[0012] The technical solution of this invention is implemented as follows: A deep learning-based intelligent analysis and standardization system for multi-source heterogeneous data from contaminated sites includes: The intelligent parser for site survey reports uses a deep learning-based document understanding model to parse site survey reports and test data in formats such as PDF, Word, and Excel. It automatically extracts key information such as pollutant concentration, sampling points, and site characteristics, and converts them into standardized structured data. The risk assessment data structuring processing unit is used to process risk assessment data from different sites or different stages of the same site. Through multi-source model mapping and standardized risk characterization algorithms, it unifies the output results of risk assessment and ensures the consistency and comparability of risk assessment results. The dynamic access mechanism for national and local technical guidelines is used to periodically and automatically detect and obtain the latest national and local technical guidelines. It extracts key parameters through a pre-trained language model, handles version differences and standard conflicts, and ensures that the standards and guidelines in the system are up-to-date. The historical remediation case database stores and manages cases of completed contaminated site remediation projects. It adopts a graph database architecture to achieve efficient case retrieval, similarity calculation, and reliability rating, providing reference decision support for new projects.
[0013] As a preferred technical solution, the intelligent parser for site survey reports is based on an improved BERT variant architecture for document understanding and uses multimodal information extraction technology to process text, tables and image content simultaneously. The key information includes pollutant names and CAS numbers, concentration values and units, sampling point coordinates, detection methods, site geological and hydrological characteristics, risk assessment parameters, etc.
[0014] As a preferred technical solution, the intelligent parser for site survey reports includes a document preprocessing submodule and a multimodal parsing submodule: the document preprocessing submodule combines the OpenCV image processing library to achieve unified conversion and image enhancement of documents of multiple formats; the multimodal parsing submodule uses the Transformer encoder to extract text features, performs table boundary detection and cell segmentation based on ResNet-50, and achieves multimodal feature fusion through an attention mechanism.
[0015] As a preferred technical solution, the risk assessment data structuring processing unit includes a multi-source model mapping module, a risk value standardization module, and an uncertainty quantification module, which are used to perform unified conversion of different risk assessment models, standardized score calculation, and risk uncertainty analysis functions, respectively.
[0016] As a preferred technical solution, the dynamic access mechanism for national and local technical guidelines includes: Network information extraction module: It uses the Scrapy framework's network information extraction technology to obtain the latest technical guidelines documents from the official website of the environmental protection department, thereby reducing the workload of manual monitoring; Parameter extraction and conflict resolution module: It uses a RoBERTa-based pre-trained language model to extract key parameters from the guidelines and handles parameter conflicts through a five-level priority strategy to improve the system's adaptability to the new standard. Knowledge graph update module: It updates the extracted parameter information to the system knowledge base, and ensures the continuous optimization and efficient operation of the system through an incremental update strategy.
[0017] A method for intelligent analysis and standardization of multi-source heterogeneous data from contaminated sites, utilizing the aforementioned deep learning-based intelligent analysis and standardization system for multi-source heterogeneous data from contaminated sites, includes the following steps: receiving various contaminated site investigation reports and monitoring data; performing multimodal document parsing based on a deep learning model; transmitting the extracted structured information to a risk assessment data processing unit; the risk assessment data processing unit unifying the risk assessment output results and standardizing risk scores; dynamically acquiring the latest technical guideline parameters, updating the system standard library, and retrieving similar cases from a historical case database to provide remediation solution recommendations.
[0018] As a preferred technical solution, the intelligent parser for site survey reports includes a document preprocessing submodule and a multimodal parsing submodule. The workflow of the document preprocessing submodule is as follows: input multi-format documents, perform format conversion according to document type, apply image enhancement algorithms to improve OCR recognition quality, and perform paragraph segmentation according to the structure of the contaminated site report. The workflow of the multimodal parsing submodule is as follows: input the preprocessed document content, extract text semantic features using a Transformer encoder, recognize table structures and extract data based on ResNet-50, fuse different types of features through a multimodal attention mechanism, and output standardized structured data.
[0019] As a preferred technical solution, the workflow of the multimodal document understanding model is as follows: input document content, perform word embedding and positional encoding, extract semantic features through a multi-head attention mechanism, apply specific task heads for entity recognition and relation extraction, and output pollutant information, sampling data, and confidence scores; the workflow of the risk assessment standardization module is as follows: input the original risk assessment results, identify the assessment model type, apply the corresponding mapping function for model transformation, perform risk value standardization calculation, and output a unified format risk score and uncertainty index; the workflow of the case retrieval module is as follows: input site features and pollution features, construct feature vectors, pre-screen candidate cases through local sensitive hashing, calculate weighted cosine similarity, combine reliability rating ranking, and output a list of similar cases and remediation plan suggestions.
[0020] A contaminated site data intelligent analysis device is used to perform an intelligent analysis and standardization method for multi-source heterogeneous data of contaminated sites, including: The storage device is used to store the original site survey report, monitoring data and analysis results, and supports high-concurrency queries and data retrieval. GPU servers provide computing resources for running deep learning models, enabling intelligent document parsing and data processing; Database: Stores standardized contaminated site data, risk assessment results, and historical case information, and supports quick queries; Interactive devices: provide a user-friendly interface and support custom data query, batch processing and result export functions.
[0021] A non-transitory storage medium for storing a program that enables a contaminated site data intelligent analysis device to perform the following action: execute a method for intelligent analysis and standardization of multi-source heterogeneous data from contaminated sites.
[0022] Compared with existing technologies, this solution has the following advantages: 1. This invention achieves automatic parsing of contaminated site reports in various formats through a deep learning document understanding model, improving data extraction efficiency by 5.8 times and achieving an F1 score of 0.85 for key information extraction, significantly reducing manual processing costs and error rates; 2. The standardized risk characterization algorithm designed in this invention enables unified characterization and comparison of results from different risk assessment models, providing an objective and consistent quantitative basis for risk assessment; 3. The dynamic access mechanism for technical guidelines developed in this invention ensures timely updates of standard parameters in the system, achieves a rule extraction accuracy rate of 92%, and solves the problem of delayed updates to technical guidelines. 4. The case graph database constructed by this invention uses multi-dimensional similarity calculation and reliability rating to achieve efficient retrieval and application of case knowledge, and improves the retrieval speed of similar cases by 15 times; 5. The microservice architecture design of this invention has good scalability and maintainability. In a standard deployment environment, the system availability reaches 97-99.5%, with an average response time of 2-5 seconds. Under high load, it can still maintain more than 90% availability of core functions. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a diagram of the overall system architecture of the present invention; Figure 2 This is a diagram of the intelligent parser architecture for the site survey report of this invention. Figure 3 This is a flowchart of the risk assessment data structuring processing unit of the present invention; Figure 4 This is a diagram illustrating the dynamic access mechanism architecture of the national and local technical guidelines for this invention. Figure 5 This is a diagram of the database structure for historical repair cases of this invention. Detailed Implementation
[0025] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0026] Reference Figure 1 This application proposes a deep learning-based intelligent analysis and standardization system for multi-source heterogeneous data from contaminated sites, comprising the following core components: an intelligent site investigation report parser, used to automate the processing of unstructured investigation reports using multimodal information extraction technology; a risk assessment data structuring unit, used to process risk assessment data from different sites or different stages of the same site, ensuring the consistency and comparability of risk assessment results; a dynamic access mechanism for national and local technical guidelines, used to ensure that standards and guidelines in the system remain up-to-date; and a historical remediation case database, used to store and manage completed remediation project cases, providing reference for new projects. The system adopts a modular microservice architecture, achieving efficient linkage and intelligent response across the entire process of data acquisition, analysis, standardization, and storage.
[0027] Unified terminology and symbol definitions: Ri: Represents the original risk value of a certain pollution indicator; R min R max These are the minimum and maximum limits for the risk indicator, respectively. Sstd: Standardization scale, defaults to 1; Smin: Minimum risk value after standardization, usually 0. Rstd: According to the formula The obtained standardized risk score.
[0028] (a) Intelligent parser for site survey report
[0029] The intelligent parser for site survey reports employs a deep learning-based document understanding model, whose structure includes an input layer, embedding layer, encoder layer, decoder layer, and output layer. This model handles mixed content in documents, using word embeddings and multi-head attention mechanisms, and features table structure recognition and key information extraction. Document preprocessing includes image enhancement and OCR recognition, employing neural networks for text recognition and professional terminology correction. Model training utilizes standard deep learning optimization methods and is performed using a suitable amount of labeled data. The intelligent parser for site survey reports employs a deep learning-based document understanding model, and its structure includes the following technical details: Detailed design of the model architecture: Basic framework: An improved BERT variant is adopted as the core architecture. This choice is based on its superior performance in bidirectional contextual understanding during the pre-training stage, and experimental results show that it improves the accuracy of contextual domain term recognition by 7-12% compared to RoBERTa and XLNet.
[0030] Model parameter configuration: Word embedding dimension: 768; Transformer layers: 6 encoder layers; Number of attention heads: 12 Feedforward network hidden layer dimension: 3072; activation function: GELU; maximum sequence length: 512 Vocabulary size: 30,522 basic vocabulary words + 5,000 environmental-specific vocabulary words Model improvements: 1) Add a domain adaptation layer: Add an environmental engineering-specific adaptation layer on top of the standard BERT. 2) Cross-modal attention mechanism: Design a three-modal attention fusion module for text, tables, and images. 3) Hierarchical document understanding: Design a three-level hierarchical understanding mechanism for documents, paragraphs, and sentences. 1.1 Text Processing Flow Algorithm 1.1.1 Word segmentation and word embedding Text segmentation is a fundamental step in converting a raw document into a processable sequence of tokens. (1) Where S is the token sequence after WordPiece word segmentation. Let be the i-th sub-word token, and n be the sequence length.
[0031] Word embedding maps discrete tokens to continuous vector representations: (2) Where E is the word embedding matrix, Let d be the d-dimensional word vector representation of the i-th token, where d is the embedding dimension (usually 768).
[0032] 1.1.2 Location Coding
[0033] Add positional information to each position in the sequence, using sine and cosine encoding: (3a) (3b) in, Let k be the position encoding vector for the i-th position, where k is the dimension index (k = 0, 1, ..., ...). d / 2 -1), 10000 is the frequency adjustment constant.
[0034] The input vector is obtained by adding word embeddings and positional encodings: (4) in, Let be the final input vector for the i-th position, containing vocabulary and positional information.
[0035] 1.1.3 Transformer Encoding
[0036] Single-head attention mechanism calculates the intra-sequence correlation: (5) Where Q, K, and V are the query, key, and value matrices, respectively, and d_k is the dimension of the key vector. Use a scaling factor to prevent gradient vanishing.
[0037] Multi-head attention enhances expressive power by computing multiple attention heads in parallel: (6) in, h represents the number of attention heads. , , Let be the projection matrix of the i-th head. This is for outputting the projection matrix.
[0038] The output representation of the L-layer Transformer encoder: (7) Where X is the input sequence, Θ is the parameter set of all layers, and H is the encoded hidden layer representation order. 1.1.4 Task-Specific Output Layer Document categories are indicated using the [CLS] tag: (8) in, Hidden layer representation for special classification labels, and For classification layer parameters, This represents the probability distribution for classification.
[0039] Named entity recognition labels each token: (9) in, Let be the hidden representation of the i-th token. and For entity recognition layer parameters, This represents the probability distribution of entity labels.
[0040] Relation extraction calculates the probability of relationships between entity pairs: (10) in, and Let represent two entities span, ⊕ is the concatenation operation, g is the pooling function, r is the relation type, and σ is the sigmoid function.
[0041] The final text features are obtained through pooling or concatenation tasks: (11) in, _text is the feature extraction function, which can be average pooling, max pooling, or [CLS].
[0042] 1.2 Table Recognition Algorithm: 1.2.1 Image Feature Extraction Extracting deep features from table images using ResNet-50: (12) in, To input a table image, F is the feature extraction function for ResNet-50, and F is the extracted feature map.
[0043] 1.2.2 Table Structure Detection
[0044] Table boundary detection uses a convolutional network to predict boundary probabilities: (13) in, Convolutional layers for boundary detection Here, σ is the bias term, and σ is the sigmoid activation function. This is a boundary probability diagram.
[0045] Cell semantic segmentation predicts the cell category for each pixel: (14) in, Cell-level convolutional layers For bias terms, Probability plots for each category (header, data cell, background, etc.).
[0046] 1.2.3 Cell Extraction
[0047] Generate a binary mask for the cells: (15) Where 1{·} is an indicator function, It is a binary mask for the cell range.
[0048] Extracting independent cells using connected component analysis: (16) Where CC stands for Connectivity Component Algorithm, and C is the set of extracted cells. This refers to the j-th cell range.
[0049] 1.2.4 OCR Recognition and Error Correction
[0050] Perform OCR recognition on each cell and calculate the confidence score: (17) in, For the recognized text, For confidence level, P(t|c j ) represents the text probability of a given cell image.
[0051] Error correction based on context information: (18) Where N(j) is the set of neighboring cells of cell j. For language model parameters, This is the corrected text.
[0052] 1.2.5 Table Structure
[0053] Construct a table structure diagram: G = (V, E), V = C, and E is determined by spatial adjacency (19). Where V is the set of cell nodes and E is the set of edges based on spatial location.
[0054] Identify header cells: (20) in, (c) represents cell features, w_h and b_h are header classifier parameters, τ_h is the threshold, and H is the set of header cells.
[0055] Generate structured tables: (twenty one) Here, ψ is a structured function that converts graph structures, text content, and header information into a standard table format.
[0056] 1.3 Multimodal Fusion Algorithm: 1.3.1 Feature Projection Alignment Projecting features from different modalities into a unified space: (22a) (22b) (22c) in, , , Features of the original text, tables, and images. , , For the projection matrix, , , This is a bias term.
[0057] 1.3.2 Cross-modal attention computation
[0058] Calculate the attention of text to the table (taking t←b as an example): (twenty three) in, , , These are the projection matrices for the query, key, and value, respectively.
[0059] Attention weight calculation: (twenty four) Where d is the feature dimension, and A_{t←b} is the attention weight matrix of the text to the table.
[0060] Similarly, define attention between other modalities: (25) Where the subscripts indicate the attention direction, t represents the text modality, b represents the table modality, and i represents the image modality.
[0061] 1.3.3 Feature Enhancement and Fusion
[0062] Enhanced features across modalities based on attention weights: (26) (27) (28) Where α, α', β, β', γ, and γ' are the fusion weight parameters. _、 , This refers to the enhanced features.
[0063] The final fused features are processed by a multilayer perceptron: (29) Where Pool represents pooling operation, ⊕ represents feature concatenation, and MLP represents multilayer perceptron.
[0064] 1.4 Specific methods for model training
[0065] 1.4.1 Training Dataset Construction: Source: 10,000 annotated contaminated site investigation reports, including 5,000 industrial site reports, 3,000 mining site reports, and 2,000 gas station reports; Annotation content: text paragraph type, key entities (pollutants, concentration values, sampling points, etc.), table structure, and image region classification; Data augmentation: report page rotation (±5°), contrast adjustment (±15%), random occlusion (5% of the area), and text synonym replacement (10% of technical terms). 1.4.2 Training Strategy Details: Pre-training phase: The masked language model was pre-trained using 50GB of literature in the field of environmental engineering; pre-training tasks: masked word prediction, document paragraph order prediction, and table cell relationship prediction; pre-training parameters: batch size 64, learning rate 5e. -5 Train for 8 epochs.
[0066] Fine-tuning phase: Fine-tuning for specific tasks: entity recognition, relation extraction, table recognition; Fine-tuning parameters: batch size 32, learning rate 3e. -5 Training for 10 epochs; Optimizer: AdamW, β1=0.9, β2=0.999, weight decay 0.01; Learning rate scheduling: linear warm-up for the first 10% of steps, followed by cosine decay.
[0067] Validation and Evaluation: Cross-validation to ensure model stability; Key metrics: F1 score for key information extraction, accuracy of table structure recognition, and accuracy of document type classification; Early stopping strategy: Stop if the F1 score on the validation set does not improve for 3 consecutive epochs.
[0068] Mathematical representation of model training: Total loss: (30) in, , , , Non-negative weights (can be normalized); , Losses are for entities and relationships, respectively. For table segmentation / boundary detection loss; For cross-modal / cross-task consistency regularization.
[0069] Physical loss: (31) Where i is the token position; c is the entity category; y_{i,c} is the one-hot tag; P(c| ) represents the predicted probability.
[0070] Relationship loss: (32) in, Let be the set of candidate entity pairs; r be the relation category; y{ij,r} be the one-hot key label; ,e For entity representation.
[0071] Tables and Consistency (Example): (33) (34) Where p is the pixel / grid position; k is the semantic category; and m is the text / table distribution pair for the same field.
[0072] Learning rate scheduling (linear warmup + cosine annealing): (35) Where t is the current training step; η_min is the peak LR; η_min is the floor LR; T_warmup is the number of warmup steps; T_total is the total number of steps.
[0073] 1.4.3 Domain-Specific Model Optimization: Professional terminology processing: Construction of an environmental engineering professional dictionary, extracting 10,000+ professional terms from national standards, technical guidelines, and academic literature; Vocabulary weight enhancement: Adjusting the weight of professional terms in the model to improve recognition priority; Synonym network: Constructing a synonym network for professional terms such as pollutant names and detection methods, and handling expression differences; Pollutant parameter identification optimization: Unit unification identification algorithm: automatically identifies and unifies different units of measurement (such as mg / kg, ppm, μg / g, etc.); Numerical format standardization: handles numerical expression issues such as scientific notation and differences in significant figures; Parameter association identification: establishes associations between pollutant names and their concentration values, detection limits, standard limits, etc. Inference performance optimization: Model compression: Knowledge distillation: Train a small model using the full model, maintaining 95% performance while reducing the number of parameters by 50%; Weight quantization: Use INT8 quantization, reducing the model size by 75% and improving inference speed by 2.3 times; Structural pruning: Remove attention heads that contribute less, retaining 95% performance while reducing computation by 30%; Inference acceleration: Batch processing optimization: Dynamic batch processing strategy to maximize GPU utilization; Caching mechanism: Cache results for common document fragments, and directly call duplicate content; Model deployment: TensorRT optimization, ONNX format conversion, and inference speed improved by 3.5 times.
[0074] (ii) Risk assessment data structuring processing unit
[0075] like Figure 3 As shown, the risk assessment data structuring unit uses a standardized risk characterization algorithm and includes the following steps: Based on multi-source model mapping function (36) Where Si represents the output of different risk assessment models, and T represents the unified internal representation; The risk value standardization function is defined as: (37) in, Standardized risk value; : The i-th risk value; , : Minimum and maximum risk values; , Standardized parameters.
[0076] The quantitative weighting of exposure pathways, the calculation of receptor sensitivity coefficients, and the quantification of risk uncertainty are all achieved using statistical simulation methods to ensure the accuracy and comparability of risk assessment results.
[0077] 2.1 Multi-source risk model mapping algorithm: 2.1.1 Model Unified Mapping Define the source model set and mapping function: (38) Where s: source model, S is the set of supported risk assessment models (such as HJ 25.3-2019, RBCA, CalTOX). : Output of source model s Output dimensions for the source model. To unify the target dimensions.
[0078] Model-specific mapping function: (39) in, is a specific mapping function for the source model s, which can be a linear transformation or a deep neural network.
[0079] Cross-model calibration parameters: (40) in, This is the calibration matrix (usually a diagonal matrix) for model s. This is the bias vector, used to eliminate systematic biases between models.
[0080] Unified risk characterization: (41) in, This is the standardized risk representation vector.
[0081] 2.1.2 Consistency Check and Correction
[0082] Rule consistency score: (42) Where κ is the consistency scoring function, τ is the acceptable threshold, and Π is the minimum correction projection operator to ensure that the results conform to physical constraints.
[0083] Reference alignment score (when a reference standard exists): (43) in, This is the reference value for the i-th term. For consistency score.
[0084] 2.2 Detailed algorithm for risk value standardization:
[0085] 2.2.1 Determination of Standardized Parameters
[0086] Determine standardized parameters based on risk type: (44) in, Risk type (carcinogenic risk, hazard index, etc.) and These are the lower and upper limits of the risk value. and This refers to the standardized range parameters.
[0087] 2.2.2 Piecewise Standardization Transformation
[0088] Standardization is achieved using piecewise functions: (45) in, The standardized i-th risk value. : The original i-th risk value, , : Minimum and maximum risk values Standardized minimum value Scope of standardization.
[0089] Equivalent closed expression: (46) Here, clip(x; a, b) = min(max(x, a), b) is the cutoff function.
[0090] 2.3 Algorithm for Quantitative Weighting of Exposure Pathways: 2.3.1 Calculation of Exposure Indicators Calculate three key exposure indicators: (47a) (47b) (47c) in, In order to expose the possibility, As the transfer factor, For the proportion of the exposed population, For the i-th exposure pathway, Due to site characteristics, For population distribution.
[0091] 2.3.2 Weight Calculation and Normalization
[0092] Calculate the unnormalized weights: (48) Where α, β, and γ are the weight coefficients of each indicator, satisfying α + β + γ = 1.
[0093] Normalized weights: (49) Where W_i is the final weight of the i-th path, satisfying nonnegativity and normalization constraints.
[0094] 2.4 Uncertainty Quantification Algorithm (LHS + Monte Carlo + Sobol)
[0095] 2.4.1 Latin hypercube sampling
[0096] LHS sampling is performed on the k-th dimension parameter: (50) Where π_k is the random permutation of the k-th dimension, N is the number of samples, and F_k^{-1} is the inverse cumulative distribution function of the k-th dimension parameter.
[0097] Generate complete sample vectors: (51) Where p is the parameter dimension, Let be the j-th sample vector.
[0098] 2.4.2 Risk Calculation and Statistical Analysis
[0099] Calculate the risk value for each sample: (52) Where R is the risk assessment function. Let be the risk output for the j-th sample.
[0100] Calculate basic statistics: (53a) (53b) (53c) in, s is the mean, s² is the variance, and CV is the coefficient of variation.
[0101] Calculate the 95% confidence interval: (54) in, This is the k-th sequential statistic.
[0102] 2.4.3 Sobol Sensitivity Analysis
[0103] Calculate the first-order and total effect sensitivity indices: (55a) (55b) in, Let i be the first-order sensitivity index of the i-th parameter. The overall effect sensitivity index, This represents all parameters except the i-th parameter.
[0104] (III) Dynamic Access Mechanism for National and Local Technical Guidelines
[0105] The dynamic access mechanism for national and local technical guidelines adopts a layered processing architecture, including a data acquisition layer, an information extraction layer, a knowledge fusion layer, and an application interface layer. The information extraction layer employs a parameter extraction algorithm based on a pre-trained language model in the relevant professional domain. Its main functions include: guideline text segmentation, parameter identification, parameter association, and parameter conflict detection and resolution. The system automatically checks periodically for the latest guideline updates to ensure that the standards and guidelines in the system are up-to-date.
[0106] 3.1. Technical Guideline Parameter Extraction Algorithm
[0107] 3.1.1 Document Preprocessing
[0108] Document format conversion and cleaning: (56) Where D represents the original guideline document, and Tika is the Apache Tika document transformation function. For text cleaning functions, This is the cleaned plain text.
[0109] Document segmentation: (57a) Seg_doc is a document-level segmentation function that segments based on chapter titles and structural tags.
[0110] (57b)
[0111] in, For set union operation, s represents the various chapters. This is a paragraph-level splitting function.
[0112] 3.1.2 Named Entity Recognition
[0113] Parameter identification using a pre-trained model: (58) Where M is a pre-trained language model (such as RoBERTa), and p is the paragraph text. and These are the parameters for the NER layer.
[0114] Extract entities with confidence scores exceeding a threshold: (59) Where e is the extracted entity, y is the entity label, and p is the paragraph text. Set the confidence threshold for entity recognition (usually 0.85).
[0115] 3.1.3 Parameter Structuring
[0116] Construct the parameter vector: (60) Where p is the parameter quadruple, value is the parameter value, unit is the unit of measurement, cond is the applicable condition, and range is the valid range of values.
[0117] Construct a parameter relationship graph: G = (V, E), where V = entity and parameter node; E = dependency / constraint / adjacency relationship (61) Where V is the set of nodes and E is the set of relation edges.
[0118] Relational reasoning closure: (62) in, G' is an inference operator that supplements remote dependencies with transitivity rules and domain knowledge, and G' is an inference closure graph.
[0119] Generate structured parameters: P = Φ(G'), verification V(P) = True (63) Where Φ is the transformation function from graph to structured parameters, and V is the consistency check function, which checks the logical consistency of dimensions and ranges.
[0120] 3.2 Conflict Detection and Resolution Algorithm: 3.2.1 Conflict Candidate Identification Identifying potential conflict parameters based on similarity: (64) in, For the new parameters, Given an existing parameter library, sim is the semantic similarity function. This is the similarity threshold (usually set to 0.8).
[0121] 3.2.2 Conflict Type Determination
[0122] Define the conflict detection predicate: (65) Here, ∨ represents the logical OR operation, ValueConflict detects numerical conflicts, and ScopeConflict detects conflicts in the scope of application.
[0123] 3.2.3 Priority Scoring
[0124] Parameter priority in calculation: (66) Among them, Level is the hierarchical weight (national > provincial > municipal > industry > project), Recency is the timeliness weight, and Authority is the authority weight of the issuing authority.
[0125] 3.2.4 Conflict Resolution Strategies
[0126] Determine the appropriate action based on priority differences: (67) Where Action: parameter conflict resolution action, Priority: parameter priority function. , : New and old parameters, δ: Priority difference threshold.
[0127] 3.2.5 Parameter Library Update
[0128] Update the parameter library based on the solution strategy: (68) in, The updated parameter library , : Existing parameter library and new parameter set, Ω is the parameter library update operator, which performs replacement, retention, merging or pending operation according to the Action strategy set.
[0129] (iv) Historical Restoration Case Database
[0130] like Figure 5 As shown, the historical restoration case database adopts a graph database storage architecture, where nodes represent case entities and edges represent relationships between cases. Case similarity is calculated using a multidimensional weighted cosine similarity formula: ∑αi (69) Among them, C x and C y For two cases, V xi and V yi Let αi represent the vector representations of the two cases on the i-th feature, and αi be the weight of the i-th feature. The case reliability rating employs a multi-factor comprehensive scoring system, considering factors such as technological maturity, data completeness, effectiveness verification, and independent verification to ensure the quality and usability of the case library. The detailed implementation of the historical restoration case database includes: 4.1 Graph Database Architecture Design Core Engine: Employs Neo4j 4.4 graph database, configured with an 8-core CPU, 64GB of memory, and 2TB of SSD storage; Indexing Strategy: Creates composite indexes for node attributes to improve query performance, with key field index coverage >95%; Caching Configuration: 32GB heap cache, 16GB page cache, 8GB query cache, with a hot data cache hit rate >90%; Sharding Strategy: Stores data in shards based on pollution type to reduce cross-shard queries and improve local performance.
[0131] 4.2 Detailed Design of Data Model: Node types and attributes: Case node: contains 25 core attributes such as ID, name, time, location, area, and total cost; Pollutant node: contains 15 attributes including name, CAS number, category, and physical and chemical properties; Technology node: Includes 18 attributes such as name, category, principle, and applicable conditions; SiteFeature node: includes 12 attributes such as soil type and hydrological characteristics; Effect evaluation node: includes 10 attributes such as removal rate, compliance rate, and processing cost; Relationship types and attributes: APPLIED_IN (Technology → Cases): Includes attributes such as application time, scale, and parameter settings; TREATS (Technology → Pollutant): Includes attributes such as treatment efficiency and applicable concentration range; CONTAINS (Case Study → Pollutants): Includes attributes such as pollution concentration and distribution characteristics; HAS_FEATURE(Case → Site Features): Describes the specific site features of the case. RESULTED_IN (Case Study → Effect Evaluation): Records the repair effect and evaluation indicators; INFLUENCES (Site Characteristics → Effect Evaluation): Describes the impact of site characteristics on the effect.
[0132] 4.3 Case Vectorization Algorithm
[0133] 4.3.1 Pollution Feature Extraction (15 Dimensions)
[0134] Distribution characteristics of pollutant types: (70) in, It is a 5-dimensional simplex. The proportion of pollutants of type k. Let be the quantity of pollutant of type k.
[0135] Pollution concentration level distribution: (71) in The pollution concentration level distribution vector is used to divide the concentration into 5 levels (extremely low, low, medium, high, and extremely high) according to a preset threshold.
[0136] Spatial distribution characteristics: (72) in, It is a spatial distribution feature vector, containing five spatial statistics, including Moran's I spatial autocorrelation index and semivariogram features.
[0137] Pollution feature vector: (73) Where ⊕ represents the vector concatenation operation, This is a 15-dimensional pollution feature vector.
[0138] 4.3.2 Site Feature Extraction (12-dimensional) (74) in, Soil feature extraction function (4-dimensional). This is a hydrological feature extraction function (4-dimensional). This is a land use feature extraction function (4-dimensional).
[0139] 4.3.3 Extraction of technical parameters (20 dimensions) (75) in, For technical feature vectors, It is a 6-dimensional technology type feature. It is an 8-dimensional technical parameter feature. It features 6-dimensional cost efficiency characteristics.
[0140] 4.3.4 Complete Vector Construction
[0141] All features are concatenated to form a 47-dimensional vector: (76) Standardization process: (77) Where μ and σ are the mean and standard deviation of each dimension.
[0142] Similarity calculation optimization algorithm: 4.4 Similarity Calculation and LSH Pre-screening Algorithm 4.4.1 Locality Sensitive Hashing Define L hash tables, each using m random hyperplanes: (78) in, Here, k is the hash table index, and k is the hyperplane index. Let be the normal vector of the random hyperplane, and sign be the sign function. Let v be the binary representation of vector v in the hyperplane.
[0143] Constructing hash codes and candidate sets: (79a) Where concat_k is the bit-string concatenation operation along the k-th dimension, h_ (v) is the first The hash code of each table.
[0144] (79b)
[0145] in, For set union operations, _ For the first The hash table contains buckets, where V_target is the target case vector and Cand is the candidate case set.
[0146] 4.4.2 Weighted Similarity Calculation
[0147] Calculate the weighted cosine similarity: (80) Where W = diag(w j ) is the weight diagonal matrix, w j Let be the importance weight of the j-th dimension feature.
[0148] 4.4.3 Reliability Adjustment and Prioritization
[0149] The final score is calculated based on the reliability of the case studies: (81) Where λ is a monotonically increasing reliability weighting function. For the reliability rating of case i, For the target case, These are candidate cases.
[0150] Filter and sort the results and return them: (82) in, is the similarity threshold, k is the number of returned cases, Sort is the descending sorting function, and [:k] is to take the first k elements.
[0151] 4.5 Case Reliability Rating Algorithm
[0152] 4.5.1 Technology Readiness Score (83) in, For technology maturity assessment functions, The number of successful application cases, For the years of technology application, This represents the degree of proof-of-principle verification.
[0153] 4.5.2 Data Completeness Score (84) in, For the core field completeness rate, To monitor data integrity, For the completeness of technical parameters.
[0154] 4.5.3 Effectiveness Verification Scoring (85) in, For the quality of the results data, To achieve the target rate, For longer-lasting effects.
[0155] 4.5.4 Independent Validation Scoring (86) in, To verify the independence of the parties, To verify the authority of the party, To verify the standardization of the process.
[0156] 4.5.5 Weighted Total Score and Rating (87) Overall score Technical score Data scoring Effectiveness rating : Indicator score, 0.3, 0.25, 0.25, 0.2: Weight coefficient of each scoring item (the sum of weights is 1.0).
[0157] 4.6 Case Knowledge Extraction and Automated Management: Semi-automatic case construction process: Document parsing stage: Automatically extract basic information using a site survey report intelligent parser; Knowledge extraction stage: Construct core case knowledge using professional terminology identification and relationship extraction; Case template filling: Automatically map extracted information to predefined case templates, with a completion rate of >85%; Manual review and supplementation: Experts review the extraction results, supplement missing information, and confirm data accuracy; Quality control: Ensure cases meet the quality standards for inclusion in the database through 20 quality inspection rules.
[0158] Algorithm for automatically summarizing lessons learned: 4.6.1 Feature Extraction and Scoring Extracting key features from the case text: Features = Extract(Text, K) (88) Where Features are the extracted feature vectors, Text is the case text, and K is the domain knowledge base.
[0159] Calculate the importance of factors: (89) Among them, Frequency is the frequency of occurrence, Impact is the degree of influence, and Novelty is the novelty level.
[0160] 4.6.2 Confidence Level and Classification
[0161] Calculate the confidence level of the lesson: Confidence ) = P( |evidence) · Reliability(source) (90) Wherein, P( |evidence) represents the probability of a lesson learned based on evidence, and Reliability(source) represents the reliability of the source.
[0162] Lesson type classification: (91) in, For lessons learned, τ represents the type of lesson (success factors, challenges, applicable conditions, improvement suggestions), and P represents the classification probability.
[0163] 4.6.3 Merging and Sorting
[0164] Merging similar lessons: (92) in, The relationship is an equivalence relation, and Sim is a semantic similarity function. This is the merging threshold (usually set to 0.9).
[0165] Sort by importance: (93) Where L is the list of lessons learned, Sort is the sorting function, key is the sorting key, and desc indicates descending order.
[0166] (v) Data Integration and Quality Control Module
[0167] This module ensures data flow and quality assurance between various system components, and is the foundation for the normal operation of the entire system.
[0168] 5.1 Standardized Data Interface
[0169] Unified data exchange format: Adopts JSON-LD format, supports semantic annotation; defines core schema to ensure data structure consistency; version control mechanism to support smooth interface upgrades.
[0170] 5.2 Multi-source data standardization transformation algorithm: 5.2.1 Format Detection and Mapping Automatically identify data source format: (94) Among them, Format The optimal format for identification is defined by f, where f is the candidate format type, features is the data feature extraction function, and P is the format classification probability.
[0171] Define field mapping relationships: (95) Where M is the field mapping function, For the source field set, For the target field set.
[0172] 5.2.2 Value Transformation and Unit Unification
[0173] Perform value conversion and unit standardization: (96) in, For value conversion functions, For the collection of source field values, For unit conversion functions, This is the final standardized value.
[0174] 5.2.3 Semantic annotation and verification
[0175] Add semantic annotations: (97)
[0176] Establish the relationship and add traceability information: (98a) (98b) in, The data is semantically annotated. For the target data, Schema is the semantic annotation pattern. For the data after establishing the relationship, The data after adding traceability information is src, which is the data source information, actor, which is the operator information, and time, which is the timestamp information.
[0177] 5.2.4 Data Validation and Correction
[0178] Perform constraint verification: (99) Where Valid(D) is the overall validity judgment function, D is the data to be validated, and ∧ is the logical AND operation. Let be the i-th constraint, and Valid be the overall validity judgment function.
[0179] Identify and correct errors: Viol(D) = Violation data set (100a) (100b) Output standardized data: (101) Where Viol(D) is the set of non-compliant data. The data is now corrected; AutoCorrect is the automatic correction function. For data with traceability information, For the final standardized data.
[0180] 5.3 Data Quality Control Mechanism
[0181] Completeness assessment: calculation of missing rates for key fields; data coverage assessment (time and spatial dimensions); data quality scoring to guide users to prioritize the addition of important data.
[0182] Automatic consistency checks: cross-validation of cross-source data; validation of logical rules (numerical range, unit consistency, etc.); validation of time series data trends.
[0183] Outlier detection: based on statistical methods such as Z-score and interquartile range; based on domain rules such as the reasonable range defined by professional knowledge; and based on machine learning methods such as isolated forest and autoencoder.
[0184] Data repair suggestions: Smart fill suggestions: based on similar data and pattern recognition; Outlier handling options: replace, mark, or retain; Data augmentation techniques: improve the usability of sparse data.
[0185] Mathematical representation of data quality control: Integrity check: (102) Where I(·) is the indicator function, Let i be the i-th data field, and |Fields| be the total number of fields.
[0186] Accuracy verification: (103) Among them, Valid( Let |Values| be the validity judgment function for the i-th value, and let |Values| be the total number of values.
[0187] Consistency check: (104) Among them, Compatible ( ) is the function for determining the compatibility of numeric pairs, where |Pairs| is the number of all possible numeric pairs.
[0188] Outlier detection: (105) Where x is the value to be detected, μ is the sample mean, and σ is the sample standard deviation. This is the threshold for anomaly detection.
[0189] Overall quality rating: (106) in, , , , The weights for each quality indicator are denoted by , and OutlierRate represents the proportion of outliers.
[0190] (vi) System deployment and performance optimization
[0191] To ensure the system's high performance, high availability, and scalability, this invention employs a modern microservice architecture for deployment and optimization.
[0192] (vii) Case Studies of Practical Applications
[0193] The following two typical cases demonstrate the analytical capabilities and standardization effects of the deep learning-based intelligent analysis and standardization system for multi-source heterogeneous data of contaminated sites when processing data from different sources and in different formats.
[0194] Example 1: Intelligent Analysis and Standardization of Multi-Source Heterogeneous Data from Heavy Metal Contaminated Mining Sites
[0195] Input details of multi-source heterogeneous data: (1) Preliminary investigation report: PDF format, 156 pages in total, including site overview, historical background, preliminary sampling analysis, etc.; (2) Detailed investigation report: Word format, 298 pages in total, including data tables of 156 soil samples, 18 groundwater samples, geological data of 53 boreholes, hydrogeological survey, pumping test data, etc. (3) Risk assessment report: PDF format, 342 pages in total, including detailed calculation process of HJ 25.3-2019 model, sensitivity analysis, Monte Carlo simulation results, etc.; (4) Soil physicochemical property test data: 12 Excel spreadsheets, including moisture content 20.5–27.1%, natural unit weight 13.83–16.09 kN / m³, void ratio 0.650–0.921, and permeability coefficient (vertical) 7.62×1 ~6.32×1 Parameters such as cm / s; (5) Raw laboratory test data: 15 CSV format data files, including ICP-MS mass spectrometry, AAS chromatograms and other analytical results; (6) National and local technical guidelines: 8 PDF documents, including the "Soil Environmental Quality Construction Land Soil Pollution Risk Control Standard", local implementation rules, etc. (7) Historical remediation case data: 23 PDF documents, including remediation technical solutions and effect evaluation reports for similar heavy metal contaminated sites.
[0196] Detailed process of intelligent system analysis and standardized application: (1) Application of core technologies in intelligent parsing of site survey reports: a. Deep learning document understanding model architecture operation: The input layer receives a total of 796 pages of documents in PDF, Word, and Excel formats. The embedding layer converts the text into 768-dimensional word embedding vectors and adds positional encoding to process long document sequences. The encoder layer consists of 6 Transformer encoders, each containing a 12-head self-attention mechanism, to deeply extract semantic features from the documents. The decoder layer performs specific decoding for three tasks: text classification, entity recognition, and relation extraction. The output layer generates standardized structured data and identifies 312 professional terms from an environmental engineering thesaurus, such as "acidic mine wastewater," "heavy metal speciation analysis," "permeability coefficient anisotropy," and "vadose zone thickness," with an accuracy rate of 96.8%.
[0197] b. In-depth application of multimodal information extraction technology: Text processing employed WordPiece word segmentation and extracted structured text features using a Transformer encoder. Table processing utilized ResNet-50 for table boundary detection and cell segmentation, handling 38 complex merged cell tables and 15 cross-page tables. Combined with OCR recognition and contextual error correction, the cell recognition accuracy reached 97.1%, correcting 23 errors in scientific notation. Image processing used ResNet50 to extract features from 45 borehole columnar sections, 12 geological profiles, and 8 contour maps, achieving multi-scale feature fusion through a Feature Pyramid Network (FPN). Multimodal feature fusion achieved text features through projection alignment and attention weighting. Table features and image features The fusion outputs fusion features. This enables cross-modal information consistency verification.
[0198] c. Intelligent extraction and structuring of key information: Automatic extraction of pollutant concentration data: copper content 23–8,950 mg / kg (68 samples exceeded the standard, with the largest exceeding the standard by 17.9 times), cadmium content 0.08–45.6 mg / kg (32 samples exceeded the standard, with the largest exceeding the standard by 91 times), lead content 15.2–2,890 mg / kg (41 samples exceeded the standard), etc., with a data extraction completeness rate of 98.7%; extraction of soil physicochemical property parameters: key parameters such as water content, natural density, porosity, and permeability coefficient, with a parameter identification accuracy rate of 97.3%; extraction of groundwater parameters: water level depth 8.2m, aquifer thickness 12.8m, and permeability coefficient 1.2×10⁻⁶. -4 cm / s, hydrochemical type HC -Ca·Mg type, etc.; 2,456 sets of quaternary relationships of pollutant-concentration-location-depth were established, and the F1 value of key information extraction reached 0.85.
[0199] (2) Core applications of the risk assessment data structuring processing unit: a. Application of multi-source risk model mapping algorithm: The system uses the mapping function M:{ , ,…, →T converts the output of the HJ 25.3-2019 risk assessment model into a unified internal representation, supporting multiple risk assessment models such as RBCA and CalTOX; automatically extracts 28 exposure parameters: soil intake rate (100 mg / d for adults, 200 mg / d for children), exposure frequency (350 d / a), exposure period (24 years for adults, 6 years for children), body weight and other key parameters.
[0200] b. Application of standardized risk characterization algorithms: Standardized scores were calculated using the risk value standardization formula, where Rmin and Rmax are the lower and upper limits of the risk value, Sstd=1, and Smin=0. The risk values of the five heavy metals were uniformly mapped to the [0,1] interval: copper 0.823, cadmium 0.946, lead 0.715, zinc 0.687, and arsenic 0.542. The exposure pathway weights were assigned to the overall exposure probability E. Transfer factor T and the proportion of exposed population P Calculate weights The weightings for direct soil ingestion are 0.42, groundwater ingestion is 0.28, skin contact is 0.18, and particulate inhalation is 0.12.
[0201] c. Application of risk uncertainty quantification algorithms: Monte Carlo simulation combined with Latin hypercube sampling was used to generate 10,000 parameter samples for uncertainty analysis. The output 95% confidence intervals were: a comprehensive risk value of [0.756, 0.891], and individual heavy metal confidence intervals of [0.789, 0.856] for copper and [0.889, 0.998] for cadmium. Sobol sensitivity analysis identified soil uptake rate as the most sensitive parameter, with a first-order sensitivity index of 0.412 and a transfer factor sensitivity index of 0.328, providing a scientific basis for risk management optimization.
[0202] (3) Core applications of the dynamic access mechanism in national and local technical guidelines: a. Intelligent crawling at the data acquisition layer: Using Scrapy-based network information extraction technology, eight of the latest guidelines were automatically obtained from environmental protection department websites, supporting multiple levels of websites including the Ministry of Ecology and Environment and provincial environmental protection departments; the guidelines were regularly monitored for updates, and the latest revised version of the "Soil Environmental Quality Standard for Construction Land Soil Pollution Risk Control" (GB36600-2018) and local implementation rules were automatically downloaded.
[0203] b. Application of the RoBERTa model in the information extraction layer: Using a RoBERTa-based pre-trained language model, 67 key parameters, including copper screening values of 18,000 mg / kg and cadmium screening values of 65 mg / kg, were extracted through named entity recognition technology. The parameter extraction F1 score reached 93%, and the important update of adjusting the cadmium screening value from 65 mg / kg to 20 mg / kg was automatically recognized.
[0204] c. Conflict handling at the knowledge fusion layer: To handle differences between old and new versions and conflicts between different levels of guidelines, a five-level priority strategy was designed: National Standards (weight 0.4) > Industry Standards (0.3) > Local Standards (0.2) > Technical Specifications (0.1); 15 parameter conflicts were automatically resolved with a conflict resolution accuracy of 96.7%; the application interface layer provides standard parameter query APIs and rule verification interfaces, and supports RESTful APIs, WebSocket notifications, and GraphQL interfaces.
[0205] (4) Application of historical restoration case database graph storage: a. Neo4j graph database architecture retrieval: Neo4j was used to store case nodes, pollutant nodes, technology nodes, site feature nodes, and effect evaluation nodes. Similar cases of heavy metal contaminated mining areas were retrieved from 3,247 stored historical cases. Case nodes contain 25 core attributes such as ID, name, time, location, and area, while pollutant nodes contain 15 attributes such as name, CAS number, category, and physicochemical properties.
[0206] b. Case vectorization representation and LSH retrieval: The cases were vectorized into 47-dimensional feature vectors: 15-dimensional pollution features (distribution of heavy metal types, concentration levels, pH values, etc.), 12-dimensional site features (soil type, hydrological characteristics, geological conditions, etc.), and 20-dimensional technical parameters (remediation methods, cost-efficiency, time cycles, etc.). Locality Sensitive Hashing (LSH) and hierarchical indexing were used to accelerate the query. 89 candidate cases were pre-screened, and 23 cases with a similarity >0.85 were identified by calculating weighted cosine similarity.
[0207] c. Reliability rating based on a five-level scoring system: The quality of the cases was evaluated using a five-level scoring system: S / A / B / C / D. There were 8 S-level cases (high technical maturity, high data completeness, sufficient effect verification, and authoritative independent verification), 11 A-level cases, and 4 B-level cases. Based on the analysis of similar cases, the success rate of stabilization / solidification technology application was 92% (21 / 23 cases), and the success rate of phytoremediation + soil improvement was 87% (20 / 23 cases). The system recommends the combined technical solution of "zonal treatment + stabilization and solidification + phytoremediation".
[0208] (5) Core applications of the data integration and quality control module: a. Standardization and conversion of JSON-LD format data: Define unified naming and unit specifications for more than 500 core fields, and convert data from 7 different sources into JSON-LD format; establish a pollutant standard dictionary: copper / Copper / Cu are uniformly named, concentration units are uniformly mg / kg, density units are uniformly kN / m³, and coordinate system is uniformly the National 2000 coordinate system; field mapping accuracy is 98.2%, and data standardization completeness is 98.7%.
[0209] b. Multi-dimensional quality control algorithm: Integrity checks revealed a 1.3% missing rate for core fields, a 0.8% missing rate for pollutant concentration data, and a 0.5% missing rate for coordinate information. Accuracy verification passed logical rule validation, identifying 3 orders-of-magnitude errors in the permeability coefficient and 5 abnormal density values. Consistency checks found 15 data conflicts across reports, and the verification of coordinate deviations of the same sampling point being less than 2m passed, with automatic consistency checks achieving over 95%. Outlier detection employed a Z-score statistical method combined with an isolated forest machine learning method, identifying 23 outliers with a detection accuracy of 93.5% and a false alarm rate of 6.2%.
[0210] c. Version control and source tracing based on Git principles: Establish a field-level traceability chain to record all 156 data modification histories, including complete information such as data source, processing algorithm, quality assessment, operator, and timestamp; support data version rollback and change auditing, with a traceability chain integrity rate of 99.2%; and 100% version control coverage to ensure traceability throughout the entire data processing process.
[0211] Application effect evaluation: Data parsing efficiency was improved by 5.8 times, the F1 score for key information extraction reached 0.85, and the error rate was reduced by 75%. A unified and standardized dataset containing four core dimensions—pollution, geology, hydrology, and risk—was established. The accuracy rate of multimodal information fusion was 96.3%, and the accuracy rate of professional terminology recognition was 96.8%. Case retrieval matched 23 similar cases, and the accuracy rate of remediation technology recommendation reached 95.8%. The overall data quality score was 0.954, providing a high-quality and standardized data foundation for contaminated site remediation decisions.
[0212] Example 2: Intelligent Analysis and Standardization of Complex Pollution Data from Organic Chemical Sites
[0213] Input details of multi-source heterogeneous data: (1) Preliminary investigation report: PDF format, 156 pages in total, including the historical production situation of a certain pesticide chemical plant, site survey, etc.; (2) Detailed investigation report: Word format, 312 pages in total, including test data of 218 soil samples, 24 groundwater samples, geological data of 67 boreholes, hydrogeological survey, pumping test data, etc. (3) Risk assessment report: PDF format, 387 pages in total, including analysis of the synergistic effects of compound pollution, multi-scenario exposure assessment, Monte Carlo uncertainty analysis, sensitivity analysis, etc. (4) Soil physicochemical property test data: 15 Excel spreadsheets, including organic matter 1.0–3.5%, average 2.03%, pH value 6.24–8.91, average 7.45, and permeability coefficient 7.45×10 -6 ~6.40×10-5 Parameters such as cm / s; (5) Raw laboratory test data: 18 CSV format data files, including GC-MS and LC-MS detection spectral data, quality control sample results, etc. (6) National and local technical guidelines: 12 PDF documents, including the "Implementation Plan for Groundwater Pollution Prevention and Control" and the "Technical Guidelines for Soil Pollution Risk Assessment of Construction Land". (7) Historical remediation case data: 31 PDF documents, including case studies of remediation technologies for organic chemical contaminated sites, effect assessments, experience summaries, etc.
[0214] Detailed process of intelligent system analysis and standardized application: (1) Intelligent parser for site survey reports: complex document processing. a. Deep applications of deep learning document understanding model architecture: The input layer receives a total of 855 pages of complex technical documents; the embedding layer converts chemical engineering text into 768-dimensional word embedding vectors and processes special character sequences such as chemical formulas and CAS numbers; the 6-layer Transformer encoder identifies 412 chemical engineering terms through a 12-head self-attention mechanism, including "chloroaromatic biodegradation", "cometabolite pathway", "volatile organic compound migration", and "bioavailability", with a term recognition accuracy of 97.1%; the decoder layer performs specific decoding for tasks such as chemical entity recognition, chemical reaction relationship extraction, and toxicity classification; the output layer automatically constructs a 589-node knowledge graph of pollutants-CAS numbers-molecular formulas-toxicity parameters.
[0215] b. Complex applications of multimodal information extraction technology: Text processing utilizes WordPiece word segmentation to process chemical names, and a Transformer encoder to extract complex semantic features such as chemical reaction mechanisms and process flows. Table processing uses ResNet-50 to identify 156 complex tables containing chemical structural formulas, achieving a cell segmentation accuracy of 95.7%. OCR combined with context correction identifies 67 chemical molecular formulas and 218 CAS numbers, achieving an accuracy of 94.3%. Image processing uses ResNet50+FPN to process 78 GC-MS spectra, 45 LC-MS spectra, and 23 process flow diagrams, automatically annotating key information such as retention time, peak area, and quantitative ions. Multimodal feature fusion achieves consistency verification of "chemical description-detection data-spectral information," with a fusion accuracy of 96.8%.
[0216] c. Intelligent extraction of key information on complex pollution: Automatically extract organic pollutant concentration data: Chlorobenzene content < LOR~15,600 mg / kg (45 samples exceeded the standard, with the maximum exceeding 56.78 times), nitrobenzene content < LOR~894 mg / kg (32 samples exceeded the standard, with the maximum exceeding 10.76 times), aniline content < LOR~2,350 mg / kg (28 samples exceeded the standard, with the maximum exceeding 8.04 times), etc.; establish 2,847 groups of five - element relationships of pollutant - concentration - location - time - detection method; extract 589 groups of physicochemical property parameters: molecular weight, melting and boiling points, vapor pressure, water solubility, LogKow, etc.; the F1 value of key information extraction reaches 0.91, which is 38% higher than the traditional method.
[0217] (2) Structured processing unit for risk assessment data - Application in combined pollution: a. Mapping and integration of multi - source risk models: The system simultaneously uses the HJ 25.3 - 2019 and RBCA models to identify risk assessment reports, and uniformly represents the outputs of different models through the mapping function M; the carcinogenic risk output by the HJ model is 1.8×10 -4 and the hazard index is 2.34, and the total risk output by the RBCA model is 2.1×10 -4 and the hazard quotient is 2.67; after unified standardization, the comprehensive carcinogenic risk score is 0.876 (indicating a very high risk, close to the highest value of 1.0), and the non - carcinogenic risk score is 0.834; support the collaborative integration of CalTOX and the US EPA model.
[0218] b. Standardized risk characterization of combined pollution: For the combined pollution of chlorobenzene + nitrobenzene + aniline, the toxicity equivalent factor (TEF) method is used to quantify the synergistic effect; the standardized risk characterization algorithm calculates the risk values of each pollutant: chlorobenzene 0.834, nitrobenzene 0.892, aniline 0.756, 2,4 - dichlorophenol 0.923 (all are 0 - 1 standardized scores, the larger the value, the higher the risk); assign weights to exposure pathways: direct soil ingestion 0.35, groundwater ingestion 0.28, volatile inhalation 0.22, skin contact 0.15 (the sum of weights is 1.0); the toxicity enhancement factor of combined pollution is 1.23, and the accuracy of synergistic effect assessment is 94%.
[0219] c. In - depth analysis of risk uncertainty quantification: Monte Carlo simulation combined with Latin hypercube sampling generates 15,000 parameter samples, covering 47 key variables such as physicochemical parameters, toxicity parameters, and exposure parameters; output the 95% confidence interval: combined carcinogenic risk [l.6×10 -4 , 2.2×10 -4 (far exceeding the acceptable level of 10 -6), non-carcinogenic risk [2.1, 2.7] (exceeding the safety threshold of 1.0); Sobol sensitivity analysis identified key sensitive parameters: soil-groundwater partition coefficient (sensitivity index 0.385), degradation half-life (0.312), and soil uptake rate (0.267).
[0220] (3) In-depth application of dynamic access mechanisms in national and local technical guidelines: a. Intelligent acquisition and parsing of multi-level guidelines: The Scrapy framework automatically retrieves 12 of the latest guidelines, totaling 534 pages, including the Ministry of Ecology and Environment's "Implementation Plan for Groundwater Pollution Prevention and Control" and provincial-level "Technical Guidelines for Soil Pollution Risk Assessment of Construction Land." It monitors guideline updates in real time and supports national and regional environmental protection department websites. It also automatically identifies key changes such as adjustments to organic pollutant screening values and updates to detection methods.
[0221] b. Intelligent parameter extraction from the RoBERTa model: Based on the RoBERTa model pre-trained in the field of chemistry, 78 screening values for organic pollutants, 45 control values, and 23 groundwater quality standards were extracted through named entity recognition. Key parameters in the GB36600-2018 standard were identified: chlorobenzene screening value of 270 mg / kg, nitrobenzene screening value of 76 mg / kg, aniline screening value of 260 mg / kg, and 2,4-dichlorophenol screening value of 843 mg / kg. The F1 score of parameter extraction reached 93%, which is 15% higher than that of general models.
[0222] c. Five-level priority conflict resolution mechanism: 34 parameter differences between national and local standards were addressed; priority scoring algorithm: latest national standard (weight 0.4) > industry standard (0.3) > local standard (0.2) > technical specification (0.1) > project standard (0.05); automatic conflict resolution strategy: national standard priority 0.95, local implementation rules 0.73, industry technical specification 0.58; 67 parameter database items were updated, affecting the results of 89 sample points exceeding the standard.
[0223] (4) Complex retrieval applications of historical restoration case database: a. Neo4j graph database complex pollution case management: Organic chemical pollution cases were retrieved from 3,247 stored cases. The graph database includes case nodes, pollutant nodes, technology nodes, site feature nodes, and effect evaluation nodes. Through graph structure association analysis, 36 cases of "chlorinated aromatic hydrocarbons + nitro compounds" compound pollution pattern and 28 cases of "volatile + semi-volatile" combination pollution were identified.
[0224] b. 47-dimensional case vectorization and LSH intelligent retrieval: The 15-dimensional pollution characteristics highlight the synergistic effects of complex pollution, volatility, biodegradability, and toxicity levels; the 12-dimensional site characteristics emphasize groundwater pollution depth, soil-groundwater interaction, distribution of sensitive receptors, and geological permeability; the 20-dimensional technical parameters cover in-situ oxidation, bioremediation, physical barrier, and combined technologies; the local sensitive hash LSH pre-screening efficiency is improved by 15 times, and weighted cosine similarity is calculated from candidate cases to identify 26 highly similar cases (similarity > 0.8).
[0225] c. Success rate analysis of the five-level scoring system: Based on the S / A / B / C / D five-level scoring system: 9 S-level cases (complete data, sufficient effect verification, and mature technology), 12 A-level cases, and 5 B-level cases; statistical analysis of technical success rates: in-situ chemical oxidation single technology had a success rate of 89% (23 / 26 cases), bioremediation single technology had a success rate of 73% (19 / 26 cases), and combined technology had a success rate of 92% (24 / 26 cases); combined with reliability ratings, a phased remediation strategy of "chemical oxidation + bioreinforcement" is recommended.
[0226] (5) In-depth application of data integration and quality control modules: a. JSON-LD format chemical semantic annotation: 723 dedicated fields for organic pollutants were defined, and a complete correspondence between CAS number, SMILES code, molecular formula, Chinese and English names, physicochemical properties, and toxicity parameters was established; a standardized nomenclature system for chemical substances was implemented: Chlorobenzene / Cl / 108-90-7, Nitrobenzene N / 98-95-3、Aniline / N / 62-53-3 Unified expression; supports semantic annotation of chemical ontology, and is compatible with international chemical database standards such as ChEBI and PubChem.
[0227] b. Multi-dimensional quality control deep algorithm: Completeness check: Key field missing rate 0.8%, quality control sample data coverage 98.7%, chemical structure information completeness rate 97.3%; Accuracy verification: Verified through chemometric rules, 12 concentration level errors, 6 detection method mismatches, and 3 CAS number errors were identified; Consistency check: Soil-groundwater partition coefficient conforms to Kd model prediction, correlation coefficient R²=0.89, Henry constant and volatility verification consistency 96%; Outlier detection: Isolation forest algorithm combined with chemical domain knowledge rules, detection accuracy 91.7%, effectively identifying detection anomalies, data entry errors, etc.
[0228] c. Field-level traceability and chemical data version control: Establish a dedicated traceability chain for chemical data: original testing → quality control verification → unit conversion → standardized naming → toxicity association → final storage; record 2,847 data modification history entries, including testing batches, algorithm versions, parameter settings, operation time, reviewers, etc.; support molecular-level data traceability and chemical reaction path reproduction, with 100% version control coverage.
[0229] Application effect evaluation: The system processes 855 pages of complex chemical engineering documents, improving multimodal parsing efficiency by 7.3 times, achieving a 97.1% accuracy rate in identifying chemical terminology and a 94.3% accuracy rate in identifying chemical structures. Standardization of compound pollution risk achieves unified characterization with HJ 25.3-2019 and the RBCA model, with a 94% accuracy rate in quantifying synergistic effects. Dynamic updates to technical guidelines ensure timely synchronization of 78 organic pollutant parameters. Case retrieval matches 26 similar compound pollution cases, with a 92% confidence level for technical recommendations. A standardized dataset encompassing four dimensions—time, space, chemistry, and toxicology—is established, achieving a comprehensive data quality score of 0.964, providing a high-quality, standardized data foundation for remediation decisions at compound pollution sites.
[0230] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A deep learning-based intelligent analysis and standardization system for multi-source heterogeneous data of contaminated sites, characterized in that, include: The intelligent parser for site survey reports uses multimodal information extraction technology to automate the processing of unstructured survey reports, including a document understanding model based on deep learning. The risk assessment data structuring unit processes risk assessment data from different sites or stages, and ensures the consistency and comparability of risk assessment results through standardized risk characterization algorithms. The national and local technical guidelines dynamic access module adopts a layered processing architecture to periodically detect and extract the latest technical guidelines parameters. A historical restoration case database is used, employing a graph database storage architecture to calculate case similarity and reliability rating. It also includes a data integration and quality control module, which is used to realize data flow and quality assurance between various subsystems of the system; The document understanding model of the intelligent parser for the site survey report includes: The input layer is used to receive documents in various formats; The embedding layer converts the text into 768-dimensional word embedding vectors and adds positional encoding; The encoder layer consists of 6 Transformer encoders, each containing 12 self-attention mechanisms. The decoder layer performs specific decoding for different tasks such as text classification, entity recognition, and relation extraction; The output layer is used to generate standardized structured data; The multimodal processing of the document understanding model specifically includes: (a) Text processing: WordPiece word segmentation was used, and structured text features were extracted using the Transformer encoder; (b) Table processing: ResNet-50-based table boundary detection and cell segmentation, combined with OCR recognition and context correction; (c) Image processing: ResNet50 is used to extract features, and the Feature Pyramid Network (FPN) is used to achieve multi-scale feature fusion; (d) Multimodal feature fusion processing: Text features V_text, table features V_table, and image features V_image are fused through projection alignment and attention weighting, and the fused feature V_fused is output. The risk assessment data structuring unit employs a standardized risk characterization algorithm, specifically including the following steps: Step S1: Multi-source risk model mapping: based on mapping function M:{S1,S2,…, →T converts the outputs of different risk assessment models into a unified internal representation, supporting HJ 25.3-2019, RBCA, and CalTOX risk assessment models; Step S2: Risk Value Standardization: According to the formula Calculate standardized scores, where... The lower and upper limits of the risk value. =1, =0; Step S3: Weighting of exposure pathways: Calculate the weight Wᵢ by combining the exposure probability EPᵢ, the transfer factor TFᵢ, and the proportion of exposed population PEᵢ; Step S4: Quantification of risk uncertainty: Monte Carlo simulation combined with Latin hypercube sampling is used to output confidence intervals and Sobol sensitivity index.
2. The intelligent analysis and standardization system for multi-source heterogeneous data of contaminated sites based on deep learning as described in claim 1, characterized in that, The dynamic access module for national and local technical guidelines includes: The data acquisition layer uses Scrapy framework-based web information extraction technology to obtain the latest guidelines from environmental protection department websites, supporting national and regional environmental protection department websites. The information extraction layer employs a pre-trained language model based on RoBERTa, extracting key parameters through named entity recognition technology, achieving an F1 score of 93%. The knowledge fusion layer handles differences between old and new versions and conflicts between guidelines at different levels, and designs a five-level priority strategy to automatically resolve conflicts. The application interface layer provides standard parameter query APIs and rule validation interfaces, and supports RESTful APIs, WebSocket notifications, and GraphQL interfaces.
3. The intelligent analysis and standardization system for multi-source heterogeneous data of contaminated sites based on deep learning as described in claim 1, characterized in that, The historical remediation case database adopts a graph database architecture, using Neo4j to store case nodes, pollutant nodes, technology nodes, site feature nodes, and effect evaluation nodes; and vectorizes the cases, including 15-dimensional pollution features, 12-dimensional site features, and 20-dimensional technical parameters; it uses Local Sensitive Hash (LSH) and hierarchical indexing to accelerate queries; and it uses a five-level scoring system (S / A / B / C / D) to evaluate case quality.
4. The intelligent analysis and standardization system for multi-source heterogeneous data of contaminated sites based on deep learning as described in claim 1, characterized in that, The data integration and quality control module is used to achieve: a) Standardized conversion of JSON-LD format data, defining unified naming and unit specifications for more than 500 core fields; b) Multi-dimensional quality control: integrity check, accuracy verification, consistency check and outlier detection, with automatic consistency check reaching over 95%; c) Version control and field-level source chain tracing based on Git principles, recording all data modification history and the operators.
5. A method for intelligent analysis and standardization of multi-source heterogeneous data from contaminated sites based on deep learning, characterized in that, The system for intelligent analysis and standardization of multi-source heterogeneous data from contaminated sites based on deep learning, as described in any one of claims 1 to 4, includes the following steps: Step S1: Parse unstructured site survey reports using a multimodal document understanding model to extract pollutant concentrations, sampling locations, and site characteristics; Step S2: Use a standardized risk characterization algorithm to unify the results of different risk assessment models and output a risk score with confidence intervals; Step S3: Update the dynamic access technology guidelines and maintain the parameter library through conflict detection and resolution strategies; Step S4: Retrieve similar historical cases based on graph databases and recommend remediation solutions based on reliability ratings.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, When the processor executes the program, it implements the steps of the method as described in claim 5.
7. A non-transitory computer-readable storage medium storing a computer program, characterized in that, When the program is executed by the system as described in any one of claims 1 to 4, it implements the steps of the method as described in claim 5.
Citation Information
Patent Citations
Data management system for unifying multi-modal data structured processing based on GPT model
CN117235276A
Dynamic data pipeline construction method based on artificial intelligence and multi-modal data processing
CN119830200A