A method and system for constructing a conceptual model of a contaminated site

By integrating multi-source data and information decoupling techniques, key information gaps in the conceptual model of contaminated sites are identified and optimized, addressing the shortcomings of traditional methods in contaminated site management. This enables dynamic adjustment of the model and comprehensive consideration of stakeholder needs, thereby improving the model's adaptability and accuracy.

CN120408994BActive Publication Date: 2025-10-17BEIJING MUNICIPAL RES INST OF ENVIRONMENT PROTECTION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510509194.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-10-17
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

Traditional contaminated site management methods struggle to systematically depict the complex relationships between pollution sources, migration pathways, and receptors, leading to remediation plans that deviate from actual needs or waste resources. Furthermore, they lack comprehensive consideration of the needs of different stakeholders, and models are ill-suited to adapt to constantly changing realities and new requirements.

Method used

An initial conceptual model of the contaminated site was constructed by integrating multi-source heterogeneous data. Information decoupling technology was used to decompose the parameter matrix of the site prior information model into independent units. Natural language processing technology and semantic query algorithms were used to identify key information gaps such as exposure pathways not covered in the site prior information, and the contaminated site conceptual model was iteratively optimized.

Benefits of technology

It enables intelligent identification of key information gaps in the conceptual model of contaminated sites, guides the dynamic adjustment and improvement of the model, makes it more in line with the actual situation and the needs of stakeholders, and improves the model's guiding value and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408994B_ABST
    Figure CN120408994B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of contaminated site model construction, and particularly discloses a construction method and system of a contaminated site conceptual model, which integrates multi-source heterogeneous data of a contaminated site to construct an initial contaminated site conceptual model, describes the space-time coupling relationship of a pollution source, a migration path and a receptor in the form of a parameter matrix, then further decomposes the site prior information model parameter matrix into independent units, takes feedback information of stakeholders as a constraint condition, uses natural language processing technology and a semantic query algorithm, performs panoramic semantic scanning and query on the site prior information to identify key information gaps such as an exposure path not covered in the site prior information, and iteratively optimizes the contaminated site conceptual model. In this way, intelligent identification of key information gaps of the contaminated site conceptual model can be realized, dynamic adjustment and improvement of the conceptual model can be guided, and the conceptual model can be made more in line with actual conditions and the needs of stakeholders.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of contaminated site modeling, and more particularly, to a method and system for constructing a contaminated site conceptual model. BACKGROUND

[0002] With the accelerated global industrialization and urbanization, soil and groundwater pollution has become a major challenge to the ecological environment and public safety. Traditional contaminated site management relies on experience or fragmented data, which is difficult to systematically depict the complex relationship between pollution sources, migration paths and receptors, leading to deviation of remediation schemes from actual needs or waste of resources. The international environmental governance system requires a dynamic and structured contaminated site conceptual model (CSM) to integrate the three-dimensional spatiotemporal migration of pollutants and exposure risks, providing a scientific basis for environmental decision-making. Its core value lies in achieving "pollution traceability, risk quantification, and measure implementation", thereby balancing environmental governance, economic cost, and social fairness under complex site conditions.

[0003] At the same time, the concealment of contaminated sites, the incompleteness of historical data, and the diverse demands of stakeholders, put higher requirements on the completeness of the conceptual model. The differentiated needs of stakeholders (such as government, enterprises, and residents) (such as health risk, remediation cost, and land reuse) are essentially coupled constraints of the social-environmental system. However, in the process of site planning and design, traditional methods often lack comprehensive consideration of the needs of different stakeholders, leading to difficulties in fully meeting the requirements of all parties in the site conceptual model. At the same time, once the model is determined, there is a lack of effective iterative updating mechanism, making it difficult to adapt to changing actual conditions and new demands. For example, a conceptual model constructed under the condition of insufficient site information may overgeneralize the spatial location of pollutants, ignore key pollution migration paths or receptor exposure risks, etc., resulting in a large difference between the constructed conceptual model and the actual contaminated site, reducing its guiding value in actual environmental decision-making.

[0004] Therefore, an optimized method and system for constructing a contaminated site conceptual model are expected. SUMMARY

[0005] To solve the above technical problems, the present application is proposed. Embodiments of the present application provide a method and system for constructing a contaminated site conceptual model, which integrates multi-source heterogeneous data of a contaminated site to construct a structured initial contaminated site conceptual model, uses a site prior information model parameter matrix to depict the spatiotemporal coupling relationship of a pollution source-migration pathway-receptor, then further decomposes the site prior information model parameter matrix into independent units using information decoupling technology, and uses feedback information of stakeholders as a constraint condition, uses natural language processing technology and semantic query algorithm, and performs panoramic semantic scanning and querying on the site prior information to identify key information gaps such as uncovered exposure pathways in the site prior information, so as to iteratively optimize the contaminated site conceptual model. In this way, intelligent identification of key information gaps of the contaminated site conceptual model can be realized, and dynamic adjustment and improvement of the conceptual model can be guided, so that the conceptual model is more in line with the actual situation and the needs of the stakeholders.

[0006] Correspondingly, according to one aspect of the present application, a method for constructing a contaminated site conceptual model is provided, which comprises:

[0007] constructing an initial contaminated site conceptual model, the initial contaminated site conceptual model being used to describe known or presumed pollution sources, possible pollution migration pathways, and potential exposure pathways and receptors;

[0008] collecting feedback information provided by stakeholders, the feedback information including demands, expectations and concerns of the stakeholders on the site;

[0009] extracting a site prior information model parameter matrix from the initial contaminated site conceptual model, and performing information decoupling on the site prior information model parameter matrix according to a row vector to obtain a set of site prior information model parameter row vectors;

[0010] based on the feedback information provided by the stakeholders, performing site prior semantic panoramic scanning on the set of site prior information model parameter row vectors to determine key information gaps existing in the initial contaminated site conceptual model.

[0011] According to another aspect of the present application, a system for constructing a contaminated site conceptual model is provided, which comprises:

[0012] a contaminated site conceptual model construction module, configured to construct an initial contaminated site conceptual model, the initial contaminated site conceptual model being used to describe known or presumed pollution sources, possible pollution migration pathways, and potential exposure pathways and receptors;

[0013] a feedback information collection module, configured to collect feedback information provided by stakeholders, the feedback information including demands, expectations and concerns of the stakeholders on the site;

[0014] information decoupling module, configured to extract a site priori information model parameter matrix from the initial contaminated site conceptual model, and perform information decoupling on the site priori information model parameter matrix according to a row vector to obtain a set of site priori information model parameter row vectors;

[0015] key information gap identification module, configured to perform site priori semantic panoramic scanning on the set of site priori information model parameter row vectors based on the feedback information provided by the stakeholders to determine a key information gap existing in the initial contaminated site conceptual model.

[0016] Compared with the prior art, the construction method and system of the contaminated site conceptual model provided in the application can construct a structured initial contaminated site conceptual model by integrating multi-source heterogeneous data of the contaminated site, describe the spatiotemporal coupling relationship of the pollution source-migration pathway-receptor by the site priori information model parameter matrix, then further decompose the site priori information model parameter matrix into independent units by using the information decoupling technology, take the feedback information of the stakeholders as a constraint condition, and use the natural language processing technology and the semantic query algorithm to identify a key information gap such as an exposure pathway not covered in the site priori information by performing panoramic semantic scanning and querying on the site priori information, so as to iteratively optimize the contaminated site conceptual model. In this way, the intelligent identification of the key information gap of the contaminated site conceptual model can be realized, the dynamic adjustment and improvement of the conceptual model can be guided, and the conceptual model can be made more consistent with the actual situation and the needs of the stakeholders. BRIEF DESCRIPTION OF DRAWINGS

[0017] The above and other objects, features and advantages of the present application will become more apparent from the following detailed description of embodiments of the present application, taken in conjunction with the accompanying drawings. The drawings provided in the appended hereto are intended to explain further to those skilled in the art the principles of the present application and constitute a part of the specification, which serve to explain the present application together with the embodiments of the present application, but do not constitute a limitation on the present application. In the drawings, the same reference numerals generally designate the same components or steps throughout the drawings.

[0018] Figure 1 A flowchart of the construction method of the contaminated site conceptual model according to the embodiments of the present application.

[0019] Figure 2 A data flow diagram of the construction method of the contaminated site conceptual model according to the embodiments of the present application.

[0020] Figure 3 A flowchart of step S4 in the construction method of the contaminated site conceptual model according to the embodiments of the present application.

[0021] Figure 4 A flowchart of step S42 in the construction method of the contaminated site conceptual model according to the embodiments of the present application.

[0022] Figure 5 A block diagram of a system for constructing a conceptual model of a contaminated site according to an embodiment of the present application. DETAILED DESCRIPTION

[0023] Hereinafter, example embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part but not all of the embodiments of the present application. It should be understood that the present application is not limited to the described embodiments.

[0024] Figure 1 A flow chart of a method for constructing a conceptual model of a contaminated site according to an embodiment of the present application. Figure 2 A data flow diagram of a method for constructing a conceptual model of a contaminated site according to an embodiment of the present application. As shown in Figure 1 and Figure 2 The method for constructing a conceptual model of a contaminated site according to an embodiment of the present application comprises the steps of: S1, constructing an initial conceptual model of a contaminated site, the initial conceptual model of a contaminated site being used to describe known or presumed pollution sources, possible migration paths of pollutants, and potential exposure paths and receptors; S2, collecting feedback information provided by stakeholders, the feedback information including requirements, expectations and concerns of the site; S3, extracting a site prior information model parameter matrix from the initial conceptual model of a contaminated site, and decoupling information of the site prior information model parameter matrix according to a row vector to obtain a set of site prior information model parameter row vectors; S4, based on the feedback information provided by the stakeholders, performing a site prior semantic panoramic scan on the set of site prior information model parameter row vectors to determine key information gaps existing in the initial conceptual model of a contaminated site.

[0025] In the method for constructing the conceptual model of the contaminated site, the step S1 of constructing an initial contaminated site conceptual model is used to describe known or presumed pollution sources, possible pollution migration pathways, and potential exposure pathways and receptors. It should be understood that the complexity and concealment of the contaminated site make it difficult for a single data source to comprehensively describe the pollution distribution and migration rules. Traditional methods rely on scattered data or empirical assumptions, which easily ignore key pollution sources or exposure pathways. Therefore, the present application proposes to integrate multi-source heterogeneous data, such as geological survey data, environmental monitoring data, historical pollution records, etc., to construct an initial contaminated site conceptual model, to structurally express the spatio-temporal coupling relationship of pollution sources, migration pathways and receptors, and to provide a unified framework for subsequent analysis. Among them, the pollution source refers to the source of the pollutants in the site, which may include waste discharge in industrial production process, chemical leakage, garbage landfill, etc. For example, a former chemical plant site, various chemical raw materials and intermediates used and stored in the production process may be potential pollution sources. The migration pathway is used to describe the migration mode and path of the pollutants in the site environment, including the diffusion, penetration, convection, etc. of the pollutants in the soil, groundwater, air, etc. For example, heavy metals in the soil may migrate downward with the leaching action of rainwater into groundwater; volatile organic compounds may enter the atmospheric environment through volatilization, and then diffuse with air flow. The receptor refers to the object that may be exposed to the pollutants, including humans (such as residents living around the site, personnel working in the site), ecosystems (such as animals and plants in the site, surrounding water ecosystem), etc. For example, residents living near the contaminated site may be exposed to pollutants through breathing contaminated air, drinking contaminated groundwater, or eating crops planted in contaminated soil, thereby causing potential harm to their health.

[0026] In the process of building the initial conceptual model of contaminated sites, it is crucial to comprehensively and systematically describe known or suspected sources of pollution, possible migration pathways of pollutants, and potential exposure pathways and receptors. This involves the collection, processing, and integration of multi-source data. Specifically, first, the collection of multi-source data is the foundation of building the model, which requires extensive collection of various types of data related to contaminated sites. Among them, multi-source remote sensing image data is an important source. Through high-resolution satellite remote sensing or aerial remote sensing, macroscopic image information of the site and its surrounding area can be obtained. By using specific image processing software and algorithms, remote sensing images can be interpreted to identify land use types within the site, such as industrial land, agricultural land, residential land, etc. It can also find areas that may have signs of pollution, such as color anomalies, abnormal vegetation growth areas, etc., which may indicate potential sources of pollution. For example, in a former industrial cluster, if the remote sensing image shows that a certain plot has a significant color difference from the surrounding area and the vegetation coverage is sparse, further investigation may reveal that the plot is contaminated with heavy metals due to the discharge of heavy metal-containing waste during past industrial production.

[0027] Secondly, grid data is also indispensable, which can accurately represent the spatial distribution of the site's topography and related environmental parameters. Through topographic surveying equipment, the site's elevation data can be obtained, and a digital elevation model (DEM) can be generated after processing. From the DEM, the site's topographic changes can be visually understood. Topography has an important influence on the migration of pollutants. In low-lying areas, pollutants tend to accumulate, while in areas with steep slopes, pollutants may quickly migrate with water flow. In addition, grid data can also include information such as soil type and groundwater level. Different soil types have different adsorption and transmission capabilities for pollutants. Clay soils have strong adsorption for heavy metals, while sandy soils are more conducive to the diffusion of pollutants. Through analysis of grid data, the possible migration direction and range of pollutants under different topographic and soil conditions can be initially determined.

[0028] Contaminated site vector data is also an important data type, which can clearly define site boundaries, infrastructure layout, and other information. Specifically, with the help of geographic information system (GIS) technology, vector data such as site boundaries, underground pipeline distribution, building locations, etc. can be entered into the system to build the site's spatial topology, which is of great significance for analyzing the migration path of pollutants within the site and the distribution of potential receptors. For example, understanding the direction of underground pipelines can help determine whether pollutants may leak and spread through the pipelines; knowing the location of buildings can help determine whether personnel living or working in the site are potential receptors.

[0029] In addition to the structured data mentioned above, the textual descriptions, data charts, and maps in the detailed investigation report also contain rich information. For the textual descriptions, optical character recognition (OCR) technology is used to convert the text in paper documents into editable text, and then natural language processing techniques are used for semantic analysis. Key information such as the historical use of the site, past pollution incidents, types of pollutants, and emission records can be extracted from the text. Data charts may contain pollutant monitoring concentration data at different time points. By organizing and analyzing these data, we can understand the trend of pollutant concentration over time and speculate on the development of pollution. Maps may include geological cross-sections of the site, hydrogeological maps, etc. From the geological cross-sections, we can obtain information about the stratigraphic structure of the site and determine the blocking or promoting effect of different strata on the migration of pollutants. Hydrogeological maps can show the flow direction and hydraulic gradient of groundwater, providing a basis for analyzing the migration of pollutants in groundwater.

[0030] After data collection, preprocessing is required. Specifically, data cleaning techniques are used to remove noise, errors, and duplicate data from the data. For example, in pollutant monitoring data, there may be abnormal values caused by instrument failure. By setting a reasonable data range and statistical analysis method, these abnormal values can be identified and removed. Standardization of data is performed to unify data format and units. Different sources of data may use different units, such as meters and feet for length. Through unit conversion, they are unified into international standard units, facilitating subsequent analysis and integration. Next, geographic information system (GIS) technology, data mining algorithms, and other tools are used for data integration and analysis. GIS technology has powerful spatial analysis capabilities and plays an important role in this process. The preprocessed multi-source data is imported into the GIS system, and different types of data are fused through spatial overlay analysis. For example, the suspected pollution area obtained by interpreting remote sensing images is overlaid with the geological map of the site to analyze the influence of the geological conditions of the area on pollutants. The groundwater level data is overlaid with the site topographic data to determine the relationship between the flow direction of groundwater and the terrain, and then to infer the migration path of pollutants in groundwater. Data mining algorithms are used to discover potential patterns and rules from large amounts of data. For pollutant monitoring data, clustering algorithms can be used to divide the site into different pollution areas, with similar pollutant concentrations and types in each area. Association rule mining algorithms can analyze the association between different pollutants, such as the simultaneous occurrence of certain pollutants, which may indicate the same pollution source. Through decision tree algorithms, the migration trend of pollutants under different conditions can be predicted based on various characteristic parameters of the site, such as soil type, terrain slope, groundwater level, etc.

[0031] Finally, based on the above data processing and analysis results, an initial contaminated site conceptual model is constructed. Specifically, when determining the location of the pollution source, multiple data sources are integrated, such as through remote sensing image interpretation, historical pollution records, and field investigation, to accurately locate the geographic coordinates of the pollution source. For the type of pollutants, the types of pollutants present in the site are determined based on chemical analysis results, enterprise production records, etc., including heavy metals, organic matter, inorganic matter, etc. When speculating the possible migration path, the migration process of pollutants in soil, water, air, etc. is simulated by combining topography, hydrogeological conditions, and the rules obtained through data mining analysis. When determining the distribution of potential receptors, population-dense areas such as residential areas, enterprises, schools, and ecologically sensitive areas such as nature reserves, rivers, and lakes are considered, and the personnel and ecosystems in these areas are determined as potential receptors. Finally, the constructed initial contaminated site conceptual model is presented in the form of graphics, charts, or text. For example, a two-dimensional or three-dimensional map is generated using a GIS system to visually display the location of the pollution source, the migration path, and the distribution of potential receptors. The types, concentrations, and trends over time of the pollutants are presented in chart form. Through detailed textual explanations, the basis for model construction, the relationships between various elements, and the limitations of the model are explained. In addition, different colored icons on the map can represent different types of pollution sources, and arrows can represent the direction of pollutant migration; bar charts can be used to display the concentration of pollutants in different areas, and line charts can be used to display the change in pollutant concentration over time; in the text explanation, it is explained which data and techniques the model is based on, and the uncertainty of the model due to incomplete data. Through the above methods, the constructed initial contaminated site conceptual model can accurately describe known or speculated pollution sources, possible pollutant migration pathways, and potential exposure pathways and receptors, providing a solid foundation for subsequent in-depth analysis of site pollution conditions, risk assessment, and development of remediation strategies.

[0032] In the method for constructing the conceptual model of the contaminated site, the step S2 of collecting feedback information provided by the stakeholders includes demands, expectations and concerns of the site. It can be understood that different stakeholders (including government departments, enterprises and residents, etc.) have different demands for the contaminated site based on their own role positioning and interest demands. For example, the government department usually concerns about the risk of the site to public health and the strict implementation of environmental regulations, and expects to protect the social public interest through effective site management; the enterprise may be more concerned about the control of site remediation cost and the feasibility of land reuse, so as to maximize the economic benefit; the residents are highly concerned about the safety of living environment and the life quality not to be affected. However, the traditional method often ignores the comprehensive consideration of the demands of multiple stakeholders when constructing the site conceptual model. Therefore, the present application further collects the demands, expectations and concerns of the site of each relevant stakeholder through interviews, questionnaire surveys, symposiums and the like, reveals the complex coupling constraint relationship between the social-environmental system, and provides a clear direction for optimizing the conceptual model based on the feedback information of the stakeholders.

[0033] In the method for constructing the conceptual model of the contaminated site described above, the step S3 is to extract a site prior information model parameter matrix from the initial conceptual model of the contaminated site, and to decouple information of the site prior information model parameter matrix according to a row vector to obtain a set of site prior information model parameter row vectors. Specifically, although the initial conceptual model of the contaminated site has integrated multi-source data, the information therein is complex and interwoven, which is not conducive to effective comparison and analysis with the feedback information of the stakeholders. Therefore, in order to further analyze the site prior information in the initial conceptual model of the contaminated site, identify the differences between the initial conceptual model of the contaminated site and the actual demand, and find the key information gap, the complex pollution-related information in the initial conceptual model of the contaminated site is abstractly represented in the form of a mathematical matrix to form a site prior information model parameter matrix, so as to clearly depict the spatiotemporal coupling relationship among the pollution source, the migration pathway and the receptor. Specifically, first, according to the pollution-related information of the initial conceptual model of the contaminated site, the key parameters that can accurately describe the model are comprehensively sorted out and determined. For example, for the pollution source, the three-dimensional spatial coordinates (X, Y, Z coordinates) are determined, as well as the parameters such as the type of the pollutant, the initial concentration, the release rate, etc.; for the migration pathway, the geometric parameters such as the length, the width and the slope are measured or estimated, as well as the physical and chemical property parameters of the medium (such as soil, water body, air) such as the soil porosity, the water flow rate, the air diffusion coefficient, etc., which will affect the migration rate and direction of the pollutant; for the receptor, the type (such as human beings, animals and plants, etc.), the specific location coordinates and the exposure mode and frequency that may be affected by the pollution, etc. are determined. Then, the determined parameters are organized into a two-dimensional matrix according to the logical order to obtain the site prior information model parameter matrix. For example, each row of the site prior information model parameter matrix can represent a specific pollution scenario or element, and each column corresponds to a specific parameter. Then, the information decoupling technology is used to decompose the site prior information model parameter matrix according to the row vector, so that each row vector represents a relatively independent pollution information unit, which describes the specific location of the pollution source, the type and concentration of the pollutant, the geometric characteristics and physical and chemical conditions of the migration pathway, and the exposure risk of the potential receptor, etc. in a numerical way, thereby forming the set of site prior information model parameter row vectors. In this way, not only the readability and operability of the information are enhanced, but also the subsequent identification of the information gap and the optimization of the model are facilitated.

[0034] In the method for constructing the conceptual model of the contaminated site described above, the step S4 is to perform a site prior semantic panoramic scanning on the set of site prior information model parameter row vectors based on the feedback information provided by the stakeholders to determine the key information gap existing in the initial conceptual model of the contaminated site. Wherein, Figure 3The flow chart is for step S4 in the method for constructing a conceptual model of a contaminated site according to an embodiment of the present application. As shown in Figure 3 S41, using a text encoder to textually structure code the feedback information provided by the stakeholders to obtain a feedback information semantic coding vector; S42, inputting the feedback information semantic coding vector and the set of site prior information model parameter row vectors into a feedback information-model prior semantic panoramic scanning network to obtain a feedback information-site prior semantic scanning response coding vector; and S43, feature decoding the feedback information-site prior semantic scanning response coding vector to obtain the identification result of the key information gap.

[0035] Specifically, the step S41 uses a text encoder to textually structure code the feedback information provided by the stakeholders to obtain a feedback information semantic coding vector. That is, considering that the feedback information provided by the stakeholders is mostly in the form of natural language text, such an unstructured data format is quite different from the structured information represented by the site prior information model parameter matrix, and it is difficult to directly compare and analyze them effectively. Therefore, in order to compare the two in the same semantic space and identify the key information gap, the present application further uses a text encoder to convert the feedback information provided by the stakeholders into an encoded form compatible with the model information. In one specific example of the present application, the text encoder is a pre-trained language model based on the Transformer architecture. It should be understood that the Transformer architecture can efficiently handle complex semantic relationships in natural language text by virtue of its powerful self-attention mechanism, capturing long-distance dependency relationships between different words and sentences in the text, thereby more accurately understanding the semantics of the text and converting it into a corresponding vector representation. By inputting the feedback information into the text encoder, the present application utilizes the powerful semantic understanding and representation capabilities of the Transformer architecture to convert the natural language text in the feedback information into a vector representation in a high-dimensional semantic space through steps such as word segmentation processing, word embedding, position coding, and self-attention mechanism, obtaining a feedback information semantic coding vector to express the needs, expectations, and concerns of the stakeholders for the site, providing a solid semantic foundation for subsequent information gap identification and model optimization.

[0036] Specifically, the step S42 inputs the feedback information semantic encoding vector and the set of site prior information model parameter row vectors into a feedback information-model prior semantic panoramic scanning network to obtain a feedback information-site prior semantic scanning response encoding vector. It should be understood that an initial contaminated site conceptual model is constructed based on known or presumed physicochemical data (such as pollutant concentration, geological parameters), but is limited by incomplete data, ambiguous historical information or technical assumptions simplification (such as homogeneous medium assumption), and may miss hidden pollution sources (such as uninvestigated abandoned storage tanks), atypical migration pathways (such as preferential flow in fractured rock layers, atmospheric diffusion path), special exposure receptors (such as children's activity area, sensitive ecological system), and social constraints (such as repair budget limitation, land reuse planning), etc. The needs, expectations and concerns of stakeholders (such as “repair should not affect the surrounding farmland” “ensure the safety of the children's activity area”) essentially reflect the social-environmental coupling problems that are not explicitly expressed in the model. For example, residents' complaints about “odors” may indicate that the atmospheric diffusion path of volatile organic compounds (VOCs) is not fully modeled; the developer's requirement for “land development timeline” may expose the model's failure to consider the trade-off relationship between repair period and cost. Traditional model validation methods rely on physical detection or statistical testing, which are difficult to actively identify unobserved information gaps (such as unidentified pollution sources or exposure pathways). To this end, the present application designs a feedback information-model prior semantic panoramic scanning network that actively identifies the contradictory nodes of the model and the feedback information of the stakeholders' needs (such as the mismatch between the location of the pollution source and the area of the residents' complaints, the overlap between the pollutant migration path and the ecological sensitive area but not reflected in the model, the conflict between the receptor exposure risk and the community planning, etc.) by deeply mining the potential semantic association between the feedback information semantic encoding vector and the set of site prior information model parameter row vectors, to reveal the potential key information gaps in the model.

[0037] Figure 4 A flowchart of step S42 in the method for constructing a contaminated site conceptual model according to an embodiment of the present application. As shown in Figure 4 S42, it includes: S421, respectively performing semantic query response encoding on the feedback information semantic encoding vector and each site prior information model parameter row vector in the set of site prior information model parameter row vectors to obtain a set of feedback information-model prior semantic query score encoding vectors; S422, based on the feature set of the set of feedback information-model prior semantic query score encoding vectors, performing semantic matching gate aggregation on the set of feedback information-model prior semantic query score encoding vectors from the distribution characteristics to obtain the feedback information-site prior semantic scanning response encoding vector.

[0038] More specifically, the step S421 first performs feature enhancement on the feedback information semantic encoding vector based on deconvolution coding to obtain a feedback information semantic reinforced encoding vector, which has the same feature dimension as each of the venue prior information model parameter row vectors in the set of venue prior information model parameter row vectors, and is expressed by the formula as follows:

[0039]

[0040] wherein V1 represents the feedback information semantic encoding vector, V1' represents the feedback information semantic reinforced encoding vector, W deconv represents a deconvolution weight matrix, f deconv (·) represents a deconvolution coding process, and ‖·‖ represents a norm of a vector.

[0041] Here, considering that the feedback information of stakeholders (such as “repair cost needs to be lower than the budget” and “children's activity area needs to be zero risk”) usually exists in the form of unstructured text, the semantic encoding vector thereof is different from the venue prior information model parameter row vectors in feature dimension and semantic space, and direct semantic matching analysis may cause numerical instability or semantic deviation (such as interference of high-dimensional vector with redundant noise to low-dimensional parameters). In this regard, the present application further maps the feedback information semantic encoding vector to a feature space consistent with the venue prior information model parameter row vectors by performing deconvolution coding on the feedback information semantic encoding vector, so as to realize alignment of feature dimensions and improve richness of semantic expression of feedback information, to obtain the feedback information semantic reinforced encoding vector.

[0042] Then, the feedback information semantic reinforced encoding vector is respectively subjected to single-body semantic query coding with each of the venue prior information model parameter row vectors in the set of venue prior information model parameter row vectors to obtain a set of feedback information-model prior semantic query score encoding vectors. In one specific example of the present application, the feedback information semantic reinforced encoding vector and the venue prior information model parameter row vector are concatenated and fused, and then input into a neural network layer based on a tanh function to obtain the feedback information-model prior semantic query score encoding vector, which is expressed by the formula as follows:

[0043] V2={V 21 ,V 22 ,...,V 2i ,...,V en}

[0044] R i =tanh{W Ri [V1′;V 2i ]+b i}

[0045] wherein V2 represents a set of venue prior information model parameter row vectors, V 21 , V 22 , V 2i and V 2n represent the 1st, 2nd, i-th and n-th venue prior information model parameter row vector of the set of venue prior information model parameter row vectors respectively, n is the number of the venue prior information model parameter row vectors, tanh(·) represents the hyperbolic tangent function, b i represents the bias term, W Ri represents the weight matrix of the neural network layer, [·; ·] represents the concatenation operation, R i represents the feedback information-model prior semantic query score encoding vector between V1' and V 2i .

[0046] It can be understood that the relevance of the set of venue prior information model parameter row vectors and the feedback information semantic reinforcement encoding vector needs to be evaluated through fine-grained semantic matching analysis. However, the traditional cosine similarity only measures the surface feature similarity and cannot capture the deep semantic logic. Therefore, the present application introduces a deep neural network to perform joint semantic analysis on each venue prior information model parameter row vector and the feedback information semantic reinforcement encoding vector, respectively, so as to utilize the hierarchical structure and nonlinear activation function of the deep neural network to mine the deep semantic information in the high-dimensional vector and model the multi-level semantic association between the venue prior information and the related feedback information, thereby obtaining a set of feedback information-model prior semantic query score encoding vectors.

[0047] More specifically, the step S422 comprises determining the individual semantic matching degree of each feedback information-model prior semantic query score encoding vector in the set of feedback information-model prior semantic query score encoding vectors based on the feature set distribution characteristics of the set of feedback information-model prior semantic query score encoding vectors to obtain a set of feedback information-model prior semantic matching degrees. In one preferred example of the present application, first, based on the semantic feature interaction between the feedback information semantic encoding vector and each venue prior information model parameter row vector in the set of venue prior information model parameter row vectors, the symmetry constraint optimization is performed on each corresponding feedback information-model prior semantic query score encoding vector to obtain a set of optimized feedback information-model prior semantic query score encoding vectors, which is represented by the formula as follows:

[0048] V 3i = [V1'; V 2i ]

[0049] V 4i = W i V 3i - V 3i

[0050] R′ i = R i + gS i V 4i

[0051]

[0052] V 4i = S i R i

[0053] wherein, V 3i represents a feedback information-model prior semantic feature concatenated vector between V1' and V 2i , W i represents a linear mapping matrix, V 4i represents an interaction potential vector between V1' and V 2i , g represents a coupling constant, S i represents a covariant matrix, R′ i represents R i corresponds to an optimized feedback information-model prior semantic query score encoding vector.

[0054] In particular, the present application considers that the site prior information model parameter row vector is constructed based on physical and chemical data, while the feedback information semantic reinforcement encoding vector more integrates the social constraints of stakeholders. There is a natural gap in the representation logic of the two in the semantic space, which leads to potential deviation in the semantic matching relationship between the two. To further optimize the semantic matching accuracy between feedback information and site prior information, the present application utilizes the covariant symmetry of feature space and semantic space to force the interaction between the site prior information model parameter row vector and the feedback information semantic reinforcement encoding vector to follow a unified mapping specification. Specifically, first, the feedback semantic reinforcement encoding vector and the site prior information model parameter row vector are spliced to generate an interaction potential vector between the two through a linear mapping matrix, representing the correlation strength of the two in the hidden space. Then, to eliminate the matching noise caused by semantic ambiguity, a coupling constant and a covariant matrix are further introduced to constrain the interaction process, ensuring that the matching strength of forward query (feedback→ parameter) and reverse verification (parameter→ feedback) is consistent, avoiding logical contradictions caused by one-way association, obtaining an optimized feedback information-model prior semantic query score encoding vector, to improve the representation accuracy of the semantic query score between feedback information and site prior information.

[0055] Then, the context semantic correlation degree of each optimization feedback information-model prior semantic query score encoding vector in the set of optimization feedback information-model prior semantic query score encoding vectors relative to other optimization feedback information-model prior semantic query score encoding vectors is calculated as the single semantic matching degree to obtain the set of feedback information-model prior semantic matching degrees, which is expressed by the formula as follows:

[0056]

[0057] wherein R' k represents the kth optimization feedback information-model prior semantic query score encoding vector in the set of optimization feedback information-model prior semantic query score encoding vectors, (·) T represents the transpose of a vector, exp(·) represents the exponential function operation with e as the base, softmax(·) represents the normalized exponential function, a k represents the feedback information-model prior semantic matching degree corresponding to the kth optimization feedback information-model prior semantic query score encoding vector, and a k represents the feedback information-model prior semantic matching degree corresponding to the kth optimization feedback information-model prior semantic query score encoding vector. k T i i

[0058] That is, considering that a single feedback information-model prior semantic query feature may be affected by local noise, the application further measures the deviation degree of a single score vector from the overall distribution by calculating the context semantic correlation degree of each optimization feedback information-model prior semantic query score encoding vector relative to other optimization feedback information-model prior semantic query score encoding vectors to evaluate the relative importance of each feedback information-model prior semantic query feature from a global perspective to obtain the set of feedback information-model prior semantic matching degrees.

[0059] More specifically, the step S422 further comprises: based on the set of feedback information-model prior semantic matching degrees, performing global semantic gating aggregation on the set of feedback information-model prior semantic query score encoding vectors to obtain the feedback information-site prior semantic scan response encoding vector. In one specific example of the application, first, the set of feedback information-model prior semantic matching degrees is input into a relationship gating agent module to obtain a set of feedback information-model prior query semantic self-attention weights; then, based on the set of feedback information-model prior query semantic self-attention weights, the set of optimization feedback information-model prior semantic query score encoding vectors is aggregated to obtain the feedback information-site prior semantic scan response encoding vector, which is expressed by the formula as follows:

[0060]

[0061] wherein τ represents a gating threshold, mask(·) represents a mask operation, w k represents the feedback information-model prior query semantic self-attention weight corresponding to the kth optimization feedback information-model prior semantic query score encoding vector, and a k represents the feedback information-model prior semantic matching degree corresponding to the kth optimization feedback information-model prior semantic query score encoding vector. i i ​​​​​The corresponding feedback information-model priori query semantic self-attention weight v p The feedback information-site priori semantic scanning response encoding vector is represented.

[0062] That is, further according to the size of the feedback information-model priori semantic matching degree, the aggregation weight of the feedback information-model priori semantic query score encoding vector is dynamically adjusted by using a gating mechanism, so as to filter out low-quality noise information and retain the attention degree to the key contradiction node. In this way, it is helpful for the model to focus more on the core concerns of stakeholders and potential environmental risk points, and to improve the information processing efficiency and sensitivity and identification ability to unobserved information. Further, based on the feedback information-model priori query semantic self-attention weight, the optimized feedback information-model priori semantic query score encoding vector is weighted and summed to form a comprehensive feedback information-site priori semantic scanning response encoding vector. In this way, the feedback information-site priori semantic scanning response encoding vector not only integrates the specific needs and expectations of stakeholders, but also captures unexpressed social-environmental coupling problems through deep semantic matching with site priori information, reveals key information gaps and potential risk points in the model, and further provides more comprehensive and accurate support for the iterative update of the model.

[0063] Specifically, the feedback information-site priori semantic scanning response encoding vector is feature-decoded to obtain the identification result of the key information gap in step S43. It should be understood that the feedback information-site priori semantic scanning response encoding vector internally contains semantic differences and associated information between the feedback information of stakeholders and the site priori information. In order to further convert it into a specific and understandable key information gap identification result, so as to provide a clear direction for the subsequent iterative optimization of the contaminated site conceptual model, the present application further performs feature decoding processing on the feedback information-site priori semantic scanning response encoding vector. Specifically, the feature decoding process is based on the inverse operation principle corresponding to the encoding process, and through steps such as layer-by-layer upsampling, attention weight inversion, and vocabulary decoding, the vector representation in the high-dimensional semantic space is restored to the key information gap description in the form of natural language text, thereby clearly revealing the key information gaps missing or not fully considered in the model, such as environmental risk points of particular concern to stakeholders, unobserved pollutant migration paths, and key parameters not included in the model. In this way, the site conceptual model can adapt to changing actual conditions, providing a clear direction and basis for the iterative optimization of the contaminated site conceptual model, helping to improve the adaptability and vitality of the model, and gradually fitting the real situation of the site in the improvement process at different stages, providing scientific guidance for subsequent management decisions of contaminated land.

[0064] In summary, the method for constructing a contaminated site conceptual model according to the embodiment of the present application is explained. It constructs a structured initial contaminated site conceptual model by integrating multi-source heterogeneous data of the contaminated site, and uses the site prior information model parameter matrix to characterize the spatiotemporal coupling relationship of pollution source-migration pathway-receptor. Then, the site prior information model parameter matrix is ​​further decomposed into independent units using information decoupling technology, and the feedback information of stakeholders is used as a constraint condition. Natural language processing technology and semantic query algorithms are used to perform a panoramic semantic scanning query on the site prior information to identify key information gaps such as exposure pathways that are not covered in the site prior information, so as to iteratively optimize the contaminated site conceptual model. In this way, it is possible to achieve intelligent identification of key information gaps in the contaminated site conceptual model, guide the dynamic adjustment and improvement of the conceptual model, and make it more in line with the actual situation and the needs of stakeholders.

[0065] Furthermore, the present application also provides a system for constructing a contaminated site conceptual model.

[0066] Figure 5 FIG. 1 is a block diagram of a system for constructing a contaminated site conceptual model according to an embodiment of the present application. Figure 5 As shown, according to an embodiment of the present application, a system 100 for constructing a contaminated site conceptual model includes: a contaminated site conceptual model construction module 110, which is used to construct an initial contaminated site conceptual model, and the initial contaminated site conceptual model is used to describe known or inferred sources of pollution, possible pollutant migration pathways, and potential exposure pathways and receptors; a feedback information collection module 120, which is used to collect feedback information provided by stakeholders, and the feedback information includes requirements, expectations and concerns for the site; an information decoupling module 130, which is used to extract a site prior information model parameter matrix from the initial contaminated site conceptual model, and perform information decoupling on the site prior information model parameter matrix according to row vectors to obtain a set of site prior information model parameter row vectors; a key information gap identification module 140, which is used to perform a site prior semantic panoramic scan on the set of site prior information model parameter row vectors based on the feedback information provided by the stakeholders to determine the key information gaps in the initial contaminated site conceptual model.

[0067] Here, those skilled in the art will appreciate that the specific operations of each module in the above-mentioned system for constructing the conceptual model of the contaminated site have been described in the above-mentioned Figures 1 to 4 The construction method of the contaminated site conceptual model has been described in detail in the description of the contaminated site conceptual model, and therefore, its repeated description will be omitted.

[0068] Finally, it should be noted that the above examples are merely intended to illustrate the technical solutions of the present application and not to limit the present application. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present application.

Claims

1. A method for constructing a conceptual model of a contaminated site, characterized in that: include: Constructing an initial contaminated site conceptual model, which is used to describe known or suspected pollution sources, possible pollutant migration pathways, and potential exposure pathways and receptors; Gather feedback from stakeholders, including needs, expectations, and concerns about the site; Extracting a site prior information model parameter matrix from the initial contaminated site conceptual model, and performing information decoupling on the site prior information model parameter matrix according to row vectors to obtain a set of site prior information model parameter row vectors, including: comprehensively sorting out and determining key parameters that can accurately describe the model based on pollution-related information of the initial contaminated site conceptual model, organizing the determined parameters into a two-dimensional matrix in a logical order to obtain a site prior information model parameter matrix, and using information decoupling technology to decompose the site prior information model parameter matrix according to row vectors so that each row vector represents a relatively independent pollution information unit, describing in a numerical manner key details such as the specific location of the pollution source, the type and concentration of the pollutant, the geometric characteristics and physical and chemical conditions of the migration path, and the exposure risk of potential receptors, so as to form a set of site prior information model parameter row vectors; Using a text encoder to perform text structured encoding on the feedback information provided by the stakeholders to obtain a semantic encoding vector of the feedback information; Inputting the feedback information semantic coding vector and the set of the site prior information model parameter row vectors into a feedback information-model prior semantic panoramic scanning network to obtain a feedback information-site prior semantic scanning response coding vector, including: performing semantic query response coding on the feedback information semantic coding vector and each site prior information model parameter row vector in the set of the site prior information model parameter row vectors to obtain a set of feedback information-model prior semantic query score coding vectors; based on the feature set self-distribution characteristics of the set of feedback information-model prior semantic query score coding vectors, performing semantic matching gated aggregation on the set of feedback information-model prior semantic query score coding vectors to obtain the feedback information-site prior semantic scanning response coding vector; Feature decoding is performed on the feedback information-site prior semantic scan response encoding vector to obtain an identification result of a key information gap.

2. The method for constructing a contaminated site conceptual model according to claim 1, characterized in that: The text encoder is a pre-trained language model based on the Transformer architecture.

3. The method for constructing a contaminated site conceptual model according to claim 2, characterized in that: The feedback information semantic encoding vector is respectively subjected to semantic query response encoding with each site prior information model parameter row vector in the set of site prior information model parameter row vectors to obtain a set of feedback information-model prior semantic query score encoding vectors, including: performing feature enhancement based on deconvolution coding on the feedback information semantic coding vector to obtain a feedback information semantic enhancement coding vector, wherein the feedback information semantic enhancement coding vector has the same feature scale as each site prior information model parameter row vector in the set of site prior information model parameter row vectors; The feedback information semantic enhancement coding vector is respectively subjected to monomer semantic query coding with each site prior information model parameter row vector in the set of site prior information model parameter row vectors to obtain the set of feedback information-model prior semantic query score coding vectors.

4. The method for constructing a contaminated site conceptual model according to claim 3, characterized in that: The feedback information semantic enhancement encoding vector is respectively subjected to monomer semantic query encoding with each site prior information model parameter row vector in the set of site prior information model parameter row vectors to obtain a set of feedback information-model prior semantic query score encoding vectors, including: The feedback information semantic enhancement coding vector and the site prior information model parameter row vector are cascaded and fused, and then input into a neural network layer based on a tanh function to obtain the feedback information-model prior semantic query score coding vector.

5. The method for constructing a contaminated site conceptual model according to claim 4, characterized in that: Based on the feature set self-distribution property of the set of feedback information-model prior semantic query score encoding vectors, performing semantic matching gated aggregation on the set of feedback information-model prior semantic query score encoding vectors to obtain the feedback information-venue prior semantic scan response encoding vector, including: Determining, based on the feature set self-distribution property of the set of feedback information-model prior semantic query score encoding vectors, the individual semantic matching degree of each feedback information-model prior semantic query score encoding vector in the set of feedback information-model prior semantic query score encoding vectors to obtain a set of feedback information-model prior semantic matching degrees; Based on the set of feedback information-model prior semantic matching degrees, a global semantic gating aggregation is performed on the set of feedback information-model prior semantic query score encoding vectors to obtain the feedback information-venue prior semantic scanning response encoding vector.

6. The method for constructing a contaminated site conceptual model according to claim 5, characterized in that: Determining the individual semantic matching degree of each feedback information-model prior semantic query score encoding vector in the set of feedback information-model prior semantic query score encoding vectors based on the feature set self-distribution property of the set of feedback information-model prior semantic query score encoding vectors to obtain a set of feedback information-model prior semantic matching degrees, including: Based on the semantic feature interaction between the feedback information semantic encoding vector and each site prior information model parameter row vector in the set of site prior information model parameter row vectors, symmetry-constrained optimization is performed on each corresponding feedback information-model prior semantic query score encoding vector to obtain a set of optimized feedback information-model prior semantic query score encoding vectors; Calculate the contextual semantic relevance of each optimized feedback information-model prior semantic query score encoding vector in the set of the optimized feedback information-model prior semantic query score encoding vector relative to other optimized feedback information-model prior semantic query score encoding vectors as the monomer semantic matching degree to obtain the set of the feedback information-model prior semantic matching degrees.

7. The method for constructing a contaminated site conceptual model according to claim 6, characterized in that: Based on the set of feedback information-model prior semantic matching degrees, performing global semantic gating aggregation on the set of feedback information-model prior semantic query score encoding vectors to obtain the feedback information-venue prior semantic scan response encoding vector, including: Inputting the set of feedback information-model prior semantic matching degrees into the relational gating proxy module to obtain a set of feedback information-model prior query semantic self-attention weights; The set of optimized feedback information-model prior semantic query score encoding vectors is aggregated based on the set of feedback information-model prior semantic query self-attention weights to obtain the feedback information-venue prior semantic scan response encoding vector.

8. A system for constructing a contaminated site conceptual model, used to execute the method according to any one of claims 1 to 7, characterized in that: include: A contaminated site conceptual model building module is used to build an initial contaminated site conceptual model, wherein the initial contaminated site conceptual model is used to describe known or inferred pollution sources, possible pollutant migration pathways, and potential exposure pathways and receptors; A feedback information collection module is used to collect feedback information provided by stakeholders, including requirements, expectations and concerns about the site; an information decoupling module, configured to extract a site prior information model parameter matrix from the initial contaminated site conceptual model, and perform information decoupling on the site prior information model parameter matrix according to row vectors to obtain a set of site prior information model parameter row vectors; A key information gap identification module is used to perform a site prior semantic panoramic scan on the set of site prior information model parameter row vectors based on the feedback information provided by the stakeholders to determine the key information gaps in the initial contaminated site conceptual model.

Citation Information

Patent Citations

  • Constructing method and device of polluted site knowledge graph

    CN115525766A

  • Polluted site multi-source heterogeneous data fusion method

    CN117593614A