Constructing method and system for conceptual model of polluted site

By integrating multi-source data and information decoupling technology and combining natural language processing to identify key information gaps in the conceptual model of polluted sites, the shortcomings of polluted sites management in traditional methods are solved, and the dynamic optimization of the model and the satisfaction of stakeholder needs are achieved.

CN120408994AActive Publication Date: 2025-08-01BEIJING MUNICIPAL RES INST OF ENVIRONMENT PROTECTION

Patent Information

Application Number
CN202510509194.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-08-01
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

Traditional pollution site management methods are difficult to systematically characterize the complex relationship between pollution sources, migration paths and receptors, resulting in the remediation plan deviating from actual needs or waste of resources, and lacking comprehensive consideration of the needs of different stakeholders, making it difficult to adapt to changing actual situations.

Method used

By integrating multi-source heterogeneous data, the initial contaminated site concept model is constructed, and the site prior information model parameter matrix is decomposed into independent units by using information decoupling technology, and combining natural language processing technology and semantic query algorithms, the key information gaps in the site concept model are identified and optimized to meet the needs of stakeholders.

Benefits of technology

Intelligent identification and dynamic adjustment of the conceptual model of polluted sites has been achieved, ensuring that the model meets the actual situation and the needs of multiple stakeholders, and improving the scientific nature of environmental decision-making and resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408994A_ABST
    Figure CN120408994A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of construction of pollution site models, and particularly discloses a construction method and system of a conceptual model of a pollution site, an initial conceptual model of the pollution site is constructed by integrating multi-source heterogeneous data of the pollution site, a space-time coupling relation of pollution source-migration path-acceptor is described in a parameter matrix form, and the conceptual model of the pollution site is constructed. Then, further decomposing the field prior information model parameter matrix into independent units, taking feedback information of a benefit related party as a constraint condition, and performing panoramic semantic scanning query on the field prior information by utilizing a natural language processing technology and a semantic query algorithm; and key information gaps such as uncovered exposure paths in the site prior information are identified, so that the conceptual model of the polluted site is iteratively optimized. By means of the mode, intelligent recognition of key information gaps of the conceptual model of the polluted site can be achieved, dynamic adjustment and improvement of the conceptual model are guided, and the conceptual model better meets the actual situation and the requirements of interested parties.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of contaminated site model construction, and more specifically, to a method and system for constructing a conceptual model of a contaminated site. Background Art

[0002] With the accelerating advancement of global industrialization and urbanization, soil and groundwater pollution has become a major challenge threatening the ecological environment and public safety. Traditional contaminated site management relies on empirical judgment or fragmented data, making it difficult to systematically depict the complex relationships among pollution sources, migration paths, and receptors, resulting in repair plans deviating from actual needs or wasting resources. The international environmental governance system requires a dynamic and structured contaminated site conceptual model (CSM) to integrate the three-dimensional spatio-temporal migration laws and exposure risks of pollutants, providing a scientific basis for environmental decision-making. Its core value lies in achieving "pollution traceability, risk quantification, and measures implementation", thereby balancing environmental governance, economic costs, and social fairness goals under complex site conditions.

[0003] Meanwhile, the concealment of contaminated sites, the incompleteness of historical data, and the diverse demands of stakeholders pose higher requirements for the completeness of the conceptual model. The different demands of stakeholders (such as the government, enterprises, and residents) (such as health risks, repair costs, and land reuse) are essentially the coupling constraints of the social-environmental system. However, in the process of site planning and design, traditional methods often lack comprehensive consideration of the needs of different stakeholders, making it difficult for the site conceptual model to fully meet the requirements of all parties. At the same time, once the model is determined, there is a lack of an effective iterative update mechanism, making it difficult to adapt to the changing actual situation and new demands. For example, a conceptual model constructed under the condition of insufficient comprehensive site information may have problems such as over-generalizing the spatial location of pollutant occurrence, ignoring key pollution migration paths, or receptor exposure risks, resulting in a large difference between the constructed conceptual model and the actual contaminated site, reducing its guiding value in actual environmental decision-making.

[0004] Therefore, an optimized method and system for constructing a conceptual model of a contaminated site are expected. Summary of the Invention

[0005] To solve the above technical problems, the present application is proposed. Embodiments of the present application provide a method and system for constructing a conceptual model of a contaminated site, which constructs a structured initial conceptual model of the contaminated site by integrating multi-source heterogeneous data of the contaminated site, and depicts the spatio-temporal coupling relationship of the pollution source - migration pathway - receptor with the prior information model parameter matrix of the site. Then, further use information decoupling technology to decompose the prior information model parameter matrix of the site into independent units, and use the feedback information of the stakeholders as a constraint condition, and use natural language processing technology and semantic query algorithms to perform a panoramic semantic scan query on the prior information of the site to identify key information gaps such as exposure pathways not covered in the prior information of the site, so as to iteratively optimize the conceptual model of the contaminated site. In this way, it is possible to realize the intelligent identification of key information gaps in the conceptual model of the contaminated site, guide the dynamic adjustment and improvement of the conceptual model, and make it more in line with the actual situation and the needs of stakeholders.

[0006] Correspondingly, according to one aspect of the present application, there is provided a method for constructing a conceptual model of a contaminated site, which includes:

[0007] Construct an initial conceptual model of the contaminated site, where the initial conceptual model of the contaminated site is used to describe known or speculated pollution sources, possible pollutant migration pathways, and potential exposure pathways and receptors;

[0008] Collect feedback information provided by stakeholders, where the feedback information includes the needs, expectations, and concerns about the site;

[0009] Extract the prior information model parameter matrix of the site from the initial conceptual model of the contaminated site, and perform information decoupling on the prior information model parameter matrix according to row vectors to obtain a set of prior information model parameter row vectors of the site;

[0010] Based on the feedback information provided by the stakeholders, perform a panoramic semantic scan of the prior information model parameter row vectors of the site to determine the key information gaps existing in the initial conceptual model of the contaminated site.

[0011] According to another aspect of the present application, there is provided a system for constructing a conceptual model of a contaminated site, which includes:

[0012] A contaminated site conceptual model construction module for constructing an initial conceptual model of the contaminated site, where the initial conceptual model of the contaminated site is used to describe known or speculated pollution sources, possible pollutant migration pathways, and potential exposure pathways and receptors;

[0013] A feedback information collection module for collecting feedback information provided by stakeholders, where the feedback information includes the needs, expectations, and concerns about the site;

[0014] An information decoupling module, configured to extract the prior information model parameter matrix of the site from the initial contaminated site concept model, and perform information decoupling on the prior information model parameter matrix of the site by row vectors to obtain a set of prior information model parameter row vectors of the site;

[0015] A key information gap identification module, configured to perform a panoramic semantic scan of the prior information of the site on the set of prior information model parameter row vectors of the site based on the feedback information provided by the stakeholders, so as to determine the key information gaps existing in the initial contaminated site concept model.

[0016] Compared with the prior art, the method and system for constructing a contaminated site concept model provided by the present application integrate multi-source heterogeneous data of the contaminated site to construct a structured initial contaminated site concept model, characterize the spatio-temporal coupling relationship of the pollution source - migration pathway - receptor with the prior information model parameter matrix of the site. Then, further use the information decoupling technology to decompose the prior information model parameter matrix of the site into independent units, and use the feedback information of the stakeholders as a constraint condition. By using natural language processing technology and semantic query algorithms, through panoramic semantic scan query of the prior information of the site, key information gaps such as exposure pathways not covered in the prior information of the site are identified, so as to iteratively optimize the contaminated site concept model. In this way, the intelligent identification of key information gaps in the contaminated site concept model can be realized, guiding the dynamic adjustment and improvement of the concept model to make it more in line with the actual situation and the needs of stakeholders. Description of the Drawings

[0017] By describing the embodiments of the present application in more detail in conjunction with the drawings, the above and other objects, features and advantages of the present application will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0018] Figure 1 It is a flowchart of the method for constructing a contaminated site concept model according to an embodiment of the present application.

[0019] Figure 2 It is a schematic diagram of data flow of the method for constructing a contaminated site concept model according to an embodiment of the present application.

[0020] Figure 3 It is a flowchart of step S4 in the method for constructing a contaminated site concept model according to an embodiment of the present application.

[0021] Figure 4 It is a flowchart of step S42 in the method for constructing a contaminated site concept model according to an embodiment of the present application.

[0022] Figure 5 It is a block diagram of a construction system for a contaminated site conceptual model according to an embodiment of the present application. Detailed implementation manners

[0023] Next, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.

[0024] Figure 1 It is a flowchart of a construction method for a contaminated site conceptual model according to an embodiment of the present application. Figure 2 It is a schematic diagram of data flow of a construction method for a contaminated site conceptual model according to an embodiment of the present application. As Figure 1 and Figure 2 shown, the construction method for a contaminated site conceptual model according to an embodiment of the present application includes the steps of: S1, constructing an initial contaminated site conceptual model, where the initial contaminated site conceptual model is used to describe known or speculated pollution sources, possible pollutant migration pathways, and potential exposure pathways and receptors; S2, collecting feedback information provided by stakeholders, where the feedback information includes requirements, expectations, and concerns about the site; S3, extracting a site prior information model parameter matrix from the initial contaminated site conceptual model, and decoupling the information of the site prior information model parameter matrix by row vectors to obtain a set of site prior information model parameter row vectors; S4, based on the feedback information provided by the stakeholders, performing a panoramic semantic scan of the set of site prior information model parameter row vectors to determine key information gaps existing in the initial contaminated site conceptual model.

[0025] In the above method for constructing a conceptual model of a contaminated site, in step S1, an initial conceptual model of the contaminated site is constructed, and the initial conceptual model of the contaminated site is used to describe known or speculated pollution sources, possible pollutant migration pathways, and potential exposure pathways and receptors. It should be understood that the complexity and concealment of contaminated sites make it difficult for a single data source to comprehensively describe the pollution distribution and migration law. Traditional methods rely on scattered data or empirical assumptions, and it is easy to overlook key pollution sources or exposure pathways. Therefore, this application proposes to integrate multi-source heterogeneous data, such as geological exploration data, environmental monitoring data, historical pollution records, etc., to construct an initial conceptual model of the contaminated site, so as to structurally express the spatio-temporal coupling relationship of pollution sources, migration pathways, and receptors, and provide a unified framework for subsequent analysis. Among them, the pollution source refers to the source of pollutants in the site, which may include waste discharge, chemical substance leakage, landfill, etc. during industrial production processes. For example, in a former chemical plant site, various chemical raw materials and intermediate products used and stored during its production process may be potential pollution sources. The migration pathway is used to describe the migration mode and path of pollutants in the site environment, including processes such as diffusion, infiltration, and convection of pollutants in media such as soil, groundwater, and air. For example, heavy metals in the soil may migrate downward with the leaching of rainwater and enter the groundwater; volatile organic compounds may enter the atmospheric environment through volatilization and then diffuse with the air flow. The receptor refers to the object that may be affected by pollutant exposure, including humans (such as residents living around the site, personnel working in the site), ecosystems (such as animals and plants in the site, surrounding water ecosystems), etc. For example, residents living near the contaminated site may be exposed to pollutants by breathing polluted air, drinking polluted groundwater, or eating crops grown in polluted soil, thus posing a potential hazard to their physical health.

[0026] When constructing an initial conceptual model of a contaminated site, a comprehensive and systematic description of known or inferred pollution sources, possible pollutant migration pathways, and potential exposure pathways and receptors is a key step. This involves the collection, processing, and integration of multi-source data. Specifically, multi-source data collection is the foundation of model construction, requiring extensive acquisition of various types of site-related data. Multi-source remote sensing imagery is a key source. High-resolution satellite or aerial remote sensing can provide macroscopic imagery of the site and its surrounding area. Using specialized image processing software and algorithms, remote sensing imagery can be interpreted to identify land use types within the site, such as industrial, agricultural, and residential. It can also identify areas with potential signs of contamination, such as unusual color or vegetation growth, which may indicate potential sources of contamination. For example, in a former industrial cluster, if remote sensing images show a plot of land with a significantly different color from the surrounding area and sparse vegetation, further investigation may reveal heavy metal contamination due to the release of heavy metal-containing waste from past industrial production.

[0027] Secondly, raster data is equally indispensable, as it can accurately present the site's topography and the spatial distribution of relevant environmental parameters. Site elevation data is acquired through topographic surveying equipment and processed to generate a digital elevation model (DEM). The DEM provides a visual understanding of the site's topographical fluctuations. Topography has a significant impact on pollutant migration. Pollutants tend to accumulate in low-lying areas, while in areas with steeper slopes, pollutants may migrate rapidly with water flows. Furthermore, raster data can also include information such as soil type and groundwater level. Different soil types have different adsorption and transport capacities for pollutants. Clay soils have a strong adsorption capacity for heavy metals, while sandy soils are more conducive to the spread of pollutants. By analyzing raster data, we can preliminarily determine the possible migration direction and range of pollutants under different terrain and soil conditions.

[0028] Contaminated site vector data is also an important data type, which can clearly define information such as site boundaries and infrastructure layout. Specifically, with the help of geographic information system (GIS) technology, vector data such as site boundaries, underground pipeline distribution, and building locations can be entered into the system to construct the spatial topology of the site. This is of great significance for analyzing the migration paths of pollutants within the site and the distribution of potential receptors. For example, understanding the direction of underground pipelines can determine whether pollutants are likely to spread through pipeline leakage; knowing the location of buildings can determine whether people living or working in the site are potential receptors.

[0029] In addition to the above-mentioned structured data, unstructured data such as text descriptions, data charts, and maps in the detailed investigation reports also contain rich information. For the text description part, optical character recognition (OCR) technology is used to convert the text in paper documents into editable text, and then natural language processing technology is used for semantic analysis. Key information is extracted from the text, such as the historical usage of the site, past pollution incidents, types of pollutants, and emission records. The data charts may contain pollutant monitoring concentration data at different time points. By sorting out and analyzing these data, the changing trend of pollutant concentration over time can be understood, and the development process of pollution can be inferred. The maps may include geological cross-sections of the site, hydrogeological maps, etc. Information about the stratigraphic structure of the site can be obtained from the geological cross-section to judge the blocking or promoting effect of different strata on pollutant migration; the hydrogeological map can show the flow direction and hydraulic gradient of groundwater, providing a basis for analyzing the migration of pollutants in groundwater.

[0030] After the data collection is completed, preprocessing is required. Specifically, data cleaning technology is used to remove noise, error values, and duplicate data in the data. For example, in pollutant monitoring data, there may be outliers caused by instrument failures. By setting reasonable data ranges and statistical analysis methods, these outliers can be identified and removed. The data is standardized to unify the data format and units. Data from different sources may use different units. For example, length units may include meters, feet, etc. By converting the units, they are unified into international standard units to facilitate subsequent analysis and integration. Next, tools such as geographic information system (GIS) technology and data mining algorithms are used for data integration and analysis. The powerful spatial analysis function of GIS technology plays an important role in this process. The preprocessed multi-source data is imported into the GIS system, and different types of data are fused through spatial overlay analysis. For example, the suspected pollution area interpreted from remote sensing images is overlaid with the geological map of the site to analyze the impact of the geological conditions in this area on pollutants; the groundwater level data is overlaid with the topographic data of the site to judge the relationship between the flow direction of groundwater and the topography, and then infer the migration path of pollutants in groundwater. Data mining algorithms are used to discover potential patterns and rules from a large amount of data. For pollutant monitoring data, clustering algorithms can be used to divide the site into different pollution areas, where the pollutant concentration and type are similar within each area. Association rule mining algorithms can analyze the association relationships between different pollutants. For example, if certain pollutants always appear simultaneously, this may indicate that they have the same pollution source. Through decision tree algorithms, based on various characteristic parameters of the site, such as soil type, terrain slope, groundwater level, etc., the migration trend of pollutants under different conditions can be predicted.

[0031] Finally, based on the above data processing and analysis results, an initial conceptual model of the contaminated site is constructed. Specifically, when determining the location of the pollution source, multi-source data information is integrated. For example, through the interpretation of remote sensing images, historical pollution records, and on-site investigations, the geographical coordinates of the pollution source are accurately located. For the types of pollutants, based on chemical analysis results, enterprise production records, etc., the types of pollutants present in the site are determined, including heavy metals, organic substances, inorganic substances, etc. When inferring the possible migration paths, combined with the topographical and geomorphological features, hydrogeological conditions, and the laws obtained from data mining and analysis, the migration process of pollutants in media such as soil, water, and air is simulated. When determining the distribution of potential receptors, densely populated areas such as residential areas, enterprises, and schools around the site, as well as ecologically sensitive areas such as nature reserves, rivers, and lakes, are considered, and the people and ecosystems in these areas are determined as potential receptors. Finally, the constructed initial conceptual model of the contaminated site is presented in the form of graphs, charts, or texts. For example, using a GIS system to generate two-dimensional or three-dimensional maps to visually display the location of the pollution source, migration paths, and distribution of potential receptors. Present the types, concentrations, and trends of changes over time of pollutants in the form of charts. Through detailed written explanations, explain the basis for model construction, the relationships between various elements, and the limitations of the model. In addition, different types of pollution sources can be represented by icons of different colors on the map, and the migration direction of pollutants can be represented by arrows; in the chart, the concentrations of pollutants in different regions can be shown as bar charts, and the changes in pollutant concentrations over time can be shown as line charts; in the written explanation, elaborate on the data and technologies based on which the model is constructed, and the uncertainties that may exist in the model due to data incompleteness. Through the above methods, the constructed initial conceptual model of the contaminated site can more accurately describe the known or inferred pollution sources, possible pollutant migration paths, and potential exposure paths and receptors, providing a solid foundation for subsequent in-depth analysis of the site pollution situation, risk assessment, and formulation of remediation strategies.

[0032] In the above method for constructing a conceptual model of a contaminated site, in step S2, feedback information provided by stakeholders is collected, and the feedback information includes the needs, expectations, and concerns about the site. It should be understood that considering the different stakeholders (including government departments, enterprises, and residents, etc.), there are different demands for contaminated sites based on their own role positions and interest demands. For example, government departments usually focus on the risks to public health and the strict implementation of environmental regulations at the site, and expect to safeguard the public interest through effective site management; enterprises may be more concerned about the control of site remediation costs and the feasibility of land reuse to maximize economic benefits; residents highly focus on the safety of the living environment and the unaffected quality of life. However, traditional methods often ignore the comprehensive consideration of the needs of multiple stakeholders when constructing a site conceptual model. Therefore, this application further collects the needs, expectations, and concerns of relevant stakeholders about the site through interviews, questionnaires, symposiums, etc., reveals the complex coupling and constraint relationships between the social-environmental systems, and provides a clear direction for optimizing the conceptual model based on the feedback information of stakeholders in the follow-up.

[0033] In the above method for constructing a contaminated site conceptual model, in step S3, a prior information model parameter matrix of the site is extracted from the initial contaminated site conceptual model, and the prior information model parameter matrix of the site is decoupled according to row vectors to obtain a set of prior information model parameter row vectors of the site. Specifically, although the initial contaminated site conceptual model has integrated multi-source data, the information therein is complex and intertwined, which is not conducive to effectively comparing and analyzing with the feedback information of stakeholders. Therefore, in order to deeply analyze the prior information of the site in the initial contaminated site conceptual model, identify the differences between the initial contaminated site conceptual model and the actual requirements, and find the key information gaps, the present application further abstractly represents the complex pollution-related information in the initial contaminated site conceptual model in the form of a mathematical matrix to form a prior information model parameter matrix of the site, so as to clearly depict the spatio-temporal coupling relationship between the pollution source - migration pathway - receptor. Specifically, first, according to the pollution-related information of the initial contaminated site conceptual model, the key parameters that can accurately describe the model are comprehensively sorted out and determined. For example, for the pollution source, determine its three-dimensional spatial coordinates (X, Y, Z coordinates), as well as parameters such as the type of pollutant, initial concentration, and release rate; for the migration pathway, measure or estimate its geometric parameters such as length, width, and slope, as well as the physicochemical property parameters of the medium (such as soil, water body, air), such as soil porosity, water body flow rate, air diffusion coefficient, etc., which will affect the migration rate and direction of pollutants; for the receptor, clarify its type (such as humans, animals, plants, etc.), specific location coordinates, and parameters such as the exposure method and frequency that may be affected by pollution. Then, the determined parameters are organized into a two-dimensional matrix in a logical order to obtain a prior information model parameter matrix of the site. For example, each row of the prior information model parameter matrix of the site can represent a specific pollution scenario or element, and each column corresponds to a specific parameter. Next, using information decoupling technology, the prior information model parameter matrix of the site is decomposed according to row vectors, so that each row vector represents a relatively independent pollution information unit, numerically describing key details such as the specific location of the pollution source, the type and concentration of pollutants, the geometric characteristics and physicochemical conditions of the migration path, and the exposure risk of potential receptors, thereby forming a set of prior information model parameter row vectors of the site. In this way, not only the readability and operability of the information are enhanced, but also it provides convenience for subsequent information gap identification and model optimization.

[0034] In the above method for constructing a contaminated site conceptual model, in step S4, based on the feedback information provided by the stakeholders, a prior semantic panoramic scan of the set of prior information model parameter row vectors of the site is performed to determine the key information gaps existing in the initial contaminated site conceptual model. Among them, Figure 3It is a flowchart of step S4 in the method for constructing a conceptual model of a contaminated site according to an embodiment of the present application. As Figure 3 shown, the step S4 includes: S41, using a text encoder to perform text structured encoding on the feedback information provided by the stakeholders to obtain a feedback information semantic encoding vector; S42, inputting the feedback information semantic encoding vector and the set of prior information model parameter row vectors of the site into a feedback information-prior semantic panoramic scanning network to obtain a feedback information-site prior semantic scanning response encoding vector; S43, performing feature decoding on the feedback information-site prior semantic scanning response encoding vector to obtain the identification result of the key information gap.

[0035] Specifically, in step S41, a text encoder is used to perform text structured encoding on the feedback information provided by the stakeholders to obtain a feedback information semantic encoding vector. That is, considering that the feedback information provided by the stakeholders mostly presents in the form of natural language text, this unstructured data format has a large difference from the structured information represented by the prior information model parameter matrix of the site and is difficult to directly perform effective comparison and analysis. Therefore, in order to be able to deeply compare the two in the same semantic space and identify the key information gap, the present application further uses a text encoder to convert the feedback information provided by the stakeholders into a coding form compatible with the model information. In a specific example of the present application, the text encoder is a pre-trained language model based on the Transformer architecture. It should be understood that with its powerful self-attention mechanism, the Transformer architecture can efficiently process the complex semantic relationships in natural language text, capture the long-distance dependence relationships between different words and sentences in the text, thereby more accurately understand the semantics of the text and convert it into a corresponding vector representation. In the present application, by inputting the feedback information into the text encoder and utilizing the powerful semantic understanding and representation ability of the Transformer architecture, through steps such as word segmentation processing, word embedding, position encoding, and self-attention mechanism, the natural language text in the feedback information is converted into a vector representation in a high-dimensional semantic space to obtain a feedback information semantic encoding vector, so as to express the needs, expectations, and concerns of the stakeholders for the site, providing a solid semantic foundation for subsequent information gap identification and model optimization.

[0036] Specifically, in step S42, the semantic encoding vector of the feedback information and the set of row vectors of the site prior information model parameters are input into the feedback information - model prior semantic panoramic scanning network to obtain the feedback information - site prior semantic scanning response encoding vector. It should be understood that the initial contaminated site conceptual model is constructed based on known or inferred physicochemical data (such as pollutant concentration, geological parameters), but limited by incomplete data, fuzzy historical information, or simplified technical assumptions (such as homogeneous medium assumption), it may miss hidden pollution sources (such as uninvestigated abandoned storage tanks), atypical migration pathways (such as preferential flow in fractured rock formations, atmospheric diffusion paths), special exposure receptors (such as children's activity areas, sensitive ecosystems), social constraints (such as remediation budget limitations, land reuse planning), etc. The needs, expectations, and concerns of stakeholders (such as "remediation shall not affect the surrounding farmland" and "the safety of children's activity areas shall be ensured") essentially reflect the social - environmental coupling issues not explicitly expressed in the model. For example, residents' complaints about "strange smells" may point to the fact that the atmospheric diffusion path of volatile organic compounds (VOCs) has not been fully modeled; developers' requirements for the "land development timeline" may expose the trade - off relationship between the remediation cycle and cost not considered in the model. Traditional model verification methods rely on physical detection or statistical tests and are difficult to actively discover unobserved information gaps (such as unrecognized pollution sources or exposure pathways). In response to this, the present application designs a feedback information - model prior semantic panoramic scanning network, which actively identifies the conflict nodes between the model and the stakeholder demand feedback information (such as the mismatch between the pollution source location and the area complained by residents, the overlap between the pollutant migration path and the ecological sensitive area but not reflected in the model, the conflict between the receptor exposure risk and the community planning, etc.) by deeply mining the potential semantic associations between the semantic encoding vector of the feedback information and the set of row vectors of the site prior information model parameters, so as to reveal the potential key information gaps in the model.

[0037] Figure 4 It is a flowchart of step S42 in the method for constructing a contaminated site conceptual model according to an embodiment of the present application. As Figure 4 shown, step S42 includes: S421, performing semantic query response encoding on the semantic encoding vector of the feedback information and each row vector of the site prior information model parameters in the set of row vectors of the site prior information model parameters to obtain a set of feedback information - model prior semantic query score encoding vectors; S422, based on the self - distribution characteristics of the feature set of the set of feedback information - model prior semantic query score encoding vectors, performing semantic matching gating aggregation on the set of feedback information - model prior semantic query score encoding vectors to obtain the feedback information - site prior semantic scanning response encoding vector.

[0038] More specifically, in step S421, first, perform feature enhancement on the semantic encoding vector of the feedback information based on deconvolutional encoding to obtain a semantic reinforcement encoding vector of the feedback information. The semantic reinforcement encoding vector of the feedback information has the same feature scale as each site prior information model parameter row vector in the set of site prior information model parameter row vectors, which is expressed by the formula:

[0039]

[0040] where V1 represents the semantic encoding vector of the feedback information, V1' represents the semantic reinforcement encoding vector of the feedback information, W deconv represents the deconvolution weight matrix, f deconv (·) represents deconvolutional encoding processing, and ‖·‖ represents calculating the norm of a vector.

[0041] Here, considering that the feedback information of stakeholders (such as "the repair cost needs to be lower than the budget" and "the children's activity area needs to be risk-free") usually exists in the form of unstructured text, there are differences in the feature scale and semantic space between its semantic encoding vector and the site prior information model parameter row vector. Direct semantic matching analysis may lead to numerical instability or semantic deviation (such as redundant noise of high-dimensional vectors interfering with low-dimensional parameters). To this end, in this application, the semantic encoding vector of the feedback information is further deconvolutionally encoded to map it to the same feature space as the site prior information model parameter row vector, realizing the alignment of feature dimensions and at the same time enhancing the richness of the semantic expression of the feedback information to obtain a semantic reinforcement encoding vector of the feedback information.

[0042] Next, perform single-body semantic query encoding on the semantic reinforcement encoding vector of the feedback information and each site prior information model parameter row vector in the set of site prior information model parameter row vectors to obtain a set of feedback information-model prior semantic query score encoding vectors. In a specific example of this application, the semantic reinforcement encoding vector of the feedback information and the site prior information model parameter row vector are cascaded and fused and then input into a neural network layer based on the tanh function to obtain the feedback information-model prior semantic query score encoding vector, which is expressed by the formula:

[0043] V2 = {V 21 , V 22 ,..., V 2i ,..., V en}

[0044] R i = tanh{W Ri [V1'; V 2i + b i}

[0045] Among them, V2 represents the set of row vectors of the site prior information model parameters, V 21 , V 22 , V 2i and V 2n respectively represent the first, second, i-th, and n-th row vectors of the site prior information model parameters in the set of row vectors of the site prior information model parameters. n is the number of row vectors of the site prior information model parameters. tanh(·) represents the hyperbolic tangent function, b i represents the bias term, W Ri represents the weight matrix of the neural network layer, [·;·] represents the concatenation operation, R i represents the feedback information - model prior semantic query score encoding vector between V1' and V 2i .

[0046] It should be understood that the relevance between the set of row vectors of the site prior information model parameters and the feedback information semantic enhancement encoding vector needs to be evaluated through fine-grained semantic matching analysis. The traditional cosine similarity only measures the similarity of surface features and cannot capture the deep semantic logic. Therefore, in this application, by introducing a deep neural network, joint semantic analysis is performed on each row vector of the site prior information model parameters and the feedback information semantic enhancement encoding vector respectively, so as to utilize the hierarchical structure and non-linear activation function of the deep neural network to mine the deep semantic information in the high-dimensional vector, model the multi-level semantic association between the site prior information and the relevant feedback information, and thus obtain the set of feedback information - model prior semantic query score encoding vectors.

[0047] More specifically, the step S422 includes: determining the single semantic matching degree of each feedback information - model prior semantic query score encoding vector in the set of feedback information - model prior semantic query score encoding vectors based on the self-distribution characteristics of the feature set of the set of feedback information - model prior semantic query score encoding vectors to obtain the set of feedback information - model prior semantic matching degrees. In a preferred example of this application, first, based on the semantic feature interaction between the feedback information encoding vector and each row vector of the site prior information model parameters in the set of row vectors of the site prior information model parameters, symmetry constraint optimization is performed on each corresponding feedback information - model prior semantic query score encoding vector to obtain the set of optimized feedback information - model prior semantic query score encoding vectors, which is expressed by the formula as follows:

[0048] V 3i =[V1';V 2i

[0049] V 4i =W i V 3i -V 3i ​

[0050] R′ i = R i + gS i V 4i

[0051]

[0052] V 4i = S i R i

[0053] where V 3i represents the feedback information - model prior semantic feature concatenation vector between V1′ and V 2i , W i represents the linear mapping matrix, V 4i represents the interaction potential vector between V1′ and V 2i , g represents the coupling constant, S i represents the covariance matrix, and R′ i represents the optimized feedback information - model prior semantic query score coding vector corresponding to R i .

[0054] In particular, this application takes into account that the row vector of the prior information model parameters of the site is constructed based on physicochemical data, while the feedback information semantic enhancement coding vector incorporates more social constraint considerations of stakeholders. There is a natural gap in the representation logic between the two in the semantic space, which in turn leads to potential biases in the semantic matching relationship between the two. To this end, in order to further optimize the semantic matching accuracy between the feedback information and the site prior information, this application uses the covariant symmetry of the feature space and the semantic space to force the interaction between the row vector of the prior information model parameters of the site and the feedback information semantic enhancement coding vector to follow a unified mapping specification. Specifically, first, after concatenating the feedback semantic enhancement coding vector and the row vector of the prior information model parameters of the site, a linear mapping matrix is used to generate the interaction potential vector between the two, representing the association strength between the two in the hidden space. Then, to eliminate the matching noise caused by semantic ambiguity, a coupling constant and a covariance matrix are further introduced to constrain the interaction process to ensure that the matching strengths of the forward query (feedback → parameter) and the reverse verification (parameter → feedback) are consistent, avoiding logical contradictions caused by one - way association, and obtaining the optimized feedback information - model prior semantic query score coding vector to improve the representation accuracy of the semantic query score between the feedback information and the site prior information.

[0055] Next, calculate the context semantic correlation degree of each optimized feedback information - model prior semantic query score encoding vector in the set of optimized feedback information - model prior semantic query score encoding vectors relative to other optimized feedback information - model prior semantic query score encoding vectors as the single - entity semantic matching degree to obtain the set of feedback information - model prior semantic matching degrees, which is expressed by the formula:

[0056]

[0057] Among them, R′ k represents the k - th optimized feedback information - model prior semantic query score encoding vector in the set of optimized feedback information - model prior semantic query score encoding vectors, (·) T represents the transpose of the vector, exp(·) represents the exponential function operation with base e, softmax(·) represents the normalized exponential function, and a i represents the feedback information - model prior semantic matching degree corresponding to R′ i That is, considering that a single feedback information - model prior semantic query feature may be affected by local noise, therefore, in this application, further calculate the context semantic correlation degree of each optimized feedback information - model prior semantic query score encoding vector relative to other optimized feedback information - model prior semantic query score encoding vectors, measure the deviation degree of a single score vector from the overall distribution, so as to evaluate the relative importance of each feedback information - model prior semantic query feature from a global perspective and obtain the set of feedback information - model prior semantic matching degrees.

[0058] More specifically, step S422 further includes: based on the set of feedback information - model prior semantic matching degrees, perform global semantic gating aggregation on the set of feedback information - model prior semantic query score encoding vectors to obtain the feedback information - venue prior semantic scan response encoding vector. In a specific example of this application, first, input the set of feedback information - model prior semantic matching degrees into the relational gating proxy module to obtain the set of feedback information - model prior query semantic self - attention weights; then, based on the set of feedback information - model prior query semantic self - attention weights, aggregate the set of optimized feedback information - model prior semantic query score encoding vectors to obtain the feedback information - venue prior semantic scan response encoding vector, which is expressed by the formula:

[0059]

[0060]

[0061]

[0061] Among them, τ represents the gating threshold, mask(·) represents the masking operation, and w i represents a iCorresponding feedback information - The model prior queries the semantic self-attention weight, v p Indicates the feedback information - the site prior semantic scan response encoding vector.

[0062] That is, further according to the magnitude of the feedback information - model prior semantic matching degree, the gating mechanism is used to dynamically adjust the aggregation weight of the feedback information - model prior semantic query score encoding vector, so as to filter out low-quality noise information and retain the attention to key contradiction nodes. In this way, it helps the model to focus more on the core concerns of stakeholders and potential environmental risk points, improving the information processing efficiency and the sensitivity and recognition ability to unobserved information. Furthermore, based on the feedback information - model prior query semantic self-attention weight, the optimized feedback information - model prior semantic query score encoding vector is weighted and summed to form a comprehensive feedback information - site prior semantic scan response encoding vector. In this way, the feedback information - site prior semantic scan response encoding vector not only integrates the specific needs and expectations of stakeholders, but also captures the unexplicitly expressed social-environmental coupling problems through deep semantic matching with the site prior information, revealing the key information gaps and potential risk points in the model, and thus providing more comprehensive and accurate support for the iterative update of the model.

[0063] Specifically, in step S43, the feedback information - site prior semantic scan response encoding vector is feature decoded to obtain the identification result of the key information gap. It should be understood that the feedback information - site prior semantic scan response encoding vector contains the semantic differences and correlation information between the stakeholder feedback information and the site prior information. In order to further convert it into a specific and understandable key information gap identification result for providing a clear direction for the subsequent iterative optimization of the contaminated site conceptual model, the present application further performs feature decoding processing on the feedback information - site prior semantic scan response encoding vector. Specifically, the feature decoding process is based on the inverse operation principle corresponding to the encoding process. Through steps such as layer-by-layer upsampling, attention weight inversion, and vocabulary decoding, the vector representation in the high-dimensional semantic space is restored to a key information gap description in the form of natural language text, thus clearly revealing the key information gaps missing or not fully considered in the model, such as environmental risk points particularly concerned by stakeholders, unobserved pollutant migration paths, key parameters not included in the model, etc. In this way, the site conceptual model can be adapted to the changing actual situation, providing a clear direction and basis for the iterative optimization of the contaminated site conceptual model, helping to improve the adaptability and vitality of the model, making the conceptual model gradually fit the real situation of the site during the improvement process at different stages, and providing scientific guidance for the management decision-making of subsequent contaminated plots.

[0064] In summary, the method for constructing a conceptual model of a contaminated site according to an embodiment of the present application is elucidated. It constructs a structured initial conceptual model of the contaminated site by integrating multi-source heterogeneous data of the contaminated site, depicts the spatio-temporal coupling relationship of the pollution source - migration pathway - receptor with the prior information model parameter matrix of the site. Then, it further uses information decoupling technology to decompose the prior information model parameter matrix of the site into independent units, and takes the feedback information of the stakeholders as a constraint condition, and uses natural language processing technology and semantic query algorithms to perform a panoramic semantic scan query on the prior information of the site to identify key information gaps such as exposure pathways not covered in the prior information of the site, so as to iteratively optimize the conceptual model of the contaminated site. In this way, it can realize the intelligent identification of key information gaps in the conceptual model of the contaminated site, guide the dynamic adjustment and improvement of the conceptual model, and make it more in line with the actual situation and the needs of stakeholders.

[0065] Furthermore, the present application also provides a system for constructing a conceptual model of a contaminated site.

[0066] Figure 5 The block diagram of the system for constructing a conceptual model of a contaminated site according to an embodiment of the present application is as follows. Figure 5 As shown in the figure, the system 100 for constructing a conceptual model of a contaminated site according to an embodiment of the present application includes: a contaminated site conceptual model construction module 110, configured to construct an initial conceptual model of the contaminated site, where the initial conceptual model of the contaminated site is used to describe known or speculated pollution sources, possible pollutant migration pathways, and potential exposure pathways and receptors; a feedback information collection module 120, configured to collect feedback information provided by stakeholders, where the feedback information includes requirements, expectations, and concerns about the site; an information decoupling module 130, configured to extract a prior information model parameter matrix of the site from the initial conceptual model of the contaminated site, and perform information decoupling on the prior information model parameter matrix according to row vectors to obtain a set of prior information model parameter row vectors of the site; a key information gap identification module 140, configured to perform a panoramic prior semantic scan on the set of prior information model parameter row vectors of the site based on the feedback information provided by the stakeholders to determine key information gaps existing in the initial conceptual model of the contaminated site.

[0067] Here, those skilled in the art can understand that the specific operations of each module in the above system for constructing a conceptual model of a contaminated site have been described in detail in the above description of the method for constructing a conceptual model of a contaminated site according to Figures 1 to 4 and therefore, the repeated description thereof will be omitted.

[0068] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for constructing a conceptual model of a contaminated site, characterized in that Including: Constructing an initial contaminated site conceptual model for describing known or suspected pollution sources, possible pollutant migration pathways, and potential exposure pathways and receptors; Collecting feedback information provided by stakeholders, where the feedback information includes the needs, expectations, and concerns about the site; Extracting a site prior information model parameter matrix from the initial contaminated site conceptual model, and decoupling the information of the site prior information model parameter matrix by row vectors to obtain a set of site prior information model parameter row vectors; Based on the feedback information provided by the stakeholders, performing a site prior semantic panoramic scan on the set of site prior information model parameter row vectors to determine the key information gaps existing in the initial contaminated site conceptual model.

2. The method for constructing a conceptual model of a contaminated site according to claim 1, wherein Based on the feedback information provided by the stakeholders, performing a site prior semantic panoramic scan on the set of site prior information model parameter row vectors to determine the key information gaps existing in the initial contaminated site conceptual model, including: Using a text encoder to perform text structured encoding on the feedback information provided by the stakeholders to obtain a feedback information semantic encoding vector; Inputting the feedback information semantic encoding vector and the set of site prior information model parameter row vectors into a feedback information - model prior semantic panoramic scan network to obtain a feedback information - site prior semantic scan response encoding vector; Performing feature decoding on the feedback information - site prior semantic scan response encoding vector to obtain the identification result of the key information gap.

3. The method for constructing a conceptual model of a contaminated site according to claim 2, wherein The text encoder is a pre - trained language model based on the Transformer architecture.

4. The method for constructing a conceptual model of a contaminated site according to claim 3, wherein Inputting the feedback information semantic encoding vector and the set of site prior information model parameter row vectors into a feedback information - model prior semantic panoramic scan network to obtain a feedback information - site prior semantic scan response encoding vector, including: Performing semantic query response encoding on the feedback information semantic encoding vector respectively with each site prior information model parameter row vector in the set of site prior information model parameter row vectors to obtain a set of feedback information - model prior semantic query score encoding vectors; Based on the feature set self - distribution characteristics of the set of feedback information - model prior semantic query score encoding vectors, performing semantic matching gating aggregation on the set of feedback information - model prior semantic query score encoding vectors to obtain the feedback information - site prior semantic scan response encoding vector.

5. The method for constructing a conceptual model of a contaminated site according to claim 4, wherein Performing semantic query response encoding on the feedback information semantic encoding vector respectively with each site prior information model parameter row vector in the set of site prior information model parameter row vectors to obtain a set of feedback information - model prior semantic query score encoding vectors, including: Performing feature enhancement on the feedback information semantic encoding vector based on deconvolutional encoding to obtain a feedback information semantic enhanced encoding vector, where the feedback information semantic enhanced encoding vector has the same feature scale as each site prior information model parameter row vector in the set of site prior information model parameter row vectors; The semantic enhanced encoding vectors of the feedback information are respectively subjected to single semantic query encoding with each site prior information model parameter row vector in the set of site prior information model parameter row vectors to obtain the set of feedback information-model prior semantic query score encoding vectors.

6. The method for constructing a conceptual model of a contaminated site according to claim 5, characterized in that, The semantic enhanced encoding vectors of the feedback information are respectively subjected to single semantic query encoding with each site prior information model parameter row vector in the set of site prior information model parameter row vectors to obtain the set of feedback information-model prior semantic query score encoding vectors, including: The semantic enhanced encoding vector of the feedback information and the site prior information model parameter row vector are cascaded and fused and then input into a neural network layer based on the tanh function to obtain the feedback information-model prior semantic query score encoding vector.

7. The method for constructing a conceptual model of a contaminated site according to claim 6, characterized in that, Based on the self-distribution characteristic of the feature set of the set of feedback information-model prior semantic query score encoding vectors, semantic matching gating aggregation is performed on the set of feedback information-model prior semantic query score encoding vectors to obtain the feedback information-site prior semantic scan response encoding vector, including: Based on the self-distribution characteristic of the feature set of the set of feedback information-model prior semantic query score encoding vectors, the single semantic matching degree of each feedback information-model prior semantic query score encoding vector in the set of feedback information-model prior semantic query score encoding vectors is determined to obtain the set of feedback information-model prior semantic matching degrees; Based on the set of feedback information-model prior semantic matching degrees, global semantic gating aggregation is performed on the set of feedback information-model prior semantic query score encoding vectors to obtain the feedback information-site prior semantic scan response encoding vector.

8. The method for constructing a conceptual model of a contaminated site according to claim 7, characterized in that, Based on the self-distribution characteristic of the feature set of the set of feedback information-model prior semantic query score encoding vectors, the single semantic matching degree of each feedback information-model prior semantic query score encoding vector in the set of feedback information-model prior semantic query score encoding vectors is determined to obtain the set of feedback information-model prior semantic matching degrees, including: Based on the semantic feature interaction between the semantic encoding vector of the feedback information and each site prior information model parameter row vector in the set of site prior information model parameter row vectors, symmetry constraint optimization is performed on each corresponding feedback information-model prior semantic query score encoding vector to obtain the set of optimized feedback information-model prior semantic query score encoding vectors; The context semantic association degree of each optimized feedback information-model prior semantic query score encoding vector in the set of optimized feedback information-model prior semantic query score encoding vectors relative to other optimized feedback information-model prior semantic query score encoding vectors is calculated as the single semantic matching degree to obtain the set of feedback information-model prior semantic matching degrees.

9. The method for constructing a conceptual model of a contaminated site according to claim 8, characterized in that, Based on the set of the feedback information-model prior semantic matching degrees, globally semantically gate-aggregate the set of the feedback information-model prior semantic query score encoding vectors to obtain the feedback information-site prior semantic scan response encoding vectors, including: Input the set of the feedback information-model prior semantic matching degrees into a relational gating proxy module to obtain a set of feedback information-model prior query semantic self-attention weights; Aggregate the set of the optimized feedback information-model prior semantic query score encoding vectors based on the set of the feedback information-model prior query semantic self-attention weights to obtain the feedback information-site prior semantic scan response encoding vectors.

10. A construction system for a conceptual model of a contaminated site, characterized in that, Including: A contaminated site concept model construction module, configured to construct an initial contaminated site concept model, where the initial contaminated site concept model is used to describe known or speculated pollution sources, possible pollutant migration pathways, and potential exposure pathways and receptors; A feedback information collection module, configured to collect feedback information provided by stakeholders, where the feedback information includes the requirements, expectations, and concerns about the site; An information decoupling module, configured to extract a site prior information model parameter matrix from the initial contaminated site concept model, and decouple the site prior information model parameter matrix by row vectors to obtain a set of site prior information model parameter row vectors; A key information gap identification module, configured to perform a site prior semantic panoramic scan on the set of the site prior information model parameter row vectors based on the feedback information provided by the stakeholders to determine the key information gaps existing in the initial contaminated site concept model.

Citation Information

Patent Citations

  • Suspected contaminated site spatio-temporal information identification method

    CN111651432A

  • Constructing method and device of polluted site knowledge graph

    CN115525766A

  • Polluted site multi-source heterogeneous data fusion method

    CN117593614A

  • Conceptual model construction method and device for risk assessment of pollution site

    CN118246725A

  • Atmospheric pollution control scheme evaluation method, device and equipment and readable storage medium

    CN118446374A

Cited By

  • Intelligent construction method of a contaminated site remediation assessment model

    CN122508489A