New pollutant classification prediction method and related device
By collecting and standardizing the chemical information of pollutants, using Tanimoto similarity calculation and graph neural network to simulate migration paths, and combining the environmental factor correction model, a multi-mechanism comprehensive evaluation of new pollutants is achieved, which improves the classification accuracy and reliability of risk assessment.
Patent Information
- Application Number
- CN202511117013.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-08-11
AI Technical Summary
Traditional classification methods have difficulty in dynamically quantifying the migration and transformation process of new pollutants in the environment, ignore the synergistic effects of multiple mechanisms and the influence of environmental factors, and lack a comprehensive evaluation of core mechanisms such as persistence and bioaccumulation, resulting in low classification accuracy.
The chemical information of target pollutants is collected and standardized, and the migration path is simulated using Tanimoto similarity calculation and graph neural network. Combined with the environmental factor correction model, a multi-mechanism comprehensive evaluation is performed to calculate the final classification prediction value.
It improves the accuracy and reliability of the classification of new pollutants, provides a scientific basis for risk assessment, and comprehensively considers the migration, transformation and accumulation processes of pollutants under different environmental conditions.
Smart Images

Figure CN120597054A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of new pollutant classification, and in particular to a new pollutant classification prediction method and related devices. Background Art
[0002] With the acceleration of industrialization, emerging contaminants (ECs), such as drug residues and microplastics, pose a serious challenge to ecological security due to their complex structures, variable environmental behavior, and multiple mechanisms of harm. Traditional classification methods rely on molecular fingerprint matching or static physical and chemical indicators, making it difficult to dynamically quantify the migration and transformation of pollutants in the environment. They also ignore the synergistic effects of multiple mechanisms and the dynamic impact of environmental factors. Existing technologies generally have two major limitations: first, data models sever the connection between the inherent properties of pollutants and environmental parameters; second, there is a lack of a comprehensive evaluation system for core mechanisms such as persistence and bioaccumulation. Therefore, how to improve the classification accuracy of new pollutants has become a technical problem that needs to be solved urgently. Summary of the Invention
[0003] In order to improve the classification accuracy of new pollutants, the present application provides a new pollutant classification prediction method and related devices.
[0004] In the first aspect, the present application provides a new pollutant classification prediction method using the following technical solutions: A new pollutant classification prediction method, comprising: S1: Collect chemical information of target pollutants, including molecular structure description, molecular fingerprint, and physical and chemical properties, and standardize the chemical information; S2: Calculate the standardized chemical information based on the Tanimoto similarity calculation method to obtain calculation results, and perform weighted integration based on the existing classification data and the calculation results to preliminarily predict the classification of the target pollutant and generate a preliminary classification assessment result; S3: Using graph neural networks to simulate the migration path of the target pollutants in the environment and analyze the migration, transformation, and degradation behaviors of the target pollutants in different environmental media to generate analysis results; S4: Based on the analysis results, comprehensively considering the persistence, mobility and bioaccumulation of the target pollutant, performing a multi-mechanism comprehensive evaluation by weighted summation to calculate the comprehensive mechanism characteristic score corresponding to the target pollutant; S5: Introducing an environmental factor correction model to correct the intrinsic characteristic scores of target pollutants based on temperature, pH value, dissolved oxygen concentration, humidity and light intensity; S6: Calculating a final classification prediction value of the target pollutant based on the preliminary classification assessment result, the analysis result, the intrinsic characteristic score, and the environmental correction result; S7: The risk level of the target pollutant is divided according to the final classification prediction value, and the classification result is output.
[0005] Optionally, the step of collecting chemical information of the target pollutant, the chemical information including molecular structure description, molecular fingerprint and physical and chemical properties, and standardizing the chemical information includes: Collect target pollutants The chemical information is represented by the molecular structure description in the chemical information as a set consisting of a series of atoms: Each atom Has its unique chemical properties ; Use MACCS or Morgan fingerprint to convert molecules into vectors of fixed dimension. The molecular fingerprint of is a vector; Among them, each component Indicates the presence or absence of a certain structure or functional group; Standardized to express physical and chemical properties : in and are the mean and standard deviation of the features respectively.
[0006] Optionally, the Tanimoto similarity calculation formula in step S2 is: Among them, the target pollutants The molecular fingerprint of , the database The molecular fingerprint of a known classified compound is , inner product: , vector norm: .
[0007] Optionally, migration path analysis is implemented using a graph neural network, specifically including: The environment is divided into multiple nodes, each node represents a specific environmental medium, and the characteristics of each node include the migration rate, adsorption capacity, and degradation rate of pollutants; Graph neural networks are used to model the migration paths of pollutants. The network input includes the physical and chemical properties of pollutants and environmental factors to calculate the migration probability between different environmental nodes and output the migration path and expected residence time of pollutants in the environment.
[0008] Optionally, the comprehensive mechanism characteristic score is calculated using the following formula: in, 、 、 are expressed as the weights of persistence, migration and bioaccumulation mechanisms, Indicates pollutants, 、 and Represents the scores of pollutants in terms of persistence, mobility and bioaccumulation respectively.
[0009] Optionally, in step S5, the environmental correction model is calculated using the following formula: in, is the environmental correction factor, and the intrinsic characteristics of the pollutant are scored as ; Environmental parameters are recorded as , Score the inherent properties of the pollutant.
[0010] Optionally, the final classification prediction value is calculated using the following formula: in Represents the preliminary classification evaluation results based on molecular structure, : represents the analysis results considering migration path and transformation behavior, represents the intrinsic characteristic score evaluated based on multiple classification mechanisms, represents the classification characteristics after environmental correction; Weight parameter All meet the following requirements: 。
[0011] In a second aspect, the present application provides a new pollutant classification prediction system, comprising: An information collection module is used to collect chemical information of target pollutants, including molecular structure description, molecular fingerprint, and physical and chemical properties, and to standardize the chemical information; An information collection module is used to collect chemical information of target pollutants, including molecular structure description, molecular fingerprint, and physical and chemical properties, and to standardize the chemical information; A preliminary classification assessment result module is used to calculate the standardized chemical information based on the Tanimoto similarity calculation method to obtain a calculation result, and to perform weighted integration based on the existing classification data and the calculation result to preliminarily predict the classification of the target pollutant and generate a preliminary classification assessment result; An analysis result module is used to use a graph neural network to simulate the migration path of the target pollutant in the environment and analyze the migration, transformation and degradation behavior of the target pollutant in different environmental media to generate analysis results; a comprehensive mechanism characteristic scoring module, for comprehensively considering the persistence, mobility, and bioaccumulation of the target pollutant based on the analysis results, and performing a multi-mechanism comprehensive evaluation by weighted summation to calculate the comprehensive mechanism characteristic score corresponding to the target pollutant; The intrinsic property scoring module is used to introduce an environmental factor correction model to correct the intrinsic property score of the target pollutant based on temperature, pH value, dissolved oxygen concentration, humidity and light intensity; A final classification prediction value module, configured to calculate a final classification prediction value of a target pollutant based on the preliminary classification evaluation result, the analysis result, the intrinsic characteristic score, and the environmental correction result; The output module is used to classify the risk level of the target pollutants according to the final classification prediction value and output the classification results.
[0012] In a third aspect, the present application provides a computer device, comprising: a memory and a processor, wherein the processor executes the method described above when running computer instructions stored in the memory.
[0013] In a fourth aspect, the present application provides a computer-readable storage medium comprising instructions, which, when executed on a computer, enable the computer to execute the method described above.
[0014] In summary, this application has the following beneficial technical effects: This application collects and standardizes chemical information (molecular structure, molecular fingerprint, and physical and chemical properties) of target pollutants; performs preliminary classification based on Tanimoto similarity calculations; utilizes graph neural networks to simulate the migration pathways and transformation behaviors of pollutants in the environment; conducts a comprehensive multi-mechanism evaluation that considers persistence, mobility, and bioaccumulation; introduces an environmental factor correction model to modify the score based on parameters such as temperature, pH, and dissolved oxygen concentration; integrates molecular structure, migration pathway, mechanism characteristics, and environmental correction results to calculate the final classification prediction; and categorizes risk levels (high, medium, and low) based on the predictions. This comprehensive assessment of pollutant risks through machine learning technology improves assessment accuracy and reliability, providing a scientific basis for pollution prevention and control. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a schematic diagram of the computer device structure of the hardware operating environment involved in the embodiment of the present application; Figure 2 This is a flow chart of the first embodiment of the new pollutant classification prediction method of the present application; Figure 3 This is a structural block diagram of the first embodiment of the new pollutant classification prediction system of this application. DETAILED DESCRIPTION
[0016] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below through the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0017] Reference Figure 1 , Figure 1 This is a schematic diagram of the computer device structure of the hardware operating environment involved in the embodiment of the present application.
[0018] like Figure 1As shown, the computer device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display and an input unit, such as a keyboard. Optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a wireless fidelity (Wi-Fi) interface). The memory 1005 may be a high-speed random access memory (RAM) or a stable non-volatile memory (NVM), such as a disk storage device. The memory 1005 may also be a storage device independent of the processor 1001.
[0019] Those skilled in the art will understand that Figure 1 The structure shown in the figure does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0020] like Figure 1 As shown, the memory 1005 as a storage medium may include an operating system, a network communication module, a user interface module, and a new pollutant classification prediction program.
[0021] exist Figure 1 In the computer device shown, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in this application can be set in the computer device, and the computer device calls the new pollutant classification prediction program stored in the memory 1005 through the processor 1001, and executes the new pollutant classification prediction method provided in the embodiment of this application.
[0022] This application embodiment provides a new pollutant classification prediction method, referring to Figure 2 , Figure 2 This is a flow chart of the first embodiment of the new pollutant classification prediction method of this application.
[0023] In this embodiment, the new pollutant classification prediction method includes the following steps: S1: Collect chemical information of target pollutants, including molecular structure description, molecular fingerprint, and physical and chemical properties, and standardize the chemical information.
[0024] It is understandable that traditional new pollutant classification and prediction methods, such as QSAR models and molecular docking technology, although they can provide some characteristic information to a certain extent, face the following challenges: Data scarcity: For new pollutants, there is a lack of experimental data and ready-made classification models.
[0025] High cost and low efficiency: Existing classification experiments and simulation methods require a large amount of experimental data support and are usually costly and have a long experimental cycle.
[0026] Poor comprehensiveness: Existing technologies cannot effectively combine multi-dimensional data of pollutants (such as chemical structure, environmental behavior, migration path, etc.) for comprehensive analysis.
[0027] Insufficient prediction of environmental behavior: Existing methods find it difficult to fully consider the migration, transformation and accumulation processes of pollutants under different environmental conditions.
[0028] It should be noted that the step of collecting chemical information of target pollutants, including molecular structure description, molecular fingerprint and physical and chemical properties, and standardizing the chemical information includes: collecting target pollutants The chemical information is represented by the molecular structure description in the chemical information as a set consisting of a series of atoms: Each atom Has its unique chemical properties ; Use MACCS or Morgan fingerprint to convert molecules into vectors of fixed dimension. The molecular fingerprint of is a vector; Among them, each component Indicates the presence or absence of a certain structure or functional group; Standardized to express physical and chemical properties : in and are the mean and standard deviation of the features respectively.
[0029] S2: Calculate the standardized chemical information based on the Tanimoto similarity calculation method to obtain calculation results, and perform weighted integration based on the existing classification data and the calculation results to preliminarily predict the classification of the target pollutants and generate preliminary classification evaluation results.
[0030] It is understandable that the Tanimoto similarity calculation formula in step S2 is: Among them, the target pollutants The molecular fingerprint of , the database The molecular fingerprint of a known classified compound is , inner product: , vector norm: .
[0031] It should be noted that the initial classification prediction is made by using an existing classification database combined with molecular structure similarity. The key here is to use a data-driven approach to infer the potential category of the target pollutant by comparing the classification labels of compounds with similar structures.
[0032] In the specific implementation, the steps of weighted integration of classified data include: assuming that there are Individual and Structurally similar compounds, each All have known classification labels (e.g. persistent organic pollutants, endocrine disruptors, etc.), the initial predicted category of the target molecule can be weighted voting: Where I(·) is the indicator function, is the set of all possible classification labels.
[0033] S3: Use graph neural networks to simulate the migration path of the target pollutants in the environment and analyze the migration, transformation and degradation behaviors of the target pollutants in different environmental media to generate analysis results.
[0034] It should be noted that in step S3, the migration path analysis is implemented through a graph neural network, which specifically includes: dividing the environment into multiple nodes, each node represents a specific environmental medium, and the characteristics of each node include the migration rate, adsorption capacity, and degradation rate of the pollutant; using a graph neural network to model the migration path of the pollutant, and the network input includes the physical and chemical properties and environmental factors of the pollutant to calculate the migration probability between different environmental nodes, and output the migration path and expected residence time of the pollutant in the environment.
[0035] In specific implementation, the steps for evaluating environmental behavior and migration paths include: 1. Migration Path Analysis The migration paths of pollutants in the environment are primarily influenced by their physical and chemical properties (such as solubility, volatility, and hydrophilicity) and environmental conditions (such as temperature, pH, and wind speed). To accurately predict the migration paths of pollutants in different environmental media (such as water, soil, and atmosphere), this example uses a graph neural network (GNN) for simulation. The specific process is as follows: 1.1 Environmental Media Modeling: The environment is divided into multiple nodes, each representing a specific environmental medium (e.g., water, soil, atmosphere, etc.). The characteristics of each node include the migration rate, adsorption capacity, and degradation rate of pollutants.
[0036] 1.2 Migration Path Calculation: Graph Neural Networks are used to model the migration paths of pollutants. The network inputs include the physical and chemical properties of the pollutants and environmental factors (such as temperature, humidity, and pH). GNNs are used to calculate the migration probabilities of pollutants between various environmental nodes, deriving the migration paths of pollutants in the environment.
[0037] 1.3 Migration path output: The model will output the migration probability distribution of pollutants in different environmental media, as well as the expected residence time and migration range of pollutants in each environmental medium.
[0038] 2. Conversion behavior modeling: Pollutants may undergo transformation in the environment, such as photodegradation, biodegradation or chemical reaction. This embodiment defines the transformation function To describe the transformation behavior of pollutants, the following steps are used for modeling; 2.1 Definition of transformation mechanism: The transformation behavior of pollutants is determined by different transformation mechanisms, such as chemical degradation, photodegradation, microbial degradation, etc. The transformation rate of each mechanism can be estimated from experimental data or literature.
[0039] 2.2 Prediction of transformation products: pollutants It may be transformed into multiple products in the environment. Each transformation mechanism will produce a certain set of transformation products. , these transformation products may affect the final classification of pollutants.
[0040] 2.3 Transformation Impact Assessment: By establishing a dynamic model of pollutant transformation, combined with environmental conditions and the initial properties of the pollutants, the impact of transformation on the classification characteristics of the pollutants is assessed. For example, some pollutants may become less bioaccumulative after biodegradation, while some transformation products may have higher toxicity.
[0041] S4: Based on the analysis results, the persistence, mobility and bioaccumulation of the target pollutants are comprehensively considered, and a multi-mechanism comprehensive evaluation is performed through weighted summation to calculate the comprehensive mechanism characteristic score corresponding to the target pollutant.
[0042] It is understood that in step S4, the comprehensive mechanism characteristic score is calculated by the following formula: in, 、 、 are expressed as the weights of persistence, migration and bioaccumulation mechanisms, Indicates pollutants, 、 and Represents the scores of pollutants in terms of persistence, mobility and bioaccumulation respectively.
[0043] In the specific implementation, mechanism-related classification modeling; Each mechanism has a different impact on how pollutants behave during the classification process. We consider the following main mechanisms: Persistence: This indicates how long a pollutant persists in the environment, usually expressed as a half-life or degradation rate. The persistence of a pollutant is closely related to its chemical stability and environmental conditions (such as temperature and light).
[0044] Mobility: This describes the ability of a pollutant to migrate between different media. The mobility of a pollutant is affected by its physicochemical properties (such as partition coefficient and solubility). Pollutants with high mobility are more likely to diffuse in the environment.
[0045] Bioaccumulation: This refers to the ability of a pollutant to accumulate in organisms, usually measured by the bioaccumulation factor (BAF). Pollutants with high bioaccumulation potential may have long-term impacts on ecosystems.
[0046] In specific implementation, the steps of weight setting and optimization include: weight setting for mechanisms such as persistence, migration and bioaccumulation, which can be optimized through machine learning methods so that these weights can better adapt to the performance of pollutants in actual experimental data or simulated environments. The optimization goal is to minimize the error between the comprehensive score and the actual observation value. The comprehensive mechanism score is , which is defined as: in, 、 、 are the weights of persistence, migration and bioaccumulation, respectively. 、 and It is a pollutant Scoring under these mechanisms.
[0047] The goal of optimization is to adjust these weights so that the prediction results are as close as possible to the actual classification labels or experimental data of the pollutants. The specific objective function is: in, Indicates the The true classification score of each pollutant (experimental observation data or expert labels), is the number of samples.
[0048] S5: Introduce an environmental factor correction model to correct the intrinsic characteristic scores of target pollutants based on temperature, pH value, dissolved oxygen concentration, humidity and light intensity.
[0049] It should be noted that, in step S5, the environmental correction model is calculated using the following formula: in, is the environmental correction factor, and the intrinsic characteristics of the pollutant are scored as ; Environmental parameters are recorded as , Score the inherent properties of the pollutant.
[0050] S6: Calculate the final classification prediction value of the target pollutant based on the preliminary classification evaluation result, the analysis result, the intrinsic characteristic score and the environmental correction result.
[0051] It is understood that the final classification prediction value is calculated by the following formula: in Represents the preliminary classification evaluation results based on molecular structure, : represents the analysis results considering migration path and transformation behavior, represents the intrinsic characteristic score evaluated based on multiple classification mechanisms, represents the classification characteristics after environmental correction; Weight parameter All meet the following requirements: 。
[0052] S7: The risk level of the target pollutant is divided according to the final classification prediction value, and the classification result is output.
[0053] In specific implementation, according to The value of can be used to classify pollutants into different levels: in and The threshold is determined based on experimental data or industry standards.
[0054] Similarity calculation and preliminary classification: Based on the Tanimoto similarity calculation method, the standardized chemical information is calculated and weighted integration is performed in combination with existing classification data to preliminarily predict the classification of target pollutants.
[0055] Migration path simulation: Use graph neural network models to simulate the migration paths of target pollutants in the environment and analyze their migration, transformation, and degradation behaviors in different environmental media.
[0056] Multi-mechanism comprehensive evaluation: Taking into account the persistence, mobility and bioaccumulation of target pollutants, a multi-mechanism comprehensive evaluation is conducted through weighted summation to calculate the comprehensive mechanism characteristic score.
[0057] Environmental factor correction: An environmental factor correction model is introduced to correct the intrinsic characteristic scores of target pollutants based on temperature, pH value, dissolved oxygen concentration, humidity and light intensity.
[0058] Final classification prediction: Integrate molecular structure, migration pathway, mechanism characteristics and environmental correction results to calculate the final classification prediction value of the target pollutant.
[0059] Risk level classification: Based on the final classification prediction value, the target pollutant is classified into three risk levels: high risk, medium risk, and low risk.
[0060] This example collects and standardizes chemical information (molecular structure, molecular fingerprint, and physicochemical properties) of target pollutants; performs preliminary classification based on Tanimoto similarity calculations; utilizes graph neural networks to simulate the migration pathways and transformation behaviors of pollutants in the environment; conducts a comprehensive multi-mechanism evaluation that considers persistence, mobility, and bioaccumulation; introduces an environmental factor correction model to modify the score based on parameters such as temperature, pH, and dissolved oxygen concentration; integrates molecular structure, migration pathway, mechanistic characteristics, and environmental correction results to calculate a final classification prediction; and categorizes risk levels (high, medium, and low) based on the predictions. This comprehensive assessment of pollutant risk through machine learning improves assessment accuracy and reliability, providing a scientific basis for pollution prevention and control.
[0061] The contents of the second embodiment of this application include: Step 1: Collect chemical information of target pollutants, including molecular structure description, molecular fingerprint, and physical and chemical properties, and standardize the chemical information.
[0062] Specifically, the molecular structure data of the target pollutants are input into the standardization processing module for standardization. The standardization method can adopt the Z-Score standardization method or the Min-Max standardization method. The attention mechanism is used to find the substructures that contribute most to the predicted attribute value. The attention mechanism can adopt a multi-head self-attention mechanism. The problem of insufficient sample size is solved through parameter migration. The parameter migration method can adopt a parameter migration method based on domain adversarial. An interspecific extrapolation normalization model is established to increase the number of species in the derivation of soil environmental benchmark values, making the derived benchmark values more scientific and reliable, thereby protecting more species. The interspecific extrapolation normalization model can adopt an interspecific extrapolation model based on machine learning.
[0063] Step 2: Calculate the standardized chemical information based on the Tanimoto similarity calculation method to obtain calculation results, and perform weighted integration based on the existing classification data and the calculation results to preliminarily predict the classification of the target pollutant.
[0064] The similarity coefficient of the Tanimoto similarity calculation method ranges from 0 to 1. The closer the similarity coefficient is to 1, the higher the similarity. A similarity coefficient greater than 0.8 is considered highly similar. When weighted integration is performed, a linear weighting method can be used, and the weights can be determined through grid search.
[0065] Step 3: Use the graph neural network model to simulate the migration path of the target pollutant in the environment, and analyze the migration, transformation and degradation behavior of the target pollutant in different environmental media to generate analysis results.
[0066] The graph neural network model uses a graph convolutional neural network model based on the attention mechanism. The model input is the molecular structure data of pollutants and environmental parameter data, and the output is the migration path and transformation and degradation of pollutants in different environmental media. Environmental media can include the atmosphere, water, soil, etc.
[0067] Step 4: Based on the analysis results, the persistence, mobility and bioaccumulation of the target pollutants are comprehensively considered, and a multi-mechanism comprehensive evaluation is performed by weighted summation to calculate the comprehensive mechanism characteristic score corresponding to the target pollutant.
[0068] The persistence score ranges from 0 to 1, the mobility score ranges from 0 to 1, and the bioaccumulation score ranges from 0 to 1, and the sum of the weighting factors is 1. The weighting factors can be set to 0.4, 0.3, and 0.3 respectively.
[0069] Step 5: Introduce an environmental factor correction model to modify the intrinsic property score of the target pollutant based on temperature, pH, dissolved oxygen concentration, humidity, and light intensity. This environmental factor correction model uses a nonlinear regression model based on random forests, with environmental parameter data as input and correction coefficients as output. Environmental parameter data can include a temperature range of 10-35°C, a pH range of 5-9, and a dissolved oxygen concentration range of 2-10 mg / L.
[0070] Step 6: Integrate the molecular structure, migration pathway, mechanistic characteristics, and environmental correction results to calculate the final classification prediction value of the target pollutant. This integration is performed using a weighted summation method, with weights set to 0.3, 0.3, 0.2, and 0.2, respectively. The weights were determined using a 5-fold cross-validation method.
[0071] Step 7: Based on the final classification prediction value, the target pollutant is classified into three risk levels: high risk, medium risk, and low risk. The corresponding classification prediction value ranges are 0.7-1, 0.4-0.7, and 0-0.4, respectively.
[0072] This implementation achieves the following: Chemical information collection and standardization: Chemical information on target pollutants, including molecular structure descriptions, molecular fingerprints, and physical and chemical properties, is collected and standardized. An attention mechanism is used to identify substructures that contribute significantly to predicted attribute values, and parameter transfer is used to address insufficient sample size. An interspecies extrapolation normalization model is established to increase the number of species used in the derivation of soil environmental baseline values, making the derived baseline values more scientific and reliable.
[0073] In addition, an embodiment of the present application also proposes a computer-readable storage medium, on which a program for classifying and predicting new pollutants is stored. When the program for classifying and predicting new pollutants is executed by a processor, the steps of the method for classifying and predicting new pollutants as described above are implemented.
[0074] Reference Figure 3 , Figure 3 This is a structural block diagram of the first embodiment of the new pollutant classification prediction system of this application.
[0075] like Figure 3 As shown, the new pollutant classification prediction system proposed in the embodiment of the present application includes: The information collection module 10 is used to collect chemical information of the target pollutant, including molecular structure description, molecular fingerprint and physical and chemical properties, and standardize the chemical information; A preliminary classification evaluation result module 20 is configured to calculate the standardized chemical information based on the Tanimoto similarity calculation method to obtain a calculation result, perform weighted integration based on the existing classification data and the calculation result to preliminarily predict the classification of the target pollutant and generate a preliminary classification evaluation result; An analysis result module 30 is used to simulate the migration path of the target pollutant in the environment using a graph neural network and analyze the migration, transformation, and degradation behavior of the target pollutant in different environmental media to generate analysis results; a comprehensive mechanism characteristic scoring module 40 for comprehensively considering the persistence, mobility, and bioaccumulation of the target pollutant based on the analysis results, and performing a multi-mechanism comprehensive evaluation by weighted summation to calculate a comprehensive mechanism characteristic score corresponding to the target pollutant; An intrinsic characteristic scoring module 50 is used to introduce an environmental factor correction model to correct the intrinsic characteristic score of the target pollutant based on temperature, pH value, dissolved oxygen concentration, humidity and light intensity; A final classification prediction value module 60 is used to calculate a final classification prediction value of a target pollutant based on the preliminary classification evaluation result, the analysis result, the intrinsic characteristic score, and the environmental correction result; The output module 70 is used to classify the risk level of the target pollutants according to the final classification prediction value and output the classification result.
[0076] It should be understood that the above is only an example and does not constitute any limitation to the technical solution of the present application. In specific applications, technicians in this field can make settings as needed, and the present application does not impose any restrictions on this.
[0077] This example collects and standardizes chemical information (molecular structure, molecular fingerprint, and physicochemical properties) of target pollutants; performs preliminary classification based on Tanimoto similarity calculations; utilizes graph neural networks to simulate the migration pathways and transformation behaviors of pollutants in the environment; conducts a comprehensive multi-mechanism evaluation that considers persistence, mobility, and bioaccumulation; introduces an environmental factor correction model to modify the score based on parameters such as temperature, pH, and dissolved oxygen concentration; integrates molecular structure, migration pathway, mechanistic characteristics, and environmental correction results to calculate a final classification prediction; and categorizes risk levels (high, medium, and low) based on the predictions. This comprehensive assessment of pollutant risk through machine learning improves assessment accuracy and reliability, providing a scientific basis for pollution prevention and control.
[0078] It should be noted that the workflow described above is merely illustrative and does not limit the scope of protection of this application. In actual applications, technicians in this field can select part or all of it according to actual needs to achieve the purpose of this embodiment scheme, and no restrictions are imposed here.
[0079] In addition, for technical details not fully described in this embodiment, please refer to the method for new pollutant classification prediction provided in any embodiment of this application, and will not be repeated here.
[0080] In addition, it should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or system. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or system comprising the element.
[0081] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0082] Through the above description of the embodiments, those skilled in the art will clearly understand that the above-mentioned embodiments and methods can be implemented using software plus the necessary general-purpose hardware platform. Of course, hardware can also be used, but in many cases the former is a more preferred embodiment. Based on this understanding, the technical solution of this application, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory (ROM) / RAM, a magnetic disk, or an optical disk) and includes several instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of this application. The above are only preferred embodiments of this application and do not limit the scope of the patent application. Any equivalent structure or equivalent process transformation made using the contents of this application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the scope of patent protection of this application.
Claims
1. A new pollutant classification prediction method, characterized in that: include: S1: Collect chemical information of target pollutants, including molecular structure description, molecular fingerprint, and physical and chemical properties, and standardize the chemical information; S2: Calculate the standardized chemical information based on the Tanimoto similarity calculation method to obtain calculation results, and perform weighted integration based on the existing classification data and the calculation results to preliminarily predict the classification of the target pollutant and generate a preliminary classification assessment result; S3: Using graph neural networks to simulate the migration path of the target pollutants in the environment and analyze the migration, transformation, and degradation behaviors of the target pollutants in different environmental media to generate analysis results; S4: Based on the analysis results, comprehensively considering the persistence, mobility and bioaccumulation of the target pollutant, performing a multi-mechanism comprehensive evaluation by weighted summation to calculate the comprehensive mechanism characteristic score corresponding to the target pollutant; S5: Introducing an environmental factor correction model to correct the intrinsic characteristic scores of target pollutants based on temperature, pH value, dissolved oxygen concentration, humidity and light intensity; S6: Calculating a final classification prediction value of the target pollutant based on the preliminary classification assessment result, the analysis result, the intrinsic characteristic score, and the environmental correction result; S7: The risk level of the target pollutant is divided according to the final classification prediction value, and the classification result is output.
2. The method according to claim 1, characterized in that The step of collecting chemical information of the target pollutant, the chemical information including molecular structure description, molecular fingerprint and physical and chemical properties, and standardizing the chemical information includes: Collect target pollutants The chemical information is represented by the molecular structure description in the chemical information as a set consisting of a series of atoms: Each atom Has its unique chemical properties ; Use MACCS or Morgan fingerprint to convert molecules into vectors of fixed dimension. The molecular fingerprint of is a vector; Among them, each component Indicates the presence or absence of a certain structure or functional group; Standardized to express physical and chemical properties : in and are the mean and standard deviation of the features respectively.
3. The method according to claim 1, characterized in that The Tanimoto similarity calculation formula in step S2 is: Among them, the target pollutants The molecular fingerprint of , the database The molecular fingerprint of a known classified compound is , inner product: , vector norm: .
4. The method according to claim 1, wherein In step S3, migration path analysis is implemented through a graph neural network, specifically including: The environment is divided into multiple nodes, each node represents a specific environmental medium, and the characteristics of each node include the migration rate, adsorption capacity, and degradation rate of pollutants; Graph neural networks are used to model the migration paths of pollutants. The network input includes the physical and chemical properties of pollutants and environmental factors to calculate the migration probability between different environmental nodes and output the migration path and expected residence time of pollutants in the environment.
5. The method according to claim 1, wherein In step S4, the comprehensive mechanism characteristic score is calculated using the following formula: in, 、 、 are expressed as the weights of persistence, migration and bioaccumulation mechanisms, Indicates pollutants, 、 and Represents the scores of pollutants in terms of persistence, mobility and bioaccumulation respectively.
6. The method according to claim 1, wherein In step S5, the environmental correction model is calculated using the following formula: in, is the environmental correction factor, and the intrinsic characteristics of the pollutant are scored as ; Environmental parameters are recorded as , Score the inherent properties of the pollutant.
7. The method according to claim 1, characterized in that The final classification prediction value is calculated by the following formula: in Represents the preliminary classification evaluation results based on molecular structure, : represents the analysis results considering migration path and transformation behavior, represents the intrinsic characteristic score evaluated based on multiple classification mechanisms, represents the classification characteristics after environmental correction; Weight parameter All meet the following requirements: 。 8. A new pollutant classification prediction system, characterized by: include: An information collection module is used to collect chemical information of target pollutants, including molecular structure description, molecular fingerprint, and physical and chemical properties, and to standardize the chemical information; A preliminary classification assessment result module is used to calculate the standardized chemical information based on the Tanimoto similarity calculation method to obtain a calculation result, and to perform weighted integration based on the existing classification data and the calculation result to preliminarily predict the classification of the target pollutant and generate a preliminary classification assessment result; An analysis result module is used to use a graph neural network to simulate the migration path of the target pollutant in the environment and analyze the migration, transformation and degradation behavior of the target pollutant in different environmental media to generate analysis results; a comprehensive mechanism characteristic scoring module, for comprehensively considering the persistence, mobility, and bioaccumulation of the target pollutant based on the analysis results, and performing a multi-mechanism comprehensive evaluation by weighted summation to calculate the comprehensive mechanism characteristic score corresponding to the target pollutant; The intrinsic property scoring module is used to introduce an environmental factor correction model to correct the intrinsic property score of the target pollutant based on temperature, pH value, dissolved oxygen concentration, humidity and light intensity; A final classification prediction value module, configured to calculate a final classification prediction value of a target pollutant based on the preliminary classification evaluation result, the analysis result, the intrinsic characteristic score, and the environmental correction result; The output module is used to classify the risk level of the target pollutants according to the final classification prediction value and output the classification results.
9. A computer device, characterized in that: The device comprises: a memory and a processor, wherein the processor executes the method according to any one of claims 1 to 7 when running computer instructions stored in the memory.
10. A computer-readable storage medium, characterized in that The method comprises instructions which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Screening method and system for emphatically controlling new pollutants
CN117095830A
Regional pollutant assessment method, device and equipment and storage medium
CN118229090A
Deep learning-based environmental pollution risk assessment method and system, computing device and storage medium
CN119294842A
Remediation treatment method and system for simulating cadmium-polluted soil and storage medium
CN119811525A
Ecological risk assessment method and system for heavy metal contaminated soil
CN120258538A
Cited By
Machine learning combined new pollutant rapid screening and risk assessment method and system
CN121766777A