A method, device, computer device and storage medium for identifying emerging pollutants

By integrating non-target analysis and computational toxicology methods, the application sequence and execution strategies are determined, the problem of new pollutant risk assessment under the condition of lack of toxicity experimental data is solved, efficient and accurate identification and risk assessment of new pollutants are achieved, and the targetedness and efficiency of control measures are improved.

CN119400278BActive Publication Date: 2025-08-01HANGZHOU INST FOR ADVANCED STUDY UCAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510001254.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-08-01
Estimated Expiration
2045-01-02

AI Technical Summary

Technical Problem

The prior art is difficult to achieve reliable new pollutant risk assessment under the lack of toxicity experimental data, and it is difficult to obtain compound data sets for computational toxicology methods.

Method used

By integrating non-target analysis methods and computational toxicology methods, the application sequence and target execution strategies are determined, and the risk assessment of new pollutants, including sample collection, mass spectrometry analysis, data mining and toxicity prediction, obtaining multiple risk values to determine the target new pollutants.

Benefits of technology

It improves the efficiency and accuracy of identifying new pollutants, avoids waste of resources, and can identify new pollutants that are of great threat to the environment and human health in a targeted manner, improving the efficiency and effectiveness of the implementation of control measures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119400278B_ABST
    Figure CN119400278B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of environmental analysis, and discloses a method, device, computer device and storage medium for identifying emerging pollutants. The present invention determines the target execution strategy by determining the application order of each non-target analysis method and each computational toxicology method in the high-throughput fusion method. Whether starting from the toxic effect to find the corresponding chemical substances or evaluating the toxic effect from the discovered chemical substances, it can be achieved by reasonably selecting and switching the execution strategy. Further, using the target execution strategy to conduct risk assessment of emerging pollutants and calculate multiple risk values, taking into account various characteristics of emerging pollutants and their performance in the actual environment, can quantitatively reflect the potential risk level of emerging pollutants from different perspectives. Finally, combining multiple risk values to further determine various target emerging pollutants, realizing the screening of those objects that pose a greater threat to the environment and human health and need to be prioritized for control among numerous possible emerging pollutants.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of environmental analysis technology, and in particular to a new pollutant identification method, device, computer equipment and storage medium. Background Art

[0002] The large-scale production and use of synthetic chemicals has greatly improved the quality of human life, but it has also led to their widespread entry into the environment, posing a serious threat to ecological safety and human health. The production and use of toxic and hazardous chemicals are the main sources of new pollutants. Over the past few decades, the growth trend of synthetic chemicals has been obvious and the growth rate has continued to accelerate. On the one hand, the use of numerous synthetic chemicals means a huge number of new pollutants. On the other hand, toxicity experimental data for most new pollutants is lacking. Therefore, it is crucial to develop a high-throughput fusion technology process that can comprehensively identify new pollutants in complex environmental media and human matrices and complete risk assessment.

[0003] Analytical methods and evaluation means are the foundation and core of the development of the above-mentioned processes. With the development and progress of modern high-resolution mass spectrometry instruments, non-target analysis has become one of the key chemical analysis methods for the comprehensive identification of new pollutants in complex matrices. It does not rely on authentic standards and can complete the simultaneous qualitative and quantitative analysis of thousands of new pollutants. How to achieve reliable risk assessment of new pollutants through this analysis in the absence of toxicity experimental data is a key technical challenge. Computational toxicology can quickly and accurately predict the toxic effects of massive compounds at different scales. How to obtain compound data sets is the primary technical challenge faced in using computational toxicology for toxicity prediction. Summary of the Invention

[0004] In view of this, the present invention provides a method, apparatus, computer equipment and storage medium for identifying new pollutants to solve the problems in the prior art of being unable to analyze and implement reliable new pollutant risk assessments in the absence of toxicity experimental data, and of being unable to obtain compound data sets in computational toxicology methods.

[0005] In a first aspect, the present invention provides a method for identifying new pollutants, the method comprising:

[0006] Obtain a high-throughput fusion method, which is obtained by fusing at least one non-target analysis method and at least one computational toxicology method; determine the application order of each non-target analysis method and each computational toxicology method in the high-throughput fusion method; based on the application order, determine the target execution strategy of the high-throughput fusion method, which is one or more of the toxic effect to chemical substance strategy and the chemical substance to toxic effect strategy; use the target execution strategy to conduct risk assessment of new pollutants to obtain multiple risk values; based on the multiple risk values, determine multiple target new pollutants.

[0007] The new pollutant identification method provided by the present invention integrates non-target analysis methods and computational toxicology methods, breaking the limitations of single methods, fully integrating their respective resources and capabilities, and realizing the synergistic effect of multiple methods. Further, by determining the application order of each non-target analysis method and each computational toxicology method in the high-throughput integration method and further determining the target execution strategy, the workflow can be optimized according to specific circumstances, work efficiency can be improved, and unnecessary repeated operations and resource waste can be avoided. At the same time, whether starting from toxic effects to find corresponding chemical substances or evaluating the toxic effects of the discovered chemical substances, it can be achieved by reasonably selecting and switching execution strategies, enhancing the adaptability to diverse practical problems. Further, using the determined target execution strategy for new pollutant risk assessment, various characteristics of new pollutants and their performance in the actual environment are comprehensively considered, and multiple risk values are obtained through scientific calculation and analysis, which can quantitatively reflect the potential risk level of new pollutants from different perspectives. Finally, combining multiple risk values to further determine multiple target new pollutants, realizing the screening of those new pollutants that pose a greater threat to the environment and human health and need to be prioritized for control among numerous possible new pollutants, avoiding the indiscriminate control of all new pollutants, enabling resources to be concentrated on key control objects, improving the implementation efficiency and effect of control measures, and specifically reducing the adverse effects of new pollutants on the ecological environment and human health, providing clear directions and priorities for environmental monitoring, treatment, and government supervision work.

[0008] In an alternative embodiment, the target execution strategy is the toxic effect to chemical substance strategy; using the target execution strategy for new pollutant risk assessment, multiple risk values are obtained, including:

[0009] Using computational toxicology methods to screen and determine the range of new pollutants that meet the selected toxicity endpoints from a preset database; collecting samples based on the range of new pollutants to obtain multiple initial new pollutant samples; respectively processing and performing mass spectrometry analysis on the multiple initial new pollutant samples to obtain a target mass spectrometry analysis data set; based on the range of new pollutants, performing data mining analysis and toxicity prediction on the target mass spectrometry analysis data set to obtain the new pollutant concentration level and toxicity prediction results; calculating multiple risk values based on the new pollutant concentration level and toxicity prediction results.

[0010] The new pollutant identification method provided by the present invention, when the target execution strategy is the toxicity effect to chemical substance strategy, uses computational toxicology methods to screen and determine the range of new pollutants that meet the selected toxicity endpoints from a preset database. This can quickly exclude a large number of substances irrelevant to this toxicity characteristic, narrow the investigation scope, improve the efficiency of locating new pollutants among a vast number of possible substances, avoid directionless comprehensive screening, and save resource and time costs. Further, sample collection is combined with the range of new pollutants, improving the effectiveness and representativeness of the samples. Further, the initially collected new pollutant samples are processed and subjected to mass spectrometry analysis, and after data mining analysis and toxicity prediction of the obtained target mass spectrometry analysis data set, multiple risk values are calculated, making the risk assessment results closer to the actual situation, avoiding assessment biases caused by only focusing on single factors such as concentration or toxicity, and being able to measure the potential risk levels of new pollutants to the ecological environment and human health more accurately from multiple perspectives.

[0011] In an alternative embodiment, multiple initial new pollutant samples are respectively processed and subjected to mass spectrometry analysis to obtain a target mass spectrometry analysis data set, including:

[0012] Obtaining multiple sample matrix characteristics of multiple initial new pollutant samples; based on the range of new pollutants and multiple sample matrix characteristics, using a preset sample pretreatment method to process multiple initial new pollutant samples to obtain multiple target new pollutant samples; subjecting multiple target new pollutant samples to mass spectrometry analysis to obtain an initial mass spectrometry analysis data set; and preprocessing the initial mass spectrometry analysis data set to obtain a target mass spectrometry analysis data set.

[0013] The new pollutant identification method provided by the present invention uses a preset sample pretreatment method to process multiple initial new pollutant samples to obtain multiple target new pollutant samples, which can fully consider factors such as the chemical properties of the target new pollutants and the interfering components that may exist in the samples, reducing analysis errors caused by factors such as matrix effects. Further, subjecting multiple target new pollutant samples to mass spectrometry analysis can obtain an initial mass spectrometry analysis data set that more truly and clearly reflects the characteristics of the target new pollutants. Finally, preprocessing the initial mass spectrometry analysis data set improves the quality and usability of the target mass spectrometry analysis data set.

[0014] In an alternative embodiment, based on the range of new pollutants, data mining analysis and toxicity prediction are performed on the target mass spectrometry analysis data set to obtain the new pollutant concentration level and toxicity prediction results, including:

[0015] Based on the scope of emerging pollutants, data mining is performed on the target mass spectrometry analysis dataset using non-target analysis methods to obtain multiple secondary emerging pollutants; structural annotation is performed on the multiple secondary emerging pollutants to obtain multiple chemical structures; semi-quantitative analysis is performed on the multiple chemical structures to obtain the concentration levels of emerging pollutants; and the target mass spectrometry analysis dataset is processed by computational toxicology methods to obtain toxicity prediction results.

[0016] The emerging pollutant identification method provided by the present invention, based on the scope of emerging pollutants, uses non-target analysis methods to perform data mining on the target mass spectrometry analysis dataset, which can more accurately lock in the target and improve the probability of discovering true emerging pollutants. Further, performing structural annotation on multiple secondary emerging pollutants can provide a deeper understanding of key information such as the molecular composition, chemical bond types, and functional group distributions of emerging pollutants. Further, through semi-quantitative analysis of chemical structures, the actual content of emerging pollutants in specific environments or human matrices can be quantified, avoiding the one-sidedness of only evaluating risks from a qualitative perspective. Further, obtaining toxicity prediction results through computational toxicology method processing can reasonably predict the toxicity effects of emerging pollutants in the absence of actual toxicity experimental data.

[0017] In an alternative embodiment, the target execution strategy is the chemical substance to toxicity effect strategy; using the target execution strategy for emerging pollutant risk assessment to obtain multiple risk values, including:

[0018] Obtain multiple initial emerging pollutant samples; process and perform mass spectrometry analysis on the multiple initial emerging pollutant samples respectively to obtain the target mass spectrometry analysis dataset; perform data mining analysis and toxicity prediction on the target mass spectrometry analysis dataset to obtain the concentration levels of emerging pollutants and toxicity prediction results; calculate multiple risk values based on the concentration levels of emerging pollutants and toxicity prediction results.

[0019] The emerging pollutant identification method provided by the present invention, when the target execution strategy is the chemical substance to toxicity effect strategy, processes and performs mass spectrometry analysis on the multiple obtained initial emerging pollutant samples, performs data mining analysis and toxicity prediction on the obtained target mass spectrometry analysis dataset, and then calculates multiple risk values, making the risk assessment results closer to the actual situation, avoiding the evaluation deviation caused by only focusing on a single factor such as concentration or toxicity, and being able to more accurately measure the potential risk of emerging pollutants to the ecological environment and human health from multiple perspectives.

[0020] In an alternative embodiment, performing data mining analysis and toxicity prediction on the target mass spectrometry analysis dataset to obtain the concentration levels of emerging pollutants and toxicity prediction results includes:

[0021] The target mass spectrometry data set was mined and analyzed using non-target analysis methods to obtain the concentration levels of new pollutants. The target mass spectrometry data set was processed using computational toxicology methods to obtain toxicity prediction results.

[0022] The new pollutant identification method provided by the present invention uses a non-target analysis method to perform data mining analysis to obtain the concentration level of the new pollutant, and can more accurately determine the content of the new pollutant in the actual sample based on complex mass spectrometry data. Furthermore, the toxicity prediction results are obtained by processing the data set by relying on computational toxicology methods. In the absence of a large amount of actual toxicity experimental data, the toxicity of the new pollutant can be reasonably estimated based on the existing data characteristics and related models. Therefore, through the implementation of the present invention, a clear division of labor avoids the functional overlap and duplication of non-target analysis methods and computational toxicology methods, so that each link can focus on the task it is good at, and then rationally allocate resources, reduce unnecessary calculations, analysis steps and time costs, and ensure the reliability of the final concentration level and toxicity prediction results.

[0023] In an optional embodiment, multiple target new pollutants are determined based on multiple risk values, including:

[0024] Determine multiple new pollutant risk contribution values based on multiple risk values; and determine multiple target new pollutants based on multiple risk values and multiple new pollutant risk contribution values.

[0025] The new pollutant identification method provided by the present invention determines multiple new pollutant risk contribution values through multiple risk values, which can determine the contribution degree of each new pollutant to the overall risk, and thus can reflect the potential threat of each new pollutant to the ecological environment and human health. Furthermore, by combining multiple risk values and multiple new pollutant risk contribution values to determine multiple target new pollutants, it is possible to accurately identify new pollutants that contribute more to the overall risk and pose more prominent potential threats to the ecological environment and human health, avoiding indiscriminate control of all new pollutants, allowing resources to be concentrated on key control objects, improving the implementation efficiency and effectiveness of control measures, and targetedly reducing the adverse effects of new pollutants on the ecological environment and human health, providing a clear direction and focus for environmental monitoring, governance, and government supervision.

[0026] In a second aspect, the present invention provides a new pollutant identification device, the device comprising:

[0027] An acquisition module for acquiring a high-throughput fusion method obtained by fusing at least one non-target analysis method and at least one computational toxicology method; a first determination module for determining the application order of each non-target analysis method and each computational toxicology method in the high-throughput fusion method; a second determination module for determining a target execution strategy of the high-throughput fusion method based on the application order, the target execution strategy being one or more of a toxicity effect to chemical substance strategy and a chemical substance to toxicity effect strategy; an evaluation module for performing new pollutant risk assessment using the target execution strategy to obtain a plurality of risk values; and a third determination module for determining a plurality of target new pollutants based on the plurality of risk values.

[0028] In a third aspect, the present invention provides a computer device, including: a memory and a processor, which are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the new pollutant identification method according to the first aspect or any corresponding embodiment thereof.

[0029] In a fourth aspect, the present invention provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the new pollutant identification method according to the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0031] Figure 1 is a flowchart of the new pollutant identification method according to an embodiment of the present invention;

[0032] Figure 2 is a flowchart of another new pollutant identification method according to an embodiment of the present invention;

[0033] Figure 3 is a flowchart of yet another new pollutant identification method according to an embodiment of the present invention;

[0034] Figure 4 is a detailed flowchart of high-throughput non-target analysis and computational toxicology fusion according to an embodiment of the present invention;

[0035] Figure 5 is a structural block diagram of the new pollutant identification device according to an embodiment of the present invention;

[0036] Figure 6 It is a schematic diagram of the hardware structure of the computer device according to an embodiment of the present invention. Specific embodiments

[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0038] In non-target analysis, how to obtain the toxicity endpoints of new pollutants lacking toxicity experimental data is a technical difficulty in using non-target analysis for risk assessment; while computational toxicology can predict various toxicity endpoints of new pollutants, and the application difficulty lies in obtaining the dataset of new pollutants to be predicted in actual samples. Therefore, the embodiments of the present invention provide a new pollutant identification method, which develops a high-throughput fusion method that combines the two methods based on the technical characteristics of non-target analysis and the methods and means of computational toxicology to achieve the comprehensive identification and risk assessment of new pollutant mixtures in complex real environmental scenarios.

[0039] According to an embodiment of the present invention, an embodiment of a new pollutant identification method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0040] In this embodiment, a new pollutant identification method is provided, which can be used in electronic devices such as computers, mobile phones, and tablet computers. Figure 1 It is a flowchart of the new pollutant identification method according to an embodiment of the present invention, as Figure 1 shown, and the process includes the following steps:

[0041] Step S101, obtain a high-throughput fusion method.

[0042] Specifically, the high-throughput fusion method can be obtained by fusing at least one non-target analysis method and at least one computational toxicology method.

[0043] Among them, non-target analysis represents a chemical analysis method that does not rely on true reference standards and can perform synchronous qualitative and quantitative analysis of numerous substances in complex matrices, and can be completed through procedures such as sample collection, sample pretreatment, instrumental analysis, data preprocessing, data mining, structure annotation, and (semi)-quantitative analysis.

[0044] Computational toxicology methods represent a high-throughput method that uses computer technology and mathematical models to predict the toxicity of chemical substances, and may include (quantitative) structure-activity relationship models, machine learning, molecular docking, molecular dynamics simulations, etc.

[0045] Step S102, determine the application order of each non-target analysis method and each computational toxicology method in the high-throughput fusion method.

[0046] Specifically, the application order of each non-target analysis method and each computational toxicology method in the high-throughput fusion method can be determined by comprehensively considering various factors, such as research objectives, existing data basis, and the degree of prior understanding of samples and pollutants.

[0047] For example, if the research aims to quickly identify the presence of new pollutants with potentially high-risk toxicity in actual samples, and some toxicity-related theoretical models and public database resources are already available, then it may be inclined to first use computational toxicology methods for preliminary screening, that is, to carry out computational toxicology-related operations first in terms of the application order.

[0048] If there are certain monitoring clues about the chemical substances that may be contained in a certain type of environmental sample (such as soil and water samples in a specific area) in the early stage, but the toxicity information is lacking, the non-target analysis method can be preferentially selected in the application order at this time.

[0049] In this embodiment, the above application order is not specifically limited and can be determined according to the actual situation.

[0050] Step S103, based on the application order, determine the target execution strategy of the high-throughput fusion method.

[0051] Among them, the target execution strategy is one or more of the toxicity effect to chemical substance strategy and the chemical substance to toxicity effect strategy.

[0052] Specifically, the high-throughput fusion method can be divided into a "top-down" (i.e., toxicity effect to chemical substance) strategy and a "bottom-up" (i.e., chemical substance to toxicity effect) strategy according to the application order.

[0053] Furthermore, "top-down" starts from computational toxicology predictions, then samples are collected, and non-target analysis is subsequently completed to achieve the purpose of risk assessment; "bottom-up" starts from sample collection, and then non-target analysis and computational toxicology predictions are carried out in sequence to achieve the purpose of risk assessment.

[0054] In this embodiment, "top-down" and "bottom-up" are two independent execution strategies, without a sequence or logical relationship, that is, the two execution strategies can be implemented separately or used in series, and there is no prior or subsequent application order when used in series.

[0055] Specifically, using them separately means achieving comprehensive identification and risk assessment of new pollutants in a sample only by adopting the "top-down" or "bottom-up" execution strategy; using them in series means achieving comprehensive identification and risk assessment of new pollutants in a sample by separately adopting at least one sub-step of one "top-down" and one "bottom-up" fusion strategy at the same time; further, in the case of using them in series, the "top-down" strategy can be ahead or the "bottom-up" strategy can be ahead, and it can be freely selected.

[0056] Step S104: Use the target execution strategy to conduct risk assessment of new pollutants and obtain multiple risk values.

[0057] Specifically, according to the description of step S104, a determined target execution strategy can be adopted to conduct risk assessment of new pollutants. By comprehensively considering various characteristics of new pollutants and their performance in the actual environment, multiple risk values are obtained through scientific calculation and analysis, which can quantitatively reflect the potential risk levels of new pollutants from different perspectives.

[0058] Step S105: Based on multiple risk values, determine multiple target new pollutants.

[0059] Specifically, combining the obtained multiple risk values can further determine multiple target new pollutants, realizing the screening of those objects that pose a greater threat to the environment and human health and need to be preferentially controlled among numerous possible new pollutants, avoiding the indiscriminate control of all new pollutants, enabling resources to be concentrated on key control objects, improving the implementation efficiency and effect of control measures, and specifically reducing the adverse impacts of new pollutants on the ecological environment and human health.

[0060] The new pollutant identification method provided in this embodiment integrates non-target analysis methods and computational toxicology methods, breaking the limitations of single methods, fully integrating their respective resources and capabilities, and achieving the synergistic effect of multiple methods. Further, by determining the application order of each non-target analysis method and each computational toxicology method in the high-throughput integration method and further determining the target execution strategy, the workflow can be optimized according to specific situations, work efficiency can be improved, and unnecessary repeated operations and resource waste can be avoided. At the same time, whether starting from toxic effects to find corresponding chemical substances or evaluating the toxic effects of discovered chemical substances, it can be achieved by reasonably selecting and switching execution strategies, enhancing the adaptability to diverse practical problems. Further, using the determined target execution strategy for new pollutant risk assessment, considering various characteristics of new pollutants and their performance in the actual environment, multiple risk values are obtained through scientific calculations and analyses, which can quantitatively reflect the potential risk levels of new pollutants from different perspectives. Finally, combining multiple risk values to further determine multiple target new pollutants, realizing the screening of those new pollutants that pose greater threats to the environment and human health and require priority control among numerous possible new pollutants, avoiding the indiscriminate control of all new pollutants, enabling resources to be concentrated on key control objects, improving the implementation efficiency and effect of control measures, and specifically reducing the adverse impacts of new pollutants on the ecological environment and human health, providing clear directions and priorities for environmental monitoring, governance, and government supervision work.

[0061] In this embodiment, a new pollutant identification method is provided, which can be used in electronic devices such as computers, mobile phones, and tablet computers. Figure 2 It is a flowchart of the new pollutant identification method according to an embodiment of the present invention, as Figure 2 shown, and this process includes the following steps:

[0062] Step S201, obtain the high-throughput integration method. For details, please refer to Figure 1 Step S101 of the embodiment shown, which will not be elaborated here.

[0063] Step S202, determine the application order of each non-target analysis method and each computational toxicology method in the high-throughput integration method. For details, please refer to Figure 1 Step S102 of the embodiment shown, which will not be elaborated here.

[0064] Step S203, based on the application order, determine the target execution strategy of the high-throughput integration method. For details, please refer to Figure 1 Step S103 of the embodiment shown, which will not be elaborated here.

[0065] Step S204, use the target execution strategy to conduct new pollutant risk assessment to obtain multiple risk values.

[0066] Specifically, when the target execution strategy is the toxicity effect to chemical substance strategy, the above step S204 includes:

[0067] Step S2041, using computational toxicology methods to screen and determine the range of new pollutants that meet the selected toxicity endpoints from a preset database.

[0068] Among them, the selected toxicity endpoints can be determined according to the research purpose and the actual direction of concern. For example, if the focus is on the impact of new pollutants on the aquatic ecosystem, the toxicity endpoints may be set as acute toxicity to fish, interference with the reproductive ability of aquatic organisms, etc.; if the focus is on the harm to human health, the toxicity endpoints can be carcinogenicity, teratogenicity, damage to specific organs (such as the liver, kidneys, etc.).

[0069] Furthermore, the preset database can be a publicly available chemical database (such as PubChem, ChemSpider, etc., which contain basic information such as the structures and properties of a large number of chemical substances), or a database of specific types of chemical substances accumulated within an enterprise, a collection of chemical substances related to a specific industry field, etc., and can be determined according to the actual research scenario and the available data resources.

[0070] Specifically, according to the description in step S101, methods such as (quantitative) structure-activity relationship models (QSAR / QSPR), machine learning, molecular docking, and molecular dynamics simulations can be used to screen out multiple new pollutants that meet the selected toxicity endpoints from the preset database.

[0071] In an alternative embodiment, taking the (quantitative) structure-activity relationship model (QSAR / QSPR) as an example, if there is a large amount of structural and corresponding toxicity data of known chemical substances as a basis, the QSAR / QSPR model can be used. First, analyze these data to find the quantitative relationship between chemical structure features (such as the size, shape, types and distributions of functional groups of molecules, etc.) and the selected toxicity endpoints, and establish a mathematical model. Then, substitute the structural parameters of numerous chemical substances in the preset database into this model, calculate the activity prediction values related to the selected toxicity endpoints for them, and screen out those chemical substances whose predicted activities meet the requirements (indicating that they may have corresponding toxicities).

[0072] For example, it is known that certain compounds containing specific functional groups are toxic to the liver. After analyzing this structure-toxicity association through the QSAR model, for other compounds in the preset database, check whether there are similar functional groups and related structural features in their structures to determine whether they may be toxic to the liver, and then screen them out.

[0073] Furthermore, the corresponding range of new pollutants can be determined based on multiple new pollutants.

[0074] Step S2042: Based on the scope of emerging pollutants, sample collection is carried out to obtain multiple initial emerging pollutant samples.

[0075] Specifically, the possible main media types of emerging pollutants in the environment can be judged by studying the characteristics of the substances included in the determined scope of emerging pollutants. For example, if the scope of emerging pollutants is some organic compounds with low water solubility, high volatility and easy adsorption to particulate matter, then the emerging pollutants are likely to exist more in the atmospheric environment, possibly adsorbed on atmospheric particulate matter or existing in gaseous form; if the scope of emerging pollutants is defined as some substances that are difficult to degrade, hydrophilic and easily combined with soil particles, then the soil will be one of its main storage media; for those substances that can accumulate in organisms and have bioaccumulation, it is necessary to focus on collecting from biological samples (such as tissues in organisms like fish and birds).

[0076] Furthermore, according to the above analysis results, the types of samples to be collected can be reasonably determined. For example, when it is judged that the emerging pollutants mainly exist in the water environment, the sample type can be selected as water samples, and then it can be further subdivided into water samples from different sources, such as industrial sewage (for specific industrial discharge areas, which may contain relevant emerging pollutants generated during the production process), influent and effluent of sewage treatment plants (which can reflect the changes in pollutants during the sewage treatment process and the possible remaining emerging pollutants), river water (representing the pollution status of natural water bodies affected by the surrounding environment), surface water (covering the surface conditions of a wider range of natural water areas), seawater (for emerging pollutants that may exist in coastal areas affected by marine-related activities), drinking water (related to human health, detecting whether there are emerging pollutants in it), groundwater (affected by factors such as soil infiltration, and may contain some pollutants accumulated over a long time), etc.

[0077] Furthermore, after determining the sample type, a suitable area / group can be further selected for sample collection. For example, for water sample collection, if it is industrial sewage, sample collection can be carried out at positions near the corresponding industrial discharge outlet, the front end of sewage treatment facilities, etc.

[0078] Furthermore, after determining the sampling area and group, the elements such as the time, frequency, and quantity of sample collection can be carefully planned, and a complete sample collection plan can be formed.

[0079] Furthermore, according to the set sample collection plan, suitable sampling tools and methods can be selected to carry out sample collection operations and obtain the corresponding multiple initial emerging pollutant samples.

[0080] Step S2043: Process and perform mass spectrometry analysis on multiple initial emerging pollutant samples respectively to obtain a target mass spectrometry analysis data set.

[0081] In some alternative embodiments, step S2043 includes:

[0082] Step a1, obtaining the sample matrix characteristics of a plurality of initial emerging pollutant samples.

[0083] Step a2, based on the scope of emerging pollutants and the sample matrix characteristics of the plurality of samples, using a preset sample pretreatment method to process the plurality of initial emerging pollutant samples to obtain a plurality of target emerging pollutant samples.

[0084] Step a3, performing mass spectrometry analysis on the plurality of target emerging pollutant samples to obtain an initial mass spectrometry analysis data set.

[0085] Step a4, preprocessing the initial mass spectrometry analysis data set to obtain a target mass spectrometry analysis data set.

[0086] Among them, the sample matrix characteristics represent the sample type; the preset sample pretreatment method can be methods such as ultrasonic oscillation, Soxhlet extraction, solid phase extraction, liquid-liquid extraction, solid phase microextraction, liquid phase microextraction, pressurized liquid extraction, QuECHERS, etc.

[0087] Specifically, according to the environmental medium classification described in step S2042, it can be determined whether the plurality of initial emerging pollutant samples are water samples, solid samples, gas samples, biological samples, etc. For example, it is determined that the plurality of initial emerging pollutant samples collected are river water samples, sewage treatment plant sludge solid samples, fish biological samples, etc.

[0088] Furthermore, the description of the matrix characteristics can be refined. For example, for water samples, further clarify their specific sources and related characteristics. For river water samples, it is necessary to understand the flowing areas of the river (whether it passes through industrial pollution areas, agricultural planting areas, etc.), water flow, water temperature, acidity and alkalinity (pH value), etc.; for sewage treatment plant sludge solid samples, it is necessary to know the main sources of the sewage treated by the sewage treatment plant (mainly domestic sewage or a large proportion of industrial sewage, etc.) and the texture and moisture content of the sludge; for fish biological samples, information such as the type of fish, growth environment (freshwater or seawater environment, whether it is in a polluted water area, etc.), size and age of the fish should be recorded.

[0089] Furthermore, an appropriate sample pretreatment method can be selected according to the chemical properties (such as solubility, volatility, polarity, etc.) of the target emerging pollutants determined by the emerging pollutant scope and the matrix characteristics of the initial sample. For example, if the emerging pollutant scope consists of some organic compounds with relatively strong polarity and the initial sample is a river water sample, considering that the water sample may contain a large amount of impurities and the concentration of the target substances is low, the liquid-liquid extraction method can be selected. An organic solvent that is immiscible with water and has good solubility for the target organic compounds (such as dichloromethane, etc.) is chosen for extraction to transfer the target emerging pollutants from the water sample to the organic phase, achieving the purpose of enrichment and preliminary separation of impurities.

[0090] Furthermore, according to the selected sample pretreatment method, the treatment can be carried out strictly in accordance with the corresponding operation procedures and specifications. For example, when using ultrasonic oscillation, parameters such as the power and time of the ultrasonic wave need to be controlled to ensure that the target substances can be effectively desorbed without damaging their chemical structures; during the liquid-liquid extraction operation, the dosage of the extractant, the number of extraction times, and the intensity and time of oscillation need to be accurately controlled to ensure the extraction efficiency; in the processes of solid-phase extraction, solid-phase microextraction, etc., attention needs to be paid to the operation key points of the activation, sample loading, elution, etc. of the adsorbent to smoothly complete the sample pretreatment process and finally obtain multiple treated target emerging pollutant samples.

[0091] Furthermore, the precursor ions (or adduct ions) corresponding to the predefined emerging pollutants can be input into the inclusion list (online priority acquisition precursor ion library) in the instrumental analysis, so that when the instrument conducts analysis, for the substances in the list, they will be preferentially analyzed under the condition of the same signal intensity of the substances.

[0092] Meanwhile, according to the type of the target emerging pollutant samples (such as water samples, solid samples, biological samples, etc.) and the properties of the target substances to be analyzed, other parameters of the mass spectrometer are reasonably set, such as the type of ion source (electrospray ionization source ESI, atmospheric pressure chemical ionization source APCI, etc.), the scanning range (setting an appropriate scanning interval according to the approximate molecular weight range of the target substances), the resolution (selecting a high-resolution value that meets the analysis requirements), the scanning speed, etc., to ensure that the instrument is in the best working state and can accurately detect the mass spectrometry signals of the target emerging pollutants.

[0093] Furthermore, multiple target emerging pollutant samples obtained through pretreatment can be sequentially injected into the mass spectrometer for analysis. Among them, during the analysis process, different target emerging pollutants will be ionized in the ion source to form charged ions, and then in the mass analyzer, they will be separated and detected according to their different mass-to-charge ratios (m / z), generating mass spectrometry signal information containing various ions. These signal information reflect the characteristics such as the molecular weight and ion fragments of the target emerging pollutants. The mass spectrometer records these data and forms the corresponding initial mass spectrometry analysis data set.

[0094] Further, the initial mass spectrometry analysis data set can be preprocessed to obtain a target mass spectrometry analysis data set. Among them, the preprocessing can include operations such as filtering, deconvolution, peak alignment, null value filling, normalization, and feature annotation.

[0095] Step S2044: Based on the new pollutant range, perform data mining analysis and toxicity prediction on the target mass spectrometry analysis data set to obtain the new pollutant concentration level and toxicity prediction result.

[0096] In some alternative embodiments, the above step S2044 includes:

[0097] Step b1: Based on the new pollutant range, use non-target analysis methods to perform data mining on the target mass spectrometry analysis data set to obtain multiple second new pollutants.

[0098] Step b2: Perform structure annotation on the multiple second new pollutants to obtain multiple chemical structures.

[0099] Step b3: Perform semi-quantitative analysis on the multiple chemical structures to obtain the new pollutant concentration level.

[0100] Step b4: Process the target mass spectrometry analysis data set through computational toxicology methods to obtain the toxicity prediction result.

[0101] Among them, the data mining methods include but are not limited to suspected target screening, isotope labeling, mass defect filtering, homolog retrieval, case control, and characteristic fragment ion labeling; the structure annotation methods can include spectral database matching, spectral prediction, and characteristic fragmentation path analysis.

[0102] Further, the spectral databases include but are not limited to mzClould, massBank EU, MoNA, NIST, Wiley, GNPS, etc.

[0103] Specifically, using the new pollutant range as the screening database, mainly applying the suspected target screening method in non-target analysis, and at the same time, various structure-driven data mining methods such as isotope labeling, mass defect filtering, homolog retrieval, case control, and characteristic fragment ion labeling can be combined to carry out the mining work on the target mass spectrometry analysis data set.

[0104] For example, in the suspected target screening method, based on the characteristic ions (such as molecular ion peaks, characteristic fragment ion peaks, etc.) of the target substances determined in the scope of new pollutants and information such as their relative abundances in the mass spectrum, signal peaks that match or have a high similarity are searched for in the entire target mass spectrometry analysis dataset to preliminarily judge the possible new pollutants; the isotope labeling method relies on the isotope characteristics of certain elements in the target new pollutants, and by observing the signals with corresponding isotope peak ratios or difference relationships in the mass spectrometry data, it assists in determining the possible new pollutants; mass defect filtering can screen out the substances corresponding to the mass spectrometry signals that conform to the mass defect characteristics in the dataset according to the theoretical mass defect value range of the substances in the known new pollutant scope, so as to narrow the search scope and improve the mining efficiency, etc.

[0105] Furthermore, structural annotation methods such as spectral database matching, spectral prediction, and characteristic fragmentation pathway analysis can be used to determine the chemical structure of the second new pollutant. Specifically, the spectral database matching method can be utilized to compare the mass spectrometry graph of the mined second new pollutant with the standard spectra in known spectral databases (such as mzClould, massBank EU, MoNA, NIST, Wiley, GNPS, etc.), and search for the chemical structures corresponding to the spectra with high similarity as a reference to preliminarily determine its possible structure.

[0106] For example, upload the mass spectrometry graph of a certain second new pollutant to the NIST spectral database for retrieval and matching. If a standard spectrum with a high similarity is found and the chemical substance structure corresponding to the spectrum is clearly marked, then it can be preliminarily determined that the second new pollutant has a similar chemical structure.

[0107] Furthermore, for the second new pollutants that cannot find a completely matching spectrum in the spectral database, spectral prediction tools can be used to predict the possible chemical structure based on the fragment ion information in its mass spectrometry graph and some known chemical rules (such as the bond cleavage rules, characteristic cleavage methods of functional groups, etc.). At the same time, characteristic fragmentation pathway analysis can be combined to deeply study how the characteristic fragment ions generated during the mass spectrometry analysis of the substance are cleaved step by step from its parent ion, and further improve and confirm its chemical structure through reverse derivation.

[0108] Among them, the spectral prediction tools include but are not limited to CFM-ID, CSI:FingerID, MetFrag, MS-FINDER, Mass Frontier, Molecular Structure Correlator, MS-Fragmenter, MOLGEN-MS, etc.

[0109] Furthermore, based on the determined chemical structures, appropriate semi-quantitative analysis methods can be selected according to the properties of different chemical substances and the available analytical conditions to perform semi-quantitative analysis on multiple chemical structures, obtaining the corresponding concentration levels of emerging pollutants, thus avoiding the one-sidedness of only evaluating risks from a qualitative perspective.

[0110] Among them, the semi-quantitative analysis methods can be achieved based on structurally similar reference standards, chromatographic retention times close to reference standards, ionization efficiency prediction, machine learning, etc.

[0111] Finally, based on the target mass spectrometry analysis dataset, using information such as the first-order mass spectrum, second-order mass spectrum, and compound structure as input parameters, toxicity prediction is carried out through computational toxicology methods to obtain the corresponding toxicity prediction results, enabling a reasonable prediction of the toxicity effects of emerging pollutants in the absence of actual toxicity experimental data.

[0112] Step S2045: Calculate multiple risk values based on the concentration levels of emerging pollutants and the toxicity prediction results.

[0113] Specifically, by combining the obtained concentration levels of emerging pollutants and the toxicity prediction results, the risk value of each emerging pollutant can be calculated, avoiding the evaluation bias caused by only focusing on a single factor such as concentration or toxicity, and enabling a more accurate measurement of the potential risk of emerging pollutants to the ecological environment and human health from multiple perspectives.

[0114] Step S205: Determine multiple target emerging pollutants based on the multiple risk values. For details, please refer to Figure 1 Step S105 of the illustrated embodiment, which will not be elaborated here.

[0115] The new pollutant identification method provided in this embodiment, when the target execution strategy is the toxicity effect to chemical substance strategy, uses computational toxicology methods to screen and determine the range of new pollutants that meet the selected toxicity endpoints from a preset database. This can quickly exclude a large number of substances irrelevant to this toxicity characteristic, narrow the investigation scope, improve the efficiency of locating new pollutants among a vast number of possible substances, avoid unguided comprehensive screening, and save resource and time costs. Further, using the preset sample pretreatment method to process multiple initial new pollutant samples to obtain multiple target new pollutant samples can fully consider factors such as the chemical properties of the target new pollutants and the interfering components that may exist in the samples, reducing analysis errors caused by factors such as matrix effects. Further, performing mass spectrometry analysis on multiple target new pollutant samples can obtain an initial mass spectrometry analysis dataset that more truly and clearly reflects the characteristics of the target new pollutants. Finally, preprocessing the initial mass spectrometry analysis dataset improves the quality and usability of the target mass spectrometry analysis dataset. Further, based on the range of new pollutants, using non-target analysis methods to perform data mining on the target mass spectrometry analysis dataset can lock the target more accurately and increase the probability of identifying the true new pollutants. Further, annotating the structures of multiple second new pollutants can provide a deeper understanding of key information such as the molecular composition, chemical bond types, and functional group distributions of new pollutants. Further, through semi-quantitative analysis of the chemical structures, the actual content of new pollutants in specific environmental or human matrices can be quantified, avoiding the one-sidedness of only evaluating risks from a qualitative perspective. Further, by processing with computational toxicology methods to obtain toxicity prediction results, the toxicity effects of new pollutants can be reasonably predicted in the absence of actual toxicity experiment data.

[0116] In this embodiment, a new pollutant identification method is provided, which can be used in electronic devices such as computers, mobile phones, and tablet computers. Figure 3 It is a flowchart of the new pollutant identification method according to an embodiment of the present invention, as Figure 3 shown, and this process includes the following steps:

[0117] Step S301, obtain the high-throughput fusion method. For details, please refer to Figure 1 step S101 of the embodiment shown, which will not be elaborated here.

[0118] Step S302, determine the application order of each non-target analysis method and each computational toxicology method in the high-throughput fusion method. For details, please refer to Figure 1 step S102 of the embodiment shown, which will not be elaborated here.

[0119] Step S303, based on the application order, determine the target execution strategy of the high-throughput fusion method. For details, please refer to Figure 1 step S103 of the embodiment shown, which will not be elaborated here.

[0120] Step S304: Use the target execution strategy to conduct a risk assessment of emerging pollutants, obtaining multiple risk values.

[0121] Specifically, when the target execution strategy is the chemical substance-to-toxicity effect strategy, the above-mentioned step S304 includes:

[0122] Step S3041: Obtain multiple initial emerging pollutant samples.

[0123] Specifically, sample collection can be carried out through a pre-set collection plan to obtain the corresponding multiple initial emerging pollutant samples.

[0124] Step S3042: Process and perform mass spectrometry analysis on the multiple initial emerging pollutant samples respectively to obtain a target mass spectrometry analysis data set.

[0125] The specific process can refer to the description of the above step S2043 and will not be elaborated here.

[0126] Step S3043: Conduct data mining analysis and toxicity prediction on the target mass spectrometry analysis data set to obtain the emerging pollutant concentration level and toxicity prediction results.

[0127] In some optional embodiments, the above step S3043 includes:

[0128] Step c1: Use non-target analysis methods to conduct data mining analysis on the target mass spectrometry analysis data set to obtain the emerging pollutant concentration level.

[0129] The specific process can refer to the description of the above steps b1 to b3 and will not be elaborated here.

[0130] Step c2: Based on the target mass spectrometry analysis data set, through the processing of computational toxicology methods, obtain toxicity prediction results.

[0131] Specifically, key information related to the chemical structure of emerging pollutants can be extracted from the target mass spectrometry analysis data set. Further, the extracted key information can be sorted and quantified to convert it into a data format that can be used for analysis by computational toxicology methods.

[0132] Furthermore, according to the description of step S101, methods such as (quantitative) structure-activity relationship models, machine learning, molecular docking, and molecular dynamics simulations can be used to predict the toxicity of emerging pollutants from different angles and to different degrees, and corresponding toxicity prediction results can be obtained.

[0133] Step S3044: Based on the emerging pollutant concentration level and toxicity prediction results, calculate multiple risk values.

[0134] For the specific process, reference can be made to the description of step S2045 above, which will not be elaborated here.

[0135] Step S305: Based on multiple risk values, determine multiple target new pollutants.

[0136] Specifically, step S305 above includes:

[0137] Step S3051: Based on multiple risk values, determine multiple risk contribution values of new pollutants.

[0138] Specifically, by calculating the obtained multiple risk values (such as summing them up), the total risk value can be obtained. Further, this total risk value can represent the overall risk degree generated by all new pollutants in a specific environment or for a specific receptor group.

[0139] Furthermore, for each new pollutant, the proportion of its risk value in the total risk value, that is, the risk contribution value of the new pollutant, can be calculated.

[0140] Step S3052: Based on multiple risk values and multiple risk contribution values of new pollutants, determine multiple target new pollutants.

[0141] Specifically, the risk value of each new pollutant and the corresponding risk contribution value can be simultaneously taken into consideration, and all new pollutants can be comprehensively sorted according to the level of risk contribution and the magnitude of the risk value.

[0142] For example, assume there is another group of new pollutants A, B, C, and D, whose risk values are 0.2, 0.35, 0.3, and 0.15 respectively, and the corresponding risk contribution values are 20%, 35%, 30%, and 15% respectively (assuming the total risk is 1). Generally speaking, first, sort according to the risk contribution value, and the sorting is B (35%) > C (30%) > A (20%) > D (15%); if there are cases where the risk contribution values are the same, then further refer to the magnitude of the risk value for secondary sorting, so as to determine the order of the new pollutants and obtain a sorting result similar to B - C - A - D, so as to clarify the order of the degree of influence of different new pollutants on the overall risk.

[0143] Furthermore, according to the sorting result, some of the new pollutants ranked in the front can be selected as target new pollutants. These target new pollutants are usually substances with relatively large contributions to the total risk and relatively high risk values, that is, those substances that pose a more prominent potential threat to the ecological environment and human health among many new pollutants, and they will be the objects of priority control.

[0144] The new pollutant identification method provided in this embodiment, when the target execution strategy is the chemical substance to toxicity effect strategy, processes and performs mass spectrometry analysis on a plurality of acquired initial new pollutant samples to obtain a target mass spectrometry analysis data set, and then uses a non-target analysis method to perform data mining and analysis to obtain the new pollutant concentration level, and can accurately determine the content of new pollutants in actual samples based on complex mass spectrometry data. Further, by relying on computational toxicology methods to process the data set to obtain toxicity prediction results, it is possible to reasonably estimate the toxicity of new pollutants based on the existing data characteristics and relevant models even in the absence of a large amount of actual toxicity experiment data. Further, by comprehensively considering the new pollutant concentration level and toxicity prediction results to calculate multiple risk values, the risk assessment result is made closer to the actual situation, avoiding the evaluation deviation caused by only focusing on a single factor such as concentration or toxicity, and being able to measure the potential risk of new pollutants to the ecological environment and human health more accurately from multiple perspectives. Further, by determining multiple new pollutant risk contribution values through multiple risk values, it is possible to determine the contribution degree of each new pollutant to the overall risk, and thus reflect the potential threat of each new pollutant to the ecological environment and human health. Further, by combining multiple risk values and multiple new pollutant risk contribution values to determine multiple target new pollutants, it is possible to accurately identify new pollutants with a greater contribution to the overall risk and a more prominent potential threat to the ecological environment and human health, avoiding the indiscriminate control of all new pollutants, enabling resources to be concentrated on key control targets, improving the implementation efficiency and effect of control measures, and specifically reducing the adverse effects of new pollutants on the ecological environment and human health, providing a clear direction and focus for environmental monitoring, governance, and government supervision work.

[0145] In one example, the specific implementation processes of the separate implementation and serial use of two execution strategies are provided respectively.

[0146] Example 1:

[0147] Taking the in-laboratory test example, it focuses on realizing the comprehensive identification and risk assessment of new pollutants in samples by using the established "top-down" execution strategy, as Figure 4 shown, and its specific steps include:

[0148] Step 1: In order to identify endocrine disruptors in synthetic chemicals and their hazards to human health, a fusion method of computational toxicology and non-target analysis is adopted.

[0149] Step 2: After establishing the fusion method, based on the "top-down" fusion strategy, conduct a comprehensive identification and health hazard assessment of endocrine disruptors in the human body.

[0150] Step 3: First, based on the publicly available artificial synthetic chemical database, construct a structural database containing all the organic synthetic chemicals used in production (with the number of compounds being a). Use the molecular docking model to predict the compounds with endocrine disrupting effects, and obtain b compounds with different structures, which are predefined as the endocrine disruptor database (with the number of compounds being b).

[0151] Step 4: Collect serum samples from c individuals of different ages and genders in Area A, and extract the endocrine disruptors in the samples using the liquid-liquid extraction method.

[0152] Step 5: Select an ultra-high performance liquid chromatography and high-resolution mass spectrometry combined system for instrumental analysis according to the serum matrix characteristics and the physicochemical properties of the compounds in the endocrine disruptor database. The chromatographic column is Waters ACQUITY C18 (1.7 μm, 2.1×100 mm). The mass spectrometry data is acquired based on the data-dependent acquisition mode. In the mass spectrometry method, the adduct ion forms of the b compounds in the endocrine disruptor database are added to the precursor ion library for priority acquisition in the secondary spectrum to complete the data acquisition.

[0153] Step 6: Use software to preprocess the raw mass spectrometry data. Among them, the parameter settings involved in procedures such as filtering, deconvolution, peak alignment, missing value filling, normalization, and feature annotation adopt the software default values without modification. After preprocessing, a total of d mass spectrometry features are identified from the c serum samples.

[0154] Step 7: Construct a suspected target screening database (with the number of compounds being b) based on the endocrine disruptor database, and use the suspected target screening method in non-target analysis to confirm that e0 mass spectrometry features (e0 ≤ d) among the d mass spectrometry features belong to the compounds in the endocrine disruptor database. Further, perform structural annotation on each of the e0 mass spectrometry features one by one, and determine that e compounds (e ≤ e0) among them are positive matches, that is, e new pollutants of endocrine disruptor type are preliminarily identified from the serum samples.

[0155] Meanwhile, use the machine learning algorithm to identify f0 mass spectrometry features (f0 < d) belonging to endocrine disruptors from the mass spectrometry spectra corresponding to the d preprocessed mass spectrometry features. Further, perform structural annotation on each of the f0 mass spectrometry features one by one, and the definite or possible chemical structures of f1 mass spectrometry features (f1 ≤ f0) can be given. Use the molecular docking model in Step 3 to predict that f compounds (f ≤ f1) among the f1 compounds belong to endocrine disruptors, that is, f new pollutants of endocrine disruptor type are further identified from the serum samples.

[0156] By comparing the e and f endocrine disruptors identified by the two methods respectively, it was found that there were g endocrine disruptors that overlapped in the two datasets. Therefore, a total of (e + f - g) endocrine disruptors were finally identified from c serum samples.

[0157] By investigating the authentic standards available on the market, it was determined that among the new pollutants, there were species with authentic standards. Subsequently, the authentic standards were purchased, and the new pollutants were subjected to target confirmation. Among them, the species of new pollutants were confirmed to be present in the samples. Subsequently, quantitative analysis was performed on the confirmed new pollutants to give accurate concentration levels. For the new pollutants other than those without authentic standards, semi - quantitative analysis was carried out based on the ionization efficiency prediction model to give predicted concentration levels.

[0158] Step Eight: Using estrogen and androgen effects as toxicity targets, based on the quantitative and semi - quantitative concentration levels and their binding energies with the target receptors, calculate the endocrine disruption effect risks of the identified new pollutants in serum samples.

[0159] Step Nine: Since the number of identified endocrine disruptors exceeds 50, making it difficult to include all of them in daily monitoring and government supervision activities, based on the ToxPi ranking, the top 10 compounds in the ranking were selected and established as priority new pollutants for control.

[0160] Example 2:

[0161] Taking the laboratory test example, it focuses on achieving a comprehensive identification and risk assessment of new pollutants in samples using the established "bottom - up" implementation strategy. As Figure 4 shown, its specific steps include:

[0162] Step One: To identify per - and polyfluoroalkyl substances (PFAS) in the surface water of Region B and their ecological risks to local algae, daphnia, and fish, a combined method of non - target analysis and computational toxicology was used.

[0163] Step Two: After establishing the combined method, a comprehensive identification and ecological risk assessment of PFAS in the surface water of Region B were carried out based on the "bottom - up" integration strategy.

[0164] Step Three: Starting from the receiving points of surface water bodies polluted by point sources such as fluorochemical industrial parks and the effluents of domestic sewage treatment plants, a total of a surface water samples were collected along the water flow direction.

[0165] After filtering all samples through a 0.45 μm filter membrane, 500 mL of each sample was taken, and solid-phase extraction technology was used to complete the pretreatment, extraction, enrichment, and reconstitution of the samples.

[0166] According to the characteristics of the water matrix and the physical and chemical properties of PFAS, an ultra-high performance liquid chromatography and high-resolution mass spectrometry system was selected for instrumental analysis. The chromatographic column was Waters ACQUITY C18 (1.7 μm, 2.1×100 mm). Mass spectrometry data was obtained simultaneously based on two acquisition modes: data-dependent acquisition and data-independent acquisition.

[0167] The original mass spectrometry data was preprocessed using software. Among them, the parameter settings involved in procedures such as filtering, deconvolution, peak alignment, missing value filling, normalization, and feature annotation were set to the default values of the software without modification. After pretreatment, b1 mass spectrometry features were identified from the data-dependent data in a surface water samples, and b2 mass spectrometry features were identified from the data-independent data.

[0168] Step 4: Perform data mining driven by the PFAS structure on the preprocessed mass spectrometry data, including four parts.

[0169] First, a suspected target database containing c0 PFAS was constructed based on the Organization for Economic Cooperation and Development database and the PubChem database. Using the suspected target screening method in non-target analysis, c mass spectrometry features (c < b1 and c ≤ c0) were screened out from b1 mass spectrometry features that might belong to PFAS.

[0170] Second, based on the Kendrick mass defect normalized by CF2 (≤ 0.1 or ≥ 0.8), d mass spectrometry features (d < b1) that might belong to PFAS were screened out from b1 mass spectrometry features using the mass defect filtering method in non-target analysis.

[0171] Then, based on characteristic fragment ions such as C2F5−, C2F5O−, C3F7−, C3F5−, SO3F−, C2F5−, and C5F12N−, and neutral loss fragments such as HF, e0 mass spectrometry features (e0 < b2) that might belong to PFAS were marked from b2 mass spectrometry features using the characteristic fragment ion labeling method in non-target analysis. According to the retention time, e mass spectrometry features (e ≤ e0 and e < b1) with both precursor ions and secondary fragment ions were screened out from b1 mass spectrometry features in the data-dependent data, and it was considered that e mass spectrometry features might belong to PFAS.

[0172] By comparing the mass spectrometry features screened by the above three non-target analysis methods, it was found that there were g0 repeated mass spectrometry features in the three datasets.

[0173] Finally, based on the identified (c + d + e - g0) mass spectrometry features that may belong to PFAS and the PFAS homolog units of mass 49.9968 (CF2), 99.9936 (CF2CF2), 64.0125 (CH2CF2), and 65.9917 (CF2O), etc., f mass spectrometry features that may belong to PFAS (f < b1) were screened out from b1 mass spectrometry features using the homolog search method in non-target analysis.

[0174] By comparing the mass spectrometry features screened out by the above four non-target analysis methods, it was found that there were g mass spectrometry features that were repeated in the four datasets. Therefore, finally, (c + d + e + f - g) mass spectrometry features that may belong to PFAS were identified from a surface water samples.

[0175] Step Five: Structural annotation was carried out on the excavated (c + d + e + f - g) mass spectrometry features, including three parts.

[0176] First, using spectral databases such as mzClould, massBank EU, and MoNA, mass spectra were retrieved for h1' of the (c + d + e + f - g) mass spectrometry features. Based on a spectral similarity score threshold of 70%, h1 mass spectrometry features were preliminarily identified as belonging to PFAS.

[0177] Secondly, for the mass spectrometry features that may belong to PFAS but were not identified above, the spectral prediction tool CFM-ID was used to predict their mass spectra and compare them with the acquired mass spectra in data-dependent data. Based on the principle of at least 2 secondary mass spectrometry fragment matches, h2 mass spectrometry features were further identified as belonging to PFAS.

[0178] Then, for the mass spectrometry features that may belong to PFAS but were not identified in either the spectral database or the spectral prediction tool above, based on the fragmentation path of the high-resolution mass spectrometry features of PFAS-like new pollutants, h3 mass spectrometry features were manually identified as belonging to PFAS.

[0179] Through the above methods, a total of kinds of PFAS were identified from a surface water samples.

[0180] By investigating the existing authentic standards in the market, it was determined that among the kinds of PFAS, kinds had authentic standards. Subsequently, authentic standards were purchased and kinds of PFAS were target-confirmed. Among them, kinds of PFAS were confirmed

[0181] Step Six: For the confirmed Perform targeted quantitative analysis on various PFASs to give accurate concentration levels.

[0182] For PFASs lacking authentic standards, except for those, based on the principle of structural similarity (Tanimoto coefficient > 0.7), with the aid of authentic standards of structurally similar PFASs, perform semi - quantitative analysis to give predicted concentration levels.

[0183] The cumulative number of PFASs with accurate concentration levels and predicted concentration levels is h. .

[0184] Through retrieval in the NORMAN, ECOTOX, and EnviroTox ecotoxicology databases and literature retrieval, it is found that the (semi) - quantified ha, hb, and hc PFASs respectively have experimental semi - lethal concentration or half - growth inhibition concentration for algae, daphnia, and fish.

[0185] For PFASs without toxicity experimental data for algae (h - ha species), daphnia (h - hb species), and fish (h - hc species), use the quantitative structure - activity relationship model to predict their ecotoxicity data. For PFASs with predicted toxicity concentrations higher than the water solubility, replace the toxicity prediction value with the solubility.

[0186] Step Seven: Evaluate the ecological risks of each PFAS to aquatic species such as algae, daphnia, and fish using the ratio of the PFAS concentration level to the toxicity endpoint threshold.

[0187] Based on the concentration addition model, obtain the cumulative risks of h PFASs to algae, daphnia, and fish respectively.

[0188] Step Eight: For different species, calculate the contribution of each PFAS to the cumulative risk of PFASs respectively. Select the top three PFASs with the highest risk contribution in each species as the newly emerging pollutants for priority control. It can be seen that the number of PFASs for priority control does not exceed 9.

[0189] Example 3:

[0190] Take the in - house test example, focusing on implementing the comprehensive identification and risk assessment of newly emerging pollutants in samples by integrating the established "top - down" and "bottom - up" strategies in series, as Figure 4 shown, and its specific steps include:

[0191] Step One: In order to identify newly emerging pollutants with neurotoxic effects in human cerebrospinal fluid and their neurotoxic risks, adopt a fusion method of computational toxicology followed by non - targeted analysis and then computational toxicology.

[0192] Step 2: After determining the fusion method, based on the "top-down" series-connected "bottom-up" fusion strategy, conduct a comprehensive identification of new pollutants with neurotoxic effects in human cerebrospinal fluid and their neurotoxic risk assessment.

[0193] Step 3: Similar to Step 3 in Example 1, use machine learning to construct a database of neurotoxic new pollutants; subsequently, similar to Steps 4, 5, and 6 in Example 1, identify a mass spectrometry features from cerebrospinal fluid samples.

[0194] Step 4: Similar to Steps 4 and 5 in Example 2, excavate b0 mass spectrometry features that may belong to new pollutants with neurotoxic effects from a mass spectrometry features. Further, b new pollutants with neurotoxic effects are confirmed by authentic standards. At the same time, there are c mass spectrometry features that may belong to new pollutants with neurotoxic effects.

[0195] Step 5: Similar to Step 7 in Example 1, obtain the quantitative concentration and semi-quantitative concentration of (b + c) new pollutants.

[0196] Use the quantitative structure-activity relationship model to predict the 10% growth inhibition concentration of (b + c) new pollutants on human neuroblastoma cell SH-SY5Y.

[0197] Step 6: Similar to Step 7 in Example 2, obtain the neurotoxic risk of the identified new pollutants based on the concentration level and effect concentration ratio of (b + c) new pollutants.

[0198] Combining the results of the above embodiments can comprehensively show that the new pollutant identification method provided in this embodiment has high reliability, and the non-target analysis and computational toxicology methods are complementary; it can be applied to various complex matrices, with little influence from matrix types and interference effects; the types of new pollutants screened and the types of risks covered are wide, and both known and unknown new pollutants and their risks can be efficiently identified; it has high-throughput characteristics and can complete the identification and risk assessment of a large number of new pollutants in a short time. Therefore, the new pollutant identification method provided in this embodiment has universal applicability and broad application prospects.

[0199] In this embodiment, a new pollutant identification device is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0200] This embodiment provides a new pollutant identification device, as Figure 5 shown, this device includes:

[0201] An acquisition module 501, configured to acquire a high-throughput fusion method, which is obtained by fusing at least one non-target analysis method and at least one computational toxicology method.

[0202] A first determination module 502, configured to determine the application order of each non-target analysis method and each computational toxicology method in the high-throughput fusion method.

[0203] A second determination module 503, configured to determine a target execution strategy of the high-throughput fusion method based on the application order, where the target execution strategy is one or more of a toxicity effect to chemical substance strategy and a chemical substance to toxicity effect strategy.

[0204] An evaluation module 504, configured to perform new pollutant risk assessment by using the target execution strategy to obtain a plurality of risk values.

[0205] A third determination module 505, configured to determine a variety of target new pollutants based on the plurality of risk values.

[0206] The further function descriptions of the above-mentioned various modules are the same as those in the corresponding embodiments above, and will not be repeated here.

[0207] The new pollutant identification device in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0208] The embodiment of the present invention further provides a computer device having the above-mentioned Figure 5 shown new pollutant identification device.

[0209] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a computer device provided by an optional embodiment of the present invention. As Figure 6As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting the components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories. Similarly, multiple computer devices can be connected, and each device provides part of the necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 6 In [the figure], a processor 10 is taken as an example.

[0210] The processor 10 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 can further include a hardware chip. The above-mentioned hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device can be a complex programmable logic device, a field programmable gate array, a generic array logic, or any combination thereof.

[0211] Among them, the memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiments.

[0212] The memory 20 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device, etc. In addition, the memory 20 can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 can optionally include a memory remotely set relative to the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0213] The memory 20 can include a volatile memory, such as a random access memory; the memory can also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 20 can also include a combination of the above types of memories.

[0214] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or communication networks. [[ID=I8]]

[0215] Embodiments of the present invention also provide a computer-readable storage medium. The method according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored as such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.

[0216] A part of the present invention can be applied as a computer program product, such as computer program instructions. When executed by a computer, through the operation of the computer, the method and / or technical solution according to the present invention can be called or provided. Those skilled in the art should be able to understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible by the computer.

[0217] Although the embodiments of the present invention are described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for identifying emerging pollutants, characterized in that, The method includes: Obtaining a high-throughput fusion method, which is obtained by fusing at least one non-target analysis method and at least one computational toxicology method, and the computational toxicology method includes a quantitative structure-activity relationship model, machine learning, molecular docking, and molecular dynamics simulation; Determining the application order of each non-target analysis method and each computational toxicology method in the high-throughput fusion method; Based on the application order, determining the target execution strategy of the high-throughput fusion method, and the target execution strategy is one or more of a toxicity effect to chemical substance strategy and a chemical substance to toxicity effect strategy; Using the target execution strategy to conduct a risk assessment of new pollutants to obtain multiple risk values; Based on the multiple risk values, determining multiple target new pollutants; Wherein, when the target execution strategy is a toxicity effect to chemical substance strategy, using the target execution strategy to conduct a risk assessment of new pollutants to obtain multiple risk values, including: Using the computational toxicology method to screen and determine the range of new pollutants that meet the selected toxicity endpoints from a preset database; Based on the range of new pollutants, collecting samples to obtain multiple initial new pollutant samples; Obtaining multiple sample matrix characteristics of the multiple initial new pollutant samples, where the sample matrix characteristic represents the sample type; Based on the range of new pollutants and the multiple sample matrix characteristics, using a preset sample pretreatment method to process the multiple initial new pollutant samples to obtain multiple target new pollutant samples; Performing mass spectrometry analysis on the multiple target new pollutant samples to obtain an initial mass spectrometry analysis data set; Preprocessing the initial mass spectrometry analysis data set to obtain a target mass spectrometry analysis data set; Based on the range of new pollutants, using the non-target analysis method to perform data mining on the target mass spectrometry analysis data set to obtain multiple second new pollutants, and the data mining method includes one or more of suspected target screening, isotope labeling, mass defect filtering, homolog retrieval, case control, and characteristic fragment ion labeling; Performing structure annotation on the multiple second new pollutants to obtain multiple chemical structures; According to the properties of different chemical substances and the analysis conditions available, performing semi-quantitative analysis on the multiple chemical structures to obtain the concentration level of new pollutants; Processing the target mass spectrometry analysis data set through the computational toxicology method to obtain a toxicity prediction result; Based on the concentration level of new pollutants and the toxicity prediction result, calculating the multiple risk values.

2. The method according to claim 1, wherein The target execution strategy is a chemical substance to toxicity effect strategy; Using the target execution strategy to conduct a risk assessment of new pollutants to obtain multiple risk values, including: Obtaining multiple initial new pollutant samples; Respectively processing and performing mass spectrometry analysis on the multiple initial new pollutant samples to obtain a target mass spectrometry analysis data set; Performing data mining analysis and toxicity prediction on the target mass spectrometry analysis data set to obtain the concentration level of new pollutants and a toxicity prediction result; Based on the concentration level of new pollutants and the toxicity prediction result, calculating the multiple risk values.

3. The method according to claim 2, characterized in that, Perform data mining analysis and toxicity prediction on the target mass spectrometry analysis dataset to obtain the concentration level of new pollutants and the toxicity prediction result, including: Perform data mining analysis on the target mass spectrometry analysis dataset using the non-target analysis method to obtain the concentration level of the new pollutants; Based on the target mass spectrometry analysis dataset, through the calculation of toxicology method, obtain the toxicity prediction result.

4. The method according to claim 1, wherein Based on the multiple risk values, determine multiple target new pollutants, including: Determine multiple new pollutant risk contribution values based on the multiple risk values; Based on the multiple risk values and the multiple new pollutant risk contribution values, determine the multiple target new pollutants.

5. A new pollutant identification device, characterized in that, For performing the new pollutant identification method according to any one of claims 1 to 4; the device includes: An acquisition module, configured to acquire a high-throughput fusion method, which is obtained by fusing at least one non-target analysis method and at least one calculation of toxicology method; A first determination module, configured to determine the application order of each non-target analysis method and each calculation of toxicology method in the high-throughput fusion method; A second determination module, configured to determine the target execution strategy of the high-throughput fusion method based on the application order, where the target execution strategy is one or more of a toxicity effect to chemical substance strategy and a chemical substance to toxicity effect strategy; An evaluation module, configured to perform new pollutant risk assessment using the target execution strategy to obtain multiple risk values; A third determination module, configured to determine multiple target new pollutants based on the multiple risk values.

6. A computer device, characterized in that, Including: A memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the new pollutant identification method according to any one of claims 1 to 5.

7. A computer-readable storage medium, characterized in that, Computer instructions are stored on the computer-readable storage medium, and the computer instructions are used to cause a computer to execute the new pollutant identification method according to any one of claims 1 to 5.