Ecological risk evaluation method and system based on new pollutants

By setting up monitoring points in petrochemical plants and conducting high-resolution mass spectrometry scanning and secondary mass spectrometry analysis, combined with a graph neural network model, the problem of unknown sources and quantities of new pollutants was solved, ecological risk assessment of new pollutants was achieved, and the accuracy of screening and risk identification capabilities were improved.

CN121747743APending Publication Date: 2026-03-27河南省濮阳生态环境监测中心
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies lack clarity regarding the emission sources and quantities of new pollutants in petrochemical cities, as well as their presence in environmental media. This makes it impossible to effectively assess their potential ecological risks and results in a lack of effective management measures.

Method used

By setting up monitoring points within a predetermined range of petrochemical plants, collecting sample data and performing high-resolution mass spectrometry full scans, matching normal mass spectrum fingerprints based on geographical location and time, marking unknown peaks, and inferring the molecular structure of unknown peaks through secondary mass spectrometry analysis and graph neural network models, quantitative structure-activity relationships of structural analogs are constructed to assess their potential ecological risks.

Benefits of technology

It improves the accuracy of screening for new pollutants, enabling precise identification of relevant known compounds, prediction of biological targets and signaling pathways of unknown new pollutants, identification of potential risks, and provision of a basis for environmental management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121747743A_ABST
    Figure CN121747743A_ABST
Patent Text Reader

Abstract

The invention provides an ecological risk evaluation method and system based on new pollutants, and the method comprises the steps: obtaining mass spectrum data of sample data, determining a normal mass spectrum fingerprint corresponding to a geographic position and time, marking an ion peak as an unknown peak when the ion peak in the mass spectrum data deviates from a preset range of the normal mass spectrum fingerprint, and determining that the ion peak is an unknown peak; the method comprises the following steps: acquiring a fragmented spectrum of an unknown peak based on secondary mass spectrometry, inferring a possible molecular structure of the unknown peak through a graph neural network model in combination with a preset chemical database, and constructing a quantitative structure-activity relationship of possible molecular structure analogues of the unknown peak based on a preset structural similarity network. The method comprises the following steps: acquiring predicted toxicity and signal channels influenced by the predicted toxicity, performing ecological risk assessment in a preset range of a petrochemical plant based on the predicted toxicity and the signal channels influenced by the predicted toxicity, and acquiring potential ecological risks of new pollutants. And a basic basis is provided for environment management and decision making.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of ecological assessment, in particular to an ecological risk assessment method and system based on new pollutants. BACKGROUND

[0002] New pollutants refer to pollutants caused by human activities, discharged into the environment, having characteristics such as biological toxicity, environmental persistence, biological accumulation, and having great risks to the ecological environment or human health, but not yet being included in management or existing management measures being insufficient to effectively prevent and control the risks thereof. At present, typical new pollutants widely concerned at home and abroad mainly include antibiotics, microplastics, environmental hormones (endocrine disruptors, EDCs), new persistent organic pollutants (POPs) and the like.

[0003] Puyang is a petrochemical city, and the entire city is born and thrives because of oil. A large number of petrochemical and upstream and downstream supporting industries are built in the city, and a large number of new pollutant emission enterprises are involved. However, the current new pollutant data such as enterprise emission sources and emission amounts, and the occurrence in environmental media such as water, gas and soil are not clear, and the potential risks are not known. How to investigate and monitor new pollutants, so as to master the risk impact of new pollutant pollution level on ecology, and provide a basic basis for environmental management and decision-making has become a problem that needs to be solved. SUMMARY

[0004] In order to overcome the above problems in the prior art, the application provides an ecological risk assessment method and system based on new pollutants, which adopts the following technical scheme:

[0005] In the first aspect, the application provides an ecological risk assessment method based on new pollutants, comprising:

[0006] Based on the monitoring points arranged in the preset range of the petrochemical plant, sample data of the monitoring points are collected, the sample data are subjected to high-resolution mass spectrometry full scanning, and mass spectrometry data of the sample data are obtained;

[0007] The mass spectrometry data are matched with the collected geographical position and time, and a normal mass spectrum fingerprint corresponding to the geographical position and time is determined, and when the ion peak in the mass spectrometry data deviates from the preset range of the normal mass spectrum fingerprint, it is marked as an unknown peak;

[0008] The unknown peak is subjected to secondary mass spectrometry analysis, a fragmentation spectrum of the unknown peak is obtained, and based on the fragmentation spectrum and a preset chemical database, a possible molecular structure of the unknown peak is inferred through a graph neural network model;

[0009] Based on the preset structure similarity network, the structure analog of the possible molecular structure of the unknown peak is obtained, the quantitative structure-activity relationship of the structure analog is constructed, and the predicted toxicity of the structure analog is obtained; based on the possible molecular structure of the unknown peak, the signal transduction affected by the predicted toxicity is obtained;

[0010] Based on the predicted toxicity and the signal transduction affected by the predicted toxicity, the ecological risk assessment of the preset range of the petrochemical plant is carried out, and the potential ecological risk of the new pollutant is obtained.

[0011] Further, the mass spectrum data of the sample data is obtained, wherein the mass spectrum data includes ion accurate mass and relative abundance information of the sample data.

[0012] Further, the mass spectrum data of the sample data is obtained, including:

[0013] The sample data is dissolved in a predetermined solvent, and the pretreated solvent is introduced into the mass spectrometer through a syringe pump to obtain charged, gaseous ions;

[0014] The ions pass through the mass analyzer of the mass spectrometer, and full scanning is performed in the preset scanning range. The mass spectrometer obtains each detected mass-to-charge ratio. Different mass-to-charge ratio ions reach the detector. When the ions hit the detector surface, the weak ion signal is converted into a measurable electrical signal.

[0015] The electrical signal is further amplified, and the amplified signal is converted into discrete, data-based data points.

[0016] The data-based data points are associated with the corresponding scanning time to obtain an original data file. The original data file is subjected to baseline correction, peak identification and deconvolution to obtain the accurate mass of each ion. By comparing the peak intensity or peak area of different ions, the relative abundance information of the ion in the sample is obtained. The accurate mass and relative abundance information of the ion are used as the mass spectrum data of the ion.

[0017] Further, the determination of the normal mass spectrum fingerprint under the corresponding geographical position and time attribute includes:

[0018] According to the monitoring points arranged in the preset range, different monitoring areas are divided, and based on the time attribute, the time is divided into different time windows to obtain the mass spectrum of the historical data of the space-time unit meeting the monitoring area and the time window;

[0019] The average intensity and standard deviation of each mass-to-charge ratio peak under the corresponding monitoring area and time window are calculated;

[0020] For each mass-to-charge ratio peak, a normal intensity range is determined by a percentage range, and the normal intensity range of all mass-to-charge ratio peaks is obtained, which is the normal mass spectrum fingerprint of the current space-time unit.

[0021] Further, when the ion peak in the mass spectrum data deviates from the preset range of the normal mass spectrum fingerprint, it is marked as an unknown peak, including:

[0022] All potential important mass-to-charge ratio feature peaks in the mass spectrum data are obtained by peak detection and alignment, a dimension is created for each feature peak, and all mass spectrum data is converted into a fixed-length numerical vector;

[0023] The longitude and latitude in the geographical position are converted into a numerical feature, the time is represented by one-hot encoding, and the fixed-length numerical vector, numerical feature, and one-hot encoding representation are spliced together to form a joint vector;

[0024] The joint vector is taken as the input of the preset spatiotemporal spectrum autoencoder, the association between the geographical position and the ion abundance in the input data is captured as a first data set based on multiple neural networks in the encoder of the spatiotemporal spectrum autoencoder, and the regularity between the time and the ion abundance is captured as a second data set, and the first data set and the second data set are compressed into a low-dimensional latent space representation.

[0025] The decoder reconstructs the original ion abundance data based on the latent space representation, compares the reconstructed ion abundance with the actually measured ion abundance, obtains the reconstruction error, and when the reconstruction error exceeds a preset threshold, marks the ion peak exceeding the threshold as an unknown peak.

[0026] Further, the unknown peak is subjected to secondary mass spectrum analysis to obtain a fragmentation spectrum of the unknown peak, including:

[0027] Based on the preset conditions of the mass spectrometer, the ions corresponding to the unknown peak are subjected to fragmentation processing, the ions corresponding to the unknown peak enter the collision chamber, collide with inert gas molecules, decompose the ions corresponding to the unknown peak into smaller fragment ions, the fragment ions pass through the mass analyzer to obtain the mass-to-charge ratio and the relative abundance, and generate a fragmentation spectrum of the ions corresponding to the unknown peak.

[0028] Further, based on the fragmentation spectrum and the preset chemical database, the molecular structure of the unknown peak is inferred by a graph neural network model, including:

[0029] All fragmentation spectra are sorted in order of mass-to-charge ratio from small to large to obtain a fragmentation spectrum sequence;

[0030] Based on the accurate mass of the parent ion of the unknown peak, known compounds with similar accurate mass are searched in the preset chemical database to generate a preliminary candidate molecular formula. Based on the isotope pattern of the parent ion of the unknown peak, the molecular formula with inconsistent isotope abundance in the preliminary candidate molecular formula is eliminated to obtain a secondary screened molecular formula;

[0031] The fragment map is compared with a known chemical database, and the fragment map of the candidate molecule structure library is obtained by excluding the fragment map of the molecular structure difference in the secondary screening molecular formula;

[0032] The candidate molecule structure library is input into a graph neural network model, and each molecule structure is converted into a graph, in which atoms are nodes and chemical bonds are edges. Through message passing of a multi-layer graph convolution network, the interaction between atoms and neighbors is learned, and the topological structure and chemical environment information of the candidate molecule are captured;

[0033] Based on the topological structure and chemical environment information, a predicted fragment map is generated for each candidate molecule structure. The similarity score of the predicted fragment map and the unknown peak fragment map is obtained, and all candidate molecule structures are sorted based on the similarity score. The candidate molecule structure with the highest score is selected as the possible molecular structure of the unknown peak.

[0034] In a second aspect, the present application also provides an ecological risk assessment system based on new pollutants, comprising:

[0035] A mass spectrometry data acquisition module is configured to arrange monitoring points within a preset range of an oil chemical plant, collect sample data of the monitoring points, perform high-resolution mass spectrometry full scanning on the sample data, and acquire mass spectrometry data of the sample data.

[0036] An unknown peak marking module is configured to match the mass spectrometry data with the collected geographical position and time, determine a normal mass spectrometry fingerprint corresponding to the geographical position and time, and mark an ion peak in the mass spectrometry data as an unknown peak when the ion peak deviates from a preset range of the normal mass spectrometry fingerprint.

[0037] An unknown peak possible molecular structure inference module is configured to perform secondary mass spectrometry analysis on the unknown peak, acquire a fragment map of the unknown peak, and infer a possible molecular structure of the unknown peak based on the fragment map and a preset chemical database through a graph neural network model.

[0038] An unknown peak possible molecular structure analysis module is configured to acquire structural analogues of the possible molecular structure of the unknown peak based on a preset structure similarity network, construct a quantitative structure-activity relationship of the structural analogues, acquire a predicted toxicity of the structural analogues, and acquire a signal transduction affected by the predicted toxicity based on the possible molecular structure of the unknown peak.

[0039] A potential ecological risk assessment module is configured to perform ecological risk assessment on the preset range of the oil chemical plant based on the predicted toxicity and the signal transduction affected by the predicted toxicity, and acquire a potential ecological risk of the new pollutant.

[0040] In a third aspect, the present application provides an electronic device, comprising:

[0041] One or more processors; a memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory, the one or more computer programs comprising instructions that, when executed by the device, cause the device to perform the method of the first aspect.

[0042] In a fourth aspect, the present application provides a computer readable storage medium, having stored therein a computer program which, when executed on a computer, causes the computer to perform the method of the first aspect.

[0043] In a fifth aspect, the present application provides a computer program for performing the method of the first aspect when executed by a computer.

[0044] In a possible design, the program in the fifth aspect can be stored, in whole or in part, on a storage medium packaged together with the processor, or stored, in whole or in part, on a storage medium not packaged together with the processor.

[0045] The present application has the following beneficial effects:

[0046] 1. The present application arranges monitoring points within the preset range of the petrochemical plant, collects sample data of the monitoring points, performs full scan on the sample data by high-resolution mass spectrometry, obtains mass spectrum data of the sample data, matches the mass spectrum data with the collected geographical position and time, determines the normal mass spectrum fingerprint corresponding to the geographical position and time, and marks as unknown peaks when the ion peaks in the mass spectrum data deviate from the preset range of the normal mass spectrum fingerprint. The present application performs secondary mass spectrometry analysis on the unknown peaks, obtains fragmentation spectrum of the unknown peaks, and infers possible molecular structure of the unknown peaks based on the fragmentation spectrum and the preset chemical database through a graph neural network model. In the process of screening unknown peaks, the present application filters false abnormal data caused by changes in the natural environment by considering the geographical position and time factors, and further improves the accuracy in the process of screening unknown substances.

[0047] 2. The present application obtains structural analogs of possible molecular structure of unknown peaks based on a preset structure similarity network, constructs quantitative structure-activity relationship of the structural analogs, and obtains predicted toxicity of the structural analogs. Based on the possible molecular structure of the unknown peaks, the present application obtains a signal transduction affected by the predicted toxicity, and performs ecological risk assessment within the preset range of the petrochemical plant based on the predicted toxicity and the signal transduction affected by the predicted toxicity, to obtain potential ecological risk of new pollutants. The present application can accurately find the most relevant known compounds of unknown compounds through the structure similarity network, and reduces the amount of data for screening. The present application can predict biological targets and signal pathways of unknown new pollutants, identify potential risks, and provide a basic basis for environmental management and decision-making. BRIEF DESCRIPTION OF DRAWINGS

[0048] Figure 1 A flow chart of the ecological risk assessment method based on new pollutants of the embodiments of the present application;

[0049] Figure 2 A flow chart of the mass spectrum data acquisition of the ecological risk assessment method based on new pollutants of the embodiments of the present application;

[0050] Figure 3 A flow chart of the normal mass spectrum fingerprint acquisition of the ecological risk assessment method based on new pollutants of the embodiments of the present application;

[0051] Figure 4 Unknown peak acquisition of the ecological risk assessment method based on new pollutants of the embodiments of the present application;

[0052] Figure 5 A flow chart of the system of the embodiments of the present application. DETAILED DESCRIPTION

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terminology used in the description herein is for describing particular embodiments only and is not intended to be limiting of the application; the description and the drawings are to be regarded as illustrative in nature and are not intended to limit the application; the terminology used in the description and the claims of the present application and the above description of the drawings includes the terms specifically mentioned above as well as their derivatives.

[0054] Reference herein to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearances of the phrase that in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily mutually exclusive of one another. As will be apparent to those of ordinary skill in the art, embodiments described herein can be combined with other embodiments.

[0055] In order to make the technical personnel in the art better understand the scheme of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings.

[0056] Reference should be made to Figure 1 The ecological risk assessment method based on new pollutants provided by the embodiments of the present application, the specific content includes:

[0057] In step 101, monitoring points are arranged in a preset range of the petrochemical plant, sample data of the monitoring points are collected, high-resolution mass spectrometry full scanning is performed on the sample data, and ion accurate mass and relative abundance information in the sample data are obtained.

[0058] It should be noted that the pollution generated by the petrochemical plant not only exists in the soil, but also exists in the air, so when collecting sample data of the petrochemical plant, monitoring points are arranged in a preset range of the petrochemical plant, including obtaining sample data such as surface water, underground water, soil and gas near the petrochemical plant.

[0059] It should be noted that when the monitoring points are arranged in the preset range of the petrochemical plant, the monitoring points can be set according to the monitoring purpose and the environmental conditions, the monitoring points are set in the area that may be polluted, a background point is set under various wind direction conditions, and three monitoring points are set according to a fan shape at a downwind plant boundary. The monitoring points can be set by using a grid method or an azimuth method, and the monitoring points of the pollution source area, the plant boundary, the sensitive area and the background are set according to the point function of the petrochemical plant and the surrounding environment.

[0060] It should be noted that the sample data is subjected to high-resolution mass spectrometry full scanning, and ion accurate mass and relative abundance information in the sample data are obtained. Please refer to Figure 2 , including:

[0061] In step 21, the sample data is dissolved in a preset solvent, the pretreated solvent is introduced into the mass spectrometer through the injection pump, and the charged and gaseous ions are obtained. Specifically, the sample data is dissolved in a preset solvent, and the solvent is pretreated. The pretreated solvent is introduced into the mass spectrometer through the injection pump. The pretreated solvent is converted into charged and gaseous ions by the mass spectrometer. In the embodiment of the application, the pretreatment of the solvent includes purification or dilution of the solvent.

[0062] In step 22, the ions are subjected to full scanning in a preset scanning range by the mass analyzer of the mass spectrometer, and the mass spectrometer obtains each detected mass-to-charge ratio. Different mass-to-charge ratio ions reach the detector. When the ions hit the surface of the detector, weak ion signals are converted into measurable electrical signals. Specifically, the ions are subjected to full scanning in a preset scanning range by the mass analyzer of the mass spectrometer, and all existing ions in the preset range are measured. In the process of full scanning, the mass spectrometer obtains each detected mass-to-charge ratio. Different mass-to-charge ratio ions reach the detector. When the ions hit the surface of the detector, a series of electrons are generated, and then weak ion signals are converted into measurable electrical signals.

[0063] Step 23, further amplifying the electric signal, and converting the amplified signal into discrete and dataized data points, comprising: further amplifying the electric signal, the amplified signal is proportional to the number of ions hitting the detector, and converting the amplified signal into discrete and dataized data points.

[0064] Step 24, associating the dataized data points with corresponding scanning time, and obtaining an original data file. In the embodiment of the present application, the data points form a complex mass spectrum.

[0065] Step 25, performing baseline correction, peak identification and deconvolution on the original data file, obtaining the accurate mass of each ion, obtaining the relative abundance information of the ion in the sample by comparing the peak intensity or peak area of different ions, and taking the accurate mass and relative abundance information of the ion as the mass spectrum data of the ion. In the embodiment of the present application, a window based on a preset step is slid on the original data file, the minimum value in the window is taken as a baseline point, all baseline points are connected, effective peaks are screened based on a preset threshold, and peaks overlapping with each other are separated and decomposed into independent components.

[0066] In the embodiment of the present application, the accurate mass of the ion can be the possible element composition, that is, the molecular formula, and different element combinations have different accurate masses.

[0067] Step 102, matching the mass spectrum data with the collected geographical position and time, determining the normal mass spectrum fingerprint corresponding to the geographical position and time, and marking as an unknown peak when the ion peak in the mass spectrum data deviates from the normal mass spectrum fingerprint by a preset range.

[0068] It should be noted that in the process of collecting sample data, the sample data is encoded, and the encoded sample data contains position information and time information. At this time, the position information and time information of the sample data are associated with the mass spectrum data, thereby achieving the purpose of matching the mass spectrum data with the geographical position and time. The ion abundance of different compounds is different at different time points, and there is a seasonal or diurnal variation rule. Therefore, in the process of confirming the unknown peak, the geographical position and time factors are considered, so as to avoid the problem of incorrect judgment of the unknown peak due to the influence of the geographical position and time.

[0069] It should be noted that the distance from the petrochemical plant is different, and the collected sample data is different. Therefore, in the process of data collection, the geographical position is matched, and the mass spectrum data corresponding to the geographical position can be analyzed.

[0070] In the embodiment of the present application, the normal mass spectrum fingerprint corresponding to the geographical position and time attribute is determined, please refer toFigure 3 The implementation process is as follows:

[0071] Step 31, according to the preset range, the monitoring points are divided into different monitoring areas, and based on the time attribute, the time is divided into different time windows, and the mass spectrum diagram of the historical data meeting the space-time unit of the monitoring area and the time window is obtained.

[0072] Step 32, the average intensity and standard deviation of each mass-to-charge ratio peak in the corresponding monitoring area and time window are calculated. For each mass-to-charge ratio peak, a normal intensity range is determined by the percentage range, and the normal intensity range of all mass-to-charge ratio peaks is obtained, that is, the normal mass spectrum diagram fingerprint of the current space-time unit.

[0073] In the embodiment of the application, when the ion peak in the mass spectrum data deviates from the normal mass spectrum diagram fingerprint preset range, it is marked as an unknown peak, please refer to Figure 4 The implementation process is as follows:

[0074] Step 41, all potential important mass-to-charge ratio feature peaks in the mass spectrum data are obtained through peak detection and alignment, wherein the mass-to-charge ratio feature peak represents an ion with analysis value, a dimension is created for each feature peak, and all mass spectrum data are converted into a fixed-length numerical vector. The latitude and longitude in the geographic location are converted into a numerical feature, and the time is represented by one-hot encoding. The fixed-length numerical vector, numerical feature and one-hot encoding representation are spliced together to form a joint vector.

[0075] Step 42, the joint vector is taken as the input of the preset space-time spectrum diagram autoencoder, based on the multiple neural networks in the encoder of the space-time spectrum diagram autoencoder, the correlation between the geographic location and the ion abundance in the input data is captured as a first data set, and the regularity between the time and the ion abundance is captured as a second data set, and the first data set and the second data set are compressed into a low-dimensional potential space representation.

[0076] Step 43, the decoder reconstructs the original ion abundance data based on the potential space representation, compares the reconstructed ion abundance with the actually measured ion abundance, and obtains the reconstruction error. When the reconstruction error exceeds the preset threshold, the ion peak exceeding the threshold is marked as an unknown peak.

[0077] It should be noted that the preset space-time spectrum diagram autoencoder in the application is trained based on normal data, and when the mass spectrum data is input into the preset model, the possible normal mass spectrum diagram of the corresponding ion can be predicted. By considering the time and the geographic location, the application avoids false positives caused by time and geographic location, and further improves the accuracy of the identification of unknown peaks.

[0078] In step 103, secondary mass spectrometry analysis is performed on the unknown peak to obtain a fragmentation spectrum of the unknown peak, and based on the fragmentation spectrum and a preset chemical database, a molecular structure of the unknown peak is inferred through a graph neural network model.

[0079] It should be noted that the mass-to-charge ratio information of the ion obtained in step 101 cannot determine the specific structure of the unknown peak, and the compound structure of the ion corresponding to the unknown peak can be further analyzed through secondary mass spectrometry.

[0080] In the embodiments of the present application, the secondary mass spectrometry analysis on the unknown peak to obtain the fragmentation spectrum of the unknown peak comprises:

[0081] Based on the preset conditions of the mass spectrometer, the ion corresponding to the unknown peak is fragmented, the ion corresponding to the unknown peak enters the collision chamber and collides with inert gas molecules, the ion corresponding to the unknown peak is decomposed into smaller fragment ions, the fragment ions pass through the mass analyzer to obtain the mass-to-charge ratio and the relative abundance, and the fragmentation spectrum of the ion corresponding to the unknown peak is generated.

[0082] It should be noted that the mass-to-charge ratio and the relative abundance information of the ion corresponding to the unknown peak obtained after the secondary mass spectrometry analysis constitute the fragmentation spectrum of the unknown peak.

[0083] In the embodiments of the present application, the inference of the molecular structure of the unknown peak based on the fragmentation spectrum and the preset chemical database through the graph neural network model comprises:

[0084] All fragmentation spectra are sorted in order of mass-to-charge ratio from small to large to obtain a fragmentation spectrum sequence. Each fragmentation spectrum carries key information of the parent ion structure.

[0085] Based on the accurate mass of the parent ion of the unknown peak, a known compound with a similar accurate mass is searched in the preset chemical database to generate a preliminary candidate molecular formula. Based on the isotope pattern of the parent ion of the unknown peak, the molecular formula with inconsistent isotope abundance in the preliminary candidate molecular formula is eliminated to obtain a secondary screening molecular formula.

[0086] The fragmentation spectrum is compared with the known theoretical fragmentation spectrum or the published experimental fragmentation spectrum in the preset chemical database, and the fragmentation spectrum with large differences in molecular structure in the secondary screening molecular formula is excluded to obtain a candidate molecular structure library.

[0087] For example, if the accurate mass of the fragmented ion has been determined, two molecular formulas of C6H12O6 and C9H10O3 can be matched through a preset chemical database. After isotopic pattern discovery, it is found that the isotopic abundance of C6H12O6 is similar to that of the parent ion of the unknown peak, and thus C6H12O6 is taken as a secondary screening molecular formula. The isomer molecular structure is determined through the secondary screening molecular formula, and the isomers of C6H12O6 can include fructose, glucose, and the like. The molecular structure of the fragmented ion is further determined by comparing the fragmented ion spectrum with the known or theoretical fragmented ion spectrum of fructose or glucose.

[0088] The candidate molecular structure library is taken as the input of the graph neural network model, and each molecular structure is converted into a graph, in which atoms are nodes and chemical bonds are edges. Through message passing of a multi-layer graph convolutional network, the interaction between atoms and surrounding neighbors is learned, and the topological structure and chemical environment information of the candidate molecule are captured. Based on the topological structure and chemical environment information, a predicted fragmented ion spectrum is generated for each candidate molecular structure. The predicted fragmented ion spectrum is compared with the fragmented ion spectrum of the unknown peak to obtain a similarity score, and all candidate molecular structures are ranked based on the similarity score. The candidate molecular structure with the highest score is taken as the possible molecular structure of the unknown peak.

[0089] It should be noted that, by means of the graph neural network model, the present embodiment generates a predicted fragmented ion spectrum of the unknown peak based on the understanding of the topological structure and chemical environment information of the molecule, and can predict new compounds that do not exist in a preset chemical database, thereby realizing the confirmation of the unknown peak.

[0090] In step 104, a structure analog of the possible molecular structure of the unknown peak is obtained based on a preset structure similarity network, a quantitative structure-activity relationship of the structure analog is constructed, a predicted toxicity of the structure analog is obtained, and a biological enrichment pathway of the predicted toxicity is obtained based on the possible molecular structure of the unknown peak.

[0091] In the present embodiment, the obtaining of the structure analog of the possible molecular structure of the unknown peak based on the preset structure similarity network, the construction of the quantitative structure-activity relationship of the structure analog, and the obtaining of the predicted toxicity of the structure analog include:

[0092] Obtain known toxic compound data, wherein each toxic compound data includes a molecular structure of a compound and toxicity data. Convert the molecular structure of the known toxic compound data and the possible molecular structure of the unknown peak into physicochemical / topological descriptors, and construct a molecular structure similarity network based on the physicochemical / topological descriptors, wherein the nodes in the molecular structure similarity network represent compounds, and the edges represent connections with structural similarity higher than a preset threshold. In the molecular structure similarity network, all known toxic compounds connected to the possible molecular structure of the unknown peak are regarded as structural analogues.

[0093] Train a quantitative structure-activity relationship model through the physicochemical / topological descriptors and toxicity data of the structural analogues, input the physicochemical / topological descriptors of the possible molecular structure of the unknown peak into the trained quantitative structure-activity relationship model, and obtain a predicted potential toxicity value of the possible molecular structure of the unknown peak.

[0094] In the embodiments of the present application, the obtaining of the biological enrichment pathway of the predicted toxicity based on the possible molecular structure of the unknown peak comprises:

[0095] Convert the possible molecular structure of the unknown peak into a set of numerical molecular descriptions, input the molecular descriptions into a preset target prediction model, generate a prediction score for each biological target based on the relationship between the molecular structure characteristics and the target binding probability, sort the prediction scores, and screen a candidate target list that meets a preset threshold. Perform pathway enrichment analysis on each target in the candidate target list to obtain an affected signal network.

[0096] In step 105, based on the predicted toxicity and the signal network affected by the predicted toxicity, perform ecological risk assessment on a preset range of the petrochemical plant to obtain the potential ecological risk of the new pollutant.

[0097] In the embodiments of the present application, the ecological risk assessment on the preset range of the petrochemical plant based on the predicted toxicity and the signal network affected by the predicted toxicity to obtain the potential ecological risk of the new pollutant comprises:

[0098] Obtain wastewater, waste gas and solid discharge data in the preset range of the petrochemical plant. Based on the production process and discharge amount of the petrochemical plant, evaluate the concentration of the new pollutant at the discharge port, evaluate the predicted environmental concentration in different regions and media based on a preset environmental fate model. Divide the predicted potential toxicity value by a safety factor to obtain a predicted no-effect concentration. Obtain the ratio between the predicted environmental concentration and the predicted no-effect concentration. When the ratio meets a first preset threshold, it is high risk, when the ratio meets a second preset threshold, it is medium risk; when the ratio meets a third preset threshold, it is low risk.

[0099] Embodiment two is a new pollutant-based ecological risk assessment system provided by the embodiment of the application, and specific modules are referred to Figure 5 ,

[0100] The mass spectrum data acquisition module 501 is configured to arrange monitoring points within a preset range of the oil chemical plant, collect sample data of the monitoring points, perform high-resolution mass spectrum full scanning on the sample data, and acquire mass spectrum data of the sample data.

[0101] The unknown peak marking module 502 is configured to match the mass spectrum data with the collected geographical position and time, determine a normal mass spectrum fingerprint corresponding to the geographical position and time, and mark an unknown peak when an ion peak in the mass spectrum data deviates from a preset range of the normal mass spectrum fingerprint.

[0102] The unknown peak possible molecular structure inference module 503 is configured to perform secondary mass spectrum analysis on the unknown peak, acquire a fragmentation spectrum of the unknown peak, and infer a possible molecular structure of the unknown peak based on the fragmentation spectrum and a preset chemical database through a graph neural network model.

[0103] The unknown peak possible molecular structure analysis module 504 is configured to acquire a structural analog of the possible molecular structure of the unknown peak based on a preset structure similarity network, construct a quantitative structure-activity relationship of the structural analog, and acquire a predicted toxicity of the structural analog; and acquire a signal transduction affected by the predicted toxicity based on the possible molecular structure of the unknown peak.

[0104] The potential ecological risk assessment module 505 is configured to perform ecological risk assessment on the oil chemical plant within the preset range based on the predicted toxicity and the signal transduction affected by the predicted toxicity, and acquire a potential ecological risk of the new pollutant.

[0105] Embodiment three provides an electronic device, including: one or more processors; a memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions that, when executed by the device, cause the device to perform the method of embodiment one.

[0106] Embodiment four provides a computer-readable storage medium, which stores a computer program, and when the computer program is run on a computer, the computer program causes the computer to perform the method of embodiment one.

[0107] Embodiment five provides a computer program, which, when executed by a computer, is configured to perform the method of embodiment one.

[0108] In a possible implementation, the program in the fifth aspect can be stored in whole or in part on a storage medium packaged together with the processor, or stored in whole or in part on a storage medium not packaged with the processor.

[0109] Obviously, the above-described embodiments are only some embodiments but not all of the present application, and the preferred embodiments of the present application are shown in the drawings, but do not limit the patent scope of the present application. The present application can be implemented in many different forms, and conversely, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing specific embodiments, or make equivalent replacements to some technical features. Any equivalent structure made by using the content of the specification and drawings, directly or indirectly applied to other related technical fields, is also within the patent protection scope of the present application.

Claims

1. An ecological risk assessment method based on new pollutants, characterized in that, include: Based on the monitoring points set up within the preset range of the petrochemical plant, sample data from the monitoring points are collected, and the sample data is fully scanned by high-resolution mass spectrometry to obtain the mass spectrometry data of the sample data. To match the geographical location and time of mass spectrometry data acquisition, a normal mass spectrum fingerprint corresponding to the geographical location and time is determined. When the ion peaks in the mass spectrometry data deviate from the preset range of the normal mass spectrum fingerprint, they are marked as unknown peaks. Secondary mass spectrometry analysis was performed on the unknown peak to obtain its fragmented spectrum. Based on the fragmented spectrum and a pre-set chemical database, the possible molecular structure of the unknown peak was inferred using a graph neural network model. Based on a pre-defined structural similarity network, we obtain structural analogs with possible molecular structures of unknown peaks, construct quantitative structure-activity relationships of structural analogs, and obtain predicted toxicity of structural analogs; based on possible molecular structures of unknown peaks, we obtain signal pathways that affect predicted toxicity. Based on predicted toxicity and the signal network affected by predicted toxicity, an ecological risk assessment is conducted within a pre-defined area of ​​the petrochemical plant to obtain the potential ecological risks of new pollutants.

2. The ecological risk assessment method based on new pollutants according to claim 1, characterized in that, The mass spectrometry data obtained from the sample data includes the precise ion mass and relative abundance information of the sample data.

3. The ecological risk assessment method based on new pollutants according to claim 1, characterized in that, The mass spectrometry data used to acquire the sample data includes: The sample data is dissolved in a pre-set solvent, and the pre-treated solvent is introduced into the mass spectrometer via a syringe pump to obtain charged, gaseous ions. Ions are passed through the mass analyzer of a mass spectrometer and fully scanned within a preset scanning range. The mass spectrometer acquires the mass-charge ratio of each detected ion. Ions with different mass-charge ratios reach the detector. When the ions collide with the detector surface, the weak ion signal is converted into a measurable electrical signal. The electrical signal is further amplified, and the amplified signal is converted into discrete, quantified data points. The digitized data points are correlated with the corresponding scan times to obtain the original data file. The original data file is then subjected to baseline correction, peak identification, and deconvolution to obtain the precise mass of each ion. By comparing the peak intensity or peak area of ​​different ions, the relative abundance information of the ions in the sample is obtained. The precise mass and relative abundance information of the ions are used as the mass spectrometry data of the ions.

4. The ecological risk assessment method based on new pollutants according to claim 1, characterized in that, The determination of the normal mass spectrum fingerprint under the corresponding geographical location and time attributes includes: The monitoring area is divided into different monitoring areas based on the preset range of monitoring points, and the time is divided into different time windows based on the time attribute. Mass spectra of historical data that meet the spatiotemporal units of the monitoring area and time window are obtained. Calculate the average intensity and standard deviation of each mass-charge ratio peak in the corresponding monitoring region and time window; For each mass-charge ratio peak, a normal intensity range is determined by the percentage range. The normal intensity range of all mass-charge ratio peaks is obtained, which is the normal mass spectrum fingerprint of the current spatiotemporal unit.

5. The ecological risk assessment method based on new pollutants according to claim 4, characterized in that, When an ion peak in the mass spectrometry data deviates from the preset range of the normal mass spectrum fingerprint, it is marked as an unknown peak, including: All potentially important mass-charge ratio characteristic peaks in the mass spectrometry data are obtained through peak detection and alignment. A dimension is created for each characteristic peak, and all mass spectrometry data are transformed into a fixed-length numerical vector. The latitude and longitude of the geographic location are converted into numerical features, time is represented by one-hot encoding, and a fixed-length numerical vector, numerical features and one-hot encoded representation are concatenated together to form a joint vector. The joint vector is used as the input of a pre-defined spatiotemporal spectrogram autoencoder. Based on multiple neural networks in the encoder of the spatiotemporal spectrogram autoencoder, the correlation between geographical location and ion abundance in the input data is captured as the first dataset, and the pattern between time and ion abundance is used as the second dataset. The first dataset and the second dataset are compressed into a small latent space representation. The decoder reconstructs the original ion abundance data based on the latent spatial representation, compares the reconstructed ion abundance with the actual measured ion abundance, obtains the reconstruction error, and marks the ion peaks exceeding the preset threshold as unknown peaks when the reconstruction error exceeds the preset threshold.

6. The ecological risk assessment method based on new pollutants according to claim 1, characterized in that, The step of performing secondary mass spectrometry analysis on the unknown peak to obtain the fragmented spectrum of the unknown peak includes: Based on the preset conditions of the mass spectrometer, the ions corresponding to the unknown peaks are fragmented. The ions corresponding to the unknown peaks enter the collision chamber and collide with inert gas molecules, decomposing the ions corresponding to the unknown peaks into smaller fragment ions. The fragment ions are then passed through a mass analyzer to obtain the mass-charge ratio and relative abundance, generating a fragmentation spectrum of the ions corresponding to the unknown peaks.

7. The ecological risk assessment method based on new pollutants according to claim 6, characterized in that, The method of inferring the molecular structure of unknown peaks based on fragmented spectra and a pre-set chemical database using a graph neural network model includes: All fragmented spectra are sorted in order of mass-charge ratio from smallest to largest to obtain the fragmented spectra sequence; Based on the precise mass of the precursor ion of the unknown peak, a search is conducted in a pre-defined chemical database for known compounds with similar precise masses to generate preliminary candidate molecular formulas. Based on the isotopic pattern of the precursor ion of the unknown peak, molecular formulas with mismatched isotopic abundances are eliminated from the preliminary candidate molecular formulas to obtain secondary screening molecular formulas. The fragmented spectra are compared with known theoretical fragment spectra of compounds in a pre-set chemical database or published experimental fragment spectra. Fragment spectra with large differences in molecular structure in the secondary screening molecular formulas are excluded to obtain a candidate molecular structure library. A library of candidate molecular structures is used as input to a graph neural network model. Each molecular structure is transformed into a graph, where atoms are nodes and chemical bonds are edges. Through message passing in a multi-layer graph convolutional network, the interactions between atoms and their neighbors are learned, capturing the topological and chemical environment information of candidate molecules. Based on topological and chemical environment information, a predicted fragment spectrum is generated for each candidate molecular structure. The predicted fragment spectrum is compared with the fragment spectrum of the unknown peak to obtain a similarity score. Based on the similarity score, all candidate molecular structures are ranked, and the candidate molecular structure with the highest score is taken as the possible molecular structure of the unknown peak.

8. An ecological risk assessment system based on new pollutants, used to implement the ecological risk assessment methods based on new pollutants according to claims 1-7, characterized in that, include: The mass spectrometry data acquisition module is used to set up monitoring points within a preset range of the petrochemical plant, collect sample data from the monitoring points, perform a high-resolution mass spectrometry full scan on the sample data, and acquire the mass spectrometry data of the sample data. The unknown peak marking module is used to match the geographical location and time of mass spectrometry data acquisition, determine the normal mass spectrum fingerprint of the corresponding geographical location and time, and mark the ion peak in the mass spectrometry data as an unknown peak when it deviates from the preset range of the normal mass spectrum fingerprint. The module for inferring the possible molecular structure of unknown peaks is used to perform secondary mass spectrometry analysis on unknown peaks, obtain fragmented spectra of unknown peaks, and infer the possible molecular structure of unknown peaks based on fragmented spectra and a preset chemical database through a graph neural network model. The unknown peak possible molecular structure analysis module is used to obtain structural analogs of the possible molecular structures of unknown peaks based on a preset structural similarity network, construct quantitative structure-activity relationships of structural analogs, obtain predicted toxicity of structural analogs; and obtain signal pathways affected by predicted toxicity based on the possible molecular structures of unknown peaks. The potential ecological risk assessment module is used to conduct ecological risk assessments within a preset area of ​​a petrochemical plant based on predicted toxicity and signal networks that are affected by predicted toxicity, thereby obtaining the potential ecological risks of new pollutants.

9. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-7.

10. A computer-readable storage medium storing computer instructions thereon, characterized in that, When executed by the processor, this instruction implements the steps of the method as described in any one of claims 1-7.