Machine learning-based method and device for identifying halogenated emerging pollutants

Through machine learning-based methods, the characteristic information in high-resolution mass spectrometry data is extracted, the presence of halogen is predicted, and compound identification and backward detection is carried out, which solves the limitations of screening and identification of new pollutants in the existing technology, and realizes efficient screening and environmental risk assessment of new halogenated pollutants.

CN119943193BActive Publication Date: 2025-06-17HANGZHOU INST FOR ADVANCED STUDY UCAS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510404764.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-06-17
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

There are limitations in existing new pollutant screening and identification methods, especially in the rapid screening and assessment of the environmental risks of halogenated organic pollutants, and there is a lack of efficient and flexible technical means.

Method used

Using a machine learning-based method, the durability, bioaccumulative and toxic properties of the compounds are evaluated by acquiring high-resolution mass spectrometry data, extracting characteristic information, using characteristic halogen-containing prediction models to predict the presence of halogen, combining precise mass and secondary mass spectrometry information, and evaluating the durability, bioaccumulative and toxic properties of the compounds using the compound characteristic prediction model.

Benefits of technology

It has achieved efficient screening and identification of new halogenated pollutants, evaluated their environmental risks, and carried out traceability detection, improving the efficiency and accuracy of non-target screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943193B_ABST
    Figure CN119943193B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for identifying halogenated emerging pollutants based on machine learning, including: Step S1: Obtain the original high-resolution mass spectrometry data of the sample to be tested, and extract the characteristic information in the original high-resolution mass spectrometry data; Step S2: Use the characteristic halogen-containing prediction model to predict the presence of halogen in the characteristics according to the extracted characteristic information, and perform characteristic filtering; Step S3: Based on the accurate mass and tandem mass spectrometry information, perform compound identification and retrospective detection on the filtered characteristics; Step S4: Use the compound property prediction model to predict the persistence, bioaccumulation and toxicity properties of the identified compounds. The present invention realizes the screening of halogen-containing characteristics and compound identification in mass spectrometry data, evaluates the environmental risk of the compounds represented by the characteristics, and further performs retrospective detection on the characteristics in historical mass spectrometry data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of non-target screening of new pollutants, and particularly to a method and device for identifying halogenated new pollutants based on machine learning. Background Art

[0002] New pollutants refer to toxic and harmful chemical substances that are discharged into the environment due to industrialization, technological development, or the widespread use of emerging products, have characteristics such as biotoxicity, environmental persistence, and bioaccumulation, pose great risks to the ecological environment or human health, but have not been included in management or the existing management measures are insufficient.

[0003] Halogenated organic pollutants are a very important group of new pollutants. For example, among the 34 types of persistent organic pollutants listed in the Stockholm Convention, 33 types are halogenated organic pollutants. These compounds can usually migrate over long distances through various environmental media, resulting in more extensive ecological and human health problems. Therefore, it is very important and meaningful to develop technologies that can detect and analyze these pollutants, especially methods that can quickly screen and evaluate their potential hazards from environmental detection samples.

[0004] With the progress of high-resolution mass spectrometry technology, the current basic conditions for accurate qualitative and quantitative analysis of molecules have been established. In compound identification, the standard comparison method is regarded as the gold standard, and the identification result is obtained by comparing the mass spectrometry data of the standard with the mass spectrometry data of the detection sample. Although this method has extremely high accuracy, there are limitations such as high cost, limited variety, and difficulty in obtaining some of the standards, resulting in many potential pollutants in the sample not being successfully identified. Obtaining all the standards for a small number of pollutants with known information is not only costly but also often has a serious time lag. Therefore, the conventional target monitoring mode relying on standards can only cover a part of the pollutants in the actual environment, which is far from enough for a comprehensive understanding and assessment of the exposure risks of the environment and the population. In contrast, non-target screening technology can effectively make up for this deficiency.

[0005] Compared with traditional target analysis, non-target analysis no longer relies on standards. It realizes the simultaneous detection and analysis of a large number of compounds through high-resolution and high-accuracy full-scan combined with various data processing technologies to quantify known pollutants and identify unknown pollutants. The non-target screening technology screens pollutants from complex mass spectrometry data by analyzing the characteristics of pollutants shown in the mass spectrometry.

[0006] However, existing methods focus more on the judgment of hard criteria, such as screening based on fixed isotope pattern ratios, neutral loss detection, mass defect calculation, and characteristic fragment matching. At the same time, many past studies have also ignored the importance of retrospective detection and lack the review and secondary analysis of historical data.

[0007] In summary, there are still certain limitations in the existing new pollutant screening and identification methods. There is an urgent need to develop more efficient and flexible screening technologies and processes with the help of technologies such as machine learning, especially for the screening and retrospective detection of chlorinated and brominated organic pollutants, in order to better meet the increasingly complex environmental detection needs. Summary of the Invention

[0008] The present invention provides a method and device for identifying halogenated new pollutants based on machine learning, which can screen for chlorine- and bromine-containing characteristics and identify compounds in mass spectrometry data, evaluate the environmental risks of the compounds represented by the characteristics, and further perform retrospective detection on the characteristics in historical mass spectrometry data.

[0009] The technical solution of the present invention is as follows:

[0010] A method for identifying halogenated new pollutants based on machine learning includes:

[0011] Step S1: Obtain the original high-resolution mass spectrometry data of the sample to be tested, and extract the characteristic information in the original high-resolution mass spectrometry data;

[0012] Step S2: Use the characteristic halogen-containing prediction model to predict the presence of halogens in the characteristics according to the extracted characteristic information, and filter out the characteristics without halogens;

[0013] Step S3: Perform compound identification and retrospective detection on the filtered characteristics based on the accurate mass and tandem mass spectrometry information;

[0014] Step S4: Use the compound property prediction model to predict the persistence, bioaccumulation, and toxicity properties of the identified compounds.

[0015] The main task of Step S1 is to convert the original high-resolution mass spectrometry data into a data format suitable for this method.

[0016] Preferably, Step S1 includes: obtaining the original high-resolution mass spectrometry data of the sample to be tested, using a peak picking algorithm to extract and analyze the peak characteristics in the original high-resolution mass spectrometry data, assigning corresponding tandem mass spectrometry information and isotope patterns to the peak characteristics, calculating the mass-to-charge ratio difference and intensity ratio of the isotope patterns, and generating a feature vector of the peak characteristics.

[0017] The main task of Step S2 is to predict the presence of halogens in the characteristics according to the isotope patterns of the characteristics.

[0018] Preferably, the construction method of the characteristic halogen-containing prediction model includes:

[0019] Step S2-1: Collect data on the composition and structure information of molecules, classify the data based on whether the molecules contain halogens and the number of halogens, and form a data set;

[0020] Step S2-2: According to the compositional information of the molecules in the dataset, combine the natural element isotope abundance distribution law to calculate the theoretical isotope pattern of the molecules, and calculate the differences and intensity ratios between the theoretical isotope mass-to-charge ratios as feature vectors to construct a model training dataset;

[0021] Step S2-3: Use the model training dataset to train a random forest model to obtain a characteristic halogen-containing prediction model.

[0022] The composition of the molecule is characterized by the molecular formula and exact mass; the structural information is represented in the form of a SMILES string.

[0023] The halogen mentioned is chlorine element and / or bromine element.

[0024] The characteristic halogen-containing prediction model can predict the number of halogen atoms in the molecule based on the isotope pattern and be applied to the prediction of the presence of halogens in the characteristics of actual mass spectrometry data.

[0025] The calculation method of the intensity ratio of the theoretical isotope pattern is as follows:

[0026] Taking the compound with the molecular formula C m H n as an example:

[0027] ;

[0028] ;

[0029] ;

[0030] where Int M represents the intensity of the lowest mass isotope pattern, Int M+1 represents the intensity of the isotope pattern with the mass difference rounded to 1 from the lowest mass, p C 12 represents the occurrence ratio of carbon-12 isotope, p C 13 represents the occurrence ratio of carbon-13 isotope, p H 1 represents the occurrence ratio of hydrogen-1 isotope, p H 2 represents the occurrence ratio of hydrogen-2 isotope, F M+1 / M represents the intensity ratio of adjacent mass isotope patterns.

[0031] Since the intensity calculation of the isotope pattern involves the merging of adjacent patterns, the intensity accumulation and the intensity-weighted calculation of the mass-to-charge ratio are adopted in the merging process:

[0032] ;

[0033] ;

[0034] where Int merged represents the intensity after the isotope pattern merging, Int i represents the intensity of the i th adjacent isotope pattern to be merged, m / z merged represents the mass-to-charge ratio obtained by intensity-weighted calculation after the isotope pattern merging, m / z i represents the mass-to-charge ratio of the i th adjacent isotope pattern to be merged, n is the number of adjacent isotope patterns to be merged.

[0035] Preferably, step S3 includes:

[0036] Step S3-1: Compare the accurate mass and the secondary mass spectrometry information of the feature with the compounds in the accurate mass database to identify the compound;

[0037] Step S3-2: Compare the accurate mass and the secondary mass spectrometry information of the feature with the historical mass spectrometry features to perform retrospective detection.

[0038] Step S3-1 includes:

[0039] Match the current feature based on the accurate mass:

[0040] When the mass deviation between the mass-to-charge ratio of the current feature and the mass of a certain compound C in the accurate mass database is less than or equal to 0.0005%, it is considered that the current feature matches compound C, and compound C is added to the candidate queue; otherwise, it is considered that the current feature fails to match compound C;

[0041] Match the current feature based on the secondary mass spectrometry information:

[0042] When the cosine similarity between the secondary mass spectrometry information of the current feature and the secondary mass spectrometry information of a certain compound C in the database is greater than or equal to 0.7, it is considered that the current feature matches compound C successfully, and the priority of compound C in the candidate queue is increased; otherwise, it is considered that the current feature fails to match compound C, and the priority of compound C in the candidate queue remains unchanged;

[0043] Select the compound with the highest priority in the candidate queue as the compound that matches the current feature.

[0044] Step S3-2 includes:

[0045] When the deviation of the mass-to-charge ratio of the current feature from the mass-to-charge ratio of feature D in the historical mass spectrometry data is less than or equal to 0.0005%; and at the same time, the deviation of the mass-to-charge ratio of at least two peaks in the tandem mass spectrometry of the current feature and feature D in the historical mass spectrometry data is less than or equal to 0.0005%, and the cosine similarity is greater than 0.7, it is considered that the current feature matches feature D in the historical mass spectrometry data successfully; otherwise, the backtracking match of the current feature fails.

[0046] Furthermore, the calculation method of the mass deviation between the mass-to-charge ratio of the current feature and the mass of a certain compound C in the exact mass database is as follows:

[0047]

[0048] where m t represents the exact mass of a certain compound C in the database, m / z s represents the mass-to-charge ratio of the current feature, and df represents the mass deviation between m t and m / z s .

[0049] Furthermore, the calculation method of the cosine similarity between the tandem mass spectrometry information of the current feature and the tandem mass spectrometry information of a certain compound C in the database is as follows:

[0050]

[0051] where A represents the tandem mass spectrometry information of the current feature represented in vector form, B represents the tandem mass spectrometry information in the exact mass database represented in vector form, A i and B i respectively represent the individual values in the two tandem mass spectrometry vectors, and cos similarity represents the cosine similarity between the tandem mass spectrometry information A and B.

[0052] Furthermore, before calculating the cosine similarity between the two tandem mass spectrometry information, it is necessary to perform a length alignment operation on the vectors of the two tandem mass spectrometry information according to the mass-to-charge ratio information, and fill the vacant positions with the value 0.

[0053] The construction method of the compound characteristic prediction model includes:

[0054] Step S4-1: Collect the compound data that are publicly and clearly stated to have or not have the characteristics of persistence, bioaccumulation, and toxicity in the literature to form a data set; the compound data include the structural information, persistence, bioaccumulation, and toxicity characteristics of the compound;

[0055] Step S4-2: Perform feature transformation and standardization processing on the structural information of the compounds in the data set, and randomly divide the data set into a training set, a test set, and a validation set to form the training data set of the model;

[0056] Step S4-3: Using the data of the persistence, bioaccumulation and toxicity characteristics of the compound as the target variables and the features transformed from the compound structure information as the input variables, train a multi-modal model to obtain a compound characteristic prediction model.

[0057] The structure information of the compound in the compound data is the SMILES string of the compound.

[0058] The feature transformation of the structure information of the compounds in the dataset includes: transforming the SMILES string into three features: molecular graph, Morgan molecular fingerprint, and molecular descriptor.

[0059] The compound structure information is transformed into three types of features that are homologous but have great differences in the generation method and information representation method: molecular graph, Morgan molecular fingerprint, and molecular descriptor. Among them, the molecular graph is used to extract the constituent atoms of the molecule and their chemical bond information, and the local features of the molecule are captured through convolution operations. The Morgan molecular fingerprint is used to describe the global information of the molecule. The molecular descriptor is used to characterize the local structure and physical and chemical properties of the molecule.

[0060] Furthermore, based on the SMILES string, construct the two-dimensional topological structure of the molecule, and represent the types and related characteristics of atoms and chemical bonds in the molecular graph in numerical form to obtain the molecular graph features; based on the SMILES string, use the RDKit calculation tool to generate a binary Morgan molecular fingerprint, and use 0 or 1 to indicate whether a certain molecular structure fragment exists in the molecule. If it does not exist, it is 0, and if it exists, it is 1, to obtain the Morgan molecular fingerprint; based on the SMILES string, use the Mordred or RDKit calculation tool to generate the molecular descriptor features.

[0061] Furthermore, input the three types of transformed features into a multi-modal model with a GINEConv network as the core, which can effectively fuse the features of the molecular graph, Morgan molecular fingerprint, and molecular descriptor, and perform model training to obtain a compound characteristic prediction model.

[0062] The multi-modal model described above includes:

[0063] A molecular graph feature extraction module, which takes the molecular graph as the input and extracts the molecular graph features of the compound;

[0064] A Morgan molecular fingerprint feature extraction module, which takes the Morgan molecular fingerprint as the input and extracts the Morgan molecular fingerprint features of the compound;

[0065] A molecular descriptor feature extraction module, which takes the molecular descriptor as the input and extracts the molecular descriptor features of the compound;

[0066] The feature fusion and prediction module concatenates the molecular graph features, Morgan molecular fingerprint features, and molecular descriptor features to predict compound properties.

[0067] Further, the molecular graph feature extraction module includes a GINE convolutional layer, a batch normalization layer, a GINE convolutional layer, a batch normalization layer, and a global pooling layer connected in sequence;

[0068] The Morgan molecular fingerprint feature extraction module includes a feature shaping layer and multiple fully connected layers connected in sequence;

[0069] The molecular descriptor feature extraction module includes a feature shaping layer and a fully connected layer connected in sequence;

[0070] The feature fusion and prediction module includes a fully connected layer, a batch normalization layer, a fully connected layer, and a softmax activation function connected in sequence.

[0071] The evaluation metrics for model performance are accuracy, macro-average precision, macro-average recall, and macro-average F1 score:

[0072] ;

[0073] ;

[0074] ;

[0075] ;

[0076] where Accuracy represents accuracy, Macro Precision represents macro-average precision, MacroRecall represents macro-average recall, Macro F1 Score represents macro-average F1 score, TP represents the number of positive class samples correctly predicted as positive by the model, TN represents the number of negative class samples correctly predicted as negative by the model, FP represents the number of negative class samples incorrectly predicted as positive by the model, FN represents the number of positive class samples incorrectly predicted as negative by the model, and in macro-average precision and macro-average recall, N represents the total number of classes predicted by the model, TP i represents the number of correct predictions when the model predicts the i th class, FP i represents the number of other classes incorrectly predicted as the i th class when the model predicts, FN i represents the number of the i th class incorrectly predicted as other classes when the model predicts the th class.

[0077] Based on the same inventive concept, the present invention also provides a machine learning-based halogenated emerging pollutant identification device, comprising:

[0078] A data acquisition and preprocessing module that acquires high-resolution mass spectrometry data of a sample to be tested and extracts characteristic information from the high-resolution mass spectrometry data using a peak picking algorithm;

[0079] A characteristic halogen-containing prediction module that uses a characteristic halogen-containing prediction model to predict the presence of halogen in the characteristics based on the extracted characteristic information and performs characteristic filtering;

[0080] A characteristic matching module that performs compound identification and retrospective detection on the filtered characteristics based on accurate mass and tandem mass spectrometry information;

[0081] A compound property prediction module that uses a compound property prediction model to predict the persistence, bioaccumulation, and toxicity properties of the identified compounds.

[0082] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0083] By introducing multiple machine learning models, the present invention can predict the presence of halogen in the characteristics based on the isotope pattern, reducing the interference of human factors while effectively distinguishing the interference of low-abundance elements on the isotope pattern, thereby improving the efficiency and accuracy of non-target screening. By developing a new multi-modal model to predict the persistence, bioaccumulation, and toxicity properties of the structure of the identified compounds, on the basis that traditional screening methods only focus on the compound identification results, the potential environmental hazards of the compounds are further revealed.

[0084] Secondly, the present invention integrates an accurate mass database and a tandem mass spectrometry information database, and through a program, it can automatically perform efficient and accurate compound identification by matching the mass-to-charge ratio and calculating the cosine similarity between tandem mass spectrometry information.

[0085] Finally, the present invention realizes the retrospective detection function, enabling the method to not only screen the current mass spectrometry data but also perform secondary analysis on historical mass spectrometry data, and with the increase in the amount of data, the retrospective detection effect will be further improved. Description of the Drawings

[0086] Figure 1 is a flowchart of the machine learning-based high environmental risk halogenated emerging pollutant identification method proposed by the present invention;

[0087] Figure 2 is an architecture diagram of the multi-modal model for realizing compound property prediction proposed by the present invention;

[0088] Figure 3This is a processing example diagram of the method for identifying highly environmentally risky halogenated emerging pollutants based on machine learning proposed by the present invention. Detailed implementation manners

[0089] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be noted that the following embodiments are intended to facilitate the understanding of the present invention and do not impose any limitation on it.

[0090] As Figure 1 shown, the process of the method for identifying highly environmentally risky halogenated emerging pollutants based on machine learning of the present invention mainly includes five steps: S1. Obtain the high-resolution mass spectrometry data of the sample, convert the file format of the original mass spectrometry data, and use the peak picking algorithm to extract the characteristic information in the mass spectrometry data; S2. Use a machine learning model to predict the presence of chlorine and bromine elements in the characteristics according to the characteristic information, and perform characteristic filtering; S3. Identify the compound based on the accurate mass and the secondary mass spectrometry information of the filtered characteristics; S4. Use a machine learning model to predict the persistence, bioaccumulation and toxicity characteristics of the compound represented by the characteristics; S5. Perform retrospective detection on the processed mass spectrometry characteristics.

[0091] The implementation of the steps of this method process is mainly based on the following two units and two models: a mass spectrometry data preprocessing unit, a feature matching unit, a feature chlorine and bromine prediction model, and a compound property prediction model.

[0092] The main task of the mass spectrometry data preprocessing unit is to convert the original mass spectrometry data into a data format suitable for this method. The specific process includes using the mzConvert program to convert the original mass spectrometry file into the mzML format, using the OpenMS tool to extract the peak characteristics in the mass spectrometry data, and further parsing the data using pymzml to assign the corresponding secondary mass spectrometry information and isotope pattern to the target characteristics. Finally, calculate the mass-to-charge ratio difference and intensity ratio of the isotope pattern to generate a feature vector. The features include mass-to-charge ratio, retention time, secondary mass spectrometry information, isotope pattern, mass-to-charge ratio difference and intensity ratio of the isotope pattern, etc.

[0093] The task of the feature matching unit is to measure the similarity between the current mass spectrometry feature and the data in the database or the features in other mass spectrometries through matching calculations, so as to achieve compound identification or retrospective detection. This unit includes three parts: accurate mass matching, secondary mass spectrometry peak matching, and secondary mass spectrometry cosine similarity calculation.

[0094] The task of the chlorine- and bromine-containing feature prediction model is to predict the presence of chlorine and bromine elements in the feature based on the isotope pattern of the feature. The specific process is to collect the composition and structure data of molecules through the Pubchem and CAS Common Chemistry platforms, calculate the theoretical isotope pattern in combination with the isotope abundance distribution law, and use the mass-to-charge ratio difference and intensity division quotient of the theoretical pattern as feature vectors to train a machine learning model that can predict the presence of chlorine and bromine elements in the feature. Here, the model architecture is a random forest.

[0095] The task of the compound property prediction model is to predict whether the compound that the feature may represent has persistence, bioaccumulation, and toxicity properties based on the identified compound structure information of the feature. This process collects the structure information of compounds that are publicly known to have or not have persistence, bioaccumulation, and toxicity properties in the literature, and uses the molecular graph, Morgan molecular fingerprint, and molecular descriptors transformed from the structure information as input features to train a machine learning model that can predict compound properties. Here, the model architecture is a self-developed multimodal model.

[0096] The machine learning model for predicting the chlorine and bromine situation in step S2 has a random forest architecture, and the specific implementation method is as follows:

[0097] First, collect the composition and structure data of molecules through the Pubchem and CAS Common Chemistry platforms. The molecular composition is characterized in the form of molecular formula, exact mass, etc., and the structure information is represented in the form of a SMILES (Simplified molecular input line entry system) string. Classify the data based on whether the molecular composition contains chlorine or bromine and the number of specific chlorine and bromine atoms in the molecular formula to form a data set.

[0098] Then, according to the composition information of the molecules in the data set, calculate the theoretical isotope pattern of the molecules in combination with the natural element isotope abundance distribution law, and calculate the difference between the theoretical isotope mass-to-charge ratios and the quotient of the intensities as feature vectors to construct a model training data set.

[0099] Finally, use the processed training data set to train a random forest model. This model can predict the number of chlorine and bromine atoms in the molecule based on the isotope pattern and is applied to the prediction of the presence of chlorine and bromine elements in the actual mass spectrometry data features.

[0100] Partial examples of the theoretical isotope intensity and the calculation method of the feature are for compounds with the molecular formula C m H n are as follows:

[0101] ;

[0102] ;

[0103] ;

[0104] wherein, Int M represents the intensity of the lowest mass isotope pattern, Int M+1 represents the intensity of the isotope pattern with the mass difference from the lowest mass rounded to 1 isotope, p C 12 represents the occurrence ratio of carbon-12 isotope, p C 13 represents the occurrence ratio of carbon-13 isotope, p H 1 represents the occurrence ratio of hydrogen-1 isotope, p H 2 represents the occurrence ratio of hydrogen-2 isotope, F M+1 / M represents the intensity ratio of adjacent mass isotope patterns.

[0105] Since the intensity calculation of the isotope pattern involves the merging of locally adjacent isotope patterns, the method of intensity accumulation and intensity-weighted calculation of the mass-to-charge ratio is adopted in the merging process:

[0106] ;

[0107] ;

[0108] wherein, Int merged represents the intensity after the merging of isotope patterns, Int i represents the i th intensity of the adjacent isotope pattern to be merged, m / z merged represents the mass-to-charge ratio obtained by intensity-weighted calculation after the merging of isotope patterns, m / z i represents the i th mass-to-charge ratio of the adjacent isotope pattern to be merged, n is the number of adjacent isotope patterns to be merged.

[0109] As a preferred embodiment of the present invention, the specific method for identifying a compound based on the exact mass in step S3 is:

[0110] When the mass deviation between the mass-to-charge ratio of the current feature and the mass of a certain compound in the exact mass database is less than or equal to 0.0005%, it is considered that the feature matches the compound and is added to the candidate queue;

[0111] When the mass deviation between the mass-to-charge ratio of the current feature and the mass of a certain compound in the exact mass database is greater than 0.0005%, it is considered that the feature fails to match the compound.

[0112] The calculation method of the mass deviation is as follows:

[0113]

[0114] Among them, m t represents the exact mass of a certain compound C in the database, m / z s represents the mass-to-charge ratio of the current feature, and df represents the mass deviation between m t and m / z s .

[0115] As a preferred solution of the present invention, the specific method for implementing compound identification based on the secondary mass spectrometry information in step S3 is as follows:

[0116] When the cosine similarity between the secondary mass spectrometry information of the current feature and the secondary mass spectrometry information of a certain compound in the database is greater than or equal to 0.7, it is considered that the match is successful, and the priority of the candidate is increased;

[0117] When the cosine similarity between the secondary mass spectrometry information of the current feature and the secondary mass spectrometry of a certain compound in the database is less than 0.7, it is considered that the match fails, and the priority of the candidate remains unchanged.

[0118] The calculation method of the cosine similarity between secondary mass spectrometry information is as follows:

[0119]

[0120] Among them, A represents the secondary mass spectrometry information represented in vector form of the current feature, B represents the secondary mass spectrometry information represented in vector form in the database, A i and B i respectively represent individual values in the two secondary mass spectrometry vectors, and cos similarity represents the cosine similarity between the secondary mass spectrometry information A and B. Before calculation, the two secondary mass spectrometry vectors need to be length-aligned according to the mass-to-charge ratio information, and the vacant positions are filled with the value 0.

[0121] As a preferred solution of the present invention, the machine learning model for predicting the persistence, bioaccumulation, and toxicity characteristics of compounds in step S4 is a multi-modal model, and the specific implementation method is as follows:

[0122] First, collect the data of compounds that are explicitly disclosed in the literature as having or not having persistence, bioaccumulation, and toxicity characteristics, and form a data set in the form of SMILES strings and corresponding characteristic data;

[0123] Then, convert the SMILES strings into three types of features with the same origin but significantly different chemical information representation forms, namely molecular graphs, Morgan fingerprints, and molecular descriptors, and perform standardization processing on the features. Randomly divide the constructed dataset into a training set, a test set, and a validation set to prepare data for the subsequent property prediction model;

[0124] Finally, using the data of persistence, bioaccumulation, and toxicity properties as the target variables, and the features (molecular graphs, Morgan fingerprints, molecular descriptors) transformed from the compound structure information as the input variables, train a multi-modal model. Obtain the best-performing model through the optimization methods of cross-validation and model integration, and use it to predict the properties of the identified compounds.

[0125] The architecture of the multi-modal model is as Figure 2 shown. This multi-modal model aims to fuse three feature representation methods, namely molecular graphs, Morgan fingerprints, and molecular descriptors, to comprehensively extract the structural information and physicochemical properties of molecules. The model adopts a parallel feature extraction strategy to process data of different modalities separately and perform feature fusion in the final stage to obtain more comprehensive molecular structure information, thereby improving the prediction ability and stability of the model. The overall architecture mainly consists of four main modules: a molecular graph feature extraction module, a Morgan fingerprint feature extraction module, a molecular descriptor feature extraction module, and a feature fusion and prediction module.

[0126] Molecular graph feature extraction module: This module is used to learn the topological structure of molecules and their local chemical environment. The module uses GINE convolutional layers to extract features and combines batch normalization layers to standardize the output of the GINE convolutional layers. Since the molecular structure is relatively complex, two operations of GINE convolution and batch normalization are adopted to enhance the model's perception ability of complex molecular structures. Subsequently, all atomic features are aggregated through a global pooling layer to generate molecular-level features with a fixed dimension.

[0127] Morgan fingerprint feature extraction module: This module is used to extract the global structure information of molecules. The module first adjusts the Morgan fingerprint to a shape of 2048 dimensions through a feature shaping layer to adapt to subsequent neural network layers. Subsequently, a multi-layer fully connected network is used to perform linear and non-linear transformations on the Morgan fingerprint, thereby converting it from a sparse representation to a more compact embedded representation to enhance the feature expression ability.

[0128] Molecular descriptor feature extraction module: This module is used to characterize the physical and chemical properties of molecules to supplement the structural description of molecules by molecular graphs and Morgan molecular fingerprints. First, the module adjusts the molecular descriptors to 217 dimensions through a feature shaping layer to adapt to the subsequent neural network layers. Subsequently, a fully connected layer is used to reduce the dimensionality of the molecular descriptors to extract more representative and effective molecular features, thereby enhancing the expression ability of the model.

[0129] Feature fusion and prediction module: This module concatenates the extracted molecular graph features, Morgan molecular fingerprint features, and molecular descriptor features to integrate information from different modalities and predict compound properties. First, a fully connected layer is used to perform a non-linear mapping on the concatenated features to fully learn the molecular features and the interaction relationships between their modalities. Subsequently, a batch normalization layer is used to normalize the features to prevent the features of a certain modality from dominating during training. Finally, a fully connected layer is used to transform the fused features, and a softmax activation function is used to predict the final result.

[0130] The specific process of input feature transformation is as follows: The molecular graph features construct the two-dimensional topological structure of the molecule through the SMILES string and represent the types and related properties of atoms and chemical bonds in the molecular graph in numerical form. Among them, there are 8 types of atomic features and 4 types of chemical bond features. The Morgan molecular fingerprint is based on the SMILES string and uses the RDKit calculation tool to generate a 2048-dimensional binary Morgan molecular fingerprint, using 0 or 1 to indicate whether a certain molecular structure fragment exists in the molecule. If it does not exist, it is 0; if it exists, it is 1. The molecular descriptor features are also based on the SMILES string and use the Mordred or RDKit calculation tool to generate 217-dimensional molecular descriptors to quantitatively represent the physical and chemical properties of the compound in numerical form.

[0131] The specific implementation principles of each network layer in the model are as follows:

[0132] Based on the traditional graph convolutional network, the GINE convolutional layer further enhances the model's ability to learn edge information in the graph structure through the weighting of edge features. The specific processing formula is as follows:

[0133]

[0134] Among them, h v and h u respectively represent the feature vectors of nodes v and u ; N ( v ) represents the set of adjacent nodes of node v ; e uvRepresents a node u and v The edge feature vector between; M represents a function for weighted calculation of node feature vectors and edge feature vectors; MLP represents a multi-layer perceptron for non-linearly transforming the feature vectors of each node; h’ v Represents the updated node v Feature.

[0135] The fully connected layer is used for linear and non-linear transformation of the input features, and the formula is as follows:

[0136] y = W • x + b

[0137] Among them, x Represents the input feature vector, usually the output of the previous layer; W Represents the weight matrix of the fully connected layer; b Represents the bias term; y Represents the output of the fully connected layer.

[0138] The batch normalization layer is used to standardize the data of each batch during training to help accelerate training and stabilize the model. The formula is as follows:

[0139]

[0140] Among them, x i Represents the input feature of the i th sample; μ B Represents the batch mean; σ B Represents the batch standard deviation; Represents the normalized feature.

[0141] The global pooling layer is used to aggregate the features of all nodes in the graph into a graph-level feature vector of a fixed size. The formula is as follows:

[0142]

[0143] Among them, h i Is the feature vector of node i ; N Is the total number of nodes in the graph; h global Is the aggregated global feature vector.

[0144] The evaluation metrics for model performance are Accuracy, Macro Precision, Macro Recall, and Macro F1 Score:

[0145] ;

[0146] ;

[0147] ;

[0148] ;

[0149] Among them, Accuracy represents accuracy, Macro Precision represents macro-average precision, Macro Recall represents macro-average recall, Macro F1 Score represents macro-average F1-score, TP represents the number of positive class samples correctly predicted as positive by the model, TN represents the number of negative class samples correctly predicted as negative by the model, FP represents the number of negative class samples wrongly predicted as positive by the model, FN represents the number of positive class samples wrongly predicted as negative by the model. In macro-average precision and macro-average recall, N represents the total number of classes predicted by the model, and TP i represents the number of correct predictions when the model predicts the i th class, and FP i represents the number of other classes wrongly predicted as the i th class during the model's prediction, and FN i represents the number of the i th class wrongly predicted as other classes when the model predicts the class.

[0150] Finally, the multi-modal model was tested on the test set and achieved an accuracy of 97%, a macro-average precision of 90%, a macro-average recall of 89%, and a macro-average F1-score of 89%, with the current optimal prediction performance for persistence, bioaccumulation, and toxicity characteristics.

[0151] As a preferred embodiment of the present invention, the specific method for implementing feature retrospective detection in step S5 is:

[0152] When the mass-to-charge ratio of a certain feature deviates from the mass-to-charge ratio of the feature in the historical mass spectrometry data by less than or equal to 0.0005%, and at the same time, the mass-to-charge ratios of at least two peaks in the secondary mass spectrometry deviate by less than or equal to 0.0005%, and the cosine similarity is greater than 0.7, it is considered that the feature matches the feature in the historical data successfully and is added to the retrospective detection queue. Otherwise, it is regarded as a feature matching failure.

[0153] The process of the present invention for processing the mass spectrometry data of specific samples is as follows: Figure 3 as shown below:

[0154] First, the mass spectrometry data preprocessing unit converts the original mass spectrometry file into the required feature and feature vector formats. The processing result is in the form: a certain feature ID is 2, the mass-to-charge ratio is 450.9261, the retention time is 1010.84, and the secondary mass spectrometry information is [50.1039, 59.0138, 68.9957, ……, 450.9257, 451.3118], etc.

[0155] Next, the trained random forest model is used to predict the feature vector to determine whether it contains chlorine or bromine elements, and corresponding feature filtering is performed. The processing result is in the form: the feature with ID 2 is predicted to contain chlorine elements, and 2 chlorine atoms are contained in 1 molecule. The feature can be retained for further analysis.

[0156] Then, accurate mass matching and secondary mass spectrometry cosine similarity calculation are performed on the features containing chlorine and bromine elements that are retained after filtering, so as to achieve accurate identification of compounds. The processing result is in the form: the feature with ID 2 undergoes accurate mass matching, and the candidate molecular formulas include C 12 HCl2F 11 N2, C 18 H 10 ClIO4, C 25 H4Cl2NO2S, C 12 H4Cl2F6N4O2S, etc.; after the cosine similarity calculation of the secondary mass spectrometry information, the highest score is 0.8914, and the identification result is fipronil sulfone, and the molecular formula is C 12 H4Cl2F6N4O2S, which is consistent with the accurate mass matching result.

[0157] Next, the structure information of the identified compound is input into the trained multi-modal model for predicting the persistence, bioaccumulation, and toxicity characteristics to evaluate whether the compound has potential environmental risks. The processing result is in the form: the model predicts that fipronil sulfone that the feature with ID 2 may represent does not simultaneously possess the persistence, bioaccumulation, and toxicity characteristics and may not have high environmental risk.

[0158] Finally, the features that have undergone feature screening, identification, and characteristic prediction are compared with the features in the historical mass spectrometry data, and the feature matching unit is used to find whether there are similar features, so as to perform retrospective detection of pollutants. The processing result is in the form: similar features of this feature are found in the historical mass spectrometry data file, indicating that halogenated new pollutants with high environmental risks also exist in the source environmental samples of this historical mass spectrometry data.

[0159] The above-described embodiments have elaborated in detail the technical solutions and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, supplements, equivalent replacements, etc. made within the scope of the principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for identifying new halogenated pollutants based on machine learning, characterized in that: include: Step S1: obtaining original high-resolution mass spectrum data of the sample to be tested, and extracting characteristic information from the original high-resolution mass spectrum data; Step S2: using a feature halogen-containing prediction model to predict the presence of halogen in the feature based on the extracted feature information, and filtering out features that do not contain halogen; Step S3: Compound identification and retrospective detection of the filtered features based on accurate mass and secondary mass spectrum information, including: Step S3-1: Compare the accurate mass and secondary mass spectrum information of the feature with the compounds in the accurate mass database to identify the compound: Match the current feature based on exact mass: When the mass-to-charge ratio of the current feature and the mass deviation of a compound C in the accurate mass database are less than or equal to 0.0005%, the current feature is considered to match compound C and compound C is added to the candidate queue; otherwise, the current feature is considered to fail to match compound C; Match the current feature based on the secondary mass spectrum information: When the cosine similarity between the secondary mass spectrum information of the current feature and the secondary mass spectrum information of a compound C in the database is greater than or equal to 0.7, it is considered that the current feature successfully matches compound C, and the priority of compound C in the candidate queue is increased; otherwise, it is considered that the current feature fails to match compound C, and the priority of compound C in the candidate queue remains unchanged; Select the compound with the highest priority in the candidate queue as the compound matching the current feature; Step S3-2: Compare the accurate mass and secondary mass spectrum information of the feature with the historical mass spectrum feature to achieve retrospective detection: The mass-to-charge ratio deviation of the current feature and the mass-to-charge ratio deviation of feature C in the historical mass spectrum data is less than or equal to 0.0005%; at the same time, when the mass-to-charge ratio deviation of at least two peaks in the secondary mass spectra of the current feature and feature C in the historical mass spectrum data is less than or equal to 0.0005%, and the cosine similarity is greater than 0.7, it is considered that the current feature matches the feature C in the historical mass spectrum data successfully; otherwise, the backtracking matching of the current feature fails; Step S4: Use the compound property prediction model to predict the persistence, bioaccumulation and toxicity properties of the identified compounds.

2. The method for identifying new halogenated pollutants based on machine learning according to claim 1, characterized in that: Step S1 includes: obtaining the original high-resolution mass spectrometry data of the sample to be tested, using a peak-lifting algorithm to extract and analyze the peak features in the original high-resolution mass spectrometry data, assigning corresponding secondary mass spectrometry information and isotope patterns to the peak features, calculating the mass-to-charge ratio difference and intensity ratio of the isotope pattern, and generating a characteristic vector of the peak features.

3. The method for identifying new halogenated pollutants based on machine learning according to claim 1, characterized in that: The method for constructing the characteristic halogen-containing prediction model includes: Step S2-1: Collecting the composition and structure information data of the molecule, classifying the data based on whether the composition of the molecule contains halogens and the number of halogens contained, to form a data set; Step S2-2: According to the composition information of the molecules in the data set, the theoretical isotope pattern of the molecules is calculated in combination with the isotope abundance distribution law of natural elements, and the difference between the theoretical isotope mass-to-charge ratios and the intensity ratio are calculated as feature vectors to construct a model training data set; Step S2-3: Use the model training data set to train the random forest model to obtain a characteristic halogen-containing prediction model.

4. The method for identifying new halogenated pollutants based on machine learning according to claim 1, characterized in that: The methods for constructing compound property prediction models include: Step S4-1: Collect data on compounds that are publicly known in the literature to have or not have persistence, bioaccumulation and toxicity characteristics to form a data set; the compound data includes structural information and persistence, bioaccumulation and toxicity characteristics of the compound; Step S4-2: performing feature conversion and standardization processing on the structural information of the compounds in the data set, and randomly dividing the data set into a training set, a test set, and a validation set to form a training data set for the model; Step S4-3: Using the persistence, bioaccumulation and toxicity characteristics data of the compound as the target variables and the characteristics of the compound structure information conversion as the input variables, the multimodal model is trained to obtain a compound characteristic prediction model.

5. The method for identifying new halogenated pollutants based on machine learning according to claim 4, characterized in that: The structural information of the compound in the compound data is the SMILES string of the compound; The feature conversion of the structural information of the compounds in the data set includes: converting the SMILES string into three features: molecular graph, Morgan molecular fingerprint, and molecular descriptor.

6. The method for identifying new halogenated pollutants based on machine learning according to claim 4, characterized in that: The multimodal model includes: The molecular graph feature extraction module takes the molecular graph as input and extracts the molecular graph features of the compound; Morgan molecular fingerprint feature extraction module, which takes Morgan molecular fingerprint as input and extracts the Morgan molecular fingerprint features of the compound; The molecular descriptor feature extraction module takes the molecular descriptor as input and extracts the molecular descriptor features of the compound; The feature fusion and prediction module combines molecular graph features, Morgan molecular fingerprint features and molecular descriptor features to predict compound properties.

7. A device for identifying new halogenated pollutants based on machine learning, characterized in that: include: The data acquisition and preprocessing module obtains the high-resolution mass spectrometry data of the sample to be tested and uses the peak-lifting algorithm to extract the characteristic information in the high-resolution mass spectrometry data; The feature halogen content prediction module uses the feature halogen content prediction model to predict the presence of halogens in the feature based on the extracted feature information and perform feature filtering; The feature matching module performs compound identification and retrospective detection on filtered features based on accurate mass and secondary mass spectrometry information, including: The accurate mass and secondary mass spectrum information of the features are compared with the compounds in the accurate mass database to identify the compounds: Match the current feature based on exact mass: When the mass-to-charge ratio of the current feature and the mass deviation of a compound C in the accurate mass database are less than or equal to 0.0005%, the current feature is considered to match compound C and compound C is added to the candidate queue; otherwise, the current feature is considered to fail to match compound C; Match the current feature based on the secondary mass spectrum information: When the cosine similarity between the secondary mass spectrum information of the current feature and the secondary mass spectrum information of a compound C in the database is greater than or equal to 0.7, it is considered that the current feature successfully matches compound C, and the priority of compound C in the candidate queue is increased; otherwise, it is considered that the current feature fails to match compound C, and the priority of compound C in the candidate queue remains unchanged; Select the compound with the highest priority in the candidate queue as the compound matching the current feature; Based on the accurate mass and secondary mass spectrum information of the features, retrospective detection is achieved by comparing with the historical mass spectrum features: The mass-to-charge ratio deviation of the current feature and the mass-to-charge ratio deviation of feature C in the historical mass spectrum data is less than or equal to 0.0005%; at the same time, when the mass-to-charge ratio deviation of at least two peaks in the secondary mass spectra of the current feature and feature C in the historical mass spectrum data is less than or equal to 0.0005%, and the cosine similarity is greater than 0.7, it is considered that the current feature matches the feature C in the historical mass spectrum data successfully; otherwise, the backtracking matching of the current feature fails; The compound property prediction module uses the compound property prediction model to predict the persistence, bioaccumulation and toxicity properties of identified compounds.

Citation Information

Patent Citations

  • Processing method and device of stacked integrated model for molecular attribute prediction

    CN116343950A

  • Soil organic pollutant identification method and system based on combination of AI and high-throughput screening

    CN119580881A