Halogenated new pollutant identification method and device based on machine learning

Through machine learning-based methods, characteristic information in high-resolution mass spectrometry data is extracted, the presence of halogen is predicted, and compound identification and backtrack detection is carried out in combination with accurate mass and secondary mass spectrometry information, which solves the shortcomings in the environmental risk assessment of halogenated organic pollutants in the existing technology, and achieves more efficient and accurate screening and identification of new pollutants.

CN119943193AActive Publication Date: 2025-05-06HANGZHOU INST FOR ADVANCED STUDY UCAS

Patent Information

Application Number
CN202510404764.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-05-06
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

There are limitations in existing new pollutant screening and identification methods, and it is difficult to effectively identify and evaluate the environmental risks of halogenated organic pollutants, especially in the lack of effective methods in the review of historical data and secondary analysis.

Method used

Using a machine learning-based method, the durability, bioaccumulative and toxic properties of the compounds are evaluated by acquiring high-resolution mass spectrometry data, extracting characteristic information, using characteristic halogen-containing prediction models to predict the presence of halogen, combining precise mass and secondary mass spectrometry information, and using compound characteristic prediction models to evaluate the durability, bioaccumulative and toxic properties of the compounds.

Benefits of technology

It improves the efficiency and accuracy of non-target screening, can more comprehensively identify and evaluate the environmental risks of new halogenated pollutants, and achieves efficient backtracking detection of historical mass spectrometry data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943193A_ABST
    Figure CN119943193A_ABST
Patent Text Reader

Abstract

The invention discloses a machine learning-based halogenated new pollutant identification method and device, and the method comprises the steps: S1, obtaining the original high-resolution mass spectrum data of a to-be-detected sample, and extracting the feature information in the original high-resolution mass spectrum data; s2, predicting the existence condition of halogen in the features according to the extracted feature information by using a feature halogen-containing prediction model, and filtering the features; s3, performing compound identification and traceability detection on the filtered characteristics based on accurate mass and secondary mass spectrum information; and S4, predicting the durability, the biological accumulation property and the toxicity property of the identified compound by using the compound property prediction model. While halogen-containing characteristic screening and compound identification in the mass spectrum data are realized, the environmental risk of the compound represented by the characteristic is evaluated, and the characteristics in the historical mass spectrum data are further subjected to backtracking detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of non-target screening of new pollutants, and in particular to a method and device for identifying new halogenated pollutants based on machine learning. Background Art

[0002] New pollutants refer to toxic and hazardous chemical substances that are discharged into the environment due to industrialization, technological development or the widespread use of emerging products, have characteristics such as biological toxicity, environmental persistence, and bioaccumulation, and pose great risks to the ecological environment or human health, but have not yet been included in management or the existing management measures are insufficient.

[0003] Halogenated organic pollutants are a very important group of new pollutants. For example, among the 34 types of persistent organic pollutants listed in the Stockholm Convention, 33 are halogenated organic pollutants. These compounds are usually able to migrate over long distances through various environmental media, leading to more extensive ecological and human health problems. Therefore, it is very important and meaningful to develop technologies that can detect and analyze these pollutants, especially methods that can quickly screen and evaluate their potential hazards from environmental test samples.

[0004] With the advancement of high-resolution mass spectrometry technology, the basic conditions for accurate qualitative and quantitative analysis of molecules are now available. In compound identification, the standard comparison method is regarded as the gold standard. The identification result is obtained by comparing the mass spectrometry data of the standard with the mass spectrometry data of the test sample. Although this method is extremely accurate, it has limitations such as the high price, limited variety and difficulty in obtaining some standards, resulting in many potential pollutants in the samples not being successfully identified. The full acquisition of standards for a few pollutants with known information is not only costly, but also often has serious time lags. Therefore, the conventional target monitoring mode relying on standards can only cover part of the pollutants in the actual environment, which is far from enough for a comprehensive understanding and assessment of the exposure risks of the environment and the population. In contrast, non-target screening technology can effectively make up for this deficiency.

[0005] Compared with traditional targeted analysis, non-targeted analysis no longer relies on standards. It uses high-resolution, high-accuracy full scans combined with multiple data processing technologies to simultaneously detect and analyze a large number of compounds, quantify known pollutants, and identify unknown pollutants. Non-targeted screening technology screens out pollutants from complex mass spectrometry data by analyzing the characteristics of pollutants in mass spectrometry.

[0006] However, existing methods focus more on the judgment of hard standards, such as screening based on fixed isotope pattern ratios, neutral loss detection, mass defect calculation, and characteristic fragment matching. At the same time, many previous studies have also ignored the importance of retrospective detection and lacked review and secondary analysis of historical data.

[0007] In summary, the existing methods for screening and identifying new pollutants still have certain limitations, and there is an urgent need to develop more efficient and flexible screening technologies and processes with the help of technologies such as machine learning, especially for the screening and retrospective detection of chlorinated and brominated organic pollutants, in order to better meet the increasingly complex environmental testing needs. Summary of the invention

[0008] The present invention provides a method and device for identifying new halogenated pollutants based on machine learning. While realizing the screening of chlorine- and bromine-containing features and compound identification in mass spectrometry data, it evaluates the environmental risk of the compounds represented by the features and further performs retrospective detection of the features in historical mass spectrometry data.

[0009] The technical solution of the present invention is as follows: A method for identifying new halogenated pollutants based on machine learning, comprising: Step S1: obtaining original high-resolution mass spectrum data of the sample to be tested, and extracting characteristic information from the original high-resolution mass spectrum data; Step S2: using a feature halogen-containing prediction model to predict the presence of halogen in the feature based on the extracted feature information, and filtering out features that do not contain halogen; Step S3: performing compound identification and retrospective detection on the filtered features based on accurate mass and secondary mass spectrum information; Step S4: Use the compound property prediction model to predict the persistence, bioaccumulation and toxicity properties of the identified compounds.

[0010] The main task of step S1 is to convert the original high-resolution mass spectrometry data into a data format suitable for this method.

[0011] Preferably, step S1 includes: obtaining original high-resolution mass spectrometry data of the sample to be tested, extracting and analyzing peak features in the original high-resolution mass spectrometry data using a peak lifting algorithm, assigning corresponding secondary mass spectrometry information and isotope patterns to the peak features, calculating the mass-to-charge ratio difference and intensity ratio of the isotope pattern, and generating a characteristic vector of the peak features.

[0012] The main task of step S2 is to predict the presence of halogens in the features based on their isotope patterns.

[0013] Preferably, the method for constructing the characteristic halogen-containing prediction model comprises: Step S2-1: Collecting the composition and structure information data of the molecule, classifying the data based on whether the composition of the molecule is halogen and the number of halogens contained, to form a data set; Step S2-2: According to the composition information of the molecules in the data set, the theoretical isotope pattern of the molecules is calculated in combination with the isotope abundance distribution law of natural elements, and the difference between the theoretical isotope mass-to-charge ratios and the intensity ratio are calculated as feature vectors to construct a model training data set; Step S2-3: Use the model training data set to train the random forest model to obtain a characteristic halogen-containing prediction model.

[0014] The composition of the molecule is represented by molecular formula and accurate mass; the structural information is represented by SMILES string.

[0015] The halogen is chlorine and / or bromine.

[0016] The characteristic halogen-containing prediction model can predict the number of halogen atoms in a molecule based on the isotope pattern and can be applied to the prediction of the presence of halogens in the actual mass spectrometry data features.

[0017] The intensity ratio of the theoretical isotope pattern is calculated as follows: In molecular formula form, C m H n For example, the compound: ; ; ;

[0018] Among them, Int M Int represents the intensity of the lowest mass isotope pattern. M+1 Indicates the intensity of the isotopic pattern with the mass difference from the lowest mass rounded to 1. p C 12 represents the appearance ratio of carbon-12 isotope, p C 13 represents the appearance ratio of carbon-13 isotope, p H 1 represents the occurrence ratio of hydrogen-1 isotope, p H 2 Indicates the occurrence ratio of hydrogen-2 isotope, F M+1 / M Represents the intensity ratio of adjacent mass isotope patterns.

[0019] Since the intensity calculation of isotope patterns involves the merging of adjacent patterns, the merging process adopts the intensity accumulation and mass-to-charge ratio based on intensity weighted calculation method: ; ;

[0020] Among them, Int merged Int represents the intensity of the combined isotope pattern. i Indicates i The intensity of the adjacent isotope pattern to be merged, m / z merged represents the mass-to-charge ratio calculated by intensity weighting after isotope pattern merging, m / z i Indicates i The mass-to-charge ratios of adjacent isotope patterns to be merged, n is the number of adjacent isotope patterns to be merged.

[0021] Preferably, step S3 includes: Step S3-1: comparing the accurate mass and secondary mass spectrum information of the feature with the compounds in the accurate mass database to identify the compound; Step S3-2: Based on the accurate mass and secondary mass spectrum information of the feature, the historical mass spectrum feature is compared to achieve retrospective detection.

[0022] Step S3-1 includes: Match the current feature based on exact mass: When the mass-to-charge ratio of the current feature and the mass deviation of a compound C in the accurate mass database are less than or equal to 0.0005%, the current feature is considered to match compound C and compound C is added to the candidate queue; otherwise, the current feature is considered to fail to match compound C; Match the current feature based on the secondary mass spectrum information: When the cosine similarity between the secondary mass spectrum information of the current feature and the secondary mass spectrum information of a compound C in the database is greater than or equal to 0.7, it is considered that the current feature successfully matches compound C, and the priority of compound C in the candidate queue is increased; otherwise, it is considered that the current feature fails to match compound C, and the priority of compound C in the candidate queue remains unchanged; The compound with the highest priority in the candidate queue is selected as the compound matching the current feature.

[0023] Step S3-2 includes: The mass-to-charge ratio deviation of the current feature and the mass-to-charge ratio of feature D in the historical mass spectrum data is less than or equal to 0.0005%; at the same time, when the mass-to-charge ratio deviation of at least two peaks in the secondary mass spectra of the current feature and feature D in the historical mass spectrum data is less than or equal to 0.0005%, and the cosine similarity is greater than 0.7, it is considered that the current feature matches the feature D in the historical mass spectrum data successfully; otherwise, the backtracking matching of the current feature fails.

[0024] Furthermore, the mass deviation between the mass-to-charge ratio of the current feature and a compound C in the accurate mass database is calculated as follows:

[0025] Among them, m t Indicates the exact mass of a compound C in the database, m / z s Indicates the mass-to-charge ratio of the current feature, df indicates m t and m / z s quality deviation.

[0026] Furthermore, the calculation method of the cosine similarity between the secondary mass spectrum information of the current feature and the secondary mass spectrum information of a compound C in the database is as follows:

[0027] Among them, A represents the secondary mass spectrum information of the current feature in the form of a vector, B represents the secondary mass spectrum information in the accurate mass database in the form of a vector, and A i and B i Represents a single value in two secondary mass spectrum vectors, cos similarity Represents the cosine similarity between secondary mass spectrometry information A and B.

[0028] Furthermore, before calculating the cosine similarity between two secondary mass spectrometry information, the vectors of the two secondary mass spectrometry information need to be length aligned according to the mass-to-charge ratio information, and the vacant positions are filled with the value 0.

[0029] The methods for constructing compound property prediction models include: Step S4-1: Collect data on compounds that are publicly known in the literature to have or not have persistence, bioaccumulation and toxicity characteristics to form a data set; the compound data includes structural information, persistence, bioaccumulation and toxicity characteristics of the compound; Step S4-2: performing feature conversion and standardization processing on the structural information of the compounds in the data set, and randomly dividing the data set into a training set, a test set, and a validation set to form a training data set for the model; Step S4-3: Using the persistence, bioaccumulation and toxicity characteristics data of the compound as the target variables and the characteristics of the compound structure information conversion as the input variables, the multimodal model is trained to obtain a compound characteristic prediction model.

[0030] The structural information of the compound in the compound data is the SMILES string of the compound.

[0031] The feature conversion of the structural information of the compounds in the data set includes: converting the SMILES string into three features: molecular graph, Morgan molecular fingerprint, and molecular descriptor.

[0032] The compound structure information is converted into three types of features that are homologous but have very different generation methods and information representation methods: molecular graphs, Morgan molecular fingerprints, and molecular descriptors. The molecular graph is used to extract the constituent atoms and their chemical bond information of the molecule, and the local features of the molecule are captured through convolution operations. The Morgan molecular fingerprint is used to describe the global information of the molecule. The molecular descriptor is used to characterize the local structure and physical and chemical properties of the molecule.

[0033] Furthermore, a two-dimensional topological structure of the molecule is constructed based on the SMILES string, and the types and related characteristics of atoms and chemical bonds in the molecular graph are represented in numerical form to obtain the molecular graph features; based on the SMILES string, the RDKit calculation tool is used to generate a binary Morgan molecular fingerprint, using 0 or 1 to indicate whether a certain molecular structure fragment exists in the molecule, if not, the value is 0, if present, the value is 1, and the Morgan molecular fingerprint is obtained; based on the SMILES string, the Mordred or RDKit calculation tool is used to generate molecular descriptor features.

[0034] Furthermore, the three types of transformed features are input into a multimodal model with the GINEConv network as the core, which can effectively integrate the molecular graph, Morgan molecular fingerprint and molecular descriptor features, for model training to obtain a compound property prediction model.

[0035] The multimodal model includes: The molecular graph feature extraction module takes the molecular graph as input and extracts the molecular graph features of the compound; Morgan molecular fingerprint feature extraction module, which takes Morgan molecular fingerprint as input and extracts the Morgan molecular fingerprint features of the compound; The molecular descriptor feature extraction module takes the molecular descriptor as input and extracts the molecular descriptor features of the compound; The feature fusion and prediction module combines molecular graph features, Morgan molecular fingerprint features and molecular descriptor features to predict compound properties.

[0036] Furthermore, the molecular graph feature extraction module includes a GINE convolution layer, a batch normalization layer, a GINE convolution layer, a batch normalization layer, and a global pooling layer connected in sequence; The Morgan molecular fingerprint feature extraction module includes a feature shaping layer and a multi-layer fully connected layer connected in sequence; The molecular descriptor feature extraction module includes a feature shaping layer and a fully connected layer connected in sequence; The feature fusion and prediction module includes a fully connected layer, a batch normalization layer, a fully connected layer and a softmax activation function which are connected in sequence.

[0037] The evaluation indicators of model performance are accuracy, macro-average precision, macro-average recall, and macro-average F1 score: ; ; ; ;

[0038] Among them, Accuracy represents accuracy, Macro Precision represents macro average precision, MacroRecall represents macro average recall, Macro F1 Score represents macro average F1 score, TP represents the number of positive samples correctly predicted as positive by the model, TN represents the number of negative samples correctly predicted as negative by the model, FP represents the number of negative samples incorrectly predicted as positive by the model, FN represents the number of positive samples incorrectly predicted as negative by the model, N in macro average precision and macro average recall represents the total number of categories predicted by the model, TP represents the number of positive samples correctly predicted as positive by the model, TN represents the number of negative samples correctly predicted as negative by the model, FP represents the number of negative samples incorrectly predicted as positive by the model, FN represents the number of positive samples incorrectly predicted as negative by the model, N represents the total number of categories predicted by the model, and TP represents the number of positive samples correctly predicted as positive by the model. i Indicates that the model is predicting i The number of correct predictions for the class, FP i Indicates that the model mistakenly predicts other classes as i The number of classes, FN i It means that the model mistakenly predicts the i The number of other classes that the class was predicted to be.

[0039] Based on the same inventive concept, the present invention also provides a device for identifying new halogenated pollutants based on machine learning, comprising: The data acquisition and preprocessing module obtains the high-resolution mass spectrometry data of the sample to be tested and uses the peak-lifting algorithm to extract the characteristic information in the high-resolution mass spectrometry data; The feature halogen content prediction module uses the feature halogen content prediction model to predict the presence of halogens in the feature based on the extracted feature information and perform feature filtering; Feature matching module, which performs compound identification and retrospective detection on filtered features based on accurate mass and secondary mass spectrometry information; The compound property prediction module uses the compound property prediction model to predict the persistence, bioaccumulation and toxicity properties of identified compounds.

[0040] Compared with the prior art, the present invention has the following beneficial effects: The present invention introduces multiple machine learning models to predict the presence of halogens in features based on isotope patterns. While reducing interference from human factors, it can also effectively distinguish the interference of low-abundance elements on isotope patterns, thereby improving the efficiency and accuracy of non-target screening. By developing a new multimodal model to predict the persistence, bioaccumulation and toxicity characteristics of the structures of identified compounds, the potential harm of compounds to the environment is further revealed on the basis of traditional screening methods that only focus on compound identification results.

[0041] Secondly, the present invention integrates an accurate mass database and a secondary mass spectrum information database, and can automatically perform efficient and accurate compound identification through matching of mass-to-charge ratios and calculation of cosine similarity between secondary mass spectrum information through a program.

[0042] Finally, the present invention realizes the retrospective detection function, so that the method can not only screen the current mass spectrometry data, but also perform secondary analysis on the historical mass spectrometry data, and as the amount of data increases, the retrospective detection effect will be further improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 A flowchart of a method for identifying new halogenated pollutants with high environmental risks based on machine learning is proposed for the present invention; Figure 2 The present invention proposes a multimodal model for predicting compound properties. Figure 3 A processing sample diagram of the method for identifying new halogenated pollutants with high environmental risks based on machine learning proposed in the present invention. DETAILED DESCRIPTION

[0044] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be pointed out that the embodiments described below are intended to facilitate the understanding of the present invention and do not have any limiting effect on the present invention.

[0045] like Figure 1 As shown, the process of the method for identifying new halogenated pollutants with high environmental risks based on machine learning of the present invention mainly includes five steps: S1. Obtain high-resolution mass spectrometry data of the sample, convert the file format of the original mass spectrometry data, and use the peak-lifting algorithm to extract the characteristic information in the mass spectrometry data; S2. Use the machine learning model to predict the presence of chlorine and bromine elements in the feature according to the characteristic information, and perform feature filtering; S3. Perform compound identification on the filtered features based on the accurate mass and secondary mass spectrometry information; S4. Use the machine learning model to predict the persistence, bioaccumulation and toxicity characteristics of the compound represented by the feature; S5. Perform retrospective detection on the processed mass spectrometry features.

[0046] The implementation of the process steps of this method is mainly based on the following two units and two models: mass spectrometry data preprocessing unit, feature matching unit, feature chlorine and bromine prediction model and compound property prediction model.

[0047] The main task of the mass spectrometry data preprocessing unit is to convert the original mass spectrometry data into a data format suitable for this method. The specific process includes using the mzConvert program to convert the original mass spectrometry file into the mzML format, extracting the peak features in the mass spectrometry data with the help of the OpenMS tool, and further parsing the data using pymzml, assigning the corresponding secondary mass spectrometry information and isotope pattern to the target feature, and finally calculating the mass-to-charge ratio difference and intensity ratio of the isotope pattern to generate a feature vector. The features include mass-to-charge ratio, retention time, secondary mass spectrometry information, isotope pattern, mass-to-charge ratio difference and intensity ratio of the isotope pattern, etc.

[0048] The task of the feature matching unit is to measure the similarity between the current mass spectrum feature and the data in the database or other mass spectra through matching calculation, so as to achieve compound identification or retrospective detection. This unit includes three parts: accurate mass matching, secondary mass spectrum peak matching, and secondary mass spectrum cosine similarity calculation.

[0049] The task of the feature chlorine and bromine prediction model is to predict the presence of chlorine and bromine elements in the feature based on the isotope pattern of the feature. The specific process is to collect the composition and structure data of the molecule through the Pubchem and CAS Common Chemistry platforms, calculate the theoretical isotope pattern based on the isotope abundance distribution law, and use the mass-to-charge ratio difference and intensity quotient of the theoretical pattern as feature vectors to train a machine learning model that can predict the presence of chlorine and bromine elements in the feature. Here, the model architecture is a random forest.

[0050] The task of the compound property prediction model is to predict whether the compound represented by the feature has persistence, bioaccumulation and toxicity based on the structural information of the identified compound. This process collects the structural information of compounds that have been publicly disclosed in the literature and have or do not have persistence, bioaccumulation and toxicity, and uses the molecular graph, Morgan molecular fingerprint and molecular descriptor converted from the structural information as input features to train a machine learning model that can predict the properties of the compound. The model architecture here is a self-developed multimodal model.

[0051] The machine learning model for realizing the prediction of chlorine and bromine conditions in step S2 is a random forest architecture, and the specific implementation method is as follows: First, the composition and structure data of molecules were collected through Pubchem and CAS Common Chemistry platforms. The molecular composition was represented by molecular formula, accurate mass, etc., and the structural information was represented by SMILES (Simplified molecular input line entry system) string. The data were classified based on whether the molecular composition contained chlorine or bromine and the number of specific chlorine and bromine atoms in the molecular formula to form a data set. Then, based on the composition information of the molecules in the data set, the theoretical isotope pattern of the molecules is calculated in combination with the isotope abundance distribution law of natural elements, and the difference between the theoretical isotope mass-to-charge ratios and the quotient between the intensities are calculated as feature vectors to construct a model training data set; Finally, the processed training dataset was used to train a random forest model, which was able to predict the number of chlorine and bromine atoms in a molecule based on its isotope pattern and was applied to predict the presence of chlorine and bromine elements in actual mass spectrometry data features.

[0052] Theoretical isotope intensities and characteristics are calculated in the form of molecular formula C m H n Some examples of compounds are as follows: ; ; ; Among them, Int M Int represents the intensity of the lowest mass isotope pattern. M+1 Indicates the intensity of the isotopic pattern with the mass difference from the lowest mass rounded to 1. p C 12 represents the appearance ratio of carbon-12 isotope, p C 13 represents the appearance ratio of carbon-13 isotope, p H 1 represents the occurrence ratio of hydrogen-1 isotope, p H 2 Indicates the occurrence ratio of hydrogen-2 isotope, F M+1 / M Represents the intensity ratio of adjacent mass isotope patterns.

[0053] Since the intensity calculation of isotope patterns involves the merging of local adjacent isotope patterns, the merging process adopts the intensity accumulation and mass-to-charge ratio based on intensity weighted calculation method: ; ; Among them, Int merged Int represents the intensity of the combined isotope pattern. i Indicates i The intensity of the adjacent isotope pattern to be merged, m / z merged represents the mass-to-charge ratio calculated by intensity weighting after isotope pattern merging, m / z i Indicates i The mass-to-charge ratios of adjacent isotope patterns to be merged, n is the number of adjacent isotope patterns to be merged.

[0054] As a preferred embodiment of the present invention, the specific method for achieving compound identification based on accurate mass in step S3 is: When the mass-to-charge ratio of the current feature deviates from the mass of a compound in the accurate mass database by less than or equal to 0.0005%, the feature is considered to match the compound and is added to the candidate queue; When the mass-to-charge ratio of the current feature deviates from the mass of a compound in the accurate mass database by more than 0.0005%, it is considered that the feature fails to match the compound.

[0055] The mass deviation is calculated as follows:

[0056] Among them, m t Indicates the exact mass of a compound C in the database, m / z s Indicates the mass-to-charge ratio of the current feature, df indicates m t and m / z s quality deviation.

[0057] As a preferred embodiment of the present invention, the specific method for realizing compound identification based on secondary mass spectrometry information in step S3 is: When the cosine similarity between the secondary mass spectrum information of the current feature and the secondary mass spectrum information of a compound in the database is greater than or equal to 0.7, the match is considered successful and the priority of the candidate is increased; When the cosine similarity between the secondary mass spectrum information of the current feature and the secondary mass spectrum of a compound in the database is less than 0.7, it is considered that the match fails and the priority of the candidate remains unchanged.

[0058] The calculation method of cosine similarity between secondary mass spectrometry information is as follows:

[0059] Among them, A represents the secondary mass spectrometry information of the current feature in the form of a vector, B represents the secondary mass spectrometry information in the database in the form of a vector, and A i and B iRepresents a single value in two secondary mass spectrum vectors, cos similarity Indicates the cosine similarity of the secondary mass spectrum information A and B. Before calculation, the two secondary mass spectrum vectors need to be aligned according to the mass-to-charge ratio information, and the empty positions are filled with the value 0.

[0060] As a preferred embodiment of the present invention, the machine learning model for predicting the persistence, bioaccumulation and toxicity characteristics of the compound in step S4 is a multimodal model, and the specific implementation method is as follows: First, we collected data on compounds that were publicly known in the literature to have or not have persistence, bioaccumulation, and toxicity characteristics, and constructed a dataset in the form of SMILES strings and corresponding characteristic data. Then, the SMILES string was converted into three types of features: molecular graph, Morgan fingerprint, and molecular descriptor, which are homologous but have very different chemical information representation forms. The features were standardized and the constructed data set was randomly divided into training set, test set, and validation set to prepare data for the subsequent property prediction model. Finally, the persistence, bioaccumulation and toxicity property data were used as target variables, and the features of compound structure information conversion (molecular graph, Morgan molecular fingerprint, molecular descriptor) were used as input variables to train a multimodal model, and the best performance model was obtained through cross-validation and model integration optimization methods to predict the properties of identified compounds.

[0061] The architecture of the multimodal model is as follows Figure 2 As shown in the figure, the multimodal model aims to integrate three feature representation methods: molecular graph, Morgan molecular fingerprint, and molecular descriptor, to comprehensively extract the structural information and physical and chemical properties of molecules. The model adopts a parallel feature extraction strategy to process data of different modalities separately, and performs feature fusion in the final stage to obtain more comprehensive molecular structure information, thereby improving the prediction ability and stability of the model. The overall architecture is mainly composed of four main modules: molecular graph feature extraction module, Morgan molecular fingerprint feature extraction module, molecular descriptor feature extraction module, and feature fusion and prediction module.

[0062] Molecular graph feature extraction module: This module is used to learn the topological structure of molecules and their local chemical environment. The module uses the GINE convolution layer to extract features and combines the batch normalization layer to normalize the output of the GINE convolution layer. Due to the complex molecular structure, two GINE convolutions and batch normalization operations are used to enhance the model's perception of complex molecular structures. Subsequently, all atomic features are aggregated through the global pooling layer to generate molecular-level features of fixed dimensions.

[0063] Morgan Molecular Fingerprint Feature Extraction Module: This module is used to extract the global structural information of molecules. The module first adjusts the Morgan molecular fingerprint to a 2048-dimensional shape through a feature shaping layer to adapt to the subsequent neural network layer. Subsequently, a multi-layer fully connected network is used to perform linear and nonlinear transformations on the Morgan molecular fingerprint, thereby converting it from a sparse representation to a more compact embedded representation to improve feature expression capabilities.

[0064] Molecular descriptor feature extraction module: This module is used to characterize the physical and chemical properties of molecules to supplement the structural description of molecules by molecular graphs and Morgan molecular fingerprints. The module first adjusts the molecular descriptors to 217 dimensions through the feature shaping layer to adapt to the subsequent neural network layer. Subsequently, the molecular descriptors are reduced in dimension using a fully connected layer to extract more representative and effective molecular features, thereby enhancing the expressive power of the model.

[0065] Feature fusion and prediction module: This module splices the extracted molecular graph features, Morgan molecular fingerprint features, and molecular descriptor features to integrate information from different modalities and predict compound properties. First, a fully connected layer is used to perform nonlinear mapping on the spliced ​​features to fully learn the interactions between molecular features and their modalities. Subsequently, a batch normalization layer is used to normalize the features to prevent the features of a certain modality from dominating the training process. Finally, the fused features are transformed through a fully connected layer, and the softmax activation function is used to predict the final result.

[0066] The specific process of input feature conversion is as follows: The molecular graph feature constructs a two-dimensional topological structure of the molecule through the SMILES string, and represents the types and related characteristics of atoms and chemical bonds in the molecular graph in numerical form. Among them, atomic features include 8 categories and chemical bond features include 4 categories. The Morgan molecular fingerprint is based on the SMILES string and uses the RDKit calculation tool to generate a 2048-dimensional binary Morgan molecular fingerprint. 0 or 1 is used to indicate whether a molecular structure fragment exists in the molecule. If it does not exist, it is 0, and if it exists, it is 1. The molecular descriptor feature is also based on the SMILES string. The Mordred or RDKit calculation tool is used to generate a 217-dimensional molecular descriptor to quantify the physical and chemical properties of the compound in numerical form.

[0067] The specific implementation principles of each network layer in the model are as follows: Based on the traditional graph convolutional network, the GINE convolutional layer further enhances the model's ability to learn information about edges in the graph structure by weighting edge features. The specific processing formula is as follows:

[0068] in, h v andh u Respectively represent nodes v and u The eigenvector of N ( v ) represents a node v The set of adjacent nodes of ; e uv Representation Node u and v The edge feature vector between them; M represents a function that performs weighted calculations on node feature vectors and edge feature vectors; MLP represents a multi-layer perceptron, which is used to perform nonlinear transformation on the feature vector of each node; h’ v Indicates the updated node v characteristics.

[0069] The fully connected layer is used to perform linear and nonlinear transformations on the input features. The formula is as follows: y = W • x + b in, x The feature vector representing the input, usually the output of the previous layer; W Represents the weight matrix of the fully connected layer; b represents the bias term; y represents the output of the fully connected layer.

[0070] The batch normalization layer is used to standardize the data of each batch during the training process, helping to speed up the training and stabilize the model. The formula is as follows:

[0071] in, x i Indicates i Input features of samples; μ B represents the batch mean; σ B represents batch standard deviation; Represents the normalized features.

[0072] The global pooling layer is used to aggregate the features of all nodes in the graph into a fixed-size graph-level feature vector. The formula is as follows:

[0073] in, h i Is a node i The eigenvector of N is the total number of nodes in the graph; hglobal is the aggregated global feature vector.

[0074] The evaluation indicators of model performance are Accuracy, Macro Precision, Macro Recall, and Macro F1 Score: ; ; ; ; Among them, Accuracy represents accuracy, Macro Precision represents macro average precision, MacroRecall represents macro average recall, Macro F1 Score represents macro average F1 score, TP represents the number of positive samples correctly predicted as positive by the model, TN represents the number of negative samples correctly predicted as negative by the model, FP represents the number of negative samples incorrectly predicted as positive by the model, FN represents the number of positive samples incorrectly predicted as negative by the model, N in macro average precision and macro average recall represents the total number of categories predicted by the model, TP represents the number of positive samples correctly predicted as positive by the model, TN represents the number of negative samples correctly predicted as negative by the model, FP represents the number of negative samples incorrectly predicted as positive by the model, FN represents the number of positive samples incorrectly predicted as negative by the model, N represents the total number of categories predicted by the model, and TP represents the number of positive samples correctly predicted as positive by the model. i Indicates that the model is predicting i The number of correct predictions for the class, FP i Indicates that the model mistakenly predicts other classes as i The number of classes, FN i It means that the model mistakenly predicts the i The number of other classes that the class was predicted to be.

[0075] Finally, the multimodal model was tested on the test set and achieved 97% accuracy, 90% macro-average precision, 89% macro-average recall, and 89% macro-average F1-score, which is the current best performance in predicting persistence, bioaccumulation, and toxicity characteristics.

[0076] As a preferred solution of the present invention, the specific method for implementing feature retrospective detection in step S5 is: When the mass-to-charge ratio of a feature deviates from that of the feature in the historical mass spectrometry data by less than or equal to 0.0005%, and the mass-to-charge ratio deviation of at least two peaks in the secondary mass spectrum is less than or equal to 0.0005%, and the cosine similarity is greater than 0.7, the feature is considered to have successfully matched the feature in the historical data and added to the retrospective detection queue. Otherwise, the feature is considered to have failed to match.

[0077] The process of processing specific sample mass spectrometry data of the present invention is as follows Figure 3As shown: First, the mass spectrometry data preprocessing unit converts the original mass spectrum file into the required feature and feature vector format. The processing result is in the form of: a feature ID is 2, the mass-to-charge ratio is 450.9261, the retention time is 1010.84, and the secondary mass spectrum information is [50.1039, 59.0138, 68.9957, … , 450.9257, 451.3118], etc.

[0078] Next, the trained random forest model is used to predict the feature vector to determine whether it contains chlorine or bromine elements, and the corresponding feature filtering is performed. The processing result is in the form of: the feature with ID 2 is predicted to contain chlorine elements, and 1 molecule contains 2 chlorine atoms. The feature can be retained and further analyzed.

[0079] Then, the features containing chlorine and bromine elements that have been filtered and retained are subjected to accurate mass matching and secondary mass spectrum cosine similarity calculation to achieve accurate identification of the compounds. The processing results are in the form of accurate mass matching of the feature with ID 2, and the candidate molecular formula includes C 12 HCl 2 F 11 N 2 , C 18 H 10 Clio 4 , C 25 H 4 Cl 2 NO 2 S, C 12 H 4 Cl 2 F 6 N 4 O 2 S, etc.; after the secondary mass spectrum information cosine similarity calculation, the highest score was 0.8914, and the identification result was fipronil sulfone, and the molecular formula was C 12 H 4 Cl 2 F 6 N 4 O 2 S is consistent with the accurate mass match result.

[0080] Next, the structural information of the identified compounds was input into the trained multimodal model to predict the persistence, bioaccumulation and toxicity characteristics to assess whether the compound has potential environmental risks. The processing results are in the form of: the model predicts that the feature with ID 2 may represent the fipronil sulfone, which does not have the persistence, bioaccumulation and toxicity characteristics at the same time and may not have a high environmental risk.

[0081] Finally, the features that have been screened, identified, and predicted are compared with the features in the historical mass spectrometry data, and the feature matching unit is used to find out whether there are similar features, so as to perform retrospective detection of pollutants. The processing result is in the form of: the feature finds similar features in the historical mass spectrometry data file, indicating that new halogenated pollutants with high environmental risks are also present in the source environmental samples of the historical mass spectrometry data.

[0082] The embodiments described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for identifying new halogenated pollutants based on machine learning, characterized in that: include: Step S1: obtaining original high-resolution mass spectrum data of the sample to be tested, and extracting characteristic information from the original high-resolution mass spectrum data; Step S2: using a feature halogen-containing prediction model to predict the presence of halogen in the feature based on the extracted feature information, and filtering out features that do not contain halogen; Step S3: performing compound identification and retrospective detection on the filtered features based on accurate mass and secondary mass spectrum information; Step S4: Use the compound property prediction model to predict the persistence, bioaccumulation and toxicity properties of the identified compounds.

2. The method for identifying new halogenated pollutants based on machine learning according to claim 1, characterized in that: Step S1 includes: obtaining the original high-resolution mass spectrometry data of the sample to be tested, using a peak-lifting algorithm to extract and analyze the peak features in the original high-resolution mass spectrometry data, assigning corresponding secondary mass spectrometry information and isotope patterns to the peak features, calculating the mass-to-charge ratio difference and intensity ratio of the isotope pattern, and generating a characteristic vector of the peak features.

3. The method for identifying new halogenated pollutants based on machine learning according to claim 1, characterized in that: The method for constructing the characteristic halogen-containing prediction model includes: Step S2-1: Collecting the composition and structure information data of the molecule, classifying the data based on whether the composition of the molecule is halogen and the number of halogens contained, to form a data set; Step S2-2: According to the composition information of the molecules in the data set, the theoretical isotope pattern of the molecules is calculated in combination with the isotope abundance distribution law of natural elements, and the difference between the theoretical isotope mass-to-charge ratios and the intensity ratio are calculated as feature vectors to construct a model training data set; Step S2-3: Use the model training data set to train the random forest model to obtain a characteristic halogen-containing prediction model.

4. The method for identifying new halogenated pollutants based on machine learning according to claim 1, characterized in that: Step S3 includes: Step S3-1: comparing the accurate mass and secondary mass spectrum information of the feature with the compounds in the accurate mass database to identify the compound; Step S3-2: Based on the accurate mass and secondary mass spectrum information of the feature, the historical mass spectrum feature is compared to achieve retrospective detection.

5. The method for identifying new halogenated pollutants based on machine learning according to claim 4, characterized in that: Step S3-1 includes: Match the current feature based on exact mass: When the mass-to-charge ratio of the current feature and the mass deviation of a compound C in the accurate mass database are less than or equal to 0.0005%, the current feature is considered to match compound C and compound C is added to the candidate queue; otherwise, the current feature is considered to fail to match compound C; Match the current feature based on the secondary mass spectrum information: When the cosine similarity between the secondary mass spectrum information of the current feature and the secondary mass spectrum information of a compound C in the database is greater than or equal to 0.7, it is considered that the current feature successfully matches compound C, and the priority of compound C in the candidate queue is increased; otherwise, it is considered that the current feature fails to match compound C, and the priority of compound C in the candidate queue remains unchanged; The compound with the highest priority in the candidate queue is selected as the compound matching the current feature.

6. The method for identifying new halogenated pollutants based on machine learning according to claim 4, characterized in that: Step S3-2 includes: The mass-to-charge ratio deviation of the current feature and the mass-to-charge ratio of feature D in the historical mass spectrum data is less than or equal to 0.0005%; at the same time, when the mass-to-charge ratio deviation of at least two peaks in the secondary mass spectra of the current feature and feature D in the historical mass spectrum data is less than or equal to 0.0005%, and the cosine similarity is greater than 0.7, it is considered that the current feature matches the feature D in the historical mass spectrum data successfully; otherwise, the backtracking matching of the current feature fails.

7. The method for identifying new halogenated pollutants based on machine learning according to claim 1, characterized in that: The methods for constructing compound property prediction models include: Step S4-1: Collect data on compounds that are publicly known in the literature to have or not have persistence, bioaccumulation and toxicity characteristics to form a data set; the compound data includes structural information and persistence, bioaccumulation and toxicity characteristics of the compound; Step S4-2: performing feature conversion and standardization processing on the structural information of the compounds in the data set, and randomly dividing the data set into a training set, a test set, and a validation set to form a training data set for the model; Step S4-3: Using the persistence, bioaccumulation and toxicity characteristics data of the compound as the target variables and the characteristics of the compound structure information conversion as the input variables, the multimodal model is trained to obtain a compound characteristic prediction model.

8. The method for identifying new halogenated pollutants based on machine learning according to claim 7, characterized in that: The structural information of the compound in the compound data is the SMILES string of the compound; The feature conversion of the structural information of the compounds in the data set includes: converting the SMILES string into three features: molecular graph, Morgan molecular fingerprint, and molecular descriptor.

9. The method for identifying new halogenated pollutants based on machine learning according to claim 7, characterized in that: The multimodal model includes: The molecular graph feature extraction module takes the molecular graph as input and extracts the molecular graph features of the compound; Morgan molecular fingerprint feature extraction module, which takes Morgan molecular fingerprint as input and extracts the Morgan molecular fingerprint features of the compound; The molecular descriptor feature extraction module takes the molecular descriptor as input and extracts the molecular descriptor features of the compound; The feature fusion and prediction module combines molecular graph features, Morgan molecular fingerprint features and molecular descriptor features to predict compound properties.

10. A device for identifying new halogenated pollutants based on machine learning, characterized in that: include: The data acquisition and preprocessing module obtains the high-resolution mass spectrometry data of the sample to be tested and uses the peak-lifting algorithm to extract the characteristic information in the high-resolution mass spectrometry data; The feature halogen content prediction module uses the feature halogen content prediction model to predict the presence of halogens in the feature based on the extracted feature information and perform feature filtering; Feature matching module, which performs compound identification and retrospective detection on filtered features based on accurate mass and secondary mass spectrometry information; The compound property prediction module uses the compound property prediction model to predict the persistence, bioaccumulation and toxicity properties of identified compounds.

Citation Information

Patent Citations

  • Processing method and device of stacked integrated model for molecular attribute prediction

    CN116343950A

  • Method and device for identifying or assisting in identifying new pollutants in water based on deep learning and computer readable storage medium

    CN118067897A

  • Brominated pollutant and metabolite multi-strategy annotation method thereof based on liquid chromatography-mass spectrometry

    CN118566391A

  • Soil organic pollutant identification method and system based on combination of AI and high-throughput screening

    CN119580881A

  • System and method for simulation of marine pollution dispersion using numerical tracer technique

    US20240363201A1

Cited By

  • Hazardous chemical substance feature extraction method and device based on mass spectrum image, terminal and medium

    CN120766256A

  • Perfluorinated compound identification method and device based on machine learning

    CN122024934A

  • Comprehensive detection method and system for neonicotinoid compounds in environment

    CN122109413A