Toxicity prediction method based on compound secondary mass spectrometry data

CN117409871BActive Publication Date: 2026-08-07RES CENT FOR ECO ENVIRONMENTAL SCI THE CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
RES CENT FOR ECO ENVIRONMENTAL SCI THE CHINESE ACAD OF SCI
Filing Date
2023-10-25
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

近期,一项水样本非靶向分析提出从质谱信息预测生态毒理学指标——半数致死浓度值(Lethal Concentration,LC50),但是其模型学习过程仍然基于已知结构的分子指纹特征,只在验证模型过程中使用二级质谱数据作为输入,其建立的仍然是基于结构的化学品毒性预测模型

Benefits of technology

[0020] Based on the above technical solution, the toxicity prediction method and model establishment method based on compound secondary mass spectrometry data provided in this disclosure have at least one of the following beneficial effects:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117409871B_ABST
    Figure CN117409871B_ABST
Patent Text Reader

Abstract

The disclosure provides a toxicity prediction model establishment method based on compound secondary mass spectrum data, comprising: performing chemical structure cleaning on the obtained secondary mass spectrum data of known compounds and the compounds involved in the toxicity data of known chemicals, matching the compounds involved in the secondary mass spectrum data and the toxicity data according to the standardized chemical structure, obtaining the secondary mass spectrum data of the common compounds and the corresponding binary labels of the presence or absence of toxicity; for the common compounds, converting each secondary mass spectrum data into a molecular structure feature probability vector, establishing a total data set containing the molecular structure feature probability vector and the binary label of the presence or absence of toxicity; dividing the total data set into a training set, a validation set and a test set, and constructing a toxicity prediction model with the molecular structure feature probability vector as the input and the presence or absence of toxicity as the output. The disclosure also provides a toxicity prediction method, comprising: inputting the secondary mass spectrum data of the compound to be predicted into the toxicity prediction model for toxicity prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure belongs to the field of environmental sample safety assessment technology, and more specifically relates to a toxicity prediction method and model building method based on compound secondary mass spectrometry data. Background Technology

[0002] The explosive growth of the chemical industry has led to a dramatic increase in the use of agricultural chemicals, daily chemical products, and food additives. Humans are exposed to a wide variety of compounds through multiple pathways, including environmental pollution, nutritional intake, and the use of cosmetics and pharmaceuticals. Many chemicals used in daily life, while not acutely toxic, still pose potential risks. These risks include the formation of highly toxic metabolites and transformation products, or their persistence in the environment and accumulation in the food chain, threatening ecosystems and human health. The ecological damage and health effects of compounds are assessed through in vivo or in vitro tests, but these tests often use the toxic effects of a single compound as the evaluation endpoint.

[0003] In recent years, chemical toxicity prediction models based on quantitative structure-activity relationships (QSARs) have seen rapid development in the toxicity screening and risk assessment of known chemicals. This is thanks to highly integrated public toxicity reference databases and the continuously expanding database of high-throughput in vitro testing (qHTS) of chemicals in relevant toxicity pathways. Chemical toxicity prediction models typically use molecular descriptors or molecular fingerprints based on prior knowledge as learning and prediction objects to predict toxicity endpoints.

[0004] In practical food safety monitoring and environmental safety assessments, samples often contain complex coexisting contaminants and matrices, significantly increasing the difficulty of risk assessment. Gas chromatography or liquid chromatography combined with high-resolution mass spectrometry (GC- / LC-HRS) has become a common method for non-targeted analysis of complex environmental samples, aiming to discover, identify, and quantify unknown contaminants in these samples. However, to achieve Level 1 confidence in identifying unknown components in complex samples, different complementary methods are needed to progressively determine the molecular formula and structure, followed by validation using standards—a very time-consuming and labor-intensive process. Furthermore, data obtained from non-targeted methods typically need to be combined with targeted and suspected substance screening before analysis can proceed. This means that only a small portion of the molecular characteristics obtained from non-targeted analysis can be identified, leaving a large number of potential risk components unidentified.

[0005] In recent years, numerous studies have employed in-silico methods for structural annotation of secondary mass spectra, aiming to discover the compound structures corresponding to these spectra, particularly in the fields of untargeted metabolomics and drug design. Machine learning and deep learning have been widely applied to structural annotation due to their ability to handle complex features, strong learning capabilities, and self-optimization. Furthermore, the continuously expanding libraries of annotated spectra for small molecules contribute to improving the sensitivity and specificity of machine learning models.

[0006] The primary and secondary mass spectrometry data obtained by non-targeted methods contain rich information, such as retention time, precise mass number, and ion fragmentation information. It is known that mass spectra obtained by scanning ion fragments from ion source ionization contain structural fragmentation information, and the structure-activity relationship model has been widely validated. Therefore, it is feasible to directly determine the toxicity of corresponding compounds from secondary mass spectrometry data. Methodological research on structural annotation of secondary mass spectra using computational methods provides valuable prior knowledge for the characteristic representation of secondary mass spectra, while existing public mass spectrometry libraries and toxicity databases provide sufficient sample support for statistical learning. Recently, a non-targeted analysis of water samples proposed predicting ecotoxicological indicators—the median lethal concentration (LC50)—from mass spectrometry information. 50 However, its model learning process is still based on the molecular fingerprint characteristics of known structures, and only uses secondary mass spectrometry data as input during the model validation process. The model it establishes is still a structure-based chemical toxicity prediction model.

[0007] In summary, when using non-targeted methods to analyze complex environmental samples, the structural annotation process of the spectra relies on existing spectral libraries or suspected substance libraries. Only a small number of molecular features in the sample can be clearly identified, and the identification process is cumbersome. Therefore, it is essential to effectively determine the priority of the analysis and quickly assess the environmental risk of the mixture sample before conducting experimental analysis of the pollutants in the sample. Summary of the Invention

[0008] In view of this, this disclosure addresses the problem of how to achieve toxicity prediction using high-resolution secondary mass spectrometry data obtained from non-targeted analysis, and proposes a toxicity prediction method and model building method based on compound secondary mass spectrometry data, in order to at least partially solve at least one of the above-mentioned technical problems.

[0009] As the first aspect of this disclosure, a method for establishing a toxicity prediction model based on secondary mass spectrometry data of compounds is proposed, including:

[0010] Obtain secondary mass spectrometry data of known compounds and toxicity data of known chemicals;

[0011] Chemical structure cleaning is performed on the secondary mass spectrometry data of known compounds and the toxicity data of known chemicals to obtain the standardized chemical structures of the compounds. Based on the standardized chemical structures, the secondary mass spectrometry data and the compounds involved in the toxicity data are matched to obtain the secondary mass spectrometry data of common compounds and the binary label of toxicity of interest for the corresponding common compounds.

[0012] For common compounds, each secondary mass spectrometry data is converted into a molecular structure feature probability vector S, and a total dataset containing the molecular structure feature probability vector S and a binary label of toxicity is established.

[0013] The total dataset is divided into training, validation, and test sets. A toxicity prediction model is constructed, taking the molecular structure feature probability vector S as input and the presence or absence of toxicity as output, including:

[0014] Based on multiple sets of preset hyperparameters of the prediction model used, the prediction model is trained using the training set, and the multiple sets of preset hyperparameters of the prediction model are optimized using the validation set to obtain the toxicity prediction model of the toxicity of interest and determine the toxicity judgment threshold.

[0015] The generalization performance of the toxicity prediction model was evaluated using a test set.

[0016] This disclosure also provides a method for predicting toxicity based on secondary mass spectrometry data of compounds, including:

[0017] Obtain secondary mass spectrometry data of the compound to be predicted;

[0018] After converting the secondary mass spectrometry data to be predicted into a molecular structure feature probability vector S, it is input into the toxicity prediction model and outputs the toxicity prediction probability value p of the corresponding compound for the secondary mass spectrometry data to be predicted. The toxicity prediction model is obtained by the above-mentioned toxicity prediction model establishment method.

[0019] When p is greater than or equal to the toxicity threshold, the compound corresponding to the predicted secondary mass spectrometry data is toxic; when p is less than the toxicity threshold, the compound corresponding to the predicted secondary mass spectrometry data is not toxic.

[0020] Based on the above technical solution, the toxicity prediction method and model establishment method based on compound secondary mass spectrometry data provided in this disclosure have at least one of the following beneficial effects:

[0021] (1) In the embodiments of this disclosure, this disclosure establishes a method for directly predicting the toxicity of a compound corresponding to a high-resolution secondary mass spectrometry data (hereinafter referred to as secondary mass spectrometry data). It fully combines existing mass spectrometry databases and toxicity databases, and combines them with machine learning algorithms (XGBoost) to establish a toxicity prediction model, thereby realizing the direct prediction of toxicity from high-resolution secondary mass spectrometry data of a compound.

[0022] (2) In the embodiments of this disclosure, the high-resolution secondary mass spectrometry data (hereinafter referred to as secondary mass spectrometry data) and the compounds involved in the toxicity data are chemically cleaned to obtain the standardized chemical structures of the compounds. Then, the compounds involved in the secondary mass spectrometry data and the toxicity data are matched according to the standardized chemical structures to obtain the secondary mass spectrometry data of common compounds and the binary labels of toxicity of interest for the corresponding common compounds. By converting mass spectrometry information containing various information and of varying lengths into a fixed-length molecular structure feature probability vector, a feature representation of mass spectrometry data is completed, providing the possibility of applying compound secondary mass spectrometry data to machine learning and toxicity prediction models.

[0023] (3) In the embodiments of this disclosure, a toxicity prediction model is established using known secondary mass spectrometry data of compounds and toxicity labels. The obtained toxicity prediction model can effectively predict the toxicity of compounds corresponding to the secondary mass spectrometry data. It is expected to provide a basis for determining the analysis priority in non-targeted analysis, which is conducive to quickly judging the toxicity of pollutants in complex environmental samples. It has broad application prospects in the fields of toxicity prediction of complex samples, environmental safety assessment, and health risk assessment. Attached Figure Description

[0024] Figure 1 This is a flowchart illustrating the method for establishing a toxicity prediction model based on compound secondary mass spectrometry data in this embodiment of the present disclosure.

[0025] Figure 2 This is a schematic diagram of the toxicity prediction method based on compound secondary mass spectrometry data in the embodiments of this disclosure;

[0026] Figure 3 This is a schematic diagram illustrating the principle of predicting the toxicity of a compound corresponding to a given secondary mass spectrometry data based on the predicted secondary mass spectrometry data in an embodiment of this disclosure. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.

[0028] A review of current literature on non-targeted methods for assessing the environmental risk of complex samples reveals several drawbacks: non-targeted analysis of environmental samples typically requires integration with targeted and suspected substance screening; the scope of mass spectrometry comparison is limited; and the annotation of mass spectra is extremely cumbersome. This results in only a small fraction of the molecular features obtained from non-targeted analysis being clearly identified, leaving a large number of potential risk components unidentified. The primary objective of this disclosure is to provide a novel method for prioritizing mass spectrometry data analysis in non-targeted analysis, establishing a toxicity prediction model from high-resolution secondary mass spectrometry of compounds to toxicity. This model is expected to be used to rapidly assess the environmental risk of mixture samples, identifying secondary mass spectrometry features with higher environmental risk from the non-targeted analysis data for further explicit annotation.

[0029] The basic principle of this disclosure is to convert high-resolution secondary mass spectrometry data of compounds into fixed-length molecular structure feature probability vectors, and to use machine learning methods to establish a toxicity prediction model between the molecular structure feature probability vectors and corresponding toxicity tags. This method can directly predict the toxicity of compounds obtained by non-targeting methods and is expected to serve as a useful tool for quickly assessing the environmental risk of mixture samples and determining analytical priorities.

[0030] Figure 1 This is a flowchart illustrating the method for establishing a toxicity prediction model based on compound secondary mass spectrometry data in this embodiment of the present disclosure.

[0031] Specifically, such as Figure 1 As shown, the present disclosure provides a method for establishing a toxicity prediction model based on compound secondary mass spectrometry data, including steps S101-S112.

[0032] Steps S101-S102: Obtain secondary mass spectrometry data of known compounds and toxicity data of known chemicals.

[0033] Steps S103-S105: Perform chemical structure cleaning on the secondary mass spectrometry data of known compounds and the toxicity data of known chemicals to obtain the standardized chemical structures of the compounds. Match the secondary mass spectrometry data and the compounds involved in the toxicity data according to the standardized chemical structures to obtain the secondary mass spectrometry data of common compounds and the binary label of toxicity of interest for the corresponding common compounds.

[0034] Steps S106-S107: For common compounds, each secondary mass spectrometry data is converted into a molecular structure feature probability vector S, and a total dataset containing the molecular structure feature probability vector S and binary labels of toxicity is established.

[0035] Steps S108-S112: Divide the total dataset into training, validation, and test sets, and construct a toxicity prediction model that takes the molecular structure feature probability vector S as input and toxicity status as output, including:

[0036] Steps S108-S111: Based on multiple sets of preset hyperparameters of the prediction model used, train the prediction model using the training set, optimize the multiple sets of preset hyperparameters of the prediction model using the validation set, obtain the toxicity prediction model of the toxicity of interest, and determine the toxicity judgment threshold.

[0037] Step S112: Evaluate the generalization performance of the toxicity prediction model using the test set.

[0038] In the embodiments of this disclosure, before chemical structure cleaning of the compounds involved, the acquired high-resolution secondary mass spectrometry data is screened to retain secondary mass spectrometry data containing linear molecular characterization information of compound structures and sufficient secondary mass spectrum data. Subsequently, the screened high-resolution secondary mass spectrometry data (hereinafter referred to as secondary mass spectrometry data) and the compounds involved in the toxicity data are chemically cleaned to obtain the standardized chemical structures of the compounds. Then, based on the standardized chemical structures, the compounds involved in the secondary mass spectrometry data and the toxicity data are matched to obtain secondary mass spectrometry data of multiple common compounds and binary labels indicating toxicity. By converting mass spectrometry information containing various information and of varying lengths into a fixed-length molecular structure feature probability vector, and inputting it into a machine learning algorithm for training a toxicity prediction model, the possibility of applying the secondary mass spectrometry data of compounds to machine learning and toxicity prediction models is provided.

[0039] According to embodiments of this disclosure, in step S101, the acquired secondary mass spectrometry data of the known compound is the annotated secondary mass spectrometry data of the corresponding compound, which includes linear molecular characterization information of the compound structure, secondary mass spectrum data of the compound, the precise mass number and charge number of the precursor ions measured by the mass spectrometer, the ionization mode of the mass spectrometer, and the instrument type of the mass spectrometer. The secondary mass spectrum data of the compound is the spectrum obtained after the compound is ionized, where ions with different mass-to-charge ratios are analyzed by a mass analyzer and then detected and recorded, including information such as the mass-to-charge ratio of ion fragments and peak intensity. In step S102, the toxicity data of the known chemical includes linear molecular characterization information of the compound structure and a binary label indicating whether the compound is toxic. It should be noted that the compound in the known chemical may be the same as or different from the known compound.

[0040] According to an embodiment of this disclosure, in step S103, chemical structure cleaning is performed on the compounds involved in the secondary mass spectrometry data of known compounds and the toxicity data of known chemicals to obtain the standardized chemical structure of the compounds. Specifically, the chemical structure cleaning includes standardization, solvent removal, charge correction, and deionization.

[0041] Furthermore, before matching the toxicity data obtained after cleaning with the compounds involved in the secondary mass spectrometry data according to the standardized chemical structure, the process includes: removing compounds with opposite toxicity labels from the toxicity data based on the standardized chemical structure of the compounds to eliminate their influence and ensure the accuracy of toxicity prediction.

[0042] According to embodiments of this disclosure, in step S104, for a common compound, each secondary mass spectrometry data point is converted into a molecular structure feature probability vector S. This includes: inputting each secondary mass spectrometry data point into a mass spectrometry calculation model for calculation, and outputting a molecular structure feature probability vector S, where the length of the molecular structure feature probability vector S is M. For example, the open-source software SIRIUS (mass spectrometry calculation model) is used to calculate the molecular structure feature probability vector S of the secondary mass spectrometry data. Specifically, all information of each secondary mass spectrometry data point, including the precise mass number, charge number, adduct ion, instrument and / or molecular formula, is used as input. Each secondary mass spectrometry data point is input into the open-source software SIRIUS for calculation and conversion, and a molecular structure feature probability vector S of length M is output. It should be noted that when using the open-source software SIRIUS to calculate the molecular structure feature probability vector S of the secondary mass spectrometry data, there are situations where mass spectra with multiple charges cannot be calculated, secondary mass spectrometry data points with the same compound and the same precise mass number are combined for calculation, and the mass spectrometry information is insufficient for calculation, resulting in some secondary mass spectrometry data points failing to yield a molecular structure feature probability vector.

[0043] According to an embodiment of this disclosure, in step S107, a total dataset is established based on the molecular structure feature probability vector S and the binary label of toxicity. The established total dataset includes the secondary mass spectrometry feature matrix D of the common compounds and the binary label vector of toxicity of interest for the corresponding common compounds; wherein, the size of the secondary mass spectrometry feature matrix D is N×M, where N is the number of molecular structure feature probability vectors S calculated from the secondary mass spectrometry data of the common compounds, M is the length of the molecular structure feature probability vector S, and each element D in D is... i,j Let S represent the probability that the shared compound corresponding to the molecular structure feature probability vector S calculated from the i-th secondary mass spectrometry data contains a specific molecular structure feature j. The length of the binary label vector T for the toxicity of interest of the shared compound is N, and each element T in T... i ∈{0,1} indicates whether the common compound corresponding to the molecular structure feature probability vector S calculated from the i-th secondary mass spectrometry data has the toxicity of interest. The label "0" indicates non-toxicity and the label "1" indicates toxicity.

[0044] According to embodiments of this disclosure, in steps S108 to S110, the total dataset consisting of the secondary mass spectrometry feature matrix D of the common compounds and the label vector T of the toxicity of interest is randomly divided into a training set, a validation set, and a test set using a stratified sampling method in an a:b:c ratio, ensuring that the proportion of binary toxicity category labels remains consistent within each dataset. Then, the prediction model is trained using the training set, and multiple preset hyperparameters of the prediction model are optimized using the validation set to obtain a toxicity prediction model for the toxicity of interest and determine the toxicity determination threshold. The generalization performance of the toxicity prediction model is evaluated using the test set, thereby constructing a toxicity prediction model that takes the molecular structure feature probability vector S as input and the presence or absence of toxicity as output.

[0045] According to embodiments of this disclosure, XGBoost is selected as the prediction model. The XGBoost prediction model is based on Gradient Boosting Decision Tree (GBDT). The XGBoost prediction model uses an additive model and a forward distribution algorithm. Its base model is a decision tree model. It iterates num_boost_round times. The fitting objective of each new decision tree is the negative gradient of the objective function of the previous tree. The objective function of the XGBoost prediction model is the loss function plus a regularization term. The final prediction result is the sum of all decision trees.

[0046] According to an embodiment of this disclosure, in step S111, before training the prediction model, multiple sets of hyperparameters are preset for the prediction model so that the multiple sets of preset hyperparameters of the prediction model can be optimized using a validation set to obtain the optimized hyperparameters of the prediction model. The preset hyperparameters of the prediction model are {booster, objective, num_boost_round, learning_rate, gamma, max_depth, min_child_weight, subsample, colsample_bytree, alpha, lambda};

[0047] Here, booster defines the type of base learner; objective defines the loss function to be minimized; num_boost_round is the number of iterations of the decision tree; learning_rate is the shrinking step size during the update process; gamma is the minimum loss function decrease value required for node splitting; max_depth is the maximum depth of the decision tree; min_child_weight is the minimum sum of sample weights of leaf nodes; subsample is the proportion of random sampling for each tree; colsample_bytree is the proportion of columns randomly sampled for each tree; alpha is the weight coefficient of the L1 regularization term; and lambda is the weight coefficient of the L2 regularization term.

[0048] According to embodiments of this disclosure, in step S111, during the optimization of the hyperparameters of the prediction model using a validation set, for each set of preset hyperparameters, a prediction model is trained using a training set, and the prediction results of each prediction model are validated using a validation set. The performance of each prediction model is evaluated based on statistical parameters to determine the optimized hyperparameters (i.e., the optimal hyperparameters). Based on the optimized hyperparameters, the prediction model trained using the training set is the toxicity prediction model for the toxicity of interest. The statistical parameter used is the area under the curve (AUC) of the receiver operating characteristic curve (ROC curve), and the optimized hyperparameters are the parameter combinations that satisfy preset conditions, where satisfying the preset conditions represents the optimal situation.

[0049] According to an embodiment of this disclosure, in step S111, determining the toxicity judgment threshold of the toxicity prediction model includes: inputting the secondary mass spectrometry feature matrix of the validation set into the toxicity prediction model to obtain a set of toxicity prediction probability values ​​P. Based on the set of toxicity prediction probability values ​​P obtained by validating the toxicity prediction model using the validation set, and the binary label vectors for presence and absence of toxicity in the validation set, a subject performance characteristic curve is plotted. The specific steps for determining the toxicity judgment threshold are as follows: sorting the set P from largest to smallest, with each probability value serving as a binary classification threshold; calculating the true positive rate (TPR) and false positive rate (FPR) corresponding to each threshold based on the binary label vectors for presence and absence of toxicity in the validation set and the set of toxicity prediction probability values ​​P; plotting the subject performance characteristic curve with the false positive rate on the horizontal axis and the true positive rate on the vertical axis; and taking the threshold corresponding to the point where the geometric mean (G-mean) of sensitivity and specificity meets the preset conditions as the toxicity judgment threshold P. t The preset condition is the maximum case.

[0050] Among them, the toxicity determination threshold P t =argmax(G-mean) (1);

[0051]

[0052]

[0053]

[0054] TP: The number of positive samples predicted as positive by the toxicity prediction model;

[0055] TN: The number of negative samples predicted as negative by the toxicity prediction model;

[0056] FP: The number of negative samples predicted as positive by the toxicity prediction model;

[0057] FN: The number of positive samples predicted as negative by the toxicity prediction model.

[0058] According to an embodiment of this disclosure, in step S112, the generalization performance of the toxicity prediction model is evaluated using a test set, including: inputting the secondary mass spectrometry feature matrix of the test set into the toxicity prediction model, calculating statistical parameters based on the binary label vectors of presence and absence of toxicity in the test set and the set of toxicity prediction probability values ​​in the test set, and confirming the model performance based on the statistical parameters.

[0059] Figure 2This is a schematic diagram of the toxicity prediction method based on compound secondary mass spectrometry data in the embodiments of this disclosure.

[0060] According to embodiments of this disclosure, such as Figure 2 As shown, this disclosure also provides a toxicity prediction method based on compound secondary mass spectrometry data, including steps S201-S207.

[0061] Step S201: Obtain the secondary mass spectrometry data (i.e., secondary mass spectrometry data) of the compound to be predicted.

[0062] Step S202: Convert the secondary mass spectrometry data to be predicted into a molecular structure feature probability vector S.

[0063] Steps S203-S204: After converting the secondary mass spectrometry data to be predicted into a molecular structure feature probability vector S, it is input into the toxicity prediction model and outputs the toxicity prediction probability value p of the corresponding compound for the secondary mass spectrometry data to be predicted. The toxicity prediction model is obtained by training the toxicity prediction model based on the compound secondary mass spectrometry data established in the above embodiment.

[0064] Step S205: Determine whether p is greater than or equal to the toxicity determination threshold.

[0065] Step S206: If p is greater than or equal to the toxicity determination threshold, the compound corresponding to the secondary mass spectrometry data to be predicted is toxic.

[0066] Step S207: If p is less than the toxicity determination threshold, the compound corresponding to the secondary mass spectrometry data to be predicted is not toxic.

[0067] In the embodiments of this disclosure, the secondary mass spectrometry data to be predicted is input into the toxicity prediction model trained in the above embodiments for calculation, and the toxicity type of the compound corresponding to the secondary mass spectrometry data to be predicted is output. The toxicity prediction model provided in this disclosure is suitable for predicting the toxicity of secondary mass spectrometry data obtained from non-targeted analysis. The method is simple and fast, and has important application prospects for safety assessment of complex environments and food samples.

[0068] According to an embodiment of this disclosure, in step S201, the secondary mass spectrometry data to be predicted can be obtained from a publicly available secondary mass spectrometry database or from a high-resolution secondary mass spectrometer using a non-targeted analysis method. The compound corresponding to the secondary mass spectrometry data to be predicted is a known or unknown compound. The secondary mass spectrometry data to be predicted includes: the secondary mass spectrum data of the compound, the precise mass number and charge number of the precursor ion measured by the mass spectrometer, the ionization mode of the mass spectrometer, and the instrument type of the mass spectrometer.

[0069] According to embodiments of this disclosure, in steps S202-S204, the secondary mass spectrometry data to be predicted is input into a mass spectrometry calculation model (open-source software SIRIUS) for calculation, and the molecular structure feature probability vector S corresponding to the secondary mass spectrometry data to be predicted is output. The length of the molecular structure feature probability vector S is M. Then, the molecular structure feature probability vector S is used as input to a toxicity prediction model, and the model outputs the toxicity prediction probability value of the compound corresponding to the secondary mass spectrometry data to be predicted.

[0070] According to the embodiments of this disclosure, in steps S205-S207, the toxicity prediction probability value of the compound corresponding to the secondary mass spectrometry data to be predicted is compared with the toxicity determination threshold to determine whether the compound corresponding to the secondary mass spectrometry data to be predicted is toxic, thereby realizing the direct prediction of toxicity from secondary mass spectrometry data.

[0071] Figure 3 This is a schematic diagram illustrating the principle of predicting the toxicity of a compound corresponding to a given secondary mass spectrometry data based on the predicted secondary mass spectrometry data in an embodiment of this disclosure.

[0072] like Figure 3 As shown, secondary mass spectrometry data to be predicted are acquired. A mass spectrometry computational model is used to transform the secondary mass spectrometry data into a molecular structure feature probability vector S of length M. This vector S is then input into a toxicity prediction model (XGBoost toxicity prediction model) for prediction, outputting whether the compound corresponding to the secondary mass spectrometry data is toxic. This method can quickly and efficiently predict the toxicity of compounds corresponding to secondary mass spectrometry data, providing a basis for determining analytical priorities in non-targeted analysis. It is beneficial for rapidly determining the toxicity of pollutants in complex environmental samples and has broad application prospects in the fields of toxicity prediction of complex samples, environmental safety assessment, and health risk assessment.

[0073] The technical solutions of this disclosure are further illustrated below with reference to the embodiments and accompanying drawings, so as to provide a clearer understanding of the technical content of this disclosure.

[0074] Combination Figure 1 and Figure 2 This embodiment predicts the toxicity of compounds based on high-resolution secondary mass spectrometry data using a secondary mass spectrometry data conversion and XGBoost toxicity prediction model. Specifically, the toxicity of interest is aromatic hydrocarbon receptor activation activity. The steps include:

[0075] (1) Acquisition and preprocessing of high-resolution secondary mass spectrometry data and toxicity data of the compound, as detailed below:

[0076] All publicly available mass spectrometry data were downloaded from the GNPS website (https: / / gnps-external.ucsd.edu / gnpslibrary), which includes over 580,000 annotated and unannotated secondary mass spectrometry data submitted by various laboratories and mass spectrometry libraries. After obtaining the secondary mass spectrometry data, it was filtered to retain data containing linear molecular characterization information of the compound structure and sufficient secondary mass spectrum information. The filtered secondary mass spectrometry data includes linear molecular characterization information of the compound structure, secondary mass spectrum data of the compound, the precise mass number and charge number of the precursor ion measured by the mass spectrometer, the ionization mode of the mass spectrometer, and the instrument type of the mass spectrometer. The secondary mass spectrometry data can be in mgf format. The linear molecular characterization information of the compound structure can be in accordance with the Simplified Molecular Input Line Entry System (SMILES) or the International Chemical Identifier (InChI).

[0077] Download in vitro high-throughput test data on the activation activity of aryl hydrocarbon receptors (AhR) from the Tox21 challenge website (https: / / tripod.nih.gov / tox21 / challenge / data.jsp). The data includes linear molecular characterization information of the chemical structure of the compounds and corresponding binary category labels for the activation activity of aryl hydrocarbon receptors (the number "1" represents that the compound has activation activity, and the number "0" represents that the compound has no activation activity). A total of 8159 data entries were downloaded.

[0078] Chemical structure cleaning was performed on the compound information in the aforementioned secondary mass spectrometry (MS / MS) data and aromatic hydrocarbon receptor (ARH) test data. This cleaning included standardization, solvent removal, charge correction, and deionization to obtain standardized chemical structure data for the compounds. To ensure data reliability, based on the standardized chemical structures, data records with the same compound structure but different activities were removed from the ARH receptor activation activity dataset. The compounds in the two datasets (MS / MS and ARH receptor activation activity) were then matched according to their standardized chemical structures, resulting in 91,387 MS / MS data points and 91,387 ARH receptor activation activity tags corresponding to common compounds in both datasets.

[0079] For common compounds, secondary mass spectrometry (MS / MS) data were batch-input into SIRIUS software. For each MS / MS data set, the molecular structure feature probability vector S corresponding to the molecular formula with the highest SIRIUS Score was selected as the input for the subsequent prediction model. The length of the molecular structure feature probability vector is 4456, meaning it contains 4456 molecular structure features. The number on each molecular structure feature represents the probability that the compound corresponding to that MS / MS data set contains that molecular structure feature, and the numbers range from 0 to 1. During SIRIUS processing, there were issues such as the inability to calculate secondary MS / MS data with multiple charges, merging secondary MS / MS data sets with the same compound and the same exact mass number, and insufficient mass spectrometry information for calculation. Therefore, a total of 44942 molecular structure feature probability vectors and corresponding 44942 aromatic hydrocarbon receptor activation activity tags were obtained. This resulted in a total dataset for training, validation, and testing: a 44942×4456 secondary MS / MS feature matrix D and an aromatic hydrocarbon receptor activation activity tag vector T. The tag vector T has a length of 44942, and each element T in T... i ∈{0,1} indicates whether the compound corresponding to the i-th secondary mass spectrometry data has the toxicity of interest, with label '0' indicating non-toxicity and label '1' indicating toxicity.

[0080] (2) Training and hyperparameter optimization of the XGBoost toxicity prediction model

[0081] The total dataset, consisting of the 44942×4456 secondary mass spectrometry feature matrix D and the label vector T representing the activation activity of aromatic hydrocarbon receptors, was randomly divided into a training set, a validation set, and a test set using a stratified sampling method in a 6:2:2 ratio. The training set was used to train the prediction model, the validation set was used to optimize the hyperparameters of the prediction model, obtain the toxicity prediction model, and determine the toxicity judgment threshold, and the test set was used to evaluate the generalization performance of the toxicity prediction model.

[0082] The selected prediction model is the XGBoost prediction model based on gradient boosting decision trees. This model uses an additive model and a forward distribution algorithm. Its base model is a decision tree model. The fitting objective of each new decision tree is the negative gradient of the objective function of the previous tree. The objective function of the XGBoost prediction model is the loss function plus a regularization term. The final prediction result is the sum of all decision trees.

[0083] First, set the general hyperparameters and learning task hyperparameters of the XGBoost prediction model {booster: 'gbtree', objective: 'binary: logistic'}, that is, define the base learner type as a decision tree model, and the problem to be solved is binary classification.

[0084] Other hyperparameters of the XGBoost prediction model include {num_boost_round, learning_rate, max_depth, min_child_weight, gamma, subsample, colsample_bytree, lambda, alpha}, where num_boost_round is the number of iterations of the decision tree, learning_rate is the shrinkage step size during the update process, gamma is the minimum loss function decrease value required for node splitting, max_depth is the maximum depth of the decision tree, min_child_weight is the minimum sum of sample weights for leaf nodes, subsample is the proportion of random sampling for each tree, colsample_bytree is the proportion of columns randomly sampled for each tree, alpha is the weight coefficient of the L1 regularization term, and lambda is the weight coefficient of the L2 regularization term.

[0085] Secondly, an XGBoost prediction model was established. The XGBoost prediction model was trained using a training set, and the preset hyperparameters of the XGBoost prediction model were optimized using a validation set to obtain the XGBoost toxicity prediction model and determine the toxicity judgment threshold. During the hyperparameter optimization process, for each hyperparameter, several hyperparameters were selected within a certain range and with a certain step size. For each set of hyperparameters, an XGBoost prediction model was trained using the training set, and the XGBoost prediction model was validated using a validation set. The optimized hyperparameters (i.e., the optimal hyperparameters) were determined based on statistical parameters.

[0086] The selected statistical parameter was the area under the curve (AUC) of the receiver operating characteristic (ROC) curve. The XGBoost prediction model fitted with the optimal AUC value on the validation set was taken as the aryl hydrocarbon receptor activation activity prediction model (i.e., the XGBoost toxicity prediction model).

[0087] Hyperparameter optimization consists of four steps:

[0088] The first step is to select a relatively high learning rate (learning_rate), typically 0.1. Within a certain range and with a certain step size, obtain the AUC values ​​corresponding to different combinations of hyperparameters. Then, select the number of iterations for the ideal decision tree corresponding to this learning rate: num_boost_round = 200.

[0089] The second step is to select max_depth, min_child_weight, gamma, subsample, and colsample_bytree within a certain range and step size, given the learning rate and the number of iterations of the decision tree, and obtain the AUC values ​​corresponding to different combinations of hyperparameters. Then, select the optimized hyperparameters (i.e., the best hyperparameters) {max_depth: 11, min_child_weight: 2, gamma: 0, subsample: 0.9, colsample_bytree: 0.8}.

[0090] The third step is to tune the regularization hyperparameters of the XGBoost toxicity prediction model. Within a certain range and with a certain step size, lambda and alpha are selected to obtain the AUC values ​​corresponding to different combinations of hyperparameters. The optimized hyperparameters (i.e., the best hyperparameters) {lambda: 1, alpha: 1e-05} are selected.

[0091] The fourth step is to reduce the learning rate. Within a certain range and with a certain step size, select the learning rate and obtain the AUC values ​​corresponding to different hyperparameters. Select {learning rate: 0.1}.

[0092] The optimized hyperparameters (i.e. the final best hyperparameters) are {num boost round: 200, max_depth: 11, min_child_weight: 2, gamma: 0, subsample: 0.9, colsample_bytree: 0.8, learning_rate: 0.1}.

[0093] While determining the optimal hyperparameters (i.e., the best hyperparameters) and the toxicity prediction model, based on the set P of toxicity prediction probability values ​​of the validation set on the toxicity prediction model, and the binary label vectors of presence and absence of toxicity in the validation set, a subject performance characteristic curve is plotted, and a toxicity judgment threshold is determined. The specific steps are as follows: sort the set P from largest to smallest, and use each probability value as a binary classification threshold. Calculate the true positive rate (TPR) and false positive rate (FPR) corresponding to the threshold. Plot the subject performance characteristic curve with the false positive rate on the horizontal axis and the true positive rate on the vertical axis. Take the threshold P corresponding to the point on the curve with the largest geometric mean (G-mean) of sensitivity and specificity (i.e., meeting the preset conditions). t As the final threshold for determining the toxicity of compounds corresponding to the secondary mass spectrometry data, the point with the largest G-mean corresponds to the aromatic hydrocarbon receptor activation activity threshold P. tIt is 0.0642.

[0094] (3) The generalization performance evaluation of the aromatic hydrocarbon acceptor activation activity prediction model (XGBoost toxicity prediction model) corresponding to high-resolution secondary mass spectrometry is as follows:

[0095] The test set was input into the XGBoost toxicity prediction model. Based on the set of predicted toxicity probabilities and the binary label vectors indicating the presence or absence of toxicity in the test set, the AUC value of the ROC curve was calculated. The obtained AUC value was compared with the AUC value obtained by inputting the validation set data into the XGBoost toxicity prediction model. The results show that the AUC value of the XGBoost toxicity prediction model on the test set is comparable to that on the validation set, indicating that the toxicity prediction model has good generalization ability.

[0096] (4) Determination of the aromatic hydrocarbon acceptor activation activity of the compound corresponding to the high-resolution secondary mass spectrometry data of the compound to be evaluated.

[0097] Genistein (4',5,7-Trihydroxyisoflavone) is an isoflavone phytoestrogen with antioxidant properties, and its natural sources include legumes, sandalwood, and banyan trees. Studies have shown that genistein can inhibit protein-tyrosine kinase and DNA topoisomerase-II activity, and it is currently being investigated in clinical trials for cancer treatment. However, it poses both acute and persistent hazards to aquatic environments. Therefore, genistein was selected as a test subject.

[0098] High-resolution secondary mass spectrometry (MS / MS) data of genistein (4',5,7-Trihydroxyisoflavone) were obtained from a publicly available MS / MS library. This MS / MS data file was submitted to SIRIUS software. After calculation, the molecular structure feature vector S, along with its probability value, corresponding to the molecular formula with the highest SIRIUS Score, was identified in the results, as shown in Table 1. This molecular structure feature probability vector S was input into the XGBoost toxicity prediction model for prediction. The predicted probability value of the aryl hydrocarbon receptor activation activity of the compound corresponding to this MS / MS data was 0.990. Since the predicted probability value is greater than the aryl hydrocarbon receptor activation activity threshold of 0.0642, the compound corresponding to this MS / MS data is considered to have aryl hydrocarbon receptor activation activity, i.e., the compound corresponding to this MS / MS data possesses this type of toxicity.

[0099] Table 1. The molecular structure feature probability vector S of length M after transformation of secondary mass spectrometry data by the mass spectrometry calculation model in the embodiments of this disclosure.

[0100] Table 1

[0101]

[0102] The compound was found to be active in the Tox21 project database for aromatic hydrocarbon receptor activation, and the prediction results of the XGBoost toxicity prediction model were consistent with the results.

[0103] The specific embodiments described above further illustrate the purpose, technical solutions, and beneficial effects of this disclosure. It should be understood that the above descriptions are merely specific embodiments of this disclosure and are not intended to limit this disclosure. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the protection scope of this disclosure.

Claims

1. A method for establishing a toxicity prediction model based on secondary mass spectrometry data of compounds, comprising: Obtain secondary mass spectrometry data of known compounds and toxicity data of known chemicals; Chemical structure cleaning is performed on the secondary mass spectrometry data of the known compounds and the compounds involved in the toxicity data of the known chemicals to obtain the standardized chemical structures of the compounds. Based on the standardized chemical structures, the secondary mass spectrometry data and the compounds involved in the toxicity data are matched to obtain the secondary mass spectrometry data of common compounds and the binary label of toxicity of interest for the corresponding common compounds. For the common compound, each of the secondary mass spectrometry data is converted into a molecular structure feature probability vector S, and a total dataset containing the molecular structure feature probability vector S and a binary label of toxicity is established. The total dataset is divided into a training set, a validation set, and a test set. A toxicity prediction model is constructed, taking the molecular structure feature probability vector S as input and the presence or absence of toxicity as output, including: Based on multiple sets of preset hyperparameters of the prediction model used, the prediction model is trained using the training set, and the multiple sets of preset hyperparameters of the prediction model are optimized using the validation set to obtain a toxicity prediction model for the toxicity of interest and determine the toxicity judgment threshold. The generalization performance of the toxicity prediction model was evaluated using the test set.

2. The method according to claim 1, wherein, The secondary mass spectrometry data of the known compounds are the annotated secondary mass spectrometry data of the corresponding compounds. The secondary mass spectrometry data of the known compounds include linear molecular characterization information of the compound structure, secondary mass spectrum data of the compound, precise mass number and charge number of the precursor ion measured by the mass spectrometer, ionization mode of the mass spectrometer, and instrument type of the mass spectrometer. The toxicity data for the known chemicals include linear molecular characterization information of the compound structure and binary labeling of whether the compound is toxic or not.

3. The method according to claim 1, wherein, The chemical structure cleaning includes: standardization, solvent removal, charge correction, and deionization. Before matching the compounds involved in the toxicity data and secondary mass spectrometry data according to their standardized chemical structures, the following steps are included: Based on the standardized chemical structure of the compounds, compounds with opposite toxicity labels are removed from the toxicity data.

4. The method according to claim 1, wherein, The step of converting each of the secondary mass spectrometry data into a molecular structure feature probability vector S, and establishing a total dataset containing the molecular structure feature probability vector S and a binary label indicating whether or not it is toxic, includes: Each of the secondary mass spectrometry data is input into the mass spectrometry calculation model for calculation, and the molecular structure feature probability vector S is output. The length of the molecular structure feature probability vector S is M. The N molecular structure feature probability vectors S obtained by converting the secondary mass spectrometry data constitute the secondary mass spectrometry feature matrix D of the common compound. The total dataset is constructed based on the secondary mass spectrometry feature matrix D of the common compounds and the binary label vector T for the toxicity of the common compounds of interest.

5. The method according to claim 4, wherein, The size of the secondary mass spectrometry feature matrix D is N×M, where N is the number of molecular structure feature probability vectors S calculated from the secondary mass spectrometry data of the common compounds, and each element D in D... i,j This represents the probability that the common compound corresponding to the probability vector S of the molecular structure features calculated from the i-th secondary mass spectrometry data contains a specific molecular structure feature j; The length of the binary label vector T for the toxicity of the shared compound of interest is N, and each element T in T is... i ∈{0,1} indicates whether the common compound corresponding to the molecular structure feature probability vector S calculated from the i-th secondary mass spectrometry data has the toxicity of interest. The label "0" indicates non-toxicity, and the label "1" indicates toxicity.

6. The method according to claim 1, wherein, The preset hyperparameters of the prediction model are the externally set hyperparameters of the prediction model itself. The method of optimizing multiple preset hyperparameters of the prediction model using the validation set to obtain a toxicity prediction model for the toxicity of interest and determining the toxicity determination threshold includes: For each set of preset hyperparameters, a prediction model is trained using the training set, the prediction results of each prediction model are verified using the validation set, and the optimized hyperparameters are determined based on statistical parameters. Based on the optimized hyperparameters, the prediction model trained using the training set is used as the toxicity prediction model for the toxicity of interest. Wherein, the statistical parameter is the area under the subject's operating characteristic curve, and the optimized hyperparameter is the hyperparameter when the statistical parameter satisfies preset conditions.

7. The method according to claim 6, wherein, The determination of the toxicity threshold includes: The secondary mass spectrometry feature matrix in the validation set is input into the toxicity prediction model to obtain the toxicity prediction probability value set P; The set of toxicity prediction probability values ​​P is sorted from largest to smallest, and each toxicity prediction probability value is used as a binary classification threshold. Based on the binary label vector of presence or absence of toxicity in the validation set and the set of toxicity prediction probability values ​​P, the true positive rate and false positive rate corresponding to each binary classification threshold are calculated. Plot the subject performance characteristic curve with the false positive rate on the horizontal axis and the true positive rate on the vertical axis. Take the threshold corresponding to the point where the geometric mean of sensitivity and specificity meets the preset conditions as the toxicity determination threshold P. t ; Wherein, the toxicity determination threshold P t =argmax(G-mean); TP: The number of positive samples predicted as positive by the toxicity prediction model; TN: The number of negative samples predicted as negative by the toxicity prediction model; FP: The number of negative samples predicted as positive by the toxicity prediction model; FN: The number of positive samples predicted as negative by the toxicity prediction model; TPR: True positive rate; FPR: False positive rate; sensitivity: sensitivity value; specificity: specificity value; G-mean: Geometric mean of sensitivity and specificity.

8. The method according to claim 6, wherein, The generalization performance evaluation of the toxicity prediction model using the test set includes: The secondary mass spectrometry feature matrix of the test set is input into the toxicity prediction model to obtain the toxicity prediction probability value set of the test set. Based on the binary label vector of presence or absence of toxicity in the test set and the toxicity prediction probability value set of the test set, the statistical parameters are calculated, and the model performance is confirmed based on the statistical parameters.

9. A toxicity prediction method based on secondary mass spectrometry data of compounds, comprising: Obtain secondary mass spectrometry data of the compound to be predicted; After the secondary mass spectrometry data to be predicted is converted into a molecular structure feature probability vector S, it is input into the toxicity prediction model and outputs the toxicity prediction probability value p of the compound corresponding to the secondary mass spectrometry data to be predicted. The toxicity prediction model is trained by the method of any one of claims 1-8. When p is greater than or equal to the toxicity determination threshold, the compound corresponding to the secondary mass spectrometry data to be predicted is toxic; when p is less than the toxicity determination threshold, the compound corresponding to the secondary mass spectrometry data to be predicted is not toxic.

10. The method according to claim 9, wherein, The compounds corresponding to the secondary mass spectrometry data to be predicted are known or unknown compounds; The secondary mass spectrometry data to be predicted includes the secondary mass spectrum data of the compound, the precise mass number and charge number of the precursor ion measured by the mass spectrometer, the ionization mode of the mass spectrometer, and the instrument type of the mass spectrometer.

Citation Information

Patent Citations

  • Modeling method and device of compound toxicity prediction model and application of compound toxicity prediction model

    CN110890137A

  • Method for identifying androgen active substances in environmental sample based on characteristic liquid fragment matching

    CN111413444A