Bitterness compound prediction method based on machine learning
Through machine learning-based methods, molecular fingerprint and LightGBM algorithms are used to achieve rapid and accurate prediction of the bitter taste characteristics of the compounds, solving the problems of time-consuming and low accuracy of traditional screening methods, and improving the efficiency and accuracy of drug research and development.
Patent Information
- Application Number
- CN202510319840.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-01
AI Technical Summary
Traditional bitter compound screening methods take a long time and have low accuracy, making it difficult to meet the needs of modern drug research and development to efficiently and accurately identify bitter compounds.
Using machine learning-based bitter compound prediction method, samples are obtained from the compound database, converted to SMILES format, and converted them into different types of molecular fingerprints for integration, mutual information is calculated, and characteristics are sorted. The LightGBM algorithm is used to establish a prediction model to quickly predict the bitter properties of compounds.
It improves the screening efficiency of bitter drugs, accurately locates the compound structure with therapeutic potential, saves the time and cost of drug development, and significantly reduces the development cost of bitter drugs.
Smart Images

Figure CN120236684A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of compound prediction, and particularly relates to a method for predicting bitter compounds based on machine learning. Background Art
[0002] The theory that "good medicine tastes bitter and is beneficial to the disease" is widely recognized in traditional medicine and modern pharmacology. Bitterness is often closely related to the effectiveness of drugs. Many bitter compounds, such as alkaloids, flavonoids, and glycosides, exhibit antioxidant, anti-inflammatory, antibacterial, and immunomodulatory effects and are used in the prevention and treatment of various diseases. During the drug development process, the bitter property often indicates the pharmacological activity of a compound, which can provide guidance in new drug screening and efficacy evaluation, narrow the range of potential drug selections, and save R & D costs. However, traditional methods for screening bitter compounds have problems such as long time consumption and strong subjectivity, and the strong subjectivity further leads to low screening accuracy, making it difficult to meet the requirements of modern drug R & D for efficient and accurate identification of bitter compounds. The method for predicting bitter compounds based on machine learning can achieve rapid prediction of the bitter properties of new compounds through the learning and analysis of known bitter compounds, thereby improving the screening efficiency of bitter drugs, providing scientific support for drug R & D teams to accurately locate the compound structures with therapeutic potential, and facilitating the further application and development of bitter drugs in disease treatment. Summary of the Invention
[0003] The purpose of the present invention is to solve the problems of long time consumption and low accuracy existing in traditional methods for screening bitter compounds, and a method for predicting bitter compounds based on machine learning is proposed.
[0004] The technical solution adopted by the present invention to solve the above technical problems is: A method for predicting bitter compounds based on machine learning, the method specifically includes the following steps:
[0005] Step S1: Respectively obtain bitter compound samples and sweet compound samples from a compound database, and divide all the obtained compound samples into two parts: a training set and a test set;
[0006] Step S2: Respectively obtain the SMILES formats of each compound sample in the training set and the test set;
[0007] For the SMILES format of any compound sample, respectively convert the SMILES format of this compound sample into different types of molecular fingerprints, and then integrate the different types of molecular fingerprints of this compound sample to obtain the molecular fingerprint combination corresponding to this compound sample;
[0008] Similarly, respectively obtain the molecular fingerprint combinations corresponding to each compound sample;
[0009] Step S3. For any molecular fingerprint combination, calculate the mutual information between each feature and the compound label under this molecular fingerprint combination according to the feature vectors of various compounds under this molecular fingerprint combination, and then sort the calculated mutual information in descending order to obtain the sorted mutual information; obtain the feature sorting result corresponding to this molecular fingerprint combination according to the sorted mutual information.
[0010] Similarly, obtain the feature sorting results corresponding to each molecular fingerprint combination respectively.
[0011] Step S4. For any compound, sort the feature values in the feature vector of this compound according to the feature sorting results of each molecular fingerprint combination respectively to obtain the sorted feature vector of this compound under each molecular fingerprint combination.
[0012] Similarly, obtain the sorted feature vectors of each compound under each molecular fingerprint combination respectively.
[0013] Step S5. Train the LightGBM algorithm according to the sorted feature vectors of each compound in the training set under each molecular fingerprint combination to obtain the optimal number of features corresponding to each molecular fingerprint combination.
[0014] Step S6. According to the optimal number of features of each molecular fingerprint combination obtained in Step S5, each compound in the test set performs feature screening on the sorted feature vectors under each molecular fingerprint combination, and use the feature vector composed of the screened features as the input of the LightGBM algorithm. Output the prediction accuracy of the test set samples under each molecular fingerprint combination through the LightGBM algorithm. Take the molecular fingerprint combination corresponding to the highest prediction accuracy obtained by the test set samples as the optimal molecular fingerprint combination a, and record the optimal number of features corresponding to the molecular fingerprint combination a as A.
[0015] Step S7. Obtain the SMILES format of the compound to be predicted, and then obtain the molecular fingerprint combination a of the compound to be predicted. After sorting the features in the feature vector of the compound to be predicted under the molecular fingerprint combination a according to the mutual information sorting, select the first A features after sorting, and use the feature vector composed of the selected first A features as the input of the LightGBM algorithm. Obtain the bitter taste prediction result of the compound to be predicted according to the output of the LightGBM algorithm.
[0016] The beneficial effects of the present invention are:
[0017] The present invention converts each compound sample into the SMILES format, then converts the SMILES format into different types of molecular fingerprints, and integrates different molecular fingerprints to obtain 15 molecular fingerprint combinations. The minimum redundancy maximum correlation algorithm is used to obtain the mutual information between each feature of each molecular fingerprint combination and the compound sample label. All the mutual information of each molecular fingerprint combination of each compound sample is sorted according to the mutual information value from large to small, and the molecular fingerprint combination features corresponding to the sorted mutual information of each compound sample are obtained, so that the chemical structure information of the compound sample can be characterized by the feature vector. Based on the LightGBM algorithm, a bitter compound prediction model is established. The molecular fingerprint combination features corresponding to the sorted mutual information of each compound sample are sequentially input into the bitter compound prediction model by using the feature incremental selection method, and the prediction accuracy of each molecular fingerprint combination is calculated. According to the features of each molecular fingerprint combination and its prediction accuracy, a feature incremental selection curve graph of the corresponding molecular fingerprint combination is generated, and all the features before the feature with the highest accuracy in the feature incremental selection curve graph are selected to obtain the number A of features corresponding to the molecular fingerprint combination, and then the best molecular fingerprint combination and the best number of features are obtained.
[0018] The present invention combines molecular fingerprint features with machine learning algorithms. Molecular fingerprints can reflect information on molecular structure, properties, and functional groups. By integrating multiple molecular fingerprint types of the compound to be tested and extracting its key features, and using the machine learning algorithm LightGBM for bitterness prediction, it can more quickly and accurately predict the bitterness characteristics of small molecule compounds, helping to screen out bitter compounds with therapeutic potential in the early stage of drug research and development, saving a large amount of time and cost for the development of bitter drugs, improving the success rate of drug lead screening, accelerating the screening and testing process of candidate drugs, significantly reducing the development cost of bitter drugs, and solving the problems of the small number of current bitter drugs, low screening efficiency, low screening accuracy, and high development cost. Brief Description of the Drawings
[0019] Figure 1 is the flow chart of the method of the present invention;
[0020] Figure 2 is the schematic diagram of the LightGBM algorithm;
[0021] Figure 3 is the feature incremental selection curve graph of the E+F molecular fingerprint combination of the training set;
[0022] EF training feature number represents different numbers of features under the E+F molecular fingerprint combination. Detailed Embodiments
[0023] Embodiment 1: A method for predicting bitter compounds based on machine learning. The method specifically includes the following steps:
[0024] Step S1: Obtain bitter compound samples and sweet compound samples from the compound database respectively, and divide all the obtained compound samples into two parts: a training set and a test set.
[0025] Step S2: Obtain the SMILES format of each compound sample in the training set and the test set respectively.
[0026] For the SMILES format of any compound sample, convert the SMILES format of this compound sample into different types of molecular fingerprints respectively, and then integrate the different types of molecular fingerprints of this compound sample to obtain the molecular fingerprint combination corresponding to this compound sample.
[0027] Similarly, obtain the molecular fingerprint combination corresponding to each compound sample respectively.
[0028] Step S3: For any molecular fingerprint combination, calculate the mutual information between each feature in this molecular fingerprint combination and the compound label (the compound sample label is whether the compound is bitter) according to the feature vectors of various compounds under this molecular fingerprint combination, and then sort the calculated mutual information in descending order to obtain the sorted mutual information; obtain the feature sorting result corresponding to this molecular fingerprint combination according to the sorted mutual information (that is, the feature corresponding to the largest mutual information is ranked first in the feature sorting result, the feature corresponding to the second largest mutual information is ranked second in the feature sorting result, and so on);
[0029] Similarly, obtain the feature sorting result corresponding to each molecular fingerprint combination respectively.
[0030] Step S4: For any compound, sort the feature values in the feature vector of this compound according to the feature sorting result of each molecular fingerprint combination respectively (that is, re - sort the feature values in the feature vector of this compound according to the feature sorting result obtained in step S3. For example, for the feature ranked first in the feature sorting result, the value of this feature is also ranked first in the sorted feature vector, and so on), to obtain the sorted feature vector of this compound under each molecular fingerprint combination.
[0031] Similarly, obtain the sorted feature vectors of each compound under each molecular fingerprint combination respectively.
[0032] Step S5: Train the LightGBM algorithm according to the sorted feature vectors of each compound in the training set under each molecular fingerprint combination to obtain the optimal number of features corresponding to each molecular fingerprint combination.
[0033] Step S6: According to the optimal number of features for each molecular fingerprint combination obtained in Step S5, for each compound in the test set, feature screening is performed on the sorted feature vectors under each molecular fingerprint combination (for any molecular fingerprint combination, if the optimal number of features for this molecular fingerprint combination is n, then for any compound sample in the test set, the first n features of this test sample under this molecular fingerprint combination are selected to obtain the updated feature vector of this compound sample. Then, the updated feature vectors of all compound samples in the test set are used as the input of the LightGBM algorithm). The feature vectors composed of the screened features are used as the input of the LightGBM algorithm. Through the LightGBM algorithm, the prediction accuracy of the test set samples under each molecular fingerprint combination is output. The molecular fingerprint combination corresponding to the highest prediction accuracy obtained by the test set samples is used as the optimal molecular fingerprint combination a, and the optimal number of features corresponding to the molecular fingerprint combination a is denoted as A;
[0034] Step S7: Obtain the SMILES format of the compound to be predicted, and then obtain the molecular fingerprint combination a of the compound to be predicted. After sorting the features in the feature vector of the compound to be predicted under the molecular fingerprint combination a according to the mutual information, the first A features after sorting are selected, and the feature vector composed of the selected first A features is used as the input of the LightGBM algorithm. According to the output of the LightGBM algorithm, the bitter taste prediction result of the compound to be predicted is obtained.
[0035] Specific Embodiment 2: The difference between this embodiment and Specific Embodiment 1 is that the bitter compound samples are obtained from the BitterDB database, PubMed database, FoodDB database, FlavorDB2 database, FSBI-DB database, ChemTastesDB database, and The Good Scents Company database;
[0036] The sweet compound samples are obtained from the SweetenersDB database, SuperSweet database, PubMed database, FoodDB database, FlavorDB2 database, FSBI-DB database, ChemTastesDB database, and The Good Scents Company database.
[0037] Other steps and parameters are the same as those in Specific Embodiment 1.
[0038] The present invention obtains compound samples with a bitter taste and compound samples without a bitter taste from existing major databases, more comprehensively covering all compound sample information, providing a high-quality dataset of bitter compounds, and ensuring the accuracy and reliability of the results of the present invention.
[0039] Specific Embodiment 3: The difference between this embodiment and Specific Embodiment 1 or 2 is that the specific process of step S2 is as follows:
[0040] Step S21: Call the pubchempy package in the pandas library to obtain the SMILES format of each compound sample;
[0041] Step S22: Call the RDkit package in the pandas library to convert the SMILES format of each compound sample into different types of molecular fingerprints;
[0042] Step S23: For any compound sample obtained, integrate the molecular fingerprints corresponding to the SMILES format of this compound sample to obtain the molecular fingerprint combination corresponding to this compound sample;
[0043] Similarly, obtain the molecular fingerprint combination corresponding to each compound sample respectively.
[0044] Other steps and parameters are the same as those in Specific Embodiment 1 or 2.
[0045] Specific Embodiment 4: The difference between this embodiment and any one of Specific Embodiments 1 to 3 is that the types of the molecular fingerprints include ECFP4 fingerprint, FCFP4 fingerprint, RDKit fingerprint, and MACCS fingerprint.
[0046] Other steps and parameters are the same as those in any one of Specific Embodiments 1 to 3.
[0047] Specific Embodiment 5: The difference between this embodiment and any one of Specific Embodiments 1 to 4 is that the specific process of step S23 is as follows:
[0048] For each compound sample, 15 molecular fingerprint combinations are obtained, and the 15 molecular fingerprint combinations are respectively:
[0049] ECFP4 fingerprint, FCFP4 fingerprint, RDKit fingerprint, MACCS fingerprint;
[0050] ECFP4 fingerprint + FCFP4 fingerprint, ECFP4 fingerprint + RDKit fingerprint, ECFP4 fingerprint + MACCS fingerprint, FCFP4 fingerprint + RDKit fingerprint, FCFP4 fingerprint + MACCS fingerprint, RDKit fingerprint + MACCS fingerprint;
[0051] ECFP4 fingerprint + FCFP4 fingerprint + RDKit fingerprint, ECFP4 fingerprint + FCFP4 fingerprint + MACCS fingerprint, ECFP4 fingerprint + RDKit fingerprint + MACCS fingerprint, FCFP4 fingerprint + RDKit fingerprint + MACCS fingerprint;
[0052] ECFP4 fingerprint + FCFP4 fingerprint + RDKit fingerprint + MACCS fingerprint.
[0053] Other steps and parameters are the same as those in any one of the first to fourth specific embodiments.
[0054] Specific embodiment six: The difference between this embodiment and any one of the first to fifth specific embodiments is that the mutual information is calculated using the minimum redundancy maximum correlation algorithm.
[0055] Other steps and parameters are the same as those in any one of the first to fifth specific embodiments.
[0056] Specific embodiment seven: The difference between this embodiment and any one of the first to sixth specific embodiments is that the specific process of step S5 is as follows:
[0057] Taking any one molecular fingerprint combination as an example
[0058] Step S51: Divide all compound samples in the training set into 5 subsets, and sequentially number each subset as subset 1, subset 2, subset 3, subset 4, and subset 5;
[0059] Initialize the number of features n = 1;
[0060] Step S52: Initialize j = 1;
[0061] Step S53: Use the j-th subset as the validation sample and the remaining 4 subsets as the training samples;
[0062] Step S54: Select the first n features from the sorted feature vectors of each training sample under the current molecular fingerprint combination, and use the selected features to form a new feature vector for each training sample;
[0063] Step S55: Use the new feature vector of each training sample as the input of the LightGBM algorithm to obtain the decision tree constructed by the LightGBM algorithm;
[0064] Step S56: Select the first n features from the sorted feature vectors of each validation sample in the j-th subset under the current molecular fingerprint combination, and use the selected features to form a new feature vector for each validation sample;
[0065] Use the new feature vectors of each validation sample as the input to the decision tree constructed in step S55, output the prediction results for each validation sample through the decision tree (obtain the final prediction result by synthesizing the prediction results of 100 trees), and calculate the prediction accuracy rate of the samples within the j-th subset;
[0066] Step S57, determine whether j is equal to 5;
[0067] If j is equal to 5, take the average of the prediction accuracy rates when each subset is used as a validation sample, and use the averaged result as the prediction accuracy rate corresponding to the number of features n; and continue to execute step S58;
[0068] If j is less than 5, let j = j + 1, and return to execute step S53;
[0069] Step S58, determine whether the number of features n is equal to N, where N represents the total number of features in the feature vectors under the current molecular fingerprint combination;
[0070] If n is equal to N, obtain the number of features n' corresponding to the highest prediction accuracy rate under the current molecular fingerprint combination, and use the number of features n' as the optimal number of features under the current molecular fingerprint combination;
[0071] If n is less than N, let n = n + 1, and return to execute step S52.
[0072] Other steps and parameters are the same as those in any one of the specific embodiments one to six.
[0073] Specific embodiment eight: The difference between this embodiment and any one of the specific embodiments one to seven is that the working process of the LightGBM algorithm is as follows:
[0074] Step 1, initialize the predicted values of all samples to be F0(x i );
[0075]
[0076] Among them, In the present invention, the base of the logarithm can be 10;
[0077] Step 2, initialize the number of regression trees t = 1;
[0078] Step 3, initialize the node level l of the regression tree to 1; in the present invention, the root node is used as the first level, and other levels are obtained by node splitting subsequently;
[0079] Step 4, respectively determine whether there is a node with a sample number greater than or equal to the threshold (set to 20 in the present invention) on the l-th level of the t-th regression tree;
[0080] If there are nodes with the number of samples greater than or equal to the threshold, mark each node with the number of samples less than the threshold as a leaf node, and perform step 5 on the nodes with the number of samples greater than or equal to the threshold;
[0081] If there are no nodes with the number of samples greater than or equal to the threshold, mark all nodes at the l-th level as leaf nodes, and then perform step 7;
[0082] Step 5: Each node calculates its own maximum split gain, and marks the nodes with the maximum split gain less than the threshold as leaf nodes. For each node with the maximum split gain greater than or equal to the threshold, select the split point that maximizes its own split gain for splitting (for any node, according to the value of each feature in the feature vector of the compound, the model will divide the compounds in the current node into two parts (the specific division method is: the value corresponding to each feature in the feature vector can only be 0 or 1. When allocating according to the value of a feature, allocate the compounds with the value of 0 of this feature to the left child node, and allocate the compounds with the value of 1 of this feature to the right child node). According to the division result, calculate the split gain for splitting according to the value of the current feature, and finally obtain the feature corresponding to the maximum split gain. This step splits according to this feature), and allocate the samples in itself to the left child node and the right child node;
[0083] Among them, the calculation method of the split gain is:
[0084]
[0085] Among them, G represents the split gain, I represents the set of all compound samples in the current node to be split, I L represents the set of compound samples in the left child node after splitting, I R represents the set of compound samples in the right child node after splitting, g i (t) represents the first-order gradient of the loss function of the i-th compound sample (represents the error gradient of the i-th sample), h i (t) represents the second-order gradient of the loss function of the i-th compound sample (represents the curvature information of the i-th sample), λ is the regularization term coefficient (used to control the model complexity, smooth the calculation, and prevent overfitting), and γ is the penalty term coefficient (the minimum gain threshold for node splitting, controlling the model from performing meaningless splitting);
[0086]
[0087] Among them, σ(·) represents the sigmoid activation function, F t-1 (x i ) represents the prediction result of the (t - 1)-th regression tree for the sample x i of, yi Denote the sample as x i with the true label;
[0088] Step 6: Let the node level l = l + 1, and return to execute Step 4;
[0089] Step 7: The LightGBM algorithm continues to split until the number of samples in all nodes is too small or the gain is too low. At this time, no more nodes can be split, and all nodes are marked as leaf nodes;
[0090] Calculate the output of each leaf node respectively, and denote the output of the j-th leaf node as ω j (t) ;
[0091] When the compound sample x i is in the j-th leaf node of the t-th regression tree, then let f t (x i ) = ω j (t) , and the prediction result of the t-th regression tree for the sample x i is F t (x i ):
[0092] F t (x i ) = F t-1 (x i ) + γ t f t (x i )
[0093] where γ t is the learning rate of the t-th regression tree (controlling the update amplitude of this iteration);
[0094] Step 8: Determine whether t = M is satisfied, where M is the total number of regression trees set;
[0095] If not satisfied, then let t = t + 1, and return to execute Step 3;
[0096] If satisfied, then end and obtain T regression trees.
[0097] Other steps and parameters are the same as those in any one of the specific embodiments one to seven.
[0098] Specific Embodiment Nine: Different from any one of the specific embodiments one to eight, the new feature vectors of each validation sample are used as the input of the decision tree constructed in Step S55, and the prediction results for each validation sample are output through the decision tree; the specific process is as follows:
[0099] For the validation sample x:
[0100]
[0101] Among them, F(x) represents the prediction result of the verification sample x; M is the number of decision trees; γ t is the learning rate; f t (x) is the output of the leaf node where the verification sample x is located in the t-th tree; F0(x) = F0(x i );
[0102] Calculate the probability p' that the verification sample x is a bitter agent according to F(x):
[0103]
[0104] Among them, e represents the base of the natural logarithm;
[0105] If p' > 0.5, then the verification sample x is a bitter agent, otherwise, the verification sample x is a non-bitter agent.
[0106] Other steps and parameters are the same as those in any one of the specific embodiments one to eight.
[0107] According to this specific embodiment, the probability that each verification sample is bitter can be calculated respectively, and then the prediction accuracy rate of the overall verification sample can be obtained.
[0108] Specific embodiment ten: The difference between this specific embodiment and any one of the specific embodiments one to nine is that the output ω of the j-th leaf node j (t) is:
[0109]
[0110] Among them, I j represents the set of compound samples included in the j-th leaf node in the t-th regression tree.
[0111] Other steps and parameters are the same as those in any one of the specific embodiments one to nine.
[0112] Example
[0113] Taking the Windows system as the development environment, Jupyter Notebook as the development tool, and using Python as the development language, predict whether the taste of small molecule compounds in the drug R & D stage is bitter. As Figure 1 shown, the prediction method specifically includes the following steps:
[0114] Step S1, obtain compound samples from the compound database, and divide the compound samples into a training set and a test set.
[0115] The BitterDB database is a bioinformatics database provided by the Israel Institute of Technology that is specifically used to collect and study bitter compounds. It integrates rich bitter compound data and their corresponding taste receptor information. The present invention obtains 1041 major bitter compound samples from the BitterDB database, and obtains the remaining bitter compound samples from PubMed, FoodDB, FlavorDB2, FSBI-DB, ChemTastesDB and The Good Scents Company databases. The bitter compound samples from these sources are used as positive samples (i.e., positive samples), and a total of 1581 positive samples are obtained to represent the bitter characteristics of the compounds.
[0116] SweetenersDB and SuperSweet databases are bioinformatics databases of sweet compounds, which summarize a large amount of information such as the chemical structure, molecular properties and physiological effects of natural and artificial sweeteners. SweetenersDB mainly includes various types of sweeteners, including sugars, sugar alcohols, peptides, etc., which are suitable for food science and drug development research; while the SuperSweet database covers a wealth of information on artificial sweeteners and natural sweeteners, not only including molecular structure and physicochemical properties, but also providing detailed data on sweetness intensity and taste receptor action mechanism. Therefore, the present invention obtains the main sweet compound samples from SweetenersDB and SuperSweet databases, and obtains other sweet compound samples from PubMed, FoodDB, FlavorDB2, FSBI-DB, ChemTastesDB and The Good Scents Company databases, and obtains a total of 2294 sweet compound samples as negative samples (i.e., negative samples) to characterize compounds that are not bitter. Through the comprehensive data sources of these seven databases, the present invention constructs a more comprehensive sample set of bitter and sweet compounds, which helps to improve the accuracy and generalization ability of the bitterness prediction model.
[0117] In order to ensure the training effect and prediction performance of the model, the present invention divides 1581 bitter substances and 2294 sweet substances into training sets and test sets in a ratio of 8:2, thereby establishing a more comprehensive bitterness prediction model to improve the accuracy and generalization ability of the model. The specific distribution number of training samples is shown in Table 1.
[0118] Table 1
[0119]
[0120]
[0121] Step S2: Obtain the SMILES format of each compound sample in the training set respectively, and then obtain the ECFP4 fingerprint, FCFP4 fingerprint, RDKit fingerprint, and MACCS fingerprint corresponding to each SMILES format respectively.
[0122] Integrate the four molecular fingerprints corresponding to the same SMILES format to obtain multiple molecular fingerprint combinations corresponding to this SMILES format. Similarly, obtain multiple molecular fingerprint combinations corresponding to each compound sample in the training set.
[0123] Simplified molecular input line entry system (SMILES) is a specification that clearly describes the molecular structure using ASCII strings. In this invention, the read_csv function in the pandas library is called to read the names of each compound sample in the training set and store the names in a list. According to the order of the names, the pubchempy package in the pandas library is called to obtain the SMILES format of the corresponding compound sample, and the name of each compound sample is stored together with the corresponding SMILES format. Since machine learning algorithms can only recognize input numerical features, the structural information of bitter compounds needs to be converted into numerical vector features through an effective feature extraction method, that is, the RDkit package in the pandas library is called to convert the SMILES format of each compound sample into different molecular fingerprints, including ECFP4 fingerprint, FCFP4 fingerprint, RDKit fingerprint, and MACCS fingerprint. The above fingerprints are integrated and stored in a csv file. The integration methods include ECFP4 fingerprint, FCFP4 fingerprint, RDKit fingerprint, MACCS fingerprint, ECFP4 fingerprint + FCFP4 fingerprint, ECFP4 fingerprint + RDKit fingerprint, ECFP4 fingerprint + MACCS fingerprint, FCFP4 fingerprint + RDKit fingerprint, FCFP4 fingerprint + MACCS fingerprint, RDKit fingerprint + MACCS fingerprint, ECFP4 fingerprint + FCFP4 fingerprint + RDKit fingerprint, ECFP4 fingerprint + FCFP4 fingerprint + MACCS fingerprint, ECFP4 fingerprint + RDKit fingerprint + MACCS fingerprint, FCFP4 fingerprint + RDKit fingerprint + MACCS fingerprint, ECFP4 fingerprint + FCFP4 fingerprint + RDKit fingerprint + MACCS fingerprint, a total of 15 kinds of molecular fingerprint combinations.
[0124] Thus, 15 molecular fingerprint combinations corresponding to each compound sample in the training set are obtained respectively. Molecular fingerprints can reflect information on molecular structure, properties, and functional groups. Combined with machine learning methods, they can predict candidate compounds and active compounds faster and more accurately.
[0125] In step S3, during the feature extraction process, feature vectors of several thousand dimensions or even over ten thousand dimensions are often obtained. In the present invention, the combined feature dimension of ECFP4 fingerprint + FCFP4 fingerprint + RDKit fingerprint + MACCS fingerprint is as high as 5287. Some redundant or irrelevant features are mixed in these vectors. The huge dimension will cause the "curse of dimensionality", which will not only make the model training very slow, but also easily cause overfitting results. Therefore, for each compound sample in the training set with multiple molecular fingerprint combinations, the minimum redundancy maximum correlation algorithm is used to obtain the mutual information between each feature of each molecular fingerprint combination and the compound sample label (positive sample or negative sample). Since each molecular fingerprint contains multiple features, thus, multiple mutual informations corresponding to each molecular fingerprint combination are obtained. The minimum redundancy maximum correlation algorithm will automatically sort all the mutual informations from large to small according to the mutual information values, and obtain the sorted mutual informations corresponding to each molecular fingerprint combination;
[0126] In step S4, the sorted features corresponding to each compound are obtained according to the sorted mutual informations.
[0127] In the field of computer-aided compound development, the LightGBM algorithm is an ensemble tree model based on boosting. It constructs decision trees through the leaf growth strategy and fits the residuals of the previous prediction when building a new tree each time, thereby improving the convergence speed and accuracy of the model. Therefore, the present invention establishes a bitter compound prediction model based on the LightGBM algorithm, that is, uses the LightGBM algorithm as the core of the model to quickly and accurately predict the bitter characteristics of compounds. The features of the molecular fingerprint combinations corresponding to the sorted mutual informations of each compound sample in the training set are sequentially input into the bitter compound prediction model by using the feature incremental selection method, and the probability that the taste of each molecular fingerprint combination feature is bitter is output in turn. If the probability is greater than the set threshold (the threshold set in the present invention is 0.5), it means that the taste of the current compound sample is bitter, otherwise, it means that the taste of the current compound sample is not bitter.
[0128] For the training set, calculate the proportion of the number of compound samples with correct predictions for each molecular fingerprint combination in all compound samples respectively, that is, the prediction accuracy of each molecular fingerprint combination is obtained. According to the number of features of each molecular fingerprint combination and the prediction accuracy of the corresponding molecular fingerprint combination, a feature incremental selection curve graph of the corresponding molecular fingerprint combination is generated, as Figure 3 shown. Select all the features before the feature with the highest accuracy in the feature incremental selection curve graph to obtain the feature A of each molecular fingerprint combination. The feature A represents all the features before the feature with the highest accuracy, and the feature screening is completed. Through feature screening, features with small mutual information can be discarded and do not need to be input into the bitter compound prediction model subsequently, saving the time and cost of model calculation.
[0129] Figure 2 It shows the principle of the LightGBM algorithm. LightGBM has various advantages. For example, its leaf-wise growth strategy can maximize the gain of each node split, thereby improving the convergence speed of the model. In addition, LightGBM not only uses the first derivative but also utilizes the second derivative, which makes the model more accurate when optimizing the loss. At the same time, LightGBM adopts an efficient histogram algorithm for split point search, greatly accelerating the training speed of large-scale datasets and saving memory. Based on these advantages, the present invention uses the lightgbm.LGBMClassifier function in the LightGBM package in Python to train the bitter compound prediction model. The performance of the model in five-fold cross-validation is mainly evaluated by accuracy, and this step can help us select the optimal combination of molecular fingerprint features.
[0130] For example, if the number of compound samples in the training set is 3100, then the total number of any molecular fingerprint combination in the training set is also 3100. After training with the bitter compound prediction model, the total number of correct predictions of the taste of the compound by one molecular fingerprint combination, ECFP4, is 2635. Then the accuracy of the molecular fingerprint combination ECFP4 is 2635 / 3100 = 0.85. According to the above operation, the accuracies of the predictions of 15 molecular fingerprint combinations in the training set are obtained respectively. The abscissa of the feature increment selection curve is the number of features of a certain molecular fingerprint combination, and the ordinate is the prediction accuracy of the features under this molecular fingerprint combination.
[0131] LightGBM uses the following loss function to measure the error of the model during training:
[0132]
[0133] Among them, is the objective function, representing the error of the current model, is the loss function, describing the difference between the predicted probability and the true probability of each sample, y i is the true label value of the i-th sample, is the predicted value of the i-th sample, and N is the total number of training samples, that is, how many samples are there in the dataset. Ω(f m ) is the regularization term, used to control the complexity of the bitter compound prediction model.
[0134] For binary classification tasks, LightGBM uses Log Loss to calculate the difference between the predicted value and the true value:
[0135]
[0136] Among them, y i is the true label, is the predicted probability, and F(x i ) is the predicted value of the sample, which is obtained by accumulating the predicted values of all decision trees.
[0137] The regularization term Ω(f m ) controls the complexity of each tree and prevents overfitting. It usually consists of a leaf node number penalty term and an L2 regularization term:
[0138]
[0139] Among them, Q is the number of leaf nodes of the decision tree, γQ controls the complexity of the tree, and the more leaf nodes, the greater the penalty. ω j is the predicted value of the leaf node. λ is the L2 regularization parameter, which is used to prevent the model from overfitting and make the value of the leaf node not too large.
[0140] Step S6: Use the above feature screening to obtain the optimal number of features A for various molecular fingerprint combinations. Respectively take the optimal features of the test set under each molecular fingerprint combination as the input of the LightGBM algorithm, and the prediction accuracies of the test set under 15 molecular fingerprint combinations can be obtained respectively, as shown in Table 2. Select one molecular fingerprint combination with the highest AUC on the test set.
[0141] Table 2 AUC of the model based on the LightGBM algorithm on the test set
[0142]
[0143]
[0144] Note: E represents ECFP4, F represents FCFP4, M represents MACCS, and R represents RDKit.
[0145] According to Table 2, compared with a single molecular fingerprint, the integrated molecular fingerprint features improve the performance of the bitter compound prediction model to a certain extent, which indicates that adding more features may contain more compound structure information. However, it is not that the more molecular fingerprints are integrated, the better. According to Table 2, the integration result of the four fingerprints of ECFP4 + FCFP4 + MACCS + RDKit is not the best, because it may contain more redundant information, consume more time and computational cost, and instead lead the bitter compound prediction model to make wrong decisions. Among them, the molecular fingerprint combination of ECFP4 + FCFP4 has the highest AUC, so the present invention selects the first A features of this molecular fingerprint combination as the input of the bitter compound prediction model.
[0146] The present invention combines molecular fingerprint integration with machine learning algorithms, which can more comprehensively cover the information of all compound samples, more deeply apply feature vectors to characterize the chemical structure information of compound samples, and more quickly and accurately predict whether the taste of a small molecule compound is bitter. The machine learning algorithm not only speeds up the drug R & D process, but also reduces the R & D time and cost. Compared with traditional wet experiments, the present invention has low trial-and-error costs and high efficiency, can effectively reduce the time and financial costs required for drug development, accelerate the screening and testing process of bitter compound lead compounds, and provide strong reference and prior guidance for researchers and medical staff. The present invention can distinguish which compounds are bitter compounds and which are not, and can quickly screen out lead compounds with a bitter taste during subsequent applications, and put the compounds predicted to be bitter into the next step of research and production.
[0147] Step S7: Input the feature A of the ECFP4+FCFP4 molecular fingerprint combination corresponding to the compound sample to be predicted into the algorithm. The algorithm outputs the probability that the taste of the compound to be predicted is bitter, and compares the predicted probability with the set threshold to obtain whether the taste of the current compound sample is bitter.
[0148] There are three verification methods in the evaluation model: K-fold cross-validation, leave-one-out method, and independent test set. Since the data volume in the present invention is large, the leave-one-out method verification will greatly increase the model training time. Therefore, the present invention adopts five-fold cross-validation and independent test set verification, and uses multiple indicators to evaluate the performance of the bitter compound prediction model. The specific evaluation indicators of the model include:
[0149] (1) Sensitivity (Sn):
[0150]
[0151] (2) Specificity (Sp):
[0152]
[0153] (3) Accuracy (Acc):
[0154]
[0155] (4) Matthews Correlation Coefficient (MCC):
[0156]
[0157] (5) Area under the receiver operating characteristic curve (AUC).
[0158] Among them, TP represents the number of true positive samples, FN represents the number of false negative samples, TN represents the number of true negative samples, and FP represents the number of false positive samples. Sn represents the proportion of correctly classified positive samples among all positive samples; Sp represents the proportion of correctly classified negative samples among all negative samples; Acc is an index used to evaluate the model performance, representing the proportion of all correctly predicted compound samples among the total compound samples; MCC is an index applied in machine learning to measure the classification performance of binary classification; AUC is one of the most important evaluation indexes in machine learning. The X-axis represents sensitivity and the Y-axis represents specificity. When the AUC is relatively high, it indicates that the bitter compound prediction model has good classification ability.
[0159] The present invention uses a test set to evaluate the performance of the algorithm. The sensitivity of the algorithm is 73.32%, the specificity is 91.72%, the accuracy rate is 84.41%, the Matthews correlation coefficient is 0.672, and the AUC is 0.924. These results show that the algorithm of the present invention has excellent classification ability.
[0160] From the above results, it can be seen that the method for predicting bitter compounds constructed by the present invention by combining molecular fingerprint features with machine learning algorithms can quickly and accurately predict whether the taste of a compound is bitter, saving a large amount of time for the research and development of bitter compounds, improving the success rate of screening compound lead compounds, and reducing the development cost of bitter compounds.
[0161] The above examples of the present invention are only used to illustrate in detail the calculation model and calculation process of the present invention, rather than limiting the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or variations can be made on the basis of the above description. It is impossible to list all the implementation manners here. Any obvious changes or variations derived from the technical solutions of the present invention still fall within the protection scope of the present invention.
Claims
1. A bitter compound prediction method based on machine learning, characterized in that: The method specifically comprises the following steps: Step S1, respectively obtaining bitter compound samples and sweet compound samples from a compound database, and dividing all the obtained compound samples into a training set and a test set; Step S2, respectively obtaining the SMILES format of each compound sample in the training set and the test set; For any SMILES format of a compound sample, the SMILES format of the compound sample is converted into different types of molecular fingerprints, and then the different types of molecular fingerprints of the compound sample are integrated to obtain the molecular fingerprint combination corresponding to the compound sample; Similarly, the molecular fingerprint combination corresponding to each compound sample is obtained; Step S3, for any molecular fingerprint combination, according to the feature vectors of various compounds under the molecular fingerprint combination, calculate the mutual information between each feature and the compound label under the molecular fingerprint combination, and then sort the calculated mutual information in descending order to obtain sorted mutual information; obtain the feature sorting result corresponding to the molecular fingerprint combination according to the sorted mutual information; Similarly, the feature ranking results corresponding to each molecular fingerprint combination are obtained respectively; Step S4: for any compound, sort the characteristic values in the characteristic vector of the compound according to the characteristic sorting result of each molecular fingerprint combination, and obtain the sorted characteristic vector of the compound under each molecular fingerprint combination; Similarly, the sorted feature vectors of each compound under each molecular fingerprint combination are obtained; Step S5, training the LightGBM algorithm according to the sorted feature vectors of each compound in the training set under each molecular fingerprint combination to obtain the optimal number of features corresponding to each molecular fingerprint combination; Step S6, according to the optimal number of features for each molecular fingerprint combination obtained in step S5, each compound in the test set performs feature screening on the sorted feature vector under each molecular fingerprint combination, and uses the feature vector of the screened feature composition as the input of the LightGBM algorithm. The prediction accuracy of the test set samples under each molecular fingerprint combination is output through the LightGBM algorithm, and the molecular fingerprint combination corresponding to the highest prediction accuracy obtained by the test set samples is used as the optimal molecular fingerprint combination a, and the optimal number of features corresponding to the molecular fingerprint combination a is recorded as A; Step S7, obtain the SMILES format of the compound to be predicted, then obtain the molecular fingerprint combination a of the compound to be predicted, sort the features in the feature vector of the compound to be predicted under the molecular fingerprint combination a according to the mutual information sorting, and then select the first A features after sorting, and use the feature vector composed of the selected first A features as the input of the LightGBM algorithm, and obtain the bitterness prediction result of the compound to be predicted according to the output of the LightGBM algorithm.
2. The bitter compound prediction method based on machine learning according to claim 1, characterized in that: The bitter compound samples are obtained from BitterDB database, PubMed database, FoodDB database, FlavorDB2 database, FSBI-DB database, ChemTastesDB database and The Good Scents Company database; The sweet compound samples are obtained from SweetenersDB database, SuperSweet database, PubMed database, FoodDB database, FlavorDB2 database, FSBI-DB database, ChemTastesDB database and The GoodScents Company database.
3. The bitter compound prediction method based on machine learning according to claim 1, characterized in that: The specific process of step S2 is: Step S21, calling the pubchempy package in the pandas library to obtain the SMILES format of each compound sample; Step S22, calling the RDkit package in the pandas library to convert the SMILES format of each compound sample into different types of molecular fingerprints; Step S23: for any compound sample obtained, the molecular fingerprint corresponding to the SMILES format of the compound sample is integrated to obtain a molecular fingerprint combination corresponding to the compound sample; Similarly, the molecular fingerprint combination corresponding to each compound sample is obtained.
4. The bitter compound prediction method based on machine learning according to claim 1, characterized in that: The types of molecular fingerprints include ECFP4 fingerprint, FCFP4 fingerprint, RDKit fingerprint and MACCS fingerprint.
5. The bitter compound prediction method based on machine learning according to claim 1, characterized in that: The specific process of step S23 is as follows: For each compound sample, 15 molecular fingerprint combinations were obtained, and the 15 molecular fingerprint combinations are: ECFP4 fingerprint, FCFP4 fingerprint, RDKit fingerprint, MACCS fingerprint; ECFP4 fingerprint + FCFP4 fingerprint, ECFP4 fingerprint + RDKit fingerprint, ECFP4 fingerprint + MACCS fingerprint, FCFP4 fingerprint + RDKit fingerprint, FCFP4 fingerprint + MACCS fingerprint, RDKit fingerprint + MACCS fingerprint; ECFP4 fingerprint + FCFP4 fingerprint + RDKit fingerprint, ECFP4 fingerprint + FCFP4 fingerprint + MACCS fingerprint, ECFP4 fingerprint + RDKit fingerprint + MACCS fingerprint, FCFP4 fingerprint + RDKit fingerprint + MACCS fingerprint; ECFP4 fingerprint + FCFP4 fingerprint + RDKit fingerprint + MACCS fingerprint.
6. The bitter compound prediction method based on machine learning according to claim 1, characterized in that: The mutual information is calculated using a minimum redundancy maximum correlation algorithm.
7. The bitter compound prediction method based on machine learning according to claim 5, characterized in that: The specific process of step S5 is as follows: Take any molecular fingerprint combination as an example Step S51, dividing all compound samples in the training set into 5 subsets, and numbering each subset in sequence as subset 1, subset 2, subset 3, subset 4 and subset 5; Initialize the number of features n = 1; Step S52, initializing j=1; Step S53: Use the jth subset as a verification sample, and use the remaining four subsets as training samples; Step S54, selecting the first n features from the sorted feature vectors of each training sample under the current molecular fingerprint combination, and using the selected features to form a new feature vector for each training sample; Step S55, using the new feature vector of each training sample as the input of the LightGBM algorithm to obtain a decision tree constructed by the LightGBM algorithm; Step S56, selecting the first n features from the sorted feature vectors of each verification sample in the jth subset under the current molecular fingerprint combination, and using the selected features to form a new feature vector for each verification sample; The new feature vector of each validation sample is used as the input of the decision tree constructed in step S55, the prediction result for each validation sample is output through the decision tree, and the prediction accuracy of the samples in the jth subset is calculated; Step S57, determine whether j is equal to 5; If j is equal to 5, then the prediction accuracy of each subset when used as a validation sample is averaged, and the average result is used as the prediction accuracy corresponding to the number of features n; and continue to execute step S58; If j is less than 5, set j=j+1 and return to step S53; Step S58, determining whether the number of features n is equal to N, where N represents the total number of features in the feature vector under the current molecular fingerprint combination; If n is equal to N, the number of features n' corresponding to the highest prediction accuracy under the current molecular fingerprint combination is obtained, and the number of features n' is taken as the optimal number of features under the current molecular fingerprint combination; If n is less than N, n=n+1 is set and the process returns to step S52.
8. The bitter compound prediction method based on machine learning according to claim 7, characterized in that: The working process of the LightGBM algorithm is: Step 1: Initialize the prediction values of all samples to be F0(x i ); in, Step 2, initialize the number of regression trees t = 1; Step 3, initialize the node level l of the regression tree = 1; Step 4: Determine whether there is a node with a sample number greater than or equal to the threshold on the lth level of the tth regression tree; If there are nodes whose sample numbers are greater than or equal to the threshold, mark each node whose sample numbers are less than the threshold as a leaf node, and execute step 5 for the nodes whose sample numbers are greater than or equal to the threshold; If there is no node with a sample number greater than or equal to the threshold, all nodes on the lth level are marked as leaf nodes, and then step 7 is executed; Step 5: Each node calculates its own maximum splitting gain, marks the node whose maximum splitting gain is less than the threshold as a leaf node, and for each node whose maximum splitting gain is greater than or equal to the threshold, selects the splitting point that maximizes its own splitting gain for splitting, and distributes its own samples to the left child node and the right child node; The calculation method of split gain is: Where G represents the splitting gain, I represents the set of all compound samples in the current node to be split, and I L Represents the compound sample set in the left child node after splitting, I R represents the compound sample set in the right child node after splitting, g i (t) represents the first-order gradient of the loss function of the ith compound sample, h i (t) represents the second-order gradient of the loss function of the i-th compound sample, λ is the regularization term coefficient, and γ is the penalty term coefficient; Where σ(·) represents the sigmoid activation function, F t-1 (x i ) represents the t-1th regression tree for sample x i The prediction result, y i Represents sample x i The true label of Step 6: Set the node level l=l+1, and return to step 4; Step 7: Calculate the output of each leaf node and record the output of the jth leaf node as ω j (t) ; When compound sample x i When in the jth leaf node of the tth regression tree, let f t (x i )=ω j (t) , the tth regression tree for sample x i The prediction result is F t (x i ): F t (x i )=F t-1 (x i )+γ t f t (x i ) Among them, γ t is the learning rate of the tth regression tree; Step 8: Determine whether t=M is satisfied, where M is the total number of regression trees set; If not satisfied, set t=t+1 and return to step 3; If satisfied, the process ends and T regression trees are obtained.
9. The bitter compound prediction method based on machine learning according to claim 8, characterized in that: The new feature vector of each validation sample is used as the input of the decision tree constructed in step S55, and the prediction result of each validation sample is output through the decision tree; the specific process is: For the validation sample x: Among them, F(x) represents the prediction result of the validation sample x; M is the number of decision trees; γ t is the learning rate; f t (x) is the output of the leaf node where the verification sample x is located in the tth tree; Calculate the probability p' that the verification sample x is a bittering agent based on F(x): Where, e represents the base of natural logarithm; If p'>0.5, then the verification sample x is a bitter agent, otherwise, the verification sample x is a non-bitter agent.
10. The bitter compound prediction method based on machine learning according to claim 9, characterized in that: The output ω of the jth leaf node j (t) for: Among them, I j Represents the set of compound samples contained in the jth leaf node in the tth regression tree.