Artificial intelligence-based traditional chinese medicine-derived trna, trfs, and trna halves activity prediction method
Patent Information
- Application Number
- PCT/CN2025/104550
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-21
- Filing Date
- 2025-06-27
- Publication Date
- 2026-09-24
Smart Images

Figure CN2025104550_24092026_PF_FP_ABST
Abstract
Description
Artificial Intelligence-Based Prediction Method for the Activity of tRNA, tRFs, and tRNA Halves in Traditional Chinese Medicine Technical Field
[0001] This invention relates to the field of computer-aided drug screening, specifically to an artificial intelligence-based method for predicting the activity of tRNA, tRFs, and tRNA halves in traditional Chinese medicine, applicable to predicting cell activity based on the tRNA sequence of traditional Chinese medicine. Background Technology
[0002] Traditional Chinese medicine is a treasure of humankind, but its specific mechanisms of action and effective components still require further research and clarification. Recent reports indicate that non-coding RNAs, such as microRNAs, exert different regulatory effects by targeting different aspects of RNA transcription or post-transcriptional processes in almost all eukaryotes. Lin Zhang et al. (Cell Research 2012, 22, 107-126) proposed that exogenous plant microRNAs in food can be absorbed by the digestive tract of mammals and transported to various organs via the bloodstream, thereby regulating mammalian gene expression. The research group of Jiang Zhihong at Macau University of Science and Technology discovered that ginseng-derived tRNA has a protective effect on cardiomyocytes (CN111419867B), and yew-derived tRNA has an inhibitory effect on ovarian cancer (Mol. Ther. Nucleic. 2022, 27, 718-732).
[0003] For pharmacodynamic studies of RNA derived from traditional Chinese medicine, the first step is to verify cell viability. Cell viability refers to the proportion of healthy cells in a sample population. Detecting cell viability is essential for observing the physiological state of cells during experiments and is also an important indicator for evaluating drug efficacy. Commonly used methods for cell viability detection include MTT and CCK-8. Dehydrogenases in living cells can reduce WST-8 in exogenous MTT and CCK-8 reagents to water-insoluble blue-purple crystalline formazan and highly water-soluble yellow formazan dye, respectively, while dead cells lack this function. MTT detection is characterized by high sensitivity and cost-effectiveness. However, because formazan crystals are insoluble in water, they must be dissolved in dimethyl sulfoxide (DMSO) before detection. This results in a large workload, lower accuracy, and DMSO can be harmful to cells and experimenters. In contrast, the CCK-8 method is convenient to use, has high detection sensitivity, and low toxicity. However, due to factors such as cell growth cycles, the detection of large numbers of samples often involves long experimental cycles and high labor costs. Therefore, developing efficient methods for evaluating drug activity is of great significance.
[0004] With the development of information technology, domestic and international research has begun to explore the use of machine learning and RNA features to establish computational models for predicting RNA drug activity. However, these methods face the following challenges: 1) The accuracy of activity prediction is often low for diverse RNA sequences; 2) The prediction results lack biological significance and cannot be systematically explained from the perspective of biological mechanisms of action; 3) Traditional machine learning methods cannot automatically learn feature information from large datasets, requiring extensive manual feature selection. Therefore, to address these limitations, this invention provides a deep learning-based method for predicting the cellular activity of tRNA, tRFs, and tRNA halves in traditional Chinese medicine. This method can rapidly, efficiently, and systematically predict the cellular activity of tRNA, tRFs, and tRNA halves. Summary of the Invention
[0005] Technical problem solved: This invention overcomes the shortcomings of existing technologies and discloses a method for predicting the cell activity of tRNA, tRFs, and tRNA halves of traditional Chinese medicine based on deep learning algorithms (DL). This method fully combines the traditional Chinese medicine tRNA database with artificial intelligence deep learning algorithms to establish an efficient and rapid method for predicting the cell activity of tRNA, tRFs, and tRNA halves of traditional Chinese medicine, providing an innovative approach for the research and development of new traditional Chinese medicine tRNA drugs.
[0006] Technical Solution: An AI-based method for predicting the activity of tRNA, tRFs, and tRNA halves in traditional Chinese medicine (TCM), comprising the following steps: Step 1: Extracting and separating tRNA from the target TCM material and sequencing it; Step 2: Designing tRFs and tRNA halves sequences based on the tRNA sequencing results, and constructing modeling data samples based on TCM tRNA; Step 3: Purifying the tRNA from the target TCM material, determining the types and sites of modifications on the tRNA using LC-MS / MS, and designing modified tRFs and tRNA halves sequences. The modifications are designed based on the types of natural tRNA modifications in TCM materials, using the stability and efficacy of tRFs and tRNA halves as indicators, and adjusting and optimizing the modification types and sites; Step 4: Performing cell activity experiments on the modified tRFs and tRNA halves, and screening the results as features for the final model samples; Step 5: Constructing a prediction model based on a deep learning algorithm, using the screened tRFs and tRNA halves with or without cell activity as feature sequences for model training and learning; Step 6: Optimizing the key parameters of the model through a grid search algorithm and cross-validation to improve the model's predictive performance.
[0007] Step one above includes the following steps: The target Chinese medicinal materials are yew, ginseng, Ganoderma lucidum, toad venom, Panax notoginseng, Salvia miltiorrhiza, or Ginkgo biloba; total RNA is extracted from the Chinese medicinal materials, and small RNAs with a length <200mer are separated and enriched; total tRNA in the small RNA sample is separated by 6% polyacrylamide gel electrophoresis containing 8M urea, stained with SYBR nucleic acid dye, and the gel separation bands are observed under blue light, and the gel region containing total tRNA is cut off; the cut gel is placed in a 3000MWCO dialysis bag, eluted by transverse electrophoresis, and the eluent is recovered; total tRNA in the eluent is purified by TRIzol method; the total tRNA sample is taken, a cDNA library is constructed using TruSeq Small RNA Sample Preparation Kit, and sequencing is performed using the Illumina next-generation sequencing platform; the tRNA next-generation sequencing data is quality controlled using bioinformatics methods, and the sequence and content of each tRNA are analyzed, finally obtaining the sequence information of tRNA contained in the target Chinese medicinal materials.
[0008] Step two above includes the following steps: Based on the sequence information of tRNA contained in the target Chinese medicinal material, design tRFs and tRNA halves sequences, including (1) 5'-t-half: refers to the sequence formed from the break point to the 5' end of tRNA after any break in the anticodon loop of any tRNA; (2) 3'-t-half: refers to the sequence formed from the break point to the 3' end of tRNA after any break in the anticodon loop of any tRNA; (3) 5'-tRF: refers to the sequence formed by the break from the 5' end of any tRNA to any position before the anticodon loop; (4) 3'-tRF: refers to the sequence formed by the break from the 3' end of any tRNA to any position before the anticodon loop; 20% of the modeling data samples are used as the test set, and the remaining data are used as the training set for model construction.
[0009] Step three above includes the following steps: Designing a biotinylated DNA probe based on the tRNA sequence information of the target medicinal material; utilizing the complementary pairing principle between the probe and the tRNA sequence of the target medicinal material to affinity capture the tRNA; finally, using streptavidin-modified magnetic beads to recover and purify the biotinylated DNA probe-tRNA complex; performing restriction hydrolysis of the purified tRNA monomer using RNA restriction endonucleases or RNA exonucleases to form tRNA fragments; performing qualitative analysis of the obtained tRNA fragments using LC-MS / MS; and analyzing the MS... 1 and MS 2Mass spectrometry was analyzed to confirm the tRNA modification type and site of the target Chinese medicinal material; based on the LC-MS / MS modification sequence sequencing results, tRFs and tRNA halves sequences containing modified nucleotides were designed to construct modeling data samples based on Chinese medicinal tRNA; the modification types include pseudouracil, dihydrouracil, base monomethylation modification, base dimethylation modification, base trimethylation modification, and nucleotide pentose sugar 2'-oxymethylation modification.
[0010] Step four above includes the following steps: Based on the disease treated by the selected single Chinese herbal medicine, relevant cells are selected, and cell activity is screened using CCK8. The cell activity experiment involves analyzing the correlation and differences in cell activity between the blank control group and the experimental group. For tumor cells, bands with IC50 ≤ 100 nM or below 50 nM and an inhibition rate exceeding 50% are selected as feature sequences. For normal cells, bands with IC50 ≤ 50 nM, cell activity 100% ± 10%, and significant differences from the modeling group after modeling with appropriate stimulation are selected as feature sequences. The sequences obtained from the cell experiment screening are converted into feature vectors as input for deep learning. The K-mers method is used to divide the sequence into at least K long base strings, K ≥ 12, and the frequency of these base strings is calculated as the feature representation of the sequence. Simultaneously, the sequence classification information—classified as active or inactive based on the cell experiment—is incorporated into the model input. A deep learning neural network is used to learn the feature vectors and corresponding classification information to obtain a model capable of predicting the activity of unknown sequences.
[0011] Step five above includes: selecting a sequential model interface; the model structure includes an input layer, hidden layers, and an output layer, where the hidden layers include fully connected layers (dense layers) and dropout layers; the model input consists of the feature vectors of the sequences processed in step four and the sequence activity classification information; the model output is a binary classification result of 0 and 1, where 0 represents inactive (IC50 > 100 nM) and 1 represents active (IC50 ≤ 100 nM); in the hidden layers of the model, through Rectified Linear... The ReLU activation function is used to activate the input layer values before passing them to the fully connected layer. The formula for this activation function is: y = ReLU(Wx + b); where x is the input data value, y is the activated value, W is the weight matrix, and b is the bias. In the model's output layer, the Sigmoid activation function is used to activate the hidden layer values before outputting the final result. The formula for this activation function is: z = sigmoid(W'y + b'); where y is the activated value from the hidden layer, z is the model output, W' is the transposed weight matrix, and b' is the transposed bias. The model's output layer further interfaces with public biomedical databases, including Array Express, Gene Expression, and Omnibus (GEO). During model training, the compile module is used to configure the model's learning process. The parameters are set as follows: optimizer is set to Root Mean Square prop (RMSprop), metrics are set to accuracy, and loss function is set to binary_crossentropy. The formula for calculating the loss function is: Among them, L H (x,z) represents the difference between the predicted value and the true value, i.e., the loss. x is the true value corresponding to the sample, z is the predicted value corresponding to the sample, and d is the number of epochs.
[0012] Step six above, which optimizes the key parameters of the model using a grid search algorithm and cross-validation, includes: setting the parameter optimization range, where the epoch number is [10, 50, 100, 200, 500], the batch size is [10, 32, 64, 128], the learning rate is [0.01, 0.001, 0.00001], the dropout rate is [0, 0.2, 0.5], and the node... The numbers are [50, 100, 300, 500, 1000]. A grid search algorithm is used to optimize 5×4×3×3×5 models. The predictive performance of the models is evaluated using a 10-fold cross-validation model and evaluation metrics, including: Sensitivity (SEN); Specificity (SPE); Accuracy (ACC); Matthews correlation coefficient (MCC); and Area Under the ROC Curve (AUC). The closer the sensitivity, specificity, and accuracy are to 100%, and the closer the Matthews correlation coefficient and AUC are to 1, the better the model's predictive performance. Conversely, the closer the sensitivity, specificity, and accuracy are to 0, and the closer the Matthews correlation coefficient and AUC are to 0.5, the worse the model's predictive performance.
[0013] Among them, TP represents true positive; TN represents true negative; FP represents false positive; and FN represents false negative.
[0014] Beneficial effects: 1. This prediction method can accurately predict the cellular activity of different Chinese herbal medicine (TCM) tRNAs and has good robustness; 2. The deep learning algorithm used in this method has a strong ability to automatically learn features, and can automatically learn important feature information from big data, avoiding a large amount of manual feature selection; 3. The TCM tRNA, tRFs, and tRNA halves activity prediction model constructed by this method has excellent prediction performance and good applicability, providing an innovative approach for the research and development of new TCM tRNA drugs. Attached Figure Description
[0015] Figure 1 is a flowchart of the deep learning-based tRNA prediction method for traditional Chinese medicine according to the present invention.
[0016] Figure 2 is a schematic diagram of the structure of the predictive model for the activity of tRNA, tRFs, and tRNA halves in traditional Chinese medicine according to the method of the present invention.
[0017] Figure 3 shows the prediction performance of the deep learning-based tRNA prediction method for traditional Chinese medicine. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be further described below in conjunction with the accompanying drawings.
[0019] Example 1
[0020] The specific technical solution of an artificial intelligence-based method for predicting the activity of tRNA, tRFs, and tRNA halves in traditional Chinese medicine is as follows:
[0021] 157 RNA sequences were obtained by isolating Ganoderma lucidum, a traditional Chinese medicine. 1058 data points were then obtained using CCK8 assays in human ovarian cancer cells (A2780), airway epithelial cells (16HBE), and airway smooth muscle cells (ASMC). All modeling sample data were randomly divided into training and test sets at a ratio of 80%:20%.
[0022] The tRNA data obtained from sequencing were subjected to quality control alignment to obtain tRNA sequence information. Through cell viability experiments, statistical analysis was performed on cells in the experimental group and the control group. For A2780, bands with IC50 ≤ 100 nM were selected as characteristic sequences. For ASMC, bands with cell viability of 100% ± 10% within 50 nM and inhibition rate > 30% within 50 nM after TGF-β1 modeling were selected as characteristic sequences. The sequences were further converted into feature vectors, and K-mers were used to divide the sequences into groups of 3 bases. The frequency of these base groups was calculated as the feature representation of the sequence. At the same time, the sequence viability classification information of the cell experiment was incorporated into the sequence features.
[0023] This approach uses Anaconda 5.1 with Python 3.6 to create a virtual environment and leverages the Keras deep learning framework and the sklearn machine learning package to build a tRNA activity prediction model. The model selects the sequential model interface to build a binary classification prediction model. The model structure includes an input layer, hidden layers, and an output layer. The hidden layers include fully connected layers (Dense layer) and Dropout layers (Figure 2). In the hidden layers, the values from the input layer are activated using the Rectified Linear Unit (ReLU) activation function, which is then passed to the fully connected layer. The formula for this activation function is: y = ReLU(Wx + b).
[0024] Where x is the value of the input data, y is the value of the data after activation, W is the weight matrix, and b is the bias;
[0025] In the model's output layer, the values of the hidden layers are activated using the sigmoid activation function, which is then output as the final result. The formula for this activation function is: z = sigmoid(W'y + b').
[0026] Where y is the activation value from the hidden layer, z is the model output, W' is the transpose weight matrix, and b' is the transpose bias;
[0027] The model's output layer further interfaces with public biomedical databases, including Array Express, Gene Expression, and Omnibus (GEO).
[0028] During model training, the `compile` module is used to configure the model's learning process. Its parameters are set as follows: the optimizer is set to Root Mean Square prop (RMSprop), the metrics list is set to accuracy, and the loss function is set to binary_crossentropy. The formula for calculating this loss function is:
[0029] Among them, L H (x,z) represents the difference between the predicted value and the true value (i.e., loss), where x is the true value corresponding to the sample, z is the predicted value corresponding to the sample, and d is the number of epochs.
[0030] Then, the feature vector data of the final processed sequence is used as the input of the model, with 80% used as the training set to train the model and 20% used as the test set to test the model performance.
[0031] According to the artificial intelligence-based method for predicting the activity of tRNA, tRFs, and tRNA halves in traditional Chinese medicine as described in claim 1, step six includes the following steps:
[0032] Set the parameter optimization range, where the epoch number is [10, 50, 100, 200, 500], the batch size is [10, 32, 64, 128], the learning rate is [0.01, 0.001, 0.00001], the dropout rate is [0, 0.2, 0.5], and the node number is [50, 100, 300, 500, 1000].
[0033] The predictive performance of the trained model was evaluated using test set samples from the modeling data, employing metrics such as Sensitivity (SEN), Specificity (SPE), Accuracy (ACC), Matthews correlation coefficient (MCC), and the area under the Receiver Operating Characteristic (ROC) curve (AUC) for performance assessment. Furthermore, the key parameters (epoch number, batch size, learning rate, dropout rate, and node number) of the 900 (5×4×3×3×5) models were optimized using a grid search algorithm and 10-fold cross-validation to achieve optimal predictive performance.
[0034] Among them, TP represents true positive; TN represents true negative; FP represents false positive; and FN represents false negative.
[0035] Finally, through parameter optimization, the optimal model was set with the following parameters: two hidden layers with 50 nodes each, a dropout rate of 0.5 to avoid overfitting, a learning rate of 0.001, a batch size of 128, and an epoch number of 50. Performance evaluation of the optimal model on the test set showed a prediction accuracy of 97.1%, an AUC of 0.989, a sensitivity of 97.4%, a specificity of 96.8%, and a Matthews correlation coefficient of 0.942. Compared to most RNA prediction models based on traditional machine learning, both domestically and internationally, this model exhibits superior predictive performance (Figure 3, Table 1).
[0036] The above example is merely a specific embodiment of the present invention, and simple modifications and substitutions thereof are also within the scope of protection of the invention.
[0037] Table 1. Prediction results of the deep learning-based tRNA prediction method for traditional Chinese medicine.
Claims
1. A method for predicting the activity of tRNA, tRFs, and tRNA halves in traditional Chinese medicine based on artificial intelligence, characterized in that, Includes the following steps: Step 1: Extract and isolate tRNA from the target Chinese medicinal material and perform sequencing; Step 2: Based on the tRNA sequencing results, design tRFs and tRNA halves sequences to construct modeling data samples based on traditional Chinese medicine tRNA; Step 3: Purify the tRNA of the target Chinese medicinal material, determine the types and sites of modifications on the tRNA using LC-MS / MS, and design modified tRFs and tRNA halves sequences. The modifications are designed based on the types of natural tRNA modifications in the Chinese medicinal material, and the stability and efficacy of tRFs and tRNA halves are used as indicators to adjust and optimize the modification types and sites. Step 4: Perform cell viability experiments on the modified tRFs and tRNA halves, and screen the results as the final model sample features; Step 5: Construct a prediction model based on deep learning algorithms, and use the selected tRFs and tRNA halves as feature sequences based on whether or not they have cell activity for model training and learning; Step Six: Optimize the key parameters of the model using grid search algorithm and cross-validation to improve the model's predictive performance.
2. The method for predicting the activity of tRNA, tRFs, and tRNA halves in traditional Chinese medicine based on artificial intelligence according to claim 1, characterized in that, Step one includes the following steps: The target Chinese medicinal materials are yew, ginseng, Ganoderma lucidum, toad venom, Panax notoginseng, Salvia miltiorrhiza, or Ginkgo biloba; total RNA is extracted from the Chinese medicinal materials, and small RNAs with a length <200mer are separated and enriched; total tRNA in the small RNA sample is separated by 6% polyacrylamide gel electrophoresis containing 8M urea, stained with SYBR nucleic acid dye, and the gel separation bands are observed under blue light, and the gel region containing total tRNA is cut off; the cut gel is placed in a 3000MWCO dialysis bag, eluted by transverse electrophoresis, and the eluent is recovered; total tRNA in the eluent is purified by TRIzol method; the total tRNA sample is taken, a cDNA library is constructed using TruSeq Small RNA Sample Preparation Kit, and sequencing is performed using the Illumina next-generation sequencing platform; the tRNA next-generation sequencing data is quality controlled using bioinformatics methods, and the sequence and content of each tRNA are analyzed, finally obtaining the sequence information of tRNA contained in the target Chinese medicinal materials.
3. The method for predicting the activity of tRNA, tRFs, and tRNA halves in traditional Chinese medicine based on artificial intelligence according to claim 1, characterized in that, Step two includes the following steps: Based on the sequence information of tRNA contained in the target Chinese medicinal material, design tRFs and tRNA halves sequences, including (1) 5'-t-half: refers to the sequence formed from the break point to the 5' end of tRNA after any break in the anticodon loop of any tRNA; (2) 3'-t-half: refers to the sequence formed from the break point to the 3' end of tRNA after any break in the anticodon loop of any tRNA; (3) 5'-tRF: refers to the sequence formed by the break from the 5' end of any tRNA to any position before the anticodon loop; (4) 3'-tRF: refers to the sequence formed by the break from the 3' end of any tRNA to any position before the anticodon loop; 20% of the modeling data samples are used as the test set, and the remaining data are used as the training set for model construction.
4. The method for predicting the activity of tRNA, tRFs, and tRNA halves in traditional Chinese medicine based on artificial intelligence according to claim 1, characterized in that, Step three includes the following steps: designing a biotinylated DNA probe based on the tRNA sequence information of the target medicinal material; utilizing the complementary pairing principle between the probe and the tRNA sequence of the target medicinal material to affinity capture the tRNA; finally, using streptavidin-modified magnetic beads to recover and purify the biotinylated DNA probe-tRNA complex; performing restriction hydrolysis of the purified tRNA monomer using RNA restriction endonucleases or RNA exonucleases to form tRNA fragments; performing qualitative analysis of the obtained tRNA fragments using LC-MS / MS; and analyzing the MS... 1 and MS 2 Mass spectrometry was analyzed to confirm the type and site of tRNA modification in the target Chinese medicinal material; Based on the LC-MS / MS modified sequence sequencing results, tRFs and tRNA halves sequences containing modified nucleotides were designed to construct modeling data samples based on traditional Chinese medicine tRNAs; the modification types include pseudouracil, dihydrouracil, base monomethylation, base dimethylation, base trimethylation, and nucleotide pentose sugar 2'-oxymethylation.
5. The method for predicting the activity of tRNA, tRFs, and tRNA halves in traditional Chinese medicine based on artificial intelligence according to claim 1, characterized in that, Step four includes the following steps: Based on the disease treated by the selected single Chinese herbal medicine, relevant cells are selected, and cell activity is screened using CCK8. The cell activity experiment involves analyzing the correlation and differences in cell activity between the blank control group and the experimental group. For tumor cells, bands with IC50 ≤ 100 nM or below 50 nM and an inhibition rate exceeding 50% are selected as feature sequences. For normal cells, bands with 50 nM, cell activity of 100% ± 10%, and significant differences from the modeling group after modeling with appropriate stimulation are selected as feature sequences. The sequences obtained from the cell experiment screening are converted into feature vectors as input for deep learning. The K-mers method is used to divide the sequence into at least K long base strings, K ≥ 12, and the frequency of these base strings is calculated as the feature representation of the sequence. Simultaneously, the sequence classification information—classified as active or inactive based on the cell experiment—is incorporated into the model input. A deep learning neural network is used to learn the feature vectors and corresponding classification information to obtain a model capable of predicting the activity of unknown sequences.
6. The method for predicting the activity of tRNA, tRFs, and tRNA halves in traditional Chinese medicine based on artificial intelligence according to claim 1, characterized in that, Step five above includes: selecting a sequential model interface; the model structure includes an input layer, hidden layers, and an output layer, where the hidden layers include a fully connected layer (Dense layer) and a dropout layer; the model input consists of the feature vector of the sequence obtained in step four and the sequence activity classification information; the model output is a binary classification result of 0 and 1, where 0 represents inactivity (IC50 > 100 nM) and 1 represents activity (IC50 ≤ 100 nM); in the hidden layers, the values of the input layer are activated using the Rectified Linear Unit (ReLU) activation function and then passed to the fully connected layer. The formula for this activation function is: y = ReLU(Wx + b). Where x is the value of the input data, y is the value of the data after activation, W is the weight matrix, and b is the bias; In the model's output layer, the hidden layer values are activated using the Sigmoid activation function, which then becomes the final output. The formula for this activation function is: z = sigmoid(W'y + b') Where y is the activation value from the hidden layer, z is the model output, W' is the transpose weight matrix, and b' is the transpose bias; The model's output layer further interfaces with public biomedical databases, including Array Express, Gene Expression, and Omnibus (GEO); During model training, the `compile` module is used to configure the model's learning process. Its parameters are set as follows: the optimizer is set to Root Mean Square prop (RMSprop), the metrics list is set to accuracy, and the loss function is set to binary_crossentropy. The formula for calculating this loss function is: Among them, L H (x,z) represents the difference between the predicted value and the true value, i.e., the loss; x is the true value corresponding to the sample, z is the predicted value corresponding to the sample, and d is the number of epochs.
7. The method for predicting the activity of tRNA, tRFs, and tRNA halves in traditional Chinese medicine based on artificial intelligence according to claim 1, characterized in that, Step six, which involves optimizing the key parameters of the model using a grid search algorithm and cross-validation, includes: setting the parameter optimization range, where the epoch number is [10, 50, 100, 200, 500], the batch size is [10, 32, 64, 128], the learning rate is [0.01, 0.001, 0.00001], the dropout rate is [0, 0.2, 0.5], and the node... The numbers are [50, 100, 300, 500, 1000]. A grid search algorithm is used to optimize 5×4×3×3×5 models. The predictive performance of the models is evaluated using a 10-fold cross-validation model and evaluation metrics, including: Sensitivity (SEN); Specificity (SPE); Accuracy (ACC); Matthews correlation coefficient (MCC); and Area Under the ROC Curve (AUC). The closer the sensitivity, specificity, and accuracy are to 100%, and the closer the Matthews correlation coefficient and AUC are to 1, the better the model's predictive performance. Conversely, the closer the sensitivity, specificity, and accuracy are to 0, and the closer the Matthews correlation coefficient and AUC are to 0.5, the worse the model's predictive performance. Among them, TP represents true positive; TN represents true negative; FP represents false positive; and FN represents false negative.