Traditional Chinese medicine tRNA, tRFs and tRNA halves activity prediction method based on artificial intelligence
Through deep learning algorithm combined with the traditional Chinese medicine tRNA database, the problem of cumbersomeness of Chinese medicine RNA cell activity detection and low prediction accuracy is solved, and rapid and accurate cell activity prediction is achieved, providing innovative methods for the research and development of new traditional Chinese medicine drugs.
Patent Information
- Application Number
- CN202510340271.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-08-15
AI Technical Summary
In the research on the efficacy of traditional Chinese medicine RNA, the cell activity detection methods are complicated and have poor accuracy. The prediction accuracy of machine learning methods is low and lacks biological explanation, so it is impossible to efficiently predict tRNA, tRFs, and tRNA halves cell activity.
Deep learning algorithms are used to combine traditional Chinese medicine tRNA database, and the tRNA of Chinese herbal medicines are extracted, sequenced and purified, characterized sequences are designed, cell activity experiments are carried out, prediction models are constructed, and model parameters are optimized through grid search and cross-validation to achieve fast and accurate activity prediction.
The efficiency and accuracy of cell activity prediction of traditional Chinese medicine tRNA, tRFs, and tRNA halves is achieved. The deep learning algorithm automatically learns feature information, avoids manual feature selection, and provides innovative ideas for the research and development of new traditional Chinese medicines.
Smart Images

Figure CN120496673A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer-assisted drug screening, and in particular to an artificial intelligence-based method for predicting the activity of tRNA, tRFs, and tRNA halves of traditional Chinese medicines, which is suitable for predicting cell activity based on tRNA sequences of traditional Chinese medicines. Background Art
[0002] Traditional Chinese medicine is a treasure of mankind, but its specific mechanism of action and effective ingredients still need to be studied and clarified. It has been recently reported that non-coding RNA (ncRNA) such as microRNA has different regulatory effects by targeting different aspects of RNA transcription or post-transcriptional processes in almost all eukaryotic organisms. Lin Zhang et al. (Cell research 2012, 22, 107-126) proposed that exogenous plant microRNAs in food can be absorbed by the digestive tract of mammals and transported to various organs through the bloodstream, thereby regulating the expression of mammalian genes. The research group of Jiang Zhihong from the Macau University of Science and Technology found that tRNA derived from ginseng has a protective effect on cardiomyocytes (CN111419867B) and tRNA derived from yew has an inhibitory effect on ovarian cancer (Mol. Ther. Nucleic. 2022, 27, 718-732).
[0003] The efficacy of RNA derived from traditional Chinese medicines (TCMs) begins with verification of cell viability. Cell viability refers to the proportion of healthy cells in a sample population. Cell viability testing is essential for observing the physiological state of cells during experiments and is an important indicator for evaluating drug efficacy. Common cell viability assays include MTT and CCK-8. Dehydrogenases in living cells reduce WST-8 in exogenous MTT and CCK-8 reagents to water-insoluble blue-purple crystalline formazan and highly water-soluble yellow formazan dye, respectively, while dead cells lack this ability. The MTT assay is characterized by high sensitivity and cost-effectiveness. However, because formazan crystals are insoluble in water, they must be dissolved in dimethyl sulfoxide (DMSO) for detection. This results in a high workload and poor accuracy, and DMSO can also be harmful to cells and experimenters. In contrast, the CCK-8 assay is easy to use, highly sensitive, and less toxic. However, due to factors such as the cell growth cycle, testing large numbers of samples often requires long experimental cycles and high labor costs. Therefore, developing efficient methods for assessing drug activity is crucial.
[0004] With the development of information technology, domestic and international research has begun to attempt to use machine learning and RNA features to establish computational models to predict RNA drug activity. However, these efforts face the following challenges: 1) Activity predictions for diverse RNA sequences often have low accuracy; 2) Prediction results lack biological significance and cannot be systematically explained from the perspective of biological mechanisms; and 3) Traditional machine learning methods are unable to automatically learn feature information from large datasets, requiring extensive manual feature selection. Therefore, to address the limitations of these methods, the present invention provides a deep learning-based method for predicting the cellular activity of tRNAs, tRFs, and tRNA halves from traditional Chinese medicines. This method is capable of rapidly, efficiently, and systematically predicting the cellular activity of tRNAs, tRFs, and tRNA halves. Summary of the Invention
[0005] Technical problem solved: The present invention overcomes the shortcomings of the existing technology and discloses a method for predicting the cellular activity of traditional Chinese medicine tRNA, tRFs, and tRNA halves based on a deep learning algorithm (DL). This method fully combines the traditional Chinese medicine tRNA database with an artificial intelligence deep learning algorithm to establish an efficient and rapid method for predicting the cellular activity of traditional Chinese medicine tRNA, tRFs, and tRNA halves, providing an innovative idea for the research and development of new traditional Chinese medicine tRNA drugs.
[0006] Technical solution: An artificial intelligence-based method for predicting the activity of tRNA, tRFs, and tRNA halves in traditional Chinese medicines, comprising the following steps: Step 1: Extract and isolate tRNA from the target traditional Chinese medicine and sequence it; Step 2: Based on the tRNA sequencing results, design tRFs and tRNA halves sequences, and construct a modeling data sample based on traditional Chinese medicine tRNA; Step 3: Purify the tRNA of the target traditional Chinese medicine, determine the modification types and sites on the tRNA by LC-MS / MS method, and design modified tRFs and tRNA halves sequences. The modification is based on the natural tRNA modification types of traditional Chinese medicines, and the stability and efficacy of tRFs and tRNA halves are used as indicators to adjust and optimize the modification types and sites; Step 4: Perform cell activity experiments on the modified tRFs and tRNA halves, and the screening results are used as the final model sample features; Step 5: Construct a prediction model based on a deep learning algorithm, and use the screened tRFs and tRNA halves as feature sequences for model training and learning based on the presence or absence of cell activity; Step 6: Optimize the key parameters of the model through a grid search algorithm and cross-validation to improve the model prediction performance.
[0007] The above-mentioned step one includes the following steps: the target Chinese medicinal material is yew, ginseng, ganoderma lucidum, toad venom, Panax notoginseng, salvia miltiorrhiza or ginkgo; total RNA of the Chinese medicinal material is extracted, and small RNA with a length of less than 200mer is separated and enriched; total tRNA in the small RNA sample is separated by 6% polyacrylamide gel electrophoresis containing 8M urea, and after staining with SYBR nucleic acid dye, the gel separation bands are observed under blue light, and the gel area containing total tRNA is cut out; the cut gel is placed in a 3000MWCO dialysis bag, eluted by horizontal electrophoresis and the eluate is recovered; the total tRNA in the eluate is purified by TRIzol method; a total tRNA sample is taken, a cDNA library is established by using TruSeq Small RNA Sample Preparation Kit, and sequencing is performed by using Illumina second-generation sequencing platform; bioinformatics methods are used to perform quality control on the tRNA second-generation sequencing data, and the sequence and content of each tRNA are analyzed, and finally the sequence information of the tRNA contained in the target Chinese medicinal material is obtained.
[0008] The above-mentioned step 2 includes the following steps: according to the sequence information of tRNA contained in the target Chinese medicinal materials, tRFs and tRNA halves sequences are designed, including (1) 5'-t-half: refers to the sequence formed by the break of any part of the anticodon loop of any tRNA from the break point to the 5' end of tRNA; (2) 3'-t-half: refers to the sequence formed by the break of any part of the anticodon loop of any tRNA from the break point to the 3' end of tRNA; (3) 5'-tRF: refers to the sequence formed by the break of any tRNA from the 5' end to any position before the anticodon loop; (4) 3'-tRF: refers to the sequence formed by the break of any tRNA from the 3' end to any position before the anticodon loop; 20% of the modeling data samples are used as the test set, and the remaining data are used as the training set to construct the model.
[0009] The above-mentioned step three includes the following steps: designing a biotinylated DNA probe according to the tRNA sequence information of the target Chinese medicinal material, affinity capturing the tRNA of the target Chinese medicinal material by using the complementary pairing principle of the probe and the tRNA sequence of the target Chinese medicinal material, and finally recovering and purifying the biotinylated DNA probe-tRNA complex by using streptavidin-modified magnetic beads; restrictively hydrolyzing the purified tRNA monomer by using RNA restriction endonuclease or RNA exonuclease to form tRNA fragments; qualitatively analyzing the obtained tRNA fragments by using LC-MS / MS method; and 1 and MS 2The mass spectra were analyzed to confirm the tRNA modification types and sites of the target Chinese medicinal materials. Based on the LC-MS / MS modification sequence sequencing results, tRFs and tRNA halves containing modified nucleotides were designed to construct a modeling data sample based on Chinese medicinal material tRNA. The modification types included pseudouracil, dihydrouracil, base monomethylation modification, base dimethylation modification, base trimethylation modification, and nucleotide pentose 2'-oxymethylation modification.
[0010] The above-mentioned step four includes the following steps: selecting relevant cells according to the disease treated by the selected single Chinese medicine, and screening cell activity by CCK8; the cell activity experiment is, based on the correlation and difference analysis of the cell activity of the blank control group and the experimental group, for tumor cells, selecting IC50≤100nM or below 50nM, and the inhibition rate exceeding 50% band as the feature sequence; for normal cells, selecting 50nM, cell activity 100%±10%, and after modeling with the corresponding stimulation, the band below 50nM with significant difference from the modeling group is selected as the feature sequence; the sequence obtained by the cell experiment screening is converted into a feature vector as the input of deep learning, and the sequence is divided into at least K long base strings by the K-mers method, K≥12, and the frequency of these base strings is calculated as the feature representation of the sequence, and the classification information of the sequence: divided into active and inactive according to the above-mentioned cell experiment is incorporated into the model input, and the feature vector and the corresponding classification information are learned by the deep learning neural network to obtain a model that can predict the activity of unknown sequences.
[0011] The above step 5 includes: the model selects the Sequential model interface, and the model structure includes an input layer, a hidden layer, and an output layer, wherein the hidden layer includes a dense layer and a dropout layer; the input of the model is the feature vector of the sequence obtained in step 4 and the sequence activity classification information; the output of the model is a binary classification result of 0 and 1, where 0 represents inactivity, IC50>100nM, and 1 represents activity, IC50≤100nM; in the hidden layer of the model, through the Rectified Linear The ReLU (Reduced Lu) activation function is used to activate the input layer values, which are then passed to the fully connected layer. The activation function formula is: y = ReLU(Wx + b); where x is the input data value, y is the value after data activation, W is the weight matrix, and b is the bias. In the output layer of the model, the Sigmoid activation function is used to activate the hidden layer values, which are then passed as the final output result. The activation function formula is: z = sigmoid(W'y + b'); where y is the activated value passed from the hidden layer, z is the model output result, W' is the transposed weight matrix, and b' is the transposed bias. The output layer of the model is further connected to public biomedical databases, including ArrayExpress, Gene Expression Omnibus (GEO). During the model training process, the compile module is used to configure the model learning process. Its parameters are set as follows: the optimizer is set to Root Mean Squareprop (RMSprop), the metric list is set to accuracy, and the loss function is set to binary_crossentropy. The loss function is calculated as follows: Among them, L H (x,z) is the difference between the predicted value and the true value, that is, the loss, x is the true value corresponding to the sample, z is the predicted value corresponding to the sample, and d is the number of epochs.
[0012] The steps of optimizing the key parameters of the model through grid search algorithm and cross validation in step 6 above include: setting the parameter optimization range, where the epoch number is [10, 50, 100, 200, 500], the batch size is [10, 32, 64, 128], the learning rate is [0.01, 0.001, 0.00001], and the dropout rate is [0, 0.2, 0.5], and nodenumber is [50, 100, 300, 500, 1000]. The constructed 5×4×3×3×5 models are optimized by the grid search algorithm. The prediction performance of the model is evaluated by a 10-fold cross-validation mode and evaluation indicators, where the performance evaluation indicators include: sensitivity SEN; specificity SPE; accuracy ACC; Matthews correlation coefficient MCC; area under the ROC curve AUC. Among them, the closer the sensitivity, specificity and accuracy are to 100%, the closer the Matthews correlation coefficient and the area under the ROC curve are to 1, indicating that the prediction performance of the model is better. On the contrary, the closer the sensitivity, specificity and accuracy are to 0, the closer the Matthews correlation coefficient and the area under the ROC curve are to 0.5, indicating that the prediction performance of the model is worse.
[0013]
[0014] Among them, TP stands for true positive; TN stands for true negative; FP stands for false positive; and FN stands for false negative.
[0015] Beneficial effects: 1. This prediction method can accurately predict the cellular activity of different traditional Chinese medicine tRNAs and has good robustness; 2. The deep learning algorithm used in this method has a strong ability to automatically learn features and can automatically learn important feature information from big data, avoiding a large amount of manual feature selection; 3. The traditional Chinese medicine tRNA, tRFs, and tRNAhalves activity prediction model constructed by this method has excellent prediction performance and good applicability, providing an innovative idea for the research and development of new traditional Chinese medicine tRNA drugs. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is the overall flow chart of the traditional Chinese medicine tRNA prediction method based on deep learning of the present invention.
[0017] Figure 2 It is a schematic diagram of the structure of the activity prediction model of traditional Chinese medicine tRNA, tRFs and tRNA halves according to the method of the present invention.
[0018] Figure 3 This is a prediction performance chart of the traditional Chinese medicine tRNA prediction method based on deep learning. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solutions and advantages of the present invention more clear, the technical solutions of the present invention are further described below with reference to the accompanying drawings.
[0020] Example 1
[0021] A specific technical solution for an artificial intelligence-based method for predicting the activity of tRNA, tRFs, and tRNA halves in traditional Chinese medicine is:
[0022] By isolating the traditional Chinese medicinal herb Ganoderma lucidum, 157 RNA sequences were obtained. CCK8 was used to measure 1058 data points in human ovarian cancer cells (A2780), airway epithelial cells (16HBE), and airway smooth muscle cells (ASMC). All modeling sample data were randomly divided into training and test sets at an 80%:20% ratio.
[0023] The tRNA data obtained by sequencing were quality-controlled and compared to obtain tRNA sequence information. Statistical analysis was performed on the cells in the experimental and control groups through cell activity experiments. For A2780, bands with IC50 ≤ 100 nM were selected as feature sequences. For ASMC, bands with cell activity of 100% ± 10% within 50 nM and an inhibition rate of > 30% within 50 nM after modeling with TGF-β1 were selected as feature sequences. The sequences were further converted into feature vectors and divided into groups of 3 bases using K-mers. The frequencies of these base combinations were calculated as feature representations of the sequences. At the same time, the sequence activity classification information of the cell experiments was incorporated into the sequence features.
[0024] This solution uses Anaconda 5.1 and Python 3.6 to create a virtual environment. The Keras deep learning framework and the sklearn machine learning package are used to build a tRNA activity prediction model. The model uses the Sequential model interface to build a binary classification prediction model. The model structure includes an input layer, a hidden layer, and an output layer. The hidden layer includes a fully connected layer and a dropout layer ( Figure 2 In the hidden layer of the model, the value of the input layer is activated by the Rectified Linear Unit (ReLU) activation function and then passed to the fully connected layer. The formula of the activation function is:
[0025] y=ReLU(Wx+b)
[0026] Among them, x is the value of the input data, y is the value after the data is activated, W is the weight matrix, and b is the bias;
[0027] In the output layer of the model, the Sigmoid activation function is used to activate the value of the hidden layer and then output it as the final output result. The formula of the activation function is:
[0028] z=sigmoid(W'y+b')
[0029] Among them, y is the activated value transmitted from the hidden layer, z is the model output result, W' is the transposed weight matrix, and b' is the transposed bias;
[0030] The output layer of the model is further connected to public biomedical databases, including Array Express, GeneExpression, and Omnibus (GEO).
[0031] During the model training process, the compile module is used to configure the model's learning process. Its parameters are set as follows: the optimizer is set to Root Mean Square prop (RMSprop), the metrics list is set to accuracy, and the loss function is set to binary_crossentropy. The loss function is calculated as follows:
[0032]
[0033] Among them, L H (x,z) is the difference between the predicted value and the true value (i.e., loss), x is the true value corresponding to the sample, z is the predicted value corresponding to the sample, and d is the number of epochs;
[0034] The feature vector data of the sequence finally obtained after processing is then used as the input of the model, 80% of which is used as a training set to train the model, and 20% is used as a test set to test the model performance.
[0035] The method for predicting the activity of tRNA, tRFs, and tRNA halves of traditional Chinese medicine based on artificial intelligence according to claim 1, wherein step six comprises the following steps:
[0036] Set the parameter optimization range, where epoch number is [10, 50, 100, 200, 500], batch size is [10, 32, 64, 128], learning rate is [0.01, 0.001, 0.00001], dropout rate is [0, 0.2, 0.5], and node number is [50, 100, 300, 500, 1000];
[0037] The predictive performance of the trained models was evaluated using a test set of samples from the modeling data. Performance was evaluated using metrics such as sensitivity (SEN), specificity (SPE), accuracy (ACC), Matthews correlation coefficient (MCC), and the area under the receiver operating characteristic (ROC) curve (AUC). Furthermore, a grid search algorithm and ten-fold cross-validation were used to optimize the key parameters (epoch number, batch size, learning rate, dropout rate, and node number) of the 900 (5×4×3×3×5) models constructed to achieve optimal predictive performance.
[0038]
[0039] Among them, TP stands for true positive; TN stands for true negative; FP stands for false positive; and FN stands for false negative.
[0040] Finally, through parameter optimization, the specific parameters of the optimal model were set as 2 hidden layers with 50 nodes per layer, a dropout rate of 0.5 to avoid overfitting, a learning rate of 0.001, a batch size of 128, and an epoch number of 50. The performance of the optimal model was evaluated on the test set, with a prediction accuracy of 97.1%, an AUC of 0.989, a sensitivity of 97.4%, a specificity of 96.8%, and a Matthews correlation coefficient of 0.942. Compared with most RNA prediction models based on traditional machine learning at home and abroad, this model has better prediction performance ( Figure 3 , Table 1).
[0041] The above example is only a specific embodiment of the present invention, and simple changes and replacements thereof are also within the scope of protection of the invention.
[0042] Table 1 Prediction results of traditional Chinese medicine tRNA prediction method based on deep learning
[0043] Confusion Matrix Actual: Active Actual: Inactive Prediction: Active 185 32 Prediction: Inactive 5 1058
Claims
1. An artificial intelligence-based method for predicting the activity of tRNA, tRFs, and tRNA halves in traditional Chinese medicine, characterized in that: The following steps are involved: Step 1: Extract and separate the tRNA of the target Chinese medicinal material and sequence it; Step 2: Based on the tRNA sequencing results, design tRFs and tRNA halves sequences and construct modeling data samples based on traditional Chinese medicine tRNA; Step 3: Purify the target TCM tRNA, determine the modification types and sites on the tRNA by LC-MS / MS, and design modified tRFs and tRNA halves. The modification is based on the natural tRNA modification types of the TCM, and the stability and efficacy of tRFs and tRNA halves are used as indicators to adjust and optimize the modification types and sites. Step 4: Conduct cell activity experiments on the modified tRFs and tRNA halves, and use the screening results as the final model sample characteristics; Step 5: Build a prediction model based on a deep learning algorithm, using the screened tRFs and tRNA halves as feature sequences based on whether they have cell activity for model training and learning; Step 6: Optimize the key parameters of the model through grid search algorithm and cross-validation to improve the model's prediction performance.
2. The method for predicting the activity of tRNA, tRFs, and tRNA halves of traditional Chinese medicine based on artificial intelligence according to claim 1, characterized in that: The step one comprises the following steps: the target Chinese medicinal material is yew, ginseng, ganoderma lucidum, toad venom, Panax notoginseng, salvia miltiorrhiza or ginkgo biloba; total RNA of the Chinese medicinal material is extracted, and small RNA with a length of less than 200mer is separated and enriched; total tRNA in the small RNA sample is separated by 6% polyacrylamide gel electrophoresis containing 8M urea, the gel separation bands are observed under blue light after staining with SYBR nucleic acid dye, and the gel area containing total tRNA is cut out; the cut gel is placed in a 3000MWCO dialysis bag, eluted by horizontal electrophoresis and the eluate is recovered; total tRNA in the eluate is purified by TRIzol method; a total tRNA sample is taken, a cDNA library is established by using TruSeq Small RNA Sample Preparation Kit, and sequencing is performed by using Illumina second-generation sequencing platform; bioinformatics methods are used to perform quality control on the tRNA second-generation sequencing data, and the sequence and content of each tRNA are analyzed, so as to finally obtain the sequence information of the tRNA contained in the target Chinese medicinal material.
3. The method for predicting the activity of tRNA, tRFs, and tRNA halves of traditional Chinese medicine based on artificial intelligence according to claim 1, characterized in that: The second step includes the following steps: designing tRFs and tRNA halves sequences based on the sequence information of tRNA contained in the target Chinese medicinal materials, including (1) 5'-t-half: refers to the sequence formed by the break of any part of the anticodon loop of any tRNA, from the breakpoint to the 5' end of tRNA; (2) 3'-t-half: refers to the sequence formed by the break of any part of the anticodon loop of any tRNA, from the breakpoint to the 3' end of tRNA; (3) 5'-tRF: refers to the sequence formed by the break from the 5' end of any tRNA to any position before the anticodon loop; (4) 3'-tRF: refers to the sequence formed by the break from the 3' end of any tRNA to any position before the anticodon loop; 20% of the modeling data samples are used as a test set, and the remaining data are used as a training set to construct the model.
4. The method for predicting the activity of tRNA, tRFs, and tRNA halves of traditional Chinese medicine based on artificial intelligence according to claim 1, characterized in that: The step three comprises the following steps: designing a biotinylated DNA probe according to the tRNA sequence information of the target Chinese medicinal material, affinity capturing the tRNA of the target Chinese medicinal material by using the principle of complementary pairing between the probe and the tRNA sequence of the target Chinese medicinal material, and finally recovering and purifying the biotinylated DNA probe-tRNA complex by using streptavidin-modified magnetic beads; restrictively hydrolyzing the purified tRNA monomer by using RNA restriction endonuclease or RNA exonuclease to form tRNA fragments; qualitatively analyzing the obtained tRNA fragments by using LC-MS / MS method; and 1 and MS 2 Mass spectra are analyzed to confirm the tRNA modification type and site of the target Chinese medicinal materials; Based on the LC-MS / MS modified sequence sequencing results, tRFs and tRNA halves containing modified nucleotides were designed to construct a modeling data sample based on traditional Chinese medicine tRNA; the modification types include pseudouracil, dihydrouracil, base monomethylation modification, base dimethylation modification, base trimethylation modification and nucleotide pentose 2'-oxymethylation modification.
5. The method for predicting the activity of tRNA, tRFs, and tRNA halves of traditional Chinese medicine based on artificial intelligence according to claim 1, characterized in that: The step four includes the following steps: selecting relevant cells according to the disease treated by the selected single Chinese medicine, and screening the cell activity by CCK8; the cell activity experiment is, based on the correlation and difference analysis of the cell activity between the blank control group and the experimental group, selecting IC50≤100nM or below 50nM for tumor cells, and the inhibition rate exceeding 50% as the feature sequence; for normal cells, selecting 50nM, cell activity 100%±10%, and after modeling with the corresponding stimulation, the band below 50nM with significant difference from the modeling group as the feature sequence; converting the sequence obtained by the cell experiment screening into a feature vector as the input of deep learning, dividing the sequence into at least K long base strings by the K-mers method, K≥12, and calculating the frequency of these base strings as the feature representation of the sequence, and at the same time incorporating the classification information of the sequence: divided into active and inactive according to the above-mentioned cell experiment into the model input, and learning the feature vector and the corresponding classification information by the deep learning neural network to obtain a model that can predict the activity of unknown sequences.
6. The method for predicting the activity of tRNA, tRFs, and tRNA halves of traditional Chinese medicine based on artificial intelligence according to claim 1, characterized in that: The above step 5 includes: selecting the Sequential model interface as the model, and the model structure includes an input layer, a hidden layer, and an output layer, where the hidden layer includes a dense layer and a dropout layer; the input of the model is the feature vector of the sequence obtained in step 4 and the sequence activity classification information; the output of the model is a binary classification result of 0 and 1, where 0 represents inactivity and IC50>100nM, and 1 represents activity and IC50≤100nM; in the hidden layer of the model, the value of the input layer is activated by the Rectified Linear Unit (ReLU) activation function and then passed to the fully connected layer. The formula of the activation function is: y=ReLU(Wx+b) Among them, x is the value of the input data, y is the value after the data is activated, W is the weight matrix, and b is the bias; In the output layer of the model, the Sigmoid activation function is used to activate the value of the hidden layer and then output it as the final output result. The formula of the activation function is: z=sigmoid(W'y+b') Among them, y is the activated value transmitted from the hidden layer, z is the model output result, W' is the transposed weight matrix, and b' is the transposed bias; The output layer of the model is further connected to public biomedical databases, including Array Express, GeneExpression, and Omnibus (GEO); During the model training process, the compile module is used to configure the model's learning process. Its parameters are set as follows: the optimizer is set to Root Mean Square prop (RMSprop), the metrics list is set to accuracy, and the loss function is set to binary_crossentropy. The loss function is calculated as follows: Among them, L H (x,z) is the difference between the predicted value and the true value, that is, the loss; x is the true value corresponding to the sample, z is the predicted value corresponding to the sample, and d is the number of epochs.
7. The method for predicting the activity of tRNA, tRFs, and tRNA halves of traditional Chinese medicine based on artificial intelligence according to claim 1, characterized in that: Step 6: Optimizing the key parameters of the model through grid search algorithm and cross validation includes: setting the parameter optimization range, where the epoch number is [10, 50, 100, 200, 500], the batch size is [10, 32, 64, 128], the learning rate is [0.01, 0.001, 0.00001], and the dropout rate is [0, 0.2, 0.5], and nodenumber is [50, 100, 300, 500, 1000]. The constructed 5×4×3×3×5 models are optimized by the grid search algorithm. The prediction performance of the model is evaluated by a 10-fold cross-validation mode and evaluation indicators, where the performance evaluation indicators include: sensitivity SEN; specificity SPE; accuracy ACC; Matthews correlation coefficient MCC; area under the ROC curve AUC. Among them, the closer the sensitivity, specificity and accuracy are to 100%, the closer the Matthews correlation coefficient and the area under the ROC curve are to 1, indicating that the prediction performance of the model is better. On the contrary, the closer the sensitivity, specificity and accuracy are to 0, the closer the Matthews correlation coefficient and the area under the ROC curve are to 0.5, indicating that the prediction performance of the model is worse. Among them, TP stands for true positive; TN stands for true negative; FP stands for false positive; and FN stands for false negative.
Citation Information
Patent Citations
Uses of transfer RNA molecules and fragments for the prevention or treatment of heart disease
CN111419867B