Encoder-based gradient boosting machine miRNA-disease association prediction method

Through the encoder-based gradient hoist method, combined with multi-source data fusion and lightweight gradient hoist classifier, the cost, time-consuming and overfitting problems in miRNA-disease association prediction are solved, and efficient and accurate miRNA-disease association prediction is achieved.

CN117198407BActive Publication Date: 2025-08-22HENAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311235759.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-22
Publication Date
2025-08-22
Estimated Expiration
2043-09-22

AI Technical Summary

Technical Problem

The prior art has problems such as high cost, long-term consumption, difficulty in high-precision prediction for imbalanced data sets or small samples, overfitting and low computational efficiency in miRNA-disease association prediction.

Method used

The encoder-based gradient hoist method is adopted to obtain miRNA-disease association data through multi-source data fusion, and the disease semantic similarity matrix and miRNA functional similarity matrix are constructed, the Gaussian interaction spectral core similarity matrix is ​​integrated, and the low-dimensional high-quality feature vector is extracted using an automatic encoder, and the lightweight gradient hoist classifier is input for prediction.

Benefits of technology

It reduces the prediction cost and time period, improves the calculation efficiency, reduces the overfitting problem, improves the prediction effect, and achieves efficient prediction of miRNA-disease associations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117198407B_ABST
    Figure CN117198407B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of bioinformatics association prediction technology, and specifically to an encoder-based gradient boosting machine miRNA-disease association prediction method. The method comprises: utilizing multi-source biological and medical information to obtain a miRNA-disease association adjacency matrix, a miRNA functional similarity matrix, and a disease semantic similarity matrix; concatenating the resulting integrated miRNA-disease similarity matrix with the miRNA-disease association adjacency matrix to obtain a more informative miRNA-disease association feature vector; utilizing an autoencoder to extract key features of the combined miRNA feature vector and disease feature vector, and utilizing a lightweight gradient boosting machine classifier to predict potential associations between miRNAs and diseases. The present invention achieves high prediction accuracy with low time and economic costs, aims to uncover potential miRNA-disease associations, and can aid in the study of the pathogenesis of complex diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bioinformatics association prediction, and in particular to an encoder-based gradient boosting machine miRNA-disease association prediction method. Background Art

[0002] As a 22-nucleotide non-coding single-stranded RNA (ribonucleic acid) molecule, miRNA plays a key role in nearly all biological processes, including cell proliferation, metabolism, and immune response. Therefore, miRNA disorders may lead to a variety of complex diseases. For example, overexpression of hsa-mir-449a in CL1-0 (human lung adenocarcinoma cells) aggravates damage and apoptosis in irradiated cells, thereby altering cell cycle distribution. Furthermore, hsa-mir-195 and hsa-mir-497 have been shown to have key inhibitory effects on breast malignancies. Therefore, using bioinformatics to discover the association between miRNAs and diseases may help in the prevention, diagnosis, and treatment of diseases.

[0003] To date, numerous biological experiments have been conducted to uncover associations between miRNAs and diseases. These associations have been used to establish publicly available online databases such as dbDEMC, HMDD3.0, and miR2Disease. While traditional biological experimental methods have the potential to uncover miRNA-disease associations, they still face challenges such as high cost and time consumption that need to be addressed.

[0004] Introducing high-performance computing and artificial intelligence into the field of miRNA-disease association prediction may be a reasonable approach to addressing the aforementioned issues. To date, AI methods related to miRNA-disease association prediction primarily fall into three categories: graph theory, traditional machine learning, and deep learning. These methods tend to be used to process a wide range of biological data, including miRNA expression profiles, miRNA sequences, protein sequences, and human phenotypic ontologies. Because deep learning techniques can better learn data representations, they have been increasingly applied to various fields, including genomics and drug discovery. For example, a multi-view multi-channel attention graph convolutional network (MMGCN) approach was proposed, in which similarity matrices containing different information are weighted to infer potential miRNA-disease associations. Compared to traditional machine learning and graph theory methods, deep learning methods can improve the accuracy of miRNA-disease association prediction. However, challenges remain, such as high-precision prediction for imbalanced datasets or small sample sizes and complex hyperparameter tuning. Recently, some progress has been made in high-precision prediction methods for imbalanced datasets. A miRNA-disease association prediction method based on autoencoders was proposed, and deep random forests were used for ensemble learning. In addition, a scalable tree enhancement method based on autoencoders was proposed to infer small molecule-miRNA associations, which further improved the prediction efficiency. However, the above methods may also face the problems of overfitting and low computational efficiency. Summary of the Invention

[0005] The Summary of the Present Invention is intended to briefly introduce concepts that will be described in detail in the Detailed Description of the Present Invention. The Summary of the Present Invention is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0006] In order to solve the technical problem of poor prediction effect, the present invention proposes an encoder-based gradient boosting machine miRNA-disease association prediction method.

[0007] The present invention provides a gradient boosting machine miRNA-disease association prediction method based on an encoder, the method comprising:

[0008] Step 1: Obtain miRNA-disease association data, disease autocorrelation data, and miRNA autocorrelation data from a preset number of external biological and medical data sources;

[0009] Step 2: Calculate the disease semantic similarity matrix based on the data obtained in step 1;

[0010] Step 3, based on the data obtained in step 1, calculate the miRNA functional similarity matrix;

[0011] Step 4: obtaining the miRNA-disease association adjacency matrix based on the miRNA-disease association data obtained in step 1, and calculating the Gaussian interaction spectrum kernel similarity matrix of miRNA and disease based on the miRNA-disease association adjacency matrix;

[0012] Step 5: Integrate the disease semantic similarity matrix obtained in step 2 and the miRNA functional similarity matrix obtained in step 3 with the miRNA and disease Gaussian interaction spectrum kernel similarity matrix calculated in step 4 to obtain an integrated similarity matrix of miRNA and disease;

[0013] Step 6, the integrated similarity matrix of miRNAs and diseases obtained in step 5 is spliced ​​with the miRNA-disease association adjacency matrix to obtain the comprehensive similarity matrix of miRNAs and diseases;

[0014] Step 7: Input the feature vector of the miRNA and disease comprehensive similarity matrix obtained in step 6 into the autoencoder for feature extraction to extract low-dimensional and high-quality miRNA-disease association feature vectors;

[0015] In step 8, the low-dimensional, high-quality miRNA-disease association feature vector obtained in step 7 is input into a lightweight gradient boosting machine classifier for miRNA-disease association prediction.

[0016] Optionally, the HMDD database is obtained to obtain miRNA-disease association data verified by biological experiments, thereby obtaining a miRNA-disease association adjacency matrix; a disease semantic similarity matrix is ​​obtained based on the MeSH database; and a miRNA functional similarity matrix is ​​obtained based on the gene ontology database.

[0017] Optionally, the calculation to obtain a disease semantic similarity matrix includes:

[0018] Calculate ancestral disease t versus disease d i The formula corresponding to the semantic value contribution of is:

[0019]

[0020] Where △ is the semantic contribution factor set to 0.5, t′ is one of the ancestral diseases, is one of the ancestral diseases t′ to disease d i The semantic value contribution of Is the ancestral disease t to the disease d i The semantic contribution of disease d to itself is set to 1. Combined with the contribution value of the ancestral disease in the directed acyclic graph DAG(d), disease d i The formula corresponding to the semantic value of is:

[0021]

[0022] Among them, T(d i ) is related to disease d i A collection of related ancestral diseases, disease d i and d j The formula for the semantic similarity is:

[0023]

[0024] Another way to calculate the effect of disease t on disease d in DAG i The formula corresponding to the contribution of semantic value is:

[0025]

[0026] Diseased i The formula corresponding to the semantic similarity value is:

[0027]

[0028] Diseased i and disease j The formula corresponding to the semantic similarity value between is:

[0029]

[0030] By integrating the two disease semantic similarity values, the formula corresponding to the disease semantic similarity matrix DS is proposed as follows:

[0031]

[0032] Optionally, the Gaussian interaction spectrum kernel similarity matrix of miRNA and disease is calculated based on the miRNA-disease association adjacency matrix, including:

[0033] Uncovering disease d through Gaussian interaction spectral kernel similarity i and disease j The relationship between IP(d i ) indicates disease d i Validated associations between each miRNA, IP(d j ), nd is the number of diseases, γ d It is a parameter used to adjust the kernel bandwidth. The formula corresponding to the Gaussian interaction spectrum kernel similarity matrix KD between diseases is:

[0034] KD(d i ,d j )=exp(-γ d PIP(d i)-IP(d j )P 2 )

[0035]

[0036] IP(m i ) indicates miRNAm i Known associations with each disease, IP(m j ) is similar to it, nm is the number of miRNA, γ m It describes the parameters used to adjust the kernel bandwidth. The formula corresponding to the Gaussian interaction spectrum kernel similarity matrix KM between miRNAs is:

[0037] KM(m i ,m j )=exp(-γ m PIP(m i )-IP(m j )P 2 )

[0038]

[0039] Optionally, the disease semantic similarity matrix obtained in step 2 and the miRNA functional similarity matrix obtained in step 3 are respectively integrated with the Gaussian interaction spectrum kernel similarity matrix of miRNA and disease calculated in step 4 to obtain an integrated similarity matrix of miRNA and disease, including:

[0040] By converting the miRNA Gaussian interaction profile kernel similarity matrix KM (m i ,m j ) and miRNA functional similarity matrix FS(m i ,m j ) are integrated together to obtain the integrated similarity matrix SM of miRNA:

[0041]

[0042] The Gaussian interaction spectrum kernel similarity matrix KD(d i ,d j ) and disease semantic similarity matrix DS(d i ,d j ) are integrated together to obtain the disease integrated similarity matrix SD corresponding to the formula:

[0043]

[0044] Optionally, the obtained miRNA and disease comprehensive similarity matrices are:

[0045]

[0046] S disease =(D1A1, L, D1A 495 , L, D 383 A1, L, D 383 A 495 ) T

[0047] Among them, M i (M i1 ,M i2 ,...,M i495 ), D i (D i1 ,D i2 ,...,D i383 ), A i , They represent the integrated similarity between the i-th miRNA and other miRNAs, the integrated similarity between the i-th disease and other diseases, the i-th row of the verified miRNA-disease association matrix, and the j-th column transposed of the verified miRNA-disease association matrix.

[0048] Optionally, the step of inputting the feature vector of the miRNA and disease comprehensive similarity matrix obtained in step 6 into an autoencoder for feature extraction to extract a low-dimensional, high-quality miRNA-disease association feature vector includes:

[0049] Use the self-encoder to encode the 495×383 rows and 495+495 columns of S miRNA Matrix S with 495×383 rows and 383+383 columns disease The matrix is ​​processed to extract their low-dimensional features, while reducing the noise caused by the redundant information hidden in the original feature vector, and two low-dimensional high-quality miRNA-disease association feature vectors are obtained.

[0050] Optionally, the step of inputting the low-dimensional, high-quality miRNA-disease association feature vector obtained in step 7 into a lightweight gradient boosting machine classifier for miRNA-disease association prediction includes:

[0051] The two types of low-dimensional, high-quality miRNA-disease association feature vectors obtained are input into the lightweight gradient boosting machine classifier for prediction. A prediction result is obtained from the perspective of disease and miRNA respectively, and the two prediction results are integrated to obtain the final miRNA-disease association prediction value.

[0052] The present invention has the following beneficial effects:

[0053] The present invention uses a multi-source data fusion method, where the data is derived from biological and medical information, resulting in a richer information content. Furthermore, when constructing the original similarity matrix between miRNAs and diseases, the integrated similarity matrix between miRNAs and diseases is spliced ​​with the miRNA-disease association adjacency matrix, thereby enriching the information content of the constructed similarity matrix. Compared with traditional biological experimental methods, the present invention is less expensive and time-consuming. Compared with other prediction methods, the present invention improves computational efficiency and reduces model overfitting. Therefore, the present invention proposes a lightweight gradient boosting machine miRNA-disease association prediction method based on an autoencoder, which improves the classification method of the prediction model, effectively avoiding the high cost and long time period of traditional biological experimental methods, improving the prediction effect, and having important theoretical significance and application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0055] Figure 1 Flowchart of the encoder-based gradient boosting machine miRNA-disease association prediction method of the present invention;

[0056] Figure 2 is another flow chart of the present invention;

[0057] Figure 3 Schematic diagram of the comparison of the ROC curves of the present invention and other prediction methods on the HMDDv2.0 imbalanced dataset using 5-fold cross validation;

[0058] Figure 4 Schematic diagram of the comparison of the 5-fold cross-validation ROC curves of the present invention and other different classifiers on the HMDDv2.0 imbalanced dataset. DETAILED DESCRIPTION

[0059] To further illustrate the technical means and effects employed by the present invention to achieve its intended objectives, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail the specific implementations, structures, features, and effects of the technical solutions proposed by the present invention. In the following description, references to "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.

[0060] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0061] The present invention provides a gradient boosting machine miRNA-disease association prediction method based on an encoder, the method comprising the following steps:

[0062] Step 1: Obtain miRNA-disease association data, disease autocorrelation data, and miRNA autocorrelation data from a preset number of external biological and medical data sources;

[0063] Step 2: Calculate the disease semantic similarity matrix based on the data obtained in step 1;

[0064] Step 3, based on the data obtained in step 1, calculate the miRNA functional similarity matrix;

[0065] Step 4: obtaining the miRNA-disease association adjacency matrix based on the miRNA-disease association data obtained in step 1, and calculating the Gaussian interaction spectrum kernel similarity matrix of miRNA and disease based on the miRNA-disease association adjacency matrix;

[0066] Step 5: Integrate the disease semantic similarity matrix obtained in step 2 and the miRNA functional similarity matrix obtained in step 3 with the miRNA and disease Gaussian interaction spectrum kernel similarity matrix calculated in step 4 to obtain an integrated similarity matrix of miRNA and disease;

[0067] Step 6, the integrated similarity matrix of miRNAs and diseases obtained in step 5 is spliced ​​with the miRNA-disease association adjacency matrix to obtain the comprehensive similarity matrix of miRNAs and diseases;

[0068] Step 7: Input the feature vector of the miRNA and disease comprehensive similarity matrix obtained in step 6 into the autoencoder for feature extraction to extract low-dimensional and high-quality miRNA-disease association feature vectors;

[0069] In step 8, the low-dimensional, high-quality miRNA-disease association feature vector obtained in step 7 is input into a lightweight gradient boosting machine classifier for miRNA-disease association prediction.

[0070] The following is a detailed explanation of each of the above steps:

[0071] refer to Figure 1 , shows the process of some embodiments of the encoder-based gradient boosting machine miRNA-disease association prediction method according to the present invention. The encoder-based gradient boosting machine miRNA-disease association prediction method includes the following steps:

[0072] Step 1: Obtain miRNA-disease association data, disease autocorrelation data, and miRNA autocorrelation data from a preset number of external biological and medical data sources.

[0073] In some embodiments, miRNA-disease association data, disease autocorrelation data, and miRNA autocorrelation data can be obtained from multiple external biological and medical data sources.

[0074] The preset number may be a pre-set number. For example, the preset number may be 3. The miRNA-disease association data is also called the miRNA-disease association matrix.

[0075] Obtaining a validated miRNA-disease association matrix may include: obtaining a dataset including 495 miRNAs, 383 diseases, and 5430 experimentally confirmed miRNA-disease associations, here collected from HMDDv2.0. nm×nd The adjacency matrix related to miRNA and disease is represented by 495 rows and 383 columns. If the correlation between the i-th miRNA and the j-th disease is known, then A(m i ,d j ) is set to 1, otherwise it is set to 0.

[0076] Figure 1 The specific steps to achieve this can be as follows Figure 2 shown.

[0077] Step 2: Based on the data obtained in step 1, the disease semantic similarity matrix is ​​calculated.

[0078] In some embodiments, a disease semantic similarity matrix can be calculated based on the information data obtained in step 1.

[0079] As an example, this step may include the following steps:

[0080] In the first step, we use the directed acyclic graph DAG(d)=(d,T d ,E d ) to describe the relationship between disease d and other diseases, where T d is the ancestral disease set associated with disease d, E d is the edge set related to disease d. Therefore, all directed acyclic graphs (DAGs) can be used together to construct disease semantic similarity networks. In addition, the relationships between diseases used to construct these networks can be obtained from the MeSH database, and the relationship between ancestral disease t and disease d is calculated. i The formula corresponding to the semantic value contribution of is:

[0081]

[0082] Where △ is the semantic contribution factor set to 0.5, t′ is one of the ancestral diseases, is one of the ancestral diseases t′ to disease d i The semantic value contribution of Is the ancestral disease t to the disease d i The semantic contribution of disease d to itself is set to 1. The contribution of a disease decreases as the distance from other related diseases increases. Combined with the contribution of its ancestral disease in the directed acyclic graph DAG(d), disease d i The formula corresponding to the semantic value of is:

[0083]

[0084] Among them, T(d i ) is related to disease d i A collection of related ancestral diseases. Traditionally, the more DAGs a pair of diseases share, the greater the similarity between them. Therefore, disease d i and d j The formula for the semantic similarity is:

[0085]

[0086] In the second step, in the disease semantic similarity network, a disease with fewer DAGs tends to have a higher probability of disease d than another disease with more DAGs. i The contribution of semantic value is greater, so a new model is introduced here to calculate the effect of disease t on disease d in DAG. i The contribution of semantic value is as follows:

[0087]

[0088] In the third step, compared with the first method of calculating the semantic value of the disease, the disease d i The formula corresponding to the semantic similarity value is:

[0089]

[0090] Step 4, disease d i and disease j The formula corresponding to the semantic similarity value between is:

[0091]

[0092] In the fifth step, using only one disease semantic similarity calculation method may not reveal the semantic similarity between diseases. Therefore, the two disease semantic similarity values ​​are integrated together, and the formula corresponding to the disease semantic similarity matrix DS is proposed as follows:

[0093]

[0094] The semantic similarity between all diseases is calculated to obtain a disease semantic similarity matrix.

[0095] Step 3: Based on the data obtained in step 1, the functional similarity matrix of miRNA is calculated.

[0096] In some embodiments, the miRNA functional similarity matrix can be calculated based on the information data obtained in step 1.

[0097] As an example, a disease semantic similarity matrix is ​​obtained based on the MeSH database; a miRNA functional similarity matrix is ​​obtained based on the gene ontology database.

[0098] For example, a method for calculating miRNA functional similarity was proposed by assuming that diseases with similar phenotypes may be associated with miRNAs with similar functions. Following this method, the miRNA functional similarity matrix FS was proposed, where FS(m i ,m j ) characterizes the miRNA functional similarity score between the i-th miRNA and the j-th miRNA.

[0099] Step 4: Based on the miRNA-disease association data obtained in step 1, a miRNA-disease association adjacency matrix is ​​obtained, and based on the miRNA-disease association adjacency matrix, the Gaussian interaction spectrum kernel similarity matrix of miRNA and disease is calculated respectively.

[0100] In some embodiments, Gaussian interaction spectrum kernel similarity matrices of miRNAs and diseases can be calculated based on the obtained miRNA-disease association adjacency matrix.

[0101] As an example, this step may include the following steps:

[0102] The first step is to obtain the HMDD database, from which the miRNA-disease association data verified by traditional biological experiments are obtained, thereby obtaining the miRNA-disease association adjacency matrix.

[0103] In the second step, we assume that changes in miRNAs with similar functions may induce some similar diseases, and reveal the disease d by Gaussian interaction spectrum kernel similarity. i and disease j The relationship between IP(d i ) indicates disease d i Validated associations between each miRNA, IP(d j ), nd is the number of diseases, γ dIt is a parameter used to adjust the kernel bandwidth. The formula corresponding to the Gaussian interaction spectrum kernel similarity matrix KD between diseases is:

[0104] KD(d i ,d j )=exp(-γ d PIP(d i )-IP(d j )P 2 )

[0105]

[0106] Similar to the calculation of Gaussian interaction spectrum kernel similarity between diseases, the Gaussian interaction spectrum kernel similarity between miRNAs is defined as follows, where IP(m i ) indicates miRNAm i Known associations with each disease, IP(m j ) is similar to it, nm is the number of miRNA, γ m It describes the parameters used to adjust the kernel bandwidth. The formula corresponding to the Gaussian interaction spectrum kernel similarity matrix KM between miRNAs is:

[0107] KM(m i ,m j )=exp(-γ m PIP(m i )-IP(m j )P 2 )

[0108]

[0109] In step 5, the disease semantic similarity matrix obtained in step 2 and the miRNA functional similarity matrix obtained in step 3 are respectively integrated with the Gaussian interaction spectrum kernel similarity matrix of miRNA and disease calculated in step 4 to obtain an integrated similarity matrix of miRNA and disease.

[0110] In some embodiments, the disease semantic similarity matrix obtained in step 2 and the miRNA functional similarity matrix obtained in step 3 can be integrated with the disease and miRNA Gaussian interaction spectrum kernel similarity matrix calculated by the miRNA-disease association adjacency matrix.

[0111] As an example, this step may include the following steps:

[0112] In the first step, the Gaussian interaction profile kernel similarity matrix KM (m i ,m j ) and miRNA functional similarity matrix FS(mi ,m j ) are integrated together to obtain the integrated similarity matrix SM of miRNA:

[0113]

[0114] The second step is similar to miRNA, and the Gaussian interaction spectrum kernel similarity matrix KD (d i ,d j ) and disease semantic similarity matrix DS(d i ,d j ) are integrated together to obtain the disease integrated similarity matrix SD corresponding to the formula:

[0115]

[0116] In step 6, the integrated similarity matrix of miRNA and disease obtained in step 5 is spliced ​​with the miRNA-disease association adjacency matrix to obtain the comprehensive similarity matrix of miRNA and disease.

[0117] In some embodiments, the miRNA and disease integrated similarity matrix obtained in step 5 can be spliced ​​with the miRNA-disease association adjacency matrix to obtain a comprehensive similarity matrix of miRNA and disease.

[0118] As an example, the obtained miRNA and disease comprehensive similarity matrices are:

[0119]

[0120] S disease= (D1A1, L, D1A 495 , L, D 383 A1, L, D 383 A 495 ) T

[0121] Among them, M i (M i1 ,M i2 ,...,M i495 ), D i (D i1 ,D i2 ,...,D i383 ), A i , They represent the integrated similarity between the i-th miRNA and other miRNAs, the integrated similarity between the i-th disease and other diseases, the i-th row of the verified miRNA-disease association matrix, and the j-th column transposed of the verified miRNA-disease association matrix.

[0122] In step 7, the feature vector of the miRNA and disease comprehensive similarity matrix obtained in step 6 is input into the autoencoder for feature extraction to extract low-dimensional and high-quality miRNA-disease association feature vectors.

[0123] In some embodiments, an autoencoder may be used to process high-dimensional miRNA-disease association feature vectors to reduce the feature vector dimension and reduce redundant information.

[0124] As an example, we can use the autoencoder to encode 495×383 rows and 495+495 columns of S miRNA Matrix S with 495×383 rows and 383+383 columns disease The matrix is ​​processed to extract their low-dimensional features (also known as important features), while reducing the noise caused by the redundant information hidden in the original feature vector, and obtaining two low-dimensional high-quality miRNA-disease association feature vectors.

[0125] In step 8, the low-dimensional, high-quality miRNA-disease association feature vector obtained in step 7 is input into a lightweight gradient boosting machine classifier for miRNA-disease association prediction.

[0126] In some embodiments, the low-dimensional, high-quality miRNA-disease association feature vector obtained in step 7 can be fed into a lightweight gradient boosting machine classifier to perform miRNA-disease association prediction.

[0127] As an example, the two low-dimensional, high-quality miRNA-disease association feature vectors obtained can be input into a lightweight gradient boosting machine classifier for prediction. A prediction result is obtained from the disease and miRNA perspectives respectively, and the two prediction results are integrated to obtain the final miRNA-disease association prediction value. The details of the lightweight gradient boosting machine classifier training are listed in Algorithm 1 included in Table 1:

[0128] Table 1

[0129]

[0130]

[0131] In the final step of the miRNA-disease association prediction method, two low-dimensional, high-quality feature vectors of miRNA-disease associations are used to iteratively train a lightweight gradient boosting machine classifier related to miRNA-disease association prediction. The lightweight gradient boosting machine classifier algorithm reduces the number of samples based on unilateral gradient sampling, and features by bundling them with dedicated features, thereby reducing computational cost and improving efficiency. Furthermore, the lightweight gradient boosting machine classifier imposes a depth limit on the direction of leaf node splitting to ensure high efficiency while preventing overfitting.

[0132] The potential association between miRNA and disease was predicted according to the proposed method. The performance of the proposed prediction method (LGBMDA) was evaluated by conducting a 5-fold cross-validation experiment. The ROC curve describing the relationship between the false positive rate (FPR) and the true positive rate (TPR) was used to evaluate the performance of the association prediction method. The area under the ROC curve is AUC. The closer the AUC value is to 1, the better the prediction performance of the method. The ROC curves of different prediction methods on the HMDDv2.0 imbalanced dataset are shown in Figure 2. Figure 3 As shown in the figure. From high to low, they are LGBMDA, ABMDA, GAEMDA, MLRDFM, and SAEMDA, with AUCs of 0.9699, 0.9428, 0.9333, 0.9311, and 0.9164, respectively. In addition, in order to prove the rationality of using the lightweight gradient boosting machine classifier in the proposed LGBMDA prediction method, Figure 4 In the 2017 study, the method was compared with other classifiers such as Naive Bayes, Multilayer Perceptron, Logistic Regression (LR), and Support Vector Machine (SVM). Clearly, the AUCs of these classifiers were 0.9179, 0.8661, 0.9282, and 0.9141, respectively. Therefore, it can be concluded that the use of the lightweight gradient boosting machine classifier in the proposed LGBMDA prediction method is reasonable, and the prediction method of the present invention achieves better results. Based on the experimental results, the present invention has higher prediction performance than other methods and can effectively explore potential miRNA-disease associations.

[0133] In summary, the present invention first uses multi-source biological and medical information to obtain the miRNA-disease association adjacency matrix, miRNA functional similarity matrix, and disease semantic similarity matrix; integrates the similarity matrix of miRNA and disease with the Gaussian interaction spectrum kernel similarity matrix of miRNA and disease calculated using the miRNA-disease association adjacency matrix; splices the integrated similarity matrix with the miRNA-disease association adjacency matrix to obtain a more informative miRNA-disease association feature vector; secondly, uses an autoencoder to extract the key features of the comprehensive miRNA feature vector and the disease feature vector, and finally uses a lightweight gradient boosting machine classifier to realize the potential association prediction of miRNA and disease.

[0134] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A gradient boosting machine miRNA-disease association prediction method based on an encoder, characterized in that: The following steps are involved: Step 1: Obtain miRNA-disease association data, disease autocorrelation data, and miRNA autocorrelation data from a preset number of external biological and medical data sources; Step 2: Calculate the disease semantic similarity matrix based on the data obtained in step 1; Step 3, based on the data obtained in step 1, calculate the miRNA functional similarity matrix; Step 4: obtaining the miRNA-disease association adjacency matrix based on the miRNA-disease association data obtained in step 1, and calculating the Gaussian interaction spectrum kernel similarity matrix of miRNA and disease based on the miRNA-disease association adjacency matrix; Step 5: Integrate the disease semantic similarity matrix obtained in step 2 and the miRNA functional similarity matrix obtained in step 3 with the miRNA and disease Gaussian interaction spectrum kernel similarity matrix calculated in step 4 to obtain an integrated similarity matrix of miRNA and disease; Step 6, the integrated similarity matrix of miRNAs and diseases obtained in step 5 is spliced ​​with the miRNA-disease association adjacency matrix to obtain the comprehensive similarity matrix of miRNAs and diseases; Step 7: Input the feature vector of the miRNA and disease comprehensive similarity matrix obtained in step 6 into the autoencoder for feature extraction to extract low-dimensional and high-quality miRNA-disease association feature vectors; Step 8: Input the low-dimensional, high-quality miRNA-disease association feature vector obtained in step 7 into the lightweight gradient boosting machine classifier to perform miRNA-disease association prediction; The obtained miRNA and disease comprehensive similarity matrices are: in, , , , Respectively represent The integration similarity of the miRNA with other miRNAs, The integrated similarity between the disease and other diseases, the first Row, the first row of the validated miRNA-disease association matrix Column transposition; The feature vector of the miRNA and disease comprehensive similarity matrix obtained in step 6 is input into the autoencoder for feature extraction to extract low-dimensional, high-quality miRNA-disease association feature vectors, including: Use the autoencoder to encode 495×383 rows and 495+495 columns Matrix and 495×383 rows 383+383 columns The matrix is ​​processed to extract their low-dimensional features, while reducing the noise caused by the redundant information hidden in the original feature vector, and two low-dimensional high-quality miRNA-disease association feature vectors are obtained.

2. The encoder-based gradient boosting machine miRNA-disease association prediction method according to claim 1, characterized in that: Obtain the HMDD database, from which the miRNA-disease association data verified by biological experiments are obtained, thereby obtaining the miRNA-disease association adjacency matrix; obtain the disease semantic similarity matrix based on the MeSH database; and obtain the miRNA functional similarity matrix based on the gene ontology database.

3. The encoder-based gradient boosting machine miRNA-disease association prediction method according to claim 1, characterized in that: The calculation obtains a disease semantic similarity matrix, including: Calculating ancestral diseases For disease The formula corresponding to the semantic value contribution of is: Among them, △ is the semantic contribution factor set to 0.5, is one of the ancestral diseases, One of the ancestral diseases For disease The semantic value contribution of It is an ancestral disease For disease The semantic value contribution of disease The semantic contribution value of itself is set to 1, and the contribution value of the ancestral disease in the directed acyclic graph DAG (d) is combined. The formula corresponding to the semantic value of is: in, Is related to disease A collection of related ancestral diseases, diseases and The formula for the semantic similarity is: Another way to calculate diseases in DAG For disease The formula corresponding to the contribution of semantic value is: disease The formula corresponding to the semantic similarity value is: disease and diseases The formula corresponding to the semantic similarity value between is: By integrating the two disease semantic similarity values, the formula corresponding to the disease semantic similarity matrix DS is proposed as follows: 。 4. The encoder-based gradient boosting machine miRNA-disease association prediction method according to claim 1, characterized in that: The Gaussian interaction spectrum kernel similarity matrix of miRNA and disease is calculated based on the miRNA-disease association adjacency matrix, including: Uncovering diseases through Gaussian interaction spectral kernel similarity and diseases The relationship between Indicates disease Validated associations between each miRNA, resemblance, is the number of diseases, is a parameter used to adjust the kernel bandwidth, and the Gaussian interaction spectrum kernel similarity matrix between diseases The corresponding formula is: KD(d i ,d j )=exp(-γ d PIP(d i )-IP(d j )P 2 ) miRNA Known associations with each disease, Similar to it, is the number of miRNAs, Describes the parameters used to adjust the kernel bandwidth, Gaussian interaction profile kernel similarity matrix between miRNAs The corresponding formula is: KM(m i ,m j )=exp(-γ m PIP(m i )-IP(m j )P 2 ) 5. The encoder-based gradient boosting machine miRNA-disease association prediction method according to claim 1, characterized in that: The disease semantic similarity matrix obtained in step 2 and the miRNA functional similarity matrix obtained in step 3 are respectively integrated with the miRNA and disease Gaussian interaction spectrum kernel similarity matrix calculated in step 4 to obtain an integrated similarity matrix of miRNA and disease, including: By converting the miRNA Gaussian interaction profile kernel similarity matrix and miRNA functional similarity matrix integrated together to obtain an integrated similarity matrix of miRNAs The corresponding formula is: The Gaussian interaction spectrum kernel similarity matrix of the disease and disease semantic similarity matrix Combined together to obtain the integrated similarity matrix of diseases The corresponding formula is: 。 6. The encoder-based gradient boosting machine miRNA-disease association prediction method according to claim 1, characterized in that: The low-dimensional, high-quality miRNA-disease association feature vector obtained in step 7 is input into a lightweight gradient boosting machine classifier to perform miRNA-disease association prediction, including: The two types of low-dimensional, high-quality miRNA-disease association feature vectors obtained are input into the lightweight gradient boosting machine classifier for prediction. A prediction result is obtained from the perspective of disease and miRNA respectively, and the two prediction results are integrated to obtain the final miRNA-disease association prediction value.

Citation Information

Patent Citations

  • MiRNA and disease incidence relation prediction method based on graph convolutional network

    CN114496092A

  • Method for predicting potentially associated circular RNA-disease pairs based on GCN and ensemble learning

    CN114582508A