m Association Prediction Method Based on Subspace Learning 7 Associations between G and Diseases
The SpBLRSR method enhances m7G-disease association prediction by combining Schatten-p norm minimization and sparse subspace clustering, addressing inefficiencies in existing algorithms to provide accurate m7G site predictions for tumor treatment strategies.
Patent Information
- Application Number
- CN202110903743.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-06
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-08-06
AI Technical Summary
The existing m7G-disease association prediction algorithms have insufficient matrix filling in the field of bioinformatics, and cannot effectively capture the association pattern, and the existing subspace learning algorithms are difficult to accurately recover missing elements in the matrix in noisy environments.
Using a method combining Schatten-p norm minimization and sparse subspace clustering, a comprehensive similarity matrix and disease similarity matrix of m7G sites were constructed to establish a complete low-dimensional subspace, and the model was optimized by using the alternating direction multipliers method to restore the high-dimensional correlation matrix, and potential disease-related m7G sites were extracted.
It significantly improves the accuracy and robustness of m7G-disease association prediction, can better predict m7G sites related to unknown diseases, and provide key information for subsequent wet experiments.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FDA0005406792480000011
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics research, specifically a method for predicting the association between m 7 G and diseases, providing a reference for exploring the tumorigenesis mechanism and pathogenic pathways at the epi-transcriptome level. The present invention is expected to provide new ideas for tumor treatment strategies targeting m 7 G modification sites. Background Art
[0002] 7-Methylguanosine (m 7 G) is an RNA modification occurring at the N7 position of guanosine (G), which is part of the 5'-capping structure of messenger RNA (mRNA) and widely exists in various RNAs of multiple species. In recent years, academic research has shown that m 7 G plays an important regulatory role in biological processes and normal physiological functions. Among them, it has been found that m 7 G is involved in the occurrence and metastasis of various diseases and even malignant tumors through processes such as RNA transcription, splicing, processing, translation, and decay. Shaheen et al. found that mutations in the m 7 G transferase complex METTL1-WDR4 can lead to a decrease in the number of m 7 G sites on tRNA, thus inducing primordial dwarfism and brain developmental malformations. Lin et al. found that knocking out the WDR4 gene would disrupt the differentiation of embryonic stem cells and lead to nervous system disorders. In addition, Deng et al. found that the silencing of the METTL1 gene would cause a decrease in the number of m 7 G sites on pluripotent stem cells, severely inhibiting the differentiation of mesoderm and vascular development. Additionally, recent reports have also shown abnormalities in m 7 G sites in lung cancer, colon cancer, and liver cancer cells. Therefore, if the m 7 G modification sites associated with a certain disease can be further determined, it will provide new ideas for tumor treatment strategies targeting m 7 G sites. Although wet experiments can provide relatively accurate m 7 G-disease association relationships, their processes are relatively complex and consume a large amount of manpower, material resources, and financial resources. Therefore, relying on existing bioinformatics databases that record m 7 G-disease associations, we use computational methods to predict unknown m 7 G-disease associations, providing key m 7 G sites for biologists to carry out subsequent wet experiments.
[0003] Although the current solution to m 7There are few algorithms for G-disease association prediction. However, from a mathematical perspective, implementing m 7 The essence of G-disease association prediction is to fill m 7The association matrix between the G locus and diseases, i.e., the matrix completion problem. In recent years, subspace learning, as a method to capture the latent structure of matrices, has provided new ideas for solving the matrix completion problem and has attracted extensive attention in the field of bioinformatics. The principal component analysis (PCA) proposed by Jolliffe et al. is the most widely used subspace learning algorithm and has achieved remarkable results under Gaussian noise. PCA assumes that the given data comes from a low-dimensional linear subspace and a noise subspace, and learns the denoised low-dimensional linear subspace by solving the nuclear norm minimization and Frobenius norm minimization problems to recover the missing elements in the matrix. However, in practice, the given observed data does not necessarily originate from the same low-dimensional complete subspace, but from multiple ones. To capture the structural characteristics of multiple low-dimensional complete subspaces from the noise, Liu et al. proposed the low-rank representation (LRR) model. LRR assumes that in a set of samples, each sample can be represented as a linear combination of the bases in the subspace. Further, LRR uses the subspace itself as a dictionary and aims to find the low-rank representation of all samples. During the process of sample partitioning, LRR divides the original high-dimensional space into multiple low-dimensional subspaces. The LRR model is a nuclear norm minimization model (NNM), which uses the nuclear norm to perform convex relaxation on the rank function. As a convex problem, NNM can obtain a global optimal solution, but it still deviates from the solution of the original rank minimization problem. This is attributed to two reasons: one is that the nuclear norm is not the best approximation of the rank function, and the other is that when using the singular value thresholding (SVT) algorithm, regardless of whether the singular values are large or small, SVT penalizes and converges them at the same rate. However, large singular values contain more information, while small singular values contain a lot of noise. Therefore, theoretically, smaller singular values should be penalized more severely to make them converge at a faster rate, and larger singular values should be penalized less severely to make them converge at a slower rate. In recent years, the Schatten-p norm has been gradually applied in the field of matrix completion. Compared with the nuclear norm, the Schatten-p norm is a tighter non-convex approximation of the rank function. Nie et al. first proposed the LRR model based on the Schatten-p norm, i.e., the Schatten-p norm minimization model. In the matrix completion scenario, it has obtained higher prediction accuracy and better robustness than the traditional LRR model. Then, Zhang et al. proposed the generalized matrix soft thresholding (GMST) for solving the Schatten-p norm minimization problem. GMST can make smaller singular values converge faster and larger singular values converge slower.Experiments show that GMST can better characterize the structure of the subspace than SVT. Although the Schatten-p norm minimization problem can more effectively mine the global structural characteristics of the subspace, it ignores the local structural characteristics of the subspace. In addition, the sparse subspace clustering algorithm (Sparse Subspace Clustering, SSC) can mine the local sparse expression characteristics of the subspace. Formally, SSC represents each data point as a linear combination of other data points in the subspace to divide the high-dimensional space, but SSC cannot capture the overall properties of the association matrix. Effective m. 7 The G-disease association prediction algorithm has yet to be proposed, and as a subspace learning algorithm that has achieved good results, how to clarify its association pattern under the premise that there are missing elements in the matrix is still a problem that needs to be improved urgently. Summary of the invention
[0004] The goal of this invention is to establish m based on subspace learning 7 G-disease association prediction method. Based on the assumption that "similar sites will cause similar diseases", the present invention uses the known m 7 G-disease related information, m 7 G-site similarity information and disease similarity information, starting from the high-dimensional heterogeneous subspace with missing structure and unclear association pattern, learn the complete low-dimensional subspace that reflects its statistical and structural characteristics, and then restore the incomplete high-dimensional association matrix to extract potential disease-related m 7 G site. The specific implementation steps are:
[0005] Step (1): From m 7 Get m from GDisAI database 7 G-disease association matrix H SD (741*177), m 7 G site chemical similarity matrix H SS_chemical (741*741), m 7 G site base cumulative frequency (Cumulative Nucleotide Frequency, CNF) similarity matrix H SS_CNF (741*741) and disease similarity matrix H DD (177*177). Among them:
[0006]
[0007] Step (2): Construct m 7 G site comprehensive similarity matrix H SS :
[0008] H SS =αH SS_chemical+(1 - α)H SS_CNF , 0 ≤ α ≤ 1 (1)
[0009] Where α is the combined coefficient of the chemical similarity of the m 7 G-site and the cumulative frequency similarity of the m 7 G-site base.
[0010] Step (3): Construct the m 7 G-disease heterogeneous matrix. The sparsity of the m 7 G-disease association matrix is very high, making it difficult to ensure the normal solution of the algorithm. At the same time, the high sparsity makes the association pattern of the matrix with missing elements less clear. To clarify the association pattern and ensure the normal solution of the algorithm, construct the m 7 G-disease heterogeneous matrix, as shown in formula (2).
[0011]
[0012] H SD T is the transpose of H SD . m is the number of m 7 G-sites, n is the number of diseases, then the dimension of the m 7 G-disease heterogeneous matrix H is (m + n) * (m + n). As can be seen from (2), H has good properties. First, H is a symmetric and positive semi-definite matrix, its singular values are positive real numbers and equal to the eigenvalues. And its left and right singular vectors are equal and equal to the eigenvectors. The special properties of the singular values and singular vectors of the positive semi-definite matrix greatly reduce the computational complexity and speed up the algorithm. Second, the sparsity of the m 7 G-disease heterogeneous matrix H is much smaller than that of the association matrix H SD , providing a guarantee for the normal solution of the algorithm. Finally, the missing values in H only appear in H SD and its transpose, so the filling problem of H SD can be transformed into the filling problem of H.
[0013] 1. Step (4): The m 7 G-disease association prediction method based on subspace learning learns a complete low-dimensional subspace that can preserve its statistical and structural characteristics from an incomplete high-dimensional subspace, so as to predict the missing associations in the matrix and provide key m 7 G-sites for the next wet experiment of biologists. Here, X is used to represent the complete low-rank heterogeneous matrix, and C is used to represent the sparse self-expression matrix. Combining the Schatten-p norm minimization model and the SSC model, the objective function (3) is established:
[0014]
[0015] Among them, X and C follow the above definitions and are respectively the predicted complete low-rank heterogeneous matrix and the sparse self-expression matrix. λ is the balance coefficient. ||X - XC|| F 2 The noise that may occur during the matrix recovery process is considered. The constraint diag(C) = 0 is to avoid the trivial solution C = I during the solution of C. P Ω is the projection operator, and its definition is as follows:
[0016]
[0017] Step (5): Introduce two auxiliary matrices M and A to obtain the equivalent form (4) of model (3). Furthermore, the Alternating Direction Method of Multiplier (ADMM) is used to solve it, and the specific process is as follows:
[0018]
[0019] The augmented Lagrangian function of the above formula is as shown in (5):
[0020]
[0021] Among them, μ is the penalty factor, and Y1 and Y2 are Lagrange multipliers. In the (k + 1)-th iteration, SpBLRSR needs to alternately solve the sub-problems X k+1 , M k+1 , A k+1 , C k+1 , Y1 k+1 and Y2 k+1 .
[0022] Update of X: Fix M k , A k , C k , Y1 k , Y2 k , and minimize the augmented Lagrangian function:
[0023]
[0024] Denote Take the partial derivative of X in f(X), and the analytical solution of the minimum value X * is as shown in (6).
[0025] X * = (Y2 k + μM k )(I - ((A k ) T + A k ) + Ak (A k ) T + μI) -1 (7)
[0026] Furthermore, with the restrictions on the numerical range and the projection matrix, the analytical solution of X k+1 is shown in (8):
[0027] X k+1 = L [0,1] (X * ), P Ω (X k+1 ) = P Ω (H) (8)
[0028] where L [0,1] is defined as:
[0029]
[0030] Update of M: Fix X k+1 , A k , C k , Y1 k , Y2 k , and minimize the augmented Lagrangian function:
[0031]
[0032] The above model is a Schatten-p norm minimization problem, and its analytical solution is given by the GMST operator:
[0033]
[0034] where is the GMST operator. U and V are the left and right singular matrices of the matrix respectively, and ξ = (ξ1, ξ2,..., ξ r ) are the singular values of . is the Generalized Vector Soft Thresholding (GVST) operator, and its basic composition is the Generalized Soft Thresholding (GST) operator:
[0035]
[0036] For each GST operator there is:
[0037]
[0038] Among them, and obtained by solving .
[0039] Update of A: Fix X k+1 , M k+1 , C k , Y1 k , Y2 k , minimize the augmented Lagrangian function, take the partial derivative with respect to A, and the analytical solution of A is shown in (11):
[0040] A k+1 = ((X k+1 ) T X k+1 + μI) -1 ((X k+1 ) T X k+1 - Y1 k + μC k ) (11)
[0041] Update of C: Fix X k+1 , M k+1 , A k+1 , Y1 k , Y2 k , minimize the augmented Lagrangian function, transform it into a problem of minimizing the L1 norm, and its analytical solution can be represented by the soft threshold operator:
[0042]
[0043] Among them, is the soft threshold operator.
[0044] Y 1 , Y 2 Update of: Fix X k+1 , M k+1 , A k+1 , C k+1 , obtained by the gradient descent method:
[0045]
[0046] Among them, δ is the learning rate.
[0047] Step (6): Repeat the above process until the algorithm converges. The complete low-rank m 7 The complete low-rank correlation matrix X in the G-disease heterogeneous matrix X SD is the final solution. Description of the Drawings
[0048] Figure 1 is based on subspace learning m 7Flowchart of the prediction model for the association between G and diseases.
[0049] Figure 2 It is a graph showing the comparison results of the AUC performance density distributions of the MF, SVT, RPCA, and SpBLRSR algorithms. Detailed implementation manners
[0050] To further explain the specific content and advantages of the present invention, the following is a detailed description of the specific implementation manners and the drawings.
[0051] To illustrate the effectiveness of the present algorithm, we compare the proposed algorithm SpBLRSR of the present invention with three other algorithms widely used in the field of matrix completion under two cross-validation frameworks, namely the Matrix Factorization (MF) algorithm, the Singular Value Thresholding (SVT) algorithm, and the Robust Principal Component Analysis (RPCA) algorithm. First, we compare their prediction accuracies in a 10-fold Cross Validation (10-fold CV) experiment. 7 GDisAI records 768 pairs of 7 associations between G loci and diseases, which involve 741 7 G loci and 177 diseases. We regard 130,389 unknown association samples as candidate samples (candidate set). Additionally, the 768 known associations are randomly divided into ten independent parts with nearly the same scale. Each time, nine parts are used as the training set (training set), and the remaining one part is used as the test set (test set). The algorithm trains the parameters on the training set and applies them to the test set and the candidate set to obtain prediction scores. Then, we use AUC (Area under ROC curve) and AUPR (Area under PR curve) to measure the prediction capabilities of the four algorithms, and the results are shown in Table 1. Finally, the Wilcoxon signed-rank test is used to test the significance of the differences in the algorithm prediction results. Table 2 shows the significance of the differences in the prediction results between the MF, SVT, and RPCA and the SpBLRSR algorithms.
[0052] Table 1: Prediction effects of each algorithm under 10-fold CV.
[0053]
[0054] Table 2: Significance of the differences between MF, SVT, and RPCA and SpBLRSR under 10-fold CV.
[0055]
[0056] Experiments show that the prediction ability of SpBLRSR is significantly better than that of MF, SVT, and RPCA.
[0057] Secondly, for a disease, it is more common that no m 7 G sites related to it are known. Here, we use the "Leave-One-Disease-Out-Cross-Validation" (LODOCV) method to verify the ability of SpBLRSR to predict m 7 G sites related to new diseases. Similar to 10-fold CV, we regard 130,389 unknown associations as the candidate set. The difference is that for disease d, we use all the known m 7 G sites related to d in the association matrix as the test samples, and regard the known associations between m 7 G sites and other diseases as the training samples. Finally, we train the algorithm with the training set and then apply the trained algorithm to the test set and the candidate set. We perform the above experiments on all 177 diseases to obtain 177 AUC values. Figure 2 Shows the results of the LODOCV experiments of each algorithm. As can be seen from Figure 2 , the AUC scores generated by MF, SVT, and RPCA are mainly distributed between 0.3 and 0.6, while the AUC scores obtained by SpBLRSR are concentrated between 0.6 and 0.85. This shows that SpBLRSR is better than MF, SVT, and RPCA in the LODOCV experiment, and also shows that SpBLRSR can more accurately predict the m 7 G sites related to a certain disease in the absence of relevant m 7 G sites.
[0058] In summary, the SpBLRSR algorithm proposed by the present invention is more effective than MF, SVT, and RPCA in predicting m 7 G-disease associations, showing higher prediction accuracy and the ability to provide relevant m 7 G sites for diseases.
[0059] Finally, it should be noted that the above embodiments are for better explaining the idea of the present invention and are by no means a limitation to the present invention. Any equivalent replacement, modification, or supplement made according to the essential content of the present invention should be included within the protection scope of the present invention.
Claims
1. m Association Prediction Method Based on Subspace Learning 7 for predicting the potential association between m 7 G modification and diseases, characterized in that Including the following steps: Step (1): Obtain m from 7 the GDisAI database 7 G-disease association matrix H SD (741 * 177), m 7 G-site chemical similarity matrix H SS_chemical (741 * 741), m 7 G-site base cumulative frequency (Cumulative Nucleotide Frequency, CNF) similarity matrix H SS_CNF (741 * 741) and disease similarity matrix H DD (177 * 177), where: Step (2): Construct m 7 Comprehensive similarity matrix H of G sites SS : H SS = αH SS_chemical + (1 - α)H SS_CNF , 0 ≤ α ≤ 1 (1) where α is m 7 the combination coefficient of the chemical similarity of the G site and m 7 the cumulative frequency similarity of the bases at the G site Step (3): Construct m 7 G-disease heterogeneous matrix, m 7 The sparsity of the G-disease association matrix is very high, making it difficult to ensure the normal solution of the algorithm. At the same time, the high sparsity makes the association pattern of the matrix with missing elements even less clear. To clarify the association pattern and ensure the normal solution of the algorithm, construct m 7 G-disease heterogeneous matrix, as shown in equation (2): H SD T is the transpose of H SD , where m is the number of 7 G loci, and n is the number of diseases. Then m 7 The dimension of the G-disease heterogeneous matrix H is (m + n)×(m + n). As can be seen from (2), H has good properties. First, H is a symmetric and positive semi-definite matrix. Its singular values are positive real numbers, equal to the eigenvalues, and its left singular vectors and right singular vectors are equal, equal to the eigenvectors. The particularity of the singular values and singular vectors of the positive semi-definite matrix greatly reduces the computational complexity and speeds up the running speed of the algorithm. Second, m 7 The sparsity of the G-disease heterogeneous matrix H is much smaller than that of the association matrix H SD , which guarantees the normal solution of the algorithm. Finally, the missing values in H only appear in H SD and its transpose. Therefore, the filling problem of H SD can be transformed into the filling problem of H.
2. The m 7 association prediction method between G and the disease based on subspace learning according to claim 1, characterized in that Learn a complete low-dimensional subspace from an incomplete high-dimensional subspace that can preserve its statistical and structural properties, thereby predicting the missing associations in the matrix and providing the key m for biologists to carry out the next wet experiment 7 For the G locus, where X is used to represent the complete low-rank heterogeneous matrix and C is used to represent the sparse self-expression matrix, the Schatten-p norm minimization model and the SSC model are combined to establish the objective function (3): where X and C follow the above definitions, and are the predicted complete low-rank heterogeneous matrix and the sparse self-expression matrix respectively, λ is the balance coefficient, and ||X - XC|| F 2 considers the possible noise in the matrix recovery process. The constraint diag(C) = 0 is to avoid the trivial solution C = I in the solution process of C, and P Ω is the projection operator, and its definition is as follows:
3. The m 7 G and disease association prediction method according to claim 2, characterized in that Introduce an auxiliary matrix, establish an equivalent model of the original model, and solve it using the Alternating Direction Method of Multiplier (ADMM). The specific process is as follows: Step (1): Introduce auxiliary matrices M and A to obtain the equivalent model (4): Step (2): Write out the augmented Lagrangian function of (4), as shown in (5): where μ is the penalty factor, and Y1 and Y2 are Lagrange multipliers. In the (k + 1)-th iteration, the sub-problems X k+1 , M k+1 , A k +1 , C k+1 , Y1 k+1 and Y2 k+1 are solved alternately. Step (3): Update X: Fix M k , A k , C k , Y1 k , Y2 k , Minimize the augmented Lagrangian function: Record Taking the partial derivative of \(X\) in \(f(X)\), the minimum value of \(X\) * The analytical solution is shown in (7). X * = (Y2 k + μM k )(I - ((A k ) T + A k ) + A k (A k ) T + μI) -1 (7) Furthermore, with the restrictions on the numerical range and the projection matrix, the analytical solution of X k+1 is shown in (8). X k+1 = L [0,1] (X * ), P Ω (X k+1 ) = P Ω (H)(8) where L [0,1] is defined as: Step (4): Update M: Fix X k+1 , A k , C k , Y1 k , Y2 k , Minimize the augmented Lagrangian function: The above model is a Schatten-p norm minimization problem, and its analytical solution is given by the Generalized Matrix Soft Thresholding (GMST): Among them, is the GMST operator, and U and V are the left and right singular matrices of the matrix , respectively. ξ = (ξ1, ξ2,..., ξ r ) are singular values, is the Generalized Vector Soft Thresholding (GVST) operator, and its basic composition is the Generalized Soft Thresholding (GST) operator: For each GST operator There is: Among them, and obtained by solving to get Step (4): Update A: Fix X k+1 , M k+1 , C k , Y1 k , Y2 k , Minimize the augmented Lagrangian function, take the partial derivative with respect to A, and the analytical solution of A is shown in (11): A k+1 = ((X k+1 )) T X k+1 + μI) -1 ((X k+1 )) T X k+1 - Y1 k + μC k )(11) Step (5): Update C and fix X k+1 , M k+1 , A k+1 , Y1 k , Y2 k , Minimize the augmented Lagrangian function and transform it into a problem of minimizing the L1 norm. Its analytical solution can be represented by a soft thresholding operator: Among them, is a soft threshold operator, Step (6): Update Y 1 , Y 2 : Fix X k+1 , M k+1 , A k+1 , C k+1 , obtained by the gradient descent method: where δ is the learning rate, and the above process is repeated until the algorithm converges, and the complete low-rank m 7 The complete low-rank correlation matrix X in the G-disease heterogeneous matrix X SD is the final solution sought.
Citation Information
Patent Citations
Non-coded RNA and disease association prediction method based on sparse subspace learning
CN110767263A
IncRNA-disease association prediction method based on high-order proximity and matrix completion algorithm
CN113160880A