SM-miRNA association prediction method based on matrix enhancement and collaborative double matrix completion
By using matrix enhancement and collaborative dual matrix completion methods in SM-miRNA association prediction, the problems of slow computing speed and low accuracy in the prior art are solved, and more efficient and accurate prediction effects are achieved.
Patent Information
- Application Number
- CN202411906873.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-12-24
AI Technical Summary
Existing SM-miRNA association prediction methods have slow calculation speed, similarity matrix may be noisy and underutilize topological information, resulting in low accuracy and incomplete prediction.
Using a method based on matrix enhancement and collaborative dual matrix completion, the similarity matrix of SM and miRNA is enhanced by the Gaussian radial basis function, combining truncated schatten p-norm and truncated matrix decomposition, the correlation matrix is completed and the prediction accuracy is improved.
It improves the accuracy and speed of SM-miRNA association prediction, enhances the accuracy of similarity metrics, and makes full use of topological information to achieve a more comprehensive prediction effect.
Smart Images

Figure CN119360951B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to predicting the association between small molecule drugs and miRNA, and in particular to a SM-miRNA association prediction method based on matrix enhancement and collaborative double matrix completion. Background Art
[0002] MicroRNAs (miRNAs) are evolutionarily conserved non-coding RNA molecules present in various organisms, typically spanning 21 to 25 nucleotides. They play crucial roles in regulating a variety of biological processes.
[0003] The expression pattern of miRNAs is tissue-specific, and any abnormal expression will affect the cell state. A large amount of clinical and experimental evidence shows that miRNAs are involved in a wide range of complex human diseases, including cardiovascular diseases, metabolic diseases, and inflammatory diseases. For example, the expression of hsa-miR-125a-3p is significantly reduced in breast cancer cells. In squamous cell lung cancer tissues, the expression pattern of miRNAs changes systematically, and increased levels of miR-21 are associated with shortened survival time. Therefore, miRNAs have the potential to serve as valuable clinical and prognostic biomarkers. More and more studies have shown that small molecule drugs (SM) can effectively intervene in the expression and function of specific miRNAs, thereby achieving therapeutic effects.
[0004] Recently, more and more SM-miRNA associations (MMAs) have been discovered. For example, miR-155 expression, which is upregulated in various cancers such as lung cancer, colorectal cancer, and breast cancer, can be downregulated by curcumin. This downregulation inhibits tumor cell protrusion and invasion. The ability of SM to precisely target disease-specific miRNA pathways makes it an important tool for personalized medicine and targeted therapy.
[0005] Various SMs exhibit different mechanisms of action and efficacy by targeting different miRNAs. Therefore, it is necessary to find specific associations between SMs and miRNAs. However, traditional biological experiments face challenges such as time, cost, and technical limitations, which hinder their application in large-scale studies and comprehensive analysis. Therefore, it is necessary to develop computational models to predict MMAs. In recent years, many computational models have been proposed to predict MMAs. These methods can be mainly divided into three categories: network inference-based methods, machine learning-based methods, and matrix completion-based methods.
[0006] Network inference-based methods use biological information to formulate heterogeneous networks and use inference algorithms for prediction. The TLHNSMMA model was introduced to integrate SM, miRNA and disease information to construct a three-layer heterogeneous network. The network information was then used to infer unknown MMAs. The SLHGISMMA model was proposed to integrate SM / miRNA similarity and MMA information in a heterogeneous graph. They used sparse learning methods to denoise the matrix in MMA to improve the final prediction accuracy. GISMMA used the interactions between 28 isoforms to measure the correlation strength between SMs and miRNAs. The SM / miRNA similarity network was calculated to obtain the association score. SMiR-NBI was proposed, which involves constructing a heterogeneous network using drugs, miRNAs and genetic information. Inference was performed based on the information within the heterogeneous network to obtain the prediction score. However, a potential disadvantage may be the over-reliance on web-based information, which leads to a lack of flexibility.
[0007] Machine learning-based methods mainly focus on latent feature extraction and classifiers. DAESTB combines SM and miRNA related information to construct a multidimensional feature matrix. Then, deep autoencoders are used for denoising and dimensionality reduction. Subsequently, a scalable tree boosting model is adopted to derive prediction results. EKRRSMMA curates a subset of SM and miRNA features and applies feature downscaling techniques to mitigate the impact of noisy data during the construction process. In the subsequent training phase, they use ensemble learning methods to significantly improve the achievement of accurate prediction results. CLDISMMA constructs a complex network consisting of SMs, diseases, and miRNAs. Then, a regularized model is used to infer unknown MMAs. RFSMMA uses SM / miRNA similarity as features to represent MMAs. Subsequently, a random forest technique is used for training to obtain prediction scores. Despite their effectiveness, these methods may have difficulty in capturing the underlying correlations of sparse matrices and lack interpretability.
[0008] The objective function of calculating the MMA score is constructed and optimized based on the matrix completion method. According to the principle of the method, the matrix completion method can be further divided into two categories: rank approximation norm minimization method and matrix decomposition method.
[0009] (1) Rank-approximation normed minimization: A kernel normed minimization method (BNNRSMMA) was introduced. First, they built a heterogeneous network and defined a matrix to represent it. Subsequently, a prediction model was designed to fill the missing values in the matrix using nuclear norm minimization. AMCSMMA of MMAs was proposed. Their method involves integrating bioinformatics to create a heterogeneous network. The neighborhood matrix of the network is treated as the target matrix, and the missing values are filled by truncated nuclear norm minimization. A TSPN method was proposed. The bioinformatics matrix was integrated into a heterogeneous network, and the neighboring matrix was used as the target. By minimizing the truncated schatten p-norm, they reconstructed the missing values in the target matrix. An iterative algorithm was designed to solve the model and obtain the correlation score. Although these methods have advantages in feature selectivity and noise robustness, their computational methods, involving heterogeneous network adjacency matrices, introduce significant complexity.
[0010] (2) Matrix decomposition: A SMANMF method was proposed to reveal MMAs using non-negative matrix factorization. Zhao et al. designed SNMFSMMA, which adopted a symmetric non-negative matrix factorization model and solved it by the Kronecker regularized least squares algorithm. A DCMF method was proposed. Initially, they used the WKNKN method for preprocessing. Subsequently, they constructed the objective function and iteratively updated the feature matrices of SMs and miRNAs. Finally, the association score matrix was obtained by iterative updating until convergence. Although matrix decomposition has been shown to be highly computationally efficient and interpretable, it is sensitive to noise and may perform poorly when dealing with sparse data.
[0011] Matrix completion is a feasible approach, but existing methods have limitations. In addition, similarity matrices may be noisy and their topological information is not fully utilized. Summary of the invention
[0012] In order to solve the problems in the related art, the present application provides an SM-miRNA association prediction method based on matrix enhancement and collaborative double matrix completion, which solves the problems of existing methods for predicting MMAs, such as slow calculation speed due to large amount of calculation, possible noise in similarity matrix and failure to fully utilize topological information, resulting in low accuracy and incomplete prediction.
[0013] The technical solution is as follows:
[0014] The SM-miRNA association prediction method based on matrix enhancement and collaborative double matrix completion is characterized by comprising the following steps:
[0015] Step 1, respectively enhancing the precision of the comprehensive SM similarity matrix and the comprehensive miRNA similarity matrix by using Gaussian radial basis function;
[0016] Step 2, obtain a probability value in the range of [0,1] by minimizing the truncated schatten p-norm, and use the obtained probability value to replace the missing data of the SM-miRNA association matrix to update the SM-miRNA association matrix;
[0017] Step 3: Perform truncated matrix decomposition on the SM-miRNA association matrix updated in step 2, and after decomposition, combine the enhanced comprehensive SM similarity matrix and comprehensive miRNA similarity matrix in step 1 to obtain a prediction score matrix, and predict the results based on the prediction score matrix.
[0018] Through the above technical scheme, matrix enhancement is performed through GRBF, which takes into account the structural information of the comprehensive similarity matrix, thereby improving the accuracy of the similarity measure; because the truncated schatten p-norm takes into account the physical properties of singular values, it can supplement the missing values more effectively than other rank approximation norms; in addition, the truncated matrix decomposition greatly improves the speed of the model by truncating the number of singular values; by cleverly combining the matrix completion method of truncated schatten p-norm minimization and truncated matrix decomposition, it fully utilizes the feature selectivity and robustness of the former method, and utilizes the efficient calculation and interpretability of the latter method, thereby showing comprehensive and flexible performance in MMA prediction.
[0019] Preferably, in step 1, the Gaussian radial basis function performs matrix enhancement on the comprehensive SM similarity in the comprehensive SM similarity matrix, and obtains the final refined SM similarity in the comprehensive SM similarity matrix after processing; the comprehensive SM similarity is obtained by four types of SM similarities, and the four types of SM similarities are obtained from four SM similarity matrices in the comprehensive SM similarity matrix, and the four types of SM similarities include similarity based on SM side effects, similarity based on chemical structure, similarity based on gene function consistency, and similarity based on indicator phenotype;
[0020] The Gaussian radial basis function performs matrix enhancement on the comprehensive miRNA similarity in the comprehensive miRNA similarity matrix, and obtains the final refined miRNA similarity in the comprehensive miRNA similarity matrix after processing; the comprehensive miRNA similarity is obtained through two types of miRNA similarities, and the two types of miRNA similarities are obtained from two miRNA similarity matrices in the comprehensive miRNA similarity matrix, and the two types of miRNA similarities include similarity based on gene function consistency and similarity based on disease phenotype.
[0021] Through the above technical solution, by performing matrix enhancement on the comprehensive SM similarity and the comprehensive SM similarity, the final refined SM similarity and the SM similarity can be obtained, thereby improving the accuracy of the similarity measurement.
[0022] Preferably, the four types of SM similarities and the two types of miRNA similarities are calculated based on a benchmark miRNA dataset, which includes 796 known MMAs, 831 SMs and 541 miRNA datasets.
[0023] Preferably, the calculation process of the comprehensive SM similarity is:
[0024] Create 4 matrices corresponding to the four types of SM similarities. Each matrix is used to store the corresponding SM similarity. The 4 matrices are and The sizes of the four matrices are ns×ns, where ns represents the number of small molecules. and Used to indicate the similarity of corresponding matrices at coordinates (i, j);
[0025] The SM similarities of the four matrices are integrated through a weighted combination strategy, and the comprehensive SM similarity matrix S is obtained after integration. sm ,
[0026]
[0027] Among them, α 1 , α 2 , α 3 and α 4 Respectively and The weight of .
[0028] Preferably, the calculation process of the comprehensive miRNA similarity is:
[0029] Create two matrices corresponding to the two types of miRNA similarities, each matrix is used to store the corresponding miRNA similarity. The two matrices are and The sizes of the two matrices are ms×ms, where ms represents the number of miRNAs. and Used to indicate the similarity of corresponding matrices at coordinates (i, j);
[0030] The miRNA similarities of the four matrices were integrated by a weighted combination strategy, and the comprehensive miRNA similarity matrix S was obtained after integration. m ,
[0031]
[0032] Among them, β 1 and β 2 Respectively and The weight of .
[0033] Preferably, after Gaussian radial basis function performs matrix enhancement on the comprehensive SM similarity, a composite similarity matrix S is obtained. SM ,
[0034]
[0035] Among them, R i and R j They are the comprehensive SM similarity matrix S sm The i-th row vector and the j-th row vector of , σ represents the parameter for adjusting the function bandwidth;
[0036] After Gaussian radial basis function is used to enhance the matrix of comprehensive miRNA similarity, the similarity matrix S is obtained. M ,
[0037]
[0038] Among them, R′ i and R′ j The comprehensive miRNA similarity matrix S m The i-th row vector and the j-th row vector of , σ represents the parameter for adjusting the function bandwidth.
[0039] Preferably, the MMA matrix is the target matrix H∈R ns×nm , by calculating and target matrix H∈R ns×nm The matrix X∈R with the same observation value ns×nm Complete the missing data of the MMA matrix;
[0040] In step 2, the process of truncated Schatten p-norm to fill missing data is:
[0041] Step 21, construct an objective function of the truncated schatten p-norm of the efficient approximate rank function;
[0042] Step 22, obtaining the missing data of the SM-miRNA association matrix through the objective function of the truncated schatten p-norm of the constructed efficient approximate rank function;
[0043] Among them, in step 21, the construction process of the objective function of the truncated schatten p-norm of the efficient approximate rank function is:
[0044] Matrix X∈R ns×nm The mathematical construction form is:
[0045] min X rank(X),
[0046] sP Ω (X) = P Ω (H) (4),
[0047] Among them, rank(·) is the rank function, is the location set corresponding to SM-miRNA, P Ω is the projection operator of Ω;
[0048]
[0049] Equation (4) is transformed into:
[0050]
[0051] sP Ω (X) = P Ω (H) (6),
[0052] By the truncated Schatten p-norm lemma in the matrix X∈R ns×nm The rank S (S ≤ min (ns, nm)) and the matrix X∈R ns×nm Singular value decomposition (SVD) asX=U△V T Solve the matrix X∈R ns×nm The optimal solution of
[0053]
[0054] AA T =I r×r ,BB T =I r×r (7)
[0055] Through a model for solving inequality constraints, the model is:
[0056]
[0057] Where λ is the balance coefficient, 0≤X i,j ≤1, 0≤i≤ns, 0≤j≤nm;
[0058] In step 22, the specific process of obtaining the missing data of the SM-miRNA association matrix through the objective function of the truncated schatten p-norm of the efficient approximate rank function is:
[0059] By alternating direction multiplication method and auxiliary matrix T∈R ns×nm Solve for missing data, matrix T∈R ns×nm The mathematical construction form is:
[0060]
[0061] stX=T,0≤X ij ≤1, 0≤i≤ns, 0≤j≤nm (10),
[0062] The augmented Lagrangian form of formula (10) is expressed as:
[0063]
[0064] Where E represents the Lagrange multiplier, η represents the penalty parameter, and the minimization of formula (11) is an iterative calculation process;
[0065] Solve T by iterative algorithm k+1 , X k+1 and E k+1 , in the k-th iteration, T k+1 , X k+1 and E k+1 Calculate in sequence, when the convergence condition and Stop the iterative calculation when the final T k+1 , X k+1 and E k+1 , and according to the final T k+1 , X k+1 and E k+1 Update the incidence matrix H * ;
[0066] Among them, the final T k+1 , X k+1 and E k+1 The iterative calculation process is:
[0067] The optimized T is obtained by iterative algorithm k+1 , X k+1 and E k+1 :
[0068] T k+1 =argminL(T,X k ,E k ,λ,η) (12),
[0069] X k+1 =argminL(T k+1 ,X,E k ,λ,η) (13),
[0070] E k+1 =argminL(T k+1 ,X k+1 ,E,λ,η) (14),
[0071] By T k+1 Compute closed-form solutions
[0072]
[0073] By applying the interval [0,1] range constraint to the value in equation (15), we get the final
[0074] Based on the singular value contraction operator, we obtain T k+1 :
[0075]
[0076] By changing the Lagrange multiplier E k+1 Update the auxiliary matrix T k+1 and the target matrix X k+1 , where E is calculated by the gradient ascent algorithm k+1 ,
[0077] E k+1 =E k +η(X k+1 -T k+1 ) (19).
[0078] Preferably, truncated matrix decomposition is used to decompose the MMA matrix into two feature matrices, the two feature matrices are: matrix A and matrix B, matrix H * ≈AB T , and extract key features by cutting off the number of singular values;
[0079] In step 3, the process of truncating the updated drug-miRNA matrix in the matrix decomposition step 2 is:
[0080] Step 31, constructing the objective function of truncated matrix decomposition through matrix A and matrix B;
[0081] Step 32: By * ≈AB T Use the SVD method to obtain the initial values of the variables;
[0082] Step 33: Optimize the objective function constructed in step 31 by the alternating least squares method. After optimization, iteratively update matrix A and matrix B until convergence. After convergence, obtain the final matrix A and matrix B, and obtain the final prediction score matrix based on the final matrix A and matrix B.
[0083] Among them, in step 31, the objective function of the truncated matrix decomposition constructed by matrix A and matrix B is:
[0084]
[0085] Here, ||·|| represents the Frobenius norm, λ t , s and λ m are all non-negative parameters, Denotes the matrix H * An approximate model of is used to identify the latent feature matrix A and matrix B; Represents the Tikhonov regularization requirement, minimizing the norm of matrix A and matrix B to prevent overfitting; and Both represent regularization requirements, which are used to ensure that the latent feature vectors of similar SMs and miRNAs are similar;
[0086] In step 32, the formula for obtaining the initial value of the variable by the SVD method is:
[0087]
[0088] Among them, S k Contains H * The first k largest singular values, U∈R k*ns and V∈R k*nm All contain corresponding related singular vectors;
[0089] In step 33, let equation (20) be L, and The iterative update formula of matrix A and matrix B is:
[0090] A=(H * B+λ s S SM A)(B T B+λ t I k +λ s A T A) -1 (twenty two),
[0091] B=(H * A+λ m S M B)(A T A+λ t I k +λ m B T B) -1 (twenty three),
[0092] Among them, I k is a diagonal matrix of dimension k.
[0093] In summary, the beneficial effects of the present invention are:
[0094] 1. Through GRBF, matrix enhancement is performed. Because the structural information of the comprehensive similarity matrix is taken into account, the accuracy of the similarity measure can be improved. Because the truncated schatten p-norm takes into account the physical properties of singular values, it can more effectively supplement the missing values than other rank approximation norms. In addition, the truncated matrix decomposition greatly improves the speed of the model by truncating the number of singular values. By cleverly combining the matrix completion method of truncated schatten p-norm minimization and truncated matrix decomposition, the feature selectivity and robustness of the former method are fully utilized, and the efficient calculation and interpretability of the latter method are utilized, thus showing comprehensive and flexible performance in MMA prediction.
[0095] 2. By performing matrix enhancement on the comprehensive SM similarity and the comprehensive SM similarity, the final refined SM similarity and the SM similarity can be obtained, thereby improving the accuracy of the similarity measure.
[0096] It is to be understood that the foregoing general description and the following detailed description are exemplary only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0097] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0098] Figure 1 It is a schematic diagram of the process of the present invention;
[0099] Figure 2 It is a schematic diagram of the enhanced integral similarity matrix in the present invention;
[0100] Figure 3 A schematic diagram of completing missing data in the MMA matrix in the present invention;
[0101] Figure 4 A schematic diagram of truncated matrix decomposition in the present invention; DETAILED DESCRIPTION
[0102] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0103] In a possible embodiment, as shown in the attached Figure 1-4 As shown, the SM-miRNA association prediction method based on matrix enhancement and collaborative double matrix completion includes the following steps:
[0104] Step 1, respectively enhancing the precision of the comprehensive SM similarity matrix and the comprehensive miRNA similarity matrix by using Gaussian radial basis function;
[0105] Step 2, obtain a probability value in the range of [0,1] by minimizing the truncated schatten p-norm, and use the obtained probability value to replace the missing data of the SM-miRNA association matrix to update the SM-miRNA association matrix;
[0106] Step 3: Perform truncated matrix decomposition on the SM-miRNA association matrix updated in step 2, and after decomposition, combine the enhanced comprehensive SM similarity matrix and comprehensive miRNA similarity matrix in step 1 to obtain a prediction score matrix, and predict the results based on the prediction score matrix.
[0107] By performing matrix enhancement through GRBF, the accuracy of similarity measurement can be improved because the structural information of the comprehensive similarity matrix is taken into account; the truncated schatten p-norm takes into account the physical properties of singular values and supplements missing values more effectively than other rank approximation norms; in addition, the truncated matrix decomposition greatly improves the speed of the model by truncating the number of singular values; in step 3, the singular value dimension is allowed to be modified, thereby reducing the rank of the prediction score matrix and improving the association probability; by cleverly combining the matrix completion method of truncated schatten p-norm minimization and truncated matrix decomposition, the feature selectivity and robustness of the former method are fully utilized, and the efficient calculation and interpretability of the latter method are utilized, thereby showing comprehensive and flexible performance in MMA prediction.
[0108] By cleverly combining the matrix completion method of truncated Schatten p-norm minimization and truncated matrix factorization, we fully utilize the feature selectivity and robustness of the former method and the efficient computation and interpretability of the latter method, thus demonstrating comprehensive and flexible performance in MMA prediction.
[0109] In step 1, the Gaussian radial basis function performs matrix enhancement on the comprehensive SM similarity in the comprehensive SM similarity matrix, and obtains the final refined SM similarity in the comprehensive SM similarity matrix after processing; the comprehensive SM similarity is obtained by four types of SM similarities, and the four types of SM similarities are obtained from the four types of SM similarity matrices in the comprehensive SM similarity matrix, and the four types of SM similarities include similarity based on SM side effects, similarity based on chemical structure, similarity based on gene function consistency, and similarity based on indicator phenotype;
[0110] Gaussian radial basis function performs matrix enhancement on the comprehensive miRNA similarity in the comprehensive miRNA similarity matrix, and obtains the final refined miRNA similarity in the comprehensive miRNA similarity matrix after processing; the comprehensive miRNA similarity is obtained through two types of miRNA similarities, and the two types of miRNA similarities are obtained from two miRNA similarity matrices in the comprehensive miRNA similarity matrix, and the two types of miRNA similarities include similarity based on gene function consistency and similarity based on disease phenotype.
[0111] The four types of SM similarities and two types of miRNA similarities were calculated based on the benchmark miRNA dataset, which contains 796 known MMAs, 831 SMs, and 541 miRNA datasets.
[0112] In this specific embodiment, three data sets were used to evaluate the predictive performance of MECDMC. Data set 1 contained 831 SMs, 541 miRNAs, and 664 known MMAs. In data set 2, SMs and miRNAs without known MMAs in data set 1 were excluded, resulting in 39 SMs, 286 miRNAs, and 664 known MMAs. These two data sets share the same 664 known MMAs from SM2miR v1.0. Because it is not clear whether the 664 known MMAs widely used in experiments are sensitive, we constructed a new data set including 796 known MMAs, 831 SMs, and 541 miRNAs.
[0113] Among them, the calculation process of comprehensive SM similarity is:
[0114] Create 4 matrices corresponding to the four types of SM similarities. Each matrix is used to store the corresponding SM similarity. The 4 matrices are and The sizes of the four matrices are ns×ns, where ns represents the number of small molecules. and Used to indicate the similarity of corresponding matrices at coordinates (i, j);
[0115] The SM similarities of the four matrices are integrated through a weighted combination strategy, and the comprehensive SM similarity matrix S is obtained after integration. sm ,
[0116]
[0117] Among them, α 1 , α 2 , α 3 and α 4 Respectively and The weight of 1 , α 2 , α 3 and α 4 Both are 1.
[0118] Among them, the calculation process of comprehensive miRNA similarity is:
[0119] Create two matrices corresponding to the two types of miRNA similarities, each matrix is used to store the corresponding miRNA similarity. The two matrices are and The sizes of the two matrices are ms×ms, where ms represents the number of miRNAs. and Used to indicate the similarity of corresponding matrices at coordinates (i, j);
[0120] The miRNA similarities of the four matrices were integrated by a weighted combination strategy, and the comprehensive miRNA similarity matrix S was obtained after integration. m ,
[0121]
[0122] Among them, β 1 and β 2 Respectively and The weight, in this specific embodiment, is 1 and β 2 Both are 1.
[0123] After Gaussian radial basis function performs matrix enhancement on the comprehensive SM similarity, the composite similarity matrix S is obtained. SM ,
[0124]
[0125] Among them, R i and R j They are the comprehensive SM similarity matrix S smThe i-th row vector and the j-th row vector of , σ represents the parameter for adjusting the function bandwidth;
[0126] After Gaussian radial basis function is used to enhance the matrix of comprehensive miRNA similarity, the similarity matrix S is obtained. M ,
[0127]
[0128] Among them, R′ i and R′ j The comprehensive miRNA similarity matrix S m The i-th row vector and the j-th row vector of , σ represents a parameter for adjusting the function bandwidth. In this specific embodiment, σ is 2.
[0129] The MMA matrix is the target matrix H∈R ns×nm , by calculating and target matrix H∈R ns×nm The matrix X∈R with the same observation value ns×nm Complete the missing data of the MMA matrix;
[0130] In step 2, the process of truncated Schatten p-norm to fill missing data is:
[0131] Step 21, construct an objective function of the truncated schatten p-norm of the efficient approximate rank function;
[0132] Step 22: Obtain the missing data of the SM-miRNA association matrix through the objective function of the truncated schatten p-norm of the constructed efficient approximate rank function.
[0133] Among them, in step 21, the construction process of the objective function of the truncated schatten p-norm of the efficient approximate rank function is:
[0134] Matrix X∈R ns×nm The mathematical construction form is:
[0135] min X rank(X),
[0136] sP Ω (X) = P Ω (H) (4),
[0137] Among them, rank(·) is the rank function, is the location set corresponding to SM-miRNA, P Ω is the projection operator of Ω;
[0138]
[0139] Equation (4) is transformed into:
[0140]
[0141] sP Ω (X) = P Ω (H) (6),
[0142] By the truncated Schatten p-norm lemma in the matrix X∈R ns×nm The rank S (S ≤ min (ns, nm)) and the matrix X∈R ns×nm Singular value decomposition (SVD) asX=U△V T Solve the matrix X∈R ns×nm The optimal solution of
[0143]
[0144] AA T =I r×r ,BB T =I r×r (7)
[0145] Through a model for solving inequality constraints, the model is:
[0146]
[0147] Where λ is the balance coefficient, 0≤X i,j ≤1, 0≤i≤ns, 0≤j≤nm;
[0148] In step 22, the specific process of obtaining the missing data of the SM-miRNA association matrix through the objective function of the truncated schatten p-norm of the efficient approximate rank function is:
[0149] By alternating direction multiplication method and auxiliary matrix T∈R ns×nm Solve for missing data, matrix T∈R ns×nm The mathematical construction form is:
[0150]
[0151] stX=T,0≤X ij ≤1, 0≤i≤ns, 0≤j≤nm (10),
[0152] The augmented Lagrangian form of formula (10) is expressed as:
[0153]
[0154] Where E represents the Lagrange multiplier, η represents the penalty parameter, and the minimization of formula (11) is an iterative calculation process;
[0155] Solve T by iterative algorithm k+1 , X k+1 and E k+1 , in the k-th iteration, T k+1 , X k+1 and E k+1 Calculate in sequence, when the convergence condition and Stop the iterative calculation when the final T k+1 , X k+1 and E k+1 , and according to the final T k+1 , X k+1 and E k+1 Update the incidence matrix H * ;
[0156] Among them, the final T k+1 , X k+1 and E k+1 The iterative calculation process is:
[0157] The optimized T is obtained by iterative algorithm k+1 , X k+1 and E k+1 :
[0158] T k+1 =argminL(T,X k ,E k ,λ,η) (12),
[0159] X k+1 =argminL(T k+1 ,X,E k ,λ,η) (13),
[0160] E k+1 =argminL(T k+1 ,X k+1 ,E,λ,η) (14),
[0161] By T k+1 Compute closed-form solutions
[0162]
[0163] By applying the interval [0,1] range constraint to the value in equation (15), we get the final
[0164] Based on the singular value contraction operator, we obtain T k+1:
[0165]
[0166] By changing the Lagrange multiplier E k+1 Update the auxiliary matrix T k+1 and the target matrix X k+1 , where E is calculated by the gradient ascent algorithm k+1 ,
[0167] E k+1 =E k +η(X k+1 -T k+1 ) (19).
[0168] Truncated matrix decomposition is used to decompose the MMA matrix into two feature matrices: matrix A and matrix B, matrix H * ≈AB T , and extract key features by cutting off the number of singular values;
[0169] In step 3, the process of truncating the updated drug-miRNA matrix in the matrix decomposition step 2 is:
[0170] Step 31, constructing the objective function of truncated matrix decomposition through matrix A and matrix B;
[0171] Step 32: By * ≈AB T Use the SVD method to obtain the initial values of the variables;
[0172] Step 33: Optimize the objective function constructed in step 31 by the alternating least squares method. After optimization, iteratively update matrix A and matrix B until convergence. After convergence, obtain the final matrix A and matrix B, and obtain the final prediction score matrix based on the final matrix A and matrix B.
[0173] Among them, in step 31, the objective function of the truncated matrix decomposition constructed by matrix A and matrix B is:
[0174]
[0175] Here, ||·|| represents the Frobenius norm, λ t , s and λ m are all non-negative parameters, Denotes the matrix H * An approximate model of is used to identify the latent feature matrix A and matrix B; Represents the Tikhonov regularization requirement, minimizing the norm of matrix A and matrix B to prevent overfitting; and Both represent regularization requirements to ensure that the latent feature vectors of similar SMs and miRNAs are similar.
[0176] In step 32, the formula for obtaining the initial value of the variable by the SVD method is:
[0177]
[0178] Among them, S k Contains H * The first k largest singular values, U∈R k*ns and V∈R k*nm All contain corresponding related singular vectors;
[0179] In step 33, let equation (20) be L, and The iterative update formula of matrix A and matrix B is:
[0180] A=(H * B+λ s S SM A)(B T B+λ t I k +λ s A T A) -1 (twenty two),
[0181] B=(H * A+λ m S M B)(A T A+λ t I k +λ m B T B) -1 (twenty three),
[0182] Among them, I k is a diagonal matrix of dimension k.
[0183] Those skilled in the art will readily appreciate other embodiments of the present invention after considering the specification and practicing the invention herein. This application is intended to cover any variations, uses or adaptations of the present invention, which follow the general principles of the present invention and include common knowledge or customary techniques in the art that the present invention did not invent. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present invention are indicated by the appended claims.
[0184] It should be understood that the present invention is not limited to the exact construction that has been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.
Claims
1. A SM-miRNA association prediction method based on matrix enhancement and collaborative double matrix completion, characterized in that: The following steps are involved: Step 1, respectively enhancing the precision of the comprehensive SM similarity matrix and the comprehensive miRNA similarity matrix by using Gaussian radial basis function; Step 2, obtain a probability value in the range of [0,1] by minimizing the truncated schatten p-norm, and use the obtained probability value to replace the missing data of the SM-miRNA association matrix to update the SM-miRNA association matrix; Step 3, performing truncated matrix decomposition on the SM-miRNA association matrix updated in step 2, and after decomposition, combining the enhanced comprehensive SM similarity matrix and the comprehensive miRNA similarity matrix in step 1 to calculate, obtain a prediction score matrix, and predict the results according to the prediction score matrix; After Gaussian radial basis function performs matrix enhancement on the comprehensive SM similarity, the composite similarity matrix S is obtained. SM , Among them, R i and R j They are the comprehensive SM similarity matrix S sm The i-th row vector and the j-th row vector of , σ represents the parameter for adjusting the function bandwidth; After Gaussian radial basis function is used to enhance the matrix of comprehensive miRNA similarity, the similarity matrix S is obtained. M , Among them, R′ i and R′ j The comprehensive miRNA similarity matrix S m The i-th row vector and the j-th row vector of , σ represents the parameter for adjusting the function bandwidth; Truncated matrix decomposition is used to decompose the MMA matrix into two feature matrices: matrix A and matrix B, matrix H * ≈AB T , and extract key features by cutting off the number of singular values; In step 3, the process of truncating the updated drug-miRNA matrix in the matrix decomposition step 2 is: Step 31, constructing the objective function of truncated matrix decomposition through matrix A and matrix B; Step 32: By * ≈AB T Use the SVD method to obtain the initial values of the variables; Step 33: Optimize the objective function constructed in step 31 by the alternating least squares method. After optimization, iteratively update matrix A and matrix B until convergence. After convergence, obtain the final matrix A and matrix B, and obtain the final prediction score matrix based on the final matrix A and matrix B. Among them, in step 31, the objective function of the truncated matrix decomposition constructed by matrix A and matrix B is: Here, ||·|| represents the Frobenius norm, λ t , s and λ m are all non-negative parameters, Denotes the matrix H * An approximate model of is used to identify the latent feature matrix A and matrix B; Represents the Tikhonov regularization requirement, minimizing the norm of matrix A and matrix B to prevent overfitting; and Both represent regularization requirements to ensure that the latent feature vectors of similar SMs and miRNAs are similar.
2. The SM-miRNA association prediction method based on matrix enhancement and collaborative double matrix completion according to claim 1, characterized in that: In step 1, the Gaussian radial basis function performs matrix enhancement on the comprehensive SM similarity in the comprehensive SM similarity matrix, and obtains the final refined SM similarity in the comprehensive SM similarity matrix after processing; The comprehensive SM similarity is obtained through four types of SM similarities, and the four types of SM similarities are obtained from four SM similarity matrices in the comprehensive SM similarity matrix, and the four types of SM similarities include similarity based on SM side effects, similarity based on chemical structure, similarity based on gene function consistency, and similarity based on indicator phenotype; The Gaussian radial basis function performs matrix enhancement on the comprehensive miRNA similarity in the comprehensive miRNA similarity matrix, and obtains the final refined miRNA similarity in the comprehensive miRNA similarity matrix after processing; the comprehensive miRNA similarity is obtained through two types of miRNA similarities, and the two types of miRNA similarities are obtained from two miRNA similarity matrices in the comprehensive miRNA similarity matrix, and the two types of miRNA similarities include similarity based on gene function consistency and similarity based on disease phenotype.
3. The SM-miRNA association prediction method based on matrix enhancement and collaborative double matrix completion according to claim 2, characterized in that: The four types of SM similarities and the two types of miRNA similarities are calculated based on a benchmark miRNA dataset, which includes 796 known MMAs, 831 SMs, and 541 miRNA datasets.
4. The SM-miRNA association prediction method based on matrix enhancement and collaborative double matrix completion according to claim 2, characterized in that: The calculation process of comprehensive SM similarity is: Create 4 matrices corresponding to the four types of SM similarities. Each matrix is used to store the corresponding SM similarity. The 4 matrices are and The sizes of the four matrices are ns×ns, where ns represents the number of small molecules. and Used to indicate the similarity of corresponding matrices at coordinates (i, j); The SM similarities of the four matrices are integrated through a weighted combination strategy, and the comprehensive SM similarity matrix S is obtained after integration. sm , Among them, α1, α2, α3 and α4 represent and The weight of .
5. The SM-miRNA association prediction method based on matrix enhancement and collaborative double matrix completion according to claim 4, characterized in that: The calculation process of comprehensive miRNA similarity is: Create two matrices corresponding to the two types of miRNA similarities, each matrix is used to store the corresponding miRNA similarity. The two matrices are and The sizes of the two matrices are ms×ms, where ms represents the number of miRNAs. and Used to indicate the similarity of corresponding matrices at coordinates (i, j); The miRNA similarities of the four matrices were integrated by a weighted combination strategy, and the comprehensive miRNA similarity matrix S was obtained after integration. m , Among them, β1 and β2 represent and The weight of .
6. The SM-miRNA association prediction method based on matrix enhancement and collaborative double matrix completion according to claim 1, characterized in that: The MMA matrix is the target matrix H∈R ns×nm , by calculating and target matrix H∈R ns×nm The matrix X∈R with the same observation value ns×nm Complete the missing data of the MMA matrix; In step 2, the process of truncated Schatten p-norm to fill missing data is: Step 21, construct an objective function of the truncated schatten p-norm of the efficient approximate rank function; Step 22, obtaining the missing data of the SM-miRNA association matrix through the objective function of the truncated schatten p-norm of the constructed efficient approximate rank function; Among them, in step 21, the construction process of the objective function of the truncated schatten p-norm of the efficient approximate rank function is: Matrix X∈R ns×nm The mathematical construction form is: min X rank(X), s.t.P Ω (X)=P Ω (H) (4), Among them, rank(·) is the rank function, is the set of positions corresponding to SM-miRNA, P Ω is the projection operator of Ω; Equation (4) is transformed into: s.t.P Ω (X)=P Ω (H) (6), By the truncated Schatten p-norm lemma in the matrix X∈R ns×nm The rank S (S ≤ min (ns, nm)) and the matrix X∈R ns×nm Singular value decomposition (SVD) as X=U△V T Solve the matrix X∈R ns×nm The optimal solution of s.t.AA T =I r×r ,BB T =I r×r (7), Through a model for solving inequality constraints, the model is: Where λ is the balance coefficient, 0≤X i,j ≤1, 0≤i≤ns, 0≤j≤nm.
7. The SM-miRNA association prediction method based on matrix enhancement and collaborative double matrix completion according to claim 6, characterized in that: In step 22, the specific process of obtaining the missing data of the SM-miRNA association matrix through the objective function of the truncated schatten p-norm of the efficient approximate rank function is: By alternating direction multiplication method and auxiliary matrix T∈R ns×nm Solve for missing data, matrix T∈R ns×nm The mathematical construction form is: s.t.X=T,0≤X ij ≤1,0≤i≤ns,0≤j≤nm (10), The augmented Lagrangian form of formula (10) is expressed as: Where E represents the Lagrange multiplier, η represents the penalty parameter, and the minimization of formula (11) is an iterative calculation process; Solve T by iterative algorithm k+1 , X k+1 and E k+1 , in the k-th iteration, T k+1 , X k+1 and E k+1 Calculate in sequence, when the convergence condition and Stop the iterative calculation when the final T k+1 , X k+1 and E k+1 , and according to the final T k+1 , X k+1 and E k+1 Update the incidence matrix H * ; Among them, the final T k+1 , X k+1 and E k+1 The iterative calculation process is: The optimized T is obtained by iterative algorithm k+1 , X k+1 and E k+1 : T k+1 =argminL(T,X k ,E k ,λ,η) (12), X k+1 =argminL(T k+1 ,X,E k ,λ,η) (13), From k+1 =argminL(T k+1 ,X k+1 ,E,λ,η) (14), By T k+1 Compute closed-form solutions By applying the interval [0,1] range constraint to the value in equation (15), we can obtain the final Based on the singular value contraction operator, we obtain T k+1 : By changing the Lagrange multiplier E k+1 Update the auxiliary matrix T k+1 and the target matrix X k+1 , where E is calculated by the gradient ascent algorithm k+1 , AND k+1 =E k +η(X k+1 -T k+1 ) (19)。 8. The SM-miRNA association prediction method based on matrix enhancement and collaborative double matrix completion according to claim 1, characterized in that: In step 32, the formula for obtaining the initial value of the variable by the SVD method is: Among them, S k Contains H * The first k largest singular values, U∈R k*ns and V∈R k*nm All contain corresponding related singular vectors; In step 33, let equation (20) be L, and The iterative update formula of matrix A and matrix B is: A=(H * B+λ s S SM A)(B T B+λ t I k +λ s A T A) -1 (22), B=(H * A+λ m S M B)(A T A+λ t I k +λ m B T B) -1 (23), Among them, I k is a diagonal matrix of dimension k.
Citation Information
Patent Citations
Micromolecule-miRNA interaction prediction method for minimizing truncated schatten-p norm
CN117153244A