A method for identifying prognostic markers based on multi-source biological information fusion
By integrating multi-source biological information, building a double-layer heterogeneous network and denoising it, the false positives and incomplete information when a single network recognizes cancer prognostic markers are solved, and higher recognition accuracy and biological explanatory ability are achieved.
Patent Information
- Application Number
- CN202211221837.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-08
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-10-08
AI Technical Summary
The prior art has problems with false positives and incomplete information when identifying cancer prognostic markers, resulting in insufficient identification accuracy.
Using a method based on multi-source biological information fusion, multiple biological networks are integrated, low-dimensional vectors of nodes are obtained through fast network embedding, a two-layer heterogeneous network is constructed, and a network enhancement method is used for denoising, so as to obtain the importance sorting of features through the network propagation algorithm.
Effectively reduces the problem of false positive and incomplete interaction in a single network, and improves the identification accuracy and biological explanatory ability of cancer prognostic markers.
Smart Images

Figure CN115662640B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of bioinformatics, and particularly relates to a method for identifying prognostic markers. Background Art
[0002] The pathogenesis of cancer comes from multiple factors, including genetics, age, and environment. The prognosis of cancer is very challenging. Even for cancer patients with the same pathological type and clinical stage, there are significant differences in prognosis even after the same treatment. Therefore, identifying cancer prognostic markers is an urgent problem to be solved. Prognostic markers can not only help predict the clinical prognosis of patients, but also contribute to the study of the molecular mechanism of cancer.
[0003] Network-based methods have been widely introduced to search for prognostic markers. Many methods integrate a single biological network and gene expression profiles to identify prognostic markers. For example, the paper "Y. Cun and H. "Network and data integration for biomarker signature discovery via network smoothed t-statistics," PloS one, vol. 8, no. 9, p. e73074, 2013.” proposed a feature selection method stSVM based on support vector machines to obtain effective biomarkers as features for distinguishing samples with different labels. However, there are problems such as a large number of false positives and incomplete network information in a single network. Therefore, identifying biomarkers based on a single network may not be accurate. Although some biomarker identification methods combine multiple biological network information, for example, the paper "X. Wang, S.-S. Wang, L. Zhou, L. Yu, and L.-m. Zhang, "A network-pathway based module identification for predicting the prognosis of ovarian cancer patients," Journal of ovarian research, vol. 9, no. 1, pp. 1-8, 2016.” constructed a Reactome functional interaction network using gene expression profile data, including pathways, PPIs, gene ontology (GO) annotations, and gene co-expression. However, there is a lot of noise in current network data, and these methods do not consider network denoising, which may limit the accuracy of these methods.
[0004] Based on the deficiencies of current research, it is necessary to provide a method for identifying prognostic markers based on the fusion of multi-source biological information. Summary of the Invention
[0005] To overcome the deficiencies of the prior art, the present invention provides a method for identifying prognostic biomarkers based on multi-source biological information fusion, which integrates multiple biological networks, obtains low-dimensional vectors of nodes in the network through fast network embedding; and constructs a two-layer heterogeneous network using the low-dimensional vectors of proteins and genes; introduces a network enhancement method on the constructed two-layer heterogeneous network to denoise the network, and initially scores the network nodes from two perspectives: prognostic ability and degree of correlation with known pathogenic genes; finally, uses the network propagation algorithm on the denoised two-layer heterogeneous network to obtain the importance ranking of features in the network; the features ranked at the top are considered prognostic-related biomarkers. The present invention makes full use of multi-source biological network information and can effectively identify biomarkers with classification ability and biological interpretability for prognostic analysis of complex diseases.
[0006] The technical solution adopted by the present invention to solve its technical problems includes the following steps:
[0007] Step 1: Use the fast network embedding algorithm to obtain low-dimensional vectors of nodes in the multi-source biological network;
[0008] Step 2: Use the low-dimensional vectors of genes and proteins to construct a two-layer heterogeneous network through the Pearson correlation coefficient;
[0009] Step 3: Use the network enhancement algorithm to denoise the constructed two-layer heterogeneous network;
[0010] Step 4: Initially score the nodes of the two-layer heterogeneous network from two perspectives: prognostic ability and degree of correlation with known pathogenic genes;
[0011] Step 5: Through the network propagation algorithm, obtain the importance ranking of the nodes of the two-layer heterogeneous network; the features ranked at the top are determined as prognostic-related biomarkers.
[0012] Further, the specific content of Step 1 is as follows:
[0013] In the fast network embedding algorithm, first introduce the target similarity function Φ(A) ∈ R of the adjacency matrix A of the n-node network n×n , which is defined as a polynomial function of A; assume that Φ(A) is a positive semi-definite function, then it is expressed as: Φ(A) = S·S T , where S = α 0 I + α 1 A 1 + α 2 A 2 + … + α p A p , α 0 , α 1 , α 2 , …, α pis a predefined weight, and p is the number of terms; thus the target similarity function Φ(A) ∈ R n×n will be decomposed into the product of two low-dimensional matrices;
[0014] The objective function is expressed as: To minimize the objective function, the Gaussian random projection method is used to obtain the embedding U: U = S·Q = (α 0 I + α 1 A + α 2 A 2 + … + α p A p )Q, Q ∈ R n×d obeys the Gaussian distribution, so S is randomly projected into the low-dimensional subspace.
[0015] Furthermore, the specific content of step 2 is as follows:
[0016] Construct a gene and protein bilayer heterogeneous network through the Pearson correlation coefficient; the Pearson correlation coefficient is expressed as:
[0017]
[0018] where X and Y respectively represent the low-dimensional vector representations of genes and proteins, and represent the average values of X and Y respectively, and ρ X,Y ∈ (-1, 1);
[0019] The bilayer heterogeneous network H is expressed as: where M p represents the constructed protein matrix, M G represents the constructed gene matrix, M A represents the connection relationship between genes and proteins, and M B is the transpose matrix of M A .
[0020] Furthermore, the specific content of step 3 is as follows:
[0021] Denoise the constructed bilayer heterogeneous network H using the network enhancement algorithm, and the network enhancement algorithm is expressed as:
[0022]
[0023] where W represents the pairwise weights between nodes, represents the set of the K nearest neighbors of the i-th node, is the local structure of the constructed input network, where α is a regularization parameter and t represents the iteration step.
[0024] Furthermore, the specific content of step 4 is as follows:
[0025] The correlation scores of known pathogenic genes and diseases obtained from the database are used as the scoring values of nodes in the network, and the scoring values of genes not in the list of known pathogenic genes are zero;
[0026] The prognostic ability score of a node is obtained through the q-statistic of the expression values of the node in different prognostic samples, and the expression formula is:
[0027]
[0028] Among them, is the average of the two types of samples, is the variance of the two types of samples, and n 1 、n 2 is the capacity of the two types of samples.
[0029] Furthermore, the specific steps of step 5 are as follows:
[0030] Through the network propagation algorithm, obtain the importance ranking of nodes in the network. The network propagation algorithm is expressed as:
[0031] S h+1 =(1 - α)T·S h +αS 0
[0032] Among them, T represents the adjacency matrix of the fusion network; S h represents the probability of reaching a node at time step h; S 0 represents the initial probability of the node; α is the restart probability, and S h+1 <10 -6 ;
[0033] Through the network propagation algorithm, obtain the scores of nodes in the network and sort them. Biomarkers ranked among the top are determined as biomarkers related to disease prognosis.
[0034] Furthermore, p = 3, α 0 , α 1 , α 2 , α 3 =[1, 10 2 , 10 4 , 10 5 , d = 128.
[0035] Furthermore, the restart probability α is equal to 0.7.
[0036] The beneficial effects of the present invention are as follows:
[0037] The present invention integrates multi-source biological information networks, which can effectively reduce the problems of false positives and incomplete interactions in a single network and can also introduce more biological information. And a network enhancement algorithm is used to optimize the biological network, effectively improving the classification accuracy of samples with different prognostic effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 Framework diagram of the method of the present invention.
[0039] Figure 2 Schematic diagram of the classification ability of the method of the present invention and the CPR, NetRank, MarkRank, and stSVM methods on six datasets.
[0040] Figure 3 Comparison diagram of the method of the present invention between a bilayer heterogeneous network and a single network.
[0041] Figure 4 Comparison diagram of the enrichment analysis of known pathogenic genes and differentially expressed genes between the method of the present invention and the CPR, NetRank, MarkRank, and stSVM methods. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The present invention will be further described below in conjunction with the drawings and embodiments.
[0043] As Figure 1 shown, a method for identifying prognostic markers based on multi-source biological information fusion includes the following steps:
[0044] Step 1: Use a fast network embedding algorithm to obtain low-dimensional vectors of nodes in a multi-source biological network;
[0045] In the fast network embedding algorithm, first introduce the target similarity function Φ(A) ∈ R of the adjacency matrix A of an n-node network n×n , which is defined as a polynomial function of A; assuming that Φ(A) is a positive semi-definite function, it is expressed as: Φ(A) = S·S T , where S = α 0 I + α 1 A 1 + α 2 A 2 + … + α p A p , α 0 , α 1 , α 2 , …, α p are predefined weights, and p is the order; therefore, the target similarity function Φ(A) ∈ R n×n will be decomposed into the product of two low-dimensional matrices;
[0046] Express the objective function as: To minimize the objective function, the Gaussian random projection method is used to obtain the embedding U: U = S·Q = (α 0 I + α 1 A + α 2 A 2 + … + α p A p )Q, where Q ∈ R n×d follows a Gaussian distribution, so S is randomly projected into a low-dimensional subspace. The fast embedding parameter p = 3, α 0 , α 1 , α 2 , α 3 = [1, 10 2 , 10 4 , 10 5 , and d = 128.
[0047] Step 2: Use the low-dimensional vectors of genes and proteins to construct a two-layer heterogeneous network through the Pearson correlation coefficient;
[0048] Construct a two-layer heterogeneous network of genes and proteins through the Pearson correlation coefficient; the Pearson correlation coefficient is expressed as:
[0049]
[0050] where X and Y represent the low-dimensional vector representations of genes and proteins respectively, and represent the means of X and Y respectively, and ρ X,Y ∈ (-1, 1);
[0051] The two-layer heterogeneous network H is expressed as: where M p represents the constructed protein matrix, M G represents the constructed gene matrix, M A represents the connection relationship between genes and proteins, and M B is the transpose matrix of M A .
[0052] Step 3: Use the network enhancement algorithm to denoise the constructed two-layer heterogeneous network;
[0053] Use the network enhancement algorithm to denoise the constructed two-layer heterogeneous network H. The network enhancement algorithm is expressed as:
[0054]
[0055] where W represents the pairwise weights between nodes, represents the set of the K nearest neighbors of the i-th node, is the local structure of the constructed input network, where α is a regularization parameter and t represents the iteration step.
[0056] Step 4: Initially score the nodes of the double-layer heterogeneous network from the perspectives of prognostic ability and correlation with known pathogenic genes;
[0057] The correlation scores between known pathogenic genes and diseases obtained from the database are used as the scores of nodes in the network, and the scores of genes not in the list of known pathogenic genes are zero;
[0058] The prognostic ability score of the node is obtained by the q statistic of the expression value of the node in different prognostic samples, and the expression is:
[0059]
[0060] in, is the average of the two types of samples, is the variance of the two types of samples, n 1 、n 2 is the capacity of the two types of samples.
[0061] Step 5: Obtain the importance ranking of the nodes in the double-layer heterogeneous network through the network propagation algorithm; the top-ranked features are determined as prognosis-related biomarkers.
[0062] The importance ranking of nodes in the network is obtained through the network propagation algorithm, which is expressed as:
[0063] S h+1 =(1-α)T·S h +αS 0
[0064] Where T represents the adjacency matrix of the fusion network; S h represents the probability of reaching the node at time step h; S 0 represents the initial probability of the node; the restart probability α is equal to 0.7 and S t+1 <10 -6 ;
[0065] The scores of the nodes in the network are obtained and sorted through the network propagation algorithm, and the top-ranked biomarkers are determined to be biomarkers related to disease prognosis. Specific embodiment:
[0067] 1. Preprocessing of gene expression data;
[0068] Read the gene expression data file and standardize the gene expression data using the Z score:
[0069]
[0070] x represents the original gene expression value of each sample; μ represents the mean of all the original gene expression data of each sample; σ is the standard deviation of all the original gene expression data of each sample.
[0071] II. Integrating multi-source biological networks and network embedding;
[0072] Integrate multi-source biological networks through the adjacency matrix, and use the fast network embedding algorithm to obtain the low-dimensional vector representation of the nodes in the network. In the fast network embedding algorithm, A is the adjacency matrix of a network with n nodes, which is defined as a polynomial function, and the target similarity function is Φ(A) ∈ R n×n Assume that Φ(A) is a positive semi-definite function and can be expressed as:
[0073] Φ(A) = S·S T (2)
[0074] S = α 0 I + α 1 A 1 + α 2 A 2 +…+ α p A p (3)
[0075] α 0 ,α 1 ,α 2 ,…,α p are predefined weights, A is a symmetric matrix, and p is the order. Therefore, the target similarity function Φ(A) ∈ R n×n will be decomposed into the product of two low-dimensional matrices. The objective function can be expressed as:
[0076]
[0077] To minimize the objective function, use the Gaussian random projection method, through which the embedding U can be obtained:
[0078] U = S·Q = (α 0 I + α 1 A + α 2 A 2 +…+ α p A p )Q (5)
[0079] Q ∈ R n×d obeys the Gaussian distribution, so the adjacency matrix S is randomly projected into a low-dimensional subspace. In the fast embedding, the parameter p = 3, α 0 ,α 1 ,α 2 ,α 3 =[1,10 2,10 4 ,10 5 , d = 128.
[0080] III. Construct a two - layer heterogeneous network;
[0081] Construct a gene and protein network using the Pearson correlation coefficient, and the Pearson correlation coefficient can be expressed as:
[0082]
[0083] where X and Y represent the low - dimensional vector representations of genes and proteins respectively, and represent the average values of X and Y respectively, and ρ X,Y ∈(-1, 1); the two - layer heterogeneous network H can be expressed as:
[0084]
[0085] where M p represents the constructed protein matrix, M G represents the constructed gene matrix, M A represents the connection relationship between genes and proteins, and M B is the transpose matrix of M A .
[0086] IV. Use the network enhancement algorithm to denoise the constructed two - layer heterogeneous network;
[0087] The network enhancement algorithm can be expressed as:
[0088]
[0089] where W represents the pairwise weights between nodes, represents the set of K - nearest neighbors (KNN) of the i - th node, is the local structure of the constructed input network, where α is a regularization parameter and t represents the iteration step.
[0090] V. Perform initial scoring on network nodes;
[0091] Based on the fusion network, score the network nodes from two perspectives: prognostic ability and degree of correlation with known pathogenic genes.
[0092] 1) The correlation score between the known pathogenic genes obtained from the database and the disease is used as the scoring value of the nodes in the network, and the weights of genes not in the list of known pathogenic genes are zero;
[0093] 2) The prognostic ability score of the nodes is obtained through the q - statistic of the expression values of the nodes in different prognostic samples, and the expression is:
[0094]
[0095] is the average of the two types of samples, is the variance of the two types of samples, n 1 、n 2 is the capacity of the two types of samples.
[0096] VI. Obtaining biomarkers through Internet communication
[0097] The importance ranking of nodes in the network is obtained through the network propagation algorithm. The network propagation algorithm can be expressed as:
[0098] S h+1 =(1-α)T·S h +αS 0 (10)
[0099] Where T represents the adjacency matrix of the fusion network; where S h represents the probability of reaching the node at time step h; S 0 Represents the initial probability of the node; define S h+1 <10 -6 And the restart probability α is equal to 0.7; the scores of the nodes in the network are obtained and sorted through the network propagation algorithm, and the top-ranked biomarkers are considered to be biomarkers related to disease prognosis.
[0100] VII. Experimental Verification
[0101] In order to verify the effectiveness of this method, it was validated on six breast cancer datasets. Among them, there are four breast cancer data from the GEO database (https: / / www.ncbi.nlm.nih.gov / geo / ), namely GSE1456, GSE2034, GSE3494, and GSE4922, one breast cancer data from the TCGA database (https: / / portal.gdc.cancer.gov / projects), and one public dataset NKI for breast cancer patient survival analysis published by Van De Vijver et al. in the New England Journal of Medicine. The complex survival problem is converted into a simple binary classification problem. When the patient survives for more than 10 years, the sample is marked as having a good prognosis (for GSE1456, the patient survives for more than 5 years and is marked as having a good prognosis). If the patient survives for no more than 5 years, it is marked as having a poor prognosis. A total of 511 breast cancer samples with a good prognosis and 360 breast cancer samples with a poor prognosis are included.
[0102] In order to evaluate the classification accuracy and biological interpretability of this method, the following three analyses were performed:
[0103] (1) Accuracy of prognosis sample classification
[0104] For each breast cancer dataset, based on the biomarkers extracted by the method of the present invention and each method among NetRank, MarkRank, stSVM, and CPR, the accuracy of the method was evaluated through a random forest classifier and a five-fold cross-validation method; the experiment was repeated 100 times with five-fold cross-validation in order to obtain stable classification results. The AUC index was used to evaluate the classification results, and the AUC values of the present method and the comparative methods on six datasets are as Figure 2 shown. As can be seen from Figure 2 , the AUC values of the method of the present invention are significantly better than those of the comparative methods in all datasets. It can be seen that the method of the present invention has superior classification accuracy.
[0105] (2) Comparison between the single-network and fusion-network methods
[0106] To further verify the effectiveness of the multi-source fusion network, the method framework of the present invention was applied to each single network. The results are as Figure 3 shown. On all six datasets, the effect of any single network is not as good as that of the double-layer heterogeneous network, which proves that combining networks with different biological meanings is effective.
[0107] (3) Biological interpretability of prognosis biomarkers
[0108] To examine the biological interpretability of the biomarkers obtained by the method, the enrichment degrees of the obtained biomarkers for known pathogenic genes and differentially expressed genes were analyzed. For each gene in the gene expression data, a t-test was used to obtain differentially expressed genes (P value less than 0.01). The P value of the enrichment degree of known pathogenic genes and differentially expressed genes in the biomarkers was calculated through a hypergeometric test:
[0109]
[0110] where N is the number of all genes, M is the number of known pathogenic genes and differentially expressed genes among all genes, n is the number of biomarkers, and m is the number of known pathogenic genes and differentially expressed genes in the biomarkers. The smaller the P value, the higher the enrichment degree of known pathogenic genes and differentially expressed genes in the biomarkers. The results of -log 10 P based on six datasets are as Figure 4 shown. As can be seen from Figure 4 , the enrichment degree of the method of the present invention is very high, all far greater than 1.3, that is, the P value is less than 0.05, indicating a significant enrichment of known pathogenic genes and differentially expressed genes in the biomarkers obtained by the method of the present invention, that is, it has good biological interpretability.
Claims
1. A prognostic marker identification method based on multi-source biological information fusion, It is characterized in that The steps include: Step 1: Use a fast network embedding algorithm to obtain low-dimensional vectors of nodes in the multi-source biological network; Step 2: Use the low-dimensional vectors of genes and proteins to construct a two-layer heterogeneous network of genes and proteins through the Pearson correlation coefficient; the Pearson correlation coefficient is expressed as: wherein respectively represent the low-dimensional vector representations of genes and proteins, and respectively represent the average value of, ; Double-layer heterogeneous network It is expressed as: , where represents the constructed protein matrix, represents the constructed gene matrix, represents the connection relationship between genes and proteins, is the transposed matrix of Step 3: Use the network enhancement algorithm to denoise the constructed two-layer heterogeneous network; Step 4: Initially score the nodes of the double-layer heterogeneous network from the perspectives of prognostic ability and correlation with known pathogenic genes; Step 5: Obtain the importance ranking of the nodes in the double-layer heterogeneous network through the network propagation algorithm; the top-ranked features are determined as prognosis-related biomarkers.
2. A prognostic marker identification method based on multi-source biological information fusion according to claim 1, It is characterized in that The step 1 is specifically as follows: In the fast network embedding algorithm, first introduce n the target similarity function of the adjacency matrix A of the node network , which is defined as a polynomial function of A; assume is a positive semi - definite function, then it is expressed as: , where , , is the predefined weight, p is the number of terms; thus the target similarity function will be decomposed into the product of two low - dimensional matrices; The objective function is expressed as: , to minimize the objective function, the Gaussian random projection method is used to obtain the embedding : , obeys the Gaussian distribution, so S is randomly projected into the low-dimensional subspace.
3. The method for identifying prognostic markers based on multi-source biological information fusion according to claim 1, It is characterized in that The step 3 is specifically as follows: For the constructed two-layer heterogeneous network Use the network enhancement algorithm for denoising, and the network enhancement algorithm is expressed as: wherein represents the pairwise weight between nodes, represents the set of K nearest neighbors of the -th node, is the local structure of the input network, wherein is a regularization parameter, represents the iteration step.
4. The method for identifying prognostic markers based on multi-source biological information fusion according to claim 1, It is characterized in that The step 4 is specifically as follows: The correlation scores between known pathogenic genes and diseases obtained from the database are used as the scores of nodes in the network, and the scores of genes not in the list of known pathogenic genes are zero; The prognostic ability score of the node is obtained by the q statistic of the expression value of the node in different prognostic samples, and the expression is: Among them, is the average of two types of samples, , is the variance of two types of samples, , is the capacity of two types of samples.
5. The method for identifying prognostic markers based on multi-source biological information fusion according to claim 1, It is characterized in that The step 5 is specifically as follows: The importance ranking of nodes in the network is obtained through the network propagation algorithm, which is expressed as: Among them represents the adjacency matrix of the fusion network; represents the probability of reaching a node at time step h ; represents the initial probability of a node; is the restart probability, <10 -6 ; The scores of the nodes in the network are obtained and sorted through the network propagation algorithm, and the top-ranked biomarkers are determined to be biomarkers related to disease prognosis.
6. A prognostic marker identification method based on multi-source biological information fusion according to claim 2, It is characterized in that The p = 3, , = [1, 10 2 , 10 4 , 10 5 , d = 128.
7. The method for identifying prognostic markers based on multi-source biological information fusion according to claim 5, It is characterized in that The restart probability is equal to 0.7.
Citation Information
Patent Citations
Prognostic biomarker identification method based on fusion network and multiple scoring strategies
CN110010204A
Network marker identification method fusing multiple omics data
CN115019884A