A method for constructing patient survival network based on gene regulatory network
By assigning co-expression stability attributes in gene regulatory networks, evaluating the stability of genes in the network and constructing a survival network, we solved the stability and accuracy problems of traditional survival analysis, identified regulators that affect patient survival, and expanded the discovery of cancer prognostic markers.
Patent Information
- Application Number
- CN202310909405.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-24
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2043-07-24
AI Technical Summary
Traditional survival analysis methods have difficulty in accurately and stably ranking patients, and it is difficult to discover regulators related to survival, especially miRNAs, which makes it difficult to discover cancer prognostic markers.
A survival network was constructed based on the gene regulatory network. By assigning the co-expression stability attribute to the GRN nodes, the co-expression stability of genes in the network was evaluated. The Kaplan-Meier survival analysis was used to evaluate the survival differences of patients and construct a survival network.
It improves the stability and accuracy of patient survival analysis, can identify regulators that affect patient survival, expands the discovery of cancer prognostic markers, and reduces the randomness of the results.
Smart Images

Figure CN117116345B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of molecular biology and systems biology, and relates to a method for constructing a patient survival network based on a gene regulatory network. Background Art
[0002] In the study of complex diseases, survival analysis is widely used to identify disease markers related to patient survival and prognosis, thereby guiding disease screening, early diagnosis, and personalized medical decisions. Traditional survival analysis is mainly divided into two steps: first, patients are ranked according to the expression level of specific genes; then the log-rank test is used to assess whether there is a significant difference in the survival time of patients in the first and last half (or 1 / 4) of the ranking. Genes that are significantly associated with patient survival are called cancer survival genes, and they are often closely related to cancer development and prognosis. However, traditional survival analysis has two limitations:
[0003] 1) It is difficult to accurately and stably rank patients using gene expression levels. First, significant individual variability leads to a lack of comparability in gene expression levels across patients. Furthermore, complex in vivo and in vitro factors lead to a lack of stability in the expression levels of single genes.
[0004] 2) It is difficult to identify survival-related regulators (transcription factors and small RNAs) based on expression levels. First, many regulators (especially miRNAs) are expressed at very low levels in tumor tissues, making it difficult to accurately quantify them and rank patients based on their expression levels. In addition, many regulators affect target gene expression through mechanisms other than changes in expression levels (such as protein structure and microenvironment), thereby affecting cancer progression.
[0005] Genes do not function independently, but rather interact and collaborate within complex gene regulatory networks (GRNs). The edges of a GRN represent a variety of interactions and functional associations, such as physical interactions (DNA-DNA interactions, protein-DNA interactions, protein-protein interactions), genetic interactions (two or more genes associated with the same trait), and participation in the same biological process or signaling pathway. Compared to gene expression levels, GRNs have the following advantages:
[0006] 1) GRN reflects the stable functional association and regulatory architecture of genes in multiple patients and is less affected by individual differences;
[0007] 2) Compared with the expression level of a single gene, the network composed of multiple genes has a higher data dimension, which reduces the randomness of the results;
[0008] 3) Based on GRN, we can ignore the expression level of the regulator and instead use the target gene of the regulator to reversely infer its relationship with patient survival.
[0009] In summary, we believe that survival analysis based on GRN can effectively address the limitations of traditional survival analysis and significantly expand the discovery of cancer prognostic markers. Summary of the Invention
[0010] In view of the technical problems existing in the existing survival analysis methods, the object of the present invention is to provide a method for constructing a survival network based on a gene regulatory network. The present invention gives a new attribute to GRN nodes, called co-expression stability. It is known that genes connected to each other in GRN often have similar expression patterns (expression levels are high or low in multiple samples), a phenomenon called co-expression. Co-expressed genes are often functionally related or participate in the same biological process. Based on this feature, the co-expression stability of a gene in GRN represents the expression difference between the gene and all its adjacent genes (based on Z-Score standardization to ensure the comparability of expression levels of different genes). The smaller the expression difference, the higher the co-expression stability of the gene, and the functional module composed of it and the adjacent genes operates normally; the greater the expression difference, the lower the co-expression stability of the gene, and the functional module composed of it and the adjacent genes is out of balance. In summary, the co-expression stability of a gene is closely related to its functional stability. When the co-expression stability of a gene in different patients is significantly correlated with the patient's survival time, the gene is considered to play an important role in cancer progression.
[0011] Based on the above principles, we developed a survival analysis strategy based on a gene-restricted network (GRN). This method uses gene expression data (microarray data, RNA sequencing data, protein mass spectrometry data) and survival information from cancer patients (obtained through medical records and follow-up surveys, as well as large-scale cancer research projects such as TCGA). The main analysis steps include GRN construction, co-expression stability assessment, patient ranking, and assessment of survival differences.
[0012] Step 1) Gene expression data (also known as a gene expression matrix, where rows represent all genes, columns represent all patients, and values represent the expression level of a gene in a specific patient, including transcribed RNA levels or translated protein levels) are obtained experimentally or directly from public databases. Experimental methods include detecting RNA levels in biological samples using high-throughput sequencing or detecting protein levels in biological samples using mass spectrometry. Public databases include Gene Expression Omnibus (GEO), The Cancer Genome Atlas Program (TCGA), and ArrayExpress.
[0013] Step 2) Construct a GRN based on the gene expression matrix. Existing GRN inference methods mainly include clustering algorithms (hierarchical clustering, graph clustering, etc.), machine learning algorithms (Bayesian algorithm, random forest, etc.), and deep learning algorithms (convolutional neural network, transfer learning, etc.).
[0014] Step 3) Optimize the GRN using experimental methods or interaction databases. The goal is to remove edges with low confidence and retain only interactions that have been experimentally verified or included in public databases, thereby ensuring the accuracy of subsequent analysis. Experimental methods that can be used to optimize the GRN include: predicting transcription factor-target gene interactions based on co-immunoprecipitation, and predicting protein-protein interactions based on techniques such as yeast two-hybrid, close-range fluorescence resonance, surface plasmon resonance, and mass spectrometry. Interaction databases that can be used to optimize the GRN include: chromatin interaction databases (4DGenome), transcription factor-target gene databases (TRRUST and hTFtarget), small RNA-target gene databases (miRDB and miRTarBase), protein-protein interaction databases (STRING and HuRI), and pathway databases (KEGG and Reactome).
[0015] Step 4) Evaluate the co-expression stability of each node in the GRN (each node corresponds to a gene) across different patients. This involves first obtaining the neighboring genes for each gene in the GRN; performing Z-score normalization on the expression levels of each gene across all patients to ensure comparability of expression levels across different genes; and finally evaluating the co-expression stability of each gene based on its neighboring genes (see "Detailed Description" for details).
[0016] Step 5) For each gene in the gene expression matrix, the patients were ranked based on their co-expression stability. The survival information of the top quartile and bottom quartile of co-expression stability were analyzed using Kaplan-Meier survival analysis. The log-rank test P value for the gene was then used to assess whether there was a statistically significant difference in survival between the two groups. A P value of ≤ 0.05 indicated a statistically significant difference, indicating that the co-expression stability of the gene significantly affected patient survival; a P value of > 0.05 indicated that the co-expression stability of the gene did not affect patient survival.
[0017] Step 6) Retain genes with a log-rank test P ≤ 0.05, along with the edges connecting these genes in the GRN and the genes connected by these edges, and use the Cytoscape tool to construct a survival network for the target cancer. This network shows that the co-expression levels of connected nodes vary between patients, and these perturbations can significantly affect patient survival, which is of great research value.
[0018] Based on the above content, the technical solution of the present invention is:
[0019] A method for constructing a patient survival network based on a gene regulatory network, comprising the following steps:
[0020] 1) obtaining a gene expression matrix, wherein the rows of the gene expression matrix represent genes, the columns of the gene expression matrix represent target cancer patient samples, and the element value in the mth row and nth column of the gene expression matrix represents the expression level of the mth gene in the nth target cancer patient; obtaining patient survival information corresponding to each target cancer patient sample;
[0021] 2) constructing a gene regulatory network based on the gene expression matrix;
[0022] 3) For each edge in the gene regulatory network, if the credibility of the edge is lower than a set threshold, the edge is deleted;
[0023] 4) evaluating the co-expression stability of each gene in the gene regulatory network optimized in step 3) in each target cancer patient sample;
[0024] 5) For each gene in the gene expression matrix, sort the target cancer patient samples based on the co-expression stability of the gene in each target cancer patient sample, taking the survival information of the target cancer patient samples in the top T% of co-expression stability ranking as a first set of information, and taking the survival information of the target cancer patient samples in the bottom T% of co-expression stability ranking as a second set of information; performing a survival analysis based on the first and second sets of information to obtain a log-rank test value P for the gene; then, based on the log-rank test value P for the gene, determining whether the gene has a statistically significant effect on the survival time of each target cancer patient in the top T% of target cancer patient samples and the bottom T% of target cancer patient samples; if there is a statistically significant difference, retaining the gene;
[0025] 6) Based on the genes retained in step 5) and the edges and genes connecting the retained genes in the gene regulatory network,
[0026] Constructing a survival network for target cancers.
[0027] Furthermore, the method for obtaining the co-expression stability of each gene in each target cancer patient sample is as follows: first, the neighboring genes of each gene in the gene regulatory network are obtained; then, the expression level of each of the neighboring genes in all patients in the gene expression matrix is obtained and Z-Score normalized; and then, the co-expression stability of the gene in each target cancer patient sample is evaluated based on the Z-Score-normalized neighboring genes of each gene.
[0028] Furthermore, the method for obtaining the co-expression stability of each gene in each target cancer patient sample is as follows: for each gene g0 in the gene expression matrix, the M adjacent genes {g1,…,g M}; Normalize gene g0 and its adjacent genes, where Z i ={v i,1 ,…,v i,n} represents the i-th gene g i The expression value after Z-Score standardization is i∈[0,M]∩Z, where n represents the number of target cancer patient samples; when g0 and g i The expression pattern of is positively correlated, then the co-expression stability of gene g0 in target cancer patient sample j is otherwise where v i,j Indicates gene g i The expression level in target cancer patient sample j, j∈{1,…,n}.
[0029] Furthermore, if the log-rank test value of the gene is P≤0.05, it is determined that the gene has a statistically significant effect on the survival time of each target cancer patient in the top T% of the target cancer patient samples and the bottom T% of the target cancer patient samples.
[0030] Furthermore, the gene regulatory network is optimized using experimental means or an interaction database, and edges in the gene regulatory network whose credibility is lower than a set threshold are deleted.
[0031] Furthermore, the experimental methods include: prediction of transcription factor-target gene interactions based on immunoprecipitation, yeast two-hybrid, close-range fluorescence resonance, surface plasmon resonance, and mass spectrometry; the interaction databases include: chromatin interaction database, transcription factor-target gene database, small RNA-target gene database, protein-protein interaction database, and pathway database.
[0032] A server, characterized in that it includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing each step in the above method.
[0033] A computer-readable storage medium stores a computer program thereon, wherein the computer program implements the steps of the above method when executed by a processor.
[0034] The present invention has the following advantages:
[0035] 1) This approach addresses the lack of stability inherent in traditional survival analysis. Compared to gene expression levels, the topological features of genes in a GRN can more stably reflect a patient's physiological state. First, GRNs reflect stable functional associations and regulatory architectures across multiple patients, making them less susceptible to individual differences. Furthermore, compared to single-gene expression levels, multi-gene networks have a higher data dimensionality, reducing the randomness of the results.
[0036] 2) It is possible to reverse engineer the regulators (transcription factors and small RNAs) that drive cancer progression based on survival genes. We know that the co-expression levels of survival genes obtained based on the new method are different in different patients. Regulators are one of the main reasons for the co-expression of target genes. In other words, whether the regulators play a role will cause the co-expression levels of target genes in different patients to be different, thereby affecting patient survival. Therefore, the transcription factors or small RNAs targeted by survival genes are closely related to cancer progression and patient survival. The more survival genes are regulated, the more important the role of the transcription factor or small RNA in cancer, and the higher the credibility. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1The figure is a flow chart of constructing a gene regulatory network according to the present invention.
[0038] Figure 2 This is a flowchart of constructing a patient survival network according to the present invention.
[0039] Figure 3 Schematic diagram of the algorithm of the present invention. DETAILED DESCRIPTION
[0040] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0041] The process of the present invention is as follows Figure 1 、 Figure 2 As shown, assuming that the gene expression data of a target cancer includes m different genes and n patient samples, and the survival information of these n patients is obtained, the present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0042] Step 1) Construct the gene regulatory network of these m genes using bioinformatics. Available GRN inference methods include clustering algorithms, machine learning algorithms, and deep learning algorithms. Clustering algorithms include hierarchical clustering (WGCNA) and graph clustering (MCL); machine learning algorithms include Bayesian algorithms (BANJO, CLR) and random forests (GENIE3, ReNI); and deep learning algorithms include convolutional neural networks (DeepInsight, DeepFeature) and transfer learning (Geneformer).
[0043] Step 2) Optimize the GRN using experimental methods or interaction databases, retaining interactions with high confidence. Experimental methods include predicting transcription factor-target gene interactions based on co-immunoprecipitation, and predicting protein-protein interactions using techniques such as yeast two-hybrid, close-range fluorescence resonance, surface plasmon resonance, and mass spectrometry. Public databases include the chromatin interaction database 4DGenome, the transcription factor-target gene databases TRRUST and hTFtarget, the small RNA target gene databases miRDB and miRTarBase, the protein-protein interaction databases STRING and HuRI, and the pathway databases KEGG and Reactome.
[0044] Step 3) Obtain the neighboring genes for each gene in the GRN. Genes and their neighbors can be connected through various interactions, including physical interactions (DNA-DNA interactions, protein-DNA interactions, protein-protein interactions), genetic interactions (two or more genes are associated with the same trait), and co-regulation (targeting the same transcription factor or small RNA, or participating in the same biological process or signaling pathway).
[0045] Step 4) Perform Z-score normalization on the gene expression levels across patients. This aims to eliminate differences in gene expression and retain only the relative changes in gene expression across patients, thus making the expression levels of different genes comparable.
[0046] Step 5) Evaluate the co-expression stability of each gene with its neighboring genes. Co-expression stability indicates the degree of correlation between a gene and all its neighboring genes, which indirectly reflects the functional stability of the gene. Assume that gene g0 has M neighboring genes, which are represented as {g1,…,g M}. Z i ={v i,1 ,…,v i,n} represents gene g i (i∈[0,M]∩Z) is the expression level after Z-Score normalization, where n represents the number of patients. j∈{1,…,n} represents the patient number. The co-expression stability of gene g0 in patient j is calculated as where v i,j Indicates gene g i The expression level in patient j. i The expression patterns of 0,j -v i,j ) 2 The smaller the value, the closer g0 and g i The more significant the positive correlation is. i The expression patterns of the two genes are negatively correlated (Pearson correlation coefficient is less than 0). Their expression levels should be high and low, and after Z-Score standardization, they appear as positive and negative. Therefore, the expression in the brackets is (v 0,j +v i,j ) 2 The smaller the value, the closer g0 and g i The more significant the negative correlation is. In summary, S 0,j The smaller it is, the stronger the positive / negative correlation between g0 and its adjacent genes, and the higher the functional stability of g0. Finally, the co-expression stability of each gene in the gene expression matrix in each patient is obtained.
[0047] Step 6) Survival analysis based on co-expression stability. For each gene in the gene expression matrix, patients were ranked based on their co-expression stability in each patient. Kaplan-Meier survival analysis was performed on the top 1 / 4 and bottom 1 / 4 of the co-expression stability rankings, and the log-rank test P value was obtained. The survival time of the two groups of patients was then evaluated based on the P value to determine whether there was a statistical difference. When P ≤ 0.05, there was a statistical difference, indicating that the co-expression stability of the gene significantly affected the patient's survival time; P > 0.05 indicated that the co-expression stability of the gene did not affect the patient's survival time.
[0048] like Figure 3 As shown, the upper path represents the traditional survival analysis algorithm, which ranks patients based on the expression level of gene g0, and compares the survival differences between the first and last 1 / 4 of patients based on the KM survival curve; the lower path represents the new survival analysis algorithm, which estimates the co-expression stability of g0 based on its adjacent genes in the gene regulatory network, ranks patients based on the co-expression stability of g0, and compares the survival differences between the first and last 1 / 4 of patients based on the KM survival curve.
[0049] Step 7) Construct a cancer survival network. Genes with a log-rank test P ≤ 0.05 and the edges connecting these genes in the GRN are retained to form the target cancer survival network. This network, where the co-expression levels of connected nodes vary across patients, can significantly affect patient survival and is therefore of great research value.
[0050] In summary, addressing the shortcomings of traditional survival analysis, this invention endows GRN nodes with a new attribute—coexpression stability—and establishes a correlation between coexpression stability and patient survival. It is worth noting that the mechanism of action of the survival genes discovered by this new method differs from that of traditional methods: traditional survival genes affect patient survival through their own expression levels, while our survival genes affect patient survival through perturbations in the GRN. Therefore, the significance of this new method lies not in replacing traditional survival analysis methods, but in expanding the discovery of cancer survival genes from a new dimension, forming a benign complement to traditional methods.
[0051] With the advancement of precision medicine research in cancer, the discovery of cancer biomarkers has reached a plateau. There is an urgent need to investigate the relationship between genes, cancer progression, and patient survival from various perspectives. Therefore, GRN-based survival analysis strategies hold broad application prospects.
[0052] While specific embodiments of the present invention have been disclosed for illustrative purposes, intended to facilitate understanding and implementation of the present invention, those skilled in the art will appreciate that various substitutions, variations, and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the disclosure of the preferred embodiments, and the scope of protection claimed in the present invention shall be determined by the scope of the claims.
Claims
1. A method for constructing a patient survival network based on a gene regulatory network, comprising the following steps: 1) obtaining a gene expression matrix, wherein the rows of the gene expression matrix represent genes, the columns of the gene expression matrix represent target cancer patient samples, and the element value in the mth row and nth column of the gene expression matrix represents the expression level of the mth gene in the nth target cancer patient; obtaining patient survival information corresponding to each target cancer patient sample; 2) constructing a gene regulatory network based on the gene expression matrix; 3) For each edge in the gene regulatory network, if the credibility of the edge is lower than a set threshold, the edge is deleted; 4) Evaluating the co-expression stability of each gene in the gene regulatory network optimized in step 3) in each target cancer patient sample, the method being: first obtaining the neighboring genes of each gene in the gene regulatory network; then obtaining the expression level of each of the neighboring genes in all patients in the gene expression matrix and performing Z-Score normalization on the expression level; then evaluating the co-expression stability of the gene in each target cancer patient sample based on the Z-Score-normalized neighboring genes of each gene; wherein, for each gene g0 in the gene expression matrix, obtaining M neighboring genes {g1,…,g M }; Gene g0 and its adjacent genes are standardized, and the expression of gene i after Z-Score standardization is expressed as Z i ={v i,1 ,…,v i,n }, where i∈[0,M], M represents the number of neighboring genes of gene g0, and n represents the number of target cancer patient samples; when g0 and g i The expression pattern of is positively correlated, then the co-expression stability of gene g0 in target cancer patient sample j is otherwise v i,j Indicates gene g i The expression level in target cancer patient sample j, j∈{1,…,n}; 5) For each gene in the gene expression matrix, sort the target cancer patient samples based on the co-expression stability of the gene in each target cancer patient sample, taking the survival information of the target cancer patient samples in the top T% of co-expression stability ranking as a first set of information, and taking the survival information of the target cancer patient samples in the bottom T% of co-expression stability ranking as a second set of information; performing a survival analysis based on the first set of information and the second set of information to obtain a log-rank test value P for the gene; then, based on the log-rank test value P for the gene, determining whether the gene has a statistically significant effect on the survival time of each target cancer patient in the top T% of target cancer patient samples and the bottom T% of target cancer patient samples; if there is a statistically significant difference, retaining the gene; 6) Constructing a survival network of the target cancer based on the genes retained in step 5) and the edges and genes connecting the retained genes in the gene regulatory network.
2. The method according to claim 1, characterized in that If the log-rank test value of the gene is P≤0.05, it is determined that the gene has a statistically significant effect on the survival time of each target cancer patient in the top T% of the target cancer patient samples and the bottom T% of the target cancer patient samples.
3. The method according to claim 1, characterized in that The gene regulatory network is optimized using experimental means or an interaction database, and edges in the gene regulatory network whose credibility is lower than a set threshold are deleted.
4. The method according to claim 3, characterized in that The experimental methods include: predicting transcription factor-target gene interactions based on immunoprecipitation, yeast two-hybrid, close-range fluorescence resonance, surface plasmon resonance and mass spectrometry; the interaction databases include: chromatin interaction database, transcription factor-target gene database, small RNA-target gene database, protein-protein interaction database and pathway database.
5. A server, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing each step of the method according to any one of claims 1 to 4.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Lung cancer prognosis comprehensive prediction model, construction method and device
CN112635063A
Method and device for generating lifetime prediction model, and storage medium
CN113782092A