Network propagation-based personalized cancer driver gene identification method
By integrating multiple genomic data to construct a personalized cancer gene regulatory network and calculating node influence scores, the problem of low accuracy in identifying personalized cancer driver genes in existing technologies is solved, and the accuracy of identifying miRNA driver genes is improved.
Patent Information
- Application Number
- CN202310671340.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-07
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2043-06-07
AI Technical Summary
Current technologies have low accuracy in identifying personalized cancer driver genes, especially with insufficient research on the driving role of non-coding RNAs. Furthermore, existing methods are susceptible to gene expression instability patterns when constructing personalized networks.
By integrating gene expression data, copy number variation data, and somatic mutation data, a personalized cancer gene regulatory network is constructed. The influence score of nodes is calculated using network propagation methods to identify personalized cancer driver genes.
It improved the accuracy of identifying personalized cancer driver genes, especially miRNA driver genes, reduced the bias of the network model, and achieved a more stable gene association pattern.
Smart Images

Figure CN116721702B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of biological information, and relates to a personalized cancer driver gene identification method, in particular to a personalized cancer driver gene identification method based on network propagation, which can provide help for targeted treatment of cancer. BACKGROUND
[0002] The occurrence of cancer is related to mutations of many genes, which are called pathogenic genes. The pathogenic genes are divided into driver genes playing a leading role and passenger genes playing a minor role. When the driver gene mutates, it will cause the activation of abnormal proteins, and this abnormality is transmitted to other genes through the interaction between proteins, which may cause the uncontrolled proliferation of cells and further lead to canceration. The passenger gene may also mutate, but it has little to do with the canceration of cells. Therefore, how to accurately identify the cancer driver gene in the case of hundreds of gene mutations has become an urgent problem to be solved, which is of great significance for the effective treatment of cancer.
[0003] It is generally believed that only the mutated gene can be a driver gene, but recent studies have shown that some non-mutated genes that can regulate the mutation of driver genes are also considered to be driver genes. In addition, the driver gene of cancer can also be non-coding RNA, because non-coding regions account for about 98% of the human genome, and non-coding RNA has been proven to be related to the occurrence of cancer.
[0004] Since the causes of each cancer patient are different, even if they suffer from the same cancer, they may be driven by different genomes. Therefore, it is necessary to develop a cancer driver gene research method for individual patients, that is, a personalized cancer driver gene research method. At present, a few researchers have also developed a personalized cancer driver gene identification method, but the effectiveness of the identified driver genes needs to be further verified. In addition, since non-coding RNA also participates in the development of cancer, but the research on the driving effect of non-coding RNA is still insufficient. The general idea of identifying cancer driver genes is to construct a gene regulatory network by using gene multi-omics data, and to extract the topological features of the gene nodes in the network and the biological features of the genes themselves to evaluate the importance of the genes. According to the importance evaluation results of the genes, the genes with high importance are selected as potential cancer driver genes. In order to identify personalized cancer driver genes, researchers need to construct a specific gene regulatory network for each patient, and then identify personalized cancer driver genes in the specific gene regulatory network.
[0005] For example, Pham et al. published an article entitled "pDriver: a novel method for unravelling personalized coding and miRNA cancer drivers" in Bioinformatics in 2021, which proposed a personalized cancer driver gene identification method. First, the method prepared the expression data of the coding genes and miRNAs of the patient, then constructed a personalized miRNA-TF-mRNA network for each cancer patient according to the LIONESS method, and then improved the information of the edges in the personalized miRNA-TF-mRNA network through various gene interaction databases; Finally, according to the network control method, find the minimum node set MDNs in the network, and take the genes corresponding to the nodes in MDNs as the driver genes. This method uses the idea of network control to simulate the functional regulation mechanism of driver genes in the whole regulatory network, and realizes the identification of coding driver genes and miRNA driver genes at the same time. However, this method only uses the expression data of coding genes and miRNAs as input data, and cannot comprehensively integrate biological information. In addition, the network construction process of this method is to reconstruct the aggregate network of all samples through gene expression first, and then estimate the personalized network of a specific sample through linear interpolation, which is easy to be affected by the "unstable" mode of gene expression, thereby affecting the accuracy of the constructed personalized network and further affecting the accuracy of the identification. SUMMARY
[0006] The purpose of the present application is to overcome the defects of the prior art, and to provide a personalized cancer driver gene identification method based on network propagation, which solves the technical problem of low accuracy of personalized driver gene identification in the prior art.
[0007] To achieve the above purpose, the technical solution adopted by the present application includes the following steps:
[0008] (1) Obtain the relevant data for cancer driver gene identification:
[0009] Obtain G genomics data and L gene interaction data D including C cancer types GGI Each genomics data includes N genes composed of P mRNA / TF type genes and Q miRNA type genes, and S samples composed of T diseased samples and O normal samples, each gene includes gene expression data with a size of 1xS Copy number variation data And somatic mutation data And pre-process each genomics data to obtain abnormal genomics data, including abnormal gene expression data with a size of Nx1 for all samples of C cancer types abnormal gene expression data of each corresponding diseased sample abnormal copy number variation data abnormal copy number variation data of each corresponding diseased sample abnormal somatic mutation data abnormal somatic mutation data of each corresponding diseased sample wherein C≥1, G≥3, L≥1, N≥2, T≥1, O≥1, S≥2;
[0010] (2) data fusion on abnormal genomics data:
[0011] abnormal gene expression data of all samples in abnormal genomics data abnormal copy number variation data and abnormal somatic mutation data data fusion on the weight W of N genes of all samples N , simultaneously on abnormal gene expression data of each corresponding diseased sample abnormal copy number variation data of each corresponding diseased sample and abnormal somatic mutation data of each corresponding diseased sample data fusion on the weight W of N genes of each diseased sample N ′;
[0012] (3) construction of population cancer gene regulatory network:
[0013] construction of population cancer gene regulatory network containing N nodes and E edges, with the weight W n of genes as nodes, and the weight W ij of gene interaction data between each two nodes as edges;
[0014] (4) construction of individualized cancer gene regulatory network:
[0015] construction of specificity matrix PSM with element value PSM ij of N×N dimensions according to the scatter plot of each pair of genes, and calculation of the adjacency matrix A′ of individualized cancer gene regulatory network of each diseased sample through PSM and the adjacency matrix A of population cancer gene regulatory network, and then construction of individualized cancer gene regulatory network with V genes as nodes and H gene pairs as edges, wherein V≤N, H≤E, and the weight W v ′ of interaction between nodes, and the weight W lk of interaction between edges;
[0016] (5) Calculate the influence score η of the node pair ji :
[0017] According to the weight W of the edge in the personalized cancer gene regulatory network lk Calculate the network propagation probability P of the lth node to the kth node lk , and based on P lk Calculate the influence score η of the kth node transmitted to the lth node kl :
[0018]
[0019]
[0020] Wherein, indicates the in-degree of the lth node in the personalized cancer gene regulatory network; φ l indicates the node set of all directly connected edges ending with the lth node in the personalized cancer gene regulatory network; Σ indicates the summation symbol; indicates that the sum of the weights of all directly connected edges ending with the lth node in the personalized cancer gene regulatory network is not equal to 0; α indicates the damping factor; Π indicates the product symbol; r l <r k indicates that only all nodes after the lth node are calculated;
[0021] (6) Obtain the personalized cancer driver gene identification result:
[0022] (6a) Initialize the iteration number t, the maximum iteration number T, the weight of the lth node in the personalized cancer gene regulatory network U l , the ranking vector obtained after sorting all node weights from large to small r0, the initial ranking vector r, and let t=0, U l =W l ′, r=r0;
[0023] (6b) According to the influence score η of the node pair in the personalized cancer gene regulatory network kl , the weight U l of the node in the network and the ranking vector r, update the weight U l , and arrange all updated node weights in descending order as the ranking vector r′, wherein the update result of U l is U l ′;
[0024] (6c) Determine whether the ranking vector r is the same as the ranking vector r′, or t=T is true, if so, select the genes corresponding to the top K gene weights in r′ as the personalized cancer driver gene identification result, otherwise let Ul = U l r' = r, t = t + 1, and step (6b) is performed.
[0025] Compared with the prior art, the present application has the following advantages:
[0026] (1) In the construction of personalized cancer gene regulatory network, the present application performs statistical analysis on the interaction between each pair of genes for each patient, which can convert the "unstable" gene expression pattern of the prior art into a "stable" gene correlation pattern, avoiding the deviation of the network model caused by the abnormal expression of a single gene in the prior art, and further accurately constructing a personalized gene regulatory network from the sample level, thereby effectively improving the accuracy of identifying personalized driver genes.
[0027] (2) In the data fusion of abnormal genomics data, the present application integrates various genomics data, including gene expression data, copy number variation data and somatic mutation data, and adds more network topology information in the process of network propagation, compared with the prior art which only uses single gene expression data to affect the model accuracy, which improves the abundance of information in the network model, can realize the description of the role of cancer driver genes in the network from multiple angles, and further improves the accuracy of identifying personalized driver genes. BRIEF DESCRIPTION OF DRAWINGS
[0028] Figure 1 The flowchart of the present application is shown in the figure. DETAILED DESCRIPTION
[0029] The present application will be further described in detail below in combination with the drawings and specific embodiments. It should be noted that the present application is not a method for diagnosing and treating diseases.
[0030] Referring to Figure 1 , the present application comprises the following steps:
[0031] Step 1) Obtain relevant data for cancer driver gene identification:
[0032] 1a) Obtain G genomics data and L gene interaction data D including C cancer types GGI Each genomics data includes N genes composed of P mRNA / TF type genes and Q miRNA type genes, and S samples composed of T diseased samples and O normal samples, each gene includes gene expression data with size 1xS copy number variation data and somatic mutation data Wherein, C≥1, G≥3, L≥1, N≥2, T≥1, O≥1, S≥2; in this embodiment, C=1, G=3, L=798168, N=5138, T=1085, O=104, S=1189;
[0033] 1b) Preprocessing of genomic data is performed using the following method:
[0034] I. Calculate gene expression data for each gene. The average expression value of O normal samples And calculate the expression value E of each gene. g and Differential expression values Then calculate the mean of the differential expression values of this gene across all samples. Then, the mean of the differential expression values of all genes is combined to form an abnormal gene expression data of all samples of size N×1. Differential expression values of all genes for each diseased sample were merged into an N×1 dataset of abnormal gene expression for each diseased sample.
[0035] II. Calculate the copy number variation data for each gene. The proportion of the number of samples in which this gene was amplified or deleted out of the total sample size (f) CNV Then calculate f for all samples of this gene. CNV The mean of all genes, then the f of all genes CNV The mean values are merged into N×1 samples of anomalous copy number variation data. Calculate copy number variation data for each gene. The proportion of the number of samples with gene amplification or deletion in the total number of diseased samples (f) CNV Then, for each diseased sample, f represents all genes. CNV The abnormal copy number variation data are merged into N×1 size for each diseased sample.
[0036] III. The allele mutation frequency “dna_vaf” provided by the Xena platform was used as the somatic mutation frequency f of each gene in each sample. SM Then calculate f for all samples of this gene. SM The mean of all genes, then the f of all genes SM The mean values are combined into N×1 data points representing abnormal somatic mutations across all samples. Somatic mutation frequency f of all genes in each diseased sample SM Abnormal somatic mutation data merged into N×1 size for each diseased sample
[0037] IV. Combine the results from I, II, and III into preprocessed anomalous genomic data.
[0038] Step 2) Perform data fusion for each anomalous genomic data point:
[0039] Abnormal gene expression data from all samples Abnormal copy number mutation data Abnormal somatic mutation data As a feature of the mRNA / TF type genes in all samples, principal component analysis (PCA) was used to estimate the weights w1, w2, and w3 of the three data features for all samples. Then, the weights W of the mRNA / TF type genes in all samples were calculated. P The weights W of miRNA type genes in all samples Q and to W P With W Q By splicing the data, we obtain the weights W of the N genes in all samples. N :
[0040]
[0041]
[0042] W N ={W P W Q}
[0043] Since only expression data is available for miRNA-type genes, the weight of miRNA-type genes is calculated using only the weight w1 of the expression data.
[0044] Abnormal gene expression data for each diseased sample Abnormal copy number mutation data Abnormal somatic mutation data As a feature of the mRNA / TF type gene in each diseased sample, principal component analysis (PCA) was used to estimate the weights w1′, w2′, and w3′ of the three data features for each diseased sample. Then, the weight W of the mRNA / TF type gene in each diseased sample was calculated. P The weights W of the miRNA type genes in each diseased sample. Q ′, and to W P ′ and W Q The data is spliced together to obtain the weights W of the N genes in each diseased sample. N ′:
[0045]
[0046]
[0047] W N ′={W P ′,W Q ′}。
[0048] Step 3) constructing the cancer gene regulatory network of the population:
[0049] The cancer gene regulatory network of the population containing N nodes and E edges is constructed, taking the genes with weight W n as nodes and the gene interaction data with weight W ij as edges, and in this embodiment, E = 40222;
[0050] In this embodiment, the gene interaction data D GGI of L edges is filtered using the gene set containing N genes, ensuring that both genes interacting with each other exist in the gene set containing N genes, and finally E edges of gene interaction are obtained; the weight of the edge is set according to the confidence of each interaction edge in the RegNetwork database, and the confidence of each interaction edge is divided into three categories: “high”, “medium” and “low”, which correspond to weights of 0.8, 0.5 and 0.2 in this embodiment.
[0051] Step 4) constructing the personalized cancer gene regulatory network:
[0052] 4a) constructing a specificity matrix PSM with element value PSM ij and dimension N x N, the specific process is as follows:
[0053] 4a1) initializing the specificity matrix PSM of each patient, with size N x N, and the value of each element is set to 0;
[0054] 4a2) merging the weight W N ′ of the N genes of each diseased sample, and the weight of all genes of the T diseased samples to form a gene weight matrix with size N x T, and then constructing a scatter plot taking each patient as a point and the weight of each gene pair as the coordinate value of the x-axis and y-axis of the point.
[0055] 4a3) calculating the element value PSM ij of the specificity matrix PSM for each patient according to the scatter plot;
[0056] Taking patient s as an example, the number of points n j i (s) i j (s) and the number of points n in the intersection area a x b of the former two xy (s) , and the statistic is set as p xy (s) ; wherein, max{W j ′} represents the maximum value of the weight of the jth gene in all diseased samples, max{W i ′} represents the maximum value of the weight of the ith gene in all diseased samples, a, b are indefinite values, which depend on the number of points; if the statistic p ij (s) is greater than the significance level, the value in the ith row and the jth column in the PSM matrix of the patient s is set to 1, indicating that there is an edge between the ith node and the jth node in the personalized network of the patient s, otherwise the value in the ith row and the jth column in the PSM matrix of the patient s is set to 0, indicating that there is no edge between the ith node and the jth node in the personalized network of the patient s, in the embodiment, n i (s) = 0.1N, n j (s) = 0.1N;
[0057] Calculate the element value PSM ij (s) , and the calculation formula is:
[0058]
[0059]
[0060] 4b) Calculate the adjacency matrix A' of the personalized cancer gene regulatory network of each diseased sample by the PSM and the adjacency matrix A of the population cancer gene regulatory network, and the calculation formula of the adjacency matrix A' is:
[0061] A' = A·PSM
[0062] Wherein, · represents dot product calculation;
[0063] 4c) Construct a personalized cancer gene regulatory network with V genes as nodes and H gene pairs as edges, wherein the weight of the interaction between the genes is W v ′. lk
[0064] V genes with the weight of the interaction W v ′ are nodes, wherein the interaction refers to the genes with the weight of the interaction not being 0 in the adjacency matrix A', the personalized cancer gene regulatory network belongs to the subnetwork of the population abnormal gene regulatory network, V≤N, H≤E.
[0065] Step 5) Calculate the influence score η of the node pairkl :
[0066] Based on the edge weights W in a personalized cancer gene regulatory network lk Calculate the network propagation probability P from node l to node k. lk And based on P lk Calculate the influence score η passed from the k-th node to the l-th node. kl :
[0067]
[0068]
[0069] Among them, DC l - represents the in-degree of the l-th node in a personalized cancer gene regulatory network; φ l This represents the set of all directly connected nodes in a personalized cancer gene regulation network, with the l-th node as the endpoint; Σ represents the summation symbol. In a personalized cancer gene regulatory network, the sum of the weights of all directly connected edges ending at the l-th node is not equal to 0; α represents the damping factor; Π represents the quadrature symbol; r l <r k This means that only all nodes after the l-th node are calculated; in this embodiment, α = 0.8.
[0070] Step 6) Obtain personalized cancer driver gene identification results:
[0071] 6a) The initial number of iterations is t, the maximum number of iterations is T, and the weight of the l-th node in the personalized cancer gene regulation network is U. l The ranking vector obtained after sorting all node weights from largest to smallest is r0. The initial ranking vector is r, and t = 0, U l =W l ′, r=r0, in this embodiment T=100;
[0072] 6b) Based on the influence score η of node pairs in the personalized cancer gene regulatory network kl The weight U of nodes in the network l And the ranking vector r with weight U l The updated weights of all nodes are then arranged in descending order to form a ranking vector r′, where U l The updated result is U l The updated formula is:
[0073]
[0074] 6c) judge whether the ranking vector r is same as the ranking vector r' or t=T is established, if yes, select the genes corresponding to the first K gene weights with the largest values in r' as the personalized cancer driver gene recognition result, otherwise let U l = U l ', r=r', t=t+1, and execute step (6b).
[0075] The technical effects of the present application are further described below in combination with simulation results:
[0076] 1. Experimental conditions and contents:
[0077] The hardware platform of the simulation experiment is: the CPU is 11th Gen Intel(R) Core(TM) i7-11700, the main frequency is 2.50 GHz, the memory is 16G, the software platform is: the operating system is Windows 10, and the python version is 3.8.10.
[0078] The data set used in the simulation experiment is collected from the TCGA database, and breast invasive carcinoma BRCA is taken as an example. The results of 1085 diseased samples of personalized driver genes are obtained, wherein the personalized driver gene is defined as a gene with no more than 0.5% of the number of diseased samples in the first 100 candidate driver genes of all diseased samples.
[0079] Experiment one, the simulation of the personalized driver gene result of the present application, the results are shown in table 1;
[0080]
[0081] Table 1
[0082] Method Encoding driver gene accuracy miRNA driver gene accuracy Prior art 27.6% 77.8% The invention 31.7% 88.5%
[0083] 2. Analysis of experimental results:
[0084] Referring to table 1, the accuracy of the coding driver gene of the method of the present application is 31.7%, and the accuracy of the miRNA driver gene is 88.5%, which are higher than those of the prior art method, proving that the method of the present application improves the recognition accuracy of the personalized cancer driver gene.
Claims
1. A personalized cancer driver gene identification method based on network propagation, characterized in that, Includes the following steps: (1) Obtain relevant data for cancer driver gene identification: Acquire G genomic data points representing C cancer types and L gene interaction data points. GGI Each genomics dataset includes N genes consisting of P mRNA / TF type genes and Q miRNA type genes, and S samples consisting of T diseased samples and O normal samples. Each gene includes gene expression data of size 1×S. Copy number variation data and somatic mutation data Each genomic data point was preprocessed to obtain aberrant genomic data, including aberrant gene expression data of size N×1 for all samples across C cancer types. and the abnormal gene expression data of each corresponding diseased sample. Abnormal copy number mutation data and the corresponding abnormal copy number variation data for each diseased sample. Abnormal somatic mutation data and the corresponding abnormal somatic mutation data for each diseased sample. Wherein, C≥1, G≥3, L≥1, N≥2, T≥1, O≥1, S≥2; (2) Data fusion of anomalous genomic data: Abnormal gene expression data of all samples in the abnormal genomics data Abnormal copy number mutation data and abnormal somatic mutation data Data fusion is performed to obtain the weights W of the N genes in all samples. N At the same time, Abnormal gene expression data for each corresponding diseased sample Abnormal copy number variation data for each diseased sample as well as Abnormal somatic mutation data for each corresponding diseased sample Data fusion was performed to obtain the weights W of N genes for each diseased sample. N ′; (3) Constructing a population-based cancer gene regulatory network: Construct a system with weight W n The genes are nodes, and the interaction weight between any two nodes is W. ij The gene interaction data is a cancer gene regulatory network consisting of a population of N nodes and E edges; (4) Constructing personalized cancer gene regulatory networks: Based on the scatter plot of each gene pair, construct an element-valued PSM. ij A specificity matrix PSM with dimensions N×N is used. The adjacency matrix A′ of the individualized cancer gene regulatory network for each diseased sample is calculated using the PSM and the adjacency matrix A of the population cancer gene regulatory network. Then, a system is constructed with weights W for interactions existing in adjacency matrix A′. v Let V genes be nodes, with W as the weight for interactions. lk A personalized cancer gene regulatory network with H gene pairs as edges, where V≤N, H≤E; (5) Calculate the influence score η of the node pair kl : Based on the edge weights W in a personalized cancer gene regulatory network lk Calculate the network propagation probability P from node l to node k. lk And based on P lk Calculate the influence score η passed from the k-th node to the l-th node. kl : in, φ represents the in-degree of the l-th node in a personalized cancer gene regulatory network; l This represents the set of all directly connected nodes in a personalized cancer gene regulation network, with the l-th node as the endpoint; Σ represents the summation symbol. In a personalized cancer gene regulatory network, the sum of the weights of all directly connected edges ending at the l-th node is not equal to 0; α represents the damping factor; Π represents the quadrature symbol; r l <r k This means that only all nodes after the l-th node are counted; (6) Obtain personalized cancer driver gene identification results: (6a) The initial number of iterations is t, the maximum number of iterations is T, and the weight of the l-th node in the personalized cancer gene regulation network is U. l The ranking vector obtained after sorting all node weights from largest to smallest is r0. The initial ranking vector is r, and t = 0, U l =W l ′,r=r0; (6b) Based on the influence score η of node pairs in personalized cancer gene regulatory networks kl The weight U of nodes in the network l And the ranking vector r with weight U l The updated weights of all nodes are then arranged in descending order to form a ranking vector r′, where U l The updated result is U l ′; (6c) Determine whether the ranking vector r is the same as the ranking vector r′, or whether t = T holds. If so, select the genes corresponding to the top K gene weights with the largest values in r′ as the personalized cancer driver gene identification results; otherwise, let U l =U l ', r = r', t = t+1, and execute step (6b).
2. The method according to claim 1, characterized in that, The preprocessing of each genomic data point described in step (1) is implemented as follows: I. Calculate gene expression data for each gene. The average expression value of O normal samples And calculate the expression value E of each gene. g and Differential expression values Then calculate the mean of the differential expression values of this gene across all samples. Then, the mean of the differential expression values of all genes is combined to form an abnormal gene expression data of all samples of size N×1. Differential expression values of all genes for each diseased sample were merged into an N×1 dataset of abnormal gene expression for each diseased sample. II. Calculate the copy number variation data for each gene. The proportion of the number of samples in which this gene was amplified or deleted out of the total sample size (f) CNV Then calculate f for all samples of this gene. CNV The mean of all genes, then the f of all genes CNV The mean values are merged into N×1 samples of anomalous copy number variation data. Calculate copy number variation data for each gene. The proportion of the number of samples with gene amplification or deletion in the total number of diseased samples (f) CNV Then, for each diseased sample, f represents all genes. CNV The abnormal copy number variation data are merged into N×1 size for each diseased sample. III. The allele mutation frequency "dna_vaf" provided by the Xena platform was used as the somatic mutation frequency f for each gene in each sample. SM Then calculate f for all samples of this gene. SM The mean of all genes, then the f of all genes SM The mean values are combined into N×1 data points representing abnormal somatic mutations across all samples. Somatic mutation frequency f of all genes in each diseased sample SM Abnormal somatic mutation data merged into N×1 size for each diseased sample IV. Combine the results from I, II, and III into anomalous genomic data, i.e., preprocessed genomic data.
3. The personalized cancer driver gene identification method based on network propagation according to claim 1, characterized in that, The weights W of the N genes in all samples mentioned in step (2) N The method for obtaining it is as follows: Abnormal gene expression data from all samples Abnormal copy number mutation data Abnormal somatic mutation data As a feature of the mRNA / TF type genes in all samples, principal component analysis (PCA) was used to estimate the weights w1, w2, and w3 of the three data features for all samples. Then, the weights W of the mRNA / TF type genes in all samples were calculated. P The weights W of miRNA type genes in all samples Q and to W P With W Q By splicing the data, we obtain the weights W of the N genes in all samples. N : IN N ={W P ,IN Q }。 4. The personalized cancer driver gene identification method based on network propagation according to claim 1, characterized in that, The weights W of the N genes in each diseased sample as described in step (2) N ′, its acquisition method is as follows: Abnormal gene expression data for each diseased sample Abnormal copy number mutation data Abnormal somatic mutation data As a feature of the mRNA / TF type gene in each diseased sample, principal component analysis (PCA) was used to estimate the weights w1′, w2′, and w3′ of the three data features for each diseased sample. Then, the weight W of the mRNA / TF type gene in each diseased sample was calculated. P The weights W of the miRNA type genes in each diseased sample. Q ′, and to W P ′ and W Q The data is spliced together to obtain the weights W of the N genes in each diseased sample. N ′: IN N ′={W P ',IN Q ′}。 5. The personalized cancer driver gene identification method based on network propagation according to claim 1, characterized in that, The scatter plot of each pair of genes mentioned in step (4) is constructed by merging the weights W of the N genes in each diseased sample. N The weights of all genes in the T disease samples are combined to form a gene weight matrix of size N×T. Then, a scatter plot is constructed with each patient as a point and the weights of each gene pair as the x-axis and y-axis coordinates of the point.
6. The personalized cancer driver gene identification method based on network propagation according to claim 5, characterized in that, The step (4) described above involves constructing an element-valued PSM based on the scatter plot of each gene pair. ij A specificity matrix PSM of dimension N×N, where the element values are PSM. ij The calculation formula is: Where, ρ ij This represents the statistics of the node pair consisting of the i-th node and the j-th node; n i In a scatter plot, a×max{W j The number of midpoints in the area; n j In a scatter plot, max{W i The number of midpoints in area ′}×b; n ij This represents the number of points within the area a×b of the intersection of the first two, max{W j '} represents the maximum weight of the j-th gene among all diseased samples, max{W i ′} represents the maximum weight of the i-th gene among all diseased samples.
7. The personalized cancer driver gene identification method based on network propagation according to claim 1, characterized in that, The adjacency matrix A′ mentioned in step (4) is calculated using the following formula: A′=A·PSM Where · represents dot product calculation.
8. The personalized cancer driver gene identification method based on network propagation according to claim 1, characterized in that, The weights of the interactions mentioned in step (4) are W. v The V genes in the adjacency matrix A′ are nodes, where the interaction refers to the interaction between the adjacency matrix A′ and other genes whose weights are not 0.
9. The personalized cancer driver gene identification method based on network propagation according to claim 1, characterized in that, The weighting of U described in step (6b) l The update is performed using the following formula:
Citation Information
Patent Citations
Traffic resource allocation method and system based on confidence map convolutional network, and medium
CN114781894A
Cancer driver gene identification method and system
CN115762631A