A cpga-smcmn method for identifying cancer single driver pathways
By combining biological networks and gene clusters using the CPGA-SMCMN method, and utilizing the SMCMN model and Partheno-Genetic algorithm, the problem of insufficient accuracy and coverage in the identification of single-drive pathways in cancer in existing technologies is solved, and more efficient identification of single-drive pathways in cancer is achieved.
Patent Information
- Application Number
- CN202210987079.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-17
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2042-08-17
AI Technical Summary
Existing technologies struggle to effectively utilize biological network information when identifying single-drive pathways in cancer, resulting in insufficient accuracy and coverage, particularly for identifying low-frequency mutated genes.
Using the CPGA-SMCMN method, which combines biological networks and gene clusters, we identify gene sets with maximum coverage, mutual exclusion, and network connectivity by designing novel distance measurements and the Partheno-Genetic algorithm, and then optimize them using the SMCMN model.
It improves the accuracy and coverage of identifying cancer-specific driver pathways, enabling the identification of more genes enriched in important driver pathways, and has strong scalability and practicality.
Smart Images

Figure CN115359839B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics and is used to identify single-drive pathways in cancer, specifically a CPGA-SMCMN method for identifying single-drive pathways in cancer. Background Technology
[0002] Cancer is a disease with a high mortality rate, generally caused by mutations in driver genes. Unlike passenger genes, whose mutations are not related to cancer, mutations in driver genes promote the unlimited proliferation and spread of cancer cells. Previous studies have shown that the difficulty in diagnosing and treating cancer is due to the inherent large amount of mutational heterogeneity in the cancer genome. That is, there are many key regulatory pathways or cell signaling pathways in the human body responsible for regulating cell proliferation, metabolism, and apoptosis, and each of them has a set of driver genes. Mutations in any one of these driver genes are enough to disrupt the regulatory function of the pathway and lead to cancer. Therefore, identifying a set of driver genes enriched in a pathway, i.e., driver pathways, is crucial for studying the pathogenesis of cancer. However, conducting research through laboratory biological experiments is both time-consuming and expensive. Fortunately, with the rapid development of sequencing technology, the accumulated rich multi-omics data has made it possible to identify driver pathways (driver gene sets) using computational methods, becoming one of the most widely concerned issues in the field of bioinformatics.
[0003] There are generally two approaches to identify cancer-driving pathways: de novo approaches and prior knowledge-based approaches. De novo approaches attempt to discover a set of genes that have two basic characteristics of driving pathways by using only genetic data: high coverage and high exclusivity. High coverage indicates that genes in a driving pathway are mutated in a large number of cancer samples, while high exclusivity indicates that two genes in a pathway are rarely mutated simultaneously in the same cancer sample. In 2012, Vandin et al. first proposed the maximum weighted submatrix model, attempting to simultaneously minimize coverage and exclusivity, and designed a Markov chain Monte Carlo-based method, Dentrix (De novo Driver Exclusivity), to solve this model. In the same year, Zhao et al. proposed a binary linear programming method and a GA (genetic algorithm) method to solve the model. Compared to the Dentrix method, these two methods have competing performance, and the GA-Z method is particularly suitable for solving comprehensive models containing gene expression profiles. In 2013, Zhang et al. integrated two weighted networks constructed from mutation matrices and expressions, and proposed a network-based method, iMCMC (identify Mutated Core Modules in...). Cancer extracts core modules from integrated networks; however, due to the uncontrollable nature of module size, this method cannot obtain modules of the expected size. In 2016, based on the GA method, a multi-objective optimization method, MOGA, was designed to balance the trade-off between high coverage and high exclusivity. In 2017, Yahya et al. proposed the QuaDMutEx method, which uses Monte Carlo optimization and binary quadratic programming to find gene sets. In 2019, Wu et al. redefined the maximum weight submatrix model, adjusting coverage and exclusivity according to the average weight of genes in a pathway; they proposed a parthenogenetic algorithm, PGA-MWS, to solve this model. In 2021, Wu et al. introduced a weighted non-binary mutation matrix composed of four omics data: somatic mutation data, copy number mutation data, gene expression data, and DNA methylation data. By redefining coverage and mutual exclusivity, a new maximum weight submatrix model was established, and a cooperative co-evolutionary algorithm, CGA-MWS, was designed to solve this model. In most cases, a set of genes was detected, containing more genes enriched in a known signaling pathway compared to previous methods.
[0004] Due to the large number of mutated gene combinations, the aforementioned de novo methods typically reduce intrinsic computational complexity by using pre-filtering based on mutation frequency, and may overlook some pathways containing rare mutations. Previous knowledge-based methods treated frequently mutated genes and their less frequently mutated neighbors as driving factors and attempted to detect them from known human genes or protein-level pathways or networks, such as HotNet, IDM-SPS, HotNet2, and MEXCOwalk. However, biological networks remain noisy and incomplete. Combining the intuition of these two methods—utilizing the fundamental features of driving pathways and the relationships between genes in biological networks—has begun to emerge. In 2020, Yahya et al. proposed the QuaDMutNetEx method, which is an extension of their QuaDMutEx method, incorporating gene connectivity into the recognition model. Experimental results show that, compared to the QuaDMutEx method, the QuaDMutNetEx method can identify some rare driving genes with low mutation frequencies, and the combination of fundamental pathway features and prior knowledge is indeed effective. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of existing technologies by providing a CPGA-SMCMN method for identifying single-drive pathways in cancer. This method can identify more biological information, is highly scalable and practical, and can identify more genes enriched in important drive pathways.
[0006] The technical solution to achieve the objective of this invention is:
[0007] A CPGA-SMCMN method for identifying single-drive pathways in cancer includes the following steps:
[0008] 1) Setting up the SMCMN model:
[0009] Suppose there is a somatic mutation matrix A copy number mutation The rows of somatic mutation matrix and copy number variation represent the same group of patient samples. Columns represent the same group of genes Each entry Indicates the first Is the gene in the first...? All mutations occurring in each patient, in the matrix , Indicates the first The gene in the first In the statistically significant variation region of each patient, otherwise ;also, and The two matrices record the correlations between genes, where This indicates the relationship between two genes recorded in literature indexed by the String network. This indicates the correlation between two genes obtained from related gene experiments collected from the String network, that is, the correlation measured by previous researchers through experiments. and The range is from 0 to 999;
[0010] Binary mutation matrix From the matrix and Composition, if and If they are not both equal to 0, then The value is 1 if it is true, otherwise it is 0; a new correlation matrix. Also through combination matrices and Generate, where as follows:
[0011] ,
[0012] Assumption yes any Submatrix, let , Defined as a matrix The proportion of samples covered, i.e., the matrix "Coverage":
[0013] ,
[0014] in Representation matrix The sample set covered;
[0015] Given a pair of genes and ,make represent and Relative Hamming distance That is, Hamming distance gene Relative to genes ,formula :
[0016] ,
[0017] in :
[0018] ,
[0019] For submatrix ,make Indicates gene and gene set The "mutual exclusion" between them:
[0020] ,
[0021] So, "mutual exclusion" To measure The average difference between each gene and the rest of the genes , The larger the matrix The higher the mutual exclusivity,
[0022] ,
[0023] Given a submatrix The genes contained therein allow The functional correlation between two genes in the M matrix is measured as follows:
[0024] ,
[0025] in Representation matrix The elements are from the correlation matrix. The submatrix extracted from it;
[0026] Based on the above definitions, construct the combinatorial model SMCMN and determine the submatrix with maximum coverage, mutual exclusion, and network connectivity: given a binary mutation matrix Correlation matrix and a parameter , mark a submatrix Maximize the weight function :
[0027] ;
[0028] 2) Setting up CPGA:
[0029] Since it is difficult to generate gene sets that satisfy mutual exclusion and correlation using crossover operations, a gene cluster-based Partheno-Genetic Algorithm (CPGA) is designed to solve the SMCMN model. The input is a... Binary mutation matrix ,one Correlation matrix and a parameter The output is submatrix The following describes the key steps of CPGA:
[0030] 2.1) Clustering:
[0031] In the preprocessing stage, two gene clusters are constructed for each gene, given a gene ,make Records and Genes The set of genes with Hamming distance, i.e. ,structure To record and genes Related gene sets, i.e. , and There are two preset parameters;
[0032] 2.2) Chromosome coding and initial population:
[0033] Chromosomes, also known as individuals, represent solutions to problems. In CPGA, a set of chromosomes is used. One gene encodes one chromosome, that is ,generate The initial chromosome is used to generate the initial population. The initial chromosome is created through the following two steps:
[0034] (2.2.1) Selecting genes using a roulette wheel strategy ,Right now The larger the value, the more likely it is to select a gene. ,make ;
[0035] (2.2.2) Iteratively select the remaining ones using a roulette wheel strategy. One gene, assuming ,let ,according to Select the next gene and satisfy , These are preset parameters. In any iteration, the chromosome cannot be created successfully;
[0036] 2.3) Fitness function:
[0037] Fitness, representing the quality of a solution, plays a crucial role in guiding evolution; express Corresponding chromosome of Submatrix, Used to measure chromosomes Adaptability, Defined as follows, a higher fitness level indicates a better solution.
[0038] ,
[0039] 2.4) Genetic Operators:
[0040] In the Partheno-Genetic method, only selection and recombination operators are applied to generate offspring. In CPGA, roulette wheel selection is used to select individuals with high fitness with a higher probability. Elite strategies are also used to maintain individuals with the highest fitness during evolution. A greedy recombination operator is designed to generate new offspring, as shown below:
[0041] (2.4.1) Given a chromosome Randomly execute one of the following two methods from Delete a gene, (1) from Delete the gene that mutates in the fewest patients, i.e. (2) From A gene is randomly deleted from the chromosome, and the new chromosome is used... express;
[0042] (2.4.2) Let In order to satisfy , Selecting under the premise of preset parameters The largest gene If from If no matching gene is found in the chromosome, then... Remain unchanged;
[0043] 2.5) CPGA Specific Process:
[0044] In CPGA, the parameters used in CPGA are set. Steps 1 to 4 in CPGA perform preprocessing, and step 5 in CPGA generates a dataset of size [size missing]. The initial population, from and The entire evolutionary process of control is executed from step 6 to step 19, in which the roulette selection and recombination operators are executed iteratively from step 9 to step 13. Finally, the best individual is transformed and output in steps 20 and 21.
[0045] The specific process is as follows:
[0046] .
[0047] This technical solution studies the identification method by combining the basic characteristics of the driving path with prior knowledge in biological networks. The main contributions are summarized as follows:
[0048] 1) A novel distance GGD (Gene and Gene set Distance) is designed to calculate the distance between genes and gene sets. Therefore, given a gene set, the average DGG value between each gene and the rest of the genes measures the mutual exclusivity of the set.
[0049] 2) By exploring submatrices with maximum coverage, mutual exclusion and network connectivity, a parameter-free identification model SMCMN was developed, which can filter a large number of mutant gene combinations based on their relationships on the biological network.
[0050] 3) A CPGA based on gene clustering and the Partheno-Genetic algorithm is proposed. The novel operator is designed to initialize and mutate individuals based on gene clustering, which greatly reduces the search space and ensures that the genes in the identified sets are highly correlated.
[0051] 4) Real biological data are used to illustrate the effectiveness of the proposed model and algorithm. Experimental results show that CPGA has strong search capabilities. The identification method CPGA-SMCMN, based on the SMCMN model and CPGA, has high identification accuracy and improves the enrichment and connectivity of detected genes. This has been proven through numerical experiments and comparison with other state-of-the-art methods.
[0052] This method can identify more biological information, is highly scalable and practical, and can identify more genes enriched in important driver pathways. Attached Figure Description
[0053] Figure 1 This is an example diagram of the mutation matrix in the embodiment;
[0054] Figure 2 A schematic diagram for implementing matrix fusion;
[0055] Figure 3 The diagram shows the operation of glioblastoma GBM with a driving pathway scale of 3.
[0056] Figure 4 The diagram shows the operation of glioblastoma GBM with a driving pathway scale of 6. Detailed Implementation
[0057] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments, but this is not intended to limit the scope of the invention.
[0058] The example experiment was conducted on a Lenovo PC equipped with an Intel(R) Core(TM) i7-7700 3.60GHz CPU and 24GB of RAM. The operating system was Windows 11, and the compiler was Python 3.0 in PyCharm 2018.1.4.
[0059] This example illustrates the problem of single-drive path identification.
[0060] Example:
[0061] A CPGA-SMCMN method for identifying single-drive pathways in cancer includes the following steps:
[0062] 1) Setting up the SMCMN model:
[0063] Suppose there is a somatic mutation matrix ,like Figure 1 As shown, a copy number mutation The rows of somatic mutation matrix and copy number variation represent the same group of patient samples. Columns represent the same group of genes Each entry Indicates the first Is the gene in the first...? All mutations occurring in each patient, in the matrix , Indicates the first The gene in the first In the statistically significant variation region of a sample, otherwise ;also, and The two matrices record the correlations between genes, where This indicates the relationship between two genes recorded in literature indexed by the String network. This indicates the correlation between two genes obtained from related gene experiments collected from the String network, that is, the correlation measured by previous researchers through experiments. and The range is from 0 to 999, such as Figure 2 As shown;
[0064] Three cancer datasets were used: glioblastoma (GBM), ovarian cancer (OVCA), and thyroid cancer (THCA). Mutation data for glioblastoma and ovarian cancer were obtained from Zhao et al.; mutation data for thyroid cancer and copy number variation data for the three cancers were obtained from TCGA (http: / / tcgadata.nci.nih.gov / tcga / ); GISTIC was used to convert the values of the raw copy number variation data into values in the set {-1, 0, 1}; the confidence values of associations between genes were obtained from literature and experiments, respectively, from the STRING network (https: / / cn.string-db.org).
[0065] In addition, genes with a mutation rate of less than 0.5% in the samples were filtered out. Genes TP53 and TTN were removed from the OVCA dataset because mutations on TP53 are more common than mutations on other genes. TP53 mutations occurred in more than 80% of the samples, while mutations in other genes occurred in less than 25% of the samples. Furthermore, mutations on TTN may be artificial genes. The processed data is shown in Table 1 below, where edges represent the number of edges between corresponding genes in the STRING network.
[0066] Table 1
[0067] ,
[0068] Binary mutation matrix From the matrix and Composition, if and If they are not both equal to 0, then The value is 1 if it is 1 otherwise it is 0, such as Figure 2 As shown in matrix A; a new correlation matrix Also through combination matrices and Generate, such as Figure 2 As shown in the W matrix, where as follows:
[0069] ,
[0070] Assumption yes any Submatrix, let , Defined as a matrix The proportion of samples covered, i.e., the matrix "Coverage":
[0071] ,
[0072] in Representation matrix The sample set covered;
[0073] Given a pair of genes and ,make represent and Relative Hamming distance That is, Hamming distance gene Relative to genes ,formula :
[0074] ,
[0075] in :
[0076] ,
[0077] For submatrix ,make Indicates gene and gene set The "mutual exclusion" between them:
[0078] ,
[0079] So, "mutual exclusion" To measure The average difference between each gene and the rest of the genes , The larger the matrix The higher the mutual exclusivity,
[0080] ,
[0081] Given a submatrix The genes contained therein allow The functional correlation between two genes in the M matrix is measured as follows:
[0082] ,
[0083] in Representation matrix The elements are from the correlation matrix. The submatrix extracted from it;
[0084] Based on the above definitions, construct the combinatorial model SMCMN and determine the submatrix with maximum coverage, mutual exclusion, and network connectivity: given a binary mutation matrix Correlation matrix and a parameter , mark a submatrix Maximize the weight function :
[0085] ;
[0086] 2) Setting up CPGA:
[0087] Since it is difficult to generate gene sets that satisfy mutual exclusion and correlation using crossover operations, a gene cluster-based Partheno-Genetic Algorithm (CPGA) is designed to solve the SMCMN model. The input is a... Binary mutation matrix ,one Correlation matrix and a parameter The output is submatrix The following describes the key steps of CPGA:
[0088] The parameters for CPGA are set as follows: population size N is set to 1 / 4 of the number of genes (the number of genes varies for different cancer types, and the final initial population size also varies); in this example, the maximum number of iterations required per iteration is maxt = 1000, the total number of runs is maxt = 100, the mutation probability rr = 0.3, and two thresholds are set. It equals 0.7. It equals 0.5;
[0089] 2.1) Clustering:
[0090] In the preprocessing stage, two gene clusters are constructed for each gene, given a gene ,make Records and Genes Gene sets with large Hamming distances, i.e. Similarly, construct To record and genes A set of genes with high correlation, i.e. , and There are two preset parameters;
[0091] 2.2) Chromosome coding and initial population:
[0092] Chromosomes, also known as individuals, represent solutions to problems. In CPGA, a set of chromosomes is used. One gene encodes one chromosome, that is ,generate The initial chromosome is used to generate the initial population. The initial chromosome is created through the following two steps:
[0093] (2.2.1) Selecting genes using a roulette wheel strategy ,Right now The larger the value, the more likely it is to select a gene. ,make ;
[0094] (2.2.2) Iteratively select the remaining ones using a roulette wheel strategy. One gene, assuming ,let ,according to Select the next gene and satisfy , These are preset parameters. In any iteration, the chromosome cannot be created successfully;
[0095] 2.3) Fitness function:
[0096] Fitness, representing the quality of a solution, plays a crucial role in guiding evolution; express Corresponding chromosome of Submatrix, Used to measure chromosomes Adaptability, It is defined by the following equation: the greater the fitness, the more feasible the solution.
[0097] ;
[0098] 2.4) Genetic Operators:
[0099] In Partheno-Genetic, only selection and recombination operators are used to generate offspring. In CPGA, roulette wheel selection is used to select individuals with high fitness with a higher probability. Elite strategies are also used to maintain individuals with the highest fitness during evolution. A greedy recombination operator is designed to generate new offspring, as shown below:
[0100] (2.4.1) Given a chromosome Randomly execute one of the following two methods from Delete a gene, (1) from Delete the gene that mutates in the fewest patients, i.e. (2) From A gene is randomly deleted from the chromosome, and the new chromosome is used... express;
[0101] (2.4.2) Let In order to satisfy , Selecting under the premise of preset parameters The largest gene If from If no matching gene is found in the chromosome, then... Remain unchanged;
[0102] 2.5) CPGA Specific Process:
[0103] In CPGA, the parameters used in CPGA are set. Steps 1 to 4 in CPGA perform preprocessing, and step 5 in CPGA generates a file of size [size missing]. The initial population, from and The entire evolutionary process of control is executed from step 6 to step 19, in which the roulette selection and recombination operators are executed iteratively from step 9 to step 13. Finally, the best individual is transformed and output in steps 20 and 21.
[0104] The specific process is as follows:
[0105] .
[0106] like Figure 3 and Figure 4 The maximum fitness value indicates the highest fitness value among all individuals obtained after CPGA is executed. This individual with the highest fitness value is the gene name corresponding to the optimal fitness index. The last time indicates the specific time consumed for one iteration, in seconds.
Claims
1. A CPGA-SMCMN method for identifying a cancer single driver pathway, characterized in that, Comprising the following steps: 1) Setting up the SMCMN model: Suppose there is a somatic mutation matrix , a copy number variation , the rows of the somatic mutation matrix and the copy number variation represent the same set of patient samples , the columns represent the same set of genes , each entry , indicates whether the th gene is mutated in the th patient, in the matrix , , indicates whether the th gene is statistically significantly variant in the th patient, otherwise ; in addition, and two matrices record the correlation between genes, where indicates the relationship between two genes recorded in the String network literature, indicates the correlation between two genes obtained from the relevant gene experiments included in the String network, and range from 0 to 999; binary variation matrix is composed of matrices and if and are not equal to 0, then the value of is 1, otherwise 0; a new correlation matrix is also generated by combining matrices and where as follows: Assume is any submatrix of , defined as the proportion of samples covered by the matrix , i.e. the "coverage" of the matrix . , wherein denotes a matrix the covered sample set; Given a pair of genes and , let represent the relative Hamming distance between and , i.e. the Hamming distance gene with respect to gene , the formula : , wherein : , For a sub-matrix , let denote the "mutual exclusivity" between a gene and a gene set : , So "mutual exclusivity" To measure the average , mutual exclusivity between each gene and the rest of the genes, we computed the matrix mutual exclusivity, , Given sub-matrix The genes contained in the matrix M are selected such that The functional correlation between two genes in the matrix M is measured as follows: , wherein denotes the element of the matrix which is a submatrix extracted from the correlation matrix Based on the above definitions, a combinatorial model SMCMN is constructed to determine a sub-matrix with maximum coverage, mutual exclusivity and network connectivity: given a binary variation matrix , a correlation matrix , and a parameter , identifying a sub-matrix maximizing a weight function : ; 2) Setting up the CPGA: A Partheno-Genetic Algorithm (CPGA) based on gene cluster is designed to solve the SMCMN model, the input is a binary mutation matrix , a correlation matrix and a parameter , and the output is a sub-matrix . The key steps of CPGA are described as follows: 2.1) Clustering: In the preprocessing stage, two gene clusters are constructed for each gene, given a gene , let record the gene set with Hamming distance to the gene , i.e. , construct to record the gene set related to the gene , i.e. , and are two preset parameters; 2.2) Chromosome encoding and initial population: A chromosome, also called an individual, represents a solution to the problem, in CPGA, a chromosome is encoded using a set of genes, i.e. , an initial chromosome is generated to generate an initial population, the initial chromosome is created by the following two steps: , generating (2.2.1) Selecting genes with the roulette strategy i.e. The larger, the more genes are selected with the probability ; (2.2.2) Select the remaining genes with Wheel Strategy Iteration assuming that the next gene is selected, and that , is a preset parameter, if at any iteration, then the chromosome cannot be successfully created; 2.3) Fitness function: Fitness, representing the quality of a solution, plays a crucial role in guiding evolution; let express Corresponding chromosome of Submatrix, Used to measure chromosomes Adaptability, Defined as follows, a higher fitness level indicates a better solution. , 2.4) Genetic operators: In Partheno-Genetic, only selection and recombination operators are applied to generate offspring, in CPGA, roulette wheel selection is used, individuals with higher fitness are selected with higher probability, elitist strategy is also used to keep the individual with the highest fitness during the evolution process, a greedy-based recombination operator is designed to generate new offspring, as follows: (2.4.1) Given a chromosome , one gene is deleted from by randomly performing one of the following two methods, (1) the gene that is mutated in the least number of patients is deleted from , i.e. , (2) one gene is randomly deleted from , the new chromosome is denoted by ; (2.4.2) Let , if the condition is met , , the gene with the largest value is selected , if no gene is found in that meets the condition, the chromosome remains unchanged; 2.5) CPGA specific process: In the CPGA, parameters used in the CPGA are set, the first step to the fourth step in the CPGA realize pre-processing, the fifth step of the CPGA generates an initial population with a size of , the whole evolution process controlled by and is executed from the sixth step to the nineteenth step, in which the roulette selection and the recombination operator are iteratively executed from the ninth step to the thirteenth step, finally, the best individual is converted and output in the twentieth step and the twenty-first step; The specific process is implemented as follows: 。
Citation Information
Patent Citations
Method of identifying cancer drive pathway
CN112270952A
Parameter-free nonlinear intelligent optimization method for identifying cancer driving pathway
CN114023383A