A cancer-driving pathway identification system

By building a cancer-driven pathway recognition system and using gene mutation data to establish an identification network, the problem of efficient identification of cancer-driven pathways is solved, and the theoretical and application value of cancer research and treatment is promoted.

CN119132392BActive Publication Date: 2025-08-08YAORONGYUN DIGITAL TECH (CHENGDU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411334298.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-24
Publication Date
2025-08-08
Estimated Expiration
2044-09-24

AI Technical Summary

Technical Problem

The prior art is difficult to identify cancer driving pathways efficiently and at low cost, resulting in a lack of theoretical basis for the study and treatment of cancer pathogenesis.

Method used

Build a cancer driver pathway recognition system, obtain cancer gene mutation data, establish a gene mutation similarity model, and use multiple parameters to build a cancer driver pathway recognition network to identify cancer driver pathways.

Benefits of technology

It has achieved efficient identification of cancer driving pathways, provided a theoretical basis for cancer pathogenesis, supported the clinical diagnosis of cancer and targeted drug development, and improved treatment effect and patient survival rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119132392B_ABST
    Figure CN119132392B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of cancer driver pathway identification, and in particular to a cancer driver pathway identification system. The system includes a memory, wherein the program commands stored in the memory can be executed by a processor to perform the following steps: obtaining cancer gene mutation data; using the cancer gene mutation data, obtaining a first parameter of a gene pair; constructing a gene mutation similarity model, combining the gene mutation similarity model and the cancer gene mutation data to obtain a gene mutation similarity parameter of the gene pair; establishing a cancer driver pathway identification network based on the first parameter, the gene mutation similarity parameter and the cancer gene mutation data; and identifying a cancer driver pathway based on the cancer driver pathway identification network. The present invention constructs multiple parameters based on cancer gene mutation data, and then constructs a cancer driver pathway identification network, and finally identifies the cancer driver pathway through this network, thereby solving the problem of high-efficiency and low-cost identification of cancer driver pathways.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cancer driver pathway identification, and in particular to a cancer driver pathway identification system. Background Art

[0002] Due to various factors such as changes in people's lifestyles, the aging of the population, and environmental pollution, the incidence and mortality of cancer are also rising, and it has become one of the most serious public health problems in the world.

[0003] The pathogenic mechanism of cancer is complex, and there are many factors that trigger cancer. Existing studies have confirmed that the occurrence and development of cancer are closely related to genomic mutations. Cancer is caused by the accumulation of abnormal gene mutations such as single nucleotide polymorphisms and copy number variations. Gene mutations may cause structural changes in proteins or differences in the expression of functional proteins, thereby causing an abnormal physiological environment and triggering cancer. Such mutations that drive the occurrence of cancer are called driver mutations, and the genes where the driver mutation sites are located are called driver genes. The corresponding mutations and genes that are not related to the occurrence of cancer are called passenger mutations and passenger genes. The accumulation of these passenger mutations / genes will not cause cancer.

[0004] Studies have found that at different biological scales, driver genes often aggregate to form driver pathways, which jointly perform certain biological functions. These driver factors can disrupt normal life activities. Therefore, identifying cancer driver pathways helps to gain a deeper understanding of the molecular mechanisms of cancer. By analyzing the genetic mutations and interactions in these pathways, we can reveal the key links in the occurrence and development of cancer, and provide a theoretical basis for the prevention and treatment of cancer, which is specifically manifested in the following aspects: helping doctors to more accurately determine the patient's cancer type and stage, so as to select more precise treatment methods and drugs, improve treatment effects and patient survival rates; providing new targets for drug development, and drug research and development targeting key genes or molecules in the pathway can develop more effective anti-cancer drugs with fewer side effects to meet the needs of clinical treatment; it is conducive to a deeper understanding of the role and regulatory mechanisms of biological molecules such as genes and proteins in life processes, and contribute to the development of life sciences; by analyzing whether there are specific driver pathway mutations in the patient's body, it can more accurately determine whether the patient has a certain type of cancer, providing strong support for early diagnosis and treatment.

[0005] However, it is difficult to accurately identify the driving factors associated with specific cancer categories from large-scale bio-omics data by relying solely on simple biological experiments or statistical analysis methods. Therefore, developing effective computational systems for large-scale cancer multi-omics genetic data to achieve accurate and efficient identification of cancer drivers and to find driving pathways that play an important role in promoting the occurrence and development of cancer is a major challenge in current cancer informatics research. Summary of the Invention

[0006] In response to the shortcomings of existing methods and the needs of practical applications, in order to solve the problem of identifying cancer driver pathways with high efficiency and low cost, the present invention provides a cancer driver pathway identification system, which includes a memory for storing a computer program, the computer program including program instructions. When the program instructions are executed by a processor, the processor performs the following steps:

[0007] Obtain cancer gene mutation data; use the cancer gene mutation data to obtain a first parameter of a gene pair; construct a gene mutation similarity model, and combine the gene mutation similarity model and the cancer gene mutation data to obtain gene mutation similarity parameters of the gene pair; establish a cancer driver pathway identification network using the first parameter, the gene mutation similarity parameter, and the cancer gene mutation data; and identify a cancer driver pathway based on the cancer driver pathway identification network.

[0008] The present invention uses cancer gene mutation data to construct multiple parameters, and then constructs a cancer driver pathway identification network. Ultimately, cancer driver pathways are identified through this network, solving the problem of high-efficiency and low-cost identification of cancer driver pathways. It provides a theoretical basis for deeper research and understanding of the pathogenesis of cancer, helps researchers strengthen their understanding of cancer, and has important theoretical and application value in many aspects such as clinical diagnosis of cancer, targeted drug development, and precise and personalized treatment of cancer patients.

[0009] Optionally, the processor further performs the following steps:

[0010] The cancer gene mutation data is subjected to a gene score; and the cancer gene mutation data is screened based on the gene score result. The present invention increases the proportion of driver genes by performing an initial screening of the gene mutation data score, which is beneficial for improving the calculation efficiency of the present invention.

[0011] Optionally, the gene scoring of the cancer gene mutation data satisfies the following formula:

[0012] ,in, Indicates gene Rating, represents the number of cancer patient samples, Indicates gene The number of gene mutation types, Indicates gene In the The first The present invention uses a model to evaluate genes based on the number of non-silent mutations in gene mutations. The results are specific and quantitative, making it easy to screen the cancer gene mutation data.

[0013] Optionally, the cancer gene mutation data is used to obtain a first parameter of a gene pair, which satisfies the following formula:

[0014] ,

[0015] in, Represents gene pairs The first parameter, represents the minimum function, represents the counting function, Indicates gene The covering set of Indicates gene The first parameter constructed by the present invention is conducive to coordinating and controlling the density of edges in the subsequent steps of establishing the network, and removing unnecessary edges further improves the reliability of the cancer driver pathway identification network.

[0016] Optionally, the constructed gene mutation similarity model satisfies the following formula: ,

[0017] in, Represents gene pairs The gene mutation similarity parameter, Indicates gene The mutation frequency, Indicates gene The mutation frequency, represents the number of cancer patient samples, Indicates the Genes in cancer patient samples to genes distance, , Indicates the Genes in cancer patient samples Mutation occurs, Indicates the Genes in cancer patient samples Mutation occurs, Indicates gene The covering set of The present invention utilizes the similarity of gene mutations between gene pairs to further remove unnecessary edges and improve the reliability of the cancer driver pathway identification network.

[0018] Optionally, establishing a cancer driver pathway identification network using the first parameter, the gene mutation similarity parameter, and the cancer gene mutation data comprises the following steps:

[0019] Based on the cancer gene mutation data, an initial cancer driver pathway identification network is established with genes as nodes; based on the first parameter and the gene mutation similarity parameter, edges of gene pairs in the initial cancer driver pathway identification network are determined; using the first parameter and the gene mutation similarity parameter, weights of the edges are determined; and combining the initial cancer driver pathway identification network, the edges, and the weights to establish the cancer driver pathway identification network. The cancer driver pathway identification network constructed by the present invention utilizes multiple parameters to determine edges and their weights, which facilitates accurate identification of driver pathways.

[0020] Optionally, determining the edges of gene pairs in the cancer driver pathway identification initial network based on the first parameter and the gene mutation similarity parameter comprises the following steps:

[0021] Setting a mutual exclusion threshold, a coverage threshold, a first parameter threshold, and a gene mutation similarity parameter threshold for the gene pair; and determining edges of the gene pairs in the cancer driver pathway identification initial network according to edge construction rules;

[0022] The edge construction rule satisfies the following conditions: when the gene pair satisfies the mutual exclusion greater than the mutual exclusion threshold, the coverage greater than the coverage threshold, the first parameter greater than the first parameter threshold, and the gene mutation similarity parameter greater than the gene mutation similarity parameter threshold, the gene pair has an undirected edge in the initial cancer driver pathway identification network; otherwise, the gene pair does not have an undirected edge in the initial cancer driver pathway identification network. By objectively determining the edges of gene pairs in the network, the present invention further improves the reliability of the cancer driver pathway identification network.

[0023] Optionally, the weight of the edge is determined by using the first parameter and the gene mutation similarity parameter, satisfying the following formula:

[0024] ,

[0025] Represents gene pairs represents the weight of the edge, Represents gene pairs The first parameter, Represents gene pairs The gene mutation similarity parameter, Represents gene pairs The mutual exclusion degree, Represents gene pairs The present invention uses multiple parameters to determine the weighting of edges, which can accurately reflect the relationship between gene pairs and further facilitate the identification of cancer driving pathways.

[0026] Optionally, identifying a cancer driver pathway based on the cancer driver pathway identification network comprises the following steps:

[0027] All maximal clusters are searched in the cancer driver pathway identification network to obtain a maximal cluster set; cluster weights of the maximal clusters are calculated, and the cluster weights are used to obtain maximal weight subclusters of the maximal clusters; and the maximal cluster set and the maximal weight subclusters are combined to obtain a cancer driver pathway set. By constructing a cancer driver pathway identification network, the present invention transforms the problem of finding driver pathways into the problem of finding maximal clusters in a base network graph, which helps improve the efficiency of identifying cancer driver pathways.

[0028] Optionally, the cancer driver pathway identification system further includes an input device, a processor, and an output device, wherein the input device, the processor, the output device, and the memory are interconnected, and the processor is configured to call and execute the program instructions. The cancer driver pathway identification system provided by the present invention has a compact structure and stable performance, further enhancing the overall applicability and practical application capabilities of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 A flowchart of program instructions for a cancer driver pathway identification system provided by an embodiment of the present invention;

[0030] Figure 2 A framework diagram of a cancer driver pathway identification system provided by an embodiment of the present invention;

[0031] Figure 3 A schematic diagram of the structure of a cancer driver pathway identification device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0032] Specific embodiments of the present invention will be described in detail below. It should be noted that the embodiments described herein are for illustrative purposes only and are not intended to limit the present invention. In the following description, numerous specific details are set forth to provide a thorough understanding of the present invention. However, it will be apparent to one of ordinary skill in the art that these specific details are not necessarily required to practice the present invention. In other instances, well-known circuits, software, or methods are not specifically described to avoid obscuring the present invention.

[0033] Throughout this specification, references to "one embodiment," "an embodiment," "an example," or "an example" mean that a particular feature, structure, or characteristic described in connection with the embodiment or example is included in at least one embodiment of the present invention. Therefore, appearances of the phrases "in one embodiment," "in an embodiment," "an example," or "an example" in various places throughout this specification are not necessarily all referring to the same embodiment or example. Furthermore, the particular features, structures, or characteristics may be combined in any suitable combinations and / or subcombinations in one or more embodiments or examples. Furthermore, those of ordinary skill in the art will appreciate that the figures provided herein are for illustrative purposes only and are not necessarily drawn to scale.

[0034] See also Figure 1 In order to solve the problem of identifying cancer driver pathways with high efficiency and low cost, the present invention provides a cancer driver pathway identification system, wherein the cancer driver pathway identification system includes a memory for storing a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the processor performs the following steps:

[0035] S1. Obtain cancer gene mutation data.

[0036] Genes are basic material units that carry genetic information and can guide protein synthesis through steps such as transcription and translation. In an embodiment, different cancer genomics data are obtained from public cancer genomics databases. Public cancer genomics databases include the Cancer Genome Atlas TCGA database, which is a large cancer genomics project jointly initiated by the National Cancer Institute and the National Human Genome Research Institute of the United States. It covers more than 30 different types of cancer and provides multi-omics data including DNA sequences, RNA sequences, DNA methylation, protein expression and clinical information; there are also some other commonly used cancer genomics databases, such as the International Cancer Genome Consortium ICGC database, the COSMIC database, the GDC database and the NCBI database.

[0037] Furthermore, after acquiring the data, it should be pre-processed to minimize data errors and achieve more accurate results. First, data cleaning is performed to fill in missing values and remove existing noise. Next, data integration is performed to integrate data from multiple databases with different sources and characteristics. Finally, data redundancy is processed to make the data integrated, standardized, and simplified.

[0038] In the embodiment, the acquired cancer gene mutation data can be expressed as Matrix ,Include cancer patient samples, mutant genes, matrix The elements in satisfy the following conditions: if a mutated gene is present, the value is 1; if a mutated gene is not present, the value is 0. It is understood that cancer patient samples can be from patients with the same cancer or patients with different cancers. Therefore, this embodiment can ultimately identify single cancer-driving pathways and synergistic pathways.

[0039] Furthermore, in order to improve the computational efficiency of the embodiment, the gene mutation data may be scored to perform a preliminary screening to eliminate the mixed passenger genes, including the following steps:

[0040] S11. Perform gene scoring on the cancer gene mutation data.

[0041] Specifically, the cancer gene mutation data is scored to satisfy the following formula:

[0042] ,in, Indicates gene Rating, represents the number of cancer patient samples, Indicates gene The number of gene mutation types, Indicates gene In the The first The number of non-silent mutations in each gene mutation type.

[0043] S12. Screening the cancer gene mutation data according to the gene scoring results.

[0044] Specifically, the greater the number of non-silent mutations in a gene, the greater the score result. Non-silent mutations can change the structure and function of gene products, and may further cause mutations that change biological traits. Therefore, the probability of a gene mutation becoming a passenger gene is reduced. The gene score result threshold can be reasonably set through experimental data to eliminate cancer gene mutation data with a gene score result less than the gene score result threshold. In the embodiment, the gene score result threshold is set to 0.1.

[0045] S2. Using the cancer gene mutation data, obtain a first parameter of a gene pair.

[0046] In the embodiment, the first parameter of the gene pair is obtained by using the cancer gene mutation data in step S2, and satisfies the following formula:

[0047] ,

[0048] in, Represents gene pairs The first parameter, represents the minimum function, represents the counting function, Indicates gene The covering set of Indicates gene The covering set of .

[0049] It should be understood that the coverage set refers to the matrix representing the gene in the cancer gene mutation data. The gene set in , that is, the patient set with gene mutations.

[0050] S3. Construct a gene mutation similarity model, and combine the gene mutation similarity model with the cancer gene mutation data to obtain gene mutation similarity parameters of gene pairs.

[0051] Specifically, the gene mutation similarity model constructed in step S3 satisfies the following formula:

[0052] ,

[0053] in, Represents gene pairs The gene mutation similarity parameter, Indicates gene The mutation frequency, Indicates gene The mutation frequency, represents the number of cancer patient samples, Indicates the Genes in cancer patient samples to genes distance, , Indicates the Genes in cancer patient samples Mutation occurs, Indicates the Genes in cancer patient samples Mutation occurs, Indicates gene The covering set of Represents a counting function. The mutation frequency of a gene is the number of genes in the matrix The proportion of patients with this gene mutation to the total number of samples.

[0054] S4. Establishing a cancer driver pathway identification network using the first parameter, the gene mutation similarity parameter, and the cancer gene mutation data.

[0055] In the embodiment, step S4, establishing a cancer driver pathway identification network using the first parameter, the gene mutation similarity parameter, and the cancer gene mutation data, includes the following steps:

[0056] S41. Based on the cancer gene mutation data, an initial network for identifying cancer driver pathways is established with genes as nodes.

[0057] In an embodiment, the matrix Using genes as nodes, an initial network for identifying boundless cancer-driving pathways was established.

[0058] S42. Based on the first parameter and the gene mutation similarity parameter, determine the edges of the gene pairs in the cancer driver pathway identification initial network.

[0059] Specifically, the mutual exclusivity threshold, coverage threshold, first parameter threshold and gene mutation similarity parameter threshold of the gene pair are set through experiments or expert discussion;

[0060] According to the edge construction rules, edges of gene pairs in the cancer driver pathway identification initial network are determined.

[0061] Furthermore, the edge construction rule satisfies:

[0062] When the gene pair satisfies the conditions that the mutual exclusivity is greater than the mutual exclusivity threshold, the coverage is greater than the coverage threshold, the first parameter is greater than the first parameter threshold, and the gene mutation similarity parameter is greater than the gene mutation similarity parameter threshold, the gene pair has an undirected edge in the cancer driver pathway identification initial network;

[0063] Otherwise, the gene pair has no undirected edge in the cancer driver pathway identification initial network.

[0064] It should be understood that the mutant gene network is constructed based on conditions such as high coverage and high mutual exclusivity satisfied between gene pairs, which can reduce the complexity of the algorithm. However, the density of the network edges constructed under different conditions is also different. In the embodiment, the mutual exclusivity threshold is 0.85, the coverage threshold is 0.09, the first parameter threshold is 0.13, and the gene mutation similarity parameter threshold is 0.382.

[0065] S43. Determine the weight of the edge using the first parameter and the gene mutation similarity parameter.

[0066] Specifically, in step S43, the weight of the edge is determined by using the first parameter and the gene mutation similarity parameter, satisfying the following formula:

[0067] ,

[0068] Represents gene pairs represents the weight of the edge, Represents gene pairs The first parameter, Represents gene pairs The gene mutation similarity parameter, Represents gene pairs The mutual exclusion degree, Represents gene pairs coverage.

[0069] Furthermore, the mutual exclusivity of gene pairs satisfies the following formula:

[0070] , Represents gene pairs The mutual exclusion degree of

[0071] The coverage of a gene pair satisfies the following formula:

[0072] , Represents gene pairs coverage.

[0073] S44. Combining the cancer driver pathway identification initial network, the edges, and the weights, to establish the cancer driver pathway identification network.

[0074] The edges are added to the initial cancer driver pathway identification network, and the edges are weighted using the weights to establish the cancer driver pathway identification network.

[0075] S5. Identify cancer driver pathways based on the cancer driver pathway identification network.

[0076] In an embodiment, step S5 of identifying cancer driver pathways based on the cancer driver pathway identification network comprises the following steps:

[0077] S51. Searching for all maximal clusters in the cancer driver pathway identification network to obtain a maximal cluster set.

[0078] Specifically, a variety of algorithms, including the Bron-Kerbosch algorithm, can be used to search for all maximal clusters in the cancer driver pathway identification network to obtain a maximal cluster set.

[0079] S52. Calculate the cluster weight of the maximum cluster, and use the cluster weight to obtain a maximum-weight sub-cluster of the maximum cluster.

[0080] Calculate the group weight of the maximum group, satisfying the following formula:

[0081] ,in, delegation The group weight, delegation The node set mutual exclusion degree, delegation The node set coverage, delegation The sum of the edge weights of the node set.

[0082] Furthermore, for each maximal cluster, if the cluster weight of the subcluster obtained by deleting a node is higher than the weight before deleting the node, the subcluster is considered to be better; iteratively delete nodes until the weight no longer increases, and a maximal weight subcluster can be obtained.

[0083] S53. Combining the maximal cluster set and the maximal weight sub-cluster to obtain a cancer driving pathway set.

[0084] It should be understood that since the number of genes included in the driver pathway is usually not less than 3, if the maximum weight subcluster meets the conditions: the number of nodes is greater than or equal to 3 and the coverage is greater than 0.3, then the maximum weight subcluster is determined to be a driver pathway. Summarizing all the maximum weight subclusters that meet the conditions is the cancer driver pathway set.

[0085] Therefore, this embodiment can identify a more complete cancer driving pathway in a shorter time, and does not require the number of genes in the driving pathway to be specified in advance.

[0086] See also Figure 2 In an embodiment, the present invention provides a cancer driver pathway identification system, comprising: an input device, an output device, a processor, and a memory, wherein the input device, the output device, the processor, and the memory are interconnected, and the memory contains program instructions. When the program instructions are executed by the processor, the following steps are performed:

[0087] Obtain cancer gene mutation data; use the cancer gene mutation data to obtain a first parameter of a gene pair; construct a gene mutation similarity model, and combine the gene mutation similarity model and the cancer gene mutation data to obtain gene mutation similarity parameters of the gene pair; establish a cancer driver pathway identification network using the first parameter, the gene mutation similarity parameter, and the cancer gene mutation data; and identify a cancer driver pathway based on the cancer driver pathway identification network.

[0088] The cancer-driving pathway identification system of the present invention has a compact structure and stable performance, further enhancing the overall applicability and practical application capabilities of the present invention.

[0089] In an embodiment, the processor may be a central processing unit (CPU), which may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc. The input device may be used to obtain data information. The output device may be used to output the results obtained by storing the program instructions contained in the computer program in the memory provided by the present invention. The memory may include a read-only memory and a random access memory, and provides instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory.

[0090] In yet another alternative embodiment, see Figure 3 ,This embodiment also provides a cancer driving pathway identification device, such as Figure 3 As shown, including:

[0091] The memory 10 is used to store computer programs; the processor 20 is used to execute the computer programs to implement the program instruction steps of the cancer driver pathway identification system described above. The memory 10, processor 20, communication interface 31, and communication bus 32 are all connected to each other via the communication bus 32.

[0092] In an embodiment, the memory 10 is used to store one or more program instructions. The memory 10 may store program instructions for implementing the following functions:

[0093] Obtain cancer gene mutation data; use the cancer gene mutation data to obtain a first parameter of a gene pair; construct a gene mutation similarity model, and combine the gene mutation similarity model and the cancer gene mutation data to obtain gene mutation similarity parameters of the gene pair; establish a cancer driver pathway identification network using the first parameter, the gene mutation similarity parameter, and the cancer gene mutation data; and identify a cancer driver pathway based on the cancer driver pathway identification network.

[0094] In one possible implementation, the memory 10 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and at least one application required for a function, etc.; the data storage area may store data created during use. In addition, the memory 10 may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include NVRAM. The memory stores an operating system and operating instructions, executable modules or data structures, or a subset thereof, or an extended set thereof, wherein the operating instructions may include various operating instructions for implementing various operations. The operating system may include various system programs for implementing various basic tasks and processing hardware-based tasks.

[0095] The processor 20 may be a central processing unit (CPU), an application-specific integrated circuit, a digital signal processor, a field programmable gate array, or other programmable logic device. The processor 20 may be a microprocessor or any conventional processor. The processor 20 may call programs stored in the memory 10. The communication interface 31 may be an interface of a communication module for connecting to other devices or systems.

[0096] Of course, it needs to be explained that Figure 3 The structure shown does not constitute a limitation on the cancer driving pathway identification device in this embodiment. In actual applications, the cancer driving pathway identification device may include Figure 3 More or fewer components than shown, or combinations of certain components.

[0097] An embodiment further provides a storage medium having a computer program stored thereon. When the computer program is executed by a processor, the program instruction steps of the cancer driver pathway identification system are implemented.

[0098] The storage medium may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., which can store program codes.

[0099] In summary, the present invention utilizes cancer gene mutation data to construct multiple parameters, thereby constructing a cancer driver pathway identification network. Ultimately, this network identifies cancer driver pathways, solving the problem of efficient and low-cost identification of cancer driver pathways. This provides a theoretical basis for deeper research and understanding of cancer pathogenesis, helping researchers to enhance their understanding of cancer. This method has important theoretical and applied value in many areas, including clinical cancer diagnosis, targeted drug development, and precise, personalized treatment for cancer patients. Therefore, the present invention effectively overcomes the shortcomings of the existing technology and possesses high industrial application value.

[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention, and they should all be included in the scope described in the present invention.

Claims

1. A cancer driver pathway identification system, characterized in that: The cancer driver pathway identification system includes a memory for storing a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the processor performs the following steps: Obtain cancer gene mutation data; Using the cancer gene mutation data, obtaining a first parameter of a gene pair; Constructing a gene mutation similarity model, and combining the gene mutation similarity model with the cancer gene mutation data to obtain gene mutation similarity parameters of gene pairs; Establishing a cancer driver pathway identification network using the first parameter, the gene mutation similarity parameter, and the cancer gene mutation data; identifying cancer driver pathways based on the cancer driver pathway identification network; The constructed gene mutation similarity model satisfies the following formula: , in, Represents gene pairs The gene mutation similarity parameter, Indicates gene The mutation frequency, Indicates gene The mutation frequency, represents the number of cancer patient samples, Indicates the Genes in cancer patient samples to genes distance, , Indicates the Genes in cancer patient samples Mutation occurs, Indicates the Genes in cancer patient samples Mutation occurs, Indicates gene The covering set of represents the counting function; The method of establishing a cancer driver pathway identification network using the first parameter, the gene mutation similarity parameter, and the cancer gene mutation data comprises the following steps: Based on the cancer gene mutation data, an initial network for identifying cancer driver pathways is established with genes as nodes; determining edges of gene pairs in the cancer driver pathway identification initial network based on the first parameter and the gene mutation similarity parameter; Determining the weight of the edge using the first parameter and the gene mutation similarity parameter; combining the cancer driver pathway identification initial network, the edges, and the weights to establish the cancer driver pathway identification network; The weight of the edge is determined by using the first parameter and the gene mutation similarity parameter, satisfying the following formula: , Represents gene pairs represents the weight of the edge, Represents gene pairs The first parameter, Represents gene pairs The gene mutation similarity parameter, Represents gene pairs The mutual exclusion degree, Represents gene pairs coverage; The cancer gene mutation data is used to obtain a first parameter of a gene pair, which satisfies the following formula: , in, Represents gene pairs The first parameter, represents the minimum function, represents the counting function, Indicates gene The covering set of Indicates gene The covering set of .

2. A cancer driver pathway identification system according to claim 1, characterized in that: The processor also performs the following steps: performing gene scoring on the cancer gene mutation data; The cancer gene mutation data are screened according to the gene scoring results.

3. The cancer driver pathway identification system according to claim 2, characterized in that: The gene scoring of the cancer gene mutation data satisfies the following formula: , in, Indicates gene Rating, represents the number of cancer patient samples, Indicates gene The number of gene mutation types, Indicates gene In the The first The number of non-silent mutations in each gene mutation type.

4. The cancer driver pathway identification system according to claim 1, characterized in that: The step of determining the edges of gene pairs in the cancer driver pathway identification initial network based on the first parameter and the gene mutation similarity parameter comprises the following steps: Set the mutual exclusivity threshold, coverage threshold, first parameter threshold and gene mutation similarity parameter threshold of gene pairs; Determining edges of gene pairs in the cancer driver pathway identification initial network according to edge construction rules; The edge construction rules satisfy: When the gene pair satisfies the conditions that the mutual exclusivity is greater than the mutual exclusivity threshold, the coverage is greater than the coverage threshold, the first parameter is greater than the first parameter threshold, and the gene mutation similarity parameter is greater than the gene mutation similarity parameter threshold, the gene pair has an undirected edge in the cancer driver pathway identification initial network; Otherwise, the gene pair has no undirected edge in the cancer driver pathway identification initial network.

5. The cancer driver pathway identification system according to claim 1, characterized in that: The method of identifying cancer driver pathways based on the cancer driver pathway identification network comprises the following steps: Searching for all maximal clusters in the cancer driver pathway identification network to obtain a maximal cluster set; Calculating the cluster weight of the maximum cluster, and using the cluster weight to obtain a maximum-weight subcluster of the maximum cluster; The maximal cluster set and the maximal weight subclusters are combined to obtain a cancer driver pathway set.

6. The cancer driver pathway identification system according to any one of claims 1 to 5, characterized in that: The system further comprises an input device, a processor and an output device, wherein the input device, the processor, the output device and the memory are connected to each other, and the processor is configured to call and execute the program instructions.