Cancer Co-Driver Module Identification System Based on Single-Cell Data
By integrating multiple biological data and building a functional association network, combining clustering and greedy search algorithms, the accuracy of the cancer synergistic driver module recognition in the existing technology is solved, and the accurate identification of the cancer synergistic driver module and the interpretability of the efficient identification results are achieved.
Patent Information
- Application Number
- CN202210724340.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-24
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-06-24
AI Technical Summary
The prior art is difficult to accurately identify cancer synergistically driven modules, especially the problems of insufficient combined utilization and modeling of commonality and specificity at the subtype level.
By integrating copy number variation data and single nucleotide polymorphism data, a functional association network for single-cell gene expression data is constructed, and the driver module is identified using overlapping Markov clustering and greedy search algorithms, and finally the coordinated driver module is determined through the inter-module distance function.
The precise identification of the cancer synergistic driver module is achieved, the system's operating efficiency and the interpretability of the recognition results are improved, and the subtype data can be used more fully for the identification of cancer-related genes.
Smart Images

Figure CN115148286B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence data mining classification and bioinformatics, and particularly relates to a system for identifying cancer co-driving modules based on single-cell data. Background Technique
[0002] The statements in this part merely provide background techniques related to the present invention and do not necessarily constitute prior art.
[0003] With the development of high-throughput biotechnology, large-scale cancer genome projects, such as The Cancer Genome Atlas (TCGA), International Cancer Genome Consortium (ICGC), etc., have generated and accumulated rich high-throughput multi-omics cancer data, providing support for researchers to deeply analyze the cancer mechanism from a systematic level. However, it is difficult to accurately identify driver mutations and driver genes related to specific cancer types from large-scale bi-omics data only by relying on biological experiments or simple statistical analysis methods. Therefore, for large-scale disease genetic data, developing effective computational methods to achieve accurate and efficient identification of cancer driver mutation / gene sets is a major challenge in current cancer informatics research. Accurately identifying carcinogenic genetic factors has important theoretical and application values in many aspects such as cancer diagnosis, targeted drug development, and precise and personalized treatment of cancer patients.
[0004] Currently, the main methods for identifying cancer co-driving modules are as follows: identifying functional modules that satisfy the mutual exclusion rule in the gene interaction network; reconstructing the gene interaction network using the heat diffusion algorithm and then detecting sub-networks with the best coverage and mutual exclusion from it; establishing a weight function on somatic mutation data and then combining greedy and MCMC methods to optimize the identification of gene sets with the highest weight as driving modules; combining the characteristic that genes with similar expressions usually jointly perform a certain biological function, introducing Binary Liner Programming (BLP) and genetic algorithms to more fully introduce genetic information to guide the identification of driving modules; using integer linear programming to simultaneously detect multiple functional modules with high weights as co-driving modules; combining high co-occurrence of mutations between modules with high coverage and high mutual exclusion within modules, and using BLP to identify driving modules with synergistic effects; using greedy search on the gene signal network to explore mutually exclusive modules with common downstream events, and then designing a double-regularized biclustering method to discriminate mutually exclusive modules within the same cluster as co-driving modules; first performing hierarchical clustering on genes to obtain co-mutated modules, mapping the relationships between co-modules to the pathway level through link prediction, supplementing potential connections between pathways, and finally identifying co-driving modules based on modules and the updated pathway interaction network, etc. However, most of the existing solutions cannot accurately identify cancer co-driving modules. Summary of the Invention
[0005] In view of the problems that genetic information with low and medium abundance is easily lost in the identification of existing driver module sets, and the commonalities and specificities of cancers at the subtype level are not fully utilized and modeled, the present invention introduces single-cell sequencing data and cancer subtype data, and proposes a collaborative driver module identification system based on single-cell sequencing data and cancer subtypes. First, copy number variation data and single nucleotide polymorphism data are integrated into mutation data, and gene alignment and preliminary screening are performed on multi-omics data to obtain available multi-omics data; then, based on single-cell gene expression data, cell-level specific gene co-expression networks are constructed for each subtype and normal cells respectively, and these networks are combined into an expression association network that can reflect the commonalities between different subtypes, and this network represents gene co-expression relationships that are prevalent in different subtypes and significantly different from normal cells; then, a gene function interaction network is introduced, and it is fused with the original network and optimized into a functional association network to strengthen the functional connections between genes, reduce the network complexity, and speed up the operation of the algorithm; subsequently, overlapping Markov clustering is applied to this network to obtain multiple functional clusters. To effectively mine gene expression data and utilize the high coverage and high mutual exclusivity of driver modules, the present invention respectively introduces differential expression analysis and gene mutation data, comprehensively constructs a module weight evaluation function, and then uses greedy search to identify driver modules on the functional clusters; finally, based on the functional connections and mutation co-occurrence between modules, a module distance function is defined, which can effectively determine driver modules as collaborative driver modules, thus realizing the accurate identification of cancer collaborative driver modules.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] The first aspect of the present invention provides a cancer collaborative driver module identification system based on single-cell data.
[0008] A cancer collaborative driver module identification system based on single-cell data, comprising:
[0009] A data preprocessing module, configured to: integrate copy number variation data and single nucleotide polymorphism data into mutation data, perform gene alignment and preliminary screening on multi-omics data, and obtain available multi-omics data, where the multi-omics data at least includes single-cell gene expression data;
[0010] A functional association network construction module, configured to: construct a functional association network based on single-cell gene expression data and a gene interaction network;
[0011] A functional cluster acquisition module, configured to: apply overlapping Markov clustering to the functional association network to obtain multiple functional clusters;
[0012] A driver module identification module, configured to: identify driver modules on the functional clusters using greedy search based on a module weight evaluation function;
[0013] A co-driving module recognition module, configured to: determine the recognized driving modules as co-driving modules according to the distance function between driving modules.
[0014] The second aspect of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the following steps are implemented:
[0015] Integrate copy number variation data and single nucleotide polymorphism data into mutation data, perform gene alignment and preliminary screening on multi-omics data to obtain available multi-omics data, and the multi-omics data at least includes single-cell gene expression data;
[0016] Construct a functional association network based on single-cell gene expression data and gene interaction network;
[0017] Apply overlapping Markov clustering to the functional association network to obtain multiple functional clusters;
[0018] Based on the module weight evaluation function, use greedy search to identify driving modules on the functional clusters;
[0019] According to the distance function between driving modules, determine the recognized driving modules as co-driving modules.
[0020] The third aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, the following steps are implemented:
[0021] Integrate copy number variation data and single nucleotide polymorphism data into mutation data, perform gene alignment and preliminary screening on multi-omics data to obtain available multi-omics data, and the multi-omics data at least includes single-cell gene expression data;
[0022] Construct a functional association network based on single-cell gene expression data and gene interaction network;
[0023] Apply overlapping Markov clustering to the functional association network to obtain multiple functional clusters;
[0024] Based on the module weight evaluation function, use greedy search to identify driving modules on the functional clusters;
[0025] According to the distance function between driving modules, determine the recognized driving modules as co-driving modules.
[0026] Compared with the prior art, the beneficial effects of the present invention are:
[0027] 1. The present invention first integrates copy number variation data and single nucleotide polymorphism data into comprehensive mutation data to avoid data loss and untrustworthiness problems that may exist in single mutation data. Then, gene alignment and preliminary screening are performed on multi-omics data including single-cell sequencing data, reducing the data volume, obtaining available multi-omics data, and improving the operating efficiency of the system.
[0028] 2. Based on single-cell gene expression data, the present invention constructs cell-level specific gene co-expression networks for each subtype and normal cells respectively, and combines these networks into an expression association network that can reflect the commonalities among different subtypes. This network represents gene co-expression relationships that are prevalent in different subtypes and significantly different from normal cells, and can be used for the identification of cancer-related genes.
[0029] 3. The present invention introduces a gene function interaction network, and fuses it with the original network to optimize it into a function association network to strengthen the functional connection between genes, retain more important gene co-expression relationships, and reduce the network complexity, further improving the operating speed of the system.
[0030] 4. The present invention applies overlapping Markov clustering to this network to obtain multiple functional clusters. These functional clusters allow gene duplication, which conforms to the general laws of life activities, and improves the high coverage and high mutual exclusivity of mining gene expression data and using driver modules.
[0031] 5. The present invention separately introduces differential expression analysis and gene mutation data, comprehensively constructs a module weight evaluation function, and then uses greedy search to identify driver modules by combining the evaluation function on the functional clusters, and selects the candidate module with the highest score as the driver module; based on the functional connection and mutation co-occurrence between modules, a module distance function is comprehensively defined, and through this distance, clustering methods can be effectively used to determine driver modules as co-driver modules, thus achieving the precise identification of cancer co-driver modules.
[0032] Advantages of additional aspects of the present invention will be partly given in the following description, partly will become obvious from the following description, or will be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0034] Figure 1 It is a schematic structural diagram of a cancer co-driver module identification system based on single-cell data provided for Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0035] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0036] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used in this embodiment have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0037] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0038] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0039] Embodiment 1:
[0040] As Figure 1 shown, Embodiment 1 of the present invention provides a cancer co-driving module identification system based on single-cell data and cancer subtypes, including:
[0041] A data preprocessing module configured to: integrate copy number variation data and single nucleotide polymorphism data into mutation data, and perform gene alignment and preliminary screening on multi-omics data;
[0042] A functional association network construction module configured to: construct a functional association network based on single-cell gene expression data and gene interaction networks;
[0043] A functional cluster acquisition module configured to: apply overlapping Markov clustering to the functional association network to obtain multiple functional clusters;
[0044] A driving module identification module configured to: comprehensively construct a module weight evaluation function and use greedy search to identify driving modules on the functional clusters;
[0045] A co-driving module identification module configured to: define a distance function between driving modules and determine the driving modules as co-driving modules.
[0046] Integrating copy number variation data and single nucleotide polymorphism data into mutation data, and performing gene alignment and preliminary screening on multi-omics data specifically includes:
[0047] Integrate the collected copy number variation data and single nucleotide polymorphism data into mutation data;
[0048] Align genes and perform preliminary screening on multi-omics data based on the obtained mutation data.
[0049] In this embodiment, the copy number variation data refers to the copy number variation data of the patient. The data is 10,000-dimensional, and each dimension corresponds to a gene locus of the patient. When there is a copy number duplication or deletion, the gene data is 1, and when the copy number is normal, it is 0.
[0050] In this embodiment, the single nucleotide polymorphism data refers to the single nucleotide polymorphism data of the patient. The data is 10,000-dimensional, and each dimension corresponds to a gene locus of the patient. When a single nucleotide polymorphism occurs, the gene data is 1, and when no single nucleotide polymorphism occurs, it is 0.
[0051] In this embodiment, the multi-omics data refers to: in addition to the above two types of data, it also includes: single-cell gene expression data divided into different cancer subtypes. The data is 10,000-dimensional, and each dimension corresponds to a gene locus of patients with different subtypes, and its value represents the expression level of the gene. The gene interaction network, the data is 10,000-dimensional, and each dimension corresponds to a gene, and its value describes the interaction intensity when genes jointly participate in life processes such as signal transduction, energy and material metabolism, and cell cycle regulation.
[0052] First, obtain the copy number variation data and the single nucleotide polymorphism data, and integrate the copy number variation data and the single nucleotide polymorphism data according to the mutation situation of the gene on the batch of samples to construct the mutation data. When a gene has a copy number variation or a single nucleotide polymorphism occurs, the value of the corresponding mutation data is one.
[0053] Secondly, the preliminary screening of the multi-omics data refers to, based on the previously obtained mutation data, comprehensively considering the mutation frequency of the gene and the actual situation of the patient, and selecting the genes with the top comprehensive mutation frequency to participate in the subsequent experiments.
[0054] Exemplarily, integrating the copy number variation data and the single nucleotide polymorphism data as the mutation data, and performing gene alignment and preliminary screening on the multi-omics data specifically includes:
[0055] S1011, integrate the collected copy number variation data and the single nucleotide polymorphism data as the mutation data. The specific steps are as follows:
[0056] Obtain the copy number variation data and the single nucleotide polymorphism data, and integrate the copy number variation data and the single nucleotide polymorphism data according to the mutation situation of the gene on the batch of samples to construct the mutation data. For any gene i and any sample p, its value is specifically defined as follows:
[0057]
[0058] Wherein, is mutation data, m represents the total number of genes investigated, b represents the total number of patients investigated, i and p correspond to gene i and patient p. When the SNP and CNV values in the above formula are greater than zero, it indicates that the corresponding gene has a single nucleotide polymorphism or copy number variation.
[0059] S1012, perform gene alignment and preliminary screening on multi-omics data based on the acquired mutation data. The specific steps are as follows:
[0060] First, since genes with lower mutation frequencies are mostly passenger genes that have no effect on cancer, in order to reduce computational costs and make the results more interpretable, only genes with mutation frequencies exceeding a certain threshold are retained. The present invention defines the comprehensive mutation frequency of each gene as follows:
[0061]
[0062] Where p represents the sample (patient) in the mutation data, |p| represents the total number of genes that have mutated in sample p, and Γ(i) represents the sample set in which gene i has mutated. The above formula comprehensively considers the two dimensions of gene and sample, reducing the impact of bad data and pathological samples;
[0063] Secondly, based on the obtained comprehensive mutation frequencies, the multi-omics data were aligned and screened to retain genes that existed in various data and had a higher comprehensive mutation frequency.
[0064] In this embodiment, a functional association network is constructed based on single-cell gene expression data and gene interaction network, specifically including:
[0065] Based on single-cell gene expression data, specific gene co-expression networks at the cell level were constructed for each subtype and normal cells.
[0066] These networks were combined into an expression association network that could reflect the commonalities among different subtypes;
[0067] The gene function interaction network was introduced and fused with the co-expression association network to optimize it into a functional association network.
[0068] The present invention will comprehensively utilize the single-cell gene expression data corresponding to multiple cancer subtypes and normal cells, enhance the gene co-expression relationship that is prevalent among multiple cancer subtypes, and reduce the adverse effects of noise gene associations that have little or no correlation with cancer development, and obtain a gene expression association network that can reflect the commonality of subtypes. On this basis, a gene interaction network is introduced to construct a gene molecular function association network that is instructive for exploring the pathogenesis of cancer, reduce the complexity of the original gene expression association network, enhance the biological function connection between genes, provide guidance for the subsequent identification of driver genes and driver modules, and provide support for the interpretability of experimental results.
[0069] Exemplarily, a functional association network is constructed based on single-cell gene expression data and a gene interaction network, which specifically includes:
[0070] S1021. Construct specific gene co-expression networks at the cellular level according to the single-cell gene expression data of different subtypes and normal cells respectively. The specific steps are as follows:
[0071] First, use the single-cell gene expression data to construct specific expression association networks corresponding to different subtypes of the target cancer. For gene pair i, j, the absolute value of the Spearman correlation coefficient is used to measure the co-expression relationship between genes:
[0072]
[0073] where and respectively represent the expression association matrix of the s-th subtype and the single-cell gene expression data matrix, m represents the total number of genes under investigation, d s represents the number of cells of the s-th cancer subtype, and i, j correspond to genes i and j. respectively represent the corresponding expression vectors of gene i (j) in the s-th subtype; Spearman() represents the Spearman correlation coefficient, which describes the non-parametric correlation strength of the expression between two genes. The larger its absolute value, the stronger the correlation.
[0074] Genes that are functionally associated in life activities often show a co-expression trend. Therefore the closer the value of is to 1, the more likely it is that genes i and j perform the same function in the process of cancer formation.
[0075] Similarly, the present invention uses the single-cell expression data of normal cells to generate its corresponding co-expression association network E n :
[0076]
[0077] Here and respectively represent the expression association matrix of normal cells and the gene expression data matrix, and c represents the number of normal cells. E n will be used to screen out the co-expression associations in E s that are not relevant to the target cancer.
[0078] S1022. Combine these networks into an expression association network that can reflect the commonalities between different subtypes. The specific steps are as follows:
[0079] First, based on the existing biological research results, the present invention proposes the following hypothesis:
[0080] 1) Gene pairs that maintain a strong association across different subtypes contribute more to cancer development;
[0081] 2) Genes that are similarly expressed in both tumor and normal cells may be related to the general activities of living organisms;
[0082] 3) Genes with a large difference in co-expression relationships between tumor and normal cells are more likely to drive cancer development.
[0083] Based on the above assumptions, the present invention fuses multiple co-expression networks corresponding to cancer subtypes and normal cells, comprehensively utilizes the rich genetic information contained in single-cell sequencing data, and defines a cancer gene co-expression association network:
[0084]
[0085] where represents the fused co-expression association network matrix, u represents the total number of subtypes, and E s -E n removes potential co-expression associations unrelated to cancer in the subtype association network. Through the above formula, the present invention can integrate the expression association networks from different subtypes and normal cells, and obtain an expression association network that can describe the gene co-expression relationships that widely exist in each subtype of a specific cancer and are significantly different from normal samples.
[0086] S1023, introduce a gene interaction network, and fuse it with the co-expression association network to optimize it into a functional association network. The specific steps are as follows:
[0087] First, since it is difficult to accurately describe the genetic associations between genes using only the Spearman correlation coefficient, and the generated co-expression network is too large in scale, it may retain a large number of low-quality gene co-expression relationships, and high spatio-temporal complexity will also be faced during subsequent computational analysis. Therefore, the present invention introduces a gene interaction network as prior knowledge. G describes the interaction strength between genes when they jointly participate in life processes such as signal transduction, energy and material metabolism, and cell cycle regulation. Using G as prior knowledge to guide the optimization of E can strengthen the functional cooperation relationship between genes in the functional association network, make the final recognition result more interpretable, and also effectively reduce the network complexity and improve the operation efficiency of the algorithm. The specific optimization equation is as follows:
[0088] F = E ⊙ G.
[0089] where ⊙ is the Hadamard product between matrices, It is an optimized cancer-associated gene functional association network matrix. By introducing the gene interaction network into the obtained expression association network through the above formula, the present invention can obtain a functional association network that can support the subsequent identification of cancer driver modules.
[0090] In this embodiment, the overlapping Markov clustering is applied to the functional association network to obtain multiple functional clusters, specifically including:
[0091] Markov Clustering (MCL), as a commonly used clustering algorithm, is widely applied to the analysis of important biological networks such as protein interaction networks and gene interaction networks. MCL is a graph clustering algorithm based on the transition probability (random walk) between different nodes. It performs a series of operations such as inflation and expansion on the transition matrix to make it converge. Nodes whose rows are not all zero in the final state are called attractor nodes, which gather all the points with positive values in their rows and can be regarded as clustering centers. Compared with traditional clustering algorithms such as K-means, Markov clustering is more tolerant of the noise in the data and can effectively discover high-quality functional clusters.
[0092] Traditional Markov clustering does not share elements between clusters, and the same biomolecule such as a gene or a protein can only be assigned to one functional cluster. However, research shows that biomolecules usually jointly perform multiple different life functions and jointly drive life activities. In response to this situation, the present invention introduces overlapping Markov clustering. This algorithm iteratively runs Markov clustering to generate a large number of clusters, and uses a penalty function to penalize the attractor nodes obtained by iteration to ensure that the clusters generated in each iteration are inconsistent with the previous ones, thereby ensuring that the same gene has the opportunity to be assigned to different functional clusters in different iterations.
[0093] First, the present invention initializes an attractor node indicator vector, which is used to represent the number of times a certain node (gene) is identified as an attractor node in multiple Markov clusterings and is used to penalize the attractor nodes:
[0094] Secondly, the Markov clustering algorithm is run in a loop on the functional association network. Each time the Markov matrix converges, the functional clusters generated in this round of clustering are collected, and the attractor nodes obtained in this round of clustering are marked.
[0095] Among them, in each Markov clustering process, the nodes are penalized through an inflation function according to the marked attractor node indicator vector. By penalizing the attractor nodes, the overlapping Markov clustering realizes that different functional clusters can be obtained in each iteration, and a single gene can be shared by multiple functional clusters, which conforms to the characteristic that genes perform different functions under different conditions in biology.
[0096] Exemplarily, the overlapping Markov clustering is applied to the functional association network to obtain multiple functional clusters, specifically including:
[0097] S103. Apply the overlapping Markov clustering to the functional association network to obtain multiple functional clusters. The specific steps are as follows:
[0098] First, initialize the attracting node indicator vector, which is used to represent the number of times a certain node (gene) is identified as an attracting node in multiple Markov clusterings, and is used to penalize the attracting nodes:
[0099] a = [0, 0, …, 0]
[0100] Second, run the Markov clustering algorithm in a loop on the functional association network. Each time the Markov matrix converges, collect the functional clusters generated in this round of clustering:
[0101]
[0102] Here is the set of all collected functional clusters, represents the functional clusters obtained in each iteration. And mark the attracting nodes obtained in this round of clustering. For the attracting node v, the formula for this mark is:
[0103] a(v) = a(v) + 1
[0104] Among them, in each Markov clustering process, according to the marked attracting node indicator vector, the nodes are penalized through the inflation function. The inflation function Inflate(P, r, a, β) represents performing an inflation operation on P. For any gene pair i, j, this operation is specifically as follows:
[0105]
[0106] Among them, the inflation coefficient r and the penalty ratio β are set to 2 and 1.25 respectively, and a(i) indicates the number of times gene i is an attracting node in the previous iteration process. Through the inflation function, the overlapping Markov clustering realizes that different functional clusters can be obtained in each iteration, and a single gene can be shared by multiple functional clusters, which conforms to the characteristic that genes perform different functions under different conditions in biology.
[0107] In this embodiment, a comprehensive construction module weight evaluation function is used, and greedy search is used to identify the driver modules on the functional clusters, specifically including:
[0108] Define the functional association weight of the module according to the functional association network; introduce differential expression analysis to define the differential expression weight of the module; introduce gene mutation data to define the mutation weight of the module. Combine the above three weights to construct a module weight evaluation function; construct a greedy algorithm to identify the driver modules on the functional clusters relying on the weight evaluation function.
[0109] First, according to the constructed gene function association network, define the function association weight of the gene subset (candidate module) in the network. The larger the value, the more likely the genes within the candidate module are to perform the same function.
[0110] Secondly, the present invention introduces differential expression analysis to define the differential expression weight of the module. The present invention first uses the fold change method to define the differential expression weight of individual genes, and then uses the sum of the differential expression weights of the corresponding gene subsets to define the differential expression weight of the candidate module.
[0111] Secondly, in order to effectively utilize the high coverage and high mutual exclusivity of the driver module, in this embodiment, the mutation data corresponding to the target cancer is used to calculate and obtain the mutation weight corresponding to the candidate module.
[0112] Secondly, the present invention constructs a comprehensive weight evaluation function of the module by integrating the above three weights;
[0113] Finally, construct a greedy algorithm to identify the driver module based on the weight evaluation function on the functional clusters. First, find the gene pair with the largest module weight on each functional cluster as the seed set (module); then, based on the greedy search strategy, cyclically select the optimal gene nodes to join the set until there are no genes that can be added within the functional cluster, and take the finally obtained gene set as the candidate module corresponding to the cluster; after all clusters have been searched, according to the comprehensive weight of each candidate module, determine the candidate modules with the top rankings as the driver modules.
[0114] Exemplarily, comprehensively construct the module weight evaluation function, and use greedy search to identify the driver module on the functional clusters, specifically including:
[0115] S1041, define the function association weight of the module according to the function association network, and the specific steps are as follows:
[0116] According to the constructed gene function association network, define the function association weight of the gene subset (candidate module) in the network:
[0117]
[0118] Where represents the candidate module under investigation, that is, the candidate gene subset, F ij *M ij represents the association strength between genes i and j, is the sum of the edge weights between all mutually exclusive gene pairs in the candidate module. The larger the value indicates the genes within are more likely to perform the same function, and the higher the possibility of being a driver module. It can help to find subsets of genes with closely related functions in the network.
[0119] S1042. Introduce the differential expression weight of the differential expression analysis definition module. The specific steps are as follows:
[0120] If a gene is significantly highly expressed or lowly expressed in tumor cells compared to normal cells, then it may play a promoting role in cancer development (proto-oncogene) or lose its original cancer-suppressing function (tumor suppressor gene). Gene differential expression analysis can identify cancer driver genes by comparing the expression levels of the same gene under different conditions in a way that can reflect biological significance.
[0121] To utilize this property, first, the present invention introduces the fold change method to define the differential expression weight for each gene:
[0122]
[0123] and respectively represent the gene expression data matrices in Counts format corresponding to tumor and normal cells, and d represents the number of all tumor cells. and respectively represent the average expression levels of gene i in tumor and normal samples.
[0124] Secondly, the differential expression weight of the candidate module is defined by the sum of the differential expression weights of the corresponding gene subset:
[0125]
[0126] S1043. Introduce the mutation weight of the gene mutation data definition module. The specific steps are as follows:
[0127] Since and are both positively correlated with the number of genes in the module, in order to limit the number of genes in the module and further utilize the high coverage and high mutual exclusivity of the driver module, this embodiment uses the mutation data Mu corresponding to the target cancer to calculate and obtain the mutation weight corresponding to the candidate module:
[0128]
[0129] where represents the set of all samples in which the investigated gene set has mutations on Mu. describes 's coverage, describes the overlapping (i.e., non-mutually exclusive) part in. Obviously, as the number of genes in the module increases, does not always maintain growth.
[0130] S1044. Construct a module weight evaluation function by combining the above three weights. The specific steps are as follows:
[0131] Combining the above three equations, the weight evaluation function of the candidate module can be defined as:
[0132]
[0133] where λ and γ are weight adjustment parameters, which directly affect the proportions of the differential expression weight and the mutation weight in respectively, and jointly act to indirectly adjust the functional association weight. and respectively emphasize the functional commonalities existing among the investigated sets at the cellular level and the expression differences under different conditions among multiple subtypes. then reflects the coverage and mutual exclusivity on the mutation data.
[0134] S1045. Construct a greedy algorithm to identify driver modules based on the weight evaluation function on the functional clusters. The specific steps are as follows:
[0135] First, find the gene pairs with the maximum module weight on each functional cluster as the seed set (module);
[0136] Secondly, based on the greedy search strategy, judge all the out-of-set nodes in the cluster. If the addition of a node makes the new module have the maximum weight and is greater than the weight of the original module, then add this node as the newly identified candidate gene to the gene set;
[0137] Repeat the above step until there are no genes that can be added in the functional cluster, and take the finally obtained gene set as the candidate module corresponding to this cluster;
[0138] After searching all the clusters, according to the comprehensive weights of each candidate module, determine the candidate modules with the top rankings as the driver modules.
[0139] In this embodiment, define the distance function between driver modules and determine the driver modules as co-driver modules, specifically including:
[0140] Define the functional associations between different modules according to the functional association network; define the mutation co-occurrence between different modules according to the mutation data; define the distance function between modules to determine the co-driver modules.
[0141] The driving module obtained in the previous step corresponds to a single gene subset in the gene set under investigation that has a potential association with cancer development. To further explore the synergistic mechanism among different modules associated with cancer, the present invention will construct a distance function between modules based on the functional association among modules and the co-occurrence of different modules in mutation data, and further identify the set of driving modules with synergistic effects.
[0142] First, the present invention defines the functional association between driving modules based on the average Hausdorff distance according to the functional association network. The greater the functional association between modules indicates that there are more functional interactions between modules, and it is possible that they have the same / similar functions or jointly perform a certain function in life activities, and they are more likely to be synergistic driving modules.
[0143] Secondly, the present invention defines the mutation co-occurrence of synergistic driving modules in mutation data based on the ratio of the number of shared samples of modules to the number of samples of the smaller module in the module pair under investigation. The larger the value, the more likely the modules belong to the same cancer-related process and are more likely to be synergistic driving modules.
[0144] Then, based on the functional association and mutation co-occurrence, the present invention defines the synergistic distance between modules. The closer the distance, the more obvious the synergistic driving between modules. Finally, based on this distance, the driving modules are divided into synergistic driving modules.
[0145] Exemplarily, a distance function between driving modules is defined, and the driving modules are determined as synergistic driving modules, which specifically includes:
[0146] S1051, defining the functional association between different modules according to the functional association network. The specific steps are as follows:
[0147] The classical Hausdorff distance is a distance metric for sets.
[0148] It describes the maximum distance from one set to the nearest point in another set. However, this distance is vulnerable to noise interference and may not reflect the distance between sets. Therefore, in this embodiment, the average Hausdorff distance is selected to define the functional association between modules. Specifically, in this embodiment, the minimum value of the average of all the maximum values of all gene pairs in the functional association network between two driving modules will be selected as the functional association between the two modules. The functional relationship between driving modules is defined based on the average Hausdorff distance as follows:
[0149]
[0150]
[0151]
[0152] Among them, and represent two driver modules (gene sets), and represent obtaining the average functional association from module to module module to module respectively. represents the average functional association after eliminating asymmetry. Obviously, the larger the value of and indicates that there are more functional interactions between
[0153] S1052. Define the mutation co-occurrence between different modules based on mutation data. The specific steps are as follows:
[0154] The research finds that there are a large number of shared samples of gene mutations in the co-driver modules in the mutation data, that is, there is mutation co-occurrence between the co-driver modules. In this embodiment, the co-occurrence between modules is defined as the ratio of the number of shared samples of modules in the mutation data Mu to the number of samples of the smaller module in the investigated module pair:
[0155]
[0156] reveals the co-occurrence degree between module and module The value of which directly affects the determination of the cooperation between and respectively.
[0157] S1053. Define the distance function between modules to determine co-driver modules. The specific steps are as follows:
[0158] First, because both and defined above belong to the weight function and cannot be directly used as a distance metric to measure the cooperation between modules, the present invention combines the functional association network and mutation data to define the distance function depicting the cooperation between modules as follows:
[0159]
[0160] where θ is a parameter variable responsible for regulating the contributions of functional association and mutation co-occurrence between modules.
[0161] Second, combine the distance function that fuses functional association and mutation co-occurrence The present invention can use the K-means method to cluster and divide all driver modules, and the obtained module clusters gather driver modules with the same / similar functions or those that cooperate to complete a certain biological function. The driver modules assigned to the same cluster are cooperative in the process of completing the corresponding biological function, and these driver modules cooperatively affect the occurrence and development of cancer. Therefore, the present invention finally identifies the driver modules assigned to the same cluster as co-driver gene modules.
[0162] In summary, the present invention first integrates copy number variation data and single nucleotide polymorphism data into mutation data, performs gene alignment and preliminary screening on multi-omics data; then constructs a functional association network based on single-cell gene expression data and gene interaction networks; secondly, applies overlapping Markov clustering to the functional association network to obtain multiple functional clusters; then comprehensively constructs a module weight evaluation function, and uses greedy search to identify driver modules on the functional clusters; finally, defines a distance function between driver modules based on functional association and mutation co-occurrence, and divides the driver modules into co-driver modules.
[0163] The present invention first integrates copy number variation data and single nucleotide polymorphism data into comprehensive mutation data to avoid data missing and untrustworthy problems that may exist in single mutation data. Then, gene alignment and preliminary screening are performed on multi-omics data including single-cell sequencing data to reduce the data volume, obtain available multi-omics data, and improve the operating efficiency of the system. Then, based on single-cell gene expression data, a cell-level specific gene co-expression network is constructed for each subtype and normal cells respectively, and these networks are combined into an expression association network that can reflect the commonalities between different subtypes. This network represents gene co-expression relationships that are prevalent in different subtypes and significantly different from normal cells, and can be used to identify cancer-related genes. Then, the present invention introduces a gene function interaction network, and fuses and optimizes it with the original network into a functional association network to strengthen the functional connection between genes, retain more important gene co-expression relationships, and reduce the network complexity, further improving the operating speed of the system. Subsequently, the present invention applies overlapping Markov clustering to this network to obtain multiple functional clusters, and this functional cluster allows gene duplication, which conforms to the general law of life activities. To effectively and fully mine gene expression data and utilize the high coverage and high mutual exclusivity of driver modules, the present invention respectively introduces differential expression analysis and gene mutation data, comprehensively constructs a module weight evaluation function, and then uses greedy search to identify driver modules on the functional clusters in combination with the evaluation function, and selects the candidate module with the highest score as the driver module. Finally, based on the functional connection and mutation co-occurrence between modules, the present invention comprehensively defines a distance function between modules, and through this distance, the clustering method can be effectively used to determine the driver modules as co-driver modules, thus realizing the accurate identification of cancer co-driver modules.
[0164] Example 2:
[0165] Embodiment 2 of the present invention provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, the following steps are implemented:
[0166] Integrate copy number variation data and single nucleotide polymorphism data into mutation data, perform gene alignment and preliminary screening on multi-omics data to obtain available multi-omics data, and the multi-omics data at least includes single-cell gene expression data;
[0167] Construct a functional association network based on single-cell gene expression data and gene interaction network;
[0168] Apply overlapping Markov clustering to the functional association network to obtain multiple functional clusters;
[0169] Based on the module weight evaluation function, use greedy search to identify driver modules on the functional clusters;
[0170] According to the distance function between driver modules, determine the identified driver modules as co-driver modules.
[0171] The detailed steps are the same as those provided in Embodiment 1 and will not be elaborated here.
[0172] Embodiment 3:
[0173] Embodiment 3 of the present invention provides an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, the following steps are implemented:
[0174] Integrate copy number variation data and single nucleotide polymorphism data into mutation data, perform gene alignment and preliminary screening on multi-omics data to obtain available multi-omics data, and the multi-omics data at least includes single-cell gene expression data;
[0175] Construct a functional association network based on single-cell gene expression data and gene interaction network;
[0176] Apply overlapping Markov clustering to the functional association network to obtain multiple functional clusters;
[0177] Based on the module weight evaluation function, use greedy search to identify driver modules on the functional clusters;
[0178] According to the distance function between driver modules, determine the identified driver modules as co-driver modules.
[0179] The detailed steps are the same as those provided in Embodiment 1 and will not be elaborated here.
[0180] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) that contain computer-usable program code.
[0181] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0182] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0183] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0184] Those of ordinary skill in the art can understand that all or part of the processes in the above-described embodiment methods can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-described method embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0185] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A cancer co-driver module identification system based on single-cell data, characterized in that: It includes: A data preprocessing module, configured to: integrate copy number variation data and single nucleotide polymorphism data into mutation data, perform gene alignment and preliminary screening on multi-omics data to obtain available multi-omics data, and the multi-omics data at least includes single-cell gene expression data; A functional association network construction module, configured to: construct a functional association network based on single-cell gene expression data and a gene interaction network; A functional cluster acquisition module, configured to: apply overlapping Markov clustering to the functional association network to obtain multiple functional clusters; A driver module identification module, configured to: identify driver modules on the functional clusters using greedy search based on a module weight evaluation function; According to the constructed functional association network, define the functional association weight of candidate modules in the network, and the larger its value indicates that genes in the candidate module are more likely to perform the same function; Introduce differential expression analysis to define the differential expression weight of candidate modules, use the fold change method to define the differential expression weight of individual genes, and use the sum of the differential expression weights of the corresponding gene subsets to define the differential expression weight of candidate modules; Use the mutation data corresponding to the target cancer to calculate and obtain the mutation weight corresponding to the candidate module; Construct a module weight evaluation function by integrating the functional association weight of the candidate module, the differential expression weight of the candidate module, and the mutation weight corresponding to the candidate module; Adopt a greedy algorithm to identify driver modules on the functional clusters relying on the module weight evaluation function; A co-driver module identification module, configured to: determine the identified driver modules as co-driver modules according to the distance function between driver modules.
2. The cancer co-driver module identification system based on single-cell data according to claim 1, characterized in that: The copy number variation data is: the copy number variation data of patients, the data is multi-dimensional, each dimension corresponds to a gene locus of a patient, and when there is a copy number duplication or deletion, the gene data is 1, and when the copy number is normal, it is 0; The single nucleotide polymorphism data is: the single nucleotide polymorphism data of patients, the data is multi-dimensional, each dimension corresponds to a gene locus of a patient, and when there is a single nucleotide polymorphism, the gene data is 1, and when there is no single nucleotide polymorphism, it is 0; The multi-omics data includes: copy number variation data and single nucleotide polymorphism data, and also includes: single-cell gene expression data divided into different cancer subtypes, the data is multi-dimensional, each dimension corresponds to a gene locus of a patient in a different subtype, and its value represents the expression level of the gene; a gene interaction network, the data is multi-dimensional, each dimension corresponds to a gene, and its value describes the interaction strength of genes participating in signal transduction, energy and material metabolism, and cell cycle regulation together.
3. The cancer co-driver module identification system based on single-cell data according to claim 1, characterized in that: Integrating copy number variation data and single nucleotide polymorphism data into mutation data, and performing gene alignment and preliminary screening on multi-omics data includes: Obtain copy number variation data and single nucleotide polymorphism data, and integrate the copy number variation data and single nucleotide polymorphism data according to the mutation situation of genes on a batch of samples to construct mutation data. When a gene has a copy number variation or a single nucleotide polymorphism occurs, the value of the corresponding mutation data is 1; Based on the obtained mutation data, select a preset number of genes with the top comprehensive mutation frequencies based on the mutation frequencies of the genes and the actual situation data of the patients.
4. The cancer co-driving module recognition system based on single-cell data according to claim 1, characterized in that: Construct a functional association network based on single-cell gene expression data and a gene interaction network, including: Construct cell-level specific gene co-expression networks for each subtype and normal cells respectively based on single-cell gene expression data, combine these networks into an expression association network that can reflect the commonalities between different subtypes, introduce a gene interaction network, and fuse and optimize it with the expression association network into a functional association network.
5. The cancer co-driving module recognition system based on single-cell data according to claim 1, characterized in that: Apply overlapping Markov clustering to the functional association network to obtain multiple functional clusters, including: Initialize the attracting node indicator vector, which is used to represent the number of times a certain node is marked as an attracting node in multiple Markov clusterings, and is used to penalize attracting nodes: Run the Markov clustering algorithm cyclically on the functional association network. Each time when the Markov matrix converges, collect the functional clusters generated in this round of clustering, and mark the attracting nodes obtained in this round of clustering; Among them, in each Markov clustering process, the nodes are penalized through an inflation function according to the marked attracting node indicator vector.
6. The cancer co-driving module recognition system based on single-cell data according to claim 1, characterized in that: Use a greedy algorithm to identify driving modules on the functional clusters relying on a module weight evaluation function, including: Find gene pairs with the largest module weights on each functional cluster as the seed set; cyclically select the optimal gene nodes to join the set based on the greedy search strategy until there are no genes that can be added in the functional cluster, and take the finally obtained gene set as the candidate module corresponding to the functional cluster; after searching all functional clusters, based on the comprehensive weights of each candidate module, determine the preset number of candidate modules with the top rankings as driving modules.
7. The cancer co-driving module recognition system based on single-cell data according to claim 1, characterized in that: Determine the identified driving modules as co-driving modules according to the distance function between driving modules, including: Define the functional association between driving modules based on the average Hausdorff distance according to the functional association network; Define the mutation co-occurrence of co-driving modules in the mutation data based on the ratio of the number of shared samples of the module in the mutation data to the number of samples of the smaller module in the inspected module pair; Based on the functional association and mutation co-occurrence, define the co-driving distance between modules. The smaller the distance, the more obvious the co-driving property between modules. Divide the driving modules into co-driving modules based on the co-driving distance.
8. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by a processor, it implements the following steps of the cancer co-driver module identification system based on single-cell data as described in any one of claims 1-7: Integrate copy number variation data and single nucleotide polymorphism data into mutation data, perform gene alignment and preliminary screening on multi-omics data to obtain available multi-omics data, and the multi-omics data includes at least single-cell gene expression data; Construct a functional association network based on single-cell gene expression data and gene interaction networks; Apply overlapping Markov clustering to the functional association network to obtain multiple functional clusters; Based on the module weight evaluation function, use greedy search to identify driver modules on the functional clusters; According to the distance function between driver modules, determine the identified driver modules as co-driver modules.
9. An electronic device, comprising a memory, a processor, and a program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the following steps of the cancer co-driver module identification system based on single-cell data as described in any one of claims 1-7: Integrate copy number variation data and single nucleotide polymorphism data into mutation data, perform gene alignment and preliminary screening on multi-omics data to obtain available multi-omics data, and the multi-omics data includes at least single-cell gene expression data; Construct a functional association network based on single-cell gene expression data and gene interaction networks; Apply overlapping Markov clustering to the functional association network to obtain multiple functional clusters; Based on the module weight evaluation function, use greedy search to identify driver modules on the functional clusters; According to the distance function between driver modules, determine the identified driver modules as co-driver modules.