System for identifying cancer co-driving pathways based on attribute heterogeneous network embedding
By integrating multi-omics data and using attribute-based heterogeneous network embedding, pathway interaction networks are reconstructed, addressing the problem of neglecting pathway-level biological information in existing technologies and enabling accurate identification and efficient analysis of cancer synergistic driving pathways.
Patent Information
- Application Number
- CN202210726322.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-24
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2042-06-24
AI Technical Summary
Existing methods for identifying cancer co-driving pathways often ignore biological and genetic information at the pathway level, leading to inaccurate identification. Furthermore, relying heavily on single mutation data may result in data gaps and unreliability issues.
By integrating copy number variation data and single nucleotide polymorphism data, and based on the attribute heterogeneous network embedding method, we construct and optimize gene pathway relationships, reconstruct pathway interaction networks using multi-omics data, and define synergistic driving pathways by combining driving pathways and synergistic driving capabilities.
This approach enables full utilization of gene-level information at the pathway level, improving the accuracy and efficiency of cancer co-driving pathway identification, reducing interference from irrelevant genes, and providing higher interpretability and accuracy.
Smart Images

Figure CN115910203B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence data mining classification and bioinformatics, and particularly relates to a cancer synergistic driving pathway identification system based on attribute heterogeneous network embedding. BACKGROUND
[0002] The statements in this section merely provide background information related to the present application and do not necessarily constitute the prior art.
[0003] In recent years, with the development of high-throughput biological technology, large-scale cancer genome projects such as the Cancer Genome Atlas (TCGA), the International Cancer Genome Consortium (ICGC) and the like have generated and accumulated rich high-throughput multi-omics cancer data, which provides support for researchers to in-depth analyze cancer mechanism from the system level. However, it is difficult to accurately identify driving mutations and driving genes related to a specific cancer category from large-scale biological omics data only by relying on biological experiments or simple statistical analysis methods. Therefore, for large-scale disease genetic data, it is a major challenge for current cancer informatics research to develop effective computing methods to accurately and efficiently identify cancer driving mutation / gene sets. Accurate identification of oncogenic genetic factors has important theoretical and application value for cancer diagnosis, targeted drug development, and precise and personalized treatment of cancer patients, and the like. Early studies were usually committed to finding genes with obvious high-frequency mutations as driving genes, but as the research deepens, it is found that genes cooperate within the pathway to drive the occurrence of cancer, and there is also a synergistic driving effect between pathways, so the current work mainly focuses on finding cancer synergistic driving pathways at the pathway level.
[0004] Current methods for identifying cancer synergistic driver pathways mainly include: identifying functional modules that meet the mutual exclusion rule in the gene interaction network as driver pathways; using heat diffusion algorithm to reconstruct the gene interaction network, and then detecting the subnetwork with the best coverage and mutual exclusion from the reconstructed network; by establishing a weight function on somatic mutation data, combining greedy and MCMC methods, and optimizing the identification of the gene set with the highest weight as the driver pathway; combining the characteristics that genes with similar expression usually perform certain biological functions together, introducing binary linear programming (BLP) and genetic algorithm, and more fully introducing genetic information to guide the identification of driver pathways; using integer linear programming to detect multiple functional modules with high weight as synergistic driver pathways; combining high co-occurrence of mutations between modules with high coverage and high mutual exclusion of mutations within modules, using BLP to identify synergistically acting driver pathways; using greedy search on gene signal network to explore mutually exclusive modules with common downstream events, and designing a double-regular double-clustering method to distinguish mutually exclusive modules in the same cluster as synergistic driver pathways; first performing hierarchical clustering on genes to obtain synergistic mutation modules, and then mapping the relationship between mutation modules to the pathway level through link prediction to supplement the potential relationship between pathways, and finally identifying synergistic driver pathways based on modules and updated pathway interaction network.
[0005] The inventors found that most of the existing methods usually use high coverage and high exclusivity at the gene level to identify synergistic driver gene sets, and directly map these sets to the pathway level as synergistic driver pathways, ignoring the verified biological pathways; in addition, some methods based on verified pathways often simply regard them as known gene sets, ignoring the genetic information contained in the pathways, which makes it difficult to fully utilize the rich information at the gene level on the pathway level; at the same time, the above identification methods are mainly based on somatic mutation data and gene expression data, so they may ignore other key molecules in the carcinogenic process. SUMMARY
[0006] In order to solve the problems in the prior art, the application provides a cancer synergistically driven pathway identification system based on attribute heterogeneous network embedding, which integrates copy number variation data and single nucleotide polymorphism data into mutation data, and performs data alignment and preliminary screening on multi-omics attribute heterogeneous data to obtain available heterogeneous data; then, in order to fully utilize the gene level information at the pathway level, the application obtains network weights of genes in different pathways according to the internal network structure of the pathways, and optimizes the gene pathway relationship according to the weights; secondly, the application generates an attributed heterogeneous network based on the relationship information between biomolecules and pathways and the respective relationship / attribute information of the biomolecules and the pathways, and optimizes the attributed heterogeneous network based on known cancer-related biomolecules; then, an embedding framework based on joint matrix decomposition is introduced to integrate the heterogeneous data, supplement incomplete inter-pathway relationships, and reconstruct the pathway interaction network; finally, based on the reconstructed pathway interaction network, the synergistically driven ability of the pathways is defined by using the driving ability of a single pathway and the cooperative relationship between the pathways, and the precise identification of the synergistically driven pathways is completed.
[0007] In order to achieve the above object, the application adopts the following technical scheme:
[0008] The first aspect of the application provides a cancer synergistically driven pathway identification system based on attribute heterogeneous network embedding.
[0009] The cancer synergistically driven pathway identification system based on attribute heterogeneous network embedding comprises:
[0010] The data preprocessing module is configured to integrate copy number variation data and single nucleotide polymorphism data into mutation data, and perform data alignment and preliminary screening on multi-omics attribute heterogeneous data to obtain available heterogeneous data.
[0011] The gene weight acquisition module is configured to obtain network weights of genes in different pathways according to the internal network structure of the pathways, and optimize the gene pathway interaction network according to the weights.
[0012] The attribute heterogeneous network embedding module is configured to construct an initial attribute heterogeneous network according to the collected attribute data and relationship data of multiple biomolecules and pathways, optimize the initial attribute heterogeneous network by collecting miRNAs related to the studied cancer and the optimized gene pathway interaction network, integrate the heterogeneous data based on an embedding framework of joint matrix decomposition, embed the optimized attribute heterogeneous network, supplement incomplete inter-pathway relationships, and reconstruct the pathway interaction network.
[0013] The synergistically driven pathway identification module is configured to: obtain mutation data corresponding to each pathway using the optimized gene pathway interaction network;Define the driving ability of the individual driving pathway according to the high coverage and high mutual exclusivity of the driving pathway on the mutation data;Define the synergistic driving ability between pathways according to the reconstructed pathway interaction network and mutation co-occurrence;Define the comprehensive driving weight by combining the driving ability of the pathway and the synergistic driving ability between the pathways, and identify the synergistically driven pathways on the reconstructed pathway interaction network.
[0014] The second aspect of the present application provides a computer readable storage medium, which stores a program, and the program is executed by a processor to realize the following steps:
[0015] Integrate the copy number variation data and the single nucleotide polymorphism data into mutation data, and perform data alignment and preliminary screening on the multi-omics attribute heterogeneous data to obtain available heterogeneous data;
[0016] Obtain the network weight of genes in different pathways according to the internal network structure of the pathway, and optimize the gene pathway interaction network according to the weight;
[0017] Construct an initial attribute heterogeneous network according to the collected attribute data and relationship data of multiple biological molecules and pathways, optimize the initial attribute heterogeneous network based on the collected miRNA related to the studied cancer and the optimized gene pathway interaction network, embed the optimized attribute heterogeneous network based on the embedding framework of joint matrix decomposition, supplement the incomplete relationship between pathways, and reconstruct the pathway interaction network;
[0018] Obtain the mutation data corresponding to each pathway using the optimized gene pathway interaction network;Define the driving ability of the individual driving pathway according to the high coverage and high mutual exclusivity of the driving pathway on the mutation data;Define the synergistic driving ability between pathways according to the reconstructed pathway interaction network and mutation co-occurrence;Define the comprehensive driving weight by combining the driving ability of the pathway and the synergistic driving ability between the pathways, and identify the synergistically driven pathways on the reconstructed pathway interaction network.
[0019] The third aspect of the present application provides an electronic device, which comprises a memory, a processor and a program stored in the memory and executable on the processor, and the processor realizes the following steps when executing the program:
[0020] Integrate the copy number variation data and the single nucleotide polymorphism data into mutation data, and perform data alignment and preliminary screening on the multi-omics attribute heterogeneous data to obtain available heterogeneous data;
[0021] Obtain the network weight of genes in different pathways according to the internal network structure of the pathway, and optimize the gene pathway interaction network according to the weight;
[0022] According to the collected attribute data and relationship data corresponding to a plurality of biological molecules and pathways, an initial attribute heterogeneous network is constructed, the initial attribute heterogeneous network is optimized through the collected miRNA related to the studied cancer and the optimized gene pathway interaction network, the optimized attribute heterogeneous network is embedded based on a joint matrix decomposition embedding framework integrating heterogeneous data, the incomplete inter-pathway relationship is supplemented, and the pathway interaction network is reconstructed;
[0023] Mutational data corresponding to each pathway is obtained using the optimized gene pathway interaction network, the driving ability of a single driving pathway is defined according to the high coverage and high exclusivity of the driving pathway on the mutational data, the collaborative driving ability between pathways is defined according to the reconstructed pathway interaction network and the mutational co-occurrence, and the comprehensive driving weight is defined by combining the driving ability of the pathway and the collaborative driving ability between the pathways, so that the collaborative driving pathways are identified on the reconstructed pathway interaction network.
[0024] Compared with the prior art, the present application has the beneficial effects that:
[0025] 1、The present application first integrates copy number variation data and single nucleotide polymorphism data into comprehensive mutational data, avoids the possible data loss and untrustworthy problems of single mutational data, performs data alignment and preliminary screening on multi-omics attribute heterogeneous data including a plurality of biological molecules (genes, miRNA, lncRNA) and pathways, screens out low-quality data, reduces the data volume and improves the running efficiency of the system, obtains the network weight of genes in different pathways according to the internal network structure of the pathway, and optimizes the gene pathway relationship according to the weight, generates an attribute heterogeneous network based on the relationship information between biological molecules and pathways and the respective relationship / attribute information, optimizes the attribute heterogeneous network based on the known cancer-related biological molecules, then introduces an embedding framework based on joint matrix decomposition to integrate heterogeneous data, supplement incomplete inter-pathway relationships, and reconstruct the pathway interaction network, define the driving ability of a single driving pathway according to the high coverage and high exclusivity of the driving pathway on the mutational data, define the collaborative driving ability between pathways according to the reconstructed pathway interaction network and the mutational co-occurrence, integrate and define the comprehensive driving weight, and identify the collaborative driving pathways on the reconstructed pathway interaction network, so as to realize accurate identification of cancer collaborative driving pathways.
[0026] 2、The application optimizes the relationship between gene pathways based on gene weight based on pathway structure, and the related information of important genes can contribute more to the identification of driving pathways, thereby reducing the interference degree of irrelevant genes on the identification result; The application uses attribute heterogeneous network embedding to provide rich genetic information related to other biological molecules for pathway level analysis, which makes up for the insufficient / missing problem of pathway level related biological information; The application defines the driving ability of single pathway and the cooperation relationship between pathways respectively, and the defined cooperative driving ability between pathways comprehensively considers the high coverage, high mutual exclusivity, mutation co-occurrence and functional association of the cooperative driving pathway, has high interpretability, and can efficiently and accurately determine the pathway as a cooperative driving pathway.
[0027] Advantages of the additional aspects of the application will be partially given in the following description, partially will become apparent from the following description, or will be understood by the practice of the application. BRIEF DESCRIPTION OF DRAWINGS
[0028] The accompanying drawings, which form a part of the specification, are included to provide a further understanding of the application and are incorporated herein by reference. The illustrations are shown for the purpose of enabling those skilled in the art to implement the application and are not intended to limit the scope of the application.
[0029] Figure 1 The structure diagram of the cancer cooperative driving pathway identification system based on attribute heterogeneous network embedding provided for embodiment 1 of the application. DETAILED DESCRIPTION
[0030] The application will be further described below in conjunction with the drawings and examples.
[0031] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the application. Unless otherwise indicated, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the application pertains.
[0032] It should be noted that the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the exemplary embodiments according to the application. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form, and in addition, it should be understood that when the terms "comprise" and / or "include" are used in the specification, they mean the presence of a feature, step, operation, device, component and / or combination thereof.
[0033] The embodiments in the application and the features in the embodiments can be combined with each other without conflict.
[0034] Embodiment 1:
[0035] As Figure 1As shown, the embodiment 1 of the present application provides a cancer synergistically driven pathway identification system based on attribute heterogeneous network embedding, comprising:
[0036] The data preprocessing module is configured to integrate the copy number variation data and the single nucleotide polymorphism data into mutation data, and perform data alignment and preliminary screening on the multi-omics attribute heterogeneous data to obtain available heterogeneous data.
[0037] The gene weight acquisition module is configured to obtain network weights of genes in different pathways according to the internal network structure of the pathways, and optimize the gene-pathway interaction network according to the weights.
[0038] The attribute heterogeneous network embedding module is configured to construct an initial attribute heterogeneous network according to the collected attribute data and relationship data corresponding to multiple biological molecules and pathways, optimize the initial attribute heterogeneous network by collecting miRNA related to the studied cancer and the optimized gene-pathway interaction network, embed the optimized attribute heterogeneous network based on a joint matrix decomposition embedding framework integrating heterogeneous data, supplement incomplete inter-pathway relationships, and reconstruct the pathway interaction network.
[0039] The synergistically driven pathway identification module is configured to obtain mutation data corresponding to each pathway using the optimized gene-pathway interaction network, define the driving ability of an individual driving pathway according to the high coverage and high exclusivity of the driving pathway on the mutation data, define the synergistically driven ability between pathways according to the reconstructed pathway interaction network and mutation co-occurrence, define the comprehensive driving weight by combining the driving ability of the pathways and the synergistically driven ability between the pathways, and identify synergistically driven pathways on the reconstructed pathway interaction network.
[0040] In this embodiment, the copy number variation data refers to the copy number variation data of a patient, which has more than 10,000 dimensions, each dimension corresponding to a gene site of the patient, and the data of the gene is 1 when copy number duplication or deletion occurs, and 0 when the copy number is normal.
[0041] In this embodiment, the single nucleotide polymorphism data refers to the single nucleotide polymorphism data of a patient, which has more than 10,000 dimensions, each dimension corresponding to a gene site of the patient, and the data of the gene is 1 when single nucleotide polymorphism occurs, and 0 when single nucleotide polymorphism does not occur.
[0042] In this embodiment, the multi-omics heterogeneous data refers to: in addition to the above two kinds of data, it also includes: gene expression (methylation) data, the data is more than 10000 dimensions, each dimension corresponds to a gene site of different patients, and the value represents the expression level (methylation level) of the gene. Gene interaction network, the data is more than 10000 dimensions, each dimension corresponds to a gene, and the value describes the interaction strength of genes participating in signal transmission, energy and material metabolism, and cell cycle regulation and other life processes together; lncRNA (miRNA) expression data, the data is more than 500 (600) dimensions, each dimension corresponds to a lncRNA (miRNA) site of different organs or tissues, and the value represents the average expression level of lncRNA (miRNA) in the organ or tissue; in addition, the multi-omics heterogeneous data also includes the interaction network between these biomolecules and the interaction network between the pathways, and the interaction network between the pathways (more than 200 dimensions), and the gene interaction network. These interaction networks describe the functional association strength between different nodes.
[0043] First, the copy number variation data and the single nucleotide polymorphism data are acquired, and the copy number variation data and the single nucleotide polymorphism data are integrated to construct mutation data according to the mutation of the gene on the batch sample. When the gene has copy number variation or single nucleotide polymorphism, the value of the corresponding mutation data is one.
[0044] Secondly, the multi-omics heterogeneous data is preliminarily screened and data aligned. Based on the previously obtained mutation data, the mutation frequency of the gene and the actual situation of the patient are comprehensively considered, the genes with high comprehensive mutation frequency are selected to participate in the subsequent experiment, and then the biomolecule (pathway) nodes with missing or incomplete data are screened out, and the nodes with available values in the related data are retained.
[0045] Exemplarily, the integrated copy number variation data and single nucleotide polymorphism data are mutation data, and gene alignment and preliminary screening are performed on multi-omics data, which specifically includes:
[0046] S1011, the collected copy number variation data and single nucleotide polymorphism data are integrated into mutation data, and the specific steps are as follows:
[0047] The copy number variation data (CNV) and the single nucleotide polymorphism (SNP) data are acquired, and the copy number variation data and the single nucleotide polymorphism data are integrated to construct mutation data according to the mutation of the gene on the batch sample. For any gene i and any sample p, the value is specifically defined as follows:
[0048]
[0049] Wherein, For mutation data, m represents the total number of genes under study, b represents the total number of batches of patients under study, i, p correspond to gene i and patient p, and when the values of SNP and CNV in the above formula are greater than zero, it indicates that single nucleotide polymorphism or copy number variation occurs in the corresponding gene.
[0050] S1012, gene alignment and preliminary screening are performed on the multi-omics data according to the obtained mutation data, and the specific steps are as follows:
[0051] Firstly, since the genes with low mutation frequency are mostly passenger genes without influence on cancer, in order to reduce the calculation cost and make the results more interpretable, only the genes with a mutation frequency higher than a certain threshold are retained, and the comprehensive mutation frequency of each gene is defined as follows:
[0052]
[0053] Wherein, p represents a sample (patient) in the mutation data, |p| represents the total number of genes in which sample p has mutations, and Γ(i) represents a sample set in which gene i has mutations. The above formula comprehensively considers the two dimensions of genes and samples, and reduces the influence of adverse data and ill-conditioned samples.
[0054] Secondly, according to the obtained comprehensive mutation frequency, the multi-omics heterogeneous data is screened, and the genes that exist in various data and have a high comprehensive mutation frequency are retained.
[0055] Finally, the biological molecule (pathway) nodes with missing or incomplete data are screened out, and the nodes with available values in the related data are retained.
[0056] Further, the gene weight is obtained, and the gene pathway relationship is optimized, specifically including:
[0057] According to the internal network structure of the pathway, the network weight of the gene on the corresponding pathway is obtained;
[0058] The gene pathway relationship is optimized by the network weight, and the gene pathway interaction network is optimized.
[0059] Firstly, in view of the problems such as lack of utilization of pathway related information in the existing method, the network weight of the gene on the pathway is defined by using the topological order and degree of the gene in the directed graph structure inside the pathway, and the gene with higher topological order and degree value has greater network weight.
[0060] Secondly, the gene pathway relationship is optimized according to the network weight, and the weighted gene pathway interaction network is standardized by the min-max normalization method. Therefore, in the subsequent steps, important genes can have greater influence on the analysis at the pathway level, and data at the gene level can be more reasonably used.
[0061] Exemplarily, the gene weight is obtained, and the gene pathway relationship is optimized, specifically including:
[0062] S1021, obtaining a network weight of a gene on a corresponding pathway according to a network structure inside the pathway, and the specific steps are as follows:
[0063] Most methods usually directly take the obtained synergistically driven gene set as the synergistically driven pathway without fully utilizing the relevant information of the verified pathway. Some other methods based on the verified pathway also have the following defects: due to the limited information at the pathway level, researchers often use other biological molecular data as a supplement to explore the genetic cooperation between pathways. For example, many methods map the rich data at the gene level to the pathway level through a gene pathway interaction network to supplement the information. However, the gene pathway relationship used by the existing methods is usually represented by a zero-one indicator matrix, and for a gene g and a pathway p, the corresponding entry is defined as follows:
[0064]
[0065] wherein is an initial gene-pathway relationship matrix, which represents the initial gene-pathway interaction network, n g and n p represent the total number of genes and pathways respectively. Although GP ini preserves the connection between genes and pathways, it implies that each gene in the pathway is equally important and plays the same role, which is inconsistent with the biological fact. For example, in the life activities related to the P53 signaling pathway, the TP53 gene obviously makes a greater contribution than other genes. This defect may cause the research at the pathway level to be disturbed by unimportant genes, making it difficult to reasonably utilize the information at the gene level.
[0066] Therefore, in order to more appropriately utilize the rich genetic information at the gene level to identify the synergistically driven pathway, the present application defines the importance of each gene in each pathway based on the relevant biological knowledge of the pathway through the following standards:
[0067] (1) The gene located at the upstream position of the pathway can interfere with the downstream gene and have a greater impact on the pathway.
[0068] (2) The gene that has interactions with more other genes in the pathway is more important.
[0069] In order to quantify the above standards, the present application introduces the network structure of the genes in the pathway, which is usually ignored by other methods. It is well known that the network structure of the genes in the pathway can be regarded as a directed graph, so the topological order and degree of the genes in the directed graph can be used to measure the two standards. In other words, for a gene g in a pathway p, its network weight is used to represent its importance:
[0070]
[0071] wherein, represents the weighted gene-pathway relationship matrix or the weighted gene-pathway interaction network, Dgree(g, p) and Order(g, p) are the degree and topological order of gene g in pathway p, respectively. Obviously, the higher the degree value and the earlier the topological order, the more consistent the gene is with the above criteria, and the greater the network weight, and thus the greater the contribution to the pathway. In other words, GP nw The relatively unimportant genes can be reduced to interfere with the identification of synergistically driven pathways.
[0072] S1022, optimizing the gene-pathway relationship through the network weight, and optimizing the gene-pathway interaction network. The specific steps are as follows:
[0073] After obtaining the weighted gene-pathway interaction network GP nw , in order to facilitate the execution of subsequent steps, the minimum-maximum normalization is performed on each column (i.e. each pathway) of GP nw , to obtain the final gene-pathway relationship matrix:
[0074] GP = Normalize(GP nw )
[0075] represents the optimized gene-pathway relationship matrix (interaction network), which retains the advantages of GP nw in distinguishing genes of different importance while facilitating calculation, and will be used to represent the optimized relationship between genes and pathways in subsequent steps.
[0076] Further, the attribute heterogeneous network is constructed, and the pathway interaction network is reconstructed, specifically including:
[0077] constructing and optimizing the attribute heterogeneous network according to relevant data;
[0078] embedding the attribute heterogeneous network to reconstruct the pathway interaction network.
[0079] Firstly, the present application constructs an initial attribute heterogeneous network according to the collected attribute data and relationship data of various biological molecules (genes, miRNA, lncRNA) and pathways. Then, the initial attribute heterogeneous network is optimized by collecting the miRNA (lncRNA) related to the studied cancer and the previously obtained optimized gene-pathway interaction network, so that it is related to the studied cancer and can reasonably utilize relevant information according to the importance of genes;
[0080] Secondly, this invention introduces a network embedding framework based on joint matrix factorization, which uses the square of the F norm of the matrix as the loss function to embed the previously obtained attribute heterogeneous network, integrates multi-omics heterogeneous data, thereby obtaining low-dimensional representations of different biological molecules (pathways), supplementing incomplete pathway interaction relationships, and obtaining a reconstructed pathway interaction network based on the latent semantic model.
[0081] Among them, attribute data is the attribute information within the same node type, such as gene expression data, methylation data, etc.; relation data is the interaction relationship between data objects of the same or different node types, such as gene pathway relationship data, pathway interaction network, etc.
[0082] For example, constructing a heterogeneous attribute network and reconstructing the pathway interaction network specifically includes:
[0083] S1031, Construct and optimize the attribute heterogeneous network based on relevant data. The specific steps are as follows:
[0084] First, in this invention, the attribute heterogeneous network includes four node types: three biomolecules (genes, miRNAs, and lncRNAs) and pathways. It consists of a relation matrix and an attribute matrix (instantiated representations of relational data and attribute data, respectively). For an attribute heterogeneous network with node type m, these two types of data matrices are defined as follows:
[0085] Relationship matrix: Describes the relationships between data objects of the same or different node types, showing their interactions. Used... It represents the set of all relation matrices. It is a relation matrix that will have n of type i. i An object and n of type j j These objects are associated. If s∈{1,2,L n, i} and t∈{1,2,L,n j If a known relation exists, then R ij (s,t)>0; otherwise, R ij (s,t)=0. For example, the optimized gene pathway interaction network GP obtained above and the reconstructed target pathway interaction network PPI in this section both belong to the relation matrix.
[0086] Attribute matrix: Describes the attribute information within data objects of the same node type. Xit is used to represent different attribute matrices of the i-th type object, where... t i d is the number of attribute matrices for the i-th object type. it It represents the number of attributes in the t-th attribute matrix. For example, for a gene, multi-omics data, including expression data and methylation data, all belong to its attribute matrix.
[0087] First, different types of nodes are connected according to the relationship matrix, and then the corresponding attribute matrix is added to obtain the initial attribute heterogeneous network. Then, based on the known association between biomolecules (miRNA, lncRNA) and the studied cancer, the association between relevant nodes and pathways is optimized. In other words, only the edges between other nodes and pathways related to cancer are retained, and then the cancer-related attribute heterogeneous network is obtained. Subsequently, according to the possibility of functional cooperative interaction of nodes with strong relationships or similar attributes, it is hoped that by integrating the relationship matrix between nodes in the entire attribute heterogeneous network and the attribute matrix of the nodes themselves, the pathway interaction network PPI can be reconstructed, so that it can fully integrate the rich information related to other biomolecules and better describe the functional interaction relationship between pathways.
[0088] S1032, embedding the attribute heterogeneous network, reconstructing the pathway interaction network, the specific steps are as follows:
[0089] The integration of attribute heterogeneous data in existing research usually has the following methods: (1) directly splicing into a high-dimensional matrix for modeling analysis, which is easy to cause matrix sparsity; (2) converting heterogeneous objects into homogeneous form for integrated modeling analysis, which may lose the genetic information contained in the original data; (3) modeling and analyzing different objects first, and integrating the results, which may lose the association information between different objects. In view of this, the present application constructs an attribute heterogeneous network embedding framework based on mean square error and collaborative depth matrix decomposition, which can realize the direct embedding of attribute matrix and relationship matrix. The specific details of the framework are as follows:
[0090] First, in order to effectively embed and integrate various data in the network and utilize the possible nonlinear relationships between objects, this paper uses a deep neural network to obtain a low-dimensional representation of the objects and their corresponding attributes:
[0091]
[0092]
[0093] wherein, and respectively represent the initialization features of the ith type of object and the tth attribute thereof, and the function is to reduce the noise existing in the network for easy calculation. i1 and W it2 are the first layer weight matrix of the ith type of object and the tth attribute thereof, respectively, W i2 and W it2 correspond to the second layer weight matrix, and so on. i1 , b it1 , b i2 , bit2 and so on are the corresponding biases and so on are the corresponding activation functions, and θ represents the model parameters. Through the above equation, we can get and as the low-dimensional representation of the ith object and its tth attribute, respectively.
[0094] Secondly, according to the defined matrix low-dimensional representation, the application re-models the relationship matrix and the attribute matrix in the attribute heterogeneous network based on the Latent Factor Model (LFM):
[0095]
[0096]
[0097] Here and represent the reconstructed relationship matrix between the ith and jth objects and the reconstructed tth attribute matrix of the ith object, respectively. Then, the application introduces the square of the F norm of the matrix to measure the difference between the initial matrix and the reconstructed matrix, and constructs the attribute heterogeneous network embedding loss function:
[0098]
[0099]
[0100] where λ is a hyperparameter used to control the proportion of the PPI-related loss in the overall loss. Each will be used to adjust the loss of the relationship matrix other than PPI. Similarly, represent the weights assigned to the biological molecule attribute heterogeneous data matrix. Their associated terms ensure that the objective equation can directly decompose the relationship matrix and the attribute matrix, avoiding the information loss caused by homogeneous transformation or homogeneous transformation. These two parameters will be adjusted adaptively with the optimization process. It should be noted that for the relationship matrix, if then Similarly, for the attribute matrix X it if t>max i t i then
[0101] Finally, as described in the foregoing, the application optimizes the objective equation using a deep neural network. After optimization, the application can use the reconstructed pathway interaction network to predict and update PPI, describe the functional cooperation relationship between pathways enhanced by other data. Moreover, since the prior knowledge of other biological molecules is used to enhance the association between the attribute heterogeneous network and the specific cancer when generating the attribute heterogeneous network, which can also be considered as related to a specific cancer, so it can represent the cooperation relationship of pathways in the life function related to the cancer, thereby helping the next step to identify the cancer synergistically driven pathways.
[0102] Further; define the synergistically driven ability between pathways, determine the synergistically driven pathways, specifically including:
[0103] Using the optimized gene pathway interaction network to obtain mutation data corresponding to each pathway;
[0104] According to the high coverage and high exclusivity of the driving pathway on the mutation data, the driving ability of the individual driving pathway is defined;
[0105] According to the reconstructed pathway interaction network and mutation co-occurrence, the synergistically driven ability between pathways is defined;
[0106] Combine the driving ability of the pathway and the synergistically driven ability between pathways to define the comprehensive driving weight, and identify the synergistically driven pathways on the reconstructed pathway interaction network.
[0107] First, the important synergistically driven pathways should have the following characteristics within the pathway and between the pathways:
[0108] High mutation coverage within a single pathway; high mutation exclusivity within a single pathway; high mutation co-occurrence between pathways; strong functional interaction between pathways.
[0109] Among them, the first three characteristics are derived based on mutation data, and the functional interaction strength between pathways can be represented by the reconstructed pathway interaction network obtained in the previous step. Therefore, the focus of this section is to combine mutation data and the optimized gene pathway relationship in the second step to calculate these characteristics, combine them into a comprehensive driving weight, and then identify the pathway pairs with larger driving weight on the reconstructed pathway interaction network as synergistically driven pathways.
[0110] Exemplarily, defining the synergistically driven ability between pathways, determining the synergistically driven pathways, specifically including:
[0111] S1041, using the optimized gene pathway interaction network to obtain mutation data corresponding to each pathway, specifically including:
[0112] The optimized gene pathway interaction network GP obtained previously can make genes contribute differently to pathway-level analysis according to their importance in the pathway, so as to use the information at the gene level more reasonably. In order to take advantage of this advantage of GP when calculating the mutation data correlation characteristics, the application first constructs a mutation data matrix specific to each pathway p:
[0113] Mu p = diag(GP .p )*Mu
[0114] Wherein is the initial gene mutation matrix (the instantiated representation of the mutation data), n g and n c represent the total number of genes and cases (patients). If the gene g is mutated on the case c, the relevant entry on Mu will be set to 1, otherwise it will be set to 0. GP .p is the column in GP corresponding to the pathway p, i.e. the optimized interaction between p and its corresponding genes. The result of the multiplication represents the mutation matrix corresponding to the pathway p, which will be used to calculate various characteristics of the mutation data correlation of the pathway p.
[0115] S1042, according to the high coverage and high exclusivity of the driving pathway on the mutation data, the driving ability of the individual driving pathway is defined, and the specific steps are as follows:
[0116] First, for a single pathway, the application defines its driving ability according to its high exclusivity and high coverage, and according to the weighted mutation matrix corresponding to the pathway calculated above, the following formula is defined to measure the two mutation characteristics of the pathway based on different gene weights:
[0117]
[0118] Γ(p) = sum(max(MU p ))
[0119]
[0120] Wherein max() represents taking the maximum value of the matrix by column, so Γ(p) is the weighted number of cases that are mutated in the pathway p, which describes the coverage of p. Similarly, is the row of gene g in MU p , so Γ(p, g) can represent the weighted number of cases that are mutated on gene g related to the pathway p. The weighted overlap coverage of p is described, and the higher the value indicates the lower the mutual exclusivity of pathway p. Obviously, Driver(p) can simultaneously consider the high mutual exclusivity and high coverage of p on the weighted mutation matrix, so it can be used to measure the driving ability of a single pathway p.
[0121] S1043, define the synergistic driving ability between pathways according to the reconstructed pathway interaction network and the high mutation co-occurrence on the mutation data, and the specific steps are as follows:
[0122] Then, for the synergistically driven pathways, the present application defines their synergistic driving ability from the perspective of strong functional interaction and high mutation co-occurrence. For pathways p and p', the specific definition of the synergistic driving ability between them is as follows:
[0123]
[0124] O(p,p')=Γ(p)+Γ(p')-Γ(p,p')
[0125] Γ(p,p')=sum(max(MU p ;MU p’ ))
[0126] Where (MU p ; MU p’ ) represents the vertical splicing of mutation matrix MU p and MU p’ , and Γ(p,p') represents the weighted coverage of pathways p and p', so O(p,p') can represent the overlapping coverage of pathways p and p', i.e. mutation co-occurrence. is the reconstructed pathway interaction network, and its value represents the functional interaction and cooperation relationship between pathways. Alpha is a regulating parameter for adjusting the influence of functional interaction and mutation co-occurrence on the overall synergistic driving ability. In order to ensure that there is a reliable functional cooperation relationship in the final identification result, the present application will not consider the pathway pairs without connected edges on .
[0127] S1044, define the comprehensive driving weight by combining the driving ability of the pathway and the synergistic driving ability between the pathways, and identify the synergistically driven pathways on the reconstructed pathway interaction network, and the specific steps are as follows:
[0128] By combining the driving ability within the pathway and the synergistic driving ability between the pathways, the present application can define the comprehensive driving weight of the synergistically driven pathways:
[0129]
[0130] In the above equation, the parameter β will be used to control the ratio of the driving ability and the collaborative driving ability to the comprehensive driving weight. The above formula respectively considers the driving ability of the collaborative driving pathway from the aspects of the pathway inside and between the pathways, integrates the characteristics of the collaborative driving pathway such as high coverage, high mutual exclusivity, high mutation co-occurrence and strong work interaction, and can more accurately determine the collaborative driving pathway. According to the above equation, the present application can directly measure the collaborative driving ability between the pathways, and in the above selection, the pathway pair with a higher driving weight value is selected as the collaborative driving pathway. In the above selection, the pathway pair with a higher driving weight value is selected as the collaborative driving pathway.
[0131] In summary, the present application firstly integrates the copy number variation data and the single nucleotide polymorphism data into mutation data, performs data alignment and preliminary screening on the attribute heterogeneous data, then calculates the gene weight based on the network structure inside the pathway to optimize the gene pathway relationship, secondly constructs and optimizes the attribute heterogeneous network, embeds the attribute heterogeneous network based on the embedding framework based on deep joint matrix decomposition, and reconstructs the pathway interaction network, finally, the present application defines the collaborative driving ability of the pathway by using the driving ability of the single pathway and the collaboration relationship between the pathways, so as to realize accurate identification of the cancer collaborative driving pathway.
[0132] The present application firstly integrates the copy number variation data and the single nucleotide polymorphism data into comprehensive mutation data, avoids the possible data loss and untrustworthy problem of single mutation data, then performs data alignment and preliminary screening on the multi-omics attribute heterogeneous data including various biological molecules (genes, miRNA, lncRNA) and pathways, filters out low-quality data, reduces the data amount and improves the system running efficiency. Then, in order to fully utilize the rich genetic information at the gene level at the pathway level, the present application obtains the network weight of the genes in different pathways according to the network structure inside the pathway, and optimizes the gene pathway relationship according to the weight. Secondly, the present application generates an attribute heterogeneous network based on the relationship information between the biological molecules and the pathways and the respective relationship / attribute information of the biological molecules and the pathways, optimizes the attribute heterogeneous network based on the known cancer-related biological molecules, then introduces an embedding framework based on joint matrix decomposition to integrate the heterogeneous data, supplement the incomplete relationship between the pathways, and reconstruct the pathway interaction network. Finally, the driving ability of the single driving pathway is defined according to the high coverage and high mutual exclusivity of the driving pathway on the mutation data, the collaborative driving ability between the pathways is defined according to the reconstructed pathway interaction network and the mutation co-occurrence, the comprehensive driving weight is integrated, and the collaborative driving pathway is identified on the reconstructed pathway interaction network, so as to realize rapid and accurate identification of the cancer collaborative driving pathway.
[0133] The application optimizes the relationship between gene pathways based on gene weight based on pathway structure, and the related information of important genes can contribute more to identifying the driving pathway, thereby reducing the interference degree of irrelevant genes on the identification result. Then, the application uses attribute heterogeneous network embedding to provide rich genetic information related to other biological molecules for pathway level analysis, making up for the insufficient / missing problem of related biological information at the pathway level. In addition, by respectively defining the driving ability of a single pathway and the cooperative relationship between pathways, the cooperative driving ability between the pathways defined by the application comprehensively considers the high coverage, high mutual exclusivity, mutation co-occurrence and functional association of the synergistically driven pathways, has high interpretability, and can efficiently and accurately determine the pathways as synergistically driven pathways.
[0134] Embodiment 2
[0135] The embodiment 2 of the application provides a computer readable storage medium, and a program is stored on the computer readable storage medium, and the program is executed by a processor to implement the following steps:
[0136] Integrate copy number variation data and single nucleotide polymorphism data into mutation data, and perform data alignment and preliminary screening on multi-omics attribute heterogeneous data to obtain available heterogeneous data;
[0137] Obtain network weights of genes in different pathways according to the network structure inside the pathways, and optimize the gene pathway interaction network according to the weights;
[0138] Construct an initial attribute heterogeneous network according to the attribute data and relationship data corresponding to the collected multiple biological molecules and pathways, optimize the initial attribute heterogeneous network by collecting the miRNA related to the studied cancer and the optimized gene pathway interaction network, embed the optimized attribute heterogeneous network based on a joint matrix decomposition embedding framework, supplement the incomplete relationship between the pathways, and reconstruct the pathway interaction network;
[0139] Obtain mutation data corresponding to each pathway using the optimized gene pathway interaction network, define the driving ability of an individual driving pathway according to the high coverage and high mutual exclusivity of the driving pathway on the mutation data, define the cooperative driving ability between the pathways according to the reconstructed pathway interaction network and mutation co-occurrence, and define a comprehensive driving weight by combining the driving ability of the pathways and the cooperative driving ability between the pathways to identify synergistically driven pathways on the reconstructed pathway interaction network.
[0140] The detailed steps are the same as the method provided in the embodiment 1, and will not be described here.
[0141] Embodiment 3
[0142] Embodiment 3 of the present application provides an electronic device, comprising a memory, a processor, and a program stored in the memory and executable on the processor, and the processor implements the following steps when executing the program:
[0143] Integrate the copy number variation data and the single nucleotide polymorphism data into mutation data, and perform data alignment and preliminary screening on the multi-omics attribute heterogeneous data to obtain available heterogeneous data.
[0144] Obtain the network weight of the genes in different pathways according to the internal network structure of the pathways, and optimize the gene pathway interaction network according to the weight.
[0145] Construct an initial attribute heterogeneous network according to the collected attribute data and relationship data of various biological molecules and pathways, optimize the initial attribute heterogeneous network through the collected miRNA related to the studied cancer and the optimized gene pathway interaction network, integrate the heterogeneous data based on the embedding framework of joint matrix decomposition, embed the optimized attribute heterogeneous network, supplement the incomplete inter-pathway relationship, and reconstruct the pathway interaction network.
[0146] Obtain the mutation data corresponding to each pathway using the optimized gene pathway interaction network, define the driving ability of the individual driving pathway according to the high coverage and high exclusivity of the driving pathway on the mutation data, define the synergistic driving ability between the pathways according to the reconstructed pathway interaction network and the mutation co-occurrence, define the comprehensive driving weight by combining the driving ability of the pathways and the synergistic driving ability between the pathways, and identify the synergistically driven pathways on the reconstructed pathway interaction network.
[0147] The detailed steps are the same as those provided in embodiment 1, and will not be repeated here.
[0148] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer usable program code.
[0149] The embodiments of methods, apparatuses (systems) and computer program products according to the present application can be described in the general context of method steps and processes, which can be implemented in one embodiment by a program of instructions on a computer-readable storage medium executed by a computer or other programmable apparatus. The apparatuses can be specially constructed for executing the embodiments of methods, apparatuses (systems) and computer program products according to the present application or can include a computer or other programmable apparatus. Figure 1 The flow or flows and / or blocks in a flowchart and / or a block diagram Figure 1 The functions specified in the flow or flows and / or blocks in a flowchart and / or a block diagram.
[0150] These computer program instructions can also be stored in a computer-readable storage medium that can direct a computer or other programmable apparatus to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instructions which implement the flow or flows and / or blocks in a flowchart and / or a block diagram. Figure 1 The flow or flows and / or blocks in a flowchart and / or a block diagram Figure 1 The functions specified in the flow or flows and / or blocks in a flowchart and / or a block diagram.
[0151] These computer program instructions can also be loaded onto a computer or other programmable apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the flow or flows and / or blocks in a flowchart and / or a block diagram. Figure 1 The flow or flows and / or blocks in a flowchart and / or a block diagram Figure 1 The functions specified in the flow or flows and / or blocks in a flowchart and / or a block diagram.
[0152] Those skilled in the art can understand that all or part of the flow of the above-mentioned embodiment method can be realized by computer program instructions to instruct the related hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the flow of the above-mentioned embodiment of each method. Among them, the storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM) or a random access memory (RAM) and the like.
[0153] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A cancer co-driving pathway identification system based on attribute heterogeneous network embedding, characterized by comprising: a data preprocessing module configured to integrate copy number variation data and single nucleotide polymorphism data into mutation data, and perform data alignment and preliminary screening on multi-omics attribute heterogeneous data to obtain available heterogeneous data; a gene weight acquisition module configured to obtain network weights of genes in different pathways according to the internal network structure of the pathways, and optimize the gene-pathway interaction network according to the weights; wherein the network weights of genes in different pathways are obtained according to the internal network structure of the pathways, specifically: the network weights of genes in the pathways are defined using the topological order and degree of genes in the directed graph structure of the pathways, and genes with earlier topological order and higher degree values have greater network weights; an attribute heterogeneous network embedding module configured to construct an initial attribute heterogeneous network according to the collected attribute data and relationship data of multiple biological molecules and pathways, optimize the initial attribute heterogeneous network by collecting miRNAs related to the studied cancer and the optimized gene-pathway interaction network, integrate the heterogeneous data based on a joint matrix decomposition embedding framework, embed the optimized attribute heterogeneous network, supplement incomplete inter-pathway relationships, and reconstruct the pathway interaction network; a co-driving pathway identification module configured to obtain mutation data corresponding to each pathway using the optimized gene-pathway interaction network; define the driving ability of an individual driving pathway according to the high coverage and high exclusivity of the driving pathway on the mutation data; define the co-driving ability between pathways according to the reconstructed pathway interaction network and mutation co-occurrence; define the comprehensive driving weight by combining the driving ability of the pathways and the co-driving ability between the pathways, and identify the co-driving pathways on the reconstructed pathway interaction network.
2. The cancer co-driving pathway identification system based on attribute heterogeneous network embedding according to claim 1, characterized in that: the copy number variation data is patient copy number variation data, which is multi-dimensional, with each dimension corresponding to a gene site of the patient; when copy number duplication or deletion occurs, the gene data is 1, and when the copy number is normal, it is 0.
3. The cancer co-driving pathway identification system based on attribute heterogeneous network embedding according to claim 1, characterized in that: the single nucleotide polymorphism data is patient single nucleotide polymorphism data, which is multi-dimensional, with each dimension corresponding to a gene site of the patient; when single nucleotide polymorphism occurs, the gene data is 1, and when it does not occur, it is 0.
4. The cancer co-driving pathway identification system based on attribute heterogeneous network embedding according to claim 1, characterized in that: the heterogeneous data includes copy number variation data and single nucleotide polymorphism data, and also includes single-cell gene expression data divided into different cancer subtypes, which is multi-dimensional, with each dimension corresponding to a gene site of a patient of a different subtype, and the value representing the expression level of the gene; the gene interaction network is multi-dimensional, with each dimension corresponding to a gene, and the value describing the interaction intensity of genes participating in signal transmission, energy and material metabolism, and cell cycle regulation together. 5. The cancer collaborative driving pathway identification system based on attribute heterogeneous network embedding as described in claim 1, characterized in that: Copy number variation data and single nucleotide polymorphism data were integrated into mutation data, and gene alignment and preliminary screening were performed on multi-omics data, including: Obtain copy number variation data and single nucleotide polymorphism data. Integrate copy number variation data and single nucleotide polymorphism data to construct mutation data based on the mutation status of genes in batch samples. When a gene has copy number variation or single nucleotide polymorphism, the corresponding mutation data value is 1. Based on previously obtained mutation data, and considering gene mutation frequencies and patients' actual conditions, a predetermined number of genes with the highest overall mutation frequencies were selected.
6. The cancer collaborative driving pathway identification system based on attribute heterogeneous network embedding as described in claim 1, characterized in that: Optimize gene pathway interaction networks based on weights, including: The gene pathway relationships were optimized based on the network weights, and the weighted gene pathway interaction network was standardized using the min-max normalization method.
7. The cancer co-driven pathway identification system based on attribute heterogeneous network embedding as described in claim 1, characterized in that: The attribute heterogeneous network embedding module also includes: a network embedding framework based on joint matrix factorization, which uses the square of the F norm of the matrix as the loss function to embed the previously obtained attribute heterogeneous network, fuses heterogeneous data, obtains low-dimensional representations of different biological molecules, supplements incomplete pathway interaction relationships, and obtains a reconstructed pathway interaction network based on the latent semantic model.
8. The cancer co-driven pathway identification system based on attribute heterogeneous network embedding as described in claim 1, characterized in that: The overall driving weight of the synergistic driving path: where p and p' represent two pathways respectively, CoDriver(p, p') is the cooperative driving capability between and Driver(p) is the driving capability of p pathway, and Driver(p') is the driving capability of p' pathway.
9. A computer-readable storage medium having stored thereon a program, characterized in that, When the program is executed by the processor, it performs the following steps: Copy number variation data and single nucleotide polymorphism data were integrated into mutation data, and data alignment and preliminary screening were performed on heterogeneous data with multiple omics attributes to obtain usable heterogeneous data; The network weights of genes in different pathways are obtained based on the internal network structure of the pathway, and the gene-pathway interaction network is optimized based on the weights. Specifically, the network weights of genes in different pathways are obtained based on the internal network structure of the pathway by defining the network weight of genes in the pathway using the topological order and degree of the genes in the directed graph structure within the pathway. Genes with higher topological order and higher degree value have greater network weight. An initial attribute heterogeneous network was constructed based on the attribute and relationship data of various biological molecules and pathways collected. The initial attribute heterogeneous network was then optimized by collecting cancer-related miRNAs and optimizing gene pathway interaction networks. The heterogeneous data was integrated based on the embedding framework of joint matrix factorization, and the optimized attribute heterogeneous network was embedded to supplement incomplete pathway relationships and reconstruct pathway interaction networks. The mutation data corresponding to each pathway is obtained using the optimized gene-pathway interaction network, the driving ability of an individual driving pathway is defined according to the high coverage and high exclusivity of the driving pathway on the mutation data, the synergistic driving ability between pathways is defined according to the reconstructed pathway interaction network and the mutation co-occurrence, and the comprehensive driving weight is defined by combining the driving ability of the pathway and the synergistic driving ability between the pathways, so as to identify the synergistically driven pathways on the reconstructed pathway interaction network.
10. An electronic device comprising a memory, a processor, and a program stored on the memory and executable on the processor, characterized in that, The processor implements the following steps when executing the program: Integrate the copy number variation data and the single nucleotide polymorphism data into mutation data, and perform data alignment and preliminary screening on the multi-omics attribute heterogeneous data to obtain available heterogeneous data; The network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the gene-pathway interaction network is optimized according to the weight, wherein the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight of genes in different pathways is obtained according to the internal network structure of the pathway, and the network weight
Citation Information
Patent Citations
Cancer related pathway identification method based on network convergence multiomics data
CN110428866A
Method and system for matching phenotype descriptions and pathogenic variants
EP3813070A1