Biomarker recognition system based on graph convolutional neural network
By constructing dynamic biomolecular networks through graph convolutional neural networks, extracting key genes and identifying their connectivity networks, the problem of insufficient accuracy in early detection of complex diseases is solved, and efficient biomarker identification and accurate positioning of disease development stages are achieved.
Patent Information
- Application Number
- CN202510881809.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-17
AI Technical Summary
Existing technologies fail to fully capture the dynamic interaction information of genes and the evolutionary characteristics of network structures in the early detection of complex diseases, resulting in insufficient recognition accuracy and robustness. In particular, in biomarker identification, there is a lack of in-depth exploration and modeling of the dynamic changes in network structures at different pathological stages.
A biomarker identification system based on graph convolutional neural networks is used to construct a dynamic biomolecular network through graph embedding module, clustering module, generation module and detection module, extract key genes and identify their connectivity network, generate a biomarker network, and calculate dynamic network indicators to determine the key time nodes in the disease development stage.
It has improved the sensitivity and specificity of early detection of complex diseases, can effectively utilize dynamic biomolecular network information, and opened up new research directions in bioinformation processing and computational biology.
Smart Images

Figure CN120808878A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of biological information processing, and particularly relates to a biomarker identification system based on a graph convolutional neural network. BACKGROUND
[0002] The occurrence of complex diseases is often accompanied by the interaction of multiple genes and the synergistic effect with external environmental factors, thereby forming a highly heterogeneous pathological state. Early detection of these diseases is crucial for improving the prognosis of patients and improving the effectiveness of treatment. However, current diagnostic methods mostly rely on clinical symptoms or invasive detection methods, which usually make accurate judgments when the disease has developed to a more serious or irreversible stage, thus severely limiting the effectiveness and timeliness of treatment. In recent years, with the rapid development of high-throughput sequencing technology and biological network analysis methods, network-based molecular marker identification methods have been widely used in complex disease research. This kind of method identifies molecular markers closely related to disease status by constructing gene regulatory networks, protein interaction networks and other multi-level biological networks.
[0003] However, most existing methods rely on static networks or simple statistical properties of gene expression data, and fail to fully capture the complex dynamic interaction information of genes and the evolution characteristics of network structure in the disease process. Especially in biomarker identification, current methods often focus on gene expression differences under specific pathological conditions, and lack in-depth mining and modeling of dynamic changes in network structure between different pathological stages. Therefore, identifying dynamic network biomarkers is of great significance to reveal the key turning points of diseases. In addition, although some studies have introduced graph analysis or machine learning techniques for biomarker identification, most methods still fail to fully utilize the topological information of high-dimensional biological networks and their evolution patterns at different pathological stages. Since the occurrence and development of complex diseases is a nonlinear, multi-scale and dynamic evolution process, the accuracy and robustness of existing methods in identifying early markers still have a lot of room for improvement. SUMMARY
[0004] In order to solve the problems of the prior art, the present application provides a biomarker identification system based on a graph convolutional neural network. By capturing the dynamic change characteristics of gene interaction in complex disease networks, a deep representation model in high-dimensional space is constructed to improve the accuracy and effectiveness of early detection of complex diseases. Key genes are extracted from dynamic biological molecular networks and their connected networks are identified to achieve early detection of complex diseases.
[0005] In one aspect, a biomarker identification system based on a graph convolutional neural network is provided, comprising: A graph embedding module is configured to: use a graph convolutional neural network to perform graph embedding on the gene regulatory networks of different pathological stages, thereby obtaining low-dimensional vector representations of gene nodes in the gene regulatory networks of different pathological stages; The clustering module is configured to: cluster gene nodes in different pathological stages using a clustering algorithm based on the low-dimensional vector representation of gene nodes; calculate the deviation score based on the distance between each node and the center of its cluster to measure the degree of abnormality of gene nodes in each pathological stage; and store genes with deviation scores exceeding a set threshold into the abnormal gene set; A generation module is configured to: use a depth-first search greedy algorithm to extract the minimum dominating set from the abnormal gene set; use a shortest path algorithm to optimize the connectivity between the nodes of the minimum dominating set to generate a biomarker network for different pathological stages; The detection module is configured to: calculate dynamic network indicators based on the biomarker network at different pathological stages; and determine the key time nodes of the disease development stage based on the turning points of the dynamic network indicators.
[0006] The above technical solution has the following advantages or beneficial effects: This invention provides an innovative method for identifying biomarkers in dynamic networks of complex diseases. It effectively utilizes information from dynamic biomolecular networks to improve the sensitivity and specificity of early detection of complex diseases. Furthermore, these technical solutions can open up new research directions in the fields of bioinformation processing and computational biology. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0008] Figure 1 Schematic diagram of the temporal gene regulatory network construction and graph embedding representation method.
[0009] Figure 2 Schematic diagram of the use of anomaly detection algorithms to identify dynamic network biomarkers and their connected networks.
[0010] Figure 3 Schematic diagram of dynamic network indicators indicating key disease stages. DETAILED DESCRIPTION
[0011] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0012] Embodiment one As Figure 1 and Figure 2 The embodiment provides a biomarker identification system based on a graph convolutional neural network, including: A graph embedding module configured to: utilize a graph convolutional neural network to perform graph embedding on gene regulatory networks at different pathological stages, to obtain low-dimensional vector representations of gene nodes in the gene regulatory networks at different pathological stages; A clustering module configured to: based on the low-dimensional vector representations of the gene nodes, adopt a clustering algorithm to cluster the gene nodes at different pathological stages; calculate a deviation score through the distance between each node and the cluster center to which it belongs, to measure the abnormality degree of the gene nodes at different pathological stages, and store genes with a deviation score exceeding a set threshold into an abnormal gene set; A generation module configured to: use a greedy algorithm of depth-first search to extract a minimal dominating set in the abnormal gene set; adopt a shortest path algorithm to optimize the connectivity between nodes in the minimal dominating set, to generate biomarker networks at different pathological stages; A detection module configured to: based on the biomarker networks at different pathological stages, calculate dynamic network indicators; based on the turning points of changes in the dynamic network indicators, determine key time nodes of disease development stages.
[0013] Further, the graph embedding module further includes: A preprocessing module configured to: obtain gene expression data of a patient with a target disease, perform standardization processing on the gene expression data; based on clinical data of the patient with the target disease and the standardized gene expression data, calculate differentially expressed genes; combine the differentially expressed genes with a gene knowledge base of the target disease, to obtain a background gene set; A mapping module configured to: map the background gene set on a prior gene regulatory network database, to obtain a background gene regulatory network; based on the background gene regulatory network and the standardized gene expression data, use a gene regulatory network construction algorithm to construct gene regulatory networks at different pathological stages.
[0014] Further, the gene expression data of the patient with the target disease is obtained by sequencing the expression value of each gene.
[0015] Further, the standardization processing on the gene expression data specifically includes: Using The standardization processing on the gene expression data, wherein, denotes the original gene expression data.
[0016] Further, the difference expression genes are calculated based on the clinical data of the target disease patients and the standardized gene expression data, and specifically include: In the calculation of the difference expression genes of different pathological stages, the ANOVA variance analysis algorithm is used to calculate the difference expression genes by the cumulative distribution function (CDF) of the distribution and the inter-group degrees of freedom , the intra-group degrees of freedom .
[0017] , wherein represents the probability of observing the current value under the distribution, represents the ratio of the inter-group variance to the intra-group variance, which is used to measure whether the variation degree between different groups is significantly greater than the natural variation within the group. If <0.01, the current gene is identified as a difference expression gene.
[0018] Further, the difference expression genes are combined with the gene knowledge base of the target disease to obtain a background gene set, specifically including: directly taking the union of the difference expression genes and the gene knowledge base of the target disease as the background gene set.
[0019] Further, the background gene set is mapped on a priori gene regulation network database to obtain a background gene regulation network, specifically including: The priori gene regulation network database includes: a priori gene regulation network. The priori gene regulation network includes a plurality of gene pairs; since there is a regulation relationship between genes in a cell, each gene pair includes a regulatory factor and a regulated gene. In the priori gene regulation network, if the regulatory factor and the regulated gene corresponding to a gene pair both exist in the background gene set, the current gene pair of the priori gene regulation network is retained; otherwise, the current gene pair of the priori gene regulation network is deleted. According to all the retained gene pairs in the priori gene regulation network, the background gene regulation network is obtained, wherein in the background gene regulation network, the regulatory factor and the regulated gene are regarded as nodes of the network, and the two nodes belonging to the same gene pair are provided with a connection edge.
[0020] It should be understood that the regulatory factor refers to a molecule capable of regulating the expression of other genes, mainly transcription factors, and also includes certain regulatory RNAs such as miRNA or lncRNA. The regulated gene refers to a gene whose expression level is affected by the regulatory factor during transcription or translation, and the expression result is directly or indirectly affected by the regulatory factor.
[0021] Exemplarily, the background gene set is (A, B, C), the prior gene regulatory network is {(A, B), (A, C), (A, D), (B, C), (E, F)}, and then after mapping, we obtain the background gene regulatory network {(A, B), (A, C), (B, C)}.
[0022] Further, based on the background gene regulatory network and the standardized gene expression data, a gene regulatory network construction algorithm is used to construct the gene regulatory network of different pathological stages, specifically including: The input value of the gene regulatory network construction algorithm is the background gene regulatory network and the standardized gene expression data; the output value of the gene regulatory network construction algorithm is the gene regulatory network of different pathological stages.
[0023] Different pathological stages are all stages from health to late disease, and the pathological stages of different diseases are different.
[0024] Exemplarily, the prior gene regulatory network database, for example, the human gene regulatory network database RegNetwork based on prior knowledge. The prior gene regulatory network database integrates the human gene regulatory network and provides an integrated resource of prior information of gene regulatory relationship.
[0025] Exemplarily, the gene regulatory network construction algorithm, for example, the path consistency algorithm based on conditional mutual information (PCA-CMI).
[0026] The beneficial effects of the above technical solutions are: using the prior gene regulatory network database and the gene regulatory network construction method to construct the specific gene regulatory network, which can generate a directional gene regulatory network with biological significance.
[0027] Further, the graph convolutional neural network is used to perform graph embedding on the gene regulatory network of different pathological stages to obtain a low-dimensional vector representation of the gene nodes in the gene regulatory network of different pathological stages, specifically including: The input value of the graph convolutional neural network is the gene regulatory network of different pathological stages; The output value of the graph convolutional neural network is the low-dimensional vector representation of the gene nodes; The node embedding of the last layer of the graph convolutional neural network is obtained as the final low-dimensional vector representation of the gene nodes, which is represented as:
[0028] wherein, is a normalized adjacency matrix, is the result of the adjacency matrix plus self-loop, is a diagonal matrix, is the weight matrix of the layer, is an activation function, in this method, Rectified Linear Unit (ReLU) is selected, and the embedding vector initialization can be represented by .
[0029] Further, the graph convolutional neural network is a trained graph convolutional neural network; the loss function used in the training process of the graph convolutional neural network is a mean square error loss function: Mean square error loss function:
[0030] wherein, is the last layer of the model.
[0031] The loss is calculated by minimizing the loss function using the Adam optimizer:
[0032] The above process is repeated several times until the loss converges. The node embedding of the last layer of the graph convolutional neural network is output as the final obtained low-dimensional vector representation of the gene node.
[0033] The beneficial effects of the above technical solution are: using the graph convolutional neural network to perform graph embedding on the gene regulatory network at different pathological stages, mapping the high-dimensional complex biological network to a low-dimensional vector space, and obtaining the low-dimensional vector representation of the gene node in the time-series biological molecular network.
[0034] Further, the low-dimensional vector representation based on the gene node is clustered using a clustering algorithm for the gene nodes at different pathological stages, specifically including: First, the cluster center of each time stage clustering analysis is calculated using the following formula:
[0035] wherein, represents the number of nodes in the clustering cluster , and represents the low-dimensional vector representation of the gene node.
[0036] For each clustering cluster , the node As the cluster center node of the cluster.
[0037] The selection of cluster center nodes is achieved by minimizing the distance: .
[0038] Furthermore, the deviation score is calculated by the distance between each node and the center of its cluster to measure the abnormality of the gene node in each pathological stage, and genes with deviation scores exceeding the set threshold are stored in the abnormal gene set, specifically including: In the time stage When each node The low-dimensional representation of , each cluster center The low-dimensional representation of , computing nodes and time phase The Euclidean distance of each cluster center point is specified, and the minimum value of all distances is taken as the node exist The deviation score at the moment is used to measure the abnormality of the gene node in the pathological stage.
[0039] .
[0040] The beneficial effect of the above technical solution is that it uses a graph convolutional neural network to generate low-dimensional vector representations of gene nodes and employs a K-means clustering algorithm to cluster gene nodes in different pathological stages. By comparing the distance between each node and its cluster center, a deviation score is calculated to measure the degree of abnormality of the gene node in each pathological stage. A higher deviation score indicates a higher degree of abnormality.
[0041] Furthermore, the method of using a depth-first search greedy algorithm to extract the minimum dominating set from the abnormal gene set specifically includes: A greedy algorithm is one that chooses the best option under the current circumstances at each step in the hope of eventually reaching the global optimal solution. In a greedy algorithm, the greedy strategy is to choose the vertex that covers the most uncovered vertices each time.
[0042] The greedy algorithm steps are as follows: Initialization: Create an empty dominating set and pictures The set of uncovered nodes (Initially all nodes in the graph ), where Figure It is obtained by mapping the abnormal gene set on the union of gene regulatory networks at all time stages;
[0043]
[0044] denotes assigning the empty set to the dominating set S; denotes assigning all nodes V in the graph to the initial uncovered node set U; denotes adding the dataset on the right side of the arrow to the dataset on the left side of the arrow; The iteration process: in the uncovered node set , select a node , so that it can be directly connected to the most points (i.e., the node with the most neighbors), refers to all neighbor nodes of node .
[0045]
[0046] Then, node is added to the dominating set .
[0047]
[0048] Update the uncovered node set , remove node and all its neighbor nodes.
[0049]
[0050] Until all nodes are covered (i.e. is empty).
[0051] After completing the entire iteration process, the minimum dominating set in the abnormal gene set is obtained, achieving the purpose of redundancy removal.
[0052] It should be understood that since the obtained minimum dominating set cannot guarantee to be a connected network, in the context of the abnormal gene set, a few abnormal gene nodes are added by the shortest path method, so that the finally obtained dynamic network biomarker is a connected network.
[0053] Further, the shortest path algorithm is used to optimize the connectivity between the nodes of the minimum dominating set, generating biomarker networks at different pathological stages, specifically including: The Dijkstra algorithm is used to find the shortest path, and the specific process is as follows: First, the input values of the Dijkstra algorithm are the graph and the dominating set , and the minimum dominating set and the set are initialized as empty sets.
[0054]
[0055]
[0056] Through an iterative process, for each pair of nodes in the dominating set , find the shortest path between them :
[0057] Add all nodes on the path to the set :
[0058] Construct a new graph , add all nodes in to :
[0059] In , for each pair of nodes in , if , add the edge to :
[0060] After the iteration for all pairs of nodes in , obtain the final dynamic network biomarker and its connected network; is the connected network of the dynamic network biomarker, and all nodes of the connected network are the dynamic network biomarker.
[0061] The beneficial effects of the above technical solutions are: based on the minimum dominating set, the gene nodes with high deviation scores are screened, and it is ensured that the selected nodes can cover the key functional modules of the entire network. At the same time, the shortest path algorithm is used to optimize the connectivity of the nodes, and it is ensured that the identified biomarkers form a connected network structure with biological significance.
[0062] Further, the biomarker network based on different pathological stages is used to calculate a dynamic network index, specifically including: Define the dynamic network index DNI:
[0063] is the community entropy in the biomarker network, which is defined as:
[0064] wherein, refers to the first pathological stage, refers to the union of all time stage networks, refers to the biomarker network at time stage , , refers to the first stage biomarker network all nodes and the sum of the absolute values of the PCC of the first-order neighborhood of the nodes other than the nodes within the network.
[0065] is expressed as a function: , wherein represents all nodes within the biomarker network, represents other nodes in the graph other than .
[0066] is expressed as a function: , wherein is the covariance, and is the standard deviation.
[0067] The function is used to calculate the sum of the degrees of all nodes in the graph.
[0068] For the graph , the function is expressed as: , refers to all edges of the graph , and refers to all points of the graph . is a function that calculates the degree of a node , that is, the number of nodes directly connected to in the graph .
[0069] Further, as shown in Figure 3 , the turning point based on the change of the dynamic network index determines the key time node of the disease development stage, specifically including: If the dynamic network index values before the T moment are all less than the index value at the T moment, and the index values after the T moment are all less than the index value at the T moment, it indicates that the T moment is a key time node of the disease development stage. Or in other words, at the T moment, the disease has worsened.
[0070] If there is a mutation turning point in the dynamic network index change of the disease dynamic network, the time stage corresponding to the point is most likely to be the disease pre-stage we want to find.
[0071] The beneficial effects of the above technical solutions are: based on the biomarker network identified under different pathological stages, a dynamic network index is proposed to measure the information change of the obtained dynamic network biomarker in the time sequence network, so as to evaluate its specificity and diagnostic value in different stages of the disease.
[0072] The above only describes the preferred embodiments of the present application and is not used to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A biomarker recognition system based on graph convolutional neural networks, characterized by: include: A graph embedding module is configured to: use a graph convolutional neural network (GCN) to perform graph embedding on the gene regulatory networks of different pathological stages, thereby obtaining low-dimensional vector representations of gene nodes in the gene regulatory networks of different pathological stages; A clustering module is configured to: cluster gene nodes of different pathological stages using a clustering algorithm based on low-dimensional vector representations of gene nodes; The deviation score is calculated by the distance between each node and the center of its cluster to measure the abnormality of the gene node in each pathological stage. Genes with deviation scores exceeding the set threshold are stored in the abnormal gene set. A generation module is configured to: use a depth-first search greedy algorithm to extract the minimum dominating set from the abnormal gene set; use a shortest path algorithm to optimize the connectivity between the nodes of the minimum dominating set to generate a biomarker network for different pathological stages; The detection module is configured to: calculate dynamic network indicators based on the biomarker network at different pathological stages; and determine the key time nodes of the disease development stage based on the turning points of the dynamic network indicators.
2. The biomarker identification system based on graph convolutional neural network according to claim 1, characterized in that: The graph embedding module previously also includes: a preprocessing module configured to: obtain gene expression data of patients with a target disease and perform normalization on the gene expression data; calculate differentially expressed genes based on the clinical data of the patients with the target disease and the normalized gene expression data; and merge the differentially expressed genes with a gene knowledge base of the target disease to obtain a background gene set; The mapping module is configured to: map the background gene set on the prior gene regulatory network database to obtain the background gene regulatory network; based on the background gene regulatory network and the standardized gene expression data, use the gene regulatory network construction algorithm to construct the gene regulatory network of different pathological stages.
3. The biomarker identification system based on graph convolutional neural network according to claim 2, characterized in that: The calculation of differentially expressed genes based on the clinical data of patients with target diseases and the normalized gene expression data specifically includes: When calculating the differentially expressed genes in different pathological stages, the ANOVA analysis of variance algorithm was used. Cumulative distribution function and between-group degrees of freedom for the distribution , within-group degrees of freedom To calculate : ; in, Indicates Under the distribution, it is observed that the current The probability of the value, It represents the ratio of the between-group variance to the within-group variance, and is used to measure whether the degree of variation between different groups is significantly greater than the natural variation within the group; if <0.01, the current gene is identified as a differentially expressed gene.
4. The biomarker identification system based on graph convolutional neural network according to claim 2, characterized in that: Mapping the background gene set on the prior gene regulatory network database to obtain the background gene regulatory network specifically includes: The a priori gene regulatory network database includes: a priori gene regulatory network; The priori gene regulatory network includes a number of gene pairs; each gene pair includes a regulatory factor and a regulated gene; In the prior gene regulatory network, if it is possible to find a gene pair whose corresponding regulatory factor and regulated gene both exist in the background gene set, the current gene pair in the prior gene regulatory network will be retained; otherwise, the current gene pair in the prior gene regulatory network will be deleted; According to all the retained gene pairs in the prior gene regulatory network, a background gene regulatory network is obtained, wherein in the background gene regulatory network, regulatory factors and regulated genes are regarded as nodes of the network, and a connecting edge is set between two nodes belonging to the same gene pair.
5. The biomarker identification system based on graph convolutional neural network according to claim 1, characterized in that: The graph convolutional neural network is used to perform graph embedding on the gene regulatory network of different pathological stages to obtain low-dimensional vector representations of gene nodes in the gene regulatory network of different pathological stages, specifically including: The input value of the graph convolutional neural network is the gene regulatory network at different pathological stages; The output value of the graph convolutional neural network is a low-dimensional vector representation of the gene node; The node embedding of the last layer of the graph convolutional neural network is used as the final low-dimensional vector representation of the gene node, which is expressed as: ; in, is the normalized adjacency matrix, is the result of adding self-loops to the adjacency matrix, is the angle matrix, It is The weight matrix of the layer, is the activation function, and the embedding vector is initialized with To express; The graph convolutional neural network is a trained graph convolutional neural network. The loss function used in the training process of the graph convolutional neural network is the mean square error loss function: Mean squared error loss function: ; in, It is the last layer of the model; The loss is calculated by minimizing the loss function using the Adam optimizer: 。 6. The biomarker identification system based on graph convolutional neural network according to claim 1, characterized in that: The low-dimensional vector representation based on gene nodes is used to cluster gene nodes of different pathological stages using a clustering algorithm, specifically including: First, the cluster center of each time stage cluster analysis is calculated using the following formula: ; in, Represents clusters The number of nodes in A low-dimensional vector representation representing a gene node; For each cluster , select the node closest to the cluster center As the cluster center node of the cluster; The selection of cluster center nodes is achieved by minimizing the distance: 。 7. The biomarker identification system based on graph convolutional neural network according to claim 1, characterized in that: The deviation score is calculated by the distance between each node and the center of its cluster to measure the abnormality of the gene node in each pathological stage. Genes with deviation scores exceeding the set threshold are stored in the abnormal gene set, specifically including: In the time stage When each node The low-dimensional representation of , each cluster center The low-dimensional representation of , computing nodes and time phase The Euclidean distance of each cluster center point is specified, and the minimum value of all distances is taken as the node exist The deviation score at the moment is used to measure the abnormality of the gene node in the pathological stage; 。 8. The biomarker identification system based on graph convolutional neural network according to claim 1, characterized in that: The method of using a depth-first search greedy algorithm to extract the minimum dominating set from the abnormal gene set specifically includes: A greedy algorithm is one that chooses the best option under the current circumstances at each step in the hope of eventually obtaining the global optimal solution. In a greedy algorithm, the greedy strategy is to choose the vertex that covers the most uncovered vertices each time. The steps of the greedy algorithm are as follows: Initialization: Create an empty dominating set and pictures The set of uncovered nodes , where It is obtained by mapping the abnormal gene set on the union of gene regulatory networks at all time stages; ; ; It means assigning the empty set to the dominating set S; It means assigning all nodes V in the graph to the initial uncovered node set U; Indicates adding the dataset on the right side of the arrow to the dataset on the left side of the arrow; Iterative process: In the uncovered node set , select a node , so that it can directly connect the most points, Refers to the node All neighbor nodes of ; Then, the node Join the dominant set ; ; Update the uncovered node set , remove the node and all its neighbor nodes; ; Until all nodes are covered; After completing all the iterative processes, the minimum dominating set in the abnormal gene set is obtained.
9. The biomarker identification system based on graph convolutional neural network according to claim 1, characterized in that: The shortest path algorithm is used to optimize the connectivity between the nodes of the minimum dominating set to generate a biomarker network at different pathological stages, specifically including: Use Dijkstra's algorithm to find the shortest path. The specific process is as follows: First, the input value of Dijkstra's algorithm is the graph and dominating set , and initialize the minimum dominating set and collection is an empty set; ; ; After the iterative process, for the dominating set Each pair of nodes in , find the shortest path between them : ; Add all nodes on the path to the collection middle: ; Build a new graph ,Will Add all nodes in middle: ; exist In, for Each pair of nodes in ,if , then add this edge to middle: ; for After all the node pairs in the iteration are completed, the final dynamic network biomarkers and their connected networks are obtained; It is a connected network of dynamic network biomarkers, and all nodes of the connected network are dynamic network biomarkers.
10. The biomarker identification system based on graph convolutional neural network according to claim 1, characterized in that: The dynamic network indicators are calculated based on the biomarker network at different pathological stages, specifically including: Define dynamic network indicators DNI: ; refers to the community entropy in the biomarker network and is defined as: ; in, It refers to the Pathological stage, is the union of all time-stage networks, refers to the time period Biomarker network , It refers to the sum of the absolute values of the weights of all edges except the edges within the biomarker network; Expressed as a function: ; in represents all nodes in the biomarker network, Representation diagram Except Other nodes; Expressed as a function: ; in is the covariance, and is the standard deviation; function It is used to calculate the sum of all node degrees in the graph; For the graph For example, the function Expressed as: ; Refers to the picture All sides, Refers to the picture All the points; Is a computing node The function of degree, that is, the calculation in the graph In, directly with The number of connected nodes.