Heterogeneous graph neural network-based multi-dimensional data disease association prediction method and device
Through the multidimensional data disease association prediction method based on heterogeneous graph neural network, the shortcomings of the existing technology in processing heterogeneous data are solved, more efficient disease association prediction is achieved, and prediction accuracy and generalization ability are improved.
Patent Information
- Application Number
- CN202510178677.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to effectively utilize complex interactions between different types of data when processing heterogeneous data, and is susceptible to data sparsity and noise, and has limited generalization ability when predicting correlations between diseases.
A multidimensional data disease association prediction method based on heterogeneous graph neural network is adopted to collect and integrate multi-source data, build heterogeneous graphs, and dynamically allocate relationship weights through attention mechanisms, optimize the heterogeneous graph neural network model to predict potential diseases.
It significantly improves the expression ability and prediction accuracy of the disease network, can comprehensively capture complex associations between diseases, reduce artificial bias, improve analysis efficiency, and predict a large number of new disease association edges.
Smart Images

Figure CN120148882A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of medicine and information technology, and particularly to a multi-dimensional data disease association prediction method, system, terminal, and medium based on a heterogeneous graph neural network. Background Art
[0002] In the field of biomedical research, exploring the complex associations between diseases is of great significance for deeply understanding disease mechanisms. However, traditional research methods usually rely on single disease-related factors, such as genes, symptoms, or molecular pathways. Although this single-dimensional method has achieved certain results in some aspects, it is often difficult to comprehensively reveal the potential connections between diseases. In recent years, the rise of network biology has provided a new perspective for studying disease associations. By constructing a disease network with diseases as nodes and their relationships as edges, the redefinition of disease classification can be achieved, and potential disease associations can be discovered. This network-based analysis method can systematically integrate various biological information, providing new directions for disease mechanism research and treatment strategy development. Currently, the research on disease networks is no longer limited to homogeneous networks (i.e., only considering direct associations between diseases), but is gradually expanding to the construction of heterogeneous networks. These heterogeneous networks not only include the relationships between diseases, but also can integrate various factors related to diseases, such as genes, bacteria, miRNAs (microRNAs), molecular functions, cellular components, and symptoms. In particular, by performing heterogeneous graph modeling on multi-dimensional data, different types of nodes and edges can be uniformly represented in the same network, revealing complex disease association patterns. For example, by analyzing the associations between diseases and genes, potential genetic causes can be discovered, and association analysis based on symptoms and miRNAs may reveal the molecular regulatory mechanisms of diseases. With the rapid development of artificial intelligence, machine learning methods and deep learning methods have also been applied to the field of disease associations.
[0003] However, the existing technologies often show deficiencies in processing heterogeneous data, are difficult to effectively utilize the complex interactions between different types of data, and are easily affected by data sparsity and noise. In addition, the generalization ability of the existing technologies in predicting the associations between diseases is limited and difficult to meet actual needs. Therefore, the existing technologies still need to be improved. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a multi-dimensional data disease association prediction method, system, terminal, and medium based on a heterogeneous graph neural network in view of the above-mentioned defects of the existing technologies. The technical solutions adopted by the present invention are as follows:
[0005] In a first aspect, the present invention provides a multi-dimensional data disease association prediction method based on a heterogeneous graph neural network, wherein the method includes:
[0006] Collect and integrate multi-source data to obtain the association relationship and association weight between diseases and each factor;
[0007] Construct a heterogeneous graph based on the association relationship and association weight between diseases and each factor, and dynamically allocate relationship weights through an attention mechanism to optimize the heterogeneous graph, obtaining an optimized heterogeneous graph neural network model;
[0008] Predict potential diseases based on the optimized heterogeneous graph neural network model to obtain prediction results, and integrate the prediction results into the heterogeneous graph neural network model to form a final disease network.
[0009] In one implementation, the collecting and integrating multi-source data to obtain the association relationship and association weight between diseases and each factor includes:
[0010] Collect the relationship types of eight factors related to diseases from multiple data sources to obtain the association relationship between diseases and each factor, and assign weights to the relationship strength between each disease and its factor to obtain the association weight;
[0011] Perform missing value supplementation, normalization processing, and data alignment processing on the collected data.
[0012] In one implementation, the constructing a heterogeneous graph based on the association relationship and association weight between diseases and each factor includes:
[0013] Screen out the relationships between high-correlation factor nodes and disease nodes with an association score greater than a preset value from a known disease association database, determine the known high-correlation associations between disease nodes, and use the known high-correlation associations as edges to jointly construct an initial disease network with the disease node set;
[0014] Integrate the association relationship and association weight between the collected diseases and each factor into the initial disease network to obtain the heterogeneous graph.
[0015] In one implementation, the integrating the association relationship and association weight between the collected diseases and each factor into the initial disease network to obtain the heterogeneous graph includes:
[0016] Integrate factor nodes and introduce factor nodes into the initial disease network;
[0017] According to the relationship types between different disease nodes and factor nodes, construct subgraphs corresponding to each relationship type, where the edge set and weight set of each subgraph corresponding to a relationship type are calculated based on the association relationship and association weight between diseases and each factor;
[0018] Merge the subgraphs corresponding to all relationship types to obtain the heterogeneous graph.
[0019] In one implementation, dynamically allocating relationship weights through the attention mechanism to optimize the heterogeneous graph to obtain an optimized heterogeneous graph neural network model includes:
[0020] Introducing a two-layer heterogeneous graph convolution with an attention mechanism and an edge prediction module, introducing relationship attention weights, and designing the network structure of the heterogeneous graph;
[0021] Designing a cross-entropy loss function, calculating errors by combining positive samples and negative samples, and training the heterogeneous graph to obtain an optimized heterogeneous graph neural network model.
[0022] In one implementation, the two-layer heterogeneous graph convolution with an attention mechanism and an edge prediction module introduce relationship attention weights and design the network structure of the heterogeneous graph, including:
[0023] In the first-layer heterogeneous graph convolution, converting the node input features into intermediate hidden features and introducing relationship attention weights;
[0024] In the first-layer heterogeneous graph convolution, converting the intermediate hidden features into final output features and introducing relationship attention weights;
[0025] In the edge prediction module, based on the node features output from the two-layer heterogeneous graph convolution, introducing edge type embedding vectors and performing edge prediction in combination with the node features.
[0026] In one implementation, predicting potential diseases based on the optimized heterogeneous graph neural network model to obtain prediction results, and integrating the prediction results into the heterogeneous graph neural network model to form a final disease network, including:
[0027] Based on the optimized heterogeneous graph neural network model, calculating the association strength of the edges composed of factor nodes and disease nodes to obtain prediction results;
[0028] Filtering out the edges with association strength less than the threshold to obtain a new edge set;
[0029] Merging the original edge set and the new edge set in the optimized heterogeneous graph neural network model to obtain an updated edge set, and obtaining the final disease network based on the updated edge set.
[0030] In a second aspect, an embodiment of the present invention further provides a multi-dimensional data disease association prediction system based on a heterogeneous graph neural network, where the system includes:
[0031] A data integration and processing module for collecting and integrating multi-source data to obtain the association relationship and association weight between diseases and each factor;
[0032] The heterogeneous graph construction and training module is used to construct a heterogeneous graph based on the association relationships and association weights between diseases and each factor, and dynamically allocate relationship weights through an attention mechanism to optimize the heterogeneous graph, obtaining an optimized heterogeneous graph neural network model;
[0033] The prediction module of the disease network is used to predict potential diseases based on the optimized heterogeneous graph neural network model, obtain prediction results, and integrate the prediction results into the heterogeneous graph neural network model to form a final disease network.
[0034] In a third aspect, an embodiment of the present invention further provides a terminal. The terminal includes a memory, a processor, and a multi-dimensional data disease association prediction program based on a heterogeneous graph neural network stored in the memory and executable on the processor. When the processor executes the multi-dimensional data disease association prediction program based on the heterogeneous graph neural network, the steps of the multi-dimensional data disease association prediction method based on the heterogeneous graph neural network in any one of the above solutions are implemented.
[0035] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium. A multi-dimensional data disease association prediction program based on a heterogeneous graph neural network is stored on the computer-readable storage medium. When the multi-dimensional data disease association prediction program based on the heterogeneous graph neural network is executed by a processor, the steps of the multi-dimensional data disease association prediction method based on the heterogeneous graph neural network in any one of the above solutions are implemented.
[0036] Beneficial effects: Compared with the prior art, the present invention provides a multi-dimensional data disease association prediction method based on a heterogeneous graph neural network. The present invention first collects and integrates multi-source data to obtain the association relationships and association weights between diseases and each factor; constructs a heterogeneous graph based on the association relationships and association weights between diseases and each factor, and dynamically allocates relationship weights through an attention mechanism to optimize the heterogeneous graph, obtaining an optimized heterogeneous graph neural network model; predicts potential diseases based on the optimized heterogeneous graph neural network model, obtains prediction results, and integrates the prediction results into the heterogeneous graph neural network model to form a final disease network.
[0037] The present invention constructs a heterogeneous graph network containing multi-dimensional information by introducing multiple disease-related factors. Compared with the traditional disease network with a single factor, the present invention can comprehensively capture the complex associations between diseases, improving the expression ability and prediction accuracy of the network.
[0038] The present invention predicts 35,069 new disease association edges through edge prediction using a heterogeneous graph neural network, and 499 of these associations are not covered by the existing network. These newly predicted associations are supported by the literature, verifying the reliability and prediction ability of the model.
[0039] The present invention provides a new understanding of disease classification by performing community partitioning on highly correlated disease sub-networks. For example, classifying diabetes with kidney diseases, and tumors with other cancer-related diseases into the same community provides a new direction for the research and development of disease mechanisms.
[0040] The present invention automatically learns multi-dimensional associations through a heterogeneous graph neural network, avoiding the cumbersome process of manually setting meta-paths, reducing the risk of human bias, and improving the analysis efficiency at the same time. Description of the Drawings
[0041] Figure 1 It is a flowchart of a preferred embodiment of the multi-dimensional data disease association prediction method based on a heterogeneous graph neural network provided by an embodiment of the present invention.
[0042] Figure 2 It is a schematic diagram of the technical route of the multi-dimensional data disease association prediction method based on a heterogeneous graph neural network provided by an embodiment of the present invention.
[0043] Figure 3 It is a schematic diagram of the heterogeneous graph neural network structure of the attention mechanism in the multi-dimensional data disease association prediction method based on a heterogeneous graph neural network provided by an embodiment of the present invention.
[0044] Figure 4 It is a visualization result diagram of the community partitioning in the highly correlated disease sub-network in the multi-dimensional data disease association prediction method based on a heterogeneous graph neural network provided by an embodiment of the present invention.
[0045] Figure 5 It is a schematic diagram of the architecture of the multi-dimensional data disease association prediction device based on a heterogeneous graph neural network provided by an embodiment of the present invention.
[0046] Figure 6 It is a schematic block diagram of the principle of the terminal provided by an embodiment of the present invention. Detailed Embodiments
[0047] To make the objectives, technical solutions and effects of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0048] The flowchart shown in the drawings is only an example illustration, and does not necessarily include all the contents, operations or steps, nor is it necessary to be executed in the described order. For example, some operations or steps can be decomposed, combined or partially merged, so the actual execution order may be changed according to the actual situation.
[0049] It should be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0050] It should be understood that, in order to facilitate a clear description of the technical solutions of the embodiments of the present invention, in the embodiments of the present invention, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and roles. For example, the first control information and the second control information are only used to distinguish different control information and do not limit their sequence.
[0051] Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and "first", "second", etc. do not necessarily mean different.
[0052] It should also be understood that the term "and / or" used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0053] With the rapid development of deep learning technology, graph neural networks (GNNs) have gradually become one of the mainstream methods for processing network data. GNNs can learn the feature representations of network nodes and their neighbors through an iterative manner, providing an effective tool for capturing complex network structures and semantic relationships. In particular, the proposal of heterogeneous graph neural networks (HGNNs) provides a new solution for processing heterogeneous networks containing multiple types of nodes and edges. Different from traditional GNNs that assume the network is a homogeneous structure, HGNNs can fully utilize the diverse relational semantic information in heterogeneous networks to achieve a deeper level of modeling of complex network structures. In the field of disease network research, the introduction of HGNNs provides the possibility of integrating multi-source heterogeneous data. For example, integrating the associations between diseases and multi-dimensional factors such as genes, symptoms, bacteria, miRNAs, etc. into a unified network can not only more accurately predict disease associations but also reveal potential connections that have not been discovered before. This method can avoid the cumbersome process of manually setting meta-paths, reduce the risk of human bias, and at the same time improve the quality and reliability of prediction results. In addition, by combining the attention mechanism and edge type embedding technology, HGNNs can dynamically assign weights to different associations, further enhancing the model's ability to capture complex relationships. Therefore, based on heterogeneous graph neural networks, this embodiment proposes a multi-dimensional data disease association prediction method based on heterogeneous graph neural networks and its application, which has high prediction accuracy and generalization ability, providing important technical support for disease research.
[0054] The multi-dimensional data disease association prediction method based on heterogeneous graph neural network in this embodiment can be applied to a terminal, which is an intelligent product device such as a mobile phone or a computer. As Figure 1 shown in
[0055] Step S100: Collect and integrate multi-source data to obtain the association relationship and association weight between diseases and each factor.
[0056] The collection and integration of multi-source data is the basis for constructing a human heterogeneous disease network, and its purpose is to comprehensively integrate various relationship types related to diseases and construct a heterogeneous graph network that can reflect the complex associations of diseases. Specifically, in this embodiment, the relationship types of eight factors related to diseases can be collected from multiple reliable data sources, including genes, bacteria, miRNAs, symptoms, biological processes, cellular components, molecular functions, and pathways. These relationship types come from different biological research fields and respectively reflect the genetic basis, microbial associations, epigenetic regulation, clinical manifestations, and molecular biological mechanisms of diseases. Among them, the association data between genes and diseases mainly come from the CTD database, which contains a large amount of gene-disease association information and their weights. The association data between bacteria and diseases is based on the research results of Wei Ma et al. The association data between symptoms and diseases comes from the symptom information extracted from electronic medical records. The association data between miRNAs and diseases comes from the research dataset of Shuting Jin et al. The selection of these data sources ensures the authority and diversity of the data and provides a reliable data basis for constructing a comprehensive heterogeneous graph network. In this embodiment, the node V factor sets of the relationship types between diseases and each factor are respectively:
[0057] V gene , V bacteria , V mirNA , V symptom , V bio_proc , V cell_comp , V mol_func and V pathway , corresponding to the above-mentioned genes, bacteria, miRNAs, symptoms, biological processes, cellular components, molecular functions, and pathways respectively. The disease node set is denoted as V disease , then the node set of the heterogeneous graph:
[0058] V = V disease ∪V gene ∪V bacteria ∪V miRNA ∪V symptom ∪V bio_proc ∪
[0059] V cell_comp ∪Vcell_comp ∪V mol_func ∪V pathway ;
[0060] The edge set of the heterogeneous graph, that is, the association relationship between nodes is as follows:
[0061] E = {(v i , v j , r ij ) | v i , v j ∈V, r ij ∈R}, where R is the relationship type and r ij represents the relationship type of the edge. Generate a feature vector h i for each node as the input for subsequent neural network model training. The feature vector is jointly determined by the attribute feature a i of the disease node v i ij and the association relationship r ij : h i = concat(a ij , r ij ). Where a ij represents the attribute feature of the disease node v i , and r ij represents the association relationship between the node and other nodes. Denote the total data set composed of the node set and the edge set as: DATA = (V, E).
[0062] Furthermore, to ensure that the feature information of the disease nodes is as comprehensive as possible, the screening criterion is set such that the disease appears in at least seven data sources, ensuring that the association features of these diseases cover most of the relationship types. Finally, 313 diseases and their related data are screened out, and these diseases are used for subsequent heterogeneous graph construction.
[0063] In the process of data integration in this embodiment, weights can be assigned to the relationship strength between each disease and its factors to obtain the association weights, so as to quantify the relationship strength between different diseases and their factors. The introduction of these weights can better reflect the contributions of different relationship types in disease associations, and at the same time provide a quantitative basis for subsequent network analysis and prediction. And this embodiment can perform missing value supplementation, standardization processing, and data alignment processing on the collected data to ensure the accuracy and integrity of the data, as Figure 2 shown.
[0064] Specifically, first, supplement the missing data set to ensure the accuracy and reliability of the data. Traverse the data set DATA. If the edge weight w i between a disease node v j and a disease node v ij(i.e., the association weight used to quantify the strength of the relationship between a disease and its factors) is missing in a certain data source, and the average value is used for filling:
[0065]
[0066] Among them, N(v i ) is the set of neighbor nodes directly connected to the disease node v i , and w ik represents the edge weight between the node v i and the neighbor node v k . For data sources with continuous variables such as gene expression values, a linear interpolation method is used for missing value processing:
[0067]
[0068] Since the dimensions of different data sources may vary greatly, it is necessary to standardize the data to eliminate the influence of magnitude and ensure the fairness of subsequent analysis. In this embodiment, Min-Max standardization is used:
[0069]
[0070] Among them, represents the standardized edge weight, min(w) represents the minimum value of all edge weights in the current data source, and max(w) represents the maximum value of all edge weights in the current data source.
[0071] Furthermore, data alignment processing is performed on multi-source data. Since there are inconsistent naming or annotation methods among multi-source data, such as different gene naming rules or inconsistent disease classification criteria. It is necessary to align these data to unify the format.
[0072] Specifically, a unified mapping rule is constructed to align the nodes in different data sources. In this embodiment, a standardized naming system (HGNC gene symbol or ICD disease classification code) is used:
[0073] Align(v i ) = Standard(v i ) (4)
[0074] Among them, Align(v i ) represents the alignment result of the node v i , and Standard(v i ) is the corresponding entry in the standard naming system. For nodes that cannot be directly matched, a fuzzy matching algorithm is used. In the present invention, the Levenshtein distance is used for fuzzy matching:
[0075]
[0076] where Lev(v i , v j ) is the Levenshtein distance between disease node v i and disease node v j .
[0077] |v i |, |v j | represents the lengths of disease node v i and disease node v j . Sim(v i , v j ) is the similarity score. If Sim(v i , v j ) > θ, where θ = 0.9, then it is considered that disease node v i and disease node v j are the same node.
[0078] The multi-source data after data supplementation, normalization, and unified mapping are fused into the same heterogeneous graph structure to form a complete data representation. By adding the edge weight set to the representation of the total data set, the final representation of the heterogeneous graph is: Data = (V, E, W).
[0079] In this embodiment, through the integration of multi-source data, the limitations of a single data source are overcome, and the complex associations of diseases are comprehensively captured. The combination of authoritative data sources and multi-dimensional information improves the reliability of the data and the expressive ability of the network. The construction of a heterogeneous graph based on unified weights significantly improves the prediction accuracy of subsequent models, providing a new perspective for disease classification and mechanism research.
[0080] Step S200: Construct a heterogeneous graph based on the association relationships and association weights between diseases and each factor, and dynamically allocate relationship weights through an attention mechanism to optimize the heterogeneous graph, obtaining an optimized heterogeneous graph neural network model.
[0081] After the integration of multi-source data, the association relationships and association weights between diseases and each factor are obtained. According to the heterogeneous graph Data = (V, E, W) obtained in the above step S100, in this embodiment, diseases and their related factors are organized into a unified graph structure, aiming to provide basic data for subsequent graph neural network models. This embodiment can screen out the relationships between high-correlation factor nodes and disease nodes with an association score greater than a preset value from a known disease association database, determine the known high-correlation associations between disease nodes, and use the known high-correlation associations as edges to jointly construct an initial disease network with the disease node set. Then, the association relationships and association weights between the collected diseases and each factor are integrated into the initial disease network to obtain the heterogeneous graph.
[0082] Combination Figure 2 As shown in Figure 2 , in the construction of the heterogeneous graph, it is first necessary to construct an initial disease network based on the known high-correlation associations between disease nodes. The purpose of this network is to lay a foundation for subsequent integration of various factor nodes and serve as the input for subsequent optimization.
[0083] Specifically, in this embodiment, high-correlation factor node-disease node relationships with an association score greater than 0.9 can be screened out from the known disease association database. These high-correlation associations are used as edges E high , which, together with the disease node set V disease , jointly constitute the initial disease network G base . At the same time, it is necessary to ensure that each disease node has at least one edge connecting to other disease nodes. If there is no connection, the node is retained for supplementation during subsequent optimization.
[0084] The initial disease network of this embodiment is defined as:
[0085] G base =(V disease ,E high ) (6)
[0086] Wherein, V disease represents the disease node set, containing 313 nodes.
[0087] E high ={(v i ,v j )|v i ,v j ∈V disease ,w ij >0.9} represents the set of high-correlation disease association edges. w ij is the association weight between disease v i and v j . The output disease network G base . Among them, it contains disease nodes and the edges of their high-correlation associations.
[0088] Next, after constructing the initial disease network, it is necessary to integrate various factor nodes (such as genes, symptoms, etc.) and their association relationships with diseases into the network to form a complete heterogeneous graph.
[0089] In this embodiment, factor nodes can be integrated. The factor nodes are introduced into the initial disease network. Then, according to the relationship types between different disease nodes and factor nodes, subgraphs corresponding to each relationship type are constructed. Among them, the edge set and weight set of the subgraph corresponding to each relationship type are calculated based on the association relationship and association weight between the disease and each factor. Finally, the subgraphs corresponding to all relationship types are merged to obtain the heterogeneous graph.
[0090] Specifically, eight factor nodes: Gene V gene , Bacteria V bacteria , miRNA V miRNA , Symptom V symptom , Biological Process V bio_proc , Cellular Component V cell_comp , Molecular Function V mol_func and Pathway V pathway are introduced into the initial disease network G base . Ensure that each factor node is connected to at least one disease node.
[0091] Further, according to different relationship types (such as disease - gene, disease - symptom, etc.), a sub - graph G r =(V r , E r , W r ) is constructed for each relationship type. V r is the node set of each sub - graph relationship type, and the edge set E r and the weight set W r of each sub - graph are calculated based on the association relationship and association weight between the disease and each factor.
[0092] Further, all sub - graphs of relationship types are merged to form a complete heterogeneous graph.
[0093] G = {G r =(V r , E r , W r )|r ∈ R} (7)
[0094] where R is the set of relationship types, and r represents the relationship type.
[0095] Finally, after constructing the complete heterogeneous graph, this embodiment can further optimize the graph structure, update the weights of existing edges, and predict potential new edges and their weights. This process is implemented through a heterogeneous graph neural network based on the attention mechanism. The final heterogeneous graph contains 313 disease nodes and 48,287 edges.
[0096] This embodiment integrates multi - dimensional data (such as genes, symptoms, miRNAs, etc.) into the heterogeneous graph through a unified weight mechanism, constructing a complex network containing multiple types of nodes and edges. It overcomes the limitations of traditional single - data - source and manual meta - path methods and realizes the automated modeling of complex associations between multi - dimensional factors.
[0097] Furthermore, the design and training of the Heterogeneous Graph Neural Network (HGNN) aim to predict potential associations between diseases through node feature and graph structure information and optimize the weights of existing edges. HGNN is suitable for processing heterogeneous graphs with multiple node and edge types. Compared with traditional Graph Neural Networks (GNN), HGNN can effectively capture diverse relationships in heterogeneous graphs. Since current heterogeneous graph convolutions adopt simple aggregation methods and fail to distinguish the importance of different neighbor nodes and relationships. In this embodiment, based on the complete heterogeneous graph obtained through the above steps, a heterogeneous graph neural network is designed, which consists of a two-layer heterogeneous graph convolution HeteroGraphConv and an edge prediction module. By introducing an attention mechanism, the weight coefficient of each relationship is calculated to achieve more accurate information aggregation, complete the prediction of potential associations between diseases, and optimize the existing network structure. Specifically, in this embodiment, a two-layer heterogeneous graph convolution with an attention mechanism and an edge prediction module are introduced, relationship attention weights are introduced, and the network structure of the heterogeneous graph is designed. Then, a cross-entropy loss function is designed to calculate the error by combining positive and negative samples, and the heterogeneous graph is trained to obtain an optimized heterogeneous graph neural network model.
[0098] Combined with Figure 2 As shown, in practical applications, first, in order to adapt to the multiple node and edge types in the heterogeneous graph. The design of the heterogeneous graph neural network includes a two-layer heterogeneous graph convolution with an attention mechanism and an edge prediction module.
[0099] Specifically, combined with Figure 3 As shown, in the first-layer heterogeneous graph convolution HeteroGraphConv, the node input features are converted into intermediate hidden features. The input features of the node:
[0100]
[0101] where R represents the set of relationship types, N r (i) represents the set of neighbor nodes of the disease node v i under the relationship r, c ij is the normalization coefficient of the disease node v i and the disease node v j , is the weight matrix corresponding to the relationship r, represents the bias vector corresponding to the relationship r, and σ(*) is the activation function.
[0102] Furthermore, by introducing the relationship attention weight α r , formula (8) can be changed to:
[0103]
[0104] where, Denoted as the relational weights calculated by the multi-layer perceptron MLP. In this way, the semantic features of different relationships are explicitly modeled, improving the expressive power of the heterogeneous graph convolution.
[0105] Secondly, design the second layer of heterogeneous graph convolution HeteroGraphConv to convert the intermediate hidden features into the final output features. The hidden features of the first layer are converted into the final node representations
[0106] Among them, the parameter definitions are similar to those of the first layer. is the weight matrix, is the bias vector.
[0107] Similarly, introduce the relational attention weight α r , and formula (10) can be changed to:
[0108]
[0109] Furthermore, design the edge prediction module. In the edge prediction module, based on the node features output from the HeteroGraphConv layer, the dot product method is used to calculate the existence probability of the edge. That is, according to the final representations of the nodes and predict the existence probability P of the edge ij and the edge weight Its corresponding formula is expressed as follows:
[0110]
[0111] Among them, is denoted as the dot product of the node feature vectors, are the feature representations of nodes i and j in the last layer of convolution, σ is the activation function Relu, which is used to map the dot product result to the range [0,1]. scale is the scaling factor used to adjust the weight range.
[0112] Furthermore, the edge prediction module uses a simple dot product method to calculate the existence probability of the edge, without considering the influence of the edge type. In this embodiment, it is proposed to introduce the edge type embedding vector and combine the node features for edge prediction:
[0113] Based on this, formula (12) is adjusted to:
[0114]
[0115] Among them, W e represents the edge weight matrix, β rRepresents the embedding vector of edge type r. By this method, the characteristics of nodes and the influence of edge types are comprehensively considered, improving the accuracy of edge prediction.
[0116] Finally, the model training process. The weight matrix of the model is assigned through the adjacency relationship between nodes. During the training process, node features are optimized so that the dot product result P ij is closer to the true value. Extract the actually existing edges from the existing edge set as positive samples. Generate non-existent edges through negative sampling (randomly select unconnected node pairs).
[0117] Furthermore, design the loss function. In this embodiment, the cross-entropy loss function is adopted, and the error is calculated by combining positive samples and negative samples. Maximize the prediction score of positive samples and minimize the prediction score of negative samples:
[0118]
[0119] where S is the sample set, E + is the positive sample edge set, E - is the negative sample edge set. P ij represents the predicted probability of edge (i, j), and l represents the layer. Optimize the loss function by using the gradient descent algorithm to update the network parameters and
[0120] Finally, obtain the optimized edge weight set W optimized , the newly predicted edge set E new and its weight W new , as well as the fully optimized heterogeneous graph G optinized =(V, E optimized , W optimized ).
[0121] In this embodiment, by constructing a heterogeneous graph, diseases and their multi-dimensional factors are integrated into a unified network, overcoming the limitations of traditional homogeneous networks. Subgraph decomposition and relationship type modeling significantly improve the scalability and expressive power of the network, providing high-quality input for subsequent graph neural network training.
[0122] In summary, this embodiment introduces a relational attention mechanism, dynamically assigns weights to different neighbor nodes, explicitly models multiple relationship semantics, and improves the accuracy of information aggregation. Compared with traditional graph convolutional networks (GNNs), it significantly enhances the ability to model diverse relationships in heterogeneous networks, avoiding information loss and over-smoothing problems.
[0123] Step S300: Predict potential diseases based on the optimized heterogeneous graph neural network model to obtain prediction results, and integrate the prediction results into the heterogeneous graph neural network model to form a final disease network.
[0124] In this embodiment, based on the optimized heterogeneous graph neural network model, the association strength of the edges composed of factor nodes and disease nodes can be calculated to obtain prediction results. Then, filter out the edges with association strength less than the threshold to obtain a new edge set. Next, merge the original edge set and the new edge set in the optimized heterogeneous graph neural network model to obtain an updated edge set, and obtain the final disease network based on the updated edge set. Finally, use the trained heterogeneous graph neural network to predict new disease associations and integrate the prediction results into the final disease network.
[0125] Specifically, according to the above step S300, the HGNN model based on the attention mechanism outputs node features h (l) Starting from these features, they are processed by two layers of heterogeneous graph convolution HeteroGraphConv, integrating various relationship information of the nodes in the graph. The node features can be used to calculate whether there is a potential association between two nodes.
[0126] The node features learned by the HGNN according to the attention mechanism and Use the dot product method to calculate the association strength of the edges, that is, formula (14). Further, for the original edges. For the edges that already exist in the training set, re-predict and update their association strengths. For new edge prediction, calculate the association strength for all unconnected node pairs in the graph. Set a threshold T, such as T = 0, filter out the edges with association strength less than the threshold, and only retain the edges with association strength P ij > T as new edges. Therefore, for the prediction of new edges:
[0127]
[0128] where, E new is the predicted new edge set, and E is the edge set in the original graph.
[0129] Further, for all edges, use the model to predict their weights (i.e., association strengths). The weight prediction formula is the same as the edge prediction formula, but no longer screen the original edges, only update the weights:
[0130]
[0131] where, W ij is the weight of the edge, representing the association strength between nodes i and j.
[0132] Further, integrate the new edges and the original edges. Integrate the predicted new edges Enew Merge with the original edge E to form an updated edge set:
[0133] E T = E new ∪ E(18)
[0134] Integrate edge weights and update the edge weights using the predicted weights:
[0135] W T ={W ij |(i,j)∈E T}}(19)
[0136] Finally, obtain a complete disease network based on the updated edge set and the updated edge weights,
[0137] G T =(V,E T ,W T ).
[0138] To further analyze disease communities, edges with an association strength higher than a specific threshold are selected to construct a sub-network.
[0139]
[0140] The final sub-network is denoted as, G sub =(V,E sub ,W sub ),W sub is the corresponding edge weight set. As Figure 4 shown, Figure 4 is the community partition of 19 sub-networks of highly correlated diseases selected by the model prediction designed in this embodiment.
[0141] In this embodiment, by introducing an attention mechanism and edge type embedding, the modeling ability of heterogeneous graph convolution for diverse relationships is significantly improved, enhancing the expressive ability and prediction accuracy of the model. Dynamically adjusting the influence weights of neighbor nodes avoids the over-smoothing problem of traditional aggregation methods and ensures the accuracy of information transmission.
[0142] To ensure the reliability and accuracy of the constructed disease network and the newly predicted disease associations by the model, a research design verification and evaluation process is carried out, covering the performance evaluation of the network, the reliability verification of the predicted associations, and the analysis of the medical significance of the new findings. Through verification and evaluation, the research has demonstrated that the HGNN prediction based on the attention mechanism is significantly scientific and reliable. The model proposed in the present invention shows a high overlap rate with existing disease networks (such as gene, miRNA, bacteria, and symptom networks), indicating its ability to integrate multi-source information to reproduce known associations. Through comparative analysis with random networks, the prediction results of this model are statistically significantly better than random noise, proving its ability to capture the true associations between diseases. In addition, 499 newly predicted edges by the model are verified through medical literature, and many of the potential associations with eczema and chondrosarcoma provide new clues for medical research. Meanwhile, the sub-network constructed by screening high-weight edges reveals 19 disease communities, and some of the non-traditional disease combinations provide new perspectives for the study of pathological mechanisms and disease classification.
[0143] Based on the above embodiments, the present invention also provides a multi-dimensional data disease association prediction system based on a heterogeneous graph neural network, as Figure 5 shown, the system includes: a data integration and processing module 10, a heterogeneous graph construction and training module 20, and a disease network prediction module 30. The data integration and processing module 10 is used to collect and integrate multi-source data to obtain the association relationship and association weight between diseases and each factor. The heterogeneous graph construction and training module 20 is used to construct a heterogeneous graph based on the association relationship and association weight between diseases and each factor, and dynamically allocate relationship weights through the attention mechanism to optimize the heterogeneous graph, obtaining an optimized heterogeneous graph neural network model. The disease network prediction module 30 is used to predict potential diseases based on the optimized heterogeneous graph neural network model to obtain prediction results, and integrate the prediction results into the heterogeneous graph neural network model to form a final disease network.
[0144] The working principles of the various modules in the multi-dimensional data disease association prediction device based on a heterogeneous graph neural network in this embodiment are the same as those of the various steps in the above method embodiment, and will not be elaborated here.
[0145] The various modules in the above multi-dimensional data disease association prediction device based on a heterogeneous graph neural network can be implemented in whole or in part by software, hardware, and their combinations. The above-mentioned modules can be embedded in the processor in the terminal in hardware form or be independent of it, or can be stored in the memory in the terminal in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.
[0146] Based on the above embodiments, the present invention also provides a terminal, and the principle block diagram of the terminal can be as Figure 6As shown. The terminal may include one or more processors 100 ( Figure 6 only one is shown in the figure), a memory 101, and a computer program 102 stored in the memory 101 and executable on the one or more processors 100.
[0147] In one embodiment, the so-called processor 100 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0148] In one embodiment, the memory 101 may be an internal storage unit of the electronic device, such as the hard disk or memory of the electronic device. The memory 101 may also be an external storage device of the electronic device, such as a plug-in hard disk equipped on the electronic device, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 101 may also include both an internal storage unit and an external storage device of the electronic device. The memory 101 is used to store computer programs and other programs and data required by the terminal. The memory 101 may also be used to temporarily store data that has been output or is to be output.
[0149] Those skilled in the art can understand that Figure 6 the principle block diagram shown in the figure is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the terminal to which the solution of the present invention is applied. The specific terminal may include more or fewer components than those shown in the figure, or combine certain components, or have a different component layout.
[0150] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, operational database, or other medium used in the embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0151] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multidimensional data disease association prediction method based on heterogeneous graph neural network, characterized in that: The method comprises: Collect and integrate multi-source data to obtain the correlation and correlation weight between the disease and each factor; A heterogeneous graph is constructed based on the correlation relationship and correlation weight between the disease and each factor, and the relationship weight is dynamically assigned through an attention mechanism to optimize the heterogeneous graph to obtain an optimized heterogeneous graph neural network model; Potential diseases are predicted based on the optimized heterogeneous graph neural network model to obtain prediction results, and the prediction results are integrated into the heterogeneous graph neural network model to form a final disease network.
2. The multidimensional data disease association prediction method based on heterogeneous graph neural network according to claim 1 is characterized in that: The collection and integration of multi-source data to obtain the correlation relationship and correlation weight between the disease and each factor includes: Collecting the relationship types of eight factors related to the disease from multiple data sources, obtaining the correlation relationship between the disease and each factor, and assigning a weight to the strength of the relationship between each disease and its factor to obtain the correlation weight; The collected data were processed with missing value supplement, standardization and data alignment.
3. The multidimensional data disease association prediction method based on heterogeneous graph neural network according to claim 1 is characterized in that: The heterogeneous graph is constructed based on the association relationship and association weight between the disease and each factor, including: Screening out the relationship between the high correlation factor node and the disease node with a correlation score greater than a preset value from the known disease association database, determining the known high correlation associations between the disease nodes, and using the known high correlation associations as edges and the disease node set to jointly construct an initial disease network; The collected association relationships and association weights between the disease and each factor are integrated into the initial disease network to obtain the heterogeneous graph.
4. The multidimensional data disease association prediction method based on heterogeneous graph neural network according to claim 3 is characterized in that: The collected association relationship and association weight between the disease and each factor are integrated into the initial disease network to obtain the heterogeneous graph, including: Integrate factor nodes and introduce factor nodes into the initial disease network; According to the relationship types between different disease nodes and factor nodes, a subgraph corresponding to each relationship type is constructed, wherein the edge set and weight set of the subgraph corresponding to each relationship type are calculated based on the association relationship and association weight between the disease and each factor; The subgraphs corresponding to all relationship types are merged to obtain the heterogeneous graph.
5. The multidimensional data disease association prediction method based on heterogeneous graph neural network according to claim 4 is characterized in that: The method of dynamically allocating relationship weights through an attention mechanism and optimizing the heterogeneous graph to obtain an optimized heterogeneous graph neural network model includes: Introduced a double-layer heterogeneous graph convolution with an attention mechanism and an edge prediction module, introduced relational attention weights, and designed a heterogeneous graph network structure; A cross entropy loss function is designed, errors are calculated by combining positive samples and negative samples, and the heterogeneous graph is trained to obtain an optimized heterogeneous graph neural network model.
6. The multidimensional data disease association prediction method based on heterogeneous graph neural network according to claim 5 is characterized in that: The double-layer heterogeneous graph convolution with the attention mechanism and an edge prediction module are introduced, the relational attention weight is introduced, and the network structure of the heterogeneous graph is designed, including: In the first layer of heterogeneous graph convolution, the node input features are converted into intermediate hidden features, and relational attention weights are introduced; In the first layer of heterogeneous graph convolution, the intermediate hidden features are converted into final output features, and relational attention weights are introduced; In the edge prediction module, based on the node features output from the two-layer heterogeneous graph convolution, the edge type embedding vector is introduced and combined with the node features for edge prediction.
7. The multidimensional data disease association prediction method based on heterogeneous graph neural network according to claim 6 is characterized in that: The method predicts potential diseases based on the optimized heterogeneous graph neural network model to obtain prediction results, and integrates the prediction results into the heterogeneous graph neural network model to form a final disease network, including: Based on the optimized heterogeneous graph neural network model, the correlation strength of the edge consisting of factor nodes and disease nodes is calculated to obtain the prediction results; Filter out the edges whose association strength is less than the threshold to obtain a new edge set; The original edge set and the new edge set in the optimized heterogeneous graph neural network model are merged to obtain an updated edge set, and the final disease network is obtained based on the updated edge set.
8. A multidimensional data disease association prediction system based on heterogeneous graph neural network, characterized in that: The system comprises: Data integration processing module, used to collect and integrate multi-source data to obtain the correlation and correlation weight between the disease and each factor; A heterogeneous graph construction and training module is used to construct a heterogeneous graph based on the correlation relationship and correlation weight between the disease and each factor, and dynamically allocate the relationship weight through an attention mechanism to optimize the heterogeneous graph to obtain an optimized heterogeneous graph neural network model; The prediction module of the disease network is used to predict potential diseases based on the optimized heterogeneous graph neural network model, obtain prediction results, and integrate the prediction results into the heterogeneous graph neural network model to form a final disease network.
9. A terminal, characterized in that: The terminal includes a memory, a processor, and a multidimensional data disease association prediction program based on a heterogeneous graph neural network stored in the memory and executable on the processor. When the processor executes the multidimensional data disease association prediction program based on a heterogeneous graph neural network, the steps of the multidimensional data disease association prediction method based on a heterogeneous graph neural network as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a multidimensional data disease association prediction program based on a heterogeneous graph neural network. When the multidimensional data disease association prediction program based on a heterogeneous graph neural network is executed by a processor, the steps of the multidimensional data disease association prediction method based on a heterogeneous graph neural network as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Method for dynamically predicting bone mass reduction risk of psoriasis patient and public platform
CN120748730A
A method and public platform for dynamically predicting the risk of osteopenia in psoriasis patients
CN120748730B
Cancer gene identification method based on multiplexing heterogeneous graph neural network
CN120895103A
Phosphorylation site and disease association prediction method based on graph neural network
CN121506261A
A method for predicting phosphorylation site and disease association based on graph neural network
CN121506261B