Drug-disease association prediction model construction method and prediction method

By constructing a multi-perspective representation learning model and a deep fusion strategy, the difficulty of integrating multi-source heterogeneous biological data was solved, the accuracy and reliability of drug-disease association prediction were improved, and a richer comprehensive representation was generated.

CN120763631AActive Publication Date: 2025-10-10JIANGSU KANION PHARMA CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510860060.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-10-10
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

In existing drug-disease association prediction methods, the integration of multi-source heterogeneous biological data is difficult, resulting in incomplete information processing and reducing the accuracy of prediction results.

Method used

A multi-perspective representation learning model is adopted to extract the representation information of drugs and diseases from different perspectives by constructing drug homogeneity graphs, disease homogeneity graphs and heterogeneous graphs. Multi-perspective comparative learning and deep fusion are then performed to establish a deep learning model, and the Transformer encoder and contrastive learning are used to optimize the parameters.

Benefits of technology

It significantly improves the accuracy and reliability of drug-disease association prediction, enhances the cross-perspective consistency and robustness of the model, reduces perspective bias, and generates a more informative comprehensive representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120763631A_ABST
    Figure CN120763631A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of deep learning, and discloses a drug-disease association prediction model construction method and a prediction method, and the method comprises the steps: obtaining drug similarity data, disease similarity data, and drug-disease-protein association data; based on the drug similarity data, the disease similarity data, the drug-disease-protein relevance data and a pre-established multi-view representation learning model, extracting representation information of the drug and the disease under different views; performing multi-view comparative learning on the representation information to obtain comparative loss information; performing heterogeneous fusion based on the representation information to obtain comprehensive drug representation information and comprehensive disease representation information; and establishing a deep learning model based on the comparison loss information, the comprehensive drug representation information and the comprehensive disease representation information, and performing parameter training to obtain a trained drug-disease association prediction model. According to the invention, the accuracy of predicting the potential drug-disease association is obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning technology, and in particular to a method for constructing a drug-disease association prediction model and a prediction method. Background Art

[0002] With the surge in data in the biomedical field, using computational methods to analyze these data to discover potential drug-disease associations has become an important research direction.

[0003] Currently, a variety of computational methods are available for predicting drug-disease associations, including network-based methods (e.g., network propagation and random walks), matrix decomposition-based methods, and machine learning or deep learning-based methods. These methods typically utilize data processing based on the drug's chemical structure, genomic data (e.g., targets and pathways), side effect information, and disease phenotypes, genetic associations, and related genes. However, this information comes from diverse sources and differs in nature, often reflecting the characteristics of the drug or disease from different perspectives. The multi-source, multi-perspective, and heterogeneous nature of the data often leads to difficulties in data integration and incomplete information processing, reducing the accuracy of drug-disease association prediction results. Summary of the Invention

[0004] In view of this, the present invention provides a method for constructing a drug-disease association prediction model and a prediction method to solve the problem of low accuracy of drug-disease association prediction results in the prior art.

[0005] In a first aspect, the present invention provides a method for constructing a drug-disease association prediction model, the method comprising:

[0006] Obtain drug similarity data, disease similarity data, and drug-disease-protein association data;

[0007] Extract representation information of drugs and diseases from different perspectives based on drug similarity data, disease similarity data, drug-disease-protein association data, and a pre-established multi-perspective representation learning model;

[0008] Perform multi-view contrast learning on the representation information to obtain contrast loss information;

[0009] Perform heterogeneous fusion based on representation information to obtain comprehensive drug representation information and comprehensive disease representation information;

[0010] Based on contrast loss information, comprehensive drug representation information, and comprehensive disease representation information, a deep learning model is established and parameter training is performed to obtain a drug-disease association prediction model after training.

[0011] This invention significantly enhances the performance of the drug-disease association prediction model and improves the accuracy of predicting potential drug-disease associations through deep mining, representation alignment and effective fusion of multi-perspective information and the use of a multi-task gradient balance optimization strategy during training.

[0012] In an optional embodiment, based on drug similarity data, disease similarity data, drug-disease-protein association data, and a pre-established multi-perspective representation learning model, the representation information of drugs and diseases from different perspectives is extracted, including:

[0013] Construct a drug homogeneity graph based on drug similarity data;

[0014] Based on disease similarity data, a disease homogeneity graph is constructed;

[0015] Construct heterogeneous graphs based on drug-disease-protein association data;

[0016] Input the drug homogeneity graph and the disease homogeneity graph into the similarity perspective representation model in the multi-perspective representation learning model to obtain drug similarity perspective representation information and disease similarity perspective representation information;

[0017] Input the heterogeneous graph into the neighborhood view representation model in the multi-view representation learning model to obtain neighborhood view drug representation information and neighborhood view disease representation information;

[0018] Inputting the heterogeneous graph into the multi-hop view representation model in the multi-view representation learning model to obtain multi-hop view drug representation information and multi-hop view disease representation information;

[0019] The heterogeneous graph is input into the meta-path view representation model in the multi-view representation learning model to obtain the meta-path view drug representation information and the meta-path view disease representation information.

[0020] In this embodiment, multiple specialized representation learning pathways are designed in parallel. By processing homogeneous similarities and multiple heterogeneous views in parallel, information about drugs and diseases can be extracted from different angles and granularities, making fuller use of multi-source, multi-modal biomedical data and improving the accuracy of drug-disease association prediction.

[0021] In an optional embodiment, heterogeneous fusion is performed based on the representation information to obtain comprehensive drug representation information and comprehensive disease representation information, including:

[0022] Fusing the neighborhood view drug representation information, the multi-hop view drug representation information, and the meta-path view drug representation information to obtain fused drug representation information;

[0023] Fusing the disease representation information of the neighborhood view, the disease representation information of the multi-hop view, and the disease representation information of the meta-path view to obtain fused disease representation information;

[0024] Integrate the drug similarity perspective representation information with the fused drug representation information to obtain a drug representation sequence containing two perspectives;

[0025] Integrate the disease similarity perspective representation information with the fused disease representation information to obtain a disease representation sequence containing two perspectives;

[0026] The drug representation sequence and disease representation sequence are input into the encoder network respectively to obtain comprehensive drug representation information and comprehensive disease representation information.

[0027] In this implementation, representations from homogeneous similarity views are effectively integrated with representations from one or more heterogeneous information views. Deep fusion strategies, such as stacking features and feeding them into a Transformer encoder, are employed. A self-attention mechanism is used to dynamically capture and weight the interactions and complementarities between features from different perspectives to generate a final integrated representation. This deep fusion mechanism can better capture the complex interactions between different information perspectives, generating a more informative integrated representation, which can effectively improve the accuracy of drug-disease association predictions.

[0028] In an optional embodiment, a deep learning model is established based on contrast loss information, comprehensive drug representation information, and comprehensive disease representation information, and parameter training is performed to obtain a trained drug-disease association prediction model, including:

[0029] obtaining drug-disease samples;

[0030] Determine the main task loss based on the drug-disease samples, the comprehensive drug representation information, and the comprehensive disease representation information;

[0031] Construct a deep learning model loss function based on the main task loss and contrast loss information;

[0032] The deep learning model loss function is trained to obtain a drug-disease association prediction model.

[0033] In this implementation, a contrastive learning objective is introduced for representations of the same drug or disease learned from different heterogeneous information processing pathways, such as neighborhood views, multi-hop views, and meta-path views. This contrastive learning objective minimizes the differences between representations of the same node under different views while maximizing the differences between different node representations, thereby forcing the model to learn node embeddings that are consistent across views and more robust. The introduction of contrastive learning significantly improves the cross-view consistency and robustness of node representations, reduces viewpoint bias, and makes the final drug-disease association prediction results more accurate and reliable.

[0034] In an optional embodiment, determining the main task loss based on the drug-disease sample, the comprehensive drug representation information, and the comprehensive disease representation information includes:

[0035] Extracting drug representation vectors and disease representation vectors corresponding to drug-disease samples from the comprehensive drug representation information and the comprehensive disease representation information;

[0036] Input the drug representation vector and the disease representation vector into a preset function to obtain a drug-disease joint representation;

[0037] The drug-disease joint representation is input into a pre-set multi-layer perceptron classifier, and the predicted drug-disease association result is output;

[0038] Based on the association prediction results and the actual association results, the main task loss is determined.

[0039] During the model training process, this embodiment simultaneously optimizes the main task loss for predicting association accuracy and the auxiliary task loss for contrastive learning to improve representation quality, effectively improving the accuracy and reliability of drug-disease association prediction results.

[0040] In a second aspect, the present invention provides a method for predicting drug-disease association, the method comprising:

[0041] Obtain drug-disease information to be predicted;

[0042] The drug-disease information to be predicted is input into the drug-disease association prediction model constructed according to the drug-disease association prediction model construction method of any of the above embodiments to obtain the drug-disease association prediction result to be predicted.

[0043] The present invention proposes a drug-disease association prediction method based on the fusion of multi-view graph neural networks and contrastive learning. This method overcomes the limitations of related technologies in processing multi-source heterogeneous biological data, learning high-quality node representations, and balancing multi-task optimization, and effectively achieves more accurate and reliable drug-disease association prediction.

[0044] In a third aspect, the present invention provides a device for constructing a drug-disease association prediction model, the device comprising:

[0045] Acquisition module, used to obtain drug similarity data, disease similarity data, and drug-disease-protein association data;

[0046] The extraction module is used to extract the representation information of drugs and diseases from different perspectives based on drug similarity data, disease similarity data, drug-disease-protein association data, and a pre-established multi-perspective representation learning model;

[0047] Contrastive learning module, used to perform multi-view contrast learning on the representation information to obtain contrast loss information;

[0048] A fusion module is used to perform heterogeneous fusion based on representation information to obtain comprehensive drug representation information and comprehensive disease representation information;

[0049] The training module is used to establish a deep learning model and perform parameter training based on contrast loss information, comprehensive drug representation information, and comprehensive disease representation information to obtain a drug-disease association prediction model after training.

[0050] In a fourth aspect, the present invention provides a computer device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to execute the method for constructing a drug-disease association prediction model according to the first aspect or any corresponding embodiment thereof.

[0051] In a fifth aspect, the present invention provides a computer-readable storage medium storing computer instructions, which are used to enable a computer to execute the method for constructing a drug-disease association prediction model according to the first aspect or any corresponding embodiment thereof.

[0052] In a sixth aspect, the present invention provides a computer program product comprising computer instructions for causing a computer to execute the method for constructing a drug-disease association prediction model according to the first aspect or any corresponding embodiment thereof.

[0053] It should be noted that the drug-disease association prediction model construction apparatus, computer device, computer-readable storage medium, and computer program product provided by the present invention correspond to the aforementioned drug-disease association prediction model construction method. Therefore, regarding the beneficial effects of the drug-disease association prediction model construction apparatus, computer device, computer-readable storage medium, and computer program product, please refer to the description of the corresponding beneficial effects of the drug-disease association prediction model construction method above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0055] Figure 1 is a flowchart of a method for constructing a drug-disease association prediction model according to an embodiment of the present invention;

[0056] Figure 2 is a flowchart of another method for constructing a drug-disease association prediction model according to an embodiment of the present invention;

[0057] Figure 3 2 is a structural block diagram of a device for constructing a drug-disease association prediction model according to an embodiment of the present invention;

[0058] Figure 4 Schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0059] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present invention.

[0060] In related technologies, the biomedical data used to predict drug-disease associations are diverse and complex, and processing this complex data presents several challenges. For example, integrating heterogeneous data from different sources and modalities (such as drug structural similarity, disease semantic similarity, drug-target interactions, and disease-gene associations) into a unified framework for effective learning is a key issue. Simple feature concatenation or early fusion may not fully capture the complex dependencies between data. For another example, many models may focus on learning representations from only one or a few information perspectives, such as relying solely on drug chemical structure or known association networks, while ignoring other perspectives that may contain complementary information, resulting in incomplete and inaccurate learned drug or disease representations. Another example is that different data sources or different processing methods may lead to learned node representations with specific "perspective biases." Learning high-quality representations that fully reflect information from each perspective, are consistent across perspectives, and are robust to noise and missing data is the key to improving predictive performance.

[0061] In view of this, according to an embodiment of the present invention, an embodiment of a method for constructing a drug-disease association prediction model is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0062] In this embodiment, a method for constructing a drug-disease association prediction model is provided, which can be executed by devices such as servers, terminals, and mobile terminals. Figure 1 FIG. 1 is a flow chart of a method for constructing a drug-disease association prediction model according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0063] Step S101: Obtain drug similarity data, disease similarity data, and drug-disease-protein association data.

[0064] In this embodiment, the drug similarity data includes: fingerprint similarity matrix S obtained by calculating drug chemical fingerprints DFP and the Gaussian kernel function similarity matrix S calculated based on the drug interaction spectrum DGIP Disease similarity data include: disease phenotypic similarity matrix S DPS and the Gaussian kernel function similarity matrix S calculated based on the disease association spectrum DGIP Drug-disease-protein association data include the validated drug-disease association matrix A DI , drug-protein association matrix A DP , disease-protein association matrix A IP , and the initial feature vectors X of drugs, diseases, and proteins D 、X I 、X P , such as Mol2Vec embedding, disease ontology embedding, ESM protein embedding, etc.

[0065] Step S102 , based on the drug similarity data, disease similarity data, drug-disease-protein association data and a pre-established multi-perspective representation learning model, extract the representation information of drugs and diseases in different perspectives.

[0066] Reference Figure 2As described, after obtaining drug similarity data, disease similarity data, and drug-disease-protein association data, graph construction begins. Based on the collected drug similarity data, a drug homogeneity graph G_DD is constructed. Based on the collected disease similarity data, a disease homogeneity graph G_II is constructed. Based on the collected drug-disease-protein association data, a heterogeneous graph G_Het containing these three types of nodes is constructed. Based on the constructed drug homogeneity graph G_DD, disease homogeneity graph G_II, and heterogeneous graph G_Het, information extraction is performed through a pre-established multi-perspective representation learning model, wherein the pre-established multi-perspective representation learning model includes a similarity perspective representation model, a neighborhood view representation model, a multi-hop view representation model, and a meta-path view representation model, which are respectively used to extract corresponding drug and disease representation information. Specifically, the drug homogeneity graph G_DD and the disease homogeneity graph G_II are input into the similarity perspective representation model to obtain the drug similarity perspective representation information H D,sim And the disease similarity perspective representation information H I,sim ; Input the heterogeneous graph G_Het into the neighborhood view representation model to obtain the neighborhood view drug representation information H D,neighbor And the neighborhood view disease representation information H I,neighbor ; Input the heterogeneous graph into the multi-hop view representation model to obtain the multi-hop view drug representation information H D,multihop And multi-hop view disease representation information H I,multihop ; Input the heterogeneous graph into the meta-path view representation model to obtain the meta-path view drug representation information H D,metapath And meta-path view disease representation information H I,metapath .

[0067] Step S103: Perform multi-view contrast learning on the representation information to obtain contrast loss information.

[0068] That is, the drug and disease representation information extracted based on the neighborhood view representation model, the multi-hop view representation model, and the meta-path view representation model are compared and aligned. Based on the selected representation group, positive sample pairs (representations of the same node in different views) and negative sample pairs (representations of different nodes) are constructed. A contrastive learning loss function, such as a function based on InfoNCE or supervised contrastive loss, is applied to calculate the similarity scores between these positive and negative sample pairs, and the contrastive loss information L is obtained accordingly. contrastive This contrastive loss information is then incorporated into the overall optimization objective of model training and optimized through backpropagation and parameter updates.

[0069] Taking drugs as an example, from the neighborhood view, drug representation information H D,neighbor , multi-hop view drug representation information H D,multihop and meta-path view drug representation information H D,metapathTwo sets of heterogeneous view representations are selected from . For each entity type, the system performs contrastive learning loss calculation. Specifically, the two selected sets of representations are first L2 normalized to obtain the vector zD a and vector zD b , where a and b represent two sets of heterogeneous view representations. Then, for each drug i in the batch, the dot product of its representation under the two views is calculated, that is, the positive sample similarity sim(zD a ,i,zD b ,i). At the same time, the dot product (all sample pair similarity) between the representation of drug i in view a and the representation of all drugs j in the batch in view b is calculated. Using these similarity scores, the contrast loss value L of the current batch is calculated according to the definition of the InfoNCE loss function. contrastive_D , which involves the temperature hyperparameter τ. The same operation is performed on the disease representation to obtain the contrast loss value L contrastive_I The weighted sum of these two loss values ​​​​is obtained to obtain the total contrast loss information L contrastive The contrast loss information is used to assist in model training.

[0070] Step S104: performing heterogeneous fusion based on the representation information to obtain comprehensive drug representation information and comprehensive disease representation information.

[0071] Specifically, we can first select the drug homogeneity graph, the disease homogeneity graph to obtain the drug similarity perspective representation information, the disease similarity perspective representation information, and at least one heterogeneous view representation. The perspective representation vectors of the selected drug (or disease) are stacked along a new dimension to form a representation sequence. This representation sequence is passed as input to a Transformer encoder model. The model processes the sequence internally through a multi-head self-attention mechanism and a feedforward network layer. After processing, the final output vector is obtained from the Transformer encoder. This vector is the final comprehensive drug representation information H that integrates multi-perspective information. D_final With comprehensive disease indication information H I_final .

[0072] Step S105 , based on the contrast loss information, the comprehensive drug representation information, and the comprehensive disease representation information, a deep learning model is established and parameter training is performed to obtain a drug-disease association prediction model after training.

[0073] Specifically, the deep learning model can be a common model such as a convolutional neural network. A multi-task gradient balance optimization strategy can be adopted to coordinate the prediction task loss and contrast loss information, and the gradient contribution from different loss terms to the shared model parameters can be evaluated and adjusted to achieve parameter training of the model.

[0074] The present invention provides a method for constructing a drug-disease association prediction model based on multi-view heterogeneous graph information fusion and contrastive representation learning. First, a graph neural network is used to process the homogeneous similarity network between drugs and diseases to extract a baseline representation reflecting their intrinsic attribute associations. Second, information is extracted in parallel from the heterogeneous association network containing drug, disease, and protein nodes through three view representation models with different focuses. Among them, the neighborhood view representation model focuses on aggregating the direct, cross-type neighbor information of the node and combines it with local structural propagation; the multi-hop view representation model uses advanced heterogeneous graph neural networks to capture the extensive context and long-range dependencies of the nodes in the entire heterogeneous network; and the meta-path view representation model uses an adaptive mechanism to mine and encode potential, semantically specific high-order connection patterns. Subsequently, a contrastive learning mechanism is introduced to force the representations learned for the same node under different heterogeneous views to align with each other to enhance the consistency and robustness of the representation. Next, the representations of the homogeneous similarity view and the heterogeneous view are deeply fused, and a more information-rich comprehensive embedding of drugs and diseases is extracted using sequence modeling technology. Finally, based on this comprehensive embedding, downstream prediction models are trained, ultimately using the prediction model to output the association probability of drug-disease pairs. This method significantly enhances the performance of drug-disease association prediction models and improves the accuracy of predicting potential drug-disease associations by deeply integrating and analyzing the multiple relationships and attribute information between drugs, diseases, and proteins. Through deep mining, representation alignment, and effective fusion of multi-perspective information, and employing a multi-task gradient balance optimization strategy during training, this method significantly improves the performance of drug-disease association prediction models and the accuracy of predicting potential drug-disease associations.

[0075] In some optional embodiments, step S102, i.e., extracting representation information of drugs and diseases from different perspectives based on drug similarity data, disease similarity data, drug-disease-protein association data, and a pre-established multi-perspective representation learning model, includes:

[0076] Step S1021: construct a drug homogeneity graph based on the drug similarity data.

[0077] The connection relationship between drugs can be determined based on drug similarity data and the K-nearest neighbor algorithm to construct a drug homogeneity graph G_DD.

[0078] Step S1022: construct a disease homogeneity graph based on the disease similarity data.

[0079] The disease comprehensive similarity matrix S can be calculated I , thereby constructing the disease homogeneity graph G_II.

[0080] Step S1023: constructing a heterogeneous graph based on the drug-disease-protein association data.

[0081] The drug-disease association matrix A can be used DI, drug-protein association matrix A DP , disease-protein association matrix A IP The node types (drug D, disease I, protein P) and the edges between them are clarified to construct the heterogeneous graph G_Het.

[0082] Step S1024 : inputting the drug homogeneity graph and the disease homogeneity graph into the similarity perspective representation model in the multi-perspective representation learning model to obtain drug similarity perspective representation information and disease similarity perspective representation information.

[0083] Specifically, the drug homogeneity graph G_DD and its normalized initial drug features X D , disease homogeneity graph G_II and its associated disease initial features X I , and input into their respective graph neural network models. The similarity perspective representation model first performs a linear transformation on the initial input features. The transformed features are then processed by a multi-layer stack of graph converter layers. In each graph converter layer, the node representation is first calculated by a multi-head self-attention mechanism. This self-attention mechanism determines the attention distribution between nodes through query (Query), key (Key), and value (Value) projection, and weightedly aggregates information from other nodes in the graph; the output of the attention mechanism will be sent to the feedforward neural network (FFN) for position-by-position nonlinear transformation. The layer contains residual connections and layer normalization operations to assist in training. Finally, the node representation vectors output after all graph converter layers are processed constitute the drug similarity perspective representation information H_D_sim and the disease similarity perspective representation information H_I_sim.

[0084] Taking drugs as an example, the system combines the drug homogeneity graph G_DD and the normalized drug feature X D Enter the Transformer architecture, a graph neural network model. The model first applies a linear transformation to the input features. Then, the data flows through L graph transformer layers. In each layer, the system performs multi-head self-attention calculations to obtain the attention-weighted node information AttnOutput; then performs the first residual connection and layer normalization operation; then inputs the result into the feedforward network (FFN) for transformation; finally, performs the second residual connection and layer normalization to obtain the output representation H of the layer. D (l+1) The output H of the final layer L is D (L) As the drug similarity perspective representation information H D,simo .

[0085] The system considers the disease homogeneity graph G_II and feature X I Perform exactly the same operation process to produce disease similarity perspective representation information H I,simo .

[0086] Step S1025 , inputting the heterogeneous graph into the neighborhood view representation model in the multi-view representation learning model to obtain neighborhood view drug representation information and neighborhood view disease representation information.

[0087] Specifically, in the neighborhood view representation model, we first identify different types of neighbor nodes directly connected to each central node (drug or disease) in the heterogeneous graph G_Het. Next, we aggregate the feature information of each neighbor type, including calculating the attention score between the neighbor node features and the central node features, and using these scores to perform weighted summation on the neighbor features to obtain the neighbor feature aggregation result based on the attention mechanism. The aggregation results from different neighbor types (including the node's own features) are stacked. Subsequently, a local graph structure propagation mechanism (graph propagation algorithm) is applied, which performs iterative feature updates on the graph structure. In each round of update, the representation of the node will be affected by the current representation of its neighbors, thereby smoothing and integrating a wider range of local structural information. After a predetermined number of rounds of propagation updates, the final neighborhood view drug representation information H is obtained. D_neighbor and neighborhood view disease representation information H I_neighbor .

[0088] The following is a detailed introduction. The system applies the neighborhood view encoding strategy to process the heterogeneous graph G_Het and the normalized initial features. First, for each type of node in the heterogeneous graph, the features aggregated from different types of neighboring nodes are calculated. For example, for the drug node, the features aggregated from the disease neighbors are calculated (by I Left multiply the normalized A DI These aggregated neighbor features are then stacked with the node's own features along the channel dimension. The system then applies an attention module consisting of two linear layers and a tanh activation function to calculate the attention weights for each channel and use these weights to perform a weighted summation of the stacked features to obtain a preliminary aggregate representation H. agg Afterwards, the preliminary aggregate representations of all node types are concatenated and used as the initial input H of the graph propagation algorithm. (0) The algorithm is based on the symmetric normalized adjacency matrix A of the heterogeneous graph G_Het. homo Perform K rounds of iterative updates on the , each round of updates combines the representation H of the previous round (k) and the initial representation H (0) , the calculation formula is H (k+1) =(1-αprop)A homo H (k) +α prop H (0) After completing K rounds of iterations, the final representation H (K)The drug and disease parts are divided by node type to obtain the neighborhood view drug representation information H D,neighbor and neighborhood view disease representation information H I,neighbor .

[0089] Step S1026 , inputting the heterogeneous graph into the multi-hop view representation model in the multi-view representation learning model to obtain multi-hop view drug representation information and multi-hop view disease representation information.

[0090] Specifically, a multi-hop view encoding strategy can be applied to process the heterogeneous graph G_Het and the initial features of the nodes. This strategy uses a heterogeneous graph neural network model, that is, the multi-hop view representation model in this embodiment, specifically the heterogeneous graph transformer (HGT). The model processes information through multi-layer stacking. In each layer, type-specific parameterized transformations (linear projections) are applied to process the node representation and relationship information for different types of nodes and edges; then, the attention scores based on the target node, source node and the edge type between them are calculated to determine the relevance of information transfer; then, the messages from multi-hop neighbors that have undergone type-specific transformations are weighted and aggregated according to the attention scores; finally, the aggregated information is used to update the representation of the central node. After processing through all layers, the output is a multi-hop view drug representation information H that can capture the global heterogeneous network context and long-range dependencies. D_multihop and multi-hop view disease representation information H I_multihop .

[0091] The following is a detailed introduction. The system applies a multi-hop view encoding strategy and inputs the heterogeneous graph G_Het and the normalized initial features into the heterogeneous graph conversion (HGT) model. The model contains L HGT In each layer, for each target node t, the system considers all source nodes s connected to it and the relationship type r = (s, φ, t) between them. For each relationship, the system calculates the attention weight Atts from the source node to the target node. s,t The calculation involves performing type-specific linear transformations on the source and target node representations (generating keys and queries) and calculating dot products, while multiplying by a relationship priority factor. The system then uses these attention weights to perform weighted aggregation, and obtains the aggregate message AggMsg after the source node representation has undergone type-specific linear transformations (generating values). t Finally, through residual connection and layer normalization, the aggregated message is used to update the representation of the target node t. HGT After layer processing, the final representation of drug and disease nodes is extracted as the multi-hop view drug representation information H D,multihop and multi-hop view disease representation information H I,multihop .

[0092] Step S1027 , inputting the heterogeneous graph into the meta-path view representation model in the multi-view representation learning model to obtain meta-path view drug representation information and meta-path view disease representation information.

[0093] Specifically, a set of basic graph operation units is first defined, each corresponding to a type of neighbor information aggregation or feature transformation based on a specific relationship type. The meta-path view representation model uses an internal learning mechanism (using learnable architectural parameters) or a sampling process to determine which basic operations to select during information propagation and how to combine them into effective computational paths. Node features are transmitted, transformed, and aggregated along these dynamically determined or learned operation paths that may represent high-order connection patterns. Ultimately, the mechanism outputs the meta-path view drug representation information H_D_metapath and the meta-path view disease representation information H_I_metapath.

[0094] The following is a detailed introduction. The system applies the meta-path view encoding strategy and uses a searchable neural network model. The model contains a Cell module that executes N step First, the normalized initial features are input into the Cell module. At each step i(1≤i <N step ), the system learns the architecture parameter α based on the current seq , α res To determine the basic graph operation Op to be used (the operation is based on the predefined adjacency matrix A corresponding to different relationship types). The system calculates a sequence connection term (based on the state H of the previous step). (i-1) obtained by the selected sequence operation) and a residual connection term (derived from the state H of all previous steps (j),j<i The state H of the current step is obtained by adding these two items. (i) . Complete N step After the first step of calculation, the final state output by the Cell module is processed (layer normalization, activation function), and then the representation of drugs and diseases is extracted according to the node type as the meta-path view drug representation information H D,metapath and meta-path view disease representation information H I,metapath .

[0095] In this embodiment, a similarity perspective representation model is adopted, that is, a graph neural network is used to process the similarity network within the drug or disease, which can effectively capture the intrinsic association pattern based on the overall attributes. The neighborhood view representation model is adopted, which can focus on the direct interaction of nodes in the heterogeneous network, aggregate the features of directly connected neighbors of different types, and consider the influence of the local graph structure. The multi-hop view representation model is adopted, and a powerful heterogeneous graph neural network is applied to perform multiple rounds of information propagation on the complete drug-disease-protein heterogeneous network, learn the deep contextual representation of the node in the global network environment, and capture long-range dependencies. The meta-path view representation model is adopted, that is, an adaptive or searchable mechanism is adopted to explicitly or implicitly learn and encode predefined or automatically discovered meaningful high-order connection paths in the heterogeneous network, which can effectively capture semantic associations beyond simple adjacency relationships.

[0096] In this embodiment, multiple specialized representation learning pathways are designed in parallel. By processing homogeneous similarities and multiple heterogeneous views in parallel, information about drugs and diseases can be extracted from different angles and granularities, making fuller use of multi-source, multi-modal biomedical data and improving the accuracy of drug-disease association prediction.

[0097] In some optional embodiments, the above step S104, i.e., performing heterogeneous fusion based on the representation information to obtain comprehensive drug representation information and comprehensive disease representation information, includes:

[0098] Step S1041 , fusing the neighborhood view drug representation information, the multi-hop view drug representation information, and the meta-path view drug representation information to obtain fused drug representation information.

[0099] Step S1042 , fusing the neighborhood view disease representation information, the multi-hop view disease representation information, and the meta-path view disease representation information to obtain fused disease representation information.

[0100] Step S1043 : Integrate the drug similarity perspective representation information and the fused drug representation information to obtain a drug representation sequence containing two perspectives.

[0101] Step S1044 : Integrate the disease similarity perspective representation information and the fused disease representation information to obtain a disease representation sequence containing two perspectives.

[0102] In step S1045 , the drug representation sequence and the disease representation sequence are input into the encoder network respectively to obtain comprehensive drug representation information and comprehensive disease representation information.

[0103] Specifically, taking drugs as an example, we first fuse the drug representations of three heterogeneous views, namely the neighborhood view drug representation information H D,neighbor , multi-hop view drug representation information H D,multihop, Meta-path view drug representation information H D,metapath Then, these three representation vectors are stacked along a new dimension and the average value is calculated to obtain a fused drug representation information H that integrates the three heterogeneous view information. D,het_fused At the same time, the disease information H of the corresponding neighborhood view is expressed I,neighbor , multi-hop view disease representation information H I,multihop , Meta-path view disease representation information H I,metapath Perform exactly the same stacking and averaging operations to obtain the fused disease representation information H I,het_fused .

[0104] Next, the obtained drug similarity perspective representation information H D,sim With fusion drug information H D,het_fused Perform the final integration. Specifically, the drug similarity perspective information H D,sim With fusion drug information H D,het_fused Stacked along a new dimension to form a representation sequence containing two perspectives D =Stack(H D,sim ,H D,het_fused ). The final representation sequence is Stacked D Input to a Transformer encoder network (drug_trans). The network processes this sequence through its internal multi-layer self-attention mechanism and feedforward network layer, deeply interacting and integrating the similarity perspective and the fused heterogeneous perspective information. After the Transformer encoder is processed, the final comprehensive drug representation information H is output. D,final . The disease similarity perspective represents H I,sim and fusion disease representation information H I,het_fused Perform exactly the same stacking average and Transformer encoding process to obtain the final comprehensive disease representation information H I,final .

[0105] In this embodiment, representations from homogeneous similarity views are effectively integrated with representations from one or more heterogeneous information views. A deep fusion strategy, such as stacking features and feeding them into a Transformer encoder, is employed. A self-attention mechanism is used to dynamically capture and weight the interactions and complementarities between features from different perspectives to generate a final integrated representation. This deep fusion mechanism can better capture the complex interactions between different information perspectives, generating a more informative integrated representation, which can effectively improve the accuracy of drug-disease association predictions.

[0106] In some optional embodiments, the step S105 of establishing a deep learning model based on the contrast loss information, the comprehensive drug representation information and the comprehensive disease representation information and performing parameter training to obtain a trained drug-disease association prediction model comprises:

[0107] In step S1051, a drug-disease sample is obtained.

[0108] In step S1052, a main task loss is determined based on the drug-disease sample, the comprehensive drug representation information and the comprehensive disease representation information.

[0109] In step S1053, a deep learning model loss function is constructed based on the main task loss and the contrast loss information.

[0110] In step S1054, the deep learning model loss function is trained to obtain the drug-disease association prediction model.

[0111] Specifically, a composite loss function L is first defined total The composite loss function is a weighted sum of the main task loss L prediction and the auxiliary task contrast loss, i.e., the contrast loss information L contrastive L = L + λL total L = L + λL prediction L = L + λL contrastive . Wherein, L prediction is calculated by inputting the obtained prediction result and the true label into a cross-entropy loss function; L contrastive is the contrast loss information calculated above; λ is a preset weight coefficient. In each batch of training, the system calculates L total the gradient of all trainable parameters of the model. Then, a gradient balancing optimization algorithm is applied. The algorithm checks the gradients of the shared parameters generated by the main task loss L prediction and the contrast loss information L contrastive respectively. If a gradient conflict is detected (judged based on norm comparison and direction inner product), the algorithm adjusts the gradient of the contrast loss information L contrastive (projects and scales) to generate an adjusted gradient Next, the balanced gradient Finally, the system uses a basic optimizer Adam to update all model parameters according to this (possibly balanced adjusted) gradient. The training process is repeated until the model converges.

[0112] In this embodiment, a contrastive learning objective is introduced for the representations of the same drug or disease learned from different heterogeneous information processing pathways, such as neighborhood view, multi-hop view, meta-path view, etc. The contrastive learning objective can maximize the difference between the representations of different nodes while minimizing the difference between the representations of the same node in different views, thereby forcing the model to learn node embeddings that are consistent across views and more robust. The introduction of contrastive learning significantly improves the cross-view consistency and robustness of node representations, reduces view bias, and makes the final drug-disease association prediction results more accurate and reliable.

[0113] In addition, in model training, the main task loss for predicting association accuracy and the contrastive learning auxiliary task loss for improving representation quality are optimized simultaneously. To solve the potential optimization conflict, a gradient balancing technique is used to intelligently adjust the contribution of the auxiliary task to the shared parameter gradient during backpropagation, effectively ensuring the stability and final performance of the overall model training.

[0114] In some optional embodiments, the step S1052 of determining the main task loss based on the drug-disease sample, the comprehensive drug representation information, and the comprehensive disease representation information comprises:

[0115] Extracting the drug representation vector and the disease representation vector corresponding to the drug-disease sample from the comprehensive drug representation information and the comprehensive disease representation information.

[0116] Inputting the drug representation vector and the disease representation vector into a preset function to obtain a drug-disease joint representation.

[0117] Inputting the drug-disease joint representation into a pre-set multi-layer perception classifier to output an association prediction result of the drug-disease to be predicted.

[0118] Determining the main task loss based on the association prediction result and the real association result.

[0119] Specifically, for a given drug-disease pair (i, j), the representation H D,final (i) of drug i and the representation H I,final (j) of disease j are taken out from the obtained comprehensive drug representation information and comprehensive disease representation information, respectively. The system then performs an interaction operation, specifically calculating the element-wise product of the two representation vectors to obtain a joint representation H pair(i, j). This joint representation is then input into a multi-layer perceptron (MLP) classifier. The MLP consists of multiple linear layers, ReLU activation functions, and dropout layers. After the data is processed sequentially through these layers, the output linear layer ultimately produces a two-dimensional vector, Output_logits. The two components of this vector correspond to the raw prediction scores for the association and non-association of the drug-disease pair, i.e., the association prediction result.

[0120] During the model training process, this embodiment simultaneously optimizes the main task loss for predicting association accuracy and the auxiliary task loss for contrastive learning to improve representation quality, effectively improving the accuracy and reliability of drug-disease association prediction results.

[0121] In order to verify the effectiveness of the present invention in drug-disease association prediction, three benchmark data sets, B-dataset, C-dataset, and F-dataset, were selected in this embodiment for verification, and the prediction results are shown in Table 1. Among them, AUC (Area Under the Curve) is the area under the ROC curve, and AUPR (Area Under the Precision-Recall Curve) is the area under the precision-recall curve, both of which are commonly used indicators for evaluating the performance of classification models. AMDGT (Attention Mechanism for Deep Graph-based Transfer Learning, attention mechanism for transfer learning based on deep graphs), SLGCN (Semi-supervised Learning Graph Convolutional Network, semi-supervised learning graph convolutional network), GCGB (Graph Convolutional Global Belief, graph convolution global belief), RGLDR (Reinforced Graph Learning with Deep Representations, reinforced graph learning and deep representation) are all common deep learning models.

[0122] Table 1

[0123]

[0124] The present invention proposes a drug-disease association prediction method based on the fusion of multi-view graph neural networks and contrastive learning. It constructs two key graph views for parallel processing, one of which is a homogeneous similarity graph view that reflects the intrinsic attributes of entities, and the other is a heterogeneous interaction graph view that reflects the complex interactions between entities. A dedicated, advanced graph neural network module is designed for each view to perform deep representation learning. In particular, for heterogeneous graph views, multiple parallel encoding strategies are adopted, such as a strategy based on attention aggregation and feature propagation, a strategy based on advanced heterogeneous graph convolution such as HGT, and a strategy based on adaptive learning information transfer using neural architecture search (NAS), to generate diverse representations. Then, a contrastive learning mechanism is introduced to calculate the consistency loss between corresponding node representations generated by different heterogeneous encoding strategies. This is used as a regularization term or auxiliary objective to drive the model to learn more consistent, robust, and discriminative feature representations. Subsequently, an effective feature fusion module (e.g., a deep fusion mechanism based on a Transformer Encoder) was designed to integrate representations from homogeneous views and multiple heterogeneous view representations (or their preliminary fusions) optimized through contrastive learning to generate a final multimodal fused representation. Finally, this high-quality fused representation is input into a prediction model for training and prediction, thereby outputting the association score or probability of the drug-disease pair. This approach overcomes the limitations of related technologies in processing multi-source heterogeneous biological data, learning high-quality node representations, and balancing multi-task optimization, effectively achieving more accurate and reliable drug-disease association prediction.

[0125] This embodiment provides a drug-disease association prediction method that can be executed by a server, terminal, mobile terminal, or other device. The process includes the following steps:

[0126] Step S201, obtaining drug-disease information to be predicted;

[0127] Step S202 : inputting the drug-disease information to be predicted into the drug-disease association prediction model constructed according to the drug-disease association prediction model construction method described in any of the above embodiments to obtain a drug-disease association prediction result to be predicted.

[0128] The present invention provides a drug-disease association prediction method based on multi-view heterogeneous graph information fusion and contrastive representation learning. First, a graph neural network is used to process the homogeneous similarity network between drugs and diseases to extract baseline representations reflecting their intrinsic attribute associations. Second, information is extracted from the heterogeneous association network containing drug, disease, and protein nodes in parallel using three view representation models with different focuses. Among them, the neighborhood view representation model focuses on aggregating the direct, cross-type neighbor information of nodes and combines local structural propagation; the multi-hop view representation model uses advanced heterogeneous graph neural networks to capture the extensive context and long-range dependencies of nodes in the entire heterogeneous network; and the meta-path view representation model uses an adaptive mechanism to mine and encode potential, semantically specific, high-order connection patterns. Subsequently, a contrastive learning mechanism is introduced to force the representations learned for the same node under different heterogeneous views to align with each other to enhance the consistency and robustness of the representation. Next, the representations of the homogeneous similarity view are deeply fused with the representations of the heterogeneous view, and sequence modeling technology is used to extract a more information-rich comprehensive embedding of drugs and diseases. Finally, based on this comprehensive embedding, the downstream prediction model is trained, and the prediction model is ultimately used to output the association probability of the drug-disease pair. This invention significantly enhances the performance of the drug-disease association prediction model and improves the accuracy of predicting potential drug-disease associations by deeply integrating and analyzing the multiple relationships and attribute information between drugs, diseases, and proteins, through deep mining, representation alignment, and effective fusion of multi-perspective information, and by adopting a multi-task gradient balance optimization strategy during the training process. This method also overcomes the limitations of related technologies in processing multi-source heterogeneous biological data, learning high-quality node representations, and balancing multi-task optimization, effectively achieving more accurate and reliable drug-disease association predictions.

[0129] The further description of the construction of the drug-disease association prediction model is detailed in the above corresponding embodiments and will not be repeated here.

[0130] In this embodiment, a device for constructing a drug-disease association prediction model is also provided. The device is used to implement the above-mentioned embodiments and preferred embodiments, and the details that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0131] This embodiment provides a device for constructing a drug-disease association prediction model. Figure 3 As shown, the device includes:

[0132] An acquisition module 301 is used to acquire drug similarity data, disease similarity data, and drug-disease-protein association data;

[0133] Extraction module 302, for extracting representation information of drugs and diseases from different perspectives based on drug similarity data, disease similarity data, drug-disease-protein association data, and a pre-established multi-perspective representation learning model;

[0134] Contrastive learning module 303, used to perform multi-view contrast learning on the representation information to obtain contrast loss information;

[0135] A fusion module 304 is configured to perform heterogeneous fusion based on the representation information to obtain comprehensive drug representation information and comprehensive disease representation information;

[0136] The training module 305 is used to establish a deep learning model and perform parameter training based on the contrast loss information, the comprehensive drug representation information, and the comprehensive disease representation information to obtain a drug-disease association prediction model after training.

[0137] The drug-disease association prediction model construction device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0138] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0139] The embodiment of the present invention also provides a computer device having the above Figure 3 The drug-disease association prediction model building device shown.

[0140] See also Figure 4 , Figure 4 is a structural diagram of a computer device provided by an optional embodiment of the present invention, such as Figure 4 As shown, the computer device includes: one or more processors 10, memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components utilize different buses to communicate with each other and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in the memory or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Equally, multiple computer devices can be connected, and each device provides part of the necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 4 A processor 10 is taken as an example.

[0141] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.

[0142] The memory 20 stores instructions that can be executed by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.

[0143] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0144] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0145] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.

[0146] The embodiment of the present invention also provides a computer-readable storage medium. The above-mentioned method according to the embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.

[0147] A portion of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium that can be accessed by the computer.

[0148] Although the embodiments of the present invention have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention. Such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A method for constructing a drug-disease association prediction model, characterized in that: The method comprises: Obtain drug similarity data, disease similarity data, and drug-disease-protein association data; Extracting representation information of drugs and diseases from different perspectives based on the drug similarity data, the disease similarity data, the drug-disease-protein association data, and a pre-established multi-perspective representation learning model; Perform multi-view contrast learning on the representation information to obtain contrast loss information; Performing heterogeneous fusion based on the representation information to obtain comprehensive drug representation information and comprehensive disease representation information; A deep learning model is established based on the contrast loss information, the comprehensive drug representation information, and the comprehensive disease representation information, and parameter training is performed to obtain a drug-disease association prediction model after training.

2. The method according to claim 1, characterized in that The extracting, based on the drug similarity data, the disease similarity data, the drug-disease-protein association data and a pre-established multi-perspective representation learning model, representation information of drugs and diseases from different perspectives includes: constructing a drug homogeneity graph based on the drug similarity data; constructing a disease homogeneity graph based on the disease similarity data; constructing a heterogeneous graph based on the drug-disease-protein association data; Inputting the drug homogeneity graph and the disease homogeneity graph into a similarity perspective representation model in a multi-perspective representation learning model to obtain drug similarity perspective representation information and disease similarity perspective representation information; Inputting the heterogeneous graph into a neighborhood view representation model in a multi-view representation learning model to obtain neighborhood view drug representation information and neighborhood view disease representation information; Inputting the heterogeneous graph into a multi-hop view representation model in a multi-view representation learning model to obtain multi-hop view drug representation information and multi-hop view disease representation information; The heterogeneous graph is input into a meta-path view representation model in a multi-view representation learning model to obtain meta-path view drug representation information and meta-path view disease representation information.

3. The method according to claim 2, characterized in that The heterogeneous fusion based on the representation information to obtain comprehensive drug representation information and comprehensive disease representation information includes: fusing the neighborhood view drug representation information, the multi-hop view drug representation information, and the meta-path view drug representation information to obtain fused drug representation information; fusing the neighborhood view disease representation information, the multi-hop view disease representation information, and the meta-path view disease representation information to obtain fused disease representation information; Integrating the drug similarity perspective representation information with the fused drug representation information to obtain a drug representation sequence containing two perspectives; Integrating the disease similarity perspective representation information with the fused disease representation information to obtain a disease representation sequence containing two perspectives; The drug representation sequence and the disease representation sequence are respectively input into an encoder network to obtain the comprehensive drug representation information and the comprehensive disease representation information.

4. The method according to claim 1, wherein The step of establishing a deep learning model based on the contrast loss information, the comprehensive drug representation information, and the comprehensive disease representation information and performing parameter training to obtain a drug-disease association prediction model after training includes: obtaining drug-disease samples; determining a primary task loss based on the drug-disease sample, the comprehensive drug representation information, and the comprehensive disease representation information; Constructing a deep learning model loss function based on the main task loss and the contrast loss information; The deep learning model loss function is trained to obtain a drug-disease association prediction model.

5. The method according to claim 4, characterized in that The determining of the main task loss based on the drug-disease sample, the comprehensive drug representation information, and the comprehensive disease representation information includes: Extracting a drug representation vector and a disease representation vector corresponding to the drug-disease sample from the comprehensive drug representation information and the comprehensive disease representation information; Inputting the drug representation vector and the disease representation vector into a preset function to obtain a drug-disease joint representation; Inputting the drug-disease joint representation into a pre-set multi-layer perceptron classifier, and outputting a predicted result of the association between the drug and the disease to be predicted; The main task loss is determined based on the association prediction result and the actual association result.

6. A drug-disease association prediction method, characterized in that: The method comprises: Obtain drug-disease information to be predicted; The drug-disease information to be predicted is input into the drug-disease association prediction model constructed according to the drug-disease association prediction model construction method according to any one of claims 1 to 5 above to obtain a drug-disease association prediction result to be predicted.

7. A device for constructing a drug-disease association prediction model, characterized in that: The device comprises: Acquisition module, used to obtain drug similarity data, disease similarity data, and drug-disease-protein association data; an extraction module, configured to extract representation information of drugs and diseases from different perspectives based on the drug similarity data, the disease similarity data, the drug-disease-protein association data, and a pre-established multi-perspective representation learning model; A contrastive learning module, configured to perform multi-view contrastive learning on the representation information to obtain contrast loss information; A fusion module, configured to perform heterogeneous fusion based on the representation information to obtain comprehensive drug representation information and comprehensive disease representation information; A training module is used to establish a deep learning model and perform parameter training based on the contrast loss information, the comprehensive drug representation information, and the comprehensive disease representation information to obtain a drug-disease association prediction model after training.

8. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the drug-disease association prediction model construction method according to any one of claims 1 to 5 or the drug-disease association prediction method according to claim 6 by executing the computer instructions.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the drug-disease association prediction model construction method according to any one of claims 1 to 5 or the drug-disease association prediction method according to claim 6.

10. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the method for constructing a drug-disease association prediction model according to any one of claims 1 to 5 or the method for predicting drug-disease association according to claim 6.

Citation Information

Patent Citations

  • Drug-disease association prediction method and system based on cross-view comparative learning

    CN119274687A