Drug relocation method fusing automatic meta-path selection

Through the method of fusion metapathic path automatic selection, combined with the self-attention mechanism and the Transformer mutual attention mechanism, the problem of limited accuracy of drug-disease association prediction in the prior art is solved, and more efficient feature expression and association prediction are achieved.

CN120072353AActive Publication Date: 2025-05-30QINGDAO UNIV

Patent Information

Application Number
CN202510552007.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-05-30
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The prior art is difficult to dig deep into the complex relationship between drugs and diseases, resulting in limited accuracy of drug relocation prediction.

Method used

The method of automatic selection of fusion metapaths is adopted to aggregate metapath features through a self-attention mechanism, calculate the cosine similarity of view embeddings, dynamically filter key views, and fuse multi-view embeddings using the Transformer mutual attention mechanism, and finally transmit protein feature information through a heterogeneous graph neural network.

Benefits of technology

It significantly improves the accuracy of drug-disease association prediction, can capture the dependencies and complex interaction characteristics between nodes more accurately, and improves the model's expression and generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072353A_ABST
    Figure CN120072353A_ABST
Patent Text Reader

Abstract

The invention provides a drug relocation method fusing automatic meta-path selection, and relates to the technical field of drug relocation, and the method comprises the steps: taking a drug entity and a disease entity as input through view embedding learned by a first meta-path, and carrying out meta-path feature aggregation based on a self-attention mechanism; calculating the cosine similarity between each view embedding and other view embedding, sorting the importance scores of the views, and taking the first k views to update the disease tensor and the drug tensor; based on a Transform mutual attention multi-view fusion module, multi-view embedding of drugs and diseases is fused; the feature information of the protein is dynamically transmitted to medicine and disease nodes through a heterogeneous graph neural network; embedding of drugs and diseases is achieved, and prediction output of association scores is conducted through an MLP layer. According to the technical scheme, the problem that in the prior art, medicine and disease association cannot be deeply excavated is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of drug repositioning, and particularly to a drug repositioning method that integrates automatic selection of meta-paths. Background Art

[0002] Drug repositioning research plays a key role in the drug development process. Existing drug repositioning prediction methods can be divided into two categories: machine learning methods and deep learning techniques.

[0003] The core of classical machine learning methods lies in the application of matrix completion and network propagation techniques. It is difficult to handle various types of nodes and edges, and cannot fully utilize the rich semantic information in heterogeneous networks, resulting in limited performance in capturing complex relationships and global patterns. Compared with classical learning methods, deep learning techniques have achieved significant improvements and breakthroughs in performance. For example, LAGCN is a graph convolutional network that adopts a hierarchical attention mechanism, which effectively learns the representations of drug or disease nodes through the hierarchical attention mechanism in the heterogeneous drug-disease network. DRHGCN is a graph convolutional network based on information fusion, which designs intra-domain and cross-domain embedding strategies to optimize the representation learning of drug or disease nodes. REDDA combines three mechanisms, namely node embedding, topological subnet embedding, graph attention, and hierarchical attention, to learn the representations of drugs and diseases during the hierarchical process of the heterogeneous graph convolutional network, but the scope of its graph convolution is limited to first-order neighbor nodes. The RGLDR method first defines a set of meta-paths based on various connectivity patterns to simulate the regulatory mechanism between drugs and diseases, performs random walks in the heterogeneous biological network through the defined meta-paths to generate corresponding views, and then uses graph representation learning to learn the embeddings of drugs and diseases from the constructed views. AMDGT calculates the similarity and biochemical information of drugs and diseases, and uses various types of Transformer encoders to predict new drug-disease associations. However, the above methods ignore the complex relationships presented in the interaction between drugs and diseases, which are crucial for predicting drug-disease associations.

[0004] Therefore, there is a need for a drug repositioning method that integrates automatic selection of meta-paths and can deeply explore the associations between drugs and diseases. Summary of the Invention

[0005] The main object of the present invention is to provide a drug repositioning method that integrates automatic selection of meta-paths to solve the problem in the prior art that the associations between drugs and diseases cannot be deeply explored.

[0006] To achieve the above object, the present invention provides a drug repositioning method that integrates automatic selection of meta-paths, specifically including the following steps: S1. Aggregate the meta-path features based on the self-attention mechanism with the view embeddings learned by the drug entity and the disease entity through the first meta-path as the input.

[0007] S2. Calculate the cosine similarity between each view embedding and other view embeddings, sum up the cosine similarities to obtain the importance score of each view, sort the importance scores of the views, and select the top k views to update the disease tensor and the drug tensor.

[0008] S3. Based on the Transformer cross-attention multi-view fusion module, fuse the multi-view embeddings of the drug and the disease.

[0009] S4. Dynamically transfer the feature information of the protein to the drug and disease nodes through the heterogeneous graph neural network.

[0010] S5. For the embeddings obtained for the drug and the disease, predict and output the association scores through the MLP layer.

[0011] Further, taking the aggregation of the original-path features of the drug entity as an example, step S1 specifically includes the following steps: S1.1. Summarize the topological information of adjacent nodes to obtain the view embedding representation of the drug entity learned through the first meta-path: ; where, is the feature of the th node, is the adjacency matrix of the th node, is the total number of nodes, represents the view embedding representation of the drug entity learned through the first meta-path, is the number of drug nodes.

[0012] S1.2. Construct a learnable key-value query triple: ; ; where, , , are the query vector, the key vector, and the value vector respectively; , , are learnable weight matrices, is the dimension of the key vector, is the normalization function, is the normalized view embedding representation of the drug entity learned through the first meta-path.

[0013] S1.3. Set K meta-paths for the drugs to obtain drug tensors of K views: ; Among them, is the drug tensor, is the view embedding representation learned by the normalized drug entity through the K-th meta-path.

[0014] Furthermore, the steps for aggregating the original path features of the disease entities are the same as those in steps S1.1 to S1.3 to obtain disease tensors of K views: ; Among them, is the disease tensor, is the view embedding representation learned by the normalized disease entity through the K-th meta-path, is the number of disease nodes.

[0015] Furthermore, taking the update of the drug tensor as an example, step S2 specifically includes the following steps: S2.1. Calculate the cosine similarity between each view embedding and other view embeddings : ; Among them, and are the -th and -th drug view embeddings respectively, is the feature of the -th node in the drug view .

[0016] S2.2. For , calculate the similarity with all other drug views, calculate the attention score , and normalize the attention score: ; ; Among them, is the normalized attention score.

[0017] S2.3. Sort the obtained in descending order, and take the top , and update the drug tensor to .

[0018] Furthermore, the update steps of the disease tensor are the same as those in steps S2.1 to S2.3, and the disease tensor is updated to :​ 。

[0019] Further, step S3 specifically includes the following steps: S3.1, in , each drug has different vector representations, denoted as ; in , each disease has different vector representations, denoted as .

[0020] S3.2, For each pair of embedding vectors, use query vectors, key vectors, and value vectors for mapping: ; ; where , , , , , are trainable weight matrices; and are the query vectors of the drug view and the disease view respectively; and are the key vectors of the drug view and the disease view respectively; and are the value vectors of the drug view and the disease view respectively.

[0021] S3.3, For each pair of and , by calculating the sum of the cross dot products of and , and , obtain the attention weight , and the process is expressed as: ; ; ; where and are the drug view set and the disease view set respectively, is the transpose of , is the updated drug feature vector, is the updated disease feature vector.

[0022] The fused drug feature , the fused disease characteristics .

[0023] Furthermore, step S4 specifically includes the following steps: S4.1, concatenate the characteristics of drugs, diseases, and proteins as input features Input into the multi-layer heterogeneous graph transformation network HGT, is the protein feature, represents the total number of drug nodes, disease nodes, and protein nodes, is the dimension of the embedding.

[0024] The calculation process of a single layer of HGT is as follows: ; Among them, , respectively represent the mapping of node types and the mapping of edge types; represents the homogeneous graph after the transformation of the heterogeneous graph, represents the matrix of the embedding of the layer.

[0025] S4.2, the output representation of the drug feature is , and the output representation of the disease feature is .

[0026] Furthermore, in step S4, the embeddings obtained for drugs and diseases are passed through the MLP layer to predict the association scores , specifically expressed as: ; Among them, represents the dot product.

[0027] The present invention has the following beneficial effects: In the process of entity feature aggregation at the meta-path level, the present invention introduces a self-attention mechanism to dynamically weight node features to enhance the feature expression ability and accurately capture the dependencies between nodes; by measuring the similarity between views, key view inputs are dynamically screened to optimize the model performance. At the same time, a mutual-attention mechanism is used to fuse multi-view drug and disease embeddings to fully capture and model the complex interaction features between entities, thereby improving the accuracy of feature interaction extraction and drug-disease association prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings: Figure 1 Shows the flowchart of a drug repositioning method that fuses automatic selection of meta-paths of the present invention.

[0029] Figure 2 Shows the result graph of the cold start comparison experiment of the dataset C-dataset.

[0030] Figure 3 Shows the result graph of the cold start comparison experiment of the dataset F-dataset. Specific embodiments

[0031] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the drawings. Obviously, the described embodiments are some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0032] As Figure 1 A drug repositioning method that fuses automatic selection of meta-paths shown, specifically includes the following steps: S1, taking the view embeddings learned by the drug entity and the disease entity through the first meta-path as the input, and based on the self-attention mechanism, performing meta-path feature aggregation.

[0033] S2, calculating the cosine similarity between each view embedding and other view embeddings, summing the cosine similarities, obtaining the importance score of each view, sorting the importance scores of the views, and taking the top k views to update the disease tensor and the drug tensor.

[0034] S3, based on the Transformer cross-attention multi-view fusion module, fusing the multi-view embeddings of drugs and diseases.

[0035] S4, dynamically transmitting the feature information of proteins to drug and disease nodes through the heterogeneous graph neural network.

[0036] S5, for the embeddings obtained for drugs and diseases, predicting and outputting the association scores through the MLP layer.

[0037] Based on the drugs, diseases, and proteins collected in the dataset, a heterogeneous network containing three types of entities is constructed. In the heterogeneous network, the diversity of nodes and edges poses a great challenge to representation learning. Traditional homogeneous network methods are difficult to capture the complex semantic relationships and topological structures in heterogeneous networks. To make full use of the rich information in the heterogeneous network, meta-path, as a high-level semantic pattern, is widely used to describe the composite relationships between nodes. However, existing meta-path methods are often limited to the feature extraction of a single path and fail to fully exploit the synergistic effects among multiple paths. Therefore, the present invention proposes a representation learning framework based on meta-path feature aggregation, which comprehensively captures the potential associations between nodes through multi-level feature fusion and self-attention mechanism.

[0038] Specifically, taking the meta-path feature aggregation of drug entities as an example, step S1 specifically includes the following steps: S1.1, considering the diversity of meta-paths, in the form of to generalize any type of meta-path of drugs, and the starting node drug of the meta-path is represented by to represent. , which represents the transfer process from the starting node to the adjacent node and is defined by the adjacency matrix. Propagate deeply on the meta-path and summarize the topological information of adjacent nodes to obtain the view embedding representation learned by the drug entity through the first meta-path. In this process, each node receives information from its neighbors and fuses this information into its function: ; Among them, is the feature of the th node, is the adjacency matrix of the th node, is the total number of nodes, represents the view embedding representation learned by the drug entity through the first meta-path, is the number of drug nodes.

[0039] S1.2, use the dynamic weight adaptive framework of the self-attention mechanism to deeply explore the multi-granularity semantic associations hidden in the drug nodes in the heterogeneous information network. This framework takes as the input, constructs a learnable key-value query triple, and realizes the accurate modeling of the multi-level topological dependence relationship between nodes: ; ; Among them, , , are the query vector, key vector, and value vector respectively; , , is a learnable weight matrix, is the dimension of the key vector, is a normalization function, is the view embedding representation learned by the normalized drug entity through the first meta-path.

[0040] The dynamic weight adaptive framework of the self-attention mechanism flexibly adjusts the view embedding representation according to the node features. To make the training process stable, the obtained features are normalized.

[0041] S1.3. Set K meta-paths for the drug to obtain the drug tensors of K views: ; where, is the drug tensor, is the view embedding representation learned by the normalized drug entity through the Kth meta-path.

[0042] Specifically, the steps for aggregating the original path features of the disease entity are the same as steps S1.1 to S1.3 to obtain the disease tensors of K views: ; where, is the disease tensor, is the view embedding representation learned by the normalized disease entity through the Kth meta-path, is the number of disease nodes.

[0043] Specifically, to comprehensively capture the potential associations between drugs and diseases, the present invention introduces a multi-view importance evaluation mechanism based on cosine similarity. This mechanism dynamically screens out the most discriminative view representations by quantifying the semantic consistency between different views, thereby effectively suppressing noise interference and eliminating redundant information between views. This design significantly improves the robustness and generalization ability of the model in key feature extraction. Specifically, for the multiple views obtained from multiple meta-paths, the tensors of drug and disease fusion are respectively represented as: and , by calculating the cosine similarity between each view embedding and other view embeddings and summing them, the importance score of each view is obtained. Based on this score, the model adaptively screens out the top k most representative views.

[0044] Taking the update of the drug tensor as an example, step S2 specifically includes the following steps: S2.1. Calculate the cosine similarity between each view embedding and other view embeddings : ; where, and They are the nd and th drug view embeddings, is the feature of the th node in the drug view .

[0045] S2.2. For , calculate the similarity with all other drug views, calculate the attention score , and normalize the attention score: ; ; where is the normalized attention score.

[0046] S2.3. Sort the obtained in descending order, and take the top ones , and update the drug tensor to .

[0047] Specifically, the update steps of the disease tensor are the same as those in steps S2.1~S2.3, and the disease tensor is updated to : .

[0048] Specifically, aiming at the problem of high-order correlation modeling of drug-disease multi-view data, the present invention introduces a new mutual attention mechanism on the Transformer architecture. The used transformer module efficiently fuses the multi-view embeddings of drugs and diseases through shared attention, and further explores the cross-correlation between feature vectors after meta-path aggregation in different view specificities. The model can not only generate independent representations for each view, but also fully capture the complex interaction features between drug and disease entities in the embeddings, thus significantly improving the overall representation learning ability. Specifically, in different views of drugs, each drug has k different vector representations, denoted as . Similarly, for the disease view , there are k different disease vector representations, denoted as . The Transformer mutual attention multi-view fusion module for drug and disease views provides mutual attention for each pair of drug and disease vectors. Step S3 specifically includes the following steps: S3.1. In , each drug has different vector representations, denoted as ; in Among them, each disease has different vector representations, denoted as .

[0049] S3.2. For each pair of embedding vectors, use query vectors, key vectors, and value vectors for mapping: ; ; where , , , , , are trainable weight matrices; and are the query vectors of the drug view and the disease view respectively; and are the key vectors of the drug view and the disease view respectively; and are the value vectors of the drug view and the disease view respectively.

[0050] S3.3. For each pair of and , by calculating and , and , obtain the attention weight , and the process is expressed as: ; ; ; where and are the drug view set and the disease view set respectively, is the transpose of , is the updated drug feature vector, is the updated disease feature vector. is the weighted sum of all the value vectors of the disease view, and a residual connection is set to ensure that drug information can be retained in the network.

[0051] S3.4. The fused drug feature , the fused disease feature . is composed of , is composed of .

[0052] Specifically, step S4 specifically includes the following steps: S4.1, based on multi-view feature fusion, the model dynamically transmits the feature information of proteins to drug and disease nodes through a heterogeneous graph neural network. Specifically, for each view of drugs and diseases, a feature dictionary containing drug, disease, and protein features is constructed respectively and loaded into the node data of the heterogeneous graph. Subsequently, the features of drugs, diseases, and proteins are concatenated as input features Input into the multi-layer heterogeneous graph transformation network HGT, is the protein feature, represents the total number of drug nodes, disease nodes, and protein nodes, is the dimension of the embedding.

[0053] The single-layer calculation process of HGT is: ; Among them, , respectively represent the mapping of node types and the mapping of edge types; represents the homogeneous graph after the transformation of the heterogeneous graph, represents the matrix of the embedding of the th layer.

[0054] Through the multi-layer heterogeneous graph transformation network (Heterogeneous Graph Transformer), the feature information of proteins propagates to drug and disease nodes through the edges of the graph, dynamically updating their feature representations.

[0055] S4.2, the output representation of the drug feature is , and the output representation of the disease feature is . This not only retains the semantic information of drugs and diseases themselves but also fuses the protein features associated with them, thus significantly enhancing the discriminative ability of the features. This process provides higher-quality feature input for drug-disease association prediction.

[0056] Specifically, the embeddings of drugs and diseases obtained in step S4 are passed through the MLP layer to predict and output the association scores , specifically expressed as: ; Among them, represents the dot product.

[0057] In the process of entity feature aggregation at the meta-path level in the present invention, in order to improve the expression ability of node features, a self-attention mechanism is introduced to accurately capture the dependence relationship between nodes and dynamically weight and aggregate features.

[0058] The present invention dynamically adjusts the key view representation input to the next module by measuring the similarity between a single view and all other views, thereby effectively optimizing the expressive ability and performance of the overall model.

[0059] The present invention proposes a mutual attention mechanism, which generates independent representations for each view by comprehensively modeling multi-view drug and disease embeddings, while fully capturing and integrating the complex interaction features between entities.

[0060] To verify the method proposed by the present invention, the following experiments are carried out: Two benchmark datasets are used to evaluate the performance of the model. Each dataset contains three entity types: drugs, diseases, and proteins. The first dataset, C-dataset, includes 2352 known associations between 663 drugs from Drugbank and 409 diseases from the OMIM database, as well as 993 proteins. The second dataset, F-dataset, contains 593 drugs, 313 diseases, and 1933 verified drug-disease associations. This dataset is derived from the research of Gottlieb et al., where drug information is extracted from the Comprehensive Drug Bank (DB) database, which provides rich data on drugs and their targets. Disease data is sourced from the Online Mendelian Inheritance in Man Phenotype Database (OMIM), a public resource dedicated to providing information on human genes and diseases. In the present invention, all unassociated drug and disease pairs are generated as negative samples. Table 1 summarizes the two datasets.

[0061] Table 1 Datasets

[0062] The 10-fold cross-validation framework is adopted to comprehensively evaluate the performance of the MAPTrans-DR model provided by the present invention. Specifically, the positive samples and all negative samples are evenly divided into 10 mutually exclusive and equal-sized subsets. In each round of validation, one of the subsets is used as the test set in turn, and the remaining 9 subsets are used as the training set, so as to ensure that the model can be evaluated on an independent test set in each round, thereby obtaining more robust performance metrics. The optimization process of the model uses the Adam Optimiser, and the learning rate is initialized to 1e-4 to balance the training efficiency and convergence stability. In addition, in the feature representation learning stage, 4 meta-paths (K = 4) are initialized for drugs and diseases, and then through the model's adaptive screening mechanism, the top k most discriminative views (k = 2) are dynamically selected to optimize the feature representation and improve the generalization ability of the model. In addition, the metrics AUC for evaluating the performance of binary classification models and AUPR for evaluating the performance of classification models are introduced to accurately evaluate the performance of the model of the present invention and other baseline models.

[0063] To evaluate the prediction performance of MAPTrans-DR, it was compared with several state-of-the-art baseline models, including NIMCGCN, LAGCN, DRWBNCF, DRHGCN, DRAGNN, and AMDGT. NIMCGCN uses graph convolutional networks to learn latent feature representations from similarity networks, and then inputs the learned features into a neural inductive matrix completion model to generate association matrix completion results. LAGCN constructs a heterogeneous network by integrating known drug-disease associations and uses graph convolution to learn node embedding representations. On this basis, an attention mechanism is introduced to adaptively fuse the structural information of multiple graph convolutional layers. DRWBNCF designs an innovative weighted bilinear graph convolution operation, which is based on the neighborhood interaction neural collaborative filtering framework and aims to efficiently predict DDAs. DRHGCN constructs a model for drug repositioning based on GCN, combining cross-domain and intra-domain feature extraction modules and a hierarchical attention mechanism. DRAGNN uses graph attention mechanisms to integrate similarity networks and drug-disease association networks, enabling node embeddings to have heterogeneous information and neighborhood homogeneous information. AMDGT constructs a comprehensive similarity network of drugs and diseases and uses Transformer to learn node embedding representations in the network.

[0064] The experiments were conducted on two recognized benchmark datasets. The performance of the models was evaluated using 10-fold CV on the datasets, and the specific results are shown in Tables 2 and 3. The experimental results show that MAPTrans-DR is significantly superior to the current state-of-the-art AMDGT model in terms of performance. Specifically, on the C-dataset, MAPTrans-DR improved by 0.9% and 10.21% in the two key metrics of AUC and AUPR, respectively; on the F-dataset, MAPTrans-DR also performed outstandingly, with AUC and AUPR increasing by 0.64% and 9.58%, respectively. The improvement of the metrics fully demonstrates the excellent ability of MAPTrans-DR in drug-disease data feature extraction and feature optimization. Therefore, it is considered that MAPTrans-DR not only has important application value in the drug repositioning prediction task, but also shows great potential in exploring new drug uses and discovering potential therapeutic targets.

[0065] Experimental analysis on the C-dataset shows that by introducing the self-attention mechanism, the MAPTrans-DR model can dynamically capture the global dependencies between nodes, significantly enhancing the comprehensiveness and distinctiveness of feature representation. Meanwhile, MAPTrans-DR optimizes the fusion and transmission of key information by dynamically adjusting multi-view representations, thus achieving remarkable improvements in prediction accuracy and model robustness. In addition, MAPTrans-DR realizes more refined feature aggregation at the meta-path level, can learn more discriminative node embeddings, effectively capture the high-order association patterns between drugs and diseases, and thus performs excellently in modeling complex non-linear relationships. Finally, MAPTrans-DR effectively reduces data noise through multi-view representation learning, and combines with the improved Transformer architecture to design an attention mechanism shared by drugs and diseases, enabling them to co-learn and update their respective embedding representations in a unified semantic space, thus achieving more comprehensive feature modeling and higher prediction performance. In summary, through the meta-path-guided multi-level feature aggregation mechanism, MAPTrans-DR can more efficiently extract the feature representations of nodes, while dynamically optimizing multi-view representations, effectively avoiding information redundancy and noise interference. In addition, MAPTrans-DR uses the Transformer-based multi-view mutual attention fusion mechanism to further enhance the information expression ability of node embeddings, thus significantly improving the prediction accuracy of the model.

[0066] Table 2 Comparison results of key indicators of each model on the C-dataset

[0067] Table 3 Comparison results of key indicators of each model on the F-dataset

[0068] To comprehensively evaluate the performance of the MAPTrans-DR model in predicting potential drugs for new diseases, a leave-one-out cross-validation (LOOCV) experiment based on the C-dataset and F-dataset was designed. The core goal of this experiment is to simulate the real-world scenario where, when faced with a brand-new disease, the model can effectively predict potential therapeutic drugs in the absence of known drug-disease association (DDA) data. Specifically, for each disease in the dataset, all its known DDAs are removed from the training set as test data, while the remaining data is used to train the model. This experimental design can strictly test the generalization ability of the model and its adaptability to unknown data. As Figure 2 and Figure 3As shown, compared with the current state-of-the-art baseline model AMDGT, the model provided by the present invention performs particularly outstandingly on the C-dataset. The two key indicators of AUC and AUPR are respectively improved by 6.83% and 6.91%, demonstrating significant performance advantages. It is worth noting that on the F-dataset, although the model of the present invention is slightly lower than the compared baseline model in terms of the AUC indicator, it still maintains considerable competitiveness in terms of the AUPR indicator. These results indicate that AMDGT has poor comprehensive performance in this task. This is mainly because AMDGT enhances the connectivity of the drug-disease network by introducing protein information to improve performance, but does not fully consider the complexity of the heterogeneous biological network. The model may be difficult to balance the biological attribute information of drugs and diseases during the embedding learning process, resulting in suboptimal performance. In contrast, the model of the present invention shows stronger adaptability and expressive ability when dealing with complex biological networks. It not only aggregates features from the self-attention mechanism at the meta-path level to capture high-order semantic information under different meta-paths, but also collaboratively optimizes multiple modules to achieve unified learning and deep fusion of multi-view embeddings.

[0069] To verify the ability of the MAPTrans-DR model to predict DDA in practical scenarios, a case study experiment was conducted on the C-dataset. The C-dataset has a larger number of drugs and diseases compared with the F-dataset, providing a wider verification space and enabling a more comprehensive evaluation of the model's performance. In this experiment, all known DDAs in the C-dataset were used to construct a training set, and the potential candidate drugs for two highly concerned diseases, Alzheimer’s disease and Lung cancer, were mainly predicted. These two diseases were selected as the research objects mainly based on their wide attention in medical research and the accumulation of rich treatment information. As a representative of neurodegenerative diseases, Alzheimer's disease, and as a typical high-incidence cancer, lung cancer, the integrity and diversity of their research data provide a solid foundation for model verification.

[0070] In Table 4, the top ten potential drugs predicted by MAPTrans-DR for Alzheimer’s disease are listed, and 6 of the candidate drugs are verified from the relevant collected literature. Taking Phenytoin as an example, this is an anticonvulsant drug that plays a neuroprotective role by regulating glutamatergic transmission and may induce beneficial changes in brain structure. Hippocampal atrophy is an early sign of Alzheimer's disease and is closely related to cognitive decline and seizures. Phenytoin delays the progression of Alzheimer's disease and reduces the risk of seizures by inhibiting hippocampal neuron atrophy. Therefore, Phenytoin has potential pharmacological effects in the treatment of Alzheimer's disease. Table 5 lists the top ten potential drugs predicted by MAPTrans-DR for the treatment of lung cancer, and according to the collected literature, 7 of these drugs have been confirmed to be effective in the treatment of lung cancer. Taking Orlistat as an example, Orlistat can reduce the expression of the key molecule GPX4 that regulates cells, induce ferroptosis-like cell death, and thus significantly inhibit the viability of lung cancer cells. This phenomenon scientifically confirms the potential of the MAPTrans-DR model in drug screening. Overall, MAPTrans-DR is a promising tool for discovering new drugs for the treatment of known diseases.

[0071] Table 4 The top ten candidate drugs for Alzheimer's disease predicted by MAPTrans DR

[0072] Table 5 The top ten potential drugs predicted by MAPTrans-DR for the treatment of lung cancer

[0073] The present invention proposes a MAPTrans-DR model. In this model, by introducing a self-attention mechanism to dynamically enhance the node feature representation and combining a multi-level metapath aggregation strategy, a diversified view representation of drugs and diseases is achieved, and the importance of different views is dynamically quantified through a learnable weight assignment mechanism, so as to screen out the most discriminative feature views. The present invention makes an innovative improvement to the Transformer architecture. By designing a cross-view interactive attention mechanism, deep semantic fusion of drugs and diseases in the multi-view space is achieved, so as to learn a more discriminative joint embedding representation, significantly improving the performance and generalization ability of the DDA prediction model.

[0074] Certainly, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by those skilled in the art within the essence of the present invention should also fall within the protection scope of the present invention.

Claims

1. A drug repositioning method integrating automatic selection of metapathways, characterized in that: The specific steps include: S1, takes the view embeddings learned by the drug entity and disease entity through the first meta-path as input, and performs meta-path feature aggregation based on the self-attention mechanism; S2, calculate the cosine similarity between each view embedding and other view embeddings, and sum the cosine similarities to obtain the importance score of each view, sort the importance scores of the views, and take the top k views to update the disease tensor and drug tensor; S3, based on the Transformer mutual attention multi-view fusion module, fuses the multi-view embeddings of drugs and diseases; S4, dynamically transmits protein feature information to drug and disease nodes through heterogeneous graph neural networks; S5, embedding of drugs and diseases is passed through the MLP layer to predict the output of association scores.

2. A drug repositioning method for automatic selection of fusion meta-paths according to claim 1, characterized in that: Taking the aggregation of original path features of drug entities as an example, step S1 specifically includes the following steps: S1.1, summarize the topological information of adjacent nodes and obtain the view embedding representation of the drug entity learned through the first meta-path: ; in, For the The characteristics of the nodes, For the The adjacency matrix of nodes, is the total number of nodes, represents the view embedding representation of the drug entity learned through the first meta-path, is the number of drug nodes; S1.2, construct learnable key-value query triples: ; ; in, , , They are query vector, key vector, and value vector respectively; , , is the learnable weight matrix, is the dimension of the key vector, is the normalization function, is the view embedding representation of the normalized drug entity learned through the first meta-path; S1.3, set K meta-paths for the drug and obtain the drug tensor of K views: ; in, is the drug tensor, is the view embedding representation of the normalized drug entity learned through the Kth meta-path.

3. The drug repositioning method for automatic selection of fusion meta-paths according to claim 2, characterized in that: The steps of aggregating the original path features of the disease entity are the same as steps S1.1 to S1.3, and the disease tensors of K views are obtained: ; in, is the disease tensor, is the view embedding representation of the normalized disease entity learned through the Kth meta-path, is the number of disease nodes.

4. The drug repositioning method for automatic selection of fusion meta-paths according to claim 1, characterized in that: Taking updating the drug tensor as an example, step S2 specifically includes the following steps: S2.1, calculate the cosine similarity between each view embedding and other view embeddings : ; in, and Respectively and drug view embeddings, View for medications Middle The characteristics of each node; S2.2, for , calculate the similarity with all other drug views and calculate the attention score , and normalize the attention score: ; ; in, is the normalized attention score; S2.3, the obtained Sort in descending order and take the top ranking indivual , the drug tensor Updated to .

5. The drug repositioning method for automatic selection of fusion meta-paths according to claim 4, characterized in that: The updating steps of the disease tensor are the same as steps S2.1 to S2.

3. The disease tensor is updated as : 。 6. The drug repositioning method for automatic selection of fusion meta-paths according to claim 1, characterized in that: Step S3 specifically includes the following steps: S3.1, in Each drug has Different vector representations, denoted by ;exist Each disease has Different vector representations, denoted by ; S3.2, for each pair of embedding vectors, use the query vector, key vector and value vector for mapping: ; ; in, , , , , , is a trainable weight matrix; and are the query vectors for drug view and disease view respectively; and are the key vectors for drug view and disease view respectively; and are the value vectors for drug view and disease view respectively; S3.3, for each pair and , by calculating and , and The sum of the cross products of , the process is expressed as: ; ; ; in, and They are respectively a drug view set and a disease view set. for The transpose of is the updated drug feature vector, is the updated disease feature vector; S3.4, Drug features after fusion , the fused disease features .

7. The drug repositioning method for automatic selection of fusion meta-paths according to claim 1, characterized in that: Step S4 specifically includes the following steps: S4.1, concatenating drug, disease, and protein features as input features Input multi-layer heterogeneous graph transformation network HGT, For protein characteristics, Represents the total number of drug nodes, disease nodes, and protein nodes, is the embedding dimension; The HGT single-layer calculation process is: ; in, , Represent the mapping of node types and edge types respectively; represents the homogeneous graph after the heterogeneous graph is transformed. Indicates The matrix of layer embeddings; S4.2, drug feature output is represented as , the output of disease features is expressed as .

8. The drug repositioning method for automatic selection of fusion meta-paths according to claim 1, characterized in that: The embeddings of drugs and diseases obtained in step S4 are output through the MLP layer to predict the association scores , specifically expressed as: ; in, Represents the dot product.

Citation Information

Patent Citations

  • Text recommendation method and related equipment

    CN110866106A

  • Drug ATCCode prediction method based on graph transformation network

    CN114420310A

  • Drug relocation method and system based on hypergraph convolutional neural network

    CN115527627A

  • Drug relocation method and system based on multi-task learning and deep cross-domain

    CN116453618A

  • Circular RNA-drug association prediction method based on HGT and random autoencoder

    CN116741308A

Cited By

  • Drug-target correlation prediction method based on hierarchical representation learning framework

    CN120431993A

  • Drug relocation method based on two-channel graph comparative learning

    CN121964192A

  • A drug repositioning method based on dual-channel graph contrastive learning

    CN121964192B