Drug-disease interaction prediction method, device, medium and product
By extracting and fusing the characteristics of drugs, diseases and targets, and using a fully connected neural network for prediction, the problem of insufficient accuracy and efficiency of drug-disease interaction prediction in the prior art is solved, and more efficient and accurate prediction effects are achieved.
Patent Information
- Application Number
- CN202510115428.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
Existing drug-disease interaction prediction methods have shortcomings in accuracy and computational efficiency, and it is difficult to effectively integrate complementary information in multi-source heterogeneous data.
By obtaining drug data, disease data and biological data, extracting drug molecular characteristics, disease characteristics and target characteristics, and using a fully connected neural network for feature fusion and prediction, constructing drug biological network characteristics and disease characteristics, and finally training the prediction model to improve prediction accuracy.
It improves the accuracy and efficiency of drug-disease interaction prediction, successfully solves the problem of insufficient feature expression, and enhances the model's expression ability and calculation speed.
Smart Images

Figure CN120048332A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent medical technology, and particularly to a method, device, medium and product for predicting drug-disease interactions. Background Art
[0002] Drug-Disease Interactions (DDIs) prediction is a key link in new drug research and development and drug repositioning. In current DDIs prediction research, three main technical routes have been formed. The first is the method based on matrix factorization, which decomposes the drug-disease association matrix into low-dimensional matrices and predicts potential associations through matrix reconstruction. Typical algorithms include non-negative matrix factorization and probabilistic matrix factorization, etc. This type of method has high computational efficiency, but it is difficult to fully utilize the biological feature information of drugs and diseases. The second is the method based on traditional machine learning. This type of method first constructs feature vectors of drugs and diseases, and then uses classical machine learning algorithms to establish a prediction model. The advantage of this method is that the model is highly interpretable, but its performance largely depends on manually designed features. The third is the method based on deep learning. This type of method can automatically learn feature representations and construct an end-to-end prediction model. Compared with traditional methods, deep learning methods can better capture the non-linear relationships in the data, but often require a large amount of training data and computing resources.
[0003] Overall, existing methods mainly rely on a single feature extraction strategy and fail to fully integrate the complementary information in multi-source heterogeneous data, resulting in limited feature representation ability. When constructing a heterogeneous network, simple splicing or fusion strategies are often adopted, ignoring the internal correlation between different types of data and affecting the expression ability of the model. Therefore, the accuracy of existing methods in predicting drug-disease interactions is generally not ideal. In addition, existing methods have high computational complexity when dealing with large-scale datasets, especially when constructing a heterogeneous network and extracting features, which requires a large amount of computing resources. This leads to a long model training time and difficulty in quickly responding to the update requirements of new data. In practical applications, a single prediction task may take several hours or even longer to complete. Summary of the Invention
[0004] Aiming at the problems pointed out in the above background art section, the purpose of this application is to provide a method, device, medium and product for predicting drug-disease interactions to improve the efficiency and accuracy of drug-disease interaction prediction.
[0005] To achieve the above purpose, this application provides the following solutions.
[0006] In the first aspect, this application provides a method for predicting drug-disease interactions, including:
[0007] Obtain drug data, disease data, and biological data; the drug data includes drug molecular structure data; the disease data includes disease phenotype information and disease ontology relationships; the biological data includes protein interaction data and cell signaling pathway data;
[0008] Preprocess and extract features from the drug data to obtain drug molecular features;
[0009] Extract target features based on the biological data;
[0010] Input the drug molecular features and target features into a first fully connected neural network to predict the drug-target interaction relationship;
[0011] Generate drug biological network features based on the drug-target interaction relationship;
[0012] Preprocess and extract features from the disease data to obtain disease features;
[0013] Train a second fully connected neural network based on the drug biological network features and disease features, and after training, obtain a drug-disease interaction prediction model;
[0014] Use the drug-disease interaction prediction model to predict the drug-disease interaction.
[0015] Optionally, the preprocessing and feature extraction of the drug data to obtain drug molecular features specifically includes:
[0016] For the drug molecular structure data, extract its molecular descriptors in SMILES format and use the RDKit toolkit to calculate the molecular fingerprint as the drug molecular graph structure;
[0017] Adopt the InfoGraph algorithm to process the drug molecular graph structure and extract the drug molecular features.
[0018] Optionally, the extraction of target features based on the biological data specifically includes:
[0019] Construct a directed protein interaction network based on the protein interaction data and cell signaling pathway data; in the directed protein interaction network, each node represents a protein, each edge represents the interaction relationship between two proteins, and the direction of the edge indicates the activation or inhibition effect of one protein on another protein;
[0020] Based on the directed protein interaction network, adopt the Node2vec algorithm to extract the target features.
[0021] Optionally, the drug-disease interaction prediction method further includes:
[0022] Obtain gene expression data from the TCGA database;
[0023] Use the gene expression data to identify differentially expressed genes;
[0024] Utilize weighted gene co-expression network analysis to analyze the correlation between differentially expressed genes and generate a disease-specific gene expression correlation matrix;
[0025] By setting the correlation threshold of the disease-specific gene expression correlation matrix, retain significantly correlated gene pairs and delete uncorrelated gene pairs;
[0026] Construct a weighted undirected graph based on the retained significantly correlated gene pairs as the disease-specific network;
[0027] Optimize the directed protein-protein interaction network based on the disease-specific network.
[0028] Optionally, generating drug biological network features based on drug-target interaction relationships specifically includes:
[0029] Based on drug-target interaction relationships, use a random walk strategy in the directed protein-protein interaction network to simulate the interaction of drugs on the biological network; during each random walk, the activation or inhibition effect of the drug on the target will affect the value of the node, and at the same time, the protein-protein interaction relationship will determine the update direction and intensity of the node value; multiple random walk results are integrated into a multi-dimensional vector to form drug biological network features.
[0030] Optionally, preprocessing and feature extraction of the disease data to obtain disease features specifically includes:
[0031] Integrate the disease phenotype information in the OMIM database and the disease ontology relationship in the MeSH database, and use a graph embedding algorithm to represent disease-related terms as nodes and represent the relationships between nodes as edges, thereby generating a disease relationship network;
[0032] Apply the Node2vec algorithm to process the disease relationship network and extract disease features.
[0033] Optionally, training a second fully connected neural network based on the drug biological network features and disease features, and obtaining a drug-disease interaction prediction model after training specifically includes:
[0034] Construct a second fully connected neural network including an input layer, a hidden layer, and an output layer;
[0035] Construct a dataset based on drug biological network features, disease features, and drug-disease interaction relationship data provided by the CTD database, and use the Synthetic Minority Over-sampling Technique (SMOTE) to balance the class imbalance problem in the dataset;
[0036] Use the dataset to train a second fully connected neural network. During the training process, use batch normalization and a dropout rate of 50%, and adopt Leaky ReLU as the activation function. The loss function uses binary cross-entropy loss, and the optimizer uses Adam. The learning rate and weight decay are both set to 0.001;
[0037] After training is completed, a drug-disease interaction prediction model is obtained; the input of the drug-disease interaction prediction model is drug biological network features and disease features, and the output is the drug-disease interaction score; a score greater than 0 indicates relevance, otherwise it is irrelevant, and the higher the score, the greater the likelihood of the association between the drug and the disease.
[0038] In a second aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the drug-disease interaction prediction method.
[0039] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the drug-disease interaction prediction method is implemented.
[0040] In a fourth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the drug-disease interaction prediction method is implemented.
[0041] According to the specific embodiments provided by the present application, the following technical effects are disclosed in the present application.
[0042] A method, device, medium and product for predicting drug-disease interactions provided by the present application extract drug molecular features, disease features and target features respectively based on the obtained drug data, disease data and biological data, and can extract valuable features at a relatively low computational cost, reducing the burden of high-dimensional feature calculation; a first fully connected neural network is used to fuse drug molecular features and target features to realize the prediction of drug-target interaction relationships, and further realize the extraction of drug biological network features, and then a second fully connected neural network is used to realize the fusion and association prediction of drug biological network features and disease features. This method successfully solves the problem of insufficient feature expression, effectively improves the expression ability of the model, enables it to better capture the complex relationships among drugs, targets and diseases, and thus improves the prediction accuracy. And by sharing the parameters in the fully connected neural network, repeated calculations and redundant storage are avoided, and the memory usage efficiency and calculation speed are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0044] Figure 1 is a schematic flowchart of a method for predicting drug-disease interactions according to the present application;
[0045] Figure 2 is a schematic diagram of the main process of a method for predicting drug-disease interactions according to the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.
[0047] The purpose of the present application is to propose a method, device, medium and product for predicting drug-disease interactions to improve the efficiency and accuracy of drug-disease interaction prediction.
[0048] In order to make the above objects, features and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0049] In an exemplary embodiment, as Figure 1As shown, a method for predicting drug-disease interactions is provided, including the following steps 1 to 8.
[0050] Step 1: Obtain drug data, disease data, and biological data.
[0051] The data obtained in this application includes drug data, disease data, and biological data. In terms of drug data, drug molecular structure data is collected from the DrugBank database. Drug molecular structure data refers to the chemical structure information of drugs, usually represented in formats such as SMILES, which contains atoms, bonds, chemical environments, etc. For disease data, mainly obtain disease phenotype information in the OMIM database and disease ontology relationships in the MeSH database. Among them, disease phenotype information refers to the clinical symptoms and manifestations related to diseases, which helps to describe how diseases affect the physiological and physical functions of patients. Disease ontology relationships define the relationships between different diseases through classification and hierarchical structures. For example, some diseases belong to larger disease categories.
[0052] In terms of biological data, protein interaction data in the STRING database, cell signaling pathway data in the KEGG database, and gene expression data in the TCGA database are collected and sorted. Among them, protein interaction data refers to the interaction relationships between different protein molecules. These interactions are important bases for biological processes within cells and can reveal how cells complete various functions, such as signal transduction and metabolism. Through the protein interaction network (such as the data provided by the STRING database), the cooperative effects of proteins in cells can be understood. Cell signaling pathway data describes the ways in which molecules (such as proteins, RNAs, and metabolites) interact within cells through a series of biochemical reactions and signal transduction processes. The KEGG database provides comprehensive biological pathway data, covering various biological processes such as metabolism, immunity, and cancer. Gene expression data refers to the expression levels of genes in specific cell types or conditions, which reflects gene transcriptional activity. Usually, gene expression data provided by databases such as TCGA is collected and analyzed. These data help to study the molecular mechanisms of diseases and their potential therapeutic targets.
[0053] Step 2: Preprocess and extract features from the drug data to obtain drug molecular features.
[0054] First, preprocess the drug molecular structure data and extract its molecular fingerprint as the drug molecular graph structure. Then, extract features from the drug molecular graph structure to obtain drug molecular features. The specific steps of Step 2 include:
[0055] Step 2.1: For the drug molecular structure data, extract its molecular descriptors in SMILES format and use the RDKit toolkit to calculate the molecular fingerprint as the drug molecular graph structure.
[0056] As described above, drug molecular structure data refers to the chemical structure information of drugs, usually represented in formats such as SMILES, which includes atoms, bonds, chemical environments, etc. By extracting molecular descriptors from these structure data, some chemical and physical properties of molecules can be quantified, such as topological, geometric, and electronic properties. Molecular fingerprints, on the other hand, are further simplified representations extracted from molecular structures, which encode the presence or absence of certain specific structural units (such as rings, bonds, groups, etc.) in the molecule in the form of a fixed-length binary vector or integer vector. These molecular fingerprints effectively represent the key features of the molecular graph structure, facilitating efficient structural similarity analysis. In actual calculations, molecular fingerprints are usually automatically generated from molecular structures by specialized tools (such as RDKit), which create fingerprints by detecting specific local structures in the molecule, thus simplifying the complex molecular structure into an easily processed form.
[0057] Step 2.2: Process the drug molecular graph structure using the InfoGraph algorithm to extract drug molecular features.
[0058] For the extraction of drug molecular features, the above-mentioned drug molecular graph structure is processed using the InfoGraph algorithm, with the feature dimension set to 300 dimensions, and finally a 300-dimensional vector is output to represent the drug molecular features. Among them, the InfoGraph algorithm has significant advantages in the field of graph neural networks (GNNs) and graph embedding, especially suitable for processing complex drug molecular graph structures. Specifically, the InfoGraph algorithm first converts the drug molecule into a graph structure, where atoms are nodes and chemical bonds are edges. Then, a graph convolutional network (GCN) is used to embed the graph, and through multi-level information transmission and relationship learning between nodes, the representation of the graph is optimized. Next, the InfoGraph algorithm obtains the embedding representation of each node by maximizing node information sharing and graph global consistency, and finally generates a low-dimensional feature vector of the entire drug molecular graph, usually 300 dimensions, as the extracted drug molecular features. These low-dimensional feature vectors effectively capture the structural features and interrelationships of drug molecules and can be used for subsequent machine learning models or drug screening tasks.
[0059] Step 3: Extract target features based on biological data.
[0060] The biological data obtained in this application can include protein-protein interaction data in the STRING database, cell signaling pathway data in the KEGG database, and gene expression data in the TCGA database. Target features can be accurately extracted based on the above biological data. The specific steps of Step 3 include:
[0061] Step 3.1: Construct a directed protein-protein interaction network based on protein-protein interaction data and cell signaling pathway data. In the directed protein-protein interaction network, each node represents a protein, each edge represents the interaction relationship between two proteins, and the direction of the edge indicates the influence of one protein on another protein (such as activation or inhibition). Among them, the activation or inhibition relationship between proteins is determined by protein-protein interaction data. At the same time, the cell signaling pathway data in the KEGG database provides the role of proteins in the intracellular signal transduction process, further clarifying the direction of the interaction between proteins. The constructed directed network (i.e., the directed protein-protein interaction network) consists of nodes (proteins) and directed edges (activation or inhibition relationships), and is used to represent the action path of proteins in the cell process.
[0062] Step 3.2: Extract target features using the Node2vec algorithm based on the directed protein-protein interaction network.
[0063] The directed protein-protein interaction network is a directed graph that represents the relationship between proteins and their influence directions; this network is used to reveal the potential interactions between drugs and target proteins and to help identify the biological processes and pathways that drugs may act on. Based on the constructed directed protein-protein interaction network, a large number of random walk sequences are generated using the Node2vec algorithm. Subsequently, the Node2vec algorithm is applied to derive the compact vector representation of these nodes to represent the semantic similarity between protein nodes, generating a 300-dimensional vector to represent the target features.
[0064] Specifically, the Node2vec algorithm performs random walks through the nodes in the network to generate a series of random walk sequences, which capture the local and global structural information between the nodes. The generated random walk sequences can be regarded as the random paths of the nodes in the graph, reflecting the similarity and relationship between the nodes. In the subsequent steps, the Word2vec algorithm is applied to train these random walk sequences to obtain the compact vector representation of each node. The "compact vector representation" refers to the low-dimensional vector learned by the Word2vec algorithm through context information, which can represent the semantic similarity between the nodes. The generated target features (300-dimensional vectors) will provide useful inputs for subsequent drug-target association prediction.
[0065] In terms of predicting disease-specific drug-target associations, the present application also innovatively introduces a disease-specific network, optimizes the directed protein interaction network based on the disease-specific network, and can extract disease-specific target features. Therefore, step 3 may further include step 3.3: optimizing the directed protein interaction network based on the disease-specific network. The disease-specific network is constructed based on gene expression data and weighted gene co-expression network analysis (WGCNA), aiming to identify gene expression patterns related to specific diseases and form a weighted undirected graph by calculating the gene expression correlation matrix. This step 3.3 can optimize the performance of the model. If there is no disease-related gene expression data, this step can be skipped.
[0066] Step 3.3 specifically includes:
[0067] Step 3.3.1: Obtain gene expression data from the TCGA database. Gene expression data refers to the expression level of genes in specific cell types or conditions, reflecting gene transcriptional activity, and is usually collected and analyzed through gene expression data provided by databases such as TCGA. These data help to study the molecular mechanisms of diseases and their potential therapeutic targets.
[0068] Step 3.3.2: Use the gene expression data to identify differentially expressed genes. Specifically, genes related to diseases are identified through differential expression analysis, and these genes are called differentially expressed genes (DEGs), and their expression levels differ significantly between the diseased and healthy groups.
[0069] Step 3.3.3: Use weighted gene co-expression network analysis to analyze the correlation between differentially expressed genes and generate a disease-specific gene expression correlation matrix.
[0070] Step 3.3.4: By setting the correlation threshold of the disease-specific gene expression correlation matrix, retain significantly correlated gene pairs and delete uncorrelated gene pairs. In an exemplary embodiment, the correlation threshold is set to 0.7 to ensure that only highly correlated gene pairs are retained.
[0071] Step 3.3.5: Construct a weighted undirected graph based on the retained significantly correlated gene pairs as the disease-specific network.
[0072] Construct a weighted undirected graph, where the nodes represent the retained significantly correlated genes, and the weight of the edge represents the strength of the correlation between genes. This disease-specific network helps to reveal the molecular mechanisms of diseases by reflecting the interrelationships of disease-specific genes, and thus provides support for drug target screening and precision medicine.
[0073] Step 3.3.6: Optimize the directed protein interaction network based on the disease-specific network.
[0074] The network optimization applies the above-mentioned disease-specific gene expression correlation matrix, retains significantly correlated protein pairs in the directed protein interaction network, deletes uncorrelated protein pairs, simulates the real situation of protein interactions in diseases or tissues, and optimizes to obtain a disease-specific protein interaction network. The optimized network can more accurately reflect the biological characteristics of the disease. Then, by reapplying the method in step 3.2, disease-specific target characteristics can be extracted. Since most drug targets are proteins, for example, when an antibody is a drug, the antigen is the target, and most antigens are proteins, so sometimes the target and protein are used interchangeably. Also, because proteins are encoded by genes and most drugs affect protein function by influencing the genes behind the proteins, it can also be considered that target = protein = gene.
[0075] Step 4: Input the drug molecular characteristics and target characteristics into the first fully connected neural network to predict the drug-target interaction relationship.
[0076] Integrate the drug molecular characteristics and target characteristics, and use the drug-target correlation data in the Drugbank database as a training set together. Use the fully connected neural network to predict the drug-target correlation, that is, the drug-target interaction relationship. Drug-target correlation data refers to the association information between a drug and its target, usually obtained through experimental data or databases (such as DrugBank), indicating the action intensity or affinity of a certain drug for a specific target, and is divided into activation, inhibition, and irrelevant relationships. The connection with drug molecular characteristics is that drug molecular characteristics reflect the chemical and structural properties of drugs, and these characteristics will affect the binding ability of drugs to targets. Therefore, drug molecular characteristics provide important inputs for the prediction of target correlation data. The characteristics of the target are extracted through the protein interaction network, and through the extraction of node characteristics in the network, the biological characteristics and functions of the target are revealed. The drug molecular characteristics and target characteristics are used as inputs for the fully connected neural network, and the drug-target correlation data is output. The fully connected neural network is a deep learning model that processes input data through multiple fully connected layers. The input of the first fully connected neural network is the drug molecular characteristics and target characteristics, and the output is the association prediction between the drug and the target, that is, the drug-target interaction relationship, which can predict the potential interaction between the drug and the target.
[0077] Step 5: Generate drug biological network characteristics based on the drug-target interaction relationship.
[0078] Based on the drug-target interaction relationship, the random walk strategy is used to simulate the interaction of drugs on the biological network in the directed protein interaction network, generating a drug biological network feature composed of vectors with a dimension of 18,977. Among them, the random walk strategy is used to simulate the interaction of drugs in the biological network (here refers to the directed protein interaction network). For a single drug, its multiple targets will be used as multiple starting points for random walks, and each target is an independent starting node, which can effectively capture the influence of the drug on its multiple targets. For drug combinations, independent random walks will be performed for each drug separately, and the initial value will not be reset between different drugs, which can reflect the synergistic effect between drug combinations. Specifically, the random walk strategy simulates the influence of drugs on targets by gradually updating the features of each node in the network. During each random walk, the effect of the drug on the target (activation or inhibition) will affect the value of the node, and the interaction relationship between proteins will determine the update direction and intensity of the node value. The change of each step size of the walk is controlled by two key coefficients: the restart coefficient α controls the probability of returning to the starting node at each step, and the attenuation coefficient β controls the attenuation speed when the influence propagates in the network. Each walk generates a path, and these paths reflect how the drug spreads its influence through the network. Finally, multiple random walk results are integrated into a multi-dimensional vector with a dimension of 18,977, and this vector represents the characteristics of the drug in the entire biological network, called the drug biological network feature. In this way, the influence of drugs on the biological network can be comprehensively captured, providing strong support for subsequent drug function analysis, mechanism research, and drug screening.
[0079] If there is disease-related gene expression data, the above-mentioned disease-specific gene expression correlation matrix can also be applied. By setting a correlation threshold, significant protein pairs are retained in the above biological network, and irrelevant protein pairs are deleted to simulate the real situation of protein interactions in diseases or tissues, optimizing the biological network. The optimized biological network can more accurately reflect the biological characteristics of diseases. Then, the above method is used to construct a vector representation of the disease-specific drug biological network with a dimension of 18,977, called the disease-specific drug biological network feature. Combining the biological network features of drugs with disease-specific features can further improve the performance of subsequent drug-disease interaction prediction models.
[0080] Step 6: Preprocess and extract features from the disease data to obtain disease features.
[0081] For disease data, this application integrates the disease phenotype information in the OMIM database and the disease ontology relationships in the MeSH database, uses graph embedding algorithms to represent disease-related terms, and converts the MeSH structure tree into a disease relationship network. Then, based on the disease relationship network, feature extraction is performed to extract disease features.
[0082] Step 6 specifically includes:
[0083] Step 6.1: Integrate the disease phenotype information in the OMIM database and the disease ontology relationships in the MeSH database, and use graph embedding algorithms to represent disease-related terms as nodes and the relationships between nodes as edges, thereby generating a disease relationship network.
[0084] Among them, disease phenotype information refers to the clinical symptoms and manifestations related to diseases, which helps to describe how diseases affect the physiological and physical functions of patients. Disease ontology relationships define the relationships between different diseases through classification and hierarchical structures. For example, some diseases belong to larger disease categories. Disease-related terms are standardized vocabularies used to identify diseases and their characteristics to ensure information consistency. The MeSH structure tree is a hierarchical medical subject term ontology system that helps to construct a network of disease concepts by organizing the parent-child relationships of diseases and other medical concepts. On this basis, by integrating the disease phenotype information in the OMIM database and the disease ontology relationships in the MeSH database, graph embedding algorithms can be used to represent disease-related terms as nodes and the relationships between nodes as edges, thereby generating a disease relationship network. This disease relationship network is a graph structure based on diseases and their characteristics, where nodes represent different diseases or medical terms, and edges represent the relationships between them. The input is the phenotype, ontology information, and hierarchical relationships of diseases, and the output is a structured relationship network for further disease prediction, analysis, and similarity search, providing support for subsequent drug discovery and drug-disease interaction prediction models.
[0085] Step 6.2: Apply the Node2vec algorithm to process the disease relationship network and extract disease features.
[0086] The extraction of disease features is based on the disease relationship network constructed from the above MeSH structure tree. A large number of walk sequences are generated using the Node2vec algorithm. Subsequently, the Node2vec algorithm is applied to derive the compact vector representations of these nodes to represent the semantic similarity between disease nodes, generating a 300-dimensional vector to represent disease features. Among them, the MeSH structure tree is a hierarchical medical concept network, where nodes represent different diseases or medical terms, and edges represent the semantic relationships between them. The Node2vec algorithm generates a series of walk sequences by randomly walking through the nodes in the disease relationship network, and these sequences capture the local and global structural information between the nodes. The generated walk sequences can be regarded as the random paths of nodes in the network, reflecting the similarity and relationships between the nodes. In the subsequent steps, the Word2vec algorithm is applied to train these walk sequences to obtain the compact vector representation of each node. Here, "nodes" refer to each disease or related term in the disease relationship network, and "compact vector representation" refers to the low-dimensional vector learned by the Word2vec algorithm through context information, which can represent the semantic similarity between nodes. The generated disease feature vectors (300-dimensional) will provide useful input for subsequent drug-disease association prediction.
[0087] Step 7: Train a second fully connected neural network based on the drug biological network features and disease features. After training is completed, a drug-disease interaction prediction model is obtained.
[0088] Step 7.1: Construct a second fully connected neural network including an input layer, hidden layers, and an output layer.
[0089] The specific structure of the second fully connected neural network is a 6-layer fully connected neural network (FCNN). The network structure includes an input layer, 4 hidden layers, and an output layer. The dimension of the input layer is 19,277, representing the vector concatenation result of the drug biological network features and disease features, which contains the comprehensive information of the drug biological network features and disease features. The hidden layers contain 4096, 1024, 256, and 64 neurons respectively. The output layer contains 1 neuron, which outputs a drug-disease matching score, that is, the drug-disease interaction score. A score greater than 0 indicates that there is an interaction between the drug and the disease, and a score less than 0 indicates that there is no interaction.
[0090] Step 7.2: Construct a dataset based on the drug biological network features, disease features, and drug-disease interaction relationship data provided by the CTD database, and use the Synthetic Minority Over-sampling Technique (SMOTE) to balance the class imbalance problem in the dataset.
[0091] The interaction prediction model between drugs (drug pairs) and diseases utilizes these extracted disease features and drug biological network features, combines the relationship between drug components and diseases in the CTD database (drug-disease interaction relationship data) as the training set dataset, and predicts the interaction between drugs (drug pairs) and diseases by training the second fully connected neural network. In a specific embodiment, the number of drugs, diseases, and their combinations in the dataset is shown in Table 1. The data collected in this application is huge, and prior knowledge can be fully utilized.
[0092] Table 1 The number of drugs, diseases, and their combinations in the dataset.
[0093] Drug Disease Drug-disease interaction Drug-disease pair 7,940 2,986 88,161 23,708,840
[0094] In terms of dataset processing, this application uses the Synthetic Minority Over-sampling Technique (SMOTE) to balance the class imbalance problem in the dataset. The SMOTE method oversamples the minority class (known drug-disease interactions) samples to generate new samples similar to the original minority class samples. The generated new samples are produced by interpolating the minority class samples in the feature space, ensuring that the new samples have similar features to the original samples but with slight changes in the multi-dimensional space. These generated minority class samples are added to the training set, thereby balancing the ratio of minority class and majority class samples in the dataset and avoiding over-biasing towards the majority class during the training process. The balance ratio refers to the ratio of minority class samples to majority class samples. In this method, the number of minority class samples is increased to 10 times the number of majority class samples through the SMOTE technique, ensuring that the dataset is more balanced during training and helping to improve the prediction performance of the model. The number of positive and negative samples in the original dataset and the dataset processed by the SMOTE method is shown in Table 2.
[0095] Table 2 The number of positive and negative samples in different datasets
[0096] Training set (original) Training set (SMOTE) Validation set Positive sample 70,553 705,264 17,595 Negative sample 705,264 705,264 176,359
[0097] The results shown in Table 2 indicate that the number of positive and negative samples in the original dataset collected in this application is unbalanced, and the SMOTE method can effectively balance the number of positive and negative samples. Further, cross-validation is performed using random sampling and SMOTE sampling in 1-fold and 10-fold datasets. The results show that the practice of downsampling negative samples to 10 times the number of positive samples before using the SMOTE method in this application can be closer to the real-world situation, and the trained model performs better.
[0098] Step 7.3: Train the second fully connected neural network using the dataset. During the training process, use batch normalization and a dropout rate of 50%, adopt Leaky ReLU as the activation function, use binary cross-entropy loss function as the loss function, use Adam as the optimizer, and set both the learning rate and weight decay to 0.001.
[0099] The disease features and drug biological network features obtained in the previous steps are used as input features in the drug-disease interaction prediction model, and the output is the predicted interaction score between the drug and the disease. The combination of these features and data helps the model understand the complex relationship between drugs and diseases from different levels, thereby improving the prediction accuracy. To prevent overfitting, batch normalization and a dropout rate of 50% are used, Leaky ReLU is adopted as the activation function, binary cross-entropy loss function (BCE With LogitsLoss) is used as the loss function, Adam is used as the optimizer, and both the learning rate and weight decay are set to 0.001.
[0100] Step 7.4: After the training is completed, obtain the drug-disease interaction prediction model; the input of the drug-disease interaction prediction model is the drug biological network feature and the disease feature, and the output is the drug-disease interaction score; a score greater than 0 indicates relevance, otherwise it is irrelevant, and the higher the score, the greater the likelihood of the drug being associated with the disease.
[0101] Step 8: Use the drug-disease interaction prediction model to predict drug-disease interactions.
[0102] The drug-disease interaction prediction model receives the 18,977-dimensional drug biological network feature of the drug components collected from the above-mentioned DrugBank database and the disease feature extracted from the MeSH network as input, passes through the above-mentioned drug-disease interaction prediction model, predicts all drug-disease pairs, and outputs the prediction results of drug-disease associations. A drug-disease interaction score greater than 0 indicates relevance, otherwise it is irrelevant, and the higher the score, the greater the likelihood of the drug being associated with the disease.
[0103] As Figure 2As shown in the figure, this application uses the InfoGraph method to extract drug molecular features, and the Node2vec method to extract target features and disease features. The extracted drug molecular features and target features are input into a fully connected neural network (the first fully connected neural network) for predicting drug-target interactions. Subsequently, genes perform random walks on a directed protein interaction network to extract drug biological network features. Finally, through another fully connected neural network (the second fully connected neural network), the interaction between drugs (drug pairs) and diseases is predicted using drug biological network features and disease features. The drug-disease interaction prediction model trained in this application demonstrates excellent performance in practical applications. In the single-drug prediction task, the model achieved an AUC value of 0.93, an accuracy of 0.89, and an F1 score of 0.6316. In terms of drug combination (drug pair) prediction, an F1 score of 0.7746 and an accuracy of 0.82 were obtained by fine-tuning with only 0.032% of the training data. In disease-specific prediction, the method of this application was verified in 18 cancer types, and the prediction performance of 13 cancers was significantly improved, with an average accuracy improvement of 15%.
[0104] A drug-disease interaction prediction method proposed in this application realizes the accurate prediction of drug-disease interactions through data preprocessing, feature extraction, network construction, model training, and result verification. The core of the method of this application lies in constructing a drug-disease interaction prediction model by integrating multi-source heterogeneous data, thereby improving the prediction accuracy. This application can not only predict the interaction between a single drug and a disease but also predict the therapeutic effect of drug combinations, providing strong support for personalized medicine.
[0105] Specifically, to address the problem of insufficient prediction accuracy, this application introduces biological prior knowledge to enhance prediction reliability, designs a fully connected neural network to enhance feature extraction capabilities, and improves the pertinence of prediction by constructing a disease-specific network. The organic combination of these technical means has significantly improved the prediction accuracy of the model.
[0106] Regarding the computational efficiency problem, this application adopts efficient feature extraction algorithms, reduces the computational complexity by optimizing the network structure, and realizes the effective sharing of model parameters. These optimization measures have significantly improved the computational efficiency of the model, making it more suitable for practical application scenarios. Specifically, this application adopts efficient feature extraction algorithms, such as the InfoGraph algorithm to extract drug molecular features and the Node2vec method to extract target and disease features. These algorithms can extract valuable features at a relatively low computational cost while maintaining high efficiency, reducing the burden of high-dimensional feature calculations. The effective sharing of model parameters is achieved by sharing the parameters in the neural network layer, avoiding repeated calculations and redundant storage, and improving the memory usage efficiency and computational speed.
[0107] In addition, in terms of solving the problem of sample imbalance, this application uses the SMOTE technique to process the training data, designs a special loss function, and introduces a transfer learning strategy. This multi-level solution effectively alleviates the model bias problem caused by sample imbalance. Among them, the special loss function is the FocalLoss loss function, which is used to balance the sample imbalance problem in drug target prediction, while the SMOTE method is used to balance the sample imbalance problem in drug-disease association prediction. The so-called transfer learning refers to using the drug combination-disease association dataset for retraining and migrating the single drug-disease association prediction model trained by single drug-disease data to the drug combination-disease association prediction task.
[0108] Compare the performance of the drug-disease interaction prediction model of this application (referred to as this model) and the comparison models (LAGCN, GFPred, CBPred, LRSSL, MBiRW, HGBI) in the drug-disease interaction prediction task. The results are shown in Table 3.
[0109] Table 3 Performance of different models in the drug-disease correlation prediction task
[0110]
[0111]
[0112] Among them, AUC (Area Under Curve) represents the area under the ROC curve. AUPR (Area Under Precision-Recall Curve) represents the area under the precision-recall curve. ACC (Accuracy) represents the accuracy rate. Fl (F1 Score) represents the F1 score. Precision represents the precision. Recall represents the recall rate. Specificity represents the specificity. It can be seen from the results in Table 3 that most of the performance of this model leads the other methods, or is almost on a par with the optimal performance of each model, showing excellent prediction performance.
[0113] In an exemplary embodiment, this application also provides a computer program product, including a computer program, which implements the drug-disease interaction prediction method when executed by a processor.
[0114] This application addresses the problems of insufficient prediction accuracy, sample imbalance, and low computational efficiency in existing drug-disease interaction prediction methods, and proposes three key technological innovations.
[0115] First, this application designs a novel feature extraction and fusion mechanism. This mechanism uses the InfoGraph algorithm to extract drug molecule features, combines the Node2vec method to obtain target features and disease features, and realizes the effective fusion of features through a multi-layer neural network, successfully solving the problem of insufficient feature expression. This multi-modal feature fusion method significantly improves the expression ability and prediction performance of the model. This mechanism extracts drug molecule features by using the InfoGraph algorithm, which captures the local and global information of molecules in the drug molecule graph structure through a graph neural network, thereby obtaining a high-dimensional drug molecule feature vector. At the same time, in combination with the Node2vec method, by performing random walks on the biological network, a low-dimensional embedding representation of genes is obtained, thereby extracting target features. Then, a fully connected neural network is used to fuse drug molecule features and target features to achieve drug target prediction, and drug biological network feature extraction is realized by performing random walks on the neural network. Furthermore, a fully connected neural network is used to fuse drug biological network features and disease features and perform association prediction. This method improves the expression ability of the model, enabling it to better capture the complex relationships between drugs, targets, and diseases, and improving the prediction performance.
[0116] Second, this application innovatively proposes a prediction method based on a disease-specific network. By integrating cancer sample data and normal tissue data in the TCGA dataset, a disease-specific network is constructed, significantly improving the prediction accuracy. In the validation of 18 types of cancers, 13 types of cancers showed obvious performance improvement, fully confirming the important role of the disease-specific network in improving the prediction accuracy. This key technological innovation corresponds to the construction of the disease-specific network mentioned in step 3.3. In this step, by integrating cancer sample data and normal tissue data in the TCGA dataset, differentially expressed genes are identified, and a weighted gene co-expression network analysis is used to construct a disease-specific network. This network can capture biological information related to disease characteristics by reflecting the differences in gene expression between cancer samples and normal tissues, thereby providing a more accurate prediction input. The disease-specific network can improve the prediction accuracy because it can highlight genes and pathways related to a specific disease (such as cancer), enabling the model to specifically learn disease-related patterns, thus significantly improving the prediction effect in the prediction of drug-disease interactions in cancer. 13 types of cancers showed performance improvement in the validation.
[0117] Third, the present application has developed an innovative drug combination prediction framework. High-precision drug combination prediction can be achieved through fine-tuning with a small number of samples (only using 0.032% of the initial training data), and the F1 score reaches 0.7746. This innovation provides a new technical route for solving the problem of sample imbalance and significantly improves the prediction efficiency at the same time. This key technical innovation realizes the feature combination between drugs by performing random walks on the same network for multiple targets of multiple drugs. At the same time, the Synthetic Minority Over-sampling Technique (SMOTE) is used to balance the class imbalance problem in the dataset. The SMOTE method oversamples the minority class (known drug-disease interactions) samples to generate new samples similar to the original minority class samples. The generated new samples are produced by interpolating the minority class samples in the feature space, ensuring that the new samples have similar features to the original samples but with slight changes in the multi-dimensional space. These generated minority class samples are added to the training set, thereby balancing the ratio of minority class and majority class samples in the dataset and avoiding over-biasing towards the majority class during the training process. The balance ratio refers to the ratio of minority class samples to majority class samples. In this method, the number of minority class samples is increased to 10 times the number of majority class samples through the SMOTE technique, ensuring that the dataset is more balanced during training and helping to improve the prediction performance of the model.
[0118] In an exemplary embodiment, the present application further provides a computer device, which may be a server or a terminal. The computer device includes a processor, a memory, an input / output interface, and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communicating with external terminals through a network connection. The computer program, when executed by the processor, implements the drug-disease interaction prediction method described above.
[0119] In an exemplary embodiment, the present application further provides a computer-readable storage medium, on which a computer program is stored. The computer program, when executed by the processor, implements the drug-disease interaction prediction method described above.
[0120] Those of ordinary skill in the art can understand that all or part of the processes in the above-described example methods can be completed by hardware related to computer program instructions. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the example methods as described above. Among them, any reference to a memory or other medium provided in the various embodiments of the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0121] It should be noted that the information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0122] Compared with traditional methods, the feature extraction and fusion mechanism designed in this application has produced significant synergistic effects. This effect can be understood from the following perspectives: The InfoGraph algorithm captures the structural information of drug molecules through a graph neural network, while the Node2vec algorithm extracts the topological features of biological networks. A fully connected neural network dynamically fuses these heterogeneous features (drug molecule features, disease features, and target features). This design enables the model to consider both molecular structure information and biological network information simultaneously, thereby making more comprehensive predictions. Experimental results show that this mechanism enables the AUC value of the model to reach 0.93, verifying the synergistic effect of feature fusion. Specifically, in this application, the drug molecule features are captured through the InfoGraph algorithm. First, the drug molecule is represented as a graph structure, and the graph neural network is used to learn the relationships between atoms (nodes) and chemical bonds (edges) in the drug molecule, thereby extracting the drug molecule features. At the same time, the Node2vec method generates neighborhood sequences of the biological network through a random walk strategy and uses the Word2vec method to convert the sequences into low-dimensional vector representations, thereby extracting the topological features of the network (disease features or target features). Subsequently, the first fully connected neural network fuses these two heterogeneous features, namely drug molecule features and target features, and effectively integrates the molecular structure information of the drug and the topological information of the gene network through a dynamic fusion mechanism, finally generating a comprehensive feature representation for drug-target prediction (i.e., realizing drug biological network features), which improves the prediction accuracy. For the feature representation of drug-disease prediction, it is also carried out using a fully connected neural network. By fusing drug network features and disease features, the association prediction between drugs and diseases is realized.
[0123] The advantages of this application in improving computational efficiency stem from the following designs: First, efficient feature extraction algorithms are adopted, such as the InfoGraph algorithm and the Node2vec method. By directly extracting key features from the data, the complexity and computational overhead of data preprocessing are reduced, thereby improving the overall computational efficiency. Second, the network structure is optimized. By selecting the most effective network architecture, unnecessary computational layers are reduced. For example, the most suitable model structure is selected through grid search, reducing the computational complexity and improving the computational efficiency. Finally, the parameter sharing mechanism adopted means that the same parameters or weights are used for multiple calculations in the model. For example, certain weights are shared in the fully connected neural network, reducing the number of parameters, and thus reducing the computational and storage costs. Experiments show that compared with traditional methods, the computational time of this application is reduced by approximately 40%, while maintaining a comparable prediction accuracy.
[0124] Furthermore, the present application significantly improves the prediction accuracy by introducing disease-specific networks. This improvement stems from the following theoretical basis: traditional methods use general biological networks for prediction, ignoring disease-specific information and resulting in less targeted prediction results. By analyzing cancer samples and normal samples in the TCGA dataset, the present application identifies differentially expressed genes related to specific diseases and constructs disease-specific networks. This method can accurately capture gene regulatory relationships in disease states, thus providing more accurate biological background information. Experimental results show that 13 out of 18 cancers perform better when using disease-specific networks, verifying the correctness of this theoretical derivation. In terms of the specific method, the present application significantly improves the prediction accuracy by introducing disease-specific networks, and the key steps include: First, by comparing cancer and normal tissue samples in the TCGA dataset, statistical methods are used to identify differentially expressed genes, which are the key features of the disease. Then, using the weighted gene co-expression network analysis algorithm, a weighted undirected graph is constructed based on the correlation between differentially expressed genes to form a disease-specific gene network.
[0125] The present application also effectively solves the problem of sample imbalance by adopting a transfer learning strategy, and its mechanism of action can be explained as follows: First, through pre-training the model by neural network on a large-scale dataset, the model learns the general feature representation of drug-disease interactions. Subsequently, only a small amount of target task data (0.032%) is used for fine-tuning, that is, re-training on the drug combination dataset with a small amount of data, and it can adapt to specific prediction tasks. The reason why this method is effective is that there are common molecular mechanism patterns in drug-disease interactions, and these common knowledge can be fully utilized through transfer learning. Experiments prove that this method obtains an F1 score of 0.7746 in the drug combination prediction task, fully verifying this theoretical derivation.
[0126] In addition, this application has excellent generalization ability, which stems from its unique network design. The introduction of disease-specific networks enables the model to adapt to different types of disease prediction tasks, while the multi-modal feature fusion mechanism ensures that the model can process data from different sources. The transfer learning strategy further enhances the adaptability of the model, enabling it to quickly transfer to new prediction tasks. Experimental verification shows that the model performs excellently in multiple cancer types, confirming its good generalization ability. Specifically, the disease-specific network analyzes cancer samples and normal samples in the TCGA dataset, identifies differentially expressed genes, and constructs a specific biological network, thereby enhancing the adaptability and accuracy of the model in specific disease prediction. The multi-modal feature fusion mechanism combines the InfoGraph algorithm to extract drug molecular features and the Node2vec method to obtain the topological features of gene networks and disease networks for effective fusion through neural networks, enabling the model to comprehensively consider feature information from different sources and improve the comprehensiveness and expressiveness of prediction. The transfer learning strategy pre-trains the model on a large-scale dataset and fine-tunes it using a small amount of target task data to quickly adapt to new prediction tasks, enhancing the generalization ability of the model.
[0127] Through the above theoretical derivation and experimental verification, the advantages of this application in terms of prediction accuracy, sample balance, feature fusion, computational efficiency, and model generalization are fully demonstrated. These advantages together constitute a complete technical solution, providing reliable technical support for drug-disease interaction prediction.
[0128] In terms of single-drug prediction, through innovative network construction methods and feature extraction strategies, combined with a sample balancing processing mechanism, the prediction accuracy has been significantly improved. Experimental results show that the method of this application reaches 0.93 in the AUC metric and 0.6316 in the F1 score, far exceeding the prediction performance of existing algorithms. In terms of drug combination prediction, this application has achieved direct prediction of drug combinations for the first time. Through fine-tuning with a small number of samples (only using 0.032% of the initial training data), the model achieved an F1 score of 0.7746 in the drug combination prediction task, significantly outperforming existing methods. In terms of disease-specific prediction, this application innovatively introduced a disease-specific biological network. In the validation experiments for 18 types of cancers, 13 performed better when using the disease-specific network, confirming the prediction advantage of this method in specific disease scenarios. In terms of experimental verification, the accuracy of the prediction results was successfully verified through in vitro cytotoxicity experiments, and two new synergistic drug combinations were discovered for the treatment of different types of cancers, fully confirming the practical value of the method of this application. These technological innovations not only address the deficiencies of existing methods in aspects such as prediction accuracy, sample imbalance, computational efficiency, and generalization ability, but also provide a reliable theoretical basis for the formulation of personalized medical plans. Especially in the application of the disease-specific network, this application pioneeringly proposed a new technical route, providing a new solution for improving prediction accuracy.
[0129] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0130] Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method of this application and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A method for predicting drug-disease interactions, characterized in that: include: Acquiring drug data, disease data and biological data; the drug data includes drug molecular structure data; the disease data includes disease phenotype information and disease ontology relationship; the biological data includes protein interaction data and cell signaling pathway data; Preprocess and extract features of drug data to obtain drug molecular features; Extract target features based on biological data; Inputting the drug molecular features and target features into the first fully connected neural network to predict the drug-target interaction relationship; Generate drug biological network features based on drug-target interaction relationships; Preprocess and extract features of disease data to obtain disease features; The second fully connected neural network is trained based on the characteristics of the drug biological network and the disease characteristics, and a drug-disease interaction prediction model is obtained after the training is completed; The drug-disease interaction prediction model is used to predict drug-disease interactions.
2. The drug-disease interaction prediction method according to claim 1, characterized in that: The drug data is preprocessed and feature extracted to obtain drug molecular features, specifically including: For drug molecular structure data, extract its molecular descriptors in SMILES format, and use the RDKit toolkit to calculate the molecular fingerprint as the drug molecular graph structure; The InfoGraph algorithm is used to process the drug molecular graph structure and extract the drug molecular features.
3. The drug-disease interaction prediction method according to claim 1, characterized in that: The target feature extraction based on biological data specifically includes: Constructing a directed protein interaction network based on protein interaction data and cell signaling pathway data; in the directed protein interaction network, each node represents a protein, each edge represents an interaction relationship between two proteins, and the direction of the edge indicates the activation or inhibition effect of one protein on another protein; Based on the directed protein interaction network, the Node2vec algorithm was used to extract target features.
4. The drug-disease interaction prediction method according to claim 3, characterized in that: Also includes: Obtain gene expression data from the TCGA database; Identify differentially expressed genes using gene expression data; The correlation between differentially expressed genes was analyzed using a weighted gene co-expression network to generate a disease-specific gene expression correlation matrix; By setting the correlation threshold of the disease-specific gene expression correlation matrix, significantly correlated gene pairs are retained and irrelevant gene pairs are deleted; A weighted undirected graph was constructed based on the retained significantly correlated gene pairs as a disease-specific network; Optimizing directed protein interaction networks based on disease-specific networks.
5. The drug-disease interaction prediction method according to claim 1, characterized in that: The generating of drug biological network features based on drug-target interaction relationship specifically includes: Based on the drug-target interaction relationship, a random walk strategy is used in a directed protein interaction network to simulate the interaction of drugs on biological networks. During each random walk, the activation or inhibition of the drug on the target will affect the value of the node, and the interaction relationship between proteins will determine the update direction and strength of the node value. Multiple random walk results are integrated into a multidimensional vector to constitute the characteristics of the drug-bionetwork.
6. The method for predicting drug-disease interactions according to claim 1, characterized in that: The preprocessing and feature extraction of disease data to obtain disease features specifically includes: Integrate the disease phenotype information in the OMIM database and the disease ontology relationships in the MeSH database, use the graph embedding algorithm to represent disease-related terms as nodes, and represent the relationships between nodes as edges, thereby generating a disease relationship network; The Node2vec algorithm is used to process the disease relationship network and extract disease characteristics.
7. The method for predicting drug-disease interactions according to claim 1, characterized in that: The second fully connected neural network is trained based on the drug biological network characteristics and the disease characteristics, and a drug-disease interaction prediction model is obtained after the training is completed, specifically including: Constructing a second fully connected neural network including an input layer, a hidden layer, and an output layer; A dataset is constructed based on drug biological network characteristics, disease characteristics, and drug-disease interaction relationship data provided by the CTD database, and a synthetic minority class oversampling technique is used to balance the class imbalance problem in the dataset. The second fully connected neural network was trained using the dataset. Batch normalization and a dropout rate of 50% were used during the training process. Leaky ReLU was used as the activation function. The loss function used the binary cross entropy loss function. The optimizer used Adam. The learning rate and weight decay were both set to 0.
001. After the training is completed, a drug-disease interaction prediction model is obtained; the input of the drug-disease interaction prediction model is the drug biological network characteristics and the disease characteristics, and the output is the drug-disease interaction score; a score greater than 0 is related, otherwise it is irrelevant, and the higher the score, the greater the possibility of association between the drug and the disease.
8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the drug-disease interaction prediction method according to any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the drug-disease interaction prediction method according to any one of claims 1 to 7 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the drug-disease interaction prediction method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
PRAK signal path analysis method and system
CN121350496A
PRAK signaling pathway analysis method and system
CN121350496B