An intelligent prediction method and device for adverse drug reactions based on multi-source data fusion

The Drug-Gene-ADR neural network (DGANet) addresses the inefficiencies in existing ADR prediction methods by integrating pharmacogenomic data sources, improving prediction accuracy and efficiency through advanced feature extraction and model design, supporting personalized medicine.

CN119517219BActive Publication Date: 2025-07-15GUANGDONG PHARMA UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411739486.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-29
Publication Date
2025-07-15
Estimated Expiration
2044-11-29

AI Technical Summary

Technical Problem

The prior art lacks effective feature analysis methods and efficient models when processing pharmacogenomic data, resulting in low predictive performance of drug adverse reactions.

Method used

Using a multi-source data fusion method, drug descriptors and adverse reaction descriptors were constructed by integrating comparative toxicology genomics data, LINCS L1000 data, PubChem data and medical topic lexicon data, and drug-gene-drug adverse reaction descriptors were constructed using convolutional neural networks and multi-layer perceptrons to improve prediction accuracy.

Benefits of technology

It improves the accuracy and efficiency of drug adverse reaction prediction, enhances the prediction ability of unknown ADRs, and verifies the effectiveness of DGANet in new ADR discovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119517219B_ABST
    Figure CN119517219B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose an intelligent prediction method and device for adverse drug reactions with multi-source data fusion, relating to the technical field of drug-gene-adverse drug reaction neural networks, aiming to solve the problem that the existing technology lacks more effective feature analysis methods and more efficient model designs, thus resulting in low performance in predicting ADRs. The method includes: obtaining multi-source data; after performing normalization processing on the multi-source data, constructing drug descriptor features to obtain a set of drug similarity vectors; constructing adverse drug reaction descriptor features to obtain a set of adverse drug reaction similarity vectors; based on the set of drug similarity vectors and the set of adverse drug reaction similarity vectors, training the target intelligent prediction model for adverse drug reactions, and after the training is completed, predicting the adverse reactions that occur when the target drug is used on the target individual based on the target intelligent prediction model for adverse drug reactions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of drug-gene-drug adverse reaction neural networks, and particularly to an intelligent prediction method and device for drug adverse reactions with multi-source data fusion. Background Art

[0002] Drug adverse reactions (ADRs) are an important issue in clinical treatment, which may pose serious health risks and even lead to death. Currently, the collection of pharmacogenomic data from clinical observations has increased significantly, providing the possibility of predicting ADRs based on existing experience. Facing a large amount of complex and diverse pharmacogenomic data, it is crucial for clinical drug vigilance to efficiently discover the laws therein and better guide the prediction of unknown ADRs. In recent years, computer technologies and related algorithms with ultra-high computing power and artificial intelligence analysis capabilities have played an increasingly significant role in assisting clinical ADR prediction and analysis. However, how to effectively process these massive data and mine the objective laws contained therein for accurate prediction remains a huge challenge. Existing studies have used source data such as the LINCS L1000 project, the STITCH database, and the Comparative Toxicogenomics Database (CTD) to build models and achieved certain success. However, there are still some deficiencies in these methods in terms of effective feature extraction and optimal model structure design. In addition, although these publicly available source databases provide a rich genomic information basis for ADR prediction, the huge amount of gene data also poses a significant challenge to feature processing and model design. Therefore, it is necessary to develop more effective feature analysis methods and design more efficient models to further improve the performance of predicting ADRs. Summary of the Invention

[0003] The main purpose of the present application is to provide an intelligent prediction method and device for drug adverse reactions with multi-source data fusion, aiming to solve the problem that the existing technology lacks more effective feature analysis methods and more efficient model designs, resulting in low performance in predicting ADRs.

[0004] To achieve the above object, the technical solutions adopted in the embodiments of the present application are as follows:

[0005] In a first aspect, an embodiment of the present application provides an intelligent prediction method for drug adverse reactions with multi-source data fusion, including the following steps:

[0006] Obtain multi-source data; wherein, the multi-source data includes comparative toxicogenomics data, LINCS L1000 data, PubChem data, Medical Subject Headings data, and SIDER data; the comparative toxicogenomics data is used to explain drug-gene interactions, the LINCS L1000 data is used to explain gene expression signals, the PubChem data is used to explain drug SMILES, the Medical Subject Headings data is used to explain gene-disease interactions, and the SIDER data is used to explain drug-side effect pairs;

[0007] After standardizing the multi-source data, construct drug descriptor features to obtain a drug similarity vector set; construct drug adverse reaction descriptor features to obtain a drug adverse reaction similarity vector set;

[0008] Based on the drug similarity vector set and the drug adverse reaction similarity vector set, train the target drug adverse reaction intelligent prediction model. After training is completed, based on the target drug adverse reaction intelligent prediction model, predict the adverse reactions that occur when the target drug is used on the target individual.

[0009] As some alternative embodiments of the present application, the drug similarity vector set includes: the chemical structure similarity of two drugs, the drug similarity based on drug-gene interactions, and the drug similarity of gene expression differences after drug perturbation;

[0010] The drug adverse reaction similarity vector set includes: the semantic similarity of different disease ontologies and the gene-disease association relationship set.

[0011] As some alternative embodiments of the present application, the chemical structure similarity value of the two drugs is obtained based on the following steps:

[0012] By using the drug Compound ID provided in SIDER, batch download the SMILES string representation of each drug structure from PubChem, and then use the Rdkit tool in Python to convert it into a topological fingerprint; the topological fingerprint is generated based on the topological structure and rotation angle of the four-membered ring in the molecule, and is used to describe the stereochemistry and interactions of the molecule; and use the Tanimoto coefficient as a metric standard to obtain the chemical structure similarity value of drug d i and drug d j The chemical structure similarity value of the two drugs ( represents the similarity of drug-drug in the drug chemical structure space, where cs(i,j) represents the chemical space structure of two different drugs).

[0013] As some optional embodiments of the present application, the drug similarity based on drug-gene interaction means that:

[0014] The drug-gene interaction is represented by One-Hot encoding, that is, each unique gene interaction is converted into a binary vector; drug d i and drug d j Similarity based on drug-gene interaction is calculated using the Jaccard index, and the formula is:

[0015]

[0016] where GT i and GT j respectively represent the gene sets having interaction relationships with drugs d i and d j .

[0017] As some optional embodiments of the present application, the drug similarity of gene expression differences after drug perturbation means that:

[0018] Drugs d i and d j Similarity based on gene expression differences after drug perturbation is calculated through the cosine similarity formula, and the formula is:

[0019]

[0020] where and represent the gene expression feature vectors after perturbation of drug i or j, and represent the norms of vectors and .

[0021] As some optional embodiments of the present application, the semantic similarity of different disease ontologies is obtained through the following steps:

[0022] For each drug adverse reaction, a directed acyclic graph is constructed based on its hierarchical descriptors in MeSH, where the nodes represent the nodes of the adverse reaction in MeSH, and the edges represent the relationship between the current node and its ancestors; the adverse reaction s is represented as the graph DAG s =(s, N s , E s ), where N s represents the set of all ancestor nodes including the node of the adverse reaction s itself, and E srepresents all the sets of edges pointing from parent nodes to child nodes; to encode the semantics of disease ontology terms in a measurable format, the research defines the semantic value SV(A) of disease A as the total contribution of all diseases in the DAG A to the semantics of disease A, where disease terms closer to disease A contribute more to its semantics; for two different disease ontologies A and B, the semantic similarity S(A, B) between A and B is defined as:

[0023]

[0024] As some alternative embodiments of the present application, the gene-disease association relationship set is obtained through the following steps:

[0025] Downloaded the compound-gene interaction table CTD_genes_diseases.csv.gz from CTD as the gene-disease association relationship set of the knowledge graph, which contains 107,911,805 gene-disease association records; the gene-disease associations in CTD include associations extracted and inferred by experts; the extracted associations are curated by CTD curators from published literature or exported from the OMIM database using the mim2gene file in the NCBI gene database; there are three types of evidence for direct gene-disease associations: M marker, mechanism, and treatment evidence; M marker refers to specific variations in the gene sequence; mechanism evidence includes how gene variations affect the structure and function of proteins and how these changes interfere with normal cell or physiological processes; treatment evidence refers to the improvement of disease symptoms or prevention of disease occurrence through treatments targeting specific genes or their products.

[0026] As some alternative embodiments of the present application, the target drug adverse reaction intelligent prediction model is trained through three branches by the original drug features combination Drug, the feature cross V FCs , and the original drug adverse reaction features Side combination respectively;

[0027] Among them, Drug and Side are learned using two linear sub-networks with the same structure, while V FCs is learned using a convolutional neural network sub-network; the architecture of the linear sub-network consists of two fully connected layers, a batch normalization layer, and an activation function layer; that is:

[0028] O m,k = Dropout (p) (Relu(FC n1 (CAT(x m , c k )))),

[0029] x′ m,k = Linear(Om,k ),

[0030] wherein, x′ m,k in LSN d and LSN s respectively represent the potential representations of m drugs and m drug adverse reactions with k different similarity features, that is, the output of LSN. Among them, LSN d and LSN s represent two linearly consistent linear sub-networks, FC (*) represents a fully connected layer, where * represents the number of neurons, Dropout represents a Dropout layer with a probability of p. Linear and Relu respectively represent a linear function and a rectified linear unit activation function, and CAT connects the given feature vectors. Through this step, the linear embedding vectors emb drugs and emb adrs can be obtained.

[0031] As some alternative embodiments of the present application, the target drug adverse reaction intelligent prediction model is obtained by concatenating emb drugs , emb cross and emb ADRs as the input vector and inputting it into a multi-label classifier for classification; the classifier is composed of two fully connected layers, two activation function layers and a BN layer; finally, a vector is output, where the numbers greater than 0 represent the association between the drug and the ADR, and the larger the number, the greater the possibility of the association; the output vector satisfies the following relational expression:

[0032]

[0033] wherein, FC (*) represents a fully connected layer, where * represents the number of neurons, Dropout represents a Dropout layer with a probability of p. Linear and Relu respectively represent a linear function and a rectified linear unit activation function, and CAT connects the given feature vectors.

[0034] In a second aspect, an embodiment of the present application provides a drug adverse reaction intelligent prediction device for multi-source data fusion, including:

[0035] A data acquisition module, configured to acquire multi-source data; wherein, the multi-source data includes comparative toxicogenomics data, LINCS L1000 data, PubChem data, Medical Subject Headings data, and SIDER data;

[0036] Construct a feature module for constructing drug descriptor features after normalizing the multi-source data to obtain a set of drug similarity vectors; construct drug adverse reaction descriptor features to obtain a set of drug adverse reaction similarity vectors.

[0037] A prediction module is used to train the target drug adverse reaction intelligent prediction model based on the set of drug similarity vectors and the set of drug adverse reaction similarity vectors. After the training is completed, based on the target drug adverse reaction intelligent prediction model, predict the adverse reactions that occur when the target drug is used on the target individual.

[0038] Compared with the prior art, the beneficial effects of this application are:

[0039] This application provides a new deep learning architecture - Drug-Gene-ADR neutral network (DGANet), which can effectively integrate pharmacogenomic source data from multiple public databases, mainly including Chemical-Gene Interactions (CGIs), Gene-Disease Associations (GDAs), and gene expression changes caused by drug perturbations. Based on Convolutional Neural Networks (CNN), expand the information related to the input and potential drug-gene-drug adverse reaction associations to form a new deep learning architecture with learning cross-features - DGANet, aiming to improve the accuracy and efficiency of predicting Adverse Drug Reactions (ADRs). In addition, based on the existing gene data, new genomic features are proposed to further enhance the model's prediction ability for unknown ADRs and verify the effectiveness of DGANet for the discovery of new ADRs. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a general diagram of the data source involved in the embodiment of this application;

[0041] Figure 2 It is a data processing flow chart involved in the embodiment of this application;

[0042] Figure 3 It is the overall structure diagram of DAGNet involved in the embodiment of this application;

[0043] Figure 4 It is the structure diagram of the convolutional neural sub-network involved in the embodiment of this application. DETAILED DESCRIPTION OF THE INVENTION

[0044] It should be understood that the specific embodiments described herein are merely for explaining the present application and are not used to limit the present application.

[0045] The present application proposes a prediction model based on deep learning technology. This model combines the convolutional neural network VGGNet (Visual Geometry Group Network) framework and the multilayer perceptron (MLP) to extract features from multi-source pharmacogenomics data and predict the adverse drug reactions that drugs may cause. The present application names the entire model the Drug-Gene-ADR neutral network (DGANet). DGANet mainly integrates three representation features of drugs (drug similarity based on drug chemical structure, drug similarity based on drug-gene interaction, drug similarity based on gene expression differences after drug perturbation) and the representation features of adverse drug reactions, learns the potential relationships among drugs, genes, and adverse drug reactions through a multi-layer convolutional neural network, predicts the possible adverse drug reactions of drugs, and gives the occurrence probabilities of various adverse drug reactions that the current drug may cause.

[0046] The present application selects multiple public databases as data sources, and the general situation of the data sources is as Figure 1 shown. It includes the Comparative Toxicogenomics Database (CTD), LINCS L1000, PubChem, the US National Library of Medicine's Medical Subject Headings (MeSH), and SIDER (Side Effect Resource) to ensure the breadth and reliability of the research source data.

[0047] Among the publicly available data sources selected for this application, CTD is a curated database that focuses on describing chemical-gene / protein interactions across species and the associations between chemicals and gene diseases, aiming to reveal the impact of susceptibility and environmental factors on diseases. This database is manually curated and annotated by professional bioinformaticians to ensure the high quality and accuracy of the data. Its content covers over 52,638,316 toxicogenomics relationships, involving more than 17,100 chemicals, 54,300 genes, 6,100 phenotypes, 7,270 diseases, and 202,000 exposure events. CTD utilizes MeSH and OMIM (Online Mendelian Inheritance in Man) terms to construct a comprehensive disease vocabulary for a better understanding of the potential mechanisms of the effects of specific drugs on the human body.

[0048] LINCS L1000 is one of the projects of The Library of Integrated Network-Based Cellular Signatures (LINCS) supported by the National Institutes of Health (NIH) in the United States. Its aim is to collect and analyze data on the responses of human cells to various biological and chemical perturbations through high-throughput technologies, so as to promote the understanding of biological functions. This database contains approximately 1,000,000 expression profiles, mainly by treating 99 cell lines with 32,855 small molecules and measuring the expression profiles of 978 landmark genes on 384-well plates, and using deep learning methods to predict the expression levels of the remaining 12,328 genes. The data preprocessing of LINCS L1000 is divided into five levels. The first level (Luminex Bead File, LXB): raw, untreated flow cytometry data from the Luminex scanner. An LXB file is generated for each well of the 384-well plate, and each file contains the fluorescence intensity values of each observed analyte in the well. The second level (Gene Expression, GEX): gene expression values of every 1000 genes after deconvolution from Luminex beads. The third level (Quantile-2 Normalized Data, Q2NORM): gene expression profiles of directly measured signature transcripts and inferred genes. Normalization is performed using invariant set scaling and quantile normalization. The fourth level (Z-Score Calculations, Z-SCORES): features of differentially expressed genes are calculated by robust Z-scoring each profile relative to the control (PC (plate control) as the blank control population; VC (vehicle control) as the solvent control). The fifth level (Signature, SIG) consists of the results of repeated experimental groups, usually with 3 replicates per group, and finally forms a single differential expression vector derived from the weighted average of individual replicates, providing a valuable resource for studying cell responses.

[0049] PubChem is an open chemical database maintained by the National Center for Biotechnology Information (NCBI) of the NIH. It contains information on millions of chemical substances from a wide range of sources, not limited to small molecule drugs, but also including larger molecules such as nucleotides, carbohydrates, etc. It provides information such as chemical structures, biological activities, patents, health, safety, and toxicity data. PubChem is convenient for association with multiple public databases such as Drugbank, STITCH, SIDER, etc., enhancing the availability and practicality of the data.

[0050] MeSH is a set of authoritative and standardized medical terminology glossaries compiled by the National Library of Medicine (NLM) of the United States for indexing and retrieving biomedical literature. MeSH organizes terms through a tree structure, including terms and qualifiers, organized hierarchically, covering 16 categories of biomedical terms, especially disease terms, organized hierarchically for easy understanding and precise retrieval, and is a key tool for medical information retrieval.

[0051] SIDER is a database on marketed drugs and their adverse reactions, also maintained by the NCBI of the NIH, aiming to support the research on drug side effects and the evaluation of drug safety. SIDER integrates data from multiple sources, including the label information of the Food and Drug Administration (FDA) of the United States, drug product labels, clinical trial results, and scientific literature, providing information such as drug names, active ingredients, dosage forms, administration routes, and descriptions, frequencies, severities of side effects, etc., which is of great value for drug development, clinical practice, and drug regulation. The SIDER 4.1 version used in this application contains 1,430 drugs, 5,880 side effects, and 139,756 drug-side effect pairs, providing a valuable resource for the identification, evaluation, and management of drug side effects. By integrating the information of these databases, the research of this application can comprehensively explore the complex relationships among drug action mechanisms, side effects, and diseases.

[0052] As Figure 2 shown in the data preprocessing process, in the data preprocessing step, this application constructs drug descriptors and drug adverse reaction descriptors, and the specific details are as follows:

[0053] Constructing drug descriptors: By using the drug Compound ID provided in SIDER, the SMILES (Simplified Molecular Input Line Entry System) string representation of each drug structure can be downloaded in batches from PubChem, and then converted into topological fingerprints using the Rdkit tool in Python. Topological fingerprints are generated based on the topological structure and rotation angles of four-membered rings in the molecule, and can effectively describe the stereochemistry and interactions of the molecule. To quantify the chemical structure similarity between two drugs, this application uses the Tanimoto coefficient as a metric, and defines the chemical structure similarity between two drugs as: Represents the similarity between drug-drug in the drug chemical structure space, where cs(i,j) represents the chemical space structures of two different drugs.

[0054] Secondly, to obtain drug similarity based on drug-gene interactions, this application downloaded the compound-gene interaction table CTD_chem_gene_ixns.csv.gz (version: updated on February 28, 2024) from CTD. This dataset contains 2,676,084 compound-gene interaction records. CTD extracts specific compound-gene or protein interactions from published literature, and most of these interactions are binary, but also include more complex nested events. Each chemical-gene interaction has a certain degree limit, such as increase, decrease, affect or not affect, etc. Drug-gene interactions are represented by One-Hot encoding, that is, converting each unique gene interaction into a binary vector. Drug d i and d j Drug similarity based on drug-gene interactions Calculated using the Jaccard index, the formula is:

[0055]

[0056] Where, GT i and GT j Respectively represent the gene sets that have interaction relationships with drug d i and d j There is an interaction relationship between the gene sets.

[0057] Finally, the drug similarity based on gene expression differences after drug perturbation was obtained from the LINCS L1000 project. This project analyzed the gene expression changes after treating cells with each small molecule compound at different doses and time points. In the GSE92742 dataset, the "distil_ss" variable quantifies the magnitude of the differential expression of signature genes, representing the gene expression intensity of each experiment. In this application's research, the differences in cell type, dose, or time point were ignored, and the "distil_ss" value with the largest gene expression difference was selected as the characteristic representation of each drug. The characteristics of a drug were defined as a vector of continuous values, with each value representing the direction and magnitude of differential expression between the control sample and the compound-treated sample. The characteristic of the gene expression difference perturbed by the drug was calculated using the Characteristic Direction (CD) method. Drug d i and d j Similarity based on gene expression differences after drug perturbation Calculated by the cosine similarity formula, the formula is:

[0058]

[0059] By integrating data from three aspects: drug chemical structure, drug-gene interaction, and gene expression differences after drug perturbation, a multi-dimensional drug descriptor feature was constructed, providing a comprehensive and in-depth perspective for drug similarity analysis. This not only helps to understand the mechanism of action of drugs but also provides new ideas and methods for drug repositioning and new drug development.

[0060] Constructing a new drug adverse reaction descriptor: To construct a more comprehensive and efficient drug adverse reaction descriptor, this application's research adopted two methods: drug adverse reaction semantic similarity based on MeSH hierarchical descriptors and gene-disease association relationships. First, for each drug adverse reaction, the research constructed a directed acyclic graph (DAG) based on its hierarchical descriptor in MeSH, where the nodes represent the nodes of the adverse reaction in MeSH (i.e., its medical descriptors or medical terms), and the edges represent the relationships between the current node and its ancestors. Specifically, the adverse reaction s can be represented as the graph DAG s =(s, N s , E s ), where N s represents the set of all ancestor nodes including the node of the adverse reaction s itself, and E s represents the set of all edges pointing from the parent node to the child node. To encode the semantics of disease ontology terms in a measurable format, the research defined the semantic value SV(A) of disease A as DAG AThe total contribution of all diseases to the semantics of Disease A, where disease terms closer to Disease A contribute more to its semantics. For two different disease ontologies A and B, the semantic similarity S(A, B) is defined as:

[0061]

[0062] This formula not only considers the position of the disease ontology in the DAG but also reflects its semantic relationship with its ancestor terms, making the calculation of semantic similarity more in line with the actual clinical significance.

[0063] In addition, the compound-gene interaction table CTD_genes_diseases.csv.gz (version: updated on February 28, 2024) was downloaded from CTD as the set of gene-disease association relationships in the knowledge graph, containing 107,911,805 gene-disease association records. The gene-disease associations in CTD include associations extracted and inferred by experts. The extracted associations were curated by CTD curators from published literature or exported from the OMIM database using the mim2gene file in the NCBI gene database. There are three types of evidence for direct gene-disease associations: M markers (genetic markers directly related to the disease), mechanisms (involving an understanding of how gene variants lead to the biological basis of the disease), and treatments (the link between gene variants and the disease can be verified through therapeutic interventions). M markers usually refer to specific variants in the gene sequence, such as single nucleotide polymorphisms (SNPs), insertions / deletions (indels), or copy number variations (CNVs), which are significantly associated with disease risk or phenotype. Mechanistic evidence includes how gene variants affect the structure and function of proteins and how these changes disrupt normal cellular or physiological processes, leading to the disease. Therapeutic evidence refers to the improvement of disease symptoms or prevention of disease occurrence through treatments targeting specific genes or their products, providing strong evidence for the direct link between genes and diseases.

[0064] In addition to direct associations, CTD also includes partially inferred associations that are established through chemically - gene interactions curated by CTD. The inference is based on inference scores from the network topology of a network composed of genes, diseases, and one or more chemicals used for reasoning. The inference score reflects the degree of similarity between the CTD chemical - gene - disease network and a scale - free random network, and is calculated as the logarithm - transformed product of two commonly used neighbor statistics, which are used to evaluate the functional relationships between proteins in a protein - protein interaction network. The inference score takes into account the connectivity of genes and diseases, the number of chemicals used for reasoning, and the connectivity of each chemical. The higher the inference score, the more likely the network is considered to have atypical connections. Many biological networks, such as disease and metabolic networks, have been shown to be scale - free random networks.

[0065] The overall network structure of the model is as Figure 3 shown, and the inputs of the model include a set of drug similarity vectors and a set of drug adverse reaction similarity vectors. The drug similarity set Drug is represented as where represents the similarity between drugs - drugs in the drug chemical structure space, represents the similarity between drugs - drugs in the space of gene mutation intensity after drug perturbation, represents the similarity between drugs - drugs in the drug - gene interaction space. The drug adverse reaction similarity set Side is represented as where represents the similarity between adverse reactions - adverse reactions in the MESH disease tree structure space, represents the similarity between adverse reactions - adverse reactions in the gene - adverse reaction association relationship space.

[0066] To capture the non - linear relationship between different features of drugs and adverse reactions, combined features of drugs and drug adverse reactions, namely feature crossing, are constructed. The combined feature of drugs and drug adverse reactions is defined as V FCs , and is expressed by the formula where represents the tensor product.

[0067] The original drug feature combination Drug, the feature crossing V FCs , and the original drug adverse reaction feature combination Side are trained through three branches respectively. Specifically, Drug and Side are learned using two structurally consistent linear sub - networks (LinearSubnetworks, LSN d and LSN s ), while V FCsLearn using a Convolutional Neural Networks Subnetwork (CSN). The architecture of the linear subnetwork consists of two fully connected layers, a Batch Normalization (BN) layer, and an activation function layer. The specific formula is as follows:

[0068] O m,k = Dropout (p) (Relu(FC n1 (CAT(x m , c k )))),

[0069] x′ m,k = Linear(O m,k ),

[0070] where x′ m,k represents the potential representations of m drugs and m adverse drug reactions with k different similarity features in LSN d and LSN s respectively, that is, the output of LSN. FC (*) represents a fully connected layer, where * represents the number of neurons, and Dropout represents a Dropout layer with a probability of p. Linear and Relu represent the linear function and the rectified linear unit activation function respectively, and CAT concatenates the given feature vectors. Through this step, the linear embedding vectors emb drugs and emb adrs of drugs and adverse reactions can be obtained.

[0071] The convolutional neural network subnetwork CSN refers to the practice of VGGNet. Each convolutional block consists of Conv2d (convolution) + BatchNorm2d (batch normalization) + ReLU (activation layer). Considering the experimental resources and time duration, the number of layers and parameters are simplified. Six convolutional layers are constructed, and the layer-to-layer connection uses a normalization layer, and the hidden layers are connected using the ReLU activation function. A 2×2 convolutional kernel is used, the stride is set to 2, and the number of channels is set to 32. After each sampling, the height and width of the feature matrix are reduced to half of the original, and the final output is named emb cross , and the CSN structure is as Figure 4 shown.

[0072] The emb drugs , emb cross and emb ADRsAfter splicing, it is used as an input vector to be input into a multi-label classifier for classification. The classifier consists of two fully connected layers, two activation function layers, and one BN layer. Finally, a vector is output, where numbers greater than 0 indicate the association between the drug and the ADR, and the larger the number, the greater the likelihood of the association. The output is expressed as a formula:

[0073]

[0074] Finally, it comes to the part of the loss function. In deep learning, the loss function determines the ability of the model to handle tasks. The multi-label classification task often faces problems such as class imbalance, label sparsity, and latent class correlation. To solve these problems, this application studies the use of ZLPR (zero-bounded log-sum-exp&pairwise rank-based) proposed by Su et al. as the loss function of DGANet to measure the error between the predicted value and the true value. ZLPR generalizes the "softmax + cross-entropy" scheme to the multi-label classification scenario without particularly adjusting the class weights and thresholds. Compared with the loss based on Binary Relevance (BR), ZLPR can better capture the label correlation and the ranking relationship between positive and negative classes. Compared with the loss based on Label Ranking (LR), ZLPR can adaptively determine the number of target classes and enhance the label ranking ability of the model. ZLPR combines the advantages of LR and BR, achieves efficient spatio-temporal performance, retains the class correlation information in the original data, and is applicable to the case where the number of classes is uncertain. The formula of the ZLPR loss function is as follows, and the learning rate is set to 0.005 in the experiment:

[0075]

[0076] where, s i and s j represent that these two variables represent the output scores of the model for the i-th or j-th class, and Ω pos and Ω pos represent the set of positive labels and the set of negative labels.

[0077] By adopting the ZLPR loss function, the multi-label classification task can be processed more effectively, improving the performance and robustness of the model.

[0078] Through the above steps, a multi-modal feature fusion model is constructed, which can effectively capture the complex relationship between drugs and adverse reactions, providing a new method for the prediction and management of drug adverse reactions.

[0079] This application's research systematically compared the effects of different pharmacogenomic cross - feature combinations in assisting the prediction of adverse drug reactions, and comprehensively evaluated the performance of the proposed new model. The experimental results show that the model proposed in this application's research, when using drug - gene interactions (CGI) as features, has the optimal prediction ability under all test conditions. Notably, when combined with the use of drug chemical structure (CS) information, the model achieved the best performance with an AUROC (Area Under the Receiver Operating Characteristic Curve) of 92.30% and an AUPRC (Area Under the Precision - Recall Curve) of 91.80%, demonstrating the optimality of this innovative feature combination.

[0080] In addition, this application's research innovatively uses MESH disease concepts to measure the semantic similarity between adverse drug reactions and combines it with gene - disease association (GDA). This method is superior to simply relying on GDA. This improvement not only enhances the model's ability to understand complex biomedical data but also provides a more accurate basis for subsequent analysis. More importantly, embedding the neighborhood similarity of known drug - adverse reaction associations further improves the overall performance of the model. Specifically, this strategy increased AUROC by 0.38% - 0.77% and AUPRC by 0.62% - 0.85%. These improvements demonstrate the importance of considering existing knowledge in improving prediction accuracy.

[0081] Finally, a class - balanced sampling technique was implemented within the DGANet framework to construct a balanced dataset for the training and validation processes. This method effectively mitigates the negative impacts brought about by sample imbalance, ensuring the consistency and reliability between evaluation metrics such as AUROC and AUPRC. Even when working on a relatively small - scale dataset and regardless of whether additional neighborhood similarity information based on drug side - effect association (DSA) is introduced, DGANet can still achieve increases of 4.6% and 5.06% in the AUROC metric, showing strong generalization ability and potential application value.

[0082] Therefore, through carefully designed feature engineering and algorithm optimization, this application has successfully developed a new model that can efficiently and accurately predict adverse drug reactions, laying a solid foundation for promoting the development of personalized medicine.

[0083] The above - mentioned are only the preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application shall be included within the protection scope of this application.

Claims

1. An intelligent prediction method for adverse drug reactions with multi-source data fusion, characterized in that, Including the following steps: Obtain multi-source data; wherein, the multi-source data includes comparative toxicogenomics data, LINCS L1000 data, PubChem data, Medical Subject Headings data, and SIDER data; the comparative toxicogenomics data is used to explain drug-gene interactions, the LINCS L1000 data is used to explain gene expression signals, the PubChem data is used to explain drug SMILES, the Medical Subject Headings data is used to explain gene-disease interactions, and the SIDER data is used to explain drug-side effect pairs; After normalizing the multi-source data, construct drug descriptor features to obtain a set of drug similarity vectors; construct drug adverse reaction descriptor features to obtain a set of drug adverse reaction similarity vectors; the set of drug similarity vectors includes: the chemical structure similarity of two drugs, the drug similarity based on drug-gene interactions, and the drug similarity of gene expression differences after drug perturbation; the set of drug adverse reaction similarity vectors includes: the semantic similarity of different disease ontologies and the set of gene-disease association relationships; Based on the set of drug similarity vectors and the set of drug adverse reaction similarity vectors, train a target drug adverse reaction intelligent prediction model. After training is completed, based on the target drug adverse reaction intelligent prediction model, predict the adverse reactions that occur when a target drug is used on a target individual; Among them, the target adverse drug reaction intelligent prediction model is trained through three branches respectively by combining the original drug features , feature crossing , and the original adverse drug reaction features combined. Among them, and use two linear sub-networks with the same structure for learning, while uses a convolutional neural network sub-network for learning; the architecture of the linear sub-network consists of two fully connected layers, a batch normalization layer, and an activation function layer; that is: ; Among them, In and respectively represent the potential representations of m drugs with k different similarity features and m drug adverse reactions, that is, the output of LSN; among them, and represent two linearly consistent sub-networks; represents a fully connected layer, where * represents the number of neurons; Dropout represents a Dropout layer with a probability of p; and represent a linear function and a rectified linear unit activation function respectively, and CAT concatenates the given feature vectors; The target adverse drug reaction intelligent prediction model is obtained by concatenating the linear embedding vector , the final output and the linear embedding vector to form an input vector, which is then input into a multi-label classifier for classification.

2. The intelligent prediction method for adverse drug reactions with multi-source data fusion according to claim 1, characterized in that, The chemical structure similarity value of the two drugs is obtained based on the following steps: By using the drug Compound ID provided in SIDER, the SMILES string representation of each drug structure was downloaded in batch from PubChem, and then it was converted into a topological fingerprint using the Rdkit tool in Python; the topological fingerprint is generated based on the topological structure and rotation angle of the four-membered ring in the molecule and is used to describe the stereochemistry and interactions of the molecule; and the Tanimoto coefficient was used as a metric to obtain the and the drug chemical structure similarity values of the two drugs .

3. The intelligent prediction method for adverse drug reactions with multi-source data fusion according to claim 1, wherein The drug similarity based on drug-gene interactions refers to: Drug-gene interactions are represented by one-hot encoding, i.e., each unique gene interaction is converted into a binary vector; drugs and drugs Based on the similarity of drug-gene interactions Calculated using the Jaccard index, the formula is: ; Among them, and respectively represent gene sets that have interaction relationships with drugs and ​ 4. The intelligent prediction method for adverse drug reactions with multi-source data fusion according to claim 1, wherein The drug similarity of gene expression differences after drug perturbation refers to: Drug and Based on the similarity of gene expression differences after drug perturbation Calculated by the cosine similarity formula, and the formula is: ; Among them, and represent the drug i or j the gene expression feature vector after perturbation, and represent the norm of the vector and .

5. The intelligent prediction method for adverse drug reactions by multi-source data fusion according to claim 1, wherein The semantic similarity of different disease ontologies is obtained through the following steps: For each adverse drug reaction, the study constructed a directed acyclic graph based on its hierarchical descriptors in MeSH, where nodes represent the nodes of the adverse reaction in MeSH and edges represent the relationship between the current node and its ancestors; the adverse reaction s is represented as a graph , where represents the set of all ancestor nodes including the node of the adverse reaction s itself, and represents the set of all edges pointing from parent nodes to child nodes included; to encode the semantics of disease ontology terms in a measurable format, the study defined the semantic value of disease A as the total contribution of all diseases to the semantics of disease A in , where disease terms closer to disease A contribute more to its semantics; for two different disease ontologies A and B, the semantic similarity between A and B is defined as: 。 6. The intelligent prediction method for adverse drug reactions with multi-source data fusion according to claim 1, characterized in that The set of gene-disease association relationships is obtained through the following steps: Downloaded the compound-gene interaction table CTD_genes_diseases.csv.gz from CTD as the set of gene-disease association relationships in the knowledge graph, containing 107,911,805 gene-disease association records; the gene-disease associations in CTD include associations extracted and inferred by experts; the extracted associations are curated by CTD curators from published literature or exported from the OMIM database using the mim2gene file in the NCBI gene database; there are three types of evidence for direct gene-disease associations: M markers, mechanism, and treatment evidence; M markers refer to specific variations in gene sequences; mechanism evidence includes how gene variations affect the structure and function of proteins and how these changes interfere with normal cell or physiological processes; treatment evidence refers to that the symptoms of the disease can be improved or the occurrence of the disease can be prevented by treating specific genes or their products.

7. The intelligent prediction method for adverse drug reactions with multi-source data fusion according to claim 1, characterized in that The classifier consists of two fully connected layers, two activation function layers, and one BN layer; finally, a vector is output, where numbers greater than 0 indicate the association between the drug and the ADR, and the larger the number, the greater the likelihood of the association; the output vector satisfies the following relationship: ; Among them, represents a fully connected layer, where * represents the number of neurons, and Dropout represents a Dropout layer with a probability of p; and represent a linear function and a rectified linear unit activation function respectively, and CAT concatenates the given feature vectors.

8. An intelligent prediction device for adverse drug reactions with multi-source data fusion as described in claim 1, characterized in that, Including: A data acquisition module for acquiring multi-source data; wherein the multi-source data includes comparative toxicogenomics data, LINCS L1000 data, PubChem data, Medical Subject Headings data, and SIDER data; A feature construction module for constructing drug descriptor features and obtaining a drug similarity vector set after standardizing the multi-source data; constructing drug adverse reaction descriptor features and obtaining a drug adverse reaction similarity vector set; the drug similarity vector set includes: the chemical structure similarity of two drugs, the drug similarity based on drug-gene interaction, and the drug similarity of gene expression differences after drug perturbation; the drug adverse reaction similarity vector set includes: the semantic similarity of different disease ontologies and the gene-disease association relationship set; A prediction module for training the target drug adverse reaction intelligent prediction model based on the drug similarity vector set and the drug adverse reaction similarity vector set, and after training is completed, predicting the adverse reaction generated when the target drug is used for the target individual based on the target drug adverse reaction intelligent prediction model; Among them, the target adverse drug reaction intelligent prediction model is trained through three branches respectively by combining the original drug features , feature crossing , and the original adverse drug reaction features combination. Among them, and use two linear sub-networks with consistent structures for learning, while uses a convolutional neural network sub-network for learning; the architecture of the linear sub-network consists of two fully connected layers, a batch normalization layer, and an activation function layer; that is: ; Among them, In and respectively represent the potential representations of m drugs with k different similarity features and m drug adverse reactions, that is, the output of LSN; among them, and represent two linearly consistent sub-networks; represents a fully connected layer, where * represents the number of neurons; Dropout represents a Dropout layer with a probability of p; and represent a linear function and a rectified linear unit activation function respectively, and CAT concatenates the given feature vectors; The intelligent prediction model for the target adverse drug reaction is formed by concatenating the linear embedding vector , the final output and the linear embedding vector and using the resulting vector as the input vector to a multi-label classifier for classification.

Citation Information

Patent Citations

  • Adverse drug reaction prediction method based on multi-source data

    CN118366685A