Adverse drug reaction prediction method and system based on multi-source heterogeneous modality dual-pathway fusion interaction
Through the multi-source heterogeneous mode dual-path fusion interaction method, the local and global characteristics of the drug chemical structure are dynamically integrated, and the problem of insufficient mode feature capture of drug adverse reaction prediction in the existing technology is solved, and efficient and accurate drug adverse reaction prediction is achieved, which is suitable for drug research and development and adverse reaction monitoring.
Patent Information
- Application Number
- CN202510624985.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-05-15
AI Technical Summary
Existing methods for predicting adverse drug reactions cannot dynamically capture different modal features in drug chemical structures, and cannot fully capture the potential semantic information of different modal features, resulting in limited comprehensiveness and accuracy of prediction.
The multi-source heterogeneous mode dual-path fusion interaction method is adopted, and efficient information interaction and synergistic prediction between modes is achieved by constructing a gated deep convolution spatial fusion module (GDConv-SF) and self-attention mechanism, combined with self-supervised learning, and dynamically fuse local and global features in the medicinal chemical structure.
It realizes lightweight drug adverse reaction prediction, improves prediction performance, and can surpass existing methods under low parameters. It is suitable for safety evaluation in the drug development stage and monitoring of potential adverse reactions of prescription drugs, meeting the requirements of real-time drug safety screening.
Smart Images

Figure CN120126815B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal data fusion and self-supervised learning, and in particular to a multi-source heterogeneous modality dual-pathway fusion interaction and a drug adverse reaction prediction method and system. Background Art
[0002] Adverse drug reactions (ADRs) are unexpected and potentially harmful reactions experienced by patients, even when a drug is used according to recommended dosage and quality. Early detection of potential adverse reactions during the drug development cycle not only significantly improves drug safety but also effectively reduces R&D risks and costs.
[0003] With the enrichment of biomedical information, research on computer-aided prediction of ADRs has made significant progress. In ADRs prediction, the extraction and characterization of drug features are key to model performance. Commonly used drug features are mainly divided into two categories: (1) Drug intrinsic information: including drug chemical substructures, molecular descriptors, etc. These features can directly reflect the chemical properties and biological activities of drugs. (2) Drug-related entity information: such as drug-target protein interactions, drug indications and drug pathways, etc. These features reveal the functions and potential mechanisms of drugs in biological systems. The scope of drug features has gradually expanded from basic chemical structures to include multi-source information such as drug-drug similarity and drug-disease association.
[0004] Computer-assisted methods are primarily categorized as those based on traditional machine learning and deep learning. Traditional machine learning methods typically rely on multiple drug features and predict adverse reactions by building correlations or similarities between these features. While traditional machine learning methods have achieved some success in ADR prediction, their reliance on prior knowledge, limitations in feature extraction, and insufficient generalization have limited their further development. In contrast, deep learning technology, with its capabilities for automatic feature extraction and complex relationship modeling, exhibits significant advantages and has gradually become a research hotspot in ADR prediction. Most existing deep learning-based ADR prediction methods rely on information about drug-associated entities, ignoring intrinsic drug characteristics and prone to cold-start challenges caused by missing exogenous data. For example, a meta-path graph neural network-based adverse drug reaction prediction model proposed in Chinese patent document CN115512857A requires constructing a heterogeneous information network using four entities: drug, protein, adverse reaction, and disease. This results in these methods being more inclined to identify potential unknown adverse reactions of known drugs rather than accurately predicting the risks of novel drugs. In the initial stages of drug discovery, prior knowledge of candidate drugs is usually relatively scarce, and researchers can often only rely on limited chemical structure properties to build predictive models, which greatly limits the performance of the models.
[0005] Meanwhile, existing methods for multimodal fusion strategies generally suffer from static fusion flaws. For example, the paper "Toward Unified AI Drug Discovery with Multimodal Knowledge" [J]. Yizhen Luo, Xing Yi Liu, et al. Health Data Science. February 23, 2024, proposes a deep learning framework, KEDD, for predicting ADRs. This framework achieves fusion by splicing features from four modalities: drug, protein, structured knowledge, and unstructured knowledge (assuming that each modality contributes equally to prediction), ignoring the inherent functional heterogeneity between different modalities.
[0006] As a key breakthrough direction in the field of artificial intelligence, multimodal fusion technology aims to achieve complementary integration and redundancy suppression of heterogeneous data features. Current research faces two core challenges in the design of cross-modal interaction mechanisms: (1) Limitations of static fusion methods: Traditional methods (such as direct concatenation, mean fusion, or weighted summation) generally adopt fixed weight distribution strategies. These static fusion methods lead to the dilution of key information during the fusion process, and are particularly difficult to handle the representation differences between sequence modalities (such as text, time series) and graph modalities (such as molecular structure, social network). For example, in the scheme of drug target affinity prediction method based on sequence modality and graph modality proposed in Chinese patent document CN119091974A, the features of the two modalities are fused by additive fusion, which cannot dynamically adapt the feature importance levels of different modalities. (2) Limitations of a single fusion mechanism: Sequence data and graph modality data have significant differences in representation form, feature dimension, and intrinsic semantic expression. Existing methods often employ a single fusion strategy (such as relying solely on attention mechanisms or tensor concatenation), which often fails to fully capture high-order interactions and deep semantic associations between modalities. This limits the expressive power of fused features and makes it difficult to adapt to the changing importance of modal features in different task scenarios. For example, the multimodal data scene recognition method based on multi-level interactive fusion proposed in Chinese patent document CN115878983A interactively fuses sequence data and video modality data. However, this method uses a multi-layer, multi-head self-attention network and a two-stage attention model, which significantly increases computational complexity. The introduction of a large number of learnable parameters makes it difficult for this method to meet millisecond-level response requirements.
[0007] In summary, existing methods for predicting adverse drug reactions (ADRs) are unable to dynamically capture the features of different modalities within a drug's chemical structure, nor can they fully capture the underlying semantic information of these modal features and conduct deep fusion and interaction. This significantly limits the comprehensiveness and accuracy of predictions. Therefore, it is urgent to develop an innovative ADR prediction method based solely on intrinsic drug information and integrating these modal features with interactive mechanisms. Summary of the Invention
[0008] The technical problem addressed by this invention is to provide a method and system for predicting adverse drug reactions using a multi-source, heterogeneous, dual-pathway fusion interaction. This system employs a novel dual-pathway fusion interaction mechanism to dynamically capture and integrate local functional groups and global features within a drug's chemical structure. This system captures the latent semantic information of adverse reaction categories through self-supervised learning, thereby enabling collaborative prediction of ADR probabilities.
[0009] In order to solve the above technical problems, the present invention adopts the following technical solutions:
[0010] In a first aspect, the present invention provides a multi-source heterogeneous modality dual-pathway fusion interaction method for predicting adverse drug reactions, which specifically includes the following steps:
[0011] S1. Multi-source heterogeneous modal feature extraction and representation optimization, constructing feature engineering for multi-source heterogeneous data; drug-related multi-source heterogeneous data include: drug SMILES sequence data, drug molecular fingerprint data and drug molecular structure graph data; in the scheme, drug SMILES sequence and drug molecular fingerprint are drug-related sequence data modalities; drug molecular structure graph is drug-related graph data modality.
[0012] S2, dual-pathway fusion interactive collaborative prediction, specifically including:
[0013] S21. Build a dual-pathway: Build a gated depthwise convolution with spatial fusion module (GDConv-SF module) as the first path to achieve local feature fusion in space; build a mean strategy as the second path to reduce the impact of noise and redundant information;
[0014] S22, Dual-pathway Interaction: First, self-attention is calculated on the features fused through the first path. The purpose is to establish long-range dependencies within the features and enhance representation capabilities. Then, two unidirectional cross-attention mechanisms are used to capture the attention weights and feedback information between the multi-source heterogeneous modal features fused through different paths, realizing interaction between modalities.
[0015] S23, collaborative prediction: The modal features obtained in S22 are layer-normalized and processed by multi-layer perceptron layers, and the results are used as the encoding after multi-source heterogeneous modal fusion interaction; the latent semantic information of the adverse drug reaction category is captured through self-supervised learning, and the adverse drug reaction embedding matrix is constructed, which is then matched with the encoding results after fusion interaction, and finally the predicted probability is output.
[0016] Furthermore, the spatial fusion module (GDConv-SF module) of the gated depth convolution in S21 includes the following structure:
[0017] The first linear layer is used to perform a preliminary linear transformation on the input features, in order to provide more accurate feature representation for the subsequent deep convolutional layer;
[0018] The feature splitting layer is used to split the linearly transformed features into two sub-feature blocks along the feature dimension. Its purpose is to enable the model to process different feature representations separately, thereby capturing finer-grained local information;
[0019] The deep convolution layer is used to perform a deep convolution operation on one of the sub-feature blocks. Its purpose is to capture the local correlation between different modal features and achieve spatial local feature fusion;
[0020] Dynamic gated residual is used to multiply the output of another sub-feature block and the output of the deep convolution layer element by element after being activated by the Silu function. Its purpose is to dynamically adjust the impact of the residual on the final output;
[0021] The second linear layer is used to linearly transform the features after the dynamic gated residual connection to generate the final output representation.
[0022] Furthermore, the spatial fusion module of gated depth convolution (GDConv-SF module) is implemented by the following expression:
[0023] SubfeatureBlock1,SubfeatureBlock2 = FeatureSplit(Linearl(Input));
[0024] Sub-feature block 1 = Activation(Reshape(depth_convolution(Reshape(sub-feature block 1))));
[0025] SubfeatureBlock2 = Activation(SubfeatureBlock2);
[0026] Output = Linear2(Multiply(Sub-feature Block 1, Sub-feature Block 2)).
[0027] In the above formula, Activation represents the Silu activation function, Multiply represents element-by-element multiplication, and Linear1 and Linear2 represent two different linear layers respectively.
[0028] Furthermore, the method for constructing the adverse drug reaction embedding matrix described in step 23 is as follows: each adverse drug reaction is embedded as a standardized random vector with an initial mean of 0 and a standard deviation of 0.1, and the final embedding matrix dimension is [number of adverse reactions, 8 × hidden layer dimension]; as a learnable parameter, the parameter can be updated during the training process.
[0029] Since the true labels of adverse reactions are not directly used for training, the embedding matrix can capture the latent semantic information of adverse reaction categories.
[0030] Furthermore, in S23, the correlation between the features after multi-source heterogeneous modality fusion and each adverse reaction is calculated through matrix multiplication, and the inner product is mapped to a decimal between 0 and 1 through the Sigmoid function, and the predicted probability is finally output.
[0031] Furthermore, the multi-source heterogeneous modal feature extraction in S1 includes:
[0032] S11. Linear symbol vectorization: Specifically, the drug SMILES sequence is linearly symbolized and vectorized using the molecular substructure vector representation method.
[0033] Furthermore, the multi-source heterogeneous modal feature extraction in S1 also includes:
[0034] S12. Molecular fingerprint embedding: refers to encoding drug molecular fingerprints and performing multi-resolution signal processing.
[0035] Specifically, a bit compression algorithm based on hexadecimal character mapping converts the traditional 1024-bit drug molecular fingerprint into a compact 256-bit code. This code is then combined with a discrete wavelet transform (DWT) for multi-resolution signal processing. To overcome the limitations of symbolic encoding in terms of feature resolution, a DWT-driven multi-resolution signal processing strategy was introduced, achieving a balanced optimization between reducing the feature space dimension and preserving key chemical information.
[0036] Furthermore, the bit compression algorithm of the hexadecimal character mapping in step S12 includes the following steps: (1) regrouping the 1024-bit binary molecular fingerprint according to 16 hexadecimal characters and compressing them into 64 hexadecimal characters; (2) performing entropy weighted encoding on each hexadecimal character to generate a 256-bit compact vector.
[0037] Furthermore, the multi-source heterogeneous modal feature extraction in S1 also includes:
[0038] S13. Use different graph embedding models to perform multi-angle topological encoding on the drug molecular structure graph. The purpose is to encode the molecular structure graph features from different angles, each focusing on a specific structural level. Preferably, the Attentive FP, MPNN, and NFGNN graph embedding models are used to perform multi-angle topological encoding on the drug molecular structure graph.
[0039] Furthermore, S1 also includes a bidirectional gated recurrent unit (B-GRU) to optimize the representation of each modality feature extracted by S11-S13, the purpose of which is to capture bidirectional contextual dependencies.
[0040] The Bidirectional Gated Recurrent Unit (B-GRU) is a bidirectional recurrent neural network architecture based on a gating mechanism. Through joint forward-backward modeling, it effectively captures potential bidirectional contextual dependencies in sequence data (such as the potential contextual relationships in molecular structures). Compared to other sequence encoders, the B-GRU significantly reduces temporal computational complexity when processing shorter sequence data through a gated state sharing mechanism.
[0041] In a second aspect, the present invention further provides a multi-source heterogeneous modality dual-pathway fusion interaction adverse drug reaction prediction system, comprising a processor and a memory, wherein the memory stores computer program code instructions;
[0042] When the computer program code instructions are called by the processor, the processor is caused to execute the above-mentioned multi-source heterogeneous modality dual-pathway fusion interaction method for predicting adverse drug reactions.
[0043] The application scenarios of the multi-source heterogeneous modality dual-pathway fusion interactive drug adverse reaction prediction method and system provided by the present invention include: safety evaluation in the drug development stage; and potential adverse reaction monitoring of prescription drugs after they are approved for marketing.
[0044] Explanation of terms in the present invention:
[0045] Mol2vec (Unsupervised Machine Learning Approach with Chemical Intuition) is an unsupervised machine learning method for generating vector representations of molecular substructures.
[0046] Attentive FP (Pushing the Boundaries of Molecular Representation for Drug Discovery with the Graph Attention Mechanism) is a graph neural network model based on a graph attention mechanism, primarily used for molecular characterization and drug discovery. This attention mechanism dynamically assigns interaction weights between atoms, enabling the localization of key pharmacophores (such as key binding groups in enzyme active centers). However, its preference for local features can weaken the generalized representation of the overall molecular topology.
[0047] MPNNs (Message Passing Neural Networks) are a type of graph neural network (GNN) framework specifically designed for processing graph-structured data. By implicitly aggregating global neighborhood information through multiple rounds of message passing, they can simulate the effects of long-range chemical interactions on physicochemical properties (such as logP). However, they can easily overlook significant substructure signals such as functional groups and ring systems.
[0048] NFGNN (Convolutional Networks on Graphs for Learning Molecular Fingerprints) is a molecular representation learning method based on graph convolutional neural networks (GCNs). This method can directly process molecular structure graphs and generate chemically meaningful molecular representations through end-to-end learning. By explicitly extracting neighborhood features of a preset radius through layered convolution, it can directly match structural alerts in known toxicity databases (such as the genotoxicity risk of nitroaromatic rings). However, its fixed-radius neighborhood partitioning makes it difficult to adapt to dynamic conformational changes.
[0049] The present invention has the following beneficial effects:
[0050] The present invention provides a method and system for predicting adverse drug reactions using dual-pathway fusion interaction of multi-source heterogeneous modalities. The method comprises two phases: multi-source heterogeneous modal feature extraction and characterization optimization, and dual-pathway fusion interaction collaborative prediction. In the multi-source heterogeneous modal feature extraction and characterization optimization phase, feature engineering is performed for multi-source heterogeneous data. In the dual-pathway fusion interaction collaborative prediction phase, hierarchical fusion and cross-modal interactive learning between multi-source heterogeneous modalities are then implemented.
[0051] This invention utilizes only the chemical structural properties of drugs as input. Through the proposed GDConv-SF module, it constructs a multi-perspective composite embedding space for linear sequence-space distribution, enabling the complementarity of intrinsic information between sequence vectors and topological encodings. The dual-pathway fusion interaction mechanism designed in this invention dynamically integrates local functional groups and global features, enabling efficient information exchange and collaborative optimization between data from different modalities. Ultimately, cross-modal collaborative prediction is achieved based on the degree of match between the fused and interactive feature representations and the embeddings for adverse reactions.
[0052] The present invention is simple to implement and easy to operate. Compared with the latest ADRs prediction method, the present invention surpasses the performance indicators with only 24.85% of the parameters. It is lightweight and greatly improves the performance indicators. The present invention has the potential to break through the limitations of traditional experimental data and can be used to explore adverse reactions that have not yet been observed in approved clinical drugs. For more than 70% of the drugs using the prediction results obtained by this method, relevant researchers can refine the results through limited manual verification links to improve the clinical applicability of the final output results. The lightweight architecture of this method can better meet the strict requirements of real-time drug safety screening on delay. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 This is a flow chart of the multi-source heterogeneous modality dual-pathway fusion interaction adverse drug reaction prediction method (abbreviated as MHMDPFI-ADRsP) provided by the present invention.
[0054] Figure 2 This is the overall framework diagram of MHMDPFI-ADRsP in an embodiment of the present invention.
[0055] Figure 3 This is a structural diagram of the spatial fusion module of the gated depthwise convolution (GDConv-SF module for short) provided in an embodiment of the present invention.
[0056] Figure 4 This is a comparison chart of the model complexity-performance balance between MHMDPFI-ADRsP of an embodiment of the present invention and other methods.
[0057] Figure 5 This is a graph showing the results of the MHMDPFI-ADRsP prediction of drug quantities within different error intervals. DETAILED DESCRIPTION
[0058] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0059] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.
[0060] Example 1:
[0061] like Figure 1-2 As shown, this embodiment provides a multi-source heterogeneous modality dual-pathway fusion interaction adverse drug reaction prediction method (hereinafter referred to as MHMDPFI-ADRsP), which specifically includes the following steps:
[0062] S1. Multi-source heterogeneous modal feature extraction and representation optimization, and feature engineering for multi-source heterogeneous data.
[0063] The multi-source heterogeneous drug-related data used in this example includes: drug SMILES sequence data, drug molecular fingerprint data, and drug molecular structure graph data. In this solution, drug SMILES sequences and drug molecular fingerprints are drug-related sequence data modalities; drug molecular structure graphs are drug-related graph data modalities.
[0064] The method for extracting multi-source heterogeneous modal features specifically includes:
[0065] S11. Linear Symbol Vectorization: As a preferred embodiment, Mol2vec, a molecular substructure vector representation method, is used to vectorize the linear symbol SMILES of the drug, thereby obtaining a vector representation of the substructure. In this way, the vector representation of the substructure is used to construct a representation of the entire drug chemical structure.
[0066] S12. Molecular Fingerprint Embedding: To address the data sparsity issues inherent in traditional high-dimensional binary vectors (1024 bits) of molecular fingerprints, this embodiment uses a bit compression algorithm based on hexadecimal character mapping to convert the traditional 1024-bit drug molecular fingerprint into a compact 256-bit code. Discrete Wavelet Transform (DWT) is also introduced for multi-resolution signal processing. To overcome the limitations of symbol encoding in feature resolution, a DWT-driven multi-resolution signal processing strategy is employed. Based on the maximum variance contribution criterion, 128 significant coefficients are retained, achieving a balanced optimization between reducing the feature space dimension and preserving key chemical information.
[0067] The bit compression algorithm of hexadecimal character mapping specifically includes the following steps: (1) regrouping the 1024-bit binary molecular fingerprint according to 16 hexadecimal characters and compressing it into 64 hexadecimal characters; (2) performing entropy weighted encoding on each hexadecimal character to generate a 256-bit compact vector.
[0068] S13. Molecular topology encoding: Use different embedding models to perform multi-angle topological encoding on the drug molecular structure diagram.
[0069] The drug molecular structure graph in this embodiment of the present invention is constructed as follows: RDKit is used to construct a molecular structure graph G = (V, E) for the molecular smile sequence. The node set V consists of atomic attributes (atom type, charge, hybridization state, etc.), and the edge set E represents chemical bond types (single bonds, double bonds, conjugated state, etc.) and stereochemical information, accurately expressing the topological connectivity and stereochemical characteristics of the molecule. RDkit (Open-source cheminformatics and machine learning) is an open-source cheminformatics toolkit for analyzing and visualizing chemical data. RDkit is primarily used in drug design, bioactivity prediction, chemical reaction prediction, and chemical data processing. As a preferred embodiment, this embodiment of the present invention employs three different graph embedding models: Attentive FP, MPNN, and NFGNN to perform multi-angle topological encoding on the constructed drug molecular structure graph.
[0070] Because the use of a single model is limited by its feature aggregation method and can only capture structural information in a certain dimension, it may lead to misjudgment of adverse reactions due to incomplete representation. Therefore, the embodiment of the present invention uses the three different graph embedding models to perform multi-angle topological encoding on the drug molecular structure graph, encoding the molecular structure graph features from different angles, each focusing on a specific structural level.
[0071] The bidirectional gated recurrent unit (B-GRU) then optimizes the representation of the features extracted from each modality in S11-S13 to capture bidirectional contextual dependencies. Compared to other sequence encoders, B-GRU can significantly reduce temporal computational complexity when processing shorter sequence data through a gated state sharing mechanism.
[0072] S2, dual-path fusion interactive collaborative prediction, realizes hierarchical fusion and cross-modal interactive learning between multi-source heterogeneous modalities. Specifically including:
[0073] S21. Dual-pathway construction: We propose a gated depthwise convolution with spatial fusion module (GDConv-SF module) and an averaging strategy to achieve fusion of different pathways. Specifically, we construct the GDConv-SF module as the first pathway to achieve spatial local feature fusion, and the averaging strategy as the second pathway to reduce the impact of noise and redundant information.
[0074] The GDConv-SF module in this embodiment constructs a multi-view composite embedding space of linear sequence-space distribution for heterogeneous data features, which can complement the intrinsic information between different modalities and generate a high-quality feature representation space. Figure 3 As shown, the GDConv-SF module of the embodiment of the present invention specifically includes the following structure:
[0075] The first linear layer is used to perform a preliminary linear transformation on the input features, in order to provide more accurate feature representation for the subsequent deep convolutional layer;
[0076] The feature splitting layer is used to split the linearly transformed features into two sub-feature blocks along the feature dimension. Its purpose is to enable the model to process different feature representations separately, thereby capturing finer-grained local information;
[0077] The deep convolution layer is used to perform a deep convolution operation on one of the sub-feature blocks. Its purpose is to capture the local correlation between different modal features and achieve spatial local feature fusion;
[0078] Dynamic gated residual is used to multiply the output of another sub-feature block and the output of the deep convolution layer element by element after being activated by the Silu function. Its purpose is to dynamically adjust the impact of the residual on the final output;
[0079] The second linear layer is used to linearly transform the features after the dynamic gated residual connection to generate the final output representation.
[0080] The GDConv-SF module is implemented by the following expression:
[0081] SubfeatureBlock1,SubfeatureBlock2 = FeatureSplit(Linearl(Input));
[0082] Sub-feature block 1 = Activation(Reshape(depth_convolution(Reshape(sub-feature block 1))));
[0083] SubfeatureBlock2 = Activation(SubfeatureBlock2);
[0084] Output = Linear2(Multiply(Sub-feature Block 1, Sub-feature Block 2)).
[0085] In the above formula, Activation represents the Silu activation function, Multiply represents element-by-element multiplication, and Linear1 and Linear2 represent two different linear layers respectively.
[0086] In the second-path fusion, mean fusion is performed on the different modal features after splicing in the channel dimension, which effectively suppresses noise and information redundancy interference and significantly improves the stability and robustness of multi-source heterogeneous modal representation.
[0087] S22. Dual-Path Interaction: A bidirectional cross-attention mechanism is introduced to achieve dual-path interaction. First, self-attention is calculated on the features after the first-path fusion in the GDConv-SF module to establish long-range dependencies within the features and enhance representation capabilities. Then, two unidirectional cross-attention mechanisms are used to capture the attention weights and feedback information between the multi-source heterogeneous modal features fused through different paths, achieving comprehensive interaction between the modalities.
[0088] S23, collaborative prediction: The modal features obtained in S22 are layer-normalized and processed by multi-layer perceptron layers, and the results are used as the encoding after multi-source heterogeneous modal fusion interaction; the latent semantic information of the adverse drug reaction category is captured through self-supervised learning, and the adverse drug reaction embedding matrix is constructed, which is then matched with the encoding results after fusion interaction, and finally the predicted probability is output.
[0089] The adverse drug reaction embedding matrix is constructed as follows: each adverse drug reaction is embedded as a normalized random vector with an initial mean of 0 and a standard deviation of 0.1. The final embedding matrix has dimensions [number of adverse reactions, 8 × hidden layer dimension]; this is used as a learnable parameter and updated during training. Since the true labels of adverse reactions are not directly used in training, the embedding matrix captures the latent semantic information of the adverse reaction category. Matrix multiplication is then used to calculate the correlation between the features fused from multi-source heterogeneous modalities and each adverse reaction. The inner product is then mapped to a decimal between 0 and 1 using a sigmoid function, and the predicted probability is finally output.
[0090] Parameter settings in this embodiment:
[0091] During the model training process, ten-fold cross-validation was used to construct the model training framework, and the training set-validation set was split based on the preset 95%-5% ratio.
[0092] During the training process, a segmented learning rate decay strategy is implemented: the initial learning rate is set to 0.001, and according to the exponential decay mechanism, the rate decay is performed by 50% every 10 epochs.
[0093] The model is trained end-to-end in each fold, and a strict dual termination criterion is implemented: (1) a hard upper limit: the maximum iteration epoch is set to 200; (2) a dynamic early stopping mechanism: training is terminated early when the performance indicators of the validation set do not improve within 70 consecutive epochs.
[0094] The embedding dimensions of Mol2vec and DWT are 200 and 128, respectively; the embedding dimensions of AFP, MPNN and NFGNN are all 617; the number of hidden units of B-GRU is 32; the feature map dimension of GDConv-SF is 128; the hidden layer dimension of MLP is 256; the batch size of the training set and the test set is 8; by analyzing the sensitivity results of the three performance indicators with the changes of L2 regularization coefficient and hidden layer dimension, the L2 regularization coefficient and hidden layer dimension are set to 1e-5 and 64, respectively; Adam is used to iteratively update the parameters in the model; the adverse reaction embedding matrix with dimension [number of adverse reaction categories, 8×hidden layer dimension] is a learnable parameter; the composite form of binary cross entropy loss and Sigmoid activation function is used as the loss value of the entire model for backpropagation.
[0095] Training module: Initializes the weights and biases of the linear layers and layer normalization in the network and performs model training. The indicator results of the validation set are used to save the optimal model weights in each iteration cycle.
[0096] The test module loads the trained model weights. The model outputs a single drug adverse reaction prediction value in the dimension [1, number of adverse reaction categories]. This value is Sigmoid activated and then compared with a threshold of 0.5. Finally, a binary sequence with the dimension of the adverse reaction category number indicates whether the value is present (1 for presence, 0 for absence).
[0097] Example 2
[0098] Based on the same inventive concept, this embodiment is a system embodiment corresponding to the above-mentioned method embodiment 1, and can be implemented in conjunction with the above-mentioned embodiment 1.
[0099] This embodiment provides a drug adverse reaction prediction system with multi-source heterogeneous modal dual-pathway fusion interaction, including a processor and a memory, wherein the memory stores computer program code instructions; when the computer program code instructions are called by the processor, the processor executes the drug adverse reaction prediction method with multi-source heterogeneous modal dual-pathway fusion interaction as described in Example 1.
[0100] The computer program code includes: a feature extraction and optimization subroutine, a dual-path cross-modal fusion interaction calculation subroutine, and a prediction subroutine; the code is written using Python 3.8 and above and the PyTorch framework and is compatible with CUDA acceleration hardware.
[0101] The relevant technical details mentioned in the above embodiment 1 are still valid in this embodiment, and the repeated parts are not repeated here.
[0102] The application scenarios of the multi-source heterogeneous modality dual-pathway fusion interactive drug adverse reaction prediction method and system provided by the present invention include: safety evaluation in the drug development stage; and potential adverse reaction monitoring of prescription drugs after they are approved for marketing.
[0103] In order to verify the feasibility and effect of the scheme of the present invention, it is further illustrated by the following examples.
[0104] Example 3 Performance Comparison
[0105] To verify the superior performance of the MHMDPFI-ADRsP of the present invention, the MHMDPFI-ADRsP provided in Example 1 was applied to a high-quality association dataset for ADRs prediction. This dataset contains 1,020 drug entities, 5,598 standardized adverse reactions, and 133,750 verified positive association samples. When the drug-adverse reaction adjacency matrix (1020x5598≈5.71x10 6 When the number of potential association combinations is quantified, the proportion of positive samples is only 2.34% (the sparsity reaches 97.66%).
[0106] In this example, the performance of the MHMDPFI-ADRsP algorithm provided by the present invention is compared with the latest baseline algorithm MSDSE and four classic algorithms: idse-HE, MNMF, MSVD, and GMF. The following is a brief introduction to each algorithm:
[0107] MSDSE (Predicting drug-side effects based on multi-scale features and deep multi-structure neural network) algorithm: adopts a multi-scale feature learning strategy, realizes local-global feature fusion through the designed multi-receptive field convolutional neural network (CNN) module Inception and the use of a multi-head self-attention mechanism.
[0108] The idse-HE (Hybrid embedding graph neural network for drug side effects prediction) algorithm combines graph embedding and node embedding to transform drug side effect prediction into a matrix completion task of low-dimensional latent factor space projection, and reconstructs the association network through bidirectional linear mapping.
[0109] MNMF (Discovering Links Between Side Effects and Drugs Using a Difusion-Based Method) algorithm: reconstructs the drug-side effect adjacency matrix through non-negative matrix factorization (NMF) and innovatively implements heat diffusion propagation in the drug semantic similarity network.
[0110] MSVD algorithm: an improved MNMF method that uses TruncatdSVD instead of NMF.
[0111] GMF (Neural Collaborative Filtering) algorithm: Generalized Matrix Factorization (GMF) reconstructs the neural form of matrix decomposition, embedding matrix factorization (MF) into the neural network framework for the first time, fusing linear factor interactions with deep nonlinear features.
[0112] A comprehensive selection of evaluation metrics that are sensitive to sparsity and clinically interpretable includes: 1. Area under the precision-recall curve (AUPR): measures the recall stability of the model under high false positive suppression requirements; 2. Matthews correlation coefficient (MCC): a balance metric based on the global statistics of the confusion matrix, suppressing the dominant effect of the majority class (negative pairs) on the score; 3. F1-score: the harmonic mean of the rising and falling trends of precision and recall.
[0113] In addition, the unit parameter effectiveness coefficient (PEC, defined as: AUPR × 100 / total number of model parameters (10 6 )) and built a quantitative evaluation system for model performance and computational cost. These four indicators have no linear correlation with each other; higher values indicate better prediction results.
[0114] The experimental results are shown in Table 1 below.
[0115]
[0116] The comparison results in Table 1 show that the proposed prediction method (MHMDPFI-ADRsP) achieves optimal values for AUPR, MCC, and F1-score compared to mainstream methods for other species. Compared to the state-of-the-art method, MSDSE, these three metrics improve by 4.8%, 6.6%, and 6.7%, respectively. This significant performance improvement is primarily attributed to the design of the GDConv-SF module and its dual-pathway fusion interaction mechanism, which enables two-fold mutual reinforcement between features from different modalities. This allows the proposed MHMDPFI-ADRsP method to more balancedly address the dual optimization challenges of balancing precision and recall (AUPR) and classification consistency (MCC and F1-score).
[0117] according to Figure 4 The parameter efficiency-performance balance results fully demonstrate that compared with the two cutting-edge methods of IDSE-HE and MSDSE, the MHMDPFI-ADRsP method provided by the present invention has significant advantages in eliminating parameter redundancy and optimizing feature encoding, achieving the optimal parameter efficiency coefficient, and better meeting the strict delay requirements of real-time drug safety screening.
[0118] Example 4 Prediction of adverse drug reactions
[0119] A test set (containing 102 drugs and involving 20-844 types of adverse reactions) contained in a single fold was randomly selected to perform method validation and visual analysis on the MHMDPFI-ADRsP provided in the embodiment of the present invention.
[0120] Figure 5 The distribution of the number of drugs in different error intervals for ADRs predicted by the MHMDPFI-ADRsP and true positive samples is shown. The Include region and Exclude region represent the number of drugs included in the statistics and excluded from the previous statistical interval at the current error threshold, respectively.
[0121] The model achieved accurate predictions for seven drugs (6.86%), with the predicted ADRs set being "identical" to the true set. The corresponding adverse reaction types for these seven drugs were 21, 71, 77, 188, 453, 453, and 823, respectively. This demonstrates that the range of adverse reaction types predicted by the MHMDPFI-ADRsP is not limited to a specific range, demonstrating its robust predictive power. Within the 40% error threshold, 73 drugs (71.57%) were effectively predicted. Of these, approximately 58.82% of the drugs had a discrepancy of less than 35%. The error distribution showed significant clustering in the 20%-30% range, encompassing 37 drugs (36.28%). Therefore, for over 70% of the drugs predicted using the MHMDPFI-ADRsP, researchers can refine the results through limited manual verification to enhance the clinical applicability of the final output.
[0122] In fact, the latest clinical studies have shown that the drug "Amlexanox (an anti-inflammatory and anti-allergic immunomodulator)" (with a variance of less than 25%) exhibits potential adverse reactions such as "skin and subcutaneous tissue disorders," "nervous system disorders," and "hepatobiliary function diagnostic procedure." Notably, the experimentally recorded ADRs dataset did not include these positive samples, yet the MHMDPFI-ADRsP proposed in this paper successfully predicted these latent adverse reactions. Similarly, the model demonstrated similar predictive power for the adverse reactions of "metipranolol (a beta-adrenergic receptor blocking drug)" (with a variance of less than 25%) and "valganciclovir" (with a variance of less than 30%). This demonstrates that the MHMDPFI-ADRsP proposed in this paper has the potential to transcend the limitations of traditional experimental data and can be used to discover undiscovered adverse reactions of known drugs.
[0123] The above descriptions are only some preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A multi-source heterogeneous modality dual-pathway fusion interaction method for predicting adverse drug reactions, characterized by: The specific steps include: S1. Multi-source heterogeneous modal feature extraction and representation optimization, building feature engineering for multi-source heterogeneous data; Drug-related multi-source heterogeneous data includes: drug SMILES sequence data, drug molecular fingerprint data, and drug molecular structure diagram data; multi-source heterogeneous modal feature extraction includes: S11. Linear symbol vectorization: Specifically, the drug SMILES sequence is linearly symbolized and vectorized using the molecular substructure vector representation method; S12. Molecular fingerprint embedding: This refers to encoding drug molecular fingerprints and performing multi-resolution signal processing. Specifically, a bit compression algorithm based on hexadecimal character mapping is used to convert the traditional 1024-bit drug molecular fingerprint into a 256-bit compact code, and multi-resolution signal processing is performed in combination with discrete wavelet transform. S13, using different graph embedding models to perform multi-angle topological encoding of drug molecular structure graphs, each focusing on a specific structural level; The bidirectional gated recurrent unit is used to characterize and optimize the modal features extracted from S11 to S13; S2, dual-pathway fusion interactive collaborative prediction, specifically including: S21, dual-path fusion: The first path fusion is achieved through the spatial fusion module of gated deep convolution, and the second path fusion is achieved using the mean strategy; the spatial fusion module of gated deep convolution includes the following structure: The first linear layer is used to perform a preliminary linear transformation on the input features; Feature splitting layer, used to split the linearly transformed features into two sub-feature blocks along the feature dimension; The depth convolution layer is used to perform a depth convolution operation on one of the sub-feature blocks; Dynamic gated residual, used to perform element-wise multiplication of another sub-feature block and the output of the deep convolutional layer after being activated by the Silu function; The second linear layer is used to linearly transform the features after the dynamic gated residual connection to generate the final output representation; In the second channel fusion, mean fusion is performed on the different modal features after channel dimension splicing to suppress noise and information redundancy interference; S22, Dual-pathway interaction: First, self-attention is calculated on the features fused by the first path; then, two unidirectional cross-attention mechanisms are used to capture the attention weights and feedback information between the multi-source heterogeneous modal features fused by different paths, thus achieving interaction between modalities. S23, collaborative prediction: The modal features obtained in S22 are layer-normalized and processed by multi-layer perceptron layers, and the results are used as the encoding after multi-source heterogeneous modal fusion interaction; the latent semantic information of the adverse drug reaction category is captured through self-supervised learning, and the adverse drug reaction embedding matrix is constructed, which is then matched with the encoding results after fusion interaction, and finally the predicted probability is output.
2. The multi-source heterogeneous modality dual-pathway fusion interaction method for predicting adverse drug reactions according to claim 1 is characterized in that: The method for constructing the adverse drug reaction embedding matrix described in step S23 is as follows: each adverse drug reaction is embedded as a standardized random vector with an initial mean of 0 and a standard deviation of 0.1, and the final embedding matrix dimension is [number of adverse reactions, 8*hidden layer dimension]; it is used as a learnable parameter and the parameter is updated during the training process.
3. The multi-source heterogeneous modality dual-pathway fusion interaction method for predicting adverse drug reactions according to claim 1 is characterized in that: Specifically, in S23, the correlation between the features after multi-source heterogeneous modal fusion and each adverse reaction is calculated through matrix multiplication, and the inner product is mapped to a decimal between 0 and 1 through the Sigmoid function, and the predicted probability is finally output.
4. The multi-source heterogeneous modality dual-pathway fusion interaction method for predicting adverse drug reactions according to claim 1, characterized in that: The bit compression algorithm for the hexadecimal character mapping in step S12 includes the following steps: (1) regrouping the 1024-bit binary molecular fingerprint according to the hexadecimal characters and compressing them into 64 hexadecimal characters; (2) performing entropy weighted encoding on each hexadecimal character to generate a 256-bit compact vector.
5. A multi-source heterogeneous modality dual-pathway fusion interaction drug adverse reaction prediction system, characterized by: comprising a processor and a memory, wherein the memory has computer program code instructions stored thereon; When the computer program code instructions are called by the processor, the processor executes the multi-source heterogeneous modality dual-pathway fusion interaction drug adverse reaction prediction method as described in any one of claims 1 to 4 above.
Citation Information
Patent Citations
Drug side effect prediction model based on meta-path graph neural network
CN115512857A
Multi-modal data scene identification method based on multi-level interactive fusion
CN115878983A
Drug target affinity prediction method based on sequence modality and graph modality
CN119091974A
Weighted fusion method for importance difference of multi-source survey data of power grid
CN118627007A
Lightweight social platform image restoration method and system based on dual-branch frequency and space fusion, and readable storage medium
CN119599915A