CrossMama-based multi-modal drug interaction intelligent prediction method and system
By combining selective state-space models and the cross-modal fusion model CrossMamba with social media text and drug SMILES molecular structures, the problems of long sequence modeling and computational complexity in drug interaction prediction are solved, achieving efficient and accurate drug interaction prediction.
Patent Information
- Application Number
- CN202510978964.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-10-31
AI Technical Summary
Existing drug interaction prediction methods struggle to efficiently model long sequence dependencies, suffer from high computational complexity, and fail to effectively integrate SMILES long sequences with textual semantics, resulting in insufficient prediction performance for small datasets and complex drug pairs.
By employing a selective state-space model (SSM) combined with the cross-modal fusion model CrossMamba, feature learning and feature fusion are performed using embedded data from social media texts and drug SMILES molecular structures. A hybrid attention mechanism is used to enhance the feature vectors, enabling intelligent prediction of drug interactions.
It improves the accuracy and generalization ability of drug interaction prediction, reduces computational complexity, enhances the ability to collaboratively analyze unstructured text and chemical information, and provides prediction results with greater clinical practical value.
Smart Images

Figure CN120877913A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the interdisciplinary fields of bioinformatics, computational biology, and natural language processing, and specifically relates to a multimodal drug interaction intelligent prediction method and system based on CrossMamba. Background Technology
[0002] Current methods for predicting drug interactions (DDI) mainly rely on drug molecular structures (such as SMILES) or biomedical text data, but they have the following limitations: SMILES-based deep learning methods (such as GNN and Transformer) are difficult to efficiently model long sequence dependencies and have high computational complexity (Shanwen Zhang, et al. SCATrans: semantic cross-attention transformer for drug–drug interaction predication through multimodal biomedical data, BMC Bioinformatics, 2025, 26, 175); while text mining methods based on social media (such as medical forums and electronic health records) can supplement real-world evidence, they face problems such as high noise and semantic sparsity. Existing methods for fusing multimodal data (such as cross-modal attention) are typically computationally expensive (Vefghi A, Rahmati Z, Akbari M. Drug-target interaction / affinity prediction: Deep learning models and advancements review. Comput Biol Med., 2025, 196: 110438.), and do not adequately address the global-local collaborative modeling problem of long SMILES sequences and textual semantics, resulting in insufficient prediction performance for small datasets and complex drug pairs.
[0003] Unlike traditional models, SSM (Short-Time Sequence Modeling) maintains linear time complexity when processing long sequences by utilizing an efficient selective state-space model (Albert Gu, TriDao. SSM: Linear-Time Sequence Modeling with Selective State Spaces.arXiv:2312.00752,2024.). It effectively handles long-range dependencies in sequential data. Currently, there is no published literature on using SSM for DDI (Driving Distance Injection) prediction methods. Therefore, there is an urgent need for an efficient and lightweight cross-modal framework to jointly optimize the complementary information of drug structural features and social media text. Summary of the Invention
[0004] To overcome the shortcomings of the prior art, the present invention aims to provide a multimodal drug interaction intelligent prediction method and system based on CrossMamba, which has better prediction performance and wider applicability, high prediction rate and low computational complexity.
[0005] To achieve the above objectives, the present invention is implemented through the following technical solution:
[0006] A multimodal drug interaction intelligent prediction method based on CrossMamba includes the following steps:
[0007] Step 1: Perform data preprocessing and data embedding on social media texts containing two drugs and SMIELS molecular structures of the two drugs to obtain DDI embedding data; the embedding data includes embedding data of social media texts and embedding data of SMIELS molecular structures of the two drugs respectively.
[0008] Step 2: Based on the embedded data of social media texts and the embedded data of the SMIELS molecular structures of the two drugs, feature learning is performed using the selective state space model (SSM) to obtain the SSM feature vectors of the social media texts and the SSM feature vectors of the two drugs respectively. Then, the SSM feature vectors of the social media texts are concatenated with the SSM feature vectors of the two drugs respectively to obtain the multimodal feature vectors of the two drugs.
[0009] Step 3: Use the CrossMamba cross-modal fusion model to fuse the multimodal features of the two drugs to obtain the fused features of the two drugs.
[0010] Step 4: Use the predictive decision model to classify the fusion characteristics of the two drugs and obtain the prediction results of DDI.
[0011] Step 1 involves the data embedding process for social media texts. First, noise is removed from the social media texts, and the data is cleaned by standardizing terms. Then, the pre-trained text model Twitter-BERT is used to embed the social media texts, as shown below:
[0012] E t =TwitterBERT(T)
[0013] Where T represents the input social media text, TwitterBERT(.) is the Twitter-BERT utility function used for embedding social media text data, and E t Let T be a low-dimensional dense vector, which serves as the input for the subsequent SSM;
[0014] The Twitter-BERT pre-trained text model is used to convert the unstructured preprocessed text into a first preliminary low-dimensional dense vector. Then, the first preliminary low-dimensional dense vector representations are linearly combined to obtain the embedding vectors for each of the two drugs. The Twitter-BERT preprocessing process is as follows:
[0015] Drug-related noise filtering: retain drug mentions, side effect keywords, and patient reports, and replace non-drug-related noise;
[0016] Standardization of pharmaceutical terminology: mapping slang to standard terms and unifying spelling variations;
[0017] Word segmentation: Processes compound drug names using a WordPiece segmenter enhanced with an expanded drug dictionary and Drugbank domain;
[0018] Add drug domain tags: [DRUG] to tag drug entities, [EFFECT] to tag effect descriptions (e.g., "taking [DRUG] aspirin [EFFECT] reduced fever"), and retain the original BERT tags ([CLS], [SEP]).
[0019] The embedding layer output of Twitter-BERT needs to be aligned with the embedding data representations of the SMILES for both drugs; therefore, cross-modal consistency needs to be considered during the embedding process.
[0020] Token embedding: The embedding of drug-related tokens is initialized as the corresponding vector in the drug knowledge graph;
[0021] Location embedding: capturing temporal relationships in drug descriptions;
[0022] Differentiate embeddings: differentiate between user descriptions and drug facts;
[0023] Cross-modal alignment: Add a lightweight adaptation layer (such as linear projection) after embedding to map the text vector space to a semantic space similar to the SMILES vectors.
[0024] Step 1 involves embedding the SMILES molecular structures of the two drugs into a second preliminary low-dimensional dense vector. This process involves several key steps, combining prior knowledge from the chemical field for structured representation learning. The specific steps include:
[0025] Chemical standardization: The RDKit tool was used to unify the SMILES representation of the two drugs, removing chiral markers and standardizing bond order, etc.
[0026] Validity check: Filter out unparseable SMILES (such as mismatched brackets);
[0027] Basic word segmentation: Segment by atom ("C", "N"), bond type ("=", "#"), ring marker ("1", "2"), merge common groups (such as "C=O", "NH2") into a single token, reduce sequence length, and add [CLS] (sequence start), [SEP] (separate two drugs SMILES), and [PAD] (padded to uniform length);
[0028] Mapping discrete symbols to continuous vectors generates initial atomic-level representations, including:
[0029] Token embedding: Each atom / subword is assigned a learnable $d_{emb}$-dimensional vector, which can be initialized by incorporating chemical priors (such as atomic number, electronegativity);
[0030] Position embedding: Encodes the position of an atom in the sequence (e.g., the first "C" → $P_1$, the second "C" → $P_2$);
[0031] Differentiate embeddings: differentiate the SMILES of two drugs (e.g., Drug_A is labeled 0, Drug_B is labeled 1);
[0032] Output the embedding vectors of the molecular structures of the two drugs SMILES; denoted as:
[0033] E d1 =RDKit(SMIELS1)
[0034] E d2 =RDKit(SMIELS2)
[0035] Here, SMIELS1 and SMIELS2 are the SMIELS molecular structures of the two drugs, RDKit(.) is the RDKit utility function used for SMIELS molecular structure data embedding, E d1 and E d2 To obtain low-dimensional dense vectors of the two drugs as input for subsequent SSM;
[0036] The text vector output by Twitter-BERT needs to interact with the embedding vectors of SMILES for the two drugs. Layer normalization is performed on the text vector and SMILES vector respectively to ensure scale consistency between modalities.
[0037] Step 2, which utilizes the Selective State-Space Model (SSM) for feature learning, can be represented as follows: Let E be the embedded data of the social media text obtained in Step 1. t The embedding data of the SMIELS molecular structures of the two drugs are respectively E d1 and E d2The embedded data is used to extract features using a selective state-space model (SSM) to obtain E. t E d1 and E d2 The feature vectors are as follows:
[0038]
[0039] in, and E respectively t E d1 and E d2 Feature vectors, Mamba t and Mamba d These are feature vectors obtained from social media text embedding data and two types of drug embedding data through the Selective State-Space Model (SSM).
[0040] Feature vector concatenation involves concatenating the SSM feature vectors of the social media text with the SSM feature vectors of the two drugs, respectively, to obtain the multimodal feature vectors of each drug, as shown below:
[0041]
[0042] Among them, E td1 and E td2 These represent the feature vectors concatenated from the social media text and the two drugs, respectively. Concat(.) represents the feature vector concatenation operation.
[0043] In step 3, the feature vectors of the two drugs are fused using the hybrid attention mechanism in the CrossMamba model. The feature fusion process is represented as follows: Feature enhancement is performed through the Selective State Space Model (SSM) and global and local scanning methods to obtain the fused feature vector. The hybrid attention mechanism layer includes a hybrid attention layer vector generation module, a bilinear transformation module, a self-attention score and output vector module, and a feature vector fusion module. After passing through the CrossMamba model, the fused feature vector of the two drugs is obtained, denoted as F. out ; E through the CrossMamba model td1 and E td2 The process of feature fusion is represented as follows:
[0044]
[0045] Where i = 1, 2, 3, 4 represent four scanning methods, Crosssscan(·) represents four-way cross-scan, and CS6(·) is the state mapping space model of SSM, used to map the high-dimensional selective state space to the low-dimensional output space to obtain the feature vector y. i reversescan(·) is yi The features of the original sequence structure need to be obtained through a cross-scanning back-scanning process.
[0046] Step 4, the DDI prediction process, involves classifying the fusion features of the two drugs using a predictive decision model to obtain the DDI prediction result. The predicted probability of DDI is then obtained through the Sigmoid function, expressed as follows:
[0047] F td =W·F out +b
[0048] P DDI =Sigmoid(F td )
[0049] Where W and b are the weight matrix and bias parameter, respectively, Sigmoid(.) is the Sigmoid function, and F td For the output of the fully connected layer, P DDI This represents the predicted probability of DDI.
[0050] A multimodal drug interaction intelligent prediction system based on CrossMamba is implemented to realize a multimodal drug interaction prediction method based on CrossMamba. The system includes a data layer, a data embedding layer, a feature extraction layer, a feature fusion layer and a DDI prediction layer connected in sequence.
[0051] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0052] By combining the Selective State-Space Model (SSM) and a cross-modal interaction mechanism, this invention innovatively integrates multimodal data from social media text and the molecular structures of drug SMILES. This overcomes the limitations of traditional methods that rely on structured databases, significantly improving the accuracy and generalization ability of drug-drug interaction (DDI) predictions and avoiding the risk of drug-drug interactions. Simultaneously, the efficient sequence modeling capability of SSM enhances computational efficiency, while the cross-modal fusion mechanism strengthens the model's ability to collaboratively analyze unstructured text and chemical information, making the prediction results more clinically valuable and providing more flexible and reliable technical support for drug safety monitoring and personalized treatment. Therefore, this invention has significant application value and social benefits. Attached Figure Description
[0053] Figure 1 This is a flowchart illustrating the method of an embodiment of the present invention.
[0054] Figure 2 This is a structural diagram of the feature fusion model CrossMamba according to an embodiment of the present invention.
[0055] Figure 3 This is an overall structural diagram of the system according to an embodiment of the present invention. Detailed Implementation
[0056] The present invention will now be described in further detail with reference to the embodiments and accompanying drawings.
[0057] Reference Figure 1 A multimodal drug interaction intelligent prediction method based on CrossMamba includes the following steps:
[0058] Step 1: Perform data preprocessing and embedding on social media texts containing the two drugs and the SMIELS molecular structures of the two drugs to obtain DDI embedding data. The embedding data includes the embedding data of the social media texts and the embedding data of the SMIELS molecular structures of the two drugs respectively. The specific process is as follows:
[0059] Data embedding of social media text; noise removal and terminology standardization for data cleaning; data embedding using the pre-trained Twitter-BERT model for social media text, represented as follows:
[0060] E t =TwitterBERT(T)
[0061] Where T represents the input social media text, TwitterBERT(.) is the Twitter-BERT utility function used for embedding social media text data, and E t Let T be a low-dimensional dense vector, which serves as the input for the subsequent SSM;
[0062] Then, the unstructured pre-trained text is converted into a first preliminary low-dimensional dense vector using the Twitter-BERT pre-trained model. This first preliminary low-dimensional dense vector representation is then linearly combined to obtain the embedding vectors for the two drugs. The Twitter-BERT preprocessing process is as follows:
[0063] Drug-related noise filtering: retain drug mentions, side effect keywords, and patient reports, and replace non-drug-related noise;
[0064] Standardization of pharmaceutical terminology: mapping slang to standard terms and unifying spelling variations;
[0065] Word segmentation: Processes compound drug names using a WordPiece segmenter enhanced with an expanded drug dictionary and Drugbank domain;
[0066] Add drug domain tags: [DRUG] to tag drug entities, [EFFECT] to tag effect descriptions (e.g., "taking [DRUG] aspirin [EFFECT] reduced fever"), and retain the original BERT tags ([CLS], [SEP]).
[0067] The embedding layer output of Twitter-BERT needs to be aligned with the embedding data representations of the SMILES for both drugs; therefore, cross-modal consistency needs to be considered during the embedding process.
[0068] Token embedding: The embedding of drug-related tokens is initialized as the corresponding vector in the drug knowledge graph;
[0069] Location embedding: capturing temporal relationships in drug descriptions;
[0070] Differentiate embeddings: differentiate between user descriptions and drug facts;
[0071] Cross-modal alignment: Add a lightweight adaptation layer (such as linear projection) after embedding to map the text vector space to a semantic space similar to the SMILES vectors;
[0072] (2) Embedding data of the SMILES molecular structures of the two drugs; The embedding process of converting the SMILES molecular structures of the two drugs into a second preliminary low-dimensional dense vector involves several key steps, combining prior knowledge in the field of chemistry for structured representation learning. The specific process includes:
[0073] Chemical standardization: The RDKit tool was used to unify the SMILES representation of the two drugs, removing chiral markers and standardizing bond order, etc.
[0074] Validity check: Filter out unparseable SMILES (such as mismatched brackets);
[0075] Basic word segmentation: Segment by atom ("C", "N"), bond type ("=", "#"), ring marker ("1", "2"), merge common groups (such as "C=O", "NH2") into a single token, reduce sequence length, and add [CLS] (sequence start), [SEP] (separate two drugs SMILES), and [PAD] (padded to uniform length);
[0076] Mapping discrete symbols to continuous vectors generates initial atomic-level representations, including:
[0077] Token embedding: Each atom / subword is assigned a learnable $d_{emb}$-dimensional vector, which can be initialized by incorporating chemical priors (such as atomic number, electronegativity);
[0078] Position embedding: Encodes the position of an atom in the sequence (e.g., the first position "C" → $P_1$, the second position "C" → $P_2$).
[0079] Differentiate embeddings: differentiate the SMILES of two drugs (e.g., Drug_A is labeled 0, Drug_B is labeled 1);
[0080] Output the embedding vectors of the molecular structures of the two drugs SMILES; denoted as:
[0081] E d1 =RDKit(SMIELS1)
[0082] E d2 =RDKit(SMIELS2)
[0083] Here, SMIELS1 and SMIELS2 are the SMIELS molecular structures of the two drugs, RDKit(.) is the RDKit utility function used for SMIELS molecular structure data embedding, E d1 and E d2 To obtain low-dimensional dense vectors of the two drugs as input for subsequent SSM;
[0084] The text vector output by Twitter-BERT needs to interact with the embedding vectors of SMILES for the two drugs. The text vector and SMILES vector are respectively normalized to ensure scale consistency between modalities.
[0085] Step 2: Based on the embedded data of social media text and the embedded data of the two drugs, feature learning is performed using the Selective State Space Model (SSM) to obtain the SSM feature vectors of the social media text and the two drugs respectively. Then, the SSM feature vectors of the social media text are concatenated with the SSM feature vectors of the two drugs to obtain the multimodal feature vectors of each drug. The specific process is described as follows:
[0086] The feature learning process using the Selective State-Space Model (SSM) is represented as follows: Let E be the embedded data of the social media text obtained in step 1. t The embedding data of the SMIELS molecular structures of the two drugs are respectively E d1 and E d2 Feature extraction was performed using the Selective State-Space Model (SSM): To reveal the interaction between two drugs, the embedded data E... t E d1 and E d2Contextual information from the embedded data captured in Mamba was used for feature extraction. Mamba employed a Selective State-Space Model (SSM) to identify the internal structure and relationships of the social media text of the drug and the SMIELS molecular structures of the two drugs using word embeddings and positional embeddings. Word embeddings assigned semantic information to each word in the input social media text of the drug and the SMIELS molecular structures of the two drugs, extracting contextual information from the sequence and aiding in understanding the relationships between words. Positional embeddings provided positional features for each word, helping the model identify the relative positions of words in the sequence, enabling the model to effectively extract the rich information contained in the social media text of the drug and the SMIELS molecular structures of the two drugs. The feature vectors of the social media text and the SMIELS molecular structures of the two drugs, obtained by feature extraction using SSM, are expressed by the following formula:
[0087]
[0088] in, and E respectively t E d1 and E d2 Feature vectors, Mamba t and Mamba d These are feature vectors obtained by SSM from social media text embedding data and two types of drug embedding data, respectively.
[0089] Feature vector concatenation involves concatenating the SSM feature vectors of the social media text with the SSM feature vectors of the two drugs, respectively, to obtain the multimodal feature vectors of each drug, as shown below:
[0090]
[0091] Among them, E td1 and E td2 These represent the feature vectors concatenated from the social media text and the two drugs, respectively, with Concat(.) representing the concatenation operation.
[0092] Step 3: Use the CrossMamba model to fuse the multimodal features of the two drugs to obtain the fused features of the two drugs. The specific process is described as follows:
[0093] Reference Figure 2In step 3, the hybrid attention mechanism in the CrossMamba model is used to fuse the feature vectors of the two drugs. The feature fusion process is represented as follows: After selective state-space model (SSM) and global and local scanning methods, feature enhancement is performed to obtain the fused feature vector. The hybrid attention mechanism layer includes a hybrid attention layer vector generation module, a bilinear transformation module, a self-attention score and output vector module, and a feature vector fusion module. After passing through the CrossMamba model, the fused feature vector of the two drugs is obtained, denoted as F. out ; E through the CrossMamba model td1 and E td2 The process of feature fusion is represented as follows:
[0094]
[0095] Where i = 1, 2, ..., 4 represent four scanning methods, Crosssscan(·) is a four-way cross-scan, and CS6(·) is the state mapping space model of SSM, used to map the high-dimensional selective state space to the low-dimensional output space to obtain the feature vector y. i reversescan(·) is y i The features of the original sequence structure need to be obtained through a cross-scanning reverse scan process;
[0096] Step 4: Classify the fusion characteristics of the two drugs using a predictive decision model to obtain the prediction results of DDI. The specific process is described below:
[0097] The final DDI prediction probability is obtained through the Sigmoid function and is expressed as:
[0098] F td =W·F out +b
[0099] P DDI =Sigmoid(F td )
[0100] Where W and b are the weight matrix and bias parameter, respectively, Sigmoid(.) is the Sigmoid function, and F td For the output of the fully connected layer, P DDI This represents the predicted probability of DDI.
[0101] Reference Figure 3 A multimodal drug interaction intelligent prediction system based on CrossMamba is proposed, which realizes a multimodal drug interaction intelligent prediction method based on CrossMamba. The system includes a data layer, a data embedding layer, a feature extraction layer, a feature fusion layer and a DDI prediction layer connected in sequence.
[0102] To verify the technical effectiveness of this invention, a comparative experiment was conducted on a public dataset to verify the technical solution of this invention with five other existing DDI prediction methods. The five methods and the method of this invention were trained and tested on Nvidia RTX 3090 GPUs. The main software environment was CUDA 11.8 and Python 3.8, the deep learning framework was PyTorch 1.13.0, and the Adam optimizer was used to train the model. The training parameters were set as follows: learning rate of 0.001, weight decay of 0.0001, batch size of 14, and epochs of 100. The experimental results are shown in Table 1.
[0103] Table 1. Experimental results of the five methods
[0104]
[0105] In Table 1, MDNN represents a multimodal deep neural network, PTBiLST represents a pre-trained bidirectional short-long-term memory network with a labeler, SRR represents a fine-grained substructure representation model, MSMDL represents a multi-layer soft-mask dual-vision learning model, and TransformerDDI represents a combined Transformer and LSTM model. The results in Table 1 show that the DDI obtained by this invention has high accuracy and a short training time. The main reason is that this invention, based on the sequence modeling capability of SSM, can efficiently capture the long-term dependencies between drug molecule structures (SMILES) and social media text, overcoming the shortcomings of traditional deep convolutional networks and Transformer networks, and improving the efficiency of DDI based on the method of this invention. The results in Table 1 demonstrate the feasibility of this invention.
Claims
1. A multimodal drug interaction intelligent prediction method based on CrossMamba, characterized in that, Includes the following steps: Step 1: Perform data preprocessing and data embedding on social media texts containing two drugs and SMIELS molecular structures of the two drugs to obtain DDI embedding data; the embedding data includes embedding data of social media texts and embedding data of SMIELS molecular structures of the two drugs respectively. Step 2: Based on the embedded data of social media texts and the embedded data of the SMIELS molecular structures of the two drugs, feature learning is performed using the selective state space model (SSM) to obtain the SSM feature vectors of the social media texts and the SSM feature vectors of the two drugs respectively. Then, the SSM feature vectors of the social media texts are concatenated with the SSM feature vectors of the two drugs respectively to obtain the multimodal feature vectors of the two drugs. Step 3: Use the CrossMamba cross-modal fusion model to fuse the multimodal features of the two drugs to obtain the fused features of the two drugs. Step 4: Use the predictive decision model to classify the fusion characteristics of the two drugs and obtain the prediction results of DDI.
2. The prediction method according to claim 1, characterized in that, Step 1, the data embedding process for social media text, includes removing noise from the social media text, standardizing terminology to clean the data, and using the pre-trained text model Twitter-BERT to embed the social media text, as shown below: E t =TwitterBERT(T) Where T represents the input social media text, TwitterBERT(.) is the Twitter-BERT utility function used for embedding social media text data, and E t To obtain a low-dimensional dense vector of T, which will serve as the input for subsequent SSM.
3. The prediction method according to claim 1, characterized in that, The data embedding process of the SMILES molecular structures of the two drugs in step 1 specifically includes: standardizing each SMILES string, then performing character-level word segmentation, mapping each atom, bond, and symbol to a unique ID, and generating a corresponding mask sequence; then converting the ID into a preliminary low-dimensional dense vector through the embedding layer, while incorporating positional encoding to maintain the sequence order, and adding segment encoding to distinguish different drugs, and then inputting the ID sequence into the embedding layer to convert it into a dense vector, represented as follows; E d1 =RDKit(SMIELS1) E d2 =RDKit(SMIELS2) Here, SMIELS1 and SMIELS2 are the SMIELS molecular structures of the two drugs d1 and d2 respectively, RDKit(.) is the RDKit utility function used for SMIELS molecular structure data embedding, E d1 and E d2 To obtain low-dimensional dense vectors for the two drugs, which will then be used as inputs for subsequent SSM.
4. The prediction method according to claim 3, characterized in that, The process of feature learning using the Selective State-Space Model (SSM) in step 2 is represented as follows: Let E be the embedded data of the social media text obtained in step 1. t The embedding data of the SMIELS molecular structures of the two drugs are respectively E d1 and E d2 The embedded data is learned through a selective state-space model (SSM) to obtain E. t E d1 and E d2 The feature vectors are as follows: in, and E respectively t E d1 and E d2 Feature vectors, Mamba t and Mamba d These are feature vectors obtained by SSM from social media text embedding data and two types of drug embedding data, respectively. Feature vector concatenation involves concatenating the SSM feature vectors of the social media text with the SSM feature vectors of the two drugs, respectively, to obtain the multimodal feature vectors of each drug, as shown below: Among them, E td1 and E td2 These represent the feature vectors concatenated from the social media text and the two drugs, respectively. Concat(.) represents the concatenation operation of the two feature vectors.
5. The prediction method according to claim 4, characterized in that, In step 3, the hybrid attention mechanism in the CrossMamba model is used to fuse the feature vectors of the two drugs, resulting in feature fusion. The process is as follows: Feature enhancement is performed through SSM, global and local scanning methods to obtain a fused feature vector; the hybrid attention mechanism layer includes a hybrid attention layer vector generation module, a bilinear transformation module, a self-attention score and output vector module, and a feature vector fusion module. After passing through the CrossMamba model, the fused feature vector of the two drugs is obtained, denoted as F. out ; E through the CrossMamba model td1 and E td2 The process of feature fusion is represented as follows: Where i = 1, 2, ..., 4 represent four scanning methods, Crosssscan(·) is a four-way cross-scan, and CS6(·) is the Selective State-Space Model (SSM) structure of the SSM, used for feature extraction to obtain feature y. i reversescan(·) is y i The features of the original sequence structure need to be obtained through a cross-scanning back-scanning process.
6. The prediction method according to claim 1, characterized in that, The DDI prediction process of the prediction decision model in step 4 is as follows: the final DDI prediction probability is obtained through the Sigmoid function, expressed as: F td =W·F out +b P DDI =Sigmoid(F td ) Where W and b are the weight matrix and bias parameter, respectively, and F out Let F be the feature vector fused from the two drugs, and let Sigmoid(.) be the Sigmoid function. td For the output of the fully connected layer, P DDI This represents the predicted probability of DDI.
7. A multimodal drug interaction intelligent prediction system based on CrossMamba, characterized in that: The system for implementing the intelligent prediction method for multimodal drug interactions based on CrossMamba as described in any one of claims 1-6 comprises a data layer, a data embedding layer, a feature extraction layer, a feature fusion layer, and a DDI prediction layer connected in sequence.