DDI prediction method and system based on multiple view angles, and storage medium

By fusion of the interactive features of SMILES mode and molecular graph modes from multiple perspectives, the DDI prediction method is optimized, which solves the problems of high cost and inaccurate prediction of traditional methods, and achieves more efficient and accurate DDI prediction.

CN120048374APending Publication Date: 2025-05-27CENT SOUTH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510016100.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Traditional DDI prediction methods rely on a huge DDI dataset, are expensive and time-consuming, and cannot cover all possible drug combinations, especially for newly launched drugs, which affects the accuracy of prediction.

Method used

The multi-view DDI prediction method is adopted to fuse the SMILES mode and molecular graph modality of the drug through the SMILES-SMILES interaction perspective, SMILES-molecular graph interaction perspective and molecular graph-molecular graph interaction perspective, and extract the interaction characteristics, and then splice them with the Morgan molecular fingerprint and input it into a multi-layer perceptron for fusing, optimizing the DDI prediction accuracy of the DDI data set of a small number of samples.

Benefits of technology

Through multi-perspective fusion, the accuracy and efficiency of DDI prediction are improved, and the dependence on a large number of DDI data sets is reduced, which is suitable for the prediction of newly marketed drugs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048374A_ABST
    Figure CN120048374A_ABST
Patent Text Reader

Abstract

The invention relates to a DDI prediction method and system based on multiple view angles and a storage medium. The method comprises the steps that four small-amount sample DDI data sets under hot start and cold start settings are constructed according to drug reaction types; extracting SMILES substructure features and molecular graph substructure features of drugs in four small-amount sample DDI data sets; respectively calculating interaction characteristics between the substructure characteristics of the two drugs from the SMILES-SMILES interaction view angle, the SMILES-molecular graph interaction view angle and the molecular graph-molecular graph interaction view angle; and splicing the interaction features of the three interaction view angles and the Morgan molecular fingerprints of the two drugs, and inputting the spliced interaction features and the Morgan molecular fingerprints into a multi-layer perceptron for fusion to obtain a DDI prediction result. The SMILES modality and the molecular graph modality of the medicine are fused through three interaction perspectives, so that the DDI prediction precision of a DDI data set of a small number of samples is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pharmacology, and in particular, to a multi-perspective-based DDI prediction method, system, and storage medium. Background Art

[0002] Drug-drug interaction (DDI) is an important and complex issue in clinical practice. When two or more drugs act on a patient simultaneously, they may affect each other's efficacy or cause adverse reactions, which pose potential risks to the patient's health and life safety.

[0003] Traditional DDI prediction relies on clinical trials and pharmacological studies, which are not only costly but also unable to cover all possible drug combinations, especially for newly marketed drugs. With the development of medical big data and artificial intelligence technologies, it has become possible to use these technologies for DDI prediction. In particular, in recent years, with the unprecedented success of deep learning and graph neural networks (GNNs), their powerful capabilities in dealing with complex data relationships have provided new opportunities to improve the accuracy and efficiency of DDI prediction.

[0004] Traditional DDI prediction methods rely on a large DDI dataset for pre-training and then use it for prediction, which are not only costly but also time-consuming. More challenging is that with the introduction of a large number of new drugs every year, many drugs and their interaction patterns have not been learned by the model, thus affecting the prediction accuracy. Summary of the Invention

[0005] The present invention provides a multi-perspective-based DDI prediction method, system, and storage medium, which fuse the SMILES modality and molecular graph modality of drugs through the SMILES-SMILES interaction perspective, SMILES-molecular graph interaction perspective, and molecular graph-molecular graph interaction perspective, thereby optimizing the DDI prediction accuracy of a small-sample DDI dataset.

[0006] To solve the above technical problems, the technical solutions adopted by the present invention are as follows: In the first aspect, a multi-perspective-based DDI prediction method is provided, including: Constructing four small-sample DDI datasets under hot start and cold start settings according to drug reaction types; Extracting the SMILES substructure features and molecular graph substructure features of drugs in the four small-sample DDI datasets; Calculating the interaction features between the substructure features of two drugs from the SMILES-SMILES interaction perspective, SMILES-molecular graph interaction perspective, and molecular graph-molecular graph interaction perspective respectively; After splicing the interaction features from the SMILES-SMILES interaction perspective, the SMILES-molecular graph interaction perspective, and the molecular graph-molecular graph interaction perspective with the Morgan molecular fingerprints of two drugs, the input is fed into a multi-layer perceptron for fusion to obtain the DDI prediction result.

[0007] Furthermore, four small-sample DDI datasets under warm-start and cold-start settings are constructed according to drug reaction types, including: Obtain the Drugbank dataset, TWOSIDES dataset, unseen DDI dataset, and unseen single-drug dataset; Statistically count the corresponding reaction drug pair sets of the Drugbank dataset, TWOSIDES dataset, unseen DDI dataset, and unseen single-drug dataset according to drug reaction types; Select N reaction drug pairs from the Drugbank dataset, unseen DDI dataset, and unseen single-drug dataset as the training set, and select N / 10 reaction drug pairs from the TWOSIDES dataset as the training set; Select N / 3 reaction drug pairs from the Drugbank dataset, unseen DDI dataset, and unseen single-drug dataset as the validation set, and select N / 30 reaction drug pairs from the TWOSIDES dataset as the validation set; For the test sets of the Drugbank dataset and TWOSIDES dataset under the warm-start setting, both drugs exist in the reaction drug pairs in the training sets of the Drugbank dataset and TWOSIDES dataset; For the test set of the unseen DDI dataset under the cold-start setting, both drugs do not exist in the reaction drug pairs in the training set of the unseen DDI dataset; For the test set of the unseen single-drug dataset under the cold-start setting, one of the two drugs exists in the reaction drug pairs in the training set of the unseen single-drug dataset.

[0008] Furthermore, extract the SMILES substructure features and molecular graph substructure features of drugs in the four small-sample DDI datasets, including: Use a SMILES encoder to extract molecular features in the SMILES modality for drugs in the four small-sample DDI datasets to obtain SMILES substructure features; Use a molecular graph encoder to extract molecular features in the molecular graph modality for drugs in the four small-sample DDI datasets to obtain molecular graph substructure features.

[0009] Furthermore, the SMILES encoder is used to extract the molecular features of drugs in the four small-sample DDI datasets in the SMILES modality, obtaining the SMILES substructure features, including: The expression for the SMILES encoder to extract the molecular features of the SMILES modality using the MoLFormer model is: ; represents the drug i , represents the extracted SMILES substructure features, represents extracting the SMILES features of drugs using the MoLFormer model.

[0010] Furthermore, the molecular graph encoder is used to extract the molecular features of drugs in the four small-sample DDI datasets in the molecular graph modality, obtaining the molecular graph substructure features, including: The expression for the molecular graph encoder to extract the molecular features of the molecular graph modality using the GROVER model is: ; represents the drug i , represents the extracted molecular graph substructure features, represents extracting the molecular features of drugs using the GROVER model.

[0011] Furthermore, from the SMILES-SMILES interaction perspective, calculate the interaction features between the substructure features of two drugs, including: Map the SMILES substructure features and the molecular graph substructure features to the feature space of the same dimension size through the multi-layer perceptron MLP respectively, obtaining the SMILES feature vector and the molecular graph feature vector, the SMILES feature vector , the molecular graph feature vector ; Extract the importance matrix between the SMILES substructure features of two drugs in the SMILES-SMILES interaction perspective through the attention module, and the calculation formula of the importance matrix is: ; represents the drug i and the drug j the importance matrix of the SMILES substructure features; represents the co-attention module; Through the bidirectional interaction module, combine to obtain the drugi Each SMILES substructure in the molecule of j and the interaction feature matrix of each SMILES substructure in the molecule of the drug ; represents the interaction feature matrix, represents the embedded feature matrix of the drug response type; According to calculate the expression of the aggregated feature vector of drug i and drug j : ; represents the interaction feature matrix, represents the embedded feature matrix of the drug response type; According to calculate the expression of the aggregated feature vector of drug i and drug j : ; represents the unidirectional feature vector mapped after the interaction of each SMILES substructure of drug i with all SMILES substructures of drug j ; represents the unidirectional feature vector mapped after the interaction of each SMILES substructure of drug j with all SMILES substructures of drug i ; represents the feature vector after the interaction of each substructure of drug i with all SMILES substructures of drug j ; represents the feature vector after the interaction of each substructure of drug j with all SMILES substructures of drug i ; Connect and to obtain the expression of the interaction feature of drug i and drug j from the SMILES-SMILES interaction perspective: ; represents the bidirectional interaction feature vector between drug i and drug j .

[0012] Further, after splicing the interaction features from the SMILES-SMILES interaction perspective, the SMILES-molecular graph interaction perspective, and the molecular graph-molecular graph interaction perspective with the Morgan molecular fingerprints of the two drugs, they are input into a multi-layer perceptron for fusion to obtain the DDI prediction result, including: The interaction features from the SMILES-SMILES interaction perspective , the interaction features from the SMILES-molecular graph interaction perspective , the interaction features from the molecular graph-molecular graph interaction perspective , the drugs i 's Morgan molecular fingerprint features and the drugs j 's Morgan molecular fingerprint features are spliced to obtain a spliced feature vector, and the expression of the spliced feature vector is: ; After inputting into the MLP, the predicted value is calculated through the Sigmoid activation function, and the expression of the predicted value is: ; represents the predicted value of the drug i and the drug j having a drug reaction; After comparing with the preset threshold, the DDI prediction result is obtained.

[0013] Second, a multi-perspective DDI prediction system is provided, including: A data acquisition module for constructing four small-sample DDI datasets under hot start and cold start settings according to drug reaction types; A feature extraction module for extracting the SMILES substructure features and molecular graph substructure features of drugs in the four small-sample DDI datasets; A multi-perspective interaction module for calculating the interaction features between the substructure features of two drugs from the SMILES-SMILES interaction perspective, the SMILES-molecular graph interaction perspective, and the molecular graph-molecular graph interaction perspective respectively; A multi-perspective fusion DDI prediction template for splicing the interaction features from the SMILES-SMILES interaction perspective, the SMILES-molecular graph interaction perspective, and the molecular graph-molecular graph interaction perspective with the Morgan molecular fingerprints of the two drugs, and then inputting them into a multi-layer perceptron for fusion to obtain the DDI prediction result.

[0014] In a third aspect, a computer-readable storage medium is provided, storing a computer program which, when called by a processor, executes the steps of the multi-perspective DDI prediction method in the first aspect above.

[0015] Advantages achieved by the present invention: Four small-sample DDI datasets under warm start and cold start settings are constructed according to drug reaction types; SMILES substructure features and molecular graph substructure features of drugs in the four small-sample DDI datasets are extracted; interaction features between substructure features of two drugs are calculated from the SMILES-SMILES interaction perspective, the SMILES-molecular graph interaction perspective, and the molecular graph-molecular graph interaction perspective respectively; after the interaction features from the three interaction perspectives are concatenated with the Morgan molecular fingerprints of the two drugs, they are input into a multi-layer perceptron for fusion to obtain the DDI prediction result. The SMILES modality and molecular graph modality of drugs are fused through the SMILES-SMILES interaction perspective, the SMILES-molecular graph interaction perspective, and the molecular graph-molecular graph interaction perspective, thereby optimizing the DDI prediction accuracy of the small-sample DDI dataset. Description of the drawings

[0016] Figure 1 It is a flowchart of the multi-perspective DDI prediction method of the present invention; Figure 2 It is a structural diagram of the multi-perspective DDI prediction system of the present invention. Detailed implementation manners

[0017] The present invention will be further described below with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention and should not be used to limit the protection scope of the present invention.

[0018] As Figure 1 shown, an embodiment of the present invention provides a multi-perspective DDI prediction method, including: 101, constructing four small-sample DDI datasets under warm start and cold start settings according to drug reaction types; In this embodiment, Drugbank is a comprehensive bioinformatics and chemoinformatics database that combines detailed drug data and comprehensive drug target information, and is widely used in drug research and development and medical research. The DrugBank database contains information on more than 5,000 approved small molecule drugs, biotech drugs, nutraceuticals, and experimental drugs. Each drug entry has a detailed description, including chemical structure, molecular weight, pharmacological properties, etc. The DrugBank database also contains 5,236 non-redundant protein (i.e., drug target / enzyme / transporter / carrier) sequences associated with these drug entries, and this information helps to study the mechanism of action and targets of drugs. Each DrugCard entry in the DrugBank database contains more than 200 data fields, half of which are for drug / chemical data and the other half are for drug target or protein data. These fields provide identification information, pharmacological information, drug interaction information, clinical information, and target information of drugs; The Twosides dataset is a dataset for DDI prediction. Specifically, it is a negative sample dataset containing drug combinations. In DDI prediction, positive samples refer to known drug combinations, while negative samples refer to those combinations that have been experimentally verified not to interact. By collecting these negative samples, the Twosides dataset provides a reliable set of negative samples for researchers, which helps to improve the accuracy and generalization ability of DDI prediction models; Obtain the Drugbank dataset, TWOSIDES dataset, unseen DDI (S1) dataset, and unseen single drug (S2) dataset; Assume that the Drugbank, S1, and S2 datasets all have a total of 86 drug reaction types, while the TWOSIDES dataset has a total of 963 drug reaction types. Statistically analyze the sets of reaction drug pairs corresponding to the Drugbank dataset, TWOSIDES dataset, S1 dataset, and S2 dataset according to the drug reaction types: ; represents the drug reaction type, The reaction drug pair representing drug reaction type R is drug i and drug j ; The set of reaction drug pairs of the Drugbank dataset is ; The set of reaction drug pairs of the TWOSIDES dataset is ; The set of reaction drug pairs of the S1 dataset is ; The set of reaction drug pairs of the S2 dataset is ; Select N (60) reaction drug pairs from the Drugbank dataset, S1 dataset, and S2 dataset as the training set. N is less than the total number of reaction drug pairs. Select N / 10 (6) reaction drug pairs from the TWOSIDES dataset as the training set; Obtain The training sets of four small-sample DDI datasets; Select N / 3 (20) reaction drug pairs from the Drugbank dataset, S1 dataset, and S2 dataset as the validation set. Select N / 30 (2) reaction drug pairs from the TWOSIDES dataset as the validation set; Obtain

[0019] The validation sets of four small-sample DDI datasets; For the test sets of the Drugbank dataset and TWOSIDES dataset under the hot-start setting, both drugs exist in the reaction drug pairs in the training sets of the Drugbank dataset and TWOSIDES dataset; for the Drugbank and TWOSIDES datasets, according to the hot-start setting, only retain the reaction drug pairs in the validation set where both drugs have been seen in the training set; the test set of the Drugbank dataset is: ; The test set of the TWOSIDES dataset is: ; For the S1 dataset, according to the S1 setting, only retain the reaction drug pairs in the test set where neither of the two drugs has been seen in the training set; the test set of the S1 dataset is: ; For the S2 dataset, according to the S2 setting, only retain the reaction drug pairs in the test set where only one of the two drugs has been seen in the training set; the test set of the S2 dataset is: .

[0020] 102. Extract the SMILES substructure features and molecular graph substructure features of the drugs in the four small-sample DDI datasets; In this embodiment, there are a SMILES encoder and a molecular graph encoder; among them, the SMILES encoder uses the Molformer network, and its network architecture is the Transformer-XL model; the molecular graph encoder uses the GROVER model, which is fine-tuned based on the graph Transformer model; The expression for using the MoLFormer model by the SMILES encoder to extract molecular features in the SMILES modality is: ; represents a drug i , represents the extracted SMILES substructure features, indicating the extraction of SMILES features of drugs using the MoLFormer model; The expression for extracting molecular features in the molecular graph modality using the molecular graph encoder with the GROVER model is: ; represents a drug i , represents the extracted molecular graph substructure features, indicating the extraction of molecular features of drugs using the GROVER model.

[0021] 103, calculate the interaction features between the substructure features of two drugs from the SMILES-SMILES interaction perspective, SMILES-molecular graph interaction perspective, and molecular graph-molecular graph interaction perspective respectively; In this embodiment, calculate the interaction features between the substructure features of two drugs from the SMILES-SMILES interaction perspective, SMILES-molecular graph interaction perspective, and molecular graph-molecular graph interaction perspective respectively. Taking the SMILES-SMILES interaction perspective as an example, the calculation process of the interaction features is described as follows: Map the SMILES substructure features and molecular graph substructure features to a feature space of the same dimension size through a multi-layer perceptron MLP to obtain the SMILES feature vector and molecular graph feature vector, the SMILES feature vector , the molecular graph feature vector ; Extract the importance matrix between the SMILES substructure features of two drugs in the SMILES-SMILES interaction perspective through the attention module. The calculation formula of the importance matrix is: ; represents drug i and drug j 's importance matrix of SMILES substructure features; represents the co-attention module; Through the bidirectional interaction module, combine to obtain each SMILES substructure in the molecule of drug i and drug jThe interaction feature matrix of each SMILES substructure in the molecule, and the calculation formula of the interaction feature matrix is: ; represents the interaction feature matrix, represents the embedded feature matrix of the drug response type; According to calculate the expression of the aggregated feature vector of drug i and drug j : ; represents the one-way feature vector mapped after the interaction of each SMILES substructure of drug i with all SMILES substructures of drug j ; represents the one-way feature vector mapped after the interaction of each SMILES substructure of drug j with all SMILES substructures of drug i ; represents the feature vector after the interaction of each substructure of drug i with all SMILES substructures of drug j ; represents the feature vector after the interaction of each substructure of drug j with all SMILES substructures of drug i ; Connect and to obtain the expression of the interaction feature of drug i and drug j from the SMILES-SMILES interaction perspective: ; represents the bidirectional interaction feature vector between drug i and drug j ; Similarly, the interaction feature from the SMILES-molecular graph interaction perspective and the interaction feature from the molecular graph-molecular graph interaction perspective can be calculated.

[0022] 104, After splicing the interaction features from the SMILES-SMILES interaction perspective, the SMILES-molecular graph interaction perspective, and the molecular graph-molecular graph interaction perspective with the Morgan molecular fingerprints of the two drugs, input them into a multi-layer perceptron for fusion to obtain the DDI prediction result.

[0023] In this embodiment, the interaction features from the SMILES-SMILES interaction perspective , the interaction features from the SMILES-molecular graph interaction perspective , the interaction features from the molecular graph-molecular graph interaction perspective , the i Morgan molecular fingerprint features of the drug and the j Morgan molecular fingerprint features of the drug are concatenated to obtain a concatenated feature vector. The Morgan molecular fingerprint features of a drug are a coding method for describing the molecular structure and are widely used in the fields of chemoinformatics and drug discovery. Its basic principle is to map the molecular structure to a fixed-length bit string for easy computer processing and comparison; The concatenated feature vector is expressed as: ; After is input into the MLP, the predicted value is calculated through the Sigmoid activation function. The predicted value is expressed as: ; represents the predicted value of the drug i and the drug j having a drug reaction; Then is compared with the preset threshold t, as shown in the following formula: ; If is greater than or equal to t, the DDI prediction result I is 1, indicating a drug reaction. Otherwise, the DDI prediction result I is 0, indicating no drug reaction.

[0024] The implementation principle and beneficial effects of the embodiments of the present invention are as follows: Four small-sample DDI datasets under hot-start and cold-start settings are constructed according to drug reaction types; the SMILES substructure features and molecular graph substructure features of drugs in the four small-sample DDI datasets are extracted; the interaction features between the substructure features of two drugs are calculated from the SMILES-SMILES interaction perspective, the SMILES-molecular graph interaction perspective, and the molecular graph-molecular graph interaction perspective respectively; after splicing the interaction features from the three interaction perspectives with the Morgan molecular fingerprints of the two drugs, they are input into a multi-layer perceptron for fusion to obtain the DDI prediction result; the SMILES modality and molecular graph modality of drugs are fused through the SMILES-SMILES interaction perspective, the SMILES-molecular graph interaction perspective, and the molecular graph-molecular graph interaction perspective, thereby optimizing the DDI prediction accuracy of the small-sample DDI dataset. By processing the Drugbank dataset, TWOSIDES dataset, unseen DDI (S1) dataset, and unseen single-drug (S2) dataset, a small-sample DDI dataset is constructed. Compared with traditional DDI prediction methods that rely on a large DDI dataset for pre-training, the data processing cost and time consumption are reduced.

[0025] Combined with the multi-perspective-based DDI prediction method described in the above embodiments, the multi-perspective-based DDI prediction system will be described below through embodiments.

[0026] As Figure 2 shown, an embodiment of the present invention provides a multi-perspective-based DDI prediction system, including: A data acquisition module 201, configured to construct four small-sample DDI datasets under hot-start and cold-start settings according to drug reaction types; A feature extraction module 202, configured to extract the SMILES substructure features and molecular graph substructure features of drugs in the four small-sample DDI datasets; A multi-perspective interaction module 203, configured to calculate the interaction features between the substructure features of two drugs from the SMILES-SMILES interaction perspective, the SMILES-molecular graph interaction perspective, and the molecular graph-molecular graph interaction perspective respectively; A multi-perspective fusion DDI prediction template 204, configured to splice the interaction features from the SMILES-SMILES interaction perspective, the SMILES-molecular graph interaction perspective, and the molecular graph-molecular graph interaction perspective with the Morgan molecular fingerprints of the two drugs, and then input them into a multi-layer perceptron for fusion to obtain the DDI prediction result.

[0027] The implementation principle and beneficial effects of the embodiments of the present invention are: The data acquisition module 201 constructs four small-sample DDI datasets under hot start and cold start settings according to the drug reaction type; the feature extraction module 202 extracts the SMILES substructure features and molecular graph substructure features of the drugs in the four small-sample DDI datasets; the multi-perspective interaction module 203 calculates the interaction features between the substructure features of two drugs from the SMILES-SMILES interaction perspective, the SMILES-molecular graph interaction perspective, and the molecular graph-molecular graph interaction perspective respectively; the multi-perspective fusion DDI prediction template 204 splices the interaction features from the three interaction perspectives with the Morgan molecular fingerprints of the two drugs and then inputs them into a multi-layer perceptron for fusion to obtain the DDI prediction result. The SMILES modality and molecular graph modality of the drugs are fused through the SMILES-SMILES interaction perspective, the SMILES-molecular graph interaction perspective, and the molecular graph-molecular graph interaction perspective, so as to optimize the DDI prediction accuracy of the small-sample DDI dataset.

[0028] This embodiment provides a computer-readable storage medium storing a computer program, which when called by a processor is used to execute the steps of the multi-perspective-based DDI prediction method shown above Figure 1 as described.

[0029] It can be understood that the same or similar parts in the above embodiments can be referred to each other, and the content not detailed in some embodiments can be seen in the same or similar content in other embodiments.

[0030] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

[0031] It should be understood that in the embodiments of the present invention, the so-called processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type. The readable storage medium is a computer-readable storage medium, which may be an internal storage unit of the controller described in any of the foregoing embodiments, such as the hard disk or memory of the controller. The readable storage medium may also be an external storage device of the controller, such as a plug-in hard disk equipped on the controller, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the readable storage medium may also include both the internal storage unit of the controller and the external storage device. The readable storage medium is used to store the computer program and other programs and data required by the controller.

[0032] The readable storage medium is a computer-readable storage medium, which may be an internal storage unit of the controller described in any of the foregoing embodiments, such as the hard disk or memory of the controller. The readable storage medium may also be an external storage device of the controller, such as a plug-in hard disk equipped on the controller, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the readable storage medium may also include both the internal storage unit of the controller and the external storage device. The readable storage medium is used to store the computer program and other programs and data required by the controller. The readable storage medium may also be used to temporarily store data that has been output or is to be output.

[0033] Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned readable storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

Claims

1. A DDI prediction method based on multiple perspectives, characterized in that: include: Four small-sample DDI datasets under hot-start and cold-start settings were constructed according to drug reaction types; Extracting the SMILES substructure features and molecular graph substructure features of the drugs in the four small sample DDI datasets; The interaction features between the substructure features of the two drugs were calculated from the SMILES-SMILES interaction perspective, the SMILES-molecular graph interaction perspective, and the molecular graph-molecular graph interaction perspective; The interactive features of the SMILES-SMILES interactive perspective, the SMILES-molecular graph interactive perspective and the molecular graph-molecular graph interactive perspective are spliced ​​with the Morgan molecular fingerprints of the two drugs and input into a multi-layer perceptron for fusion to obtain the DDI prediction result.

2. The DDI prediction method based on multiple perspectives according to claim 1, characterized in that: The four small sample DDI datasets under hot start and cold start settings are constructed according to the drug reaction type, including: Obtain Drugbank dataset, TWOSIDES dataset, unseen DDI dataset and unseen single drug dataset; According to the drug reaction type, the reaction drug pair sets corresponding to the Drugbank dataset, the TWOSIDES dataset, the unseen DDI dataset and the unseen single drug dataset are counted respectively; Selecting N reaction drug pairs from the Drugbank dataset, the unseen DDI dataset, and the unseen single drug dataset as training sets, and selecting N / 10 reaction drug pairs from the TWOSIDES dataset as training sets; Selecting N / 3 reaction drug pairs from the Drugbank dataset, the unseen DDI dataset, and the unseen single drug dataset as a validation set, and selecting N / 30 reaction drug pairs from the TWOSIDES dataset as a validation set; For the hot start setting, both drugs in the Drugbank dataset and the test set of the TWOSIDES dataset exist in the reaction drug pairs in the training sets of the Drugbank dataset and the TWOSIDES dataset; In the test set of the unseen DDI dataset in the cold start setting, both drugs do not exist in the reaction drug pairs in the training set of the unseen DDI dataset; For the cold start setting, one of the two drugs in the test set of the unseen single-drug dataset exists in the reactive drug pair in the training set of the unseen single-drug dataset.

3. The DDI prediction method based on multiple perspectives according to claim 1, characterized in that: The SMILES substructure features and molecular graph substructure features of the drugs in the four small sample DDI data sets are extracted, including: Using the SMILES encoder, the molecular features of the drugs in the four small sample DDI datasets are extracted in the SMILES mode to obtain the SMILES substructure features; The molecular graph encoder is used to extract molecular features of the molecular graph modality of the drugs in the four small sample DDI datasets to obtain the molecular graph substructure features.

4. The multi-view DDI prediction method according to claim 3, characterized in that: The SMILES encoder is used to extract the molecular features of the drugs in the four small sample DDI datasets in the SMILES modality to obtain SMILES substructure features, including: The expression of the molecular features of the SMILES modality extracted by the SMILES encoder using the MoLFormer model is: ; Said Indicates drug i , Indicates the The extracted SMILES substructure features are described Indicates the SMILES features of drugs extracted using the MoLFormer model.

5. The DDI prediction method based on multiple perspectives according to claim 3, characterized in that: The molecular graph encoder is used to extract molecular features of the drugs in the four small sample DDI datasets in the molecular graph modality to obtain molecular graph substructure features, including: The molecular graph encoder uses the GROVER model to extract the molecular features of the molecular graph modality as follows: ; Said Indicates drug i , Indicates the The extracted molecular graph substructure features, It indicates that the molecular features of drugs are extracted using the GROVER model.

6. The multi-view DDI prediction method according to claim 4 or 5, characterized in that: The calculation of the interaction features between the substructure features of the two drugs from the SMILES-SMILES interaction perspective includes: The SMILES substructure feature and the molecular graph substructure feature are respectively mapped to feature spaces of the same dimension through a multi-layer perceptron MLP to obtain a SMILES feature vector and a molecular graph feature vector. , the molecular graph feature vector ; The importance matrix between the SMILES substructure features of two drugs in the SMILES-SMILES interaction perspective is extracted through the attention module. The calculation formula of the importance matrix is: ; Said Indicates drug i and medications j The importance matrix of SMILES substructure features; represents the co-attention module; By combining the two-way interaction module Get the drug i Each SMILES substructure in the molecule and the drug j The calculation formula of the interaction feature matrix of each SMILES substructure in the molecule is: ; Said represents the interaction feature matrix, Embedded feature matrix representing drug response type; According to the The drug is calculated i and the drug j The expression for the aggregate eigenvector of is: ; Said Indicates the drug i Each SMILES substructure is associated with the drug j The one-way eigenvector mapped after all SMILES substructures interact; Indicates the drug j Each SMILES substructure is associated with the drug i The one-way eigenvector mapped after all SMILES substructures interact; Indicates the drug i Each substructure of the drug j The eigenvector after all SMILES substructures interact; Indicates the drug j Each substructure of the drug i The eigenvector after all SMILES substructures interact; The and stated Connect to get the drug from the SMILES-SMILES interaction perspective i With the drug j The expression of the interaction feature is: ; Said Indicates the drug i With the drug j The bidirectional interaction feature vector between them.

7. The multi-view DDI prediction method according to claim 6, characterized in that: The interactive features of the SMILES-SMILES interactive perspective, the SMILES-molecular graph interactive perspective and the molecular graph-molecular graph interactive perspective are spliced ​​with the Morgan molecular fingerprints of the two drugs, and then input into a multi-layer perceptron for fusion to obtain a DDI prediction result, including: The interactive features of the SMILES-SMILES interaction perspective , the interactive features of the SMILES-molecular graph interaction perspective , the interactive features of the molecular graph-molecule graph interaction perspective , the drug i Morgan molecular fingerprint characteristics And the drug j Morgan molecular fingerprint characteristics Perform splicing to obtain a splicing feature vector, wherein the splicing feature vector The expression is: ; The After being input into the MLP, the predicted value is calculated by the Sigmoid activation function. The expression is: ; Said Indicates the drug i and the drug j Predictive value of drug reactions; The Compare with the preset threshold to get the DDI prediction result.

8. A DDI prediction system based on multiple perspectives, characterized in that: include: A data acquisition module is used to construct four small-sample DDI datasets under hot start and cold start settings according to drug reaction types; A feature extraction module, used to extract SMILES substructure features and molecular graph substructure features of the drugs in the four small sample DDI datasets; A multi-perspective interaction module is used to calculate the interaction features between the substructure features of two drugs from the SMILES-SMILES interaction perspective, the SMILES-molecular graph interaction perspective, and the molecular graph-molecular graph interaction perspective; The multi-perspective fusion DDI prediction template is used to splice the interaction features of the SMILES-SMILES interaction perspective, the SMILES-molecular graph interaction perspective and the molecular graph-molecular graph interaction perspective with the Morgan molecular fingerprints of the two drugs, and input them into a multi-layer perceptron for fusion to obtain a DDI prediction result.

9. A computer-readable storage medium, characterized in that: A computer program is stored, and when the computer program is called by a processor, the computer program is used to execute: the steps of the multi-view DDI prediction method according to any one of claims 1 to 7.