Long-tail drug interaction prediction method based on uncertainty perception

By using graph neural networks and multimodal feature modeling, combined with an uncertainty perception mechanism, the problem of neglecting long-tail distribution in existing technologies is solved, improving the recall and F1-score of drug interaction prediction, and achieving better drug safety assessment and clinical decision support.

CN120929910APending Publication Date: 2025-11-11SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511024784.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing machine learning-based drug interaction prediction methods fail to adequately consider the long-tail distribution phenomenon, resulting in insufficient ability to identify tail categories and limiting their applicability in practical applications.

Method used

We employ an uncertainty-aware long-tail drug interaction prediction method. By using graph neural networks and multimodal feature modeling, combined with information entropy and cross-entropy loss functions, we improve the ability to identify tail categories. We utilize Transformer encoding of SMILES, molecular graphs, target and enzyme features to construct a high-level feature representation of the drug interaction matrix, and predict the category probability distribution through a classifier.

Benefits of technology

It significantly improves the performance of drug interaction prediction in long-tailed distribution scenarios, especially achieving breakthroughs in recall rate and F1-score, and verifying its application potential in actual clinical decision support and drug safety assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929910A_ABST
    Figure CN120929910A_ABST
Patent Text Reader

Abstract

The invention discloses a long-tail drug interaction prediction method based on uncertainty perception, and the method comprises the steps: S1, obtaining a drug interaction matrix and the features of each drug in the drug interaction matrix, the drug features comprising a simplified molecular linear input specification SMILES, a molecular map, a target spot and an enzyme; s2, coding the features of the drugs in the S1 to obtain embedded expressions of the drugs; s3, respectively inputting the embedded representations obtained in the S2 into a neural network to obtain high-level feature representations of the neural network; s4, splicing the high-level feature representations, obtained in the S3, of the four features of each medicine, and fusing the high-level feature representations through a single-layer neural network to obtain the final representation of each medicine; and S5, splicing the final expressions of the drug molecule pairs in the drug interaction matrix, inputting the final expressions into the classifier to predict the categories of the final expressions, and outputting the probability distribution of the drug molecule pair interaction categories.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the interdisciplinary research field of artificial intelligence and bioinformatics, specifically involving a long-tailed drug interaction prediction method based on uncertainty perception. Background Technology

[0002] Combination drug therapy, as a promising strategy for treating complex diseases, has significantly improved drug efficacy, but it may also be accompanied by drug side effects. Therefore, the discovery and prediction of disease-specific influencing factors (DDIs) can not only improve medication safety in clinical practice but also promote the further development of combination drug therapy in clinical settings. Traditional experimental methods for DDI prediction are costly and time-consuming. In recent years, computer-aided methods based on machine learning have significantly improved the efficiency of DDI prediction by learning models from existing data. Therefore, research on such methods has important theoretical and applied value. Existing machine learning-based methods have explored the use of various features such as simplified molecular linear input specifications (SMILES), knowledge graphs, and target points for DDI prediction, achieving significant results in multiple tasks. However, these methods often ignore the extreme long-tail phenomenon in the DDI type distribution. DDI types at the tail often have higher potential risks. Many existing methods fail to fully consider this long-tail distribution phenomenon, thus limiting their applicability in practical applications. Therefore, the trained models tend to correctly classify head categories while neglecting the ability to identify tail categories. Summary of the Invention

[0003] The purpose of this application is to provide a long-tail drug interaction prediction method based on uncertainty perception, and the specific technical solution is as follows:

[0004] A long-tail drug interaction prediction method based on uncertainty perception includes: S1, obtaining a drug interaction matrix and the features of each drug in the drug interaction matrix, wherein the drug features include simplified molecular linear input canonical SMILES, molecular graph, target and enzyme; S2, encoding the features of each drug in S1 to obtain its embedding representation; S3, inputting the embedding representation obtained in S2 into a neural network to obtain its high-level feature representation; S4, concatenating the four features of each drug in the high-level feature representation obtained in S3, and fusing them through a single-layer neural network to obtain the final representation of each drug; S5, concatenating the final representations of drug molecule pairs in the drug interaction matrix, inputting them into a classifier to predict their category, and outputting the probability distribution of the interaction category of the drug molecule pair.

[0005] S2 encodes the features of each drug, including: S2.1, encoding the simplified molecular linear input SMILES features using a pre-trained Transformer to obtain its embedding representation s. i S2.2. Encode the molecular graph using a pre-trained graph neural network to obtain its embedding representation g. i S2.3. Encode the target features using a bitwise binary vector and obtain its embedding representation t using Jaccard similarity. i S2.4. Encode the enzyme using a bitwise binary vector and calculate its embedding representation e using Jaccard similarity. i .

[0006] S3 includes inputting the embedded representations into the neural network as follows: S3.1, inputting the simplified molecular linear representation obtained in S2.1 into the normalized SMILES features. i Input a three-layer neural network module, learn its representation, and obtain its high-level features S. i S3.2. Use a three-layer neural network module to learn the molecular graph features g obtained in S2.2. i The representation of is used to obtain its high-level feature G. i The input to each layer is a simplified molecule linear input normalized SMILES feature. i and molecular diagram features g i The concatenated vector; S3.3, the enzyme feature e obtained in S2.4 i Input a three-layer neural network module to obtain its high-level feature representation E i S3.4 Utilize a three-layer neural network module to learn the target feature t obtained in S2.3. i The representation of is used to obtain its high-level feature T. i The input to each layer is the enzyme feature e. i and target features t i The concatenated vector.

[0007] S4 describes the four characteristics of each drug. i G i E i And T i The data are spliced ​​together and then fused using a single-layer neural network to obtain the final representation D of each drug. i In S5, the final representation of drug molecule pairs in the drug interaction matrix is ​​D. i and D j The concatenation is performed and fed into a classifier to predict the drug pair's category; the model outputs the category probability distribution of the drug pair. Where C is the number of categories, P cThis indicates the probability that a drug molecule pair belongs to class c interaction; it also includes:

[0008] S6. Based on the probability distribution output in S5 Its information entropy is calculated using the following expression:

[0009]

[0010] After normalization, we get:

[0011]

[0012] S7. Calculate the classification loss function of the model, expressed as:

[0013]

[0014] Where H is the normalized information entropy, and P t It is the probability value corresponding to the true class in the predicted distribution, and k is the index of the drug molecule pair;

[0015] S8. Calculate the average cross-entropy loss for each class in the current batch, denoted as . The average loss expression for all categories is:

[0016]

[0017] Define the expression for the loss variance regularization term as follows:

[0018]

[0019] S9. Calculate the total loss of the current batch, expressed as:

[0020]

[0021] Where K represents the number of drug molecule pairs in the current batch;

[0022] S10, Calculate the total loss The gradient of the model parameters θ, and the network parameters updated by the optimizer, are expressed as:

[0023]

[0024] The beneficial effects of this application are that by combining graph neural networks, multimodal feature modeling and uncertainty perception mechanisms, it significantly improves the performance of drug interaction prediction in long-tail data scenarios, especially achieving breakthroughs in Recall and F1-score, which pay more attention to tail class recognition capabilities, and verifying its application potential in actual clinical decision support and drug safety assessment. Attached Figure Description

[0025] Figure 1 This is a schematic diagram of the overall framework of this application; Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to specific embodiments and accompanying drawings. It should be understood that these descriptions are merely exemplary and not intended to limit the scope of this application. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.

[0027] A long-tail drug interaction prediction method based on uncertainty perception includes: S1, obtaining a drug interaction matrix and the features of each drug in the drug interaction matrix, wherein the drug features include simplified molecular linear input canonical SMILES, molecular graph, target and enzyme; S2, encoding the features of each drug in S1 to obtain its embedding representation; S3, inputting the embedding representation obtained in S2 into a neural network to obtain its high-level feature representation; S4, concatenating the four features of each drug in the high-level feature representation obtained in S3, and fusing them through a single-layer neural network to obtain the final representation of each drug; S5, concatenating the final representations of drug molecule pairs in the drug interaction matrix, inputting them into a classifier to predict their category, and outputting the probability distribution of the interaction category of the drug molecule pair.

[0028] S2 encodes the features of each drug, including: S2.1, encoding the simplified molecular linear input SMILES features using a pre-trained Transformer to obtain its embedding representation s. i S2.2. Encode the molecular graph using a pre-trained graph neural network to obtain its embedding representation g. i S2.3. Encode the target features using a bitwise binary vector and obtain its embedding representation t using Jaccard similarity. i S2.4. Encode the enzyme using a bitwise binary vector and calculate its embedding representation e using Jaccard similarity. i .

[0029] S3 includes inputting the embedded representations into the neural network as follows: S3.1, inputting the simplified molecular linear representation obtained in S2.1 into the normalized SMILES features. i Input a three-layer neural network module, learn its representation, and obtain its high-level features S. i S3.2. Use a three-layer neural network module to learn the molecular graph features g obtained in S2.2. i The representation of is used to obtain its high-level feature G. iThe input to each layer is a simplified molecule linear input normalized SMILES feature. i and molecular diagram features g i The concatenated vector; S3.3, the enzyme feature e obtained in S2.4 i Input a three-layer neural network module to obtain its high-level feature representation E i S3.4 Utilize a three-layer neural network module to learn the target feature t obtained in S2.3. i The representation of is used to obtain its high-level feature T. i The input to each layer is the enzyme feature e. i and target features t i The concatenated vector.

[0030] S4 describes the four characteristics of each drug. i G i E i And T i The data are spliced ​​together and then fused using a single-layer neural network to obtain the final representation D of each drug. i In S5, the final representation of drug molecule pairs in the drug interaction matrix is ​​D. i and D j The concatenation is performed and fed into a classifier to predict the drug pair's category; the model outputs the category probability distribution of the drug pair. Where C is the number of categories, P c This indicates the probability that a drug molecule pair belongs to class c interaction; it also includes:

[0031] S6. Based on the probability distribution output in S5 Its information entropy is calculated using the following expression:

[0032]

[0033] After normalization, we get:

[0034]

[0035] S7. Calculate the classification loss function of the model, expressed as:

[0036]

[0037] Where H is the normalized information entropy, and P t It is the probability value corresponding to the true class in the predicted distribution, and k is the index of the drug molecule pair;

[0038] S8. Calculate the average cross-entropy loss for each class in the current batch, denoted as . The average loss expression for all categories is:

[0039]

[0040] Define the expression for the loss variance regularization term as follows:

[0041]

[0042] S9. Calculate the total loss of the current batch, expressed as:

[0043]

[0044] Where K represents the number of drug molecule pairs in the current batch;

[0045] S10, Calculate the total loss The gradient of the model parameters θ, and the network parameters updated by the optimizer, are expressed as:

[0046]

[0047] The uncertainty-aware long-tail drug interaction prediction method proposed in this invention can improve the prediction performance of drug combination interactions in long-tail distribution scenarios. Its technical effects are specifically reflected in the following aspects:

[0048] 1. By utilizing the information entropy of the predicted category probability distribution, the prediction performance of drug interaction (DDI) in long-tailed distribution scenarios can be effectively improved.

[0049] As shown in the table below, this invention achieves leading prediction performance on two drug interaction datasets with significant long-tail distribution characteristics, particularly demonstrating significant advantages in the two key metrics of recall and F1-score. Compared with existing mainstream methods, this method improves recall by 1.56% and F1-score by 5.78% on the DDI-DB171 dataset; a similar trend is shown on the DDI-DB100 dataset.

[0050] Recall measures a model's ability to identify a significant number of true positive class samples, and is particularly useful for assessing a model's sensitivity to tail classes. In drug interaction prediction tasks, tail classes often represent rare but clinically potentially high-risk interaction types. In contrast, traditional methods often favor head classes due to sample imbalance, resulting in poor tail class identification. This invention effectively improves the model's attention to tail classes by introducing an uncertainty-aware mechanism and inter-class loss difference regularization, thereby significantly improving the recall metric. Furthermore, the F1-score, as the harmonic mean of precision and recall, is a key indicator for measuring the model's overall performance on imbalanced data.

[0051] It is worth noting that while improving Recall and F1, this invention also maintains competitiveness in other conventional metrics such as Accuracy and Precision, indicating that the method not only has advantages in tail class modeling, but also does not sacrifice head class performance, demonstrating good overall prediction ability and model generalization ability.

[0052] Table 1 Performance comparison on the long-tail dataset

[0053]

[0054] 2. Compared with existing technologies, this invention significantly improves the prediction performance of drug interaction (DDI) in long-tailed distribution scenarios.

[0055] Traditional methods such as Focal Loss or TFL can alleviate class imbalance to some extent, but they are sensitive to hyperparameters or do not adequately model the distribution between classes, making it difficult to achieve a good balance between recall and precision. This invention further enhances the model's ability to learn from "difficult-to-distinguish" samples by constructing a weighted loss related to sample uncertainty, maintaining high precision while improving recall, thus simultaneously improving the F1-score. Furthermore, this invention uses fewer hyperparameters, resulting in strong model scalability.

[0056] While existing technologies (such as FocalLoss and TFL) have alleviated the class imbalance problem to some extent, they have limitations such as sensitivity to hyperparameters and insufficient modeling of class distributions, making it difficult to achieve a good balance between recall and precision.

[0057] This invention further enhances the model's ability to distinguish "difficult to distinguish" samples by introducing a weighted loss function related to sample uncertainty, thereby improving recall while maintaining high precision and achieving a simultaneous improvement in F1-score.

[0058] Furthermore, this invention relies on fewer hyperparameters, which improves the scalability of the model and the flexibility of its practical applications.

[0059] Table 2 Comparison of model performance under different loss conditions

[0060]

[0061] In summary, this invention significantly improves the performance of drug interaction prediction in long-tail data scenarios by combining graph neural networks, multimodal feature modeling, and uncertainty perception mechanisms. In particular, it has made breakthrough progress in Recall and F1-score, which are more focused on tail class recognition capabilities, and verified its application potential in actual clinical decision support and drug safety assessment.

Claims

1. A method for predicting long-tailed drug interactions based on uncertainty perception, characterized in that, include: S1. Obtain the drug interaction matrix and the characteristics of each drug in the drug interaction matrix, wherein the drug characteristics include Simplified Molecular Linear Input Specification (SMILES), molecular diagram, target and enzyme; S2. Encode the features of each drug in S1 to obtain its embedded representation; S3. Input the embedding representations obtained in S2 into the neural network to obtain its high-level feature representations; S4. The high-level feature representations of the four features of each drug obtained in S3 are concatenated and fused through a single-layer neural network to obtain the final representation of each drug. S5. Concatenate the final representations of drug molecule pairs in the drug interaction matrix, input them into the classifier to predict their categories, and output the probability distribution of the interaction categories of the drug molecule pairs.

2. The long-tail drug interaction prediction method based on uncertainty perception as described in claim 1, characterized in that, The encoding of the characteristics of each drug in S2 includes: S2.

1. Encode the simplified molecular linear input SMILES features using a pre-trained Transformer to obtain its embedding representation s. i ; S2.2 Encode the molecular graph using a pre-trained graph neural network to obtain its embedding representation g. i ; S2.

3. Encode the target features using a bitwise binary vector and obtain its embedding representation t using Jaccard similarity. i ; S2.4 Encode the enzyme using a bitwise binary vector and calculate its embedding representation e using Jaccard similarity. i .

3. The long-tail drug interaction prediction method based on uncertainty perception as described in claim 2, characterized in that, The step S3, which involves inputting the embedded representations into the neural network, includes: S3.1, The simplified molecular linear input canonical SMILES features obtained in S2.1 are then used. i Input a three-layer neural network module, learn its representation, and obtain its high-level features S. i ; S3.

2. Use a three-layer neural network module to learn the molecular graph features g obtained in S2.

2. i The representation of is used to obtain its high-level feature G. i The input to each layer is a simplified molecule linear input normalized SMILES feature. i and molecular diagram features g i The concatenated vector; S3.3, The enzyme characteristic e obtained in S2.4 i Input a three-layer neural network module to obtain its high-level feature representation E i ; S3.

4. Use a three-layer neural network module to learn the target feature t obtained in S2.

3. i The representation of is used to obtain its high-level feature T. i The input to each layer is the enzyme feature e. i and target features t i The concatenated vector.

4. The long-tailed drug interaction prediction method based on uncertainty perception as described in claim 3, wherein in S4, the four features S of each drug are respectively... i G i E i And T i The data are spliced ​​together and then fused using a single-layer neural network to obtain the final representation D of each drug. i In step S5, the final representation D of drug molecule pairs in the drug interaction matrix is... i and D j The data is concatenated and then input into a classifier to predict its category. The model outputs the class probability distribution of drug pairs. Where C is the number of categories, P c Indicates the probability that a drug molecule pair belongs to type c interaction; characterized in that it further includes: S6. Based on the probability distribution output in S5. Its information entropy is calculated using the following expression: After normalization, we get: S7. Calculate the classification loss function of the model, expressed as: Where H is the normalized information entropy, and P t It is the probability value corresponding to the true class in the predicted distribution, and k is the index of the drug molecule pair; S8. Calculate the average cross-entropy loss for each class in the current batch, denoted as . The average loss expression for all categories is: Define the expression for the loss variance regularization term as follows: S9. Calculate the total loss of the current batch, expressed as: Where K represents the number of drug molecule pairs in the current batch; S10, Calculate the total loss The gradient of the model parameters θ, and the network parameters updated by the optimizer, are expressed as: