Drug interaction relationship prediction method based on LLMs and ICL
By applying large language models and context learning in the field of drug interaction relationship prediction, a drug relationship prediction model is constructed and weighted fusion is carried out, the problems of high data set organization cost and limited generalization ability of zero-sample scenarios are solved, and efficient drug interaction relationship prediction is achieved in zero-sample and small-sample scenarios.
Patent Information
- Application Number
- CN202510015628.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-01-06
AI Technical Summary
Large language models are expensive to organize data sets in the field of drug interaction relationship prediction, limited generalization ability in zero-sample scenarios, and better guidance is needed in integrating molecular, physiological and clinical multi-source data to improve prediction performance.
A drug interaction relationship prediction method based on large language models (LLMs) and context learning (ICL) is proposed. The final DDI prediction results are obtained by extracting drug data samples, screening drug pairs with the highest similarity, constructing a propt of the drug relationship prediction model, and using expert mixing (MOE) for weighted fusion.
The prediction performance of drug interaction relationships is significantly improved in the zero-sample and small-sample scenarios, making the prediction results more accurate, and helping clinical experts explain the interaction reactions between drugs and build new drugs.
Smart Images

Figure CN120089234A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for predicting drug - drug interaction relationships based on LLMs and ICL, belonging to the technical field of predicting drug - drug interaction relationships. Background Art
[0002] Polypharmacy, that is, the simultaneous use of multiple drugs, is very common in the treatment of patients with multiple diseases. However, due to drug - drug interactions, it may lead to adverse drug reactions (DDIs). Among all reported adverse drug reactions, DDIs account for 30%, which has a significant impact on patient safety, morbidity, mortality, and healthcare costs. Given the complexity of diseases and the limitations of monotherapy, combination therapy has the potential to improve efficacy but also increases the risk of unexpected interactions. Therefore, accurate DDI prediction is crucial for improving treatment outcomes and minimizing adverse reactions. Although DDI research has become a major focus, due to limited clinical trial resources and the rapid growth of biomedical data, the identification of DDIs remains extremely challenging.
[0003] The current state - of - the - art DDI prediction methods include traditional machine learning and deep - learning methods. Among them, deep - learning methods, with the help of technologies such as deep neural networks (DNNs), convolutional neural networks (CNNs), graph neural networks (GNNs), and self - attention mechanisms (transformer), have achieved excellent performance. However, these methods often perform poorly in zero - shot scenarios and have limitations in the ability to learn from large - scale, multi - source data integration.
[0004] Large language models (LLMs), such as GPT - 4, Claude, LLaMA, and Mistral, have achieved remarkable success in various general tasks due to their large - scale parameter configurations, pre - training methods, and advanced neural network architectures. Although LLMs perform well in general tasks, their capabilities in professional application fields are still significantly limited.
[0005] In the field of drug discovery, LLMs have shown significant potential in multiple directions, including multi - source data integration, downstream task design, and optimization of specific application prompting strategies. These advancements enable LLMs to perform tasks such as molecular property prediction and molecular transformation. However, there are still challenges in areas such as property prediction related to DDIs, including the high cost of organizing domain - specific datasets, limited generalization ability in zero - shot scenarios, and the need for better guidance in integrating multi - source data such as molecular, physiological, and clinical data to improve prediction performance. Summary of the Invention
[0006] To address the problems of high dataset collation costs in the field of predicting characteristics related to DDI for large language models, limited generalization ability in zero-shot scenarios, and the need for better guidance in integrating multi-source data such as molecular, physiological, and clinical data to improve prediction performance, the present invention proposes a method for predicting drug-drug interaction relationships based on LLMs and ICL.
[0007] The technical solution adopted by the present invention to solve the above problems is as follows: The present invention includes the following steps:
[0008] Step 1: Extract drug data samples and create a dataset;
[0009] Step 2: Screen the drug data samples in the dataset and select the top K positive and negative sample drug pairs with the highest similarity;
[0010] Step 3: Construct a prompt for the drug relationship prediction model based on in-context learning and the selected positive and negative sample drug pairs, where the prompt is a prompt dialog box;
[0011] Step 4: Use MOE to mix all drug relationship prediction models, evaluate the scores of the prompt content of the drug relationship prediction models, and perform weighted fusion according to the scores to obtain the final DDI prediction result, where MOE is a mixture of experts and DDI is adverse drug reactions.
[0012] Preferably, Step 2 specifically includes:
[0013] Calculate the Tanimoto similarity of the drug samples and select the top K positive and negative sample drug pairs with the highest Tanimoto similarity greater than a preset value;
[0014] The formula for calculating the Tanimoto similarity is:
[0015]
[0016] In formula (1), x is the feature vector of the first drug sample, y is the feature vector of the second drug sample, x·y is the dot product / inner product of the feature vectors x and y, and x·y is calculated by Σx i y i is calculated, Σx i y i is the sum of the products of the corresponding feature vectors, ||x|| 2 is the squared norm of the vector x, ||x|| 2 is calculated by , is the sum of the squared elements of the feature vector x, ||y|| 2 is the squared norm of the vector y, ||y|| 2 is calculated by and ||y||2 is the sum of the squares of the elements of the feature vector y, Sim T The value of (x, y) ranges from 0 to 1.
[0017] Preferably, the prompt of the drug relationship prediction model in step 3 follows a structured format, including an input requirement module, a prediction task module, a consideration factor module, and an example module;
[0018] The input requirement module is used to input the name of the drug and the SMILES structure of the corresponding drug. The SMILES structure is a single-line text representing the structure of the compound;
[0019] The prediction task module is used to judge the interaction between the two input drugs according to the Tanimoto similarity calculation formula. If the Tanimoto similarity is greater than the preset value, there is an interaction between the two input drugs. If the Tanimoto similarity is less than the preset value, there is no interaction between the two input drugs;
[0020] The consideration factor module is used to provide the pharmacodynamics, metabolic pathways, receptor interactions of the corresponding drug, and the analysis results of relevant clinical data;
[0021] The example module is used to combine the analysis results of the input drugs and the actual demonstration of the expected output of simulating zero-shot and few-shot scenarios for the interaction.
[0022] Preferably, each evaluation criterion is scored on a scale of 1 to 5, specifically including:
[0023] Scientific accuracy is used to judge whether the content in the prompt of the current drug relationship prediction model conforms to the current scientific knowledge and whether there are obvious errors or logical problems. If 0-24% of the content in the prompt of the current drug relationship prediction result conforms to the current scientific knowledge and logical problems, the score is 1. If only 25%-49% of the content conforms to the current scientific knowledge and logical problems, the score is 2. If only 50%-74% of the content conforms to the current scientific knowledge and logical problems, the score is 3. If only 75%-99% of the content conforms to the current scientific knowledge and logical problems, the score is 4. If all the content completely conforms to the current scientific knowledge and logical problems, the score is 5;
[0024] Clarity and coherence are used to judge whether the content in the prompt of the current drug relationship prediction model is logically clear and the language expression is coherent and easy to understand. If 0-24% of the content in the prompt of the current drug relationship prediction result is accurately expressed, logically rigorous, well-organized, easy to understand and there are no incoherences, the score is 1. If only 25%-49% of the content is accurately expressed, logically rigorous, well-organized, easy to understand and there are no incoherences, the score is 2. If only 50%-74% of the content conforms to the current scientific knowledge and logical issues, the score is 3. If only 75%-99% of the content is accurately expressed, logically rigorous, well-organized, easy to understand and there are no incoherences, the score is 4. If all the content is accurately expressed, logically rigorous, well-organized, easy to understand and there are no incoherences, the score is 5;
[0025] Evidence support is used to judge whether the content in the prompt of the current drug relationship prediction model cites sufficient evidence and there is reasonable reasoning to support the prediction result. If 0-24% of the content in the prompt of the current drug relationship prediction model completely lacks evidence support and the reasoning is baseless, the score is 4. If only 25%-49% of the content completely lacks evidence support and the reasoning is baseless, the score is 3. If only 50%-74% of the content completely lacks evidence support and the reasoning is baseless, the score is 2. If only 75%-99% of the content completely lacks evidence support and the reasoning is baseless, the score is 1. If all the content has no problem of lacking evidence support and baseless reasoning, the score is 5;
[0026] Relevance is used to judge whether the content in the prompt of the current drug relationship prediction model closely focuses on the prediction task and whether it contains redundant information. If 0-24% of the content in the prompt of the current drug relationship prediction model closely focuses on the prediction task and does not contain redundant information, the score is 1. If only 25%-49% of the content closely focuses on the prediction task and does not contain redundant information, the score is 2. If only 50%-74% of the content closely focuses on the prediction task and does not contain redundant information, the score is 3. If only 75%-99% of the content closely focuses on the prediction task and does not contain redundant information, the score is 4. If all the content closely focuses on the prediction task and does not contain redundant information, the score is 5.
[0027] Preferably, the scoring for evaluating the drug relationship prediction model in step 4 specifically includes:
[0028] Use GPT-4 as the discriminator. Based on the four evaluation criteria of the discriminator, calculate the scores for each evaluation criterion, and add up the scores for each evaluation criterion to obtain the score of the current drug relationship prediction model's prompt content. The four evaluation criteria of the discriminator include scientific accuracy, clarity and coherence, evidence support, and relevance, and each evaluation criterion is scored on a scale of 1 to 5.
[0029] Preferably, in step 4, weighted fusion is performed according to the scores to obtain the final DDI prediction result, which specifically includes:
[0030] Use the weighted fusion method to combine the scores of the prompt content of all drug relationship prediction models. Assign corresponding weights according to the scores of the prompt content of each drug relationship prediction model, multiply the scores of the prompt content of each drug relationship prediction model by the corresponding assigned weights to obtain the weighted scores, and add up the weighted scores of all drug relationship prediction models to obtain the final DDI prediction score S final ;
[0031] The final DDI prediction score S final The calculation formula is:
[0032]
[0033] In formula (2), S final is the final DDI prediction score, N is the total number of drug relationship prediction models, w i is the weight of each drug relationship prediction model, and S model is the total score of each drug relationship prediction model.
[0034] The beneficial effects of the present invention are:
[0035] 1. The present invention realizes the prediction of drug interaction relationships in zero-shot and few-shot scenarios. When compared with existing methods in the few-shot scenario, the prediction performance of the present invention is significantly higher than that of existing methods.
[0036] 2. The present invention specifically scores the drug interaction relationships and performs weighted fusion on the scores, making the prediction of drug interaction relationships more accurate, which helps clinical experts to interpret drug interaction reactions and construct new drugs. Description of the Drawings
[0037] Figure 1 is a flowchart of a method for predicting drug interaction relationships based on LLMs and ICL provided by the present invention;
[0038] Figure 2 is a structural framework diagram of a method for predicting drug interaction relationships based on LLMs and ICL provided by the present invention;
[0039] Figure 3 Schematic diagram of zero-shot scenario DDI prediction prompt provided by the present invention;
[0040] Figure 4 Schematic diagram of few-shot scenario DDI prediction prompt provided by the present invention;
[0041] Figure 5 Schematic diagram of discriminator prompt provided by the present invention. Detailed implementation manners
[0042] Combined with Figures 1-5 to illustrate this embodiment. In this embodiment, LLMs are large language models, and ICL is in-context learning. As Figure 1 and Figure 2 shown, the steps of a method for predicting drug interaction relationships based on LLMs and ICL described in this embodiment include:
[0043] S1: Extract drug data samples and make a data set;
[0044] This embodiment uses the Luo data set (Luo et al. A network integration approach for drug-target interaction prediction and computational drug repositioning from heterogeneous information.). The Luo data set contains the following content, as shown in Table 1, where there are a total of 10,036 pairs of drug-drug interactions.
[0045] Table 1
[0046]
[0047] S2: Screen the drug data samples in the data set and select the top K positive and negative sample drug pairs with the highest similarity;
[0048] In-context learning (ICL) in this embodiment is a prompting paradigm applied to large language models, which enhances the capabilities of large language models by using a small number of demonstration prompts. In DDI prediction, this embodiment needs to study how to find more suitable prompt examples. To better select prompt examples, this embodiment proposes a method for selecting positive and negative samples for ICL in DDI based on drug similarity calculation. Three widely used similarity metrics include Tanimoto similarity, cosine similarity, and Dice similarity. In drug similarity calculation, Dice similarity emphasizes shared structural features, making it suitable for identifying common substructures. Cosine similarity focuses on the angular relationship of feature vectors and is very suitable for analyzing high-dimensional molecular data. Tanimoto similarity balances shared and unique molecular features, making it particularly effective when comparing molecular fingerprints in chemoinformatics. Finally, this embodiment selects Tanimoto similarity, and Tanimoto similarity can be calculated according to the following formula:
[0049]
[0050] In formula (1), x is the feature vector of the first drug sample, y is the feature vector of the second drug sample, x·y is the dot product / inner product of feature vectors x and y, and x·y is calculated by Σx i y i The calculation is obtained, and Σx i y i is the sum of the products of the corresponding feature vectors, ||x|| 2 is the squared norm of vector x, and ||x|| 2 is calculated by The calculation is obtained, is the sum of the squared elements of feature vector x, ||y|| 2 is the squared norm of vector y, and ||y|| 2 is calculated by The calculation is obtained, and ||y|| 2 is the sum of the squared elements of feature vector y. The value range of Sim T (x, y) is from 0 to 1.
[0051] S3: Based on in-context learning and the selected positive and negative sample drug pairs, construct the prompt of the drug relationship prediction model;
[0052] In the latest progress of language models, in-context learning (ICL) has become a method for the model to learn tasks without explicit fine-tuning. ICL achieves this by providing examples in the input, enabling the model to understand the task through the context and generate accurate outputs. Based on the screened positive and negative sample examples, this embodiment constructs prompts for DDI prediction, and these prompts are divided into zero-shot and few-shot scenarios.
[0053] In the zero-shot scenario, the model makes predictions solely based on its pre-trained knowledge without relying on specific examples. This approach is suitable for predicting interactions between novel or previously unseen drug combinations. In contrast, the few-shot scenario provides a small number of examples to assist in guiding the model's predictions, especially when data is limited or there are few relevant examples; the prompt follows a structured format and consists of several key parts: input requirements, prediction task, considerations, and examples; the input requirements specify the drug names and their corresponding SMILES structures. The prediction task is to predict whether there is an interaction between two drugs, with the result being "yes" or "no". Considerations include the analysis of pharmacodynamics, metabolic pathways, receptor interactions, and relevant clinical data (including FDA labels and peer-reviewed literature). Finally, the example section provides a practical demonstration of the input format and expected prediction output to ensure clarity in model application. The prompt for the zero-shot scenario is as Figure 3 shown, and the prompt for the few-shot scenario is as Figure 4 shown.
[0054] S4: Use MOE to mix all drug relationship prediction models, evaluate the scores of the drug relationship prediction models, and perform weighted fusion according to the scores to obtain the final DDI prediction result;
[0055] In this embodiment, GPT-4 is used as the judge to evaluate the DDI predictions generated by multiple drug relationship prediction models. The discriminator evaluates the quality of the explanations provided by each drug relationship prediction model based on four key criteria, namely scientific accuracy, clarity and coherence, evidence support, and relevance.
[0056] This embodiment designs detailed prompts for the zero-shot and few-shot scenarios, as Figure 5 shown. In the zero-shot scenario, the prompt clearly lists the evaluation criteria and instructions for GPT-4 to evaluate each prediction and explanation. In the few-shot scenario, the prompt includes several examples of high-quality evaluations to help the model understand how to score. For each criterion, each explanation is scored on a scale of 1 to 5, and an overall score is given based on the evaluation; after scoring the results of all models, a weighted fusion method is used to combine the prediction results, where the score of each model is multiplied by a predetermined weight reflecting its reliability or performance, and then the weighted scores are added together to generate the final DDI prediction result.
[0057] The specific scoring of each judgment criterion on a scale of 1 to 5 includes:
[0058] Scientific accuracy is used to judge whether the content in the prompt of the current drug relationship prediction model conforms to the current scientific knowledge and whether there are obvious errors or logical problems. If 0-24% of the content in the prompt of the current drug relationship prediction result conforms to the current scientific knowledge and logical problems, the score is 1; if only 25%-49% of the content conforms to the current scientific knowledge and logical problems, the score is 2; if only 50%-74% of the content conforms to the current scientific knowledge and logical problems, the score is 3; if only 75%-99% of the content conforms to the current scientific knowledge and logical problems, the score is 4; if all the content completely conforms to the current scientific knowledge and logical problems, the score is 5;
[0059] Clarity and coherence are used to judge whether the content in the prompt of the current drug relationship prediction model is logically clear and the language expression is coherent and easy to understand. If 0-24% of the content in the prompt of the current drug relationship prediction result is accurately expressed, logically rigorous, well-organized, easy to understand and there are no incoherences, the score is 1; if only 25%-49% of the content is accurately expressed, logically rigorous, well-organized, easy to understand and there are no incoherences, the score is 2; if only 50%-74% of the content conforms to the current scientific knowledge and logical problems, the score is 3; if only 75%-99% of the content is accurately expressed, logically rigorous, well-organized, easy to understand and there are no incoherences, the score is 4; if all the content is accurately expressed, logically rigorous, well-organized, easy to understand and there are no incoherences, the score is 5;
[0060] Evidence support is used to judge whether the content in the prompt of the current drug relationship prediction model cites sufficient evidence and there is reasonable reasoning to support the prediction result. If 0-24% of the content in the prompt of the current drug relationship prediction model completely lacks evidence support and the reasoning is unfounded, the score is 4; if only 25%-49% of the content completely lacks evidence support and the reasoning is unfounded, the score is 3; if only 50%-74% of the content completely lacks evidence support and the reasoning is unfounded, the score is 2; if only 75%-99% of the content completely lacks evidence support and the reasoning is unfounded, the score is 1; if there are no problems of lacking evidence support and unfounded reasoning in all the content, the score is 5;
[0061] Relevance is used to determine whether the content in the prompt of the current drug relationship prediction model closely revolves around the prediction task and whether it contains redundant information. If 0-24% of the content in the prompt of the current drug relationship prediction model closely revolves around the prediction task and does not contain redundant information, the score is 1. If only 25%-49% of the content closely revolves around the prediction task and does not contain redundant information, the score is 2. If only 50%-74% of the content closely revolves around the prediction task and does not contain redundant information, the score is 3. If only 75%-99% of the content closely revolves around the prediction task and does not contain redundant information, the score is 4. If all the content closely revolves around the prediction task and does not contain redundant information, the score is 5.
[0062] After scoring the results of all models, a weighted fusion method is used to combine the scores of all drug relationship prediction models, where the weight of each model is w i The score S given by the discriminator for the output of model i model is determined. The final prediction score S final is calculated according to the following formula:
[0063]
[0064] In formula (2), S final is the final DDI prediction score, N is the total number of drug relationship prediction models, w i is the weight of each drug relationship prediction model, and S model is the total score of each drug relationship prediction model.
[0065] This embodiment uses AUC and AUPR scores as evaluation indicators. AUC is the abbreviation of Area under curve, which is the area under the ROC curve. ROC can reflect the classification ability. Its abscissa is the false positive rate (FPR), and its ordinate is the true positive rate (TPR). The closer the AUC is to 1, the better the model result.
[0066] AUPR is the abbreviation of Area under Precision / Recall curve, which is the area under the PR curve. The abscissa of the PR curve is the recall rate, and the ordinate is the precision rate. The PR curve is easily affected by the sample distribution (the ratio of positive and negative samples in the training samples). Therefore, AUPR can be used to measure the prediction performance for unbalanced data sets. The closer the AUPR value is to 1, the better the model performance. The calculation formulas for these AUC and AUPR indicators are as follows:
[0067] Accuracy is used to measure the proportion of the number of correctly classified samples among all positive samples. The calculation formula is as follows:
[0068]
[0069] The recall rate is used to measure the proportion of correctly classified positive samples in the actual positive samples, and the calculation formula is as follows:
[0070]
[0071] In this embodiment, the proposed method is compared with methods such as GPT-4, GPT-3.5, Davinci-003, LLAMA 2, and LLAMA 3 in the few-shot scenario. The results are shown in Table 2:
[0072] Table 2
[0073] Methods AUC AUPR GPT-4 0.681 0.643 GPT-3.5 0.632 0.622 Davinci-003 0.525 0.553 LLAMA2 0.4 0.488 LLAMA3 0.631 0.658 DDI-JUDGE 0.788 0.801
[0074] As can be seen from Table 2, compared with other LLM methods, DDI-JUDGE has improved in both AUPR and AUC in this embodiment.
[0075] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the equivalent embodiments with equivalent changes within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any simple modification, equivalent replacement, and improvement made to the above embodiments within the spirit and principle of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A method for predicting drug interaction relationships based on LLMs and ICLs, characterized in that: The steps of the drug interaction relationship prediction method based on LLMs and ICL include: Step 1: Extract drug data samples and create a dataset; Step 2: Screen the drug data samples in the dataset and select the top K positive and negative sample drug pairs with the highest similarity; Step 3: Based on context learning and the selected positive and negative sample drugs, a prompt for building a drug relationship prediction model is constructed, where prompt is a prompt dialog box; Step 4: Use MOE to mix all drug relationship prediction models, evaluate the scores of the prompt content of the drug relationship prediction models, perform weighted fusion based on the scores, and obtain the final prediction result of DDI, where MOE is expert mixing and DDI is adverse drug reaction.
2. The method for predicting drug interaction relationships based on LLMs and ICL according to claim 1, characterized in that: Step 2 specifically includes: Calculate the Tanimoto similarity of drug samples, and select the K most similar positive and negative sample drug pairs whose Tanimoto similarity is greater than a preset value; The calculation formula of Tanimoto similarity is: In formula (1), x is the feature vector of the first drug sample, y is the feature vector of the second drug sample, x·y is the dot product / inner product of the feature vectors x and y, and x·y is expressed by ∑x i y i Calculated, ∑x i y i is the sum of the products of the corresponding eigenvectors, ||x|| 2 is the square norm of the vector x and ‖x‖ 2 pass Calculated, is the sum of the squares of the elements of the eigenvector x, ||y|| 2 is the square norm of the vector y and ||y|| 2 pass Calculated, ||y|| 2 is the sum of the squares of the elements of the eigenvector y, Sim T The value of (x,y) ranges from 0 to 1.
3. The method for predicting drug interaction relationships based on LLMs and ICL according to claim 1, characterized in that: The prompt of the drug relationship prediction model in step 3 follows a structured format, including an input requirement module, a prediction task module, a consideration factor module, and an example module; The input requirement module is used to input the name of the drug and the SMILES structure of the corresponding drug, where the SMILES structure is a single line of text expressing the structure of the compound; The prediction task module is used to determine the interaction between the two input drugs according to the Tanimoto similarity calculation formula. If the Tanimoto similarity is greater than a preset value, there is an interaction between the two input drugs. If the Tanimoto similarity is less than the preset value, there is no interaction between the two input drugs. The consideration factor module is used to provide the pharmacodynamics, metabolic pathways, receptor interactions and related clinical data analysis results of the corresponding drug; The example module is used for practical demonstration of expected outputs for zero-shot and few-shot scenarios in conjunction with analytical results of input drugs and interaction simulations.
4. The method for predicting drug interaction relationships based on LLMs and ICL according to claim 1, characterized in that: The scoring of the drug relationship prediction model in step 4 specifically includes: GPT-4 was used as the discriminator. Based on the four criteria of the discriminator, the score of each criterion was calculated, and the score of each criterion was added together to obtain the score of the prompt content of the current drug relationship prediction model. The four criteria of the discriminator included scientific accuracy, clarity and coherence, evidence support, and relevance. Each criterion was scored on a scale of 1 to 5.
5. The method for predicting drug interaction relationships based on LLMs and ICL according to claim 4, characterized in that: Each criterion is scored on a scale of 1 to 5 and includes: Scientific accuracy is used to judge whether the content in the prompt of the current drug relationship prediction model conforms to current scientific knowledge and whether there are obvious errors or logical problems. If 0-24% of the content in the prompt of the current drug relationship prediction result conforms to current scientific knowledge and logical problems, the score is 1; if only 25%-49% of the content conforms to current scientific knowledge and logical problems, the score is 2; if only 50%-74% of the content conforms to current scientific knowledge and logical problems, the score is 3; if only 75%-99% of the content conforms to current scientific knowledge and logical problems, the score is 4; if all the content completely conforms to current scientific knowledge and logical problems, the score is 5; Clarity and coherence are used to judge whether the content in the prompt of the current drug relationship prediction model is logically clear and the language expression is coherent and easy to understand. If 0-24% of the content in the prompt of the current drug relationship prediction result is accurately expressed, logically rigorous, well-organized, easy to understand and without incoherence, the score is 1; if only 25%-49% of the content is accurately expressed, logically rigorous, well-organized, easy to understand and without incoherence, the score is 2; if only 50%-74% of the content is in line with current scientific knowledge and logical problems, the score is 3; if only 75%-99% of the content is accurately expressed, logically rigorous, well-organized, easy to understand and without incoherence, the score is 4; if all the content is accurately expressed, logically rigorous, well-organized, easy to understand and without incoherence, the score is 5; Evidence support is used to judge whether the content in the prompt of the current drug relationship prediction model cites sufficient evidence and whether there is reasonable reasoning to support the prediction results. If 0-24% of the content in the prompt of the current drug relationship prediction model completely lacks evidence support and the reasoning is unfounded, the score is 4; if only 25%-49% of the content completely lacks evidence support and the reasoning is unfounded, the score is 3; if only 50%-74% of the content completely lacks evidence support and the reasoning is unfounded, the score is 2; if only 75%-99% of the content completely lacks evidence support and the reasoning is unfounded, the score is 1; if all the content does not lack evidence support and the reasoning is unfounded, the score is 5; The correlation is used to determine whether the content in the prompt of the current drug relationship prediction model is closely related to the prediction task and whether it contains redundant information. If 0-24% of the content in the prompt of the current drug relationship prediction model is closely related to the prediction task and does not contain redundant information, the score is 1; if only 25%-49% of the content is closely related to the prediction task and does not contain redundant information, the score is 2; if only 50%-74% of the content is closely related to the prediction task and does not contain redundant information, the score is 3; if only 75%-99% of the content is closely related to the prediction task and does not contain redundant information, the score is 4; if all the content is closely related to the prediction task and does not contain redundant information, the score is 5.
6. The method for predicting drug interaction relationships based on LLMs and ICL according to claim 1, characterized in that: In step 4, weighted fusion is performed according to the scores to obtain the final prediction results of DDI, which include: The weighted fusion method is used to merge the scores of the prompt content of all drug relationship prediction models. The corresponding weight is assigned according to the score of the prompt content of each drug relationship prediction model. The score of the prompt content of each drug relationship prediction model is multiplied by the corresponding assigned weight to obtain the weighted score. The weighted scores of all drug relationship prediction models are added together to obtain the final prediction score of DDI S. final ; DDI final prediction score S final The calculation formula is: In formula (2), S final is the final prediction score of DDI, N is the total number of drug relationship prediction models, and w i The weight of each drug relationship prediction model, S model The total score for each drug relationship prediction model.
Citation Information
Patent Citations
Medical information element extraction method, system and device based on hybrid expert model
CN117954110A
Mathematical reasoning method for large language model based on retelling quality optimization
CN118674053A
Prediction of adverse drug reaction based on machine-learned models using protein function scores and clinical factors
US20210327553A1
Method and system for obtaining item-based recommendations
US20230267527A1