Drug effect prediction method for drug research and development
By collecting and preprocessing drug and target data from multiple data sources, and using graph neural networks and multi-task learning frameworks to establish drug-target interaction models, the time-consuming and cost-effective traditional drug efficacy prediction methods are solved, and a more accurate and efficient drug development process is achieved.
Patent Information
- Application Number
- CN202510688694.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-06-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the process of drug development, traditional drug efficacy prediction methods rely on in vitro experiments and animal models, are time-consuming and costly, and are difficult to capture the complex relationship between the drug and multiple targets, affecting the accuracy of prediction.
A drug efficacy prediction method is proposed for drug research and development. By collecting drug and target data from multiple data sources, pre-processing and feature analysis, a drug-target interaction model is established using graph neural network, and the drug efficacy is predicted based on the multi-task learning framework.
This method can quickly screen out drug-target pairs with potential drug effects, reduce R&D costs, shorten R&D cycle, improve the success rate of drug R&D, and provide more accurate drug efficacy prediction results.
Smart Images

Figure CN120199516A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pharmacodynamic analysis, and particularly relates to a method for predicting the pharmacodynamic effect for drug research and development. Background Art
[0002] Drug research and development is a highly complex and costly process. From drug discovery to market launch, it needs to go through multiple stages, including drug screening, preclinical research, clinical trials, etc. Each stage is accompanied by huge challenges. In the initial stage of drug research and development, researchers need to screen out potential candidate drugs and predict their pharmacodynamic effects. Traditional drug screening methods usually rely on in vitro experiments (such as cell experiments) and animal models. Although these experiments can provide certain pharmacodynamic information, they are often time-consuming and costly.
[0003] In the process of predicting the pharmacodynamic effect in drug research and development, drugs produce effects through interactions with multiple targets, and the complexity of the interactions will interfere with the pharmacodynamic prediction task. Therefore, how to capture the complex relationships between drugs and multiple targets and ensure that the overall pharmacodynamic prediction performance is improved under different target and disease backgrounds is the problem we need to solve. For this reason, a method for predicting the pharmacodynamic effect for drug research and development is proposed herein. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for predicting the pharmacodynamic effect for drug research and development to solve the problems raised in the above background art.
[0005] To solve the above technical problems, the technical solution adopted by the present invention is:
[0006] A method for predicting the pharmacodynamic effect for drug research and development, comprising the following steps:
[0007] Step 1: Collect drug and target data from multiple data sources, including drug chemical structures, target protein information, interaction data between known drugs and targets, drug efficacy, and disease-related data, and preprocess the drug and target data;
[0008] Step 2: Based on the preprocessed drug and target data, perform feature analysis to obtain a comprehensive feature sequence, analyze the genes and pathways of the diseases to which the drugs are applied, and perform disease feature encoding, and then model the disease background;
[0009] Step 3: Use a graph neural network to analyze the interactions between drugs and targets, generate feature representations of drug-target pairs, establish a drug-target interaction model, and analyze multi-target interactions and their relationships with drugs;
[0010] Step 4: Based on the multi-task learning framework, jointly analyze the relationships between drugs and multiple targets and disease backgrounds, and then optimize the drug-target interaction model;
[0011] Step 5: According to the optimized drug-target interaction model, conduct experiments with small batches and multiple control groups, input new drug and target pairs, predict the efficacy of the drugs, and verify the performance of the prediction results;
[0012] Step 6: Preset a discriminant threshold for performance, analyze whether the performance of the prediction results meets the standards. If it meets the standards, apply it to the research and development of new drugs, and guide drug screening and optimization according to the prediction results. If it does not meet the standards, return to the modeling process of the disease background for further improvement and optimization.
[0013] A further improvement of the technical solution of the present invention lies in that: the specific steps of Step 1 include:
[0014] Clarify the requirements for efficacy prediction in drug research and development, and then collect corresponding drug and target data from different data sources. Among them, the drug and target data include drug chemical structures, target protein information, known drug-target interaction data, drug efficacy, and disease-related data. The data sources include public databases, literature resources, and experimental data;
[0015] For drug chemical structure data, collect SMILES strings, molecular graph structures, and molecular fingerprint data;
[0016] For target protein information data, collect amino acid sequences, protein three-dimensional structures, and target annotation information data;
[0017] For known drug-target interaction data, collect binding affinity data and interaction types;
[0018] For drug efficacy data, collect efficacy indicators and disease indication data. Among them, the efficacy indicators include cure rate, remission rate, and side effect incidence rate, etc., and the disease indication indicates the specific disease types for which the drug is used for treatment;
[0019] For disease-related data, collect gene expression data, signal pathway information, and disease phenotype information. Among them, the gene expression data is used to analyze the gene expression patterns related to the disease, the signal pathway information is used to analyze the biological pathways related to the disease, and the disease phenotype information includes disease symptoms and pathogenesis, etc.;
[0020] Preprocess the collected drug and target data, including data cleaning and standardization. Specifically, check and delete duplicate compound or target information, remove records with incomplete binding affinity data or inaccurate drug structure information, convert data from different sources into a unified format, associate the drug-target interaction data with the drug chemical structure and target protein information, and integrate the disease-related gene and pathway information with the drug and target data. Then, integrate the preprocessed data into a unified dataset.
[0021] A further improvement of the technical solution of the present invention lies in: The specific steps of step two include:
[0022] Represent the drug molecular structure using SMILES strings and convert it into a molecular graph structure through a tool (RDKit). Then, extract the molecular fingerprint for quickly comparing molecular similarity, and extract the quantum chemical descriptors of the molecule (molecular weight, polar surface area, number of hydrogen bond donors / acceptors, etc.). Combine the chemical structure and physicochemical properties of the drug to generate a drug feature vector.
[0023] Analyze the amino acid sequence of the target protein, extract information such as sequence length, amino acid composition, and conserved region, and perform molecular docking analysis using the three-dimensional structure information of the protein (PDB data). Then, integrate the annotation information of the biological function of the target and the signal pathways involved to generate a functional feature vector of the target.
[0024] Use the known binding affinity data (IC50, Ki value) to evaluate the interaction strength between the drug and the target. Combine the characteristics of the drug and the target, and extract the interaction features of the drug-target pair through a graph neural network. Then, integrate the drug feature vector, the functional feature vector of the target, and the interaction features of the drug-target pair to form a comprehensive feature sequence for subsequent analysis and modeling.
[0025] Use the preprocessed disease-related gene expression data to analyze the gene expression pattern of the disease to which the drug is applied. Through differential expression analysis, identify the differentially expressed genes related to the disease, identify disease-related biomarkers through the gene expression pattern, and use pathway databases (KEGG, Reactome, etc.) to obtain disease-related signal pathway information, identify the genes and proteins in the key pathways, and construct a disease-related pathway network.
[0026] According to the disease phenotype information, formulate a disease feature coding rule, convert the expression pattern of the differentially expressed genes into a feature vector, convert the pathway information into a network graph structure, extract the key nodes and connection relationships in the pathway, and then integrate the phenotype information of the disease symptoms and pathogenesis to code the disease features and form a disease feature coding sequence.
[0027] Integrate the comprehensive feature sequence, gene and pathway analysis results, and disease feature coding sequence to form a comprehensive feature matrix containing drugs, targets, diseases, and their related biological information, and construct a disease background model by combining the BTDHDTA model and the comprehensive feature matrix.
[0028] A further improvement of the technical solution of the present invention lies in: the construction process of the disease background model:
[0029] Align the comprehensive feature sequence, gene and pathway analysis results, and disease feature coding sequence in the same dimension, and integrate them into a matrix in a row and column manner to form a comprehensive feature matrix. Among them, each row of the comprehensive feature matrix represents a sample, and each column represents a feature;
[0030] Based on the data characteristics of the comprehensive feature matrix, select the BTDHDTA model architecture, and divide the comprehensive feature matrix into a training set, a validation set, and a test set. The training set is used to train the model, the validation set is used to adjust the model parameters, and the test set is used to evaluate the model performance;
[0031] Perform standardization processing on the numerical features in the comprehensive feature matrix to make them have zero mean and unit variance, so as to improve the convergence speed and performance of the model, and construct a neural network model including a data processing module, a feature extraction module, a feature fusion module, and a prediction module according to the structure of the BTDHDTA model;
[0032] Based on the determined BTDHDTA model structure, set the model parameters including hyperparameters such as learning rate, batch size, and number of iterations, and use the training set data to train the model. Update the model parameters through the backpropagation algorithm to minimize the loss function of the model on the training set. Use the validation set data to evaluate the trained model, analyze the error between the predicted value and the true value, and then, according to the performance on the validation set, iterate multiple times to adjust the model structure and parameters to improve the generalization ability of the model;
[0033] For the model after multiple iterations, use the test set data to test the tuned model, evaluate the performance of the model in actual applications, and then obtain the trained disease background model. Among them, the input of the disease background model is the comprehensive feature matrix, and the output is the disease background representation, that is, the comprehensive representation of the disease background.
[0034] A further improvement of the technical solution of the present invention lies in: the specific steps of step three include:
[0035] Collect the structural information of drugs and the sequence information of targets, and obtain the known drug-target interaction data. Then, convert the structural information of drugs into graph data, where atoms are used as nodes and chemical bonds are used as edges, and encode the sequence information of targets;
[0036] Collect drug data involving multiple targets, including the interaction information of each drug with multiple targets, extract features for each target separately to generate the feature representation of the target, and then combine the feature representations of multiple targets to form a multi-target feature vector;
[0037] Use the GNN variant of the graph convolutional network to embed the graph data of drugs and targets to generate a low-dimensional feature representation. Through a multi-layer GNN structure, further extract the deep features of drugs and targets, and then fuse the feature representations of drugs and targets in the feature fusion layer to generate a comprehensive feature representation of the drug-target pair;
[0038] After the feature fusion layer, add a fully connected layer to predict the interaction between the drug and the target, output a binary classification result, analyze whether the drug and the target interact, select the cross-entropy loss function based on the binary classification result, and use an optimization algorithm (such as Adam, SGD, etc.) to optimize the model parameters to minimize the loss function, and then obtain a drug-target interaction model.
[0039] A further improvement of the technical solution of the present invention lies in: the specific steps of step four include:
[0040] Based on the analysis results of drug, target data, and disease background, integrate the feature data of drug features, target features, and disease background representation to obtain an optimized feature set, which provides input for multi-task learning;
[0041] Design a shared feature extraction layer in combination with the multi-task learning framework to extract the common features of drugs, targets, and disease background, and design task-specific modules including a drug-target interaction prediction module and a disease background association module. At the same time, design a fusion layer to fuse the shared features and task-specific features to generate a comprehensive feature representation. Among them, the drug-target interaction prediction module is used to predict the interaction strength between the drug and the target, and the disease background association module is used to analyze the relationship between the drug and the disease background, and then define the objective function of multi-task learning, and combine the loss functions of drug-target interaction prediction and disease background association;
[0042] Combine the optimized feature set to optimize the drug-target interaction model, and at the same time optimize the drug-target interaction prediction and disease background association tasks. Update the model parameters through backpropagation, minimize the multi-task objective function, and adjust the hyperparameters to balance the importance of different tasks. Then calculate evaluation metrics to evaluate the performance of the model to ensure that the model has good generalization ability. Among them, the evaluation metrics include the accuracy, recall rate, ROC-AUC value of drug-target interaction prediction, and the mean square error of disease background association, etc., and then obtain an optimized drug-target interaction model;
[0043] Using the optimized drug-target interaction model, predict the interactions between a drug and multiple targets, analyze the multi-target interactions and their relationships with the drug, determine whether there are synergistic or antagonistic effects, and based on the prediction results, analyze the interaction strengths between the drug and different targets.
[0044] A further improvement of the technical solution of the present invention lies in that: the process of analyzing the multi-target interactions and their relationships with the drug is as follows:
[0045] Prepare multi-target data and load the pre-trained drug-target interaction model from the storage medium;
[0046] Input the drug feature vector and the multi-target feature vector into the drug-target interaction model. The model outputs the prediction results, which are the probability values of the interactions of each drug-target pair, representing the possibility of the interaction between the drug and the target. Summarize the prediction results of all drug-target pairs to form a complete multi-target interaction matrix, where the rows of the matrix represent drugs, the columns represent targets, and each element represents the possibility of the interaction between the drug and the target;
[0047] Based on the interaction states between the drug and multiple targets, classify the interaction patterns between the drug and multiple targets. Among them, the interaction patterns include single-target interaction, multi-target synergistic interaction, multi-target antagonistic interaction, and complex interaction patterns. Single-target interaction means that the drug only interacts with one target. Multi-target synergistic interaction means that the drug interacts with multiple targets, and the interaction may produce a synergistic effect, enhancing the therapeutic effect. Multi-target antagonistic interaction means that the drug interacts with multiple targets, but the interaction may produce an antagonistic effect, weakening the therapeutic effect. Complex interaction patterns mean that the interaction relationships between the drug and multiple targets are complex and difficult to simply classify;
[0048] Based on the classification of the interaction patterns, define the synergistic and antagonistic effects, and combined with the prediction results, calculate the total effect of the interaction between the drug and multiple targets, compare the total effect with the sum of the effects of each individual action, and determine whether there are synergistic or antagonistic effects. Among them, for the analysis of the synergistic effect, identify the synergistic interaction pattern between the drug and multiple targets, and judge whether there is a synergistic effect by calculating the weighted sum of the interaction strengths of each drug-target pair. The synergistic effect is manifested as the combined action effect of the drug on multiple targets being greater than the sum of the individual actions. For the analysis of the antagonistic effect, identify the antagonistic interaction pattern between the drug and multiple targets, and judge whether there is an antagonistic effect by analyzing the difference in the interaction strengths of the drug on different targets. The antagonistic effect is manifested as the combined action effect of the drug on multiple targets being less than the sum of the individual actions;
[0049] Extract the interaction probability values between drugs and targets from the prediction results, use them as intensity indicators, compare the interaction intensities between drugs and different targets, rank the targets, and identify the main action targets of the drugs.
[0050] A further improvement of the technical solution of the present invention lies in: the specific steps of step five include:
[0051] According to the optimized model, select a new set of drug and target pairs for experimental verification, determine the targets that interact with the drugs, and the drug and target pairs cover different disease backgrounds and action mechanisms;
[0052] Set up multiple groups of control experiments, including: positive control group, negative control group and blank control group. Among them, the positive control group selects drugs known to have strong interactions with the targets as positive controls to ensure the effectiveness of the experiment. The negative control group selects drug and target pairs known to be ineffective as negative controls. The blank control group is set up as the group that does not receive any drug treatment, which is used to evaluate the biological effects at the basic level, and divide the experiment into small batches, each batch contains several drug and target pairs, and each experimental group is repeated multiple times;
[0053] According to the designed experimental plan, conduct the interaction experiment between drugs and targets. Through in vitro or in vivo experiments, observe the effects of drugs on the targets, record the results of the interaction intensities between drugs and targets, organize the experimental data in tabular form, including drug names, target names, experimental group types and experimental results, and then summarize the data of repeated experiments and calculate the average value;
[0054] Input the characteristic data of the new drugs and targets into the optimized drug-target interaction model. The model predicts the interaction intensity between drugs and targets according to the input data;
[0055] Comprehensively compare and analyze the predicted drug-target interaction intensities of the model with the experimental results, calculate the pharmacodynamic performance index, compare the differences between the experimental groups and the predicted results, and then analyze the expressiveness of the predicted results.
[0056] A further improvement of the technical solution of the present invention lies in: the process of comparing the differences between the experimental groups and the predicted results is:
[0057] Organize the experimental data of all experimental groups, including drug names, target names, experimental group types (positive control, negative control, experimental group, etc.), the number of experimental repetitions, and the results of the drug-target interaction intensities of each experiment, and obtain the predicted values of the interaction intensities of each drug-target pair output by the optimized drug-target interaction model. Among them, the input of the drug-target interaction model is the drug feature vector and the target feature vector;
[0058] Analyze the experimental data, obtain the experimental values of the interaction strength of each drug-target pair, and determine the benchmark value based on the average value of the positive control group;
[0059] According to the drug-target interaction strength predicted by the model, the drug-target interaction strength measured experimentally, and the determined benchmark value, calculate the pharmacodynamic performance index, analyze the differences between the experimental group and the prediction results, and further verify the interaction between the drug and the target. Among them, for the positive control group, the value of the pharmacodynamic performance index should be close to 1, indicating a high degree of consistency between the model prediction and the experimental results. For the negative control group, the value of the pharmacodynamic performance index should be close to 0, indicating that the model prediction is inconsistent with the experimental results.
[0060] A further improvement of the technical solution of the present invention lies in: the specific steps of step six include:
[0061] According to the drug R & D requirements, set the discrimination threshold of the pharmacodynamic performance index, and divide the prediction result expressiveness into different expressiveness levels, namely high expressiveness level, medium expressiveness level, and low expressiveness level;
[0062] For each drug-target pair, analyze the pharmacodynamic performance index by combining the model prediction value and the experimental value, compare the calculated pharmacodynamic performance index with the preset discrimination threshold, and determine the expressiveness level of each drug-target pair;
[0063] Count the number of drug-target pairs at the high expressiveness level, medium expressiveness level, and low expressiveness level, analyze the overall prediction performance of the model, and conduct a specific analysis of the drug-target pairs with low expressiveness to find out the reasons for inaccurate prediction;
[0064] If the overall expressiveness of the model prediction results meets the high standards (that is, most drug-target pairs meet the high expressiveness level), it is considered that the model is accurate and can be used to guide the R & D of new drugs. According to the prediction results, screen out drug candidate molecules with strong interaction with specific targets, and further optimize the screened drug candidate molecules, such as improving pharmacokinetic properties, reducing toxicity, etc., and then enter the clinical trial stage to verify the effectiveness and safety of the drugs;
[0065] If the overall expressiveness of the model prediction results does not meet the high standards (that is, there are a large number of drug-target pairs at the medium expressiveness level or low expressiveness level), it is considered that the model needs to be further improved and optimized. Return to the modeling process of the disease background, re-examine the disease mechanism, target selection, construction of drug feature vectors and target feature vectors, and then collect more experimental data for training and optimizing the model, conduct iterative optimization of the model to improve the prediction accuracy, and repeat the analysis, decision-making and application process until the model prediction results meet the high standards.
[0066] Due to the adoption of the above technical solution, the technical progress achieved by the present invention compared with the prior art is as follows:
[0067] 1. The present invention provides a method for predicting drug efficacy in drug research and development. Through drug efficacy prediction, it is possible to quickly screen out potential drug-target pairs at the early stage of drug research and development, avoiding further research on a large number of ineffective compounds, thus saving time and resources. Secondly, by predicting the interaction strength between drugs and targets through a model and conducting comprehensive analysis in combination with the disease background, potential drug candidate molecules can be more accurately identified, further shortening the research and development cycle.
[0068] 2. The present invention provides a method for predicting drug efficacy in drug research and development. By applying the drug efficacy prediction method, the research and development cost can be effectively reduced. On the one hand, by screening out drug candidate molecules with high drug efficacy performance indices through model prediction, the need for experimental verification of a large number of compounds is reduced. On the other hand, drugs that may fail in the subsequent research and development stage can be identified in advance, avoiding further investment in these drugs, and further improving the success rate of drug research and development. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0070] Figure 1 It is a schematic diagram of the working process of the present invention;
[0071] Figure 2 It is a schematic diagram of the method process of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0072] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0073] Example 1, as Figure 1 、 Figure 2 shown, the present invention provides a method for predicting drug efficacy in drug research and development, including the following steps:
[0074] Step 1: Collect drug and target data from multiple data sources, including drug chemical structures, target protein information, known drug-target interaction data, drug efficacy, and disease-related data, and preprocess the drug and target data to clarify the requirements for efficacy prediction in drug research and development, and then collect corresponding drug and target data from different data sources. Among them, the drug and target data include drug chemical structures, target protein information, known drug-target interaction data, drug efficacy, and disease-related data, and the data sources include public databases, literature resources, and experimental data. For drug chemical structure data, collect SMILES strings, molecular graph structures, and molecular fingerprint data. Among them, the SMILES string is a text format for representing molecular structures, which is convenient for subsequent processing. The molecular graph structure represents the molecule as a graph structure of nodes (atoms) and edges (chemical bonds). The molecular fingerprint includes MACCS fingerprints, ECFP fingerprints, etc., which are used to quickly compare molecular similarities. Obtain drug chemical structure data from public databases such as PubChem, ChEMBL, and ZINC. For target protein information data, collect amino acid sequences, protein three-dimensional structures, and target annotation information data. Among them, the amino acid sequence is used for subsequent sequence analysis and feature extraction, and the protein three-dimensional structure is used for molecular docking and structural analysis. The target annotation information includes the biological functions of the target, the signal pathways involved, etc. Obtain target protein information data from the Protein Data Bank (PDB), the sequence database (UniProt), and the target annotation database (DrugBank, Therapeutic Target Database (TTD)). For known drug-target interaction data, collect binding affinity data and interaction types. Among them, the binding affinity data is used to evaluate the binding strength between the drug and the target, and the interaction types include agonists, antagonists, and inhibitors, etc. Obtain the binding affinity data between the drug and the target from public databases including BindingDB, DTC (Drug-TargeT Commons), and ChEMBL, and extract the drug-target interaction information through literature reviews and data mining tools (Europe PMC). For drug efficacy data, collect efficacy indicators and disease indication data. Among them, the efficacy indicators include cure rate, remission rate, and side effect incidence rate, etc. The disease indication represents the specific disease type for which the drug is used. Obtain from the clinical trial database (ClinicalTrials.From the government (gov), obtain the clinical trial results of the provided drugs, obtain the efficacy information of the drugs from the drug labels of pharmaceutical manufacturers or drug regulatory agencies, obtain the research literature on drug efficacy through platforms such as PubMed and Google Scholar. For disease-related data, collect gene expression data, signaling pathway information, and disease phenotype information. Among them, the gene expression data is used to analyze the gene expression patterns related to the disease, the signaling pathway information is used to analyze the biological pathways related to the disease, and the disease phenotype information includes disease symptoms and pathogenesis, etc. Obtain the disease-related gene expression data from the Gene Expression Omnibus, obtain the disease-related signaling pathway information from pathway databases (KEGG, Reactome, etc.), and obtain the disease-related genetic information from disease databases (Online Mendelian Inheritance in Man). Preprocess the collected drug and target data, including data cleaning and standardization. Among them, check and delete duplicate compound or target information, remove records with incomplete binding affinity data or inaccurate drug structure information, convert data from different sources into a unified format, associate the interaction data of drugs and targets with the drug chemical structure and target protein information, and integrate the disease-related gene and pathway information with the drug and target data. Furthermore, integrate the preprocessed data into a unified dataset;.
[0075] Step 2: Based on the preprocessed drug and target data, conduct feature analysis to obtain a comprehensive feature sequence. Analyze the genes and pathways of the diseases to which the drugs are applied, and perform disease feature encoding. Then, model the disease background. Represent the drug molecular structure using SMILES strings and convert it into a molecular graph structure through a tool (RDKit). Furthermore, extract molecular fingerprints for quickly comparing molecular similarities, and extract the quantum chemical descriptors of the molecule (molecular weight, polar surface area, number of hydrogen bond donors / acceptors, etc.). Combine the chemical structure and physicochemical properties of the drug to generate a drug feature vector. Analyze the amino acid sequence of the target protein, extract information such as sequence length, amino acid composition, and conserved regions, and use the three-dimensional structure information of the protein (PDB data) for molecular docking analysis. Then, integrate the biological functions of the target and the annotation information of the signaling pathways involved to generate a functional feature vector of the target. Use the known binding affinity data (IC50, Ki value) to evaluate the interaction strength between the drug and the target, and combine the features of the drug and the target to extract the interaction features of the drug-target pair through a graph neural network. Furthermore, integrate the drug feature vector, the functional feature vector of the target, and the interaction features of the drug-target pair to form a comprehensive feature sequence for subsequent analysis and modeling. Use the preprocessed disease-related gene expression data to analyze the gene expression patterns of the diseases to which the drugs are applied. Through differential expression analysis, identify the differentially expressed genes related to the diseases. Identify disease-related biomarkers through gene expression patterns, and use pathway databases (KEGG, Reactome, etc.) to obtain disease-related signaling pathway information. Identify the genes and proteins in the key pathways and construct a disease-related pathway network. According to the disease phenotype information, formulate disease feature encoding rules, convert the expression patterns of the differentially expressed genes into feature vectors, convert the pathway information into a network graph structure, extract the key nodes and connection relationships in the pathway, and then integrate the phenotypic information of the disease symptoms and pathogenesis to encode the disease features and form a disease feature encoding sequence. Integrate the comprehensive feature sequence, the results of gene and pathway analysis, and the disease feature encoding sequence to form a comprehensive feature matrix containing drugs, targets, diseases, and their related biological information. Combine the BTDHDTA model and the comprehensive feature matrix to construct a disease background model. Through the model, learn the complex relationships among drugs, targets, and diseases to provide background support for drug efficacy prediction;
[0076] In addition, the construction process of the disease background model:
[0077] Align the comprehensive feature sequence, gene, and pathway analysis results, as well as the disease feature coding sequence, in the same dimension, and integrate them into a matrix in a row and column manner to form a comprehensive feature matrix. Here, each row of the comprehensive feature matrix represents a sample, and each column represents a feature. Based on the data characteristics of the comprehensive feature matrix, select the BTDHDTA model architecture, and divide the comprehensive feature matrix into a training set, a validation set, and a test set. The training set is used to train the model, the validation set is used to adjust the model parameters, and the test set is used to evaluate the model performance. Standardize the numerical features in the comprehensive feature matrix to make them have zero mean and unit variance to improve the convergence speed and performance of the model. And according to the structure of the BTDHDTA model, construct a neural network model including a data processing module, a feature extraction module, a feature fusion module, and a prediction module. Based on the determined BTDHDTA model structure, set the model parameters including hyperparameters such as learning rate, batch size, and number of iterations, and use the training set data to train the model. Update the model parameters through the backpropagation algorithm to minimize the loss function of the model on the training set. Use the validation set data to evaluate the trained model, analyze the error between the predicted value and the true value, and then, according to the performance on the validation set, iterate multiple times to adjust the model structure and parameters to improve the generalization ability of the model. For the model after multiple iterations, use the test set data to test the tuned model and evaluate the performance of the model in practical applications, and then obtain the trained disease background model. Here, the input of the disease background model is the comprehensive feature matrix, and the output is the disease background representation, that is, the comprehensive representation of the disease background;
[0078] Step 3: Use a graph neural network to analyze the interaction between drugs and targets, generate a feature representation of the drug-target pair, and establish a drug-target interaction model to analyze multi-target interactions and their relationship with drugs. Collect the structural information of drugs and the sequence information of targets, and obtain known drug-target interaction data. Then, convert the structural information of drugs into graph data, where atoms are used as nodes and chemical bonds are used as edges. Encode the sequence information of targets, collect drug data involving multiple targets, including the interaction information between each drug and multiple targets. Extract features for each target separately to generate a feature representation of the target. Then, combine the feature representations of multiple targets to form a multi-target feature vector. Use a GNN variant of the graph convolutional network to embed the graph data of drugs and targets to generate a low-dimensional feature representation. Through a multi-layer GNN structure, further extract the deep features of drugs and targets. Then, fuse the feature representations of drugs and targets in the feature fusion layer to generate a comprehensive feature representation of the drug-target pair. After the feature fusion layer, add a fully connected layer to predict the interaction between drugs and targets, and output a binary classification result to analyze whether drugs and targets interact. Based on the binary classification result, select the cross-entropy loss function and use an optimization algorithm (such as Adam, SGD, etc.) to optimize the model parameters to minimize the loss function, and then obtain a drug-target interaction model;
[0079] Step 4: Based on the multi-task learning framework, jointly analyze the relationships between drugs and multiple targets and disease backgrounds, and then optimize the drug-target interaction model. Based on the analysis results of drug, target data, and disease backgrounds, integrate the feature data representing drug features, target features, and disease backgrounds to obtain an optimized feature set, which serves as the input for multi-task learning. Design a shared feature extraction layer in combination with the multi-task learning framework to extract the common features of drugs, targets, and disease backgrounds, and design task-specific modules including a drug-target interaction prediction module and a disease background association module. At the same time, design a fusion layer to fuse the shared features with the task-specific features to generate a comprehensive feature representation. Among them, the drug-target interaction prediction module is used to predict the interaction strength between drugs and targets, and the disease background association module is used to analyze the relationship between drugs and disease backgrounds, and then define the objective function of multi-task learning. Combine the loss functions of drug-target interaction prediction and disease background association, and combine the optimized feature set to optimize the drug-target interaction model. At the same time, optimize the drug-target interaction prediction and disease background association tasks, update the model parameters through backpropagation, minimize the multi-task objective function, and adjust the hyperparameters to balance the importance of different tasks. Then calculate the evaluation metrics to evaluate the performance of the model and ensure that the model has good generalization ability. Among them, the evaluation metrics include the accuracy, recall rate, ROC-AUC value of drug-target interaction prediction, and the mean square error of disease background association, etc. Then obtain the optimized drug-target interaction model, use the optimized drug-target interaction model to predict the interactions between drugs and multiple targets, analyze the multi-target interactions and their relationships with drugs, clarify whether there are synergistic or antagonistic effects, and based on the prediction results, analyze the interaction strength between drugs and different targets;
[0080] In addition, the process of analyzing the multi-target interactions and their relationships with drugs is as follows:
[0081] Prepare multi-target data, and load the pre-trained drug-target interaction model from the storage medium. Input the drug feature vector and the multi-target feature vector into the drug-target interaction model. The model outputs the prediction results, which are the probability values of the interaction for each drug-target pair, representing the possibility of the drug interacting with the target. Aggregate the prediction results of all drug-target pairs to form a complete multi-target interaction matrix. Here, the rows of the matrix represent drugs, the columns represent targets, and each element represents the possibility of the interaction between the drug and the target. Classify the interaction patterns between the drug and multiple targets based on the interaction status between the drug and multiple targets. Among them, the interaction patterns include single-target interaction, multi-target synergistic interaction, multi-target antagonistic interaction, and complex interaction patterns. Single-target interaction means that the drug only interacts with one target. Multi-target synergistic interaction means that the drug interacts with multiple targets, and the interaction may produce a synergistic effect, enhancing the therapeutic effect. Multi-target antagonistic interaction means that the drug interacts with multiple targets, but the interaction may produce an antagonistic effect, weakening the therapeutic effect. Complex interaction patterns mean that the interaction relationship between the drug and multiple targets is complex and difficult to simply classify. Based on the classification of interaction patterns, define the synergistic effect and the antagonistic effect, and combine the prediction results to calculate the total effect of the drug interacting with multiple targets. Compare the total effect with the sum of the effects of each individual action to determine whether there is a synergistic effect or an antagonistic effect. Among them, for the analysis of the synergistic effect, identify the synergistic interaction pattern between the drug and multiple targets, and judge whether there is a synergistic effect by calculating the weighted sum of the interaction intensities of each drug-target pair. The synergistic effect is manifested as the combined effect of the drug on multiple targets being greater than the sum of the individual effects. For the analysis of the antagonistic effect, identify the antagonistic interaction pattern between the drug and multiple targets, and judge whether there is an antagonistic effect by analyzing the difference in the interaction intensities of the drug on different targets. The antagonistic effect is manifested as the combined effect of the drug on multiple targets being less than the sum of the individual effects. Extract the interaction probability value between the drug and the target from the prediction results, use it as an intensity index, and compare the interaction intensities between the drug and different targets to rank the targets and identify the main action targets of the drug;
[0082] The expression for the interaction probability value between the drug and the target is:
[0083] ;
[0084] In the formula, is the interaction probability value between the drug and the target output by the model, is the number of features, is the weight of the th feature, representing the importance of the feature, The extraction function of a feature, with the input being the optimized feature set , is the standard deviation of the th feature, used to normalize the feature, is the bias term, The value range of is [0, 1], representing the possibility of the interaction between the drug and the target. When tends to 1, it indicates a higher possibility of the interaction between the drug and the target. When is smaller, tends to 0, indicating a lower possibility of the interaction between the drug and the target;
[0085] Step Five: According to the optimized drug-target interaction model, conduct experiments with small batches and multiple control groups, input new drug and target pairs, predict the efficacy of the drug, and verify the performance of the prediction results;
[0086] Step Six: Preset a discriminant threshold for performance, analyze whether the performance of the prediction results meets the standard. If it meets the standard, apply it to the research and development of new drugs, and guide drug screening and optimization according to the prediction results. If it does not meet the standard, return to the modeling process of the disease background for further improvement and optimization.
[0087] Example 2, as Figure 1 , Figure 2 shown, based on Example 1, the present invention provides a technical solution: Preferably, Step Five specifically includes:
[0088] According to the optimized model, select a new set of drug and target pairs for experimental verification to determine the targets that interact with the drugs, and the drug and target pairs cover different disease backgrounds and mechanisms of action. Set up multiple control experiments, including: positive control group, negative control group, and blank control group. Among them, the positive control group selects drugs known to have strong interactions with the targets as positive controls to ensure the effectiveness of the experiment. The negative control group selects drug and target pairs known to be ineffective as negative controls. The blank control group is set up as the group that does not receive any drug treatment, which is used to evaluate the biological effects at the basal level. And divide the experiment into small batches, with each batch containing several drug and target pairs, and repeat each experimental group multiple times (3 - 5 times) to ensure the reliability of the results. According to the designed experimental protocol, conduct the interaction experiment between the drug and the target. Through in vitro or in vivo experiments, observe the effect of the drug on the target and record the results of the interaction intensity between the drug and the target. Organize the experimental data in tabular form, including drug name, target name, experimental group, and experimental results. Then summarize the data of the repeated experiments, calculate the average value, and input the characteristic data of the new drugs and targets into the optimized drug - target interaction model. The model predicts the interaction intensity between the drug and the target based on the input data. Conduct a comprehensive comparative analysis of the predicted drug - target interaction intensity by the model and the experimental results, calculate the pharmacodynamic performance index, compare the differences between the experimental group and the predicted results, and then analyze the performance of the predicted results;
[0089] In addition, the process of comparing the differences between the experimental group and the predicted results is as follows:
[0090] Organize the experimental data of all experimental groups, including drug name, target name, experimental group (positive control, negative control, experimental group, etc.), number of experimental repetitions, and the results of the drug - target interaction intensity for each experiment. And obtain the predicted values of the interaction intensity for each drug - target pair output by the optimized drug - target interaction model. Among them, the input of the drug - target interaction model is the drug feature vector and the target feature vector. Analyze the experimental data to obtain the experimental values of the interaction intensity for each drug - target pair. Determine the benchmark value based on the average value of the positive control group. Calculate the pharmacodynamic performance index according to the predicted drug - target interaction intensity by the model, the experimentally measured drug - target interaction intensity, and the determined benchmark value. Analyze the differences between the experimental group and the predicted results, and then verify the interaction between the drug and the target. Among them, for the positive control group, the value of the pharmacodynamic performance index should be close to 1, indicating a high consistency between the model prediction and the experimental results. For the negative control group, the value of the pharmacodynamic performance index should be close to 0, indicating an inconsistency between the model prediction and the experimental results;
[0091] The expression of the pharmacodynamic performance index is:
[0092] ;
[0093] In the formula, is the drug efficacy performance index, which is used to evaluate the consistency between the model prediction and the experimental results. is the strength of the predicted drug-target interaction of the model, that is, the interaction probability value between the drug and the target. is the strength of the drug-target interaction measured experimentally. is the benchmark value, which represents the average effect of the positive control group. is the adjustment parameter, which is used to balance the weights of the predicted value and the experimental value. The value range of is [0, 1]. When is close to , is close to 1, indicating that the model prediction is highly consistent with the experimental results. When has a large difference from , is close to 0, indicating that the model prediction is inconsistent with the experimental results. When is significantly higher than the benchmark value , the logarithmic function increases, pushing towards 1. When is close to the ratio of , the radical part is close to 1, further stabilizing the value of . The adjustment parameter can be used to calibrate the deviation of the model prediction, making
[0094] Step six specifically includes:
[0095] According to the requirements of drug research and development, a discrimination threshold for the efficacy performance index is set, and the predictive result expressiveness is divided into different expressiveness levels, namely, high expressiveness level, medium expressiveness level, and low expressiveness level. The model prediction results at the high expressiveness level are highly consistent with the experimental results, with high prediction accuracy, and can be used to guide the research and development of new drugs. The model prediction results at the medium expressiveness level have a certain degree of accuracy, but there may be certain deviations, and their application in the research and development of new drugs needs to be considered carefully. The model prediction results at the low expressiveness level are inconsistent with the experimental results, with low prediction accuracy, and cannot be directly used for the research and development of new drugs. For each drug-target pair, the efficacy performance index is analyzed by combining the model prediction value and the experimental value, and the calculated efficacy performance index is compared with the preset discrimination threshold to determine the expressiveness level of each drug-target pair. The numbers of drug-target pairs at the high expressiveness level, medium expressiveness level, and low expressiveness level are counted, the overall prediction performance of the model is analyzed, and the drug-target pairs with low expressiveness are specifically analyzed to find out the reasons for inaccurate prediction. If the overall expressiveness of the model prediction results meets the high standard (i.e., most drug-target pairs meet the high expressiveness level), the model is considered accurate and can be used to guide the research and development of new drugs. According to the prediction results, drug candidate molecules with strong interactions with specific targets are screened out, and the screened drug candidate molecules are further optimized, such as improving pharmacokinetic properties, reducing toxicity, etc., and then enter the clinical trial stage to verify the effectiveness and safety of the drugs. If the overall expressiveness of the model prediction results does not meet the high standard (i.e., there are a large number of drug-target pairs at the medium expressiveness level or low expressiveness level), the model is considered to need further improvement and optimization. Return to the modeling process of the disease background, re-examine the disease mechanism, target selection, construction links of drug feature vectors and target feature vectors, and then collect more experimental data for training and optimizing the model, perform iterative optimization on the model to improve the prediction accuracy, and repeat the analysis, decision-making, and application process until the model prediction results meet the high standard.
[0096] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for predicting drug efficacy for drug research and development, characterized in that, It includes the following steps: Step 1: Collect drug and target data from multiple data sources and preprocess the drug and target data; Step 2: Based on the preprocessed drug and target data, conduct feature analysis to obtain a comprehensive feature sequence, perform gene and pathway analysis on the diseases to which the drugs are applied, and conduct disease feature encoding, and then model the disease background; Step 3: Use a graph neural network to analyze the interaction between drugs and targets, generate a feature representation of the drug-target pair, and establish a drug-target interaction model to analyze the multi-target interaction and its relationship with drugs; Step 4: Based on the multi-task learning framework, jointly analyze the relationship between drugs and multiple targets and the disease background, and then optimize the drug-target interaction model; Step 5: According to the optimized drug-target interaction model, conduct experiments with small batches and multiple control groups, input new drug and target pairs, predict the efficacy of the drugs, and verify the performance of the prediction results; Step 6: Preset a performance discrimination threshold, analyze whether the performance of the prediction results meets the standard. If it meets the standard, apply it to the research and development of new drugs. If it does not meet the standard, return to the disease background modeling process for further improvement and optimization.
2. The method for predicting drug efficacy for drug research and development according to claim 1, wherein: The specific content of Step 1 includes: Clarify the requirements for efficacy prediction in drug research and development, and then collect the corresponding drug and target data from different data sources. Among them, the drug and target data include drug chemical structures, target protein information, known drug-target interaction data, drug efficacy, and disease-related data. The data sources include public databases, literature resources, and experimental data; For drug chemical structure data, collect SMILES strings, molecular graph structures, and molecular fingerprint data; For target protein information data, collect amino acid sequences, protein three-dimensional structures, and target annotation information data; For known drug-target interaction data, collect binding affinity data and interaction types; For drug efficacy data, collect efficacy indicators and disease indication data. Among them, the efficacy indicators include cure rate, remission rate, and side effect incidence rate, and the disease indication represents the specific disease types for which the drug is used for treatment; For disease-related data, collect gene expression data, signal pathway information, and disease phenotype information. Among them, the gene expression data is used to analyze the gene expression patterns related to the disease, the signal pathway information is used to analyze the biological pathways related to the disease, and the disease phenotype information includes disease symptoms and pathogenesis; Preprocess the collected drug and target data, including data cleaning and standardization processing, associate the drug-target interaction data with the drug chemical structure and target protein information, and integrate the disease-related gene and pathway information with the drug and target data, and then integrate the preprocessed data into a unified data set.
3. A method for predicting drug efficacy for drug research and development according to claim 2, characterized in that: The specific content of Step 2 includes: Use SMILES strings to represent drug molecular structures, and convert them into molecular graph structures through tools, and then extract molecular fingerprints for quickly comparing molecular similarities, extract the quantum chemical descriptors of the molecules, and combine the chemical structures and physicochemical properties of the drugs to generate drug feature vectors; Analyze the amino acid sequence of the target protein, extract information on sequence length, amino acid composition, and conserved regions, and perform molecular docking analysis using the three-dimensional structure information of the protein. Furthermore, integrate the annotation information on the biological functions of the target and the signaling pathways involved to generate a functional feature vector of the target; Evaluate the interaction strength between the drug and the target using the known binding affinity data, and combine the characteristics of the drug and the target. Extract the interaction features of the drug-target pair through a graph neural network. Furthermore, integrate the drug feature vector, the functional feature vector of the target, and the interaction features of the drug-target pair to form a comprehensive feature sequence; Use the preprocessed disease-related gene expression data to analyze the gene expression patterns of the diseases to which the drug is applied. Through differential expression analysis, identify the differentially expressed genes related to the diseases, identify disease-related biomarkers through gene expression patterns, and use the pathway database to obtain the signaling pathway information related to the diseases. Identify the genes and proteins in the key pathways and construct a disease-related pathway network; According to the disease phenotype information, formulate disease feature coding rules, convert the expression patterns of the differentially expressed genes into feature vectors, convert the pathway information into a network graph structure, extract the key nodes and connection relationships in the pathway, and further integrate the phenotype information on disease symptoms and pathogenesis to code the disease features and form a disease feature coding sequence; Integrate the comprehensive feature sequence, the results of gene and pathway analysis, and the disease feature coding sequence to form a comprehensive feature matrix containing drugs, targets, diseases, and their related biological information, and construct a disease background model in combination with the BTDHDTA model and the comprehensive feature matrix.
4. A method for predicting drug efficacy for drug research and development according to claim 3, characterized in that: The construction process of the disease background model: Align the comprehensive feature sequence, the results of gene and pathway analysis, and the disease feature coding sequence in the same dimension and integrate them into a matrix in a row and column manner to form a comprehensive feature matrix. Among them, each row of the comprehensive feature matrix represents a sample, and each column represents a feature; Based on the data characteristics of the comprehensive feature matrix, select the BTDHDTA model architecture and divide the comprehensive feature matrix into a training set, a validation set, and a test set; Perform standardization processing on the numerical features in the comprehensive feature matrix, and construct a neural network model including a data processing module, a feature extraction module, a feature fusion module, and a prediction module according to the structure of the BTDHDTA model; Based on the determined BTDHDTA model structure, set the model parameters including hyperparameters such as learning rate, batch size, and number of iterations, and use the training set data to train the model. Update the model parameters through the backpropagation algorithm to minimize the loss function of the model on the training set. Use the validation set data to evaluate the trained model, analyze the error between the predicted value and the true value, and then, according to the performance on the validation set, iterate multiple times to adjust the model structure and parameters; For the model after multiple iterations, use the test set data to test the tuned model, evaluate the performance of the model in actual applications, and then obtain the trained disease background model. Among them, the input of the disease background model is the comprehensive feature matrix, and the output is the disease background representation, that is, the comprehensive representation of the disease background.
5. A method for predicting drug efficacy for drug research and development according to claim 4, characterized in that: The specific steps of Step 3 include: Collect the structural information of the drug and the sequence information of the target, and obtain the known drug-target interaction data. Then, convert the structural information of the drug into graph data, where atoms are used as nodes and chemical bonds are used as edges, and encode the sequence information of the target. Collect drug data involving multiple targets, including the interaction information between each drug and multiple targets. Extract features for each target separately to generate the feature representation of the target, and then combine the feature representations of multiple targets to form a multi-target feature vector. Use the GNN variant of the graph convolutional network to embed the graph data of the drug and the target to generate a low-dimensional feature representation. Through a multi-layer GNN structure, further extract the deep features of the drug and the target. Then, fuse the feature representations of the drug and the target in the feature fusion layer to generate the comprehensive feature representation of the drug-target pair. After the feature fusion layer, add a fully connected layer to predict the interaction between the drug and the target, output the binary classification result, analyze whether the drug and the target interact, select the cross-entropy loss function based on the binary classification result, and use the optimization algorithm to optimize the model parameters to minimize the loss function, and then obtain the drug-target interaction model.
6. The pharmacodynamic prediction method for drug research and development according to claim 5, characterized in that: The specific steps of Step 4 include: Based on the analysis results of drugs, targets, and disease backgrounds, integrate the feature data of drug features, target features, and disease background representations to obtain an optimized feature set, which provides input for multi-task learning. Design a shared feature extraction layer in combination with the multi-task learning framework to extract the common features of drugs, targets, and disease backgrounds, and design task-specific modules including a drug-target interaction prediction module and a disease background association module. At the same time, design a fusion layer to fuse the shared features and task-specific features to generate a comprehensive feature representation. Among them, the drug-target interaction prediction module is used to predict the interaction strength between the drug and the target, and the disease background association module is used to analyze the relationship between the drug and the disease background. Then, define the objective function of multi-task learning and combine the loss functions of drug-target interaction prediction and disease background association. Combine the optimized feature set to optimize the drug-target interaction model, and at the same time optimize the drug-target interaction prediction and disease background association tasks. Update the model parameters through backpropagation, minimize the multi-task objective function, and adjust the hyperparameters to balance the importance of different tasks. Then, calculate the evaluation metrics to evaluate the performance of the model. The evaluation metrics include the accuracy, recall rate, ROC-AUC value of drug-target interaction prediction, and the mean squared error of disease background association. Then, obtain the optimized drug-target interaction model. Using the optimized drug-target interaction model, predict the interactions between a drug and multiple targets, analyze the multi-target interactions and their relationships with the drug, determine whether there are synergistic or antagonistic effects, and based on the prediction results, analyze the interaction strengths between the drug and different targets.
7. A method for predicting drug efficacy in drug research and development according to claim 6, characterized in that: The process of analyzing the multi-target interactions and their relationships with the drug is as follows: Prepare multi-target data and load the pre-trained drug-target interaction model from a storage medium; Input the drug feature vector and the multi-target feature vector into the drug-target interaction model. The model outputs the prediction results, which are the probability values of the interactions for each drug-target pair, and summarize the prediction results of all drug-target pairs to form a complete multi-target interaction matrix. Here, the rows of the matrix represent drugs, the columns represent targets, and each element represents the possibility of the interaction between the drug and the target; Classify the interaction patterns between the drug and multiple targets based on the interaction status between the drug and multiple targets. The interaction patterns include single-target action, multi-target synergistic action, multi-target antagonistic action, and complex interaction patterns; Based on the classification of the interaction patterns, define the synergistic and antagonistic effects, and combined with the prediction results, calculate the total effect of the interaction between the drug and multiple targets, compare the total effect with the sum of the effects of each individual action, and determine whether there are synergistic or antagonistic effects. For the analysis of the synergistic effect, identify the synergistic interaction patterns between the drug and multiple targets, and by calculating the weighted sum of the interaction strengths of each drug-target pair, determine whether there is a synergistic effect. The synergistic effect is manifested as the combined action effect of the drug on multiple targets being greater than the sum of the individual actions. For the analysis of the antagonistic effect, identify the antagonistic interaction patterns between the drug and multiple targets, and by analyzing the difference in the interaction strengths of the drug on different targets, determine whether there is an antagonistic effect. The antagonistic effect is manifested as the combined action effect of the drug on multiple targets being less than the sum of the individual actions; Extract the interaction probability values between the drug and the targets from the prediction results, use them as strength indicators, compare the interaction strengths between the drug and different targets, rank the targets, and identify the main action targets of the drug.
8. A method for predicting drug efficacy in drug research and development according to claim 7, characterized in that: The specific steps of step five include: According to the optimized model, select a new set of drug and target pairs for experimental verification, determine the targets that interact with the drug, and the drug and target pairs cover different disease backgrounds and action mechanisms; Set up multiple control experiments, including: positive control group, negative control group, and blank control group. Among them, the positive control group selects drugs known to have strong interactions with the targets as positive controls, the negative control group selects drug and target pairs known to be ineffective as negative controls, the blank control group is set as the group that does not receive any drug treatment, and the experiment is divided into small batches, each batch containing several drug and target pairs, and each experimental group is repeated multiple times; According to the designed experimental protocol, conduct the experiment on the interaction between the drug and the target. Through in vitro or in vivo experiments, observe the effect of the drug on the target, and record the results of the interaction strength between the drug and the target. Organize the experimental data in tabular form, including drug name, target name, experimental group, and experimental results. Then, summarize the data of repeated experiments and calculate the average value; Input the characteristic data of the new drug and the target into the optimized drug-target interaction model. Based on the input data, the model predicts the interaction strength between the drug and the target; Conduct a comprehensive comparative analysis of the interaction strength between the drug and the target predicted by the model and the experimental results, calculate the pharmacodynamic performance index, compare the differences between the experimental group and the predicted results, and then analyze the performance of the predicted results.
9. A method for predicting drug efficacy for drug research and development according to claim 8, characterized in that: The process of comparing the differences between the experimental group and the predicted results is as follows: Organize the experimental data of all experimental groups, including drug name, target name, experimental group, number of experimental repetitions, and the results of the interaction strength between the drug and the target for each experiment. Obtain the predicted values of the interaction strength for each drug-target pair output by the optimized drug-target interaction model. Among them, the input of the drug-target interaction model is the drug feature vector and the target feature vector; Analyze the experimental data to obtain the experimental values of the interaction strength for each drug-target pair, and determine the benchmark value based on the average value of the positive control group; According to the interaction strength between the drug and the target predicted by the model, the interaction strength between the drug and the target measured experimentally, and the determined benchmark value, calculate the pharmacodynamic performance index, analyze the differences between the experimental group and the predicted results, and then verify the interaction between the drug and the target. Among them, for the positive control group, the value of the pharmacodynamic performance index should be close to 1, indicating a high degree of consistency between the model prediction and the experimental results. For the negative control group, the value of the pharmacodynamic performance index should be close to 0, indicating that the model prediction is inconsistent with the experimental results.
10. A method for predicting drug efficacy in drug research and development according to claim 9, characterized in that: The specific content of step six includes: According to the drug R & D requirements, set the discrimination threshold of the pharmacodynamic performance index, and divide the performance of the predicted results into different performance levels, namely high performance level, medium performance level, and low performance level; For each drug-target pair, analyze the pharmacodynamic performance index by combining the model prediction value and the experimental value, compare the calculated pharmacodynamic performance index with the preset discrimination threshold, and determine the performance level of each drug-target pair; Count the number of drug-target pairs at high, medium, and low performance levels, analyze the overall prediction performance of the model, and conduct a specific analysis of the drug-target pairs with low performance to find out the reasons for inaccurate prediction; If the overall performance of the model prediction meets the high standards, it is considered that the model is accurate and can be used to guide the R & D of new drugs. According to the prediction results, screen out the drug candidate molecules with strong interaction with specific targets, further optimize the screened drug candidate molecules, and then enter the clinical trial stage to verify the effectiveness and safety of the drugs; If the overall performance of the model prediction results does not meet the high standards, it is considered that the model needs to be further improved and optimized. Return to the modeling process of the disease background, re-examine the construction links of disease mechanisms, target selection, drug feature vectors, and target feature vectors, and then perform iterative optimization on the model. Repeat the analysis, decision-making, and application processes until the model prediction results meet the high standards.
Citation Information
Patent Citations
Method for determining incidence relation between drug and drug target point
CN109493925A
Drug target action relation determination method and device based on artificial intelligence
CN114360639A
Drug molecule and target protein matching method and server
CN116052762A
Drug-target interaction prediction method fusing multi-dimensional features
CN116206775A
Target protein prediction method and device for medicine, equipment and storage medium
CN116246697A
Cited By
Intelligent evaluation method and system for drug curative effect and toxicity based on multi-modal pathological data fusion
CN121983350A