Combined medication drug response prediction model based on multi-omics data and transfer learning and application
Through a combination drug response prediction model based on multiomics data and transfer learning, the problem of insufficient accuracy and interpretability of drug combination prediction in the prior art is solved, efficient drug combination prediction and recommendation are achieved, and the credibility of the model in clinical application is improved.
Patent Information
- Application Number
- CN202510327975.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing drug combination prediction models lack the ability to model the dynamic interactions of drug combinations, and most models do not integrate multiomics data, resulting in limited prediction accuracy and lack of interpretability, limiting the application of the model in clinical decision-making.
Using a combination drug response prediction model based on multiomics data and transfer learning, the parameters of a single drug response model are transferred to the combination drug model by integrating multiomics data of the drug cell line, and the transfer learning algorithm is used to transfer the parameters of a single drug response model to enhance the interpretability of the model, and a double-tower and three-tower prediction model is constructed to predict drug response.
It improves the accuracy of drug response prediction and the robustness of the model, increases the interpretability of the model, can effectively predict the drug response of different candidate drugs in combination, recommends the best therapeutic drug combination and dosage, and expands the application of the model in clinical decision-making.
Smart Images

Figure CN120299727A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of drug response prediction, and specifically to a combined drug response prediction model and application based on multi-omics data and transfer learning. Background Art
[0002] In the field of biomedicine, drug combination is a common treatment strategy, which can enhance the efficacy and reduce drug resistance through multiple mechanisms. Traditional drug combination systems mainly focus on the prediction of drug synergistic effects. However, drug response prediction systems based on dosage are relatively rare. Traditional experimental methods, such as in vitro cell line screening, although they can provide preliminary data on drug combination, are costly, time-consuming, and cannot reflect individual patient differences, resulting in inaccurate experimental results in clinical applications.
[0003] With the development of computational biology and artificial intelligence technologies, more and more computational models have been used for drug response prediction. However, most of the existing computational models are still based on single-drug response prediction. Although they can predict the effects of drugs to a certain extent, they lack the ability to model the dynamic interactions of drug combinations. The effects of drug combinations are often not simply the sum of the effects of single drugs, but involve complex drug-drug interactions and drug-organism interactions. Therefore, prediction models based solely on single-drug responses are difficult to accurately predict the effects of drug combinations. In addition, most models do not integrate multi-omics data, resulting in limited prediction accuracy; and most models have insufficient interpretability. Many existing drug response prediction models are "black box" models based on deep learning technologies. Although these models can provide relatively high prediction accuracy in some cases, their internal mechanisms are difficult to explain. The lack of interpretability makes it difficult for doctors and researchers to understand the prediction results of the models, thus limiting the application of the models in clinical decision-making. Therefore, there is an urgent need for a new prediction model to address these challenges. Summary of the Invention
[0004] In view of the deficiencies and drawbacks existing in the prior art, the present invention provides a combined drug response prediction model based on multi-omics data and transfer learning. By integrating the multi-omics data of cell lines of drugs, the prediction model adopts a transfer learning algorithm to transfer the parameters of a single-drug response model to the combined drug model, and enhances the interpretability of the model through gene function modules to predict the drug responses of different candidate drug combinations.
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: The combined drug response prediction model based on multi-omics data and transfer learning provided by the present invention includes the following steps:
[0006] S1. Data integration module: including the original data set and the preprocessing module;
[0007] The original dataset includes digital feature data of drugs, drug doses, multi-omics data of cancer samples or cell lines, single-drug response experiment data, and combination-drug response experiment data; the preprocessing module is used to preprocess the original dataset to obtain the preprocessed original dataset;
[0008] S2. Prediction model framework: including a transfer model, an input layer, an intermediate representation layer, an output layer, and a loss function in a neural network;
[0009] The transfer model specifically includes a two-tower prediction model for single-drug response and a three-tower prediction model for combination-drug response. The two-tower prediction model for single-drug response is used as a pre-training model to train a single-drug response prediction system and obtain a dimensionality reduction function from the multi-omics data input layer of cancer samples or cell lines to the gene function module layer;
[0010] S3. Interpretability module: Introduce a gene function module as a dimensionality reduction representation of cancer sample multi-omics data. Use transfer learning algorithms to transfer the dimensionality reduction function from multi-omics data to the gene function module layer learned by the two-tower prediction model for single-drug response in step S2 to the three-tower prediction model for combination-drug response, and use data to train the model to predict the drug response of different candidate drug combinations.
[0011] S4. Model evaluation: Evaluate the predicted results according to four different cross-validation experiment methods of leaving "drug pair - cell line", leaving "drug pair", leaving "cell line", and leaving "single drug", perform hyperparameter optimization, and train the final model.
[0012] Preferably, in step S1, the multi-omics data of cancer samples or cell lines includes one or several of proteomics data, transcriptomics gene expression data, and methylation and other omics data in units of genes; the digital feature data of drugs includes one or several of drug molecular chemical formulas in isoSMILE format, one-hot encoded targets, pharmacokinetic parameters, and other data;
[0013] Both the single-drug response experiment data and the combination-drug response experiment data are high-throughput screening experiment data of drug responses and are the true label values of the prediction model samples.
[0014] Preferably, in step S1, the preprocessing of the original dataset is specifically: normalizing the multi-omics data of cancer samples or cell lines and removing batch effects, and corresponding the multi-omics data features with their gene names in a "many-to-one" or "one-to-one" manner; using chemical digital tools to convert the chemical formulas in isoSMILE format into digital features.
[0015] Preferably, in step S2, a single-drug response two-tower prediction model is used as the pre-trained model to train a single-drug response prediction system. Specifically, the digital feature data of drugs and the multi-omics data of cancer samples or cell lines are used as the input layer and input into the single-drug response two-tower prediction model; the single-drug response experimental data is used as the label value of the sample.
[0016] Starting from the input layer, the single-drug response two-tower prediction model first performs two-tower dimensionality reduction representations on the single-drug feature vector and the multi-omics data feature vector of the cancer sample or cell line respectively. The multi-omics data feature vector is sequentially reduced to the gene function module layer and the cell line hidden layer; the feature vector of the single drug is reduced to the drug hidden layer; then the above two reduced feature vectors are concatenated and further reduced to obtain the output layer of the single-drug model.
[0017] The output of the single-drug response two-tower prediction model is the single-drug response IC50 or AUC.
[0018] Preferably, in step S2, the model is optimized by introducing a loss function. The loss function is the mean square error MSE between the predicted value and the true value. On this basis, a grouped regularization norm constraint is added to effectively control the complexity of the model, avoid overfitting of the model, and improve the generalization ability and stability.
[0019] Specifically, the parameters in the model are divided into three groups and different penalty coefficients are assigned to them. The first group is the gene g i belonging to the gene function module M k , indicating that prior biological knowledge has proven the relevance between the gene and the gene function module, and the coefficient is λ1; the second group is the gene g j not belonging to the gene function module M k , indicating that the prior knowledge between the gene and the gene function module is unknown, and the coefficient is λ2; the third group is other weight parameters in the model, and the corresponding coefficient is λ3, where λ2 > λ1. The specific representation formula is:
[0020]
[0021] In the formula, y i is the true data in the experimental data; f(x i ) is the data predicted by the model; N is the number of samples; λ1, λ2, and λ3 are the regularization term penalty coefficients in the model, ω is the weight parameter in the model, and p1, p2, and p3 are the types of norm constraints, which can be set as the commonly used L1 norm constraint or L2 norm constraint.
[0022] Preferably, it is characterized in that in step S3, specifically, the digital feature data of combined medication, drug dosage, and multi-omics data of cancer samples or cell lines are used as the input layer and input into the three-tower prediction model of combined medication drug response for training, and the sample true label is the experimental value of combined medication drug response;
[0023] Starting from the input layer, the three-tower prediction model respectively performs three-tower dimensionality reduction representation on the feature vectors of the two drugs and the multi-omics data feature vectors of cancer samples or cell lines. Among them, the dimensionality reduction function from the input layer of multi-omics data of cancer samples or cell lines to the gene function module is learned and migrated from the two-tower prediction model of single drug response. Through the dimensionality reduction function, the above multi-omics data is reduced to the gene function module layer and then to the dimensionality reduction hidden layer of the subsequent cancer cell line tower;
[0024] The feature vectors of the two drugs are respectively connected with the dosages of the two drugs to form drug input feature vectors, and then reduced to the hidden layers of the two drug towers; then, the two drug towers are merged and connected and reduced to the drug tower; finally, the hidden layer of the drug and the hidden layer of the multi-omics data are merged and connected to output a normalized drug response value, thereby predicting the drug response of different candidate drug combinations;
[0025] The gene function module is used to replace the fully connected hidden layer to increase the interpretability of the model, and the norm restriction mechanism is used to prevent the model from overfitting.
[0026] Preferably, in step S3, the model is optimized by introducing a loss function. The loss function is the mean square error MSE between the predicted value and the true value, and on this basis, a regularization norm restriction is added to effectively control the complexity of the model; the specific representation formula is:
[0027]
[0028] In the formula, y i is the true data in the experimental data; f(x i ) is the data predicted by the model; N is the number of samples; λ is the regularization term penalty coefficient in the model, ω is the weight parameter in the model, and p is the type of norm restriction, which can be set as the commonly used L1 norm or L2 norm restriction.
[0029] Preferably, in step S4, the evaluation indexes of the prediction model are specifically the mean square error MSE and the coefficient of determination R 2 , jointly evaluating the prediction performance of the prediction model; the mean square error MSE is used to evaluate the error between the predicted binding affinity of the prediction model and the true value, and the coefficient of determination R 2 is used to evaluate the interpretability of the prediction model for the change of binding affinity;
[0030] The specific calculation formulas are:
[0031]
[0032] where y i is the true data in the experimental data; f(x i ) is the data predicted by the model; N is the number of samples; is the mean value of the true values in the experimental data.
[0033] The combined drug response prediction model based on multi-omics data and transfer learning constructed according to the above method is used to predict the combined drug response of cancer samples or cell lines with multi-omics data, recommend the drug combination and dosage with the best efficacy based on the drug response, and conduct organoid drug response experiments and drug recommendations.
[0034] Preferably, the prediction system for constructing the combined drug response prediction model includes a data integration module, a model prediction module, and a drug recommendation module;
[0035] The data integration module is used to obtain the original data set and perform preprocessing and feature extraction on these data;
[0036] The model prediction module migrates the dimensionality reduction function from the input layer of the multi-omics data of cancer samples or cell lines to the gene function module layer obtained by predicting the two towers of a single drug response into the three-tower prediction model of the combined drug response, so as to predict the drug response of different candidate drugs in combination.
[0037] The drug recommendation module studies the synergistic effect of drugs based on the drug response data, and conducts drug recommendation and screening for combined drugs according to the predicted drug response.
[0038] The present invention provides a combined drug response prediction model based on multi-omics data and transfer learning and its application. It has the following beneficial effects:
[0039] (1) The combined drug response prediction model based on multi-omics data and transfer learning of the present invention can effectively improve the accuracy of predicting drug response by integrating the digital feature data and drug dosage of drugs with the multi-omics data of cancer samples or cell lines; and adopts the transfer learning algorithm to migrate the dimensionality reduction function from the input layer of the multi-omics data learned by the two-tower prediction model of a single drug response to the gene function module layer into the three-tower prediction model of the combined drug response, which can effectively avoid overfitting caused by too many parameters; improve the accuracy of drug response prediction and the robustness of the model.
[0040] At the same time, introducing a gene function module to replace the fully connected hidden layer and introducing grouped regularization penalty can effectively avoid overlearning, further increase the interpretability of the model, and thus effectively predict the drug response of different candidate drugs in combination.
[0041] (2) The drug responses of different candidate drug combinations for cancer samples with multi-omics data predicted by this prediction model can be used to predict the drug responses of samples under different pairwise drug combinations and doses, recommend the drug combinations and doses with the best therapeutic effects, and further screen and recommend drugs for organoid drug response experiments. Description of the Drawings
[0042] Figure 1 It is a schematic diagram of the prediction model framework in Example 1 of the present invention;
[0043] Figure 2 It is a schematic diagram of the experimental data set designed by four cross-validation strategies in Example 1 of the present invention. Detailed Implementation Modes
[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0045] Example 1
[0046] As Figure 1 shown, the construction method of the combined drug response prediction model based on multi-omics data and transfer learning includes the following steps:
[0047] S1. Data integration module: including digital feature data of drugs, drug doses, multi-omics data of cancer samples or cell lines, single-drug response experiment data, and combined drug response experiment data; the digital feature data of drugs include chemical formulas in isoSMILE format, one-hot encoded targets, and pharmacokinetic parameters; the multi-omics data include proteomics data, transcriptomics gene expression data, and methylation data in units of genes; both the single-drug response experiment data and the combined drug response experiment data are high-throughput screening experiment data of drug responses and are the true label values of the prediction model samples.
[0048] Preprocess the obtained original data set, perform normalization operations on the multi-omics data of cancer samples or cell lines and remove batch effects, and correspond the multi-omics data features with their gene names in a "many-to-one" or "one-to-one" manner; use chemical digital tools to convert the chemical formulas in isoSMILE format into digital features. Obtain the preprocessed data set.
[0049] S2. Prediction Model Framework: It includes a transfer model, an intermediate representation layer, an output layer, and a loss function in a neural network. The transfer model specifically includes a two-tower prediction model for single-drug response and a three-tower prediction model for combination-drug response. The two-tower prediction model for single-drug response is used as a pre-training model to train a single-drug response prediction system and obtain the dimensionality reduction parameters of multi-omics data of cancer samples or cell lines.
[0050] The preprocessed original dataset is used as the input layer and input into the prediction model. The prediction model includes a two-tower prediction model for single-drug response and a three-tower prediction model for combination-drug response. The two-tower prediction model for single-drug response is used as a pre-training model to obtain the dimensionality reduction function of multi-omics data of cancer samples or cell lines. Starting from the input layer, the two-tower prediction model for single-drug response first performs two-tower dimensionality reduction representations on the single-drug feature vector and the multi-omics data feature vector of cancer samples. The multi-omics data feature vector is sequentially reduced to the gene function module layer and other hidden layers; the feature vector of the single drug is reduced to other hidden layers; then the above two dimensionally reduced feature vectors are connected in the intermediate representation layer to obtain the dimensionality reduction parameters of multi-omics data of cancer samples or cell lines.
[0051] The output of the two-tower prediction model for single-drug response is the IC50 or AUC of single-drug response.
[0052] S3. Interpretability Module: The gene function module is introduced as a dimensionality reduction representation of multi-omics data. Using the transfer learning algorithm, the dimensionality reduction function from the input layer to the gene function module layer of the multi-omics data learned by the two-tower prediction model for single-drug response is transferred to the three-tower prediction model for combination-drug response, and the model is trained using data to predict the drug response of different candidate drug combinations.
[0053] Specifically, it also includes inputting the digital feature data of combination drugs, drug doses, and multi-omics data of cancer samples or cell lines into the three-tower prediction model for combination-drug response for training.
[0054] Starting from the input layer, the three-tower prediction model performs three-tower dimensionality reduction representations on the feature vectors of two drugs and the multi-omics data feature vector of cancer samples or cell lines respectively. Among them, the multi-omics data is learned and transferred from the two-tower prediction model for single-drug response, and the above multi-omics data is reduced to the gene function module layer through the dimensionality reduction function and then reduced to the subsequent multi-omics data hidden layer.
[0055] The characteristics of the two drugs are respectively connected with the dosage, and then dimensionally reduced to the hidden layers of the two drug towers. In the prediction model, since the two drugs are in a symmetric relationship, the dimensionality reduction functions of the two drug towers share parameters. Secondly, considering the symmetry of the two drug towers in the model, the hidden layer nodes of the two drug towers can be connected by taking the mean value to form an overall drug hidden layer, and then the drug hidden layer is dimensionally reduced through several dimensionality reduction functions. Then, the drug hidden layer and the hidden layer of the multi-omics data are merged and connected, and then dimensionally reduced to represent the output of the normalized drug response value, so as to predict the drug response of different candidate drugs in combination.
[0056] The dimensionality reduction function from the m-th layer network to the m+1-th layer is expressed as:
[0057]
[0058] Among them, represents the j-th node of the m+1-th layer network, represents the i-th node of the m-th layer network.
[0059] Both of the above two models have gene function modules; the gene function module is a dimensionality reduction representation of the multi-omics data of cancer samples or cell lines; since organisms perform biological functions not independently with a single gene as a unit, but cooperate with each other with gene function modules as a unit, by using gene function modules to replace the fully connected hidden layer, the interpretability of the model is increased, and the grouped norm constraint is adopted to reduce the number of parameters from the multi-omics data to the gene function module layer, which can effectively avoid overfitting; and because there are many cell lines in the experimental data of a single drug response, using the dimensionality reduction parameters of the single drug response model to train the omics data can effectively avoid fitting caused by too many parameters, and further increase the interpretability of the model.
[0060] And the model is optimized by introducing a loss function. The loss function is the mean square error MSE between the predicted value and the true value. On this basis, the L1 norm constraint or L2 norm constraint of the weight parameter is added to effectively control the complexity of the model. At the same time, by adding a penalty term for the weight in the loss function, overfitting of the model can also be avoided, and the generalization ability and stability are improved.
[0061] In the single drug response prediction model, a grouped norm constraint strategy is adopted. The specific formula of the loss function of the single drug response prediction model is:
[0062]
[0063] In the formula, y i is the true data in the experimental data; f(x i ) is the data predicted by the model; N is the number of samples; the weight parameters of the model are divided into three groups: when gene gi Belonging to module M k When it is, the weight from the input layer of cell line multi-omics data to the functional module layer is the first group; when gene g j Does not belong to module M k When it is, the weight from the input layer of cell line multi-omics data to the functional module layer is the second group; the remaining weights in the model are the third group, and the regularization term coefficients corresponding to the three groups of weights are λ1, λ2, λ3, (λ2>λ1). The network weight parameter corresponding to each item is represented by the corresponding ω in the corresponding regularization term; p1, p2, p3 are the selected types of regularization terms, which can be L1 norm or L2 norm.
[0064] The loss function of the combined drug response prediction model is specifically represented by the formula:
[0065]
[0066] In the formula, y i Is the true data in the experimental data; f(x i ) is the data predicted by the model; N is the number of samples. λ is the regularization term penalty coefficient in the model, ω is the weight parameter in the model, and p is the type of norm limit, which can be set to the common L1 norm or L2 norm limit.
[0067] S4. Evaluation of the model: As Figure 2 Shown, four cross-validation strategies are used to evaluate the prediction model, perform hyperparameter optimization, and train the final model; the four cross-validation strategies include leaving "drug pair-cell line", leaving "drug pair", leaving "cell line", and leaving "single drug", comprehensively evaluating the prediction results output by the model in different scenarios.
[0068] The evaluation metrics of the prediction model are specifically the mean squared error MSE and the coefficient of determination R 2 , jointly evaluating the prediction performance of the prediction model; the mean squared error MSE is used to evaluate the error between the predicted binding affinity of the prediction model and the true value, and the coefficient of determination R 2 Is used to evaluate the explanatory ability of the prediction model for the change in binding affinity;
[0069] The specific calculation formulas are:
[0070]
[0071] In the formula, y i Is the true data in the experimental data; f(x i ) is the data predicted by the model; N is the number of samples; Is the mean of the true values in the experimental data.
[0072] As shown in Table 1, the prediction results of the combined drug response prediction model based on multi-omics data and transfer learning constructed in the present invention were verified with those of other existing prediction models. The prediction results of four cross-validations of the prediction model of the present invention and other prediction models on the NCI-dataset were compared. The experimental designs of the four cross-validations are as Figure 2 shown, and the comparison metrics are the mean squared error MSE and the coefficient of determination R 2 . The specific results and their analyses are as follows:
[0073]
[0074]
[0075] Table 1 Results of four cross-validations of the prediction model PreDeepDrug of the present invention, the multi-layer neural network, and the lightGBM artificial intelligence prediction method on the NCI-dataset
[0076] As can be seen from Table 1, first, the mean squared error MSE of the model PreDeepDrug of the present invention is smaller than the error values of the other two prediction methods in the four cross-validations on the NCI-dataset. Therefore, the prediction performance of the prediction model of the present invention is better. Second, the coefficient of determination R 2 of the model PreDeepDrug of the present invention is higher than the other two prediction methods in the four cross-validations on the NCI-dataset. Therefore, the interpretability of this model is stronger than that of other prediction methods. From the above comparison, it can be concluded that the combined drug response prediction model based on multi-omics data and transfer learning constructed in the present invention has higher overall performance and prediction result accuracy. Based on the prediction of the combined drug response of cancer samples with multi-omics data using this model, further research can be carried out to recommend the drug combination and dosage with the best efficacy and study the synergistic effect of drugs. According to the predicted response, drug recommendation and screening for combined drug use can also be carried out, further expanding the application of the prediction model in clinical decision-making.
[0077] In summary, the combined drug response prediction model based on multi-omics data and transfer learning of the present invention can effectively improve the detection accuracy by integrating the dataset. By using transfer learning, the dimensionality reduction function learned from the input layer to the gene function module layer of the single-drug response double-tower prediction model is transferred to the combined drug response triple-tower prediction model, and a gene function module is introduced to replace the fully connected hidden layer, increasing the interpretability of the model and reducing the number of parameters from multi-omics data to the gene function module layer, effectively avoiding overfitting; effectively improving the accuracy of drug interaction prediction and the robustness of the model.
[0078] Example 2
[0079] A combined drug response prediction model based on multi-omics data and transfer learning can predict the combined drug response of cancer samples with multi-omics data, recommend the drug combination and dosage with the best efficacy based on the drug response, and further conduct organoid drug response experiments and drug recommendations.
[0080] Adopt the combined drug response prediction model based on multi-omics data and transfer learning constructed in Example 1, and according to the steps in Example 1, obtain the original data set through the data integration module, and preprocess and extract features from these data; use the trained model to predict the combined drug response prediction of cancer samples with multi-omics data.
[0081] The drug recommendation module recommends the drug combination and dosage with the best efficacy based on the drug response, and uses the predicted data for further research, such as studying the synergistic effect of drugs, or recommending and screening drugs for combined drug use according to the predicted response, and then applying it to patients, which can be used for the prediction and optimized recommendation of the efficacy of drug combinations in personalized medicine.
[0082] On the basis of the above embodiments, the present invention continues to describe in detail the technical features involved and the functions and roles played by these technical features in the present invention to help those skilled in the art fully understand the technical solution of the present invention and reproduce it.
[0083] Finally, although this specification is described according to the embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A combined drug response prediction model based on multi-omics data and transfer learning, characterized in that The steps are as follows: S1. Data integration module: including an original data set and a preprocessing module; The original data set includes digital characteristic data of drugs, drug doses, multi-omics data of cancer samples or cell lines, single-drug response experiment data, and combination drug response experiment data; The preprocessing module is used to preprocess the original data set to obtain preprocessed input data; S2. Prediction model framework: including a transfer model, an input layer, an intermediate representation layer, an output layer, and a loss function in a neural network; The transfer model specifically includes a two-tower prediction model for single-drug response and a three-tower prediction model for combination drug response. The two-tower prediction model for single-drug response is used as a pre-training model to train a single-drug response prediction system and obtain a dimensionality reduction function from the input layer of multi-omics data of cancer samples or cell lines to the gene function module layer; S3. Interpretability module: introducing a gene function module as a dimensionality reduction representation of multi-omics data of cancer samples, using a transfer learning algorithm to transfer the dimensionality reduction function from multi-omics data to the gene function module layer learned by the two-tower prediction model for single-drug response in step S2 to the three-tower prediction model for combination drug response, and training the model with data to predict the drug response of different candidate drug combinations; S4. Model evaluation: evaluating the prediction results according to the cross-validation experiment methods of four different strategies of leaving "drug pair - cell line", leaving "drug pair", leaving "cell line", and leaving "single drug", performing hyperparameter optimization, and training the final model.
2. The combination drug response prediction model based on multi-omics data and transfer learning according to claim 1, in step S1, the multi-omics data of the cancer sample or cell line includes one or more of proteomics data, transcriptomics gene expression data, and methylation data in units of genes; the digital characteristic data of the drug includes one or more of the chemical structural formula in isoSMILE format of the drug, the one-hot encoded target, and pharmacokinetic parameters; The single-drug response experiment data and the combination drug response experiment data are both high-throughput screening experiment data of drug response and are the true label values of the prediction model samples.
3. The combination drug response prediction model based on multi-omics data and transfer learning according to claim 2, in step S1, the preprocessing of the original data set is specifically: normalizing the multi-omics data of the cancer sample or cell line and removing batch effects, and corresponding the multi-omics data features to their gene names in a "many-to-one" or "one-to-one" manner; using a chemical digitization tool to convert the drug molecular chemical formula in isoSMILE format into digital characteristics.
4. The combined drug response prediction model based on multi-omics data and transfer learning according to claim 1, wherein In step S2, the two-tower prediction model for single-drug response is used as a pre-training model to train a single-drug response prediction system. Specifically, the digital characteristic data of the drug and the multi-omics data of the cancer sample or cell line are used as the input layer and input into the two-tower prediction model for single-drug response, and the sample label is the true single-drug response experiment data; The two - tower prediction model for single - drug response starts from the input layer. The single - drug feature vector and the multi - omics data feature vectors of cancer samples or cell lines are first respectively subjected to two - tower dimensionality reduction representation. The multi - omics data feature vectors are sequentially reduced to the gene function module layer and the cell line hidden layer; the feature vector of the single drug is reduced to the drug hidden layer; then the above two dimensionality - reduced feature vectors are connected in the intermediate representation layer and further reduced to obtain the output layer of the single - drug model. The output of the two - tower prediction model for single - drug response is the single - drug sensitivity value IC50 or AUC.
5. The combined drug response prediction model based on multi-omics data and transfer learning according to claim 4, wherein In step S2, the model is optimized by introducing a loss function. The loss function is the mean squared error MSE between the predicted value and the true value. On this basis, a grouped regularization norm constraint is added to effectively control the complexity of the model, avoid overfitting of the model, and improve the generalization ability and stability. Specifically, the weight parameters in the model are divided into three groups, and different penalty coefficients are assigned to them. The first group is gene g i belongs to gene function module M k , indicating that prior biological knowledge has proven the relevance between the gene and the gene function module, with a coefficient of λ1; the second group is gene g j does not belong to gene function module M k , indicating that the prior knowledge of the gene and the gene function module is unknown, with a coefficient of λ2; the third group is other weight parameters in the model, and the regularization term coefficient is λ3, where λ2 > λ1;; The specific representation formula is: where y i is the true data in the experimental data; f(x i ) is the data predicted by the model; N is the number of samples, λ1, λ2, and λ3 are the regularization term penalty coefficients in the model, ω is the weight parameter in the model, and p1, p2, and p3 are the types of norm constraints, which can be set to the commonly used L1 norm constraint or L2 norm constraint.
6. The combined drug response prediction model based on multi-omics data and transfer learning according to claim 1, wherein In step S3, it specifically includes using the digitalized feature data of combined drugs, drug doses, and the multi - omics data of cancer samples or cell lines as the input layer and inputting them into the three - tower prediction model for combined - drug response for training. The true label of the sample is the experimental value of the combined - drug response. The three - tower prediction model starts from the input layer. The feature vectors of two drugs and the multi - omics data feature vectors of cancer samples or cell lines are respectively subjected to three - tower dimensionality reduction representation. Among them, the dimensionality reduction function from multi - omics data to the gene function module is learned and migrated from the two - tower prediction model for single - drug response. Through the dimensionality reduction function, the above - mentioned multi - omics data is reduced to the gene function module layer and then to the hidden layer of the cancer cell line tower. The feature vectors of the two drugs are respectively connected with two drug doses to form drug input feature vectors, and then reduced to the hidden layers of the two drug towers; then, the two drug towers are merged and connected, and reduced to the drug tower; finally, the hidden layer of the drug and the hidden layer of the multi - omics data are merged and connected to output a normalized drug response value, thereby predicting the drug response of different candidate drug combinations. The gene function module is used to replace the fully - connected hidden layer to increase the interpretability of the model, and the norm constraint mechanism is used to prevent the model from overfitting.
7. The combined drug response prediction model based on multi-omics data and transfer learning according to claim 6, wherein In step S3, the model is optimized by introducing a loss function. The loss function is the mean squared error MSE between the predicted value and the true value. On this basis, a regularization norm constraint is added to effectively control the complexity of the model; the specific expression formula is: where y i is the true data in the experimental data; f(x i ) is the data predicted by the model; N is the number of samples; λ is the regularization term penalty coefficient in the model, ω is the weight parameter in the model, and p is the type of norm limit, which can be set to the commonly used L1 norm or L2 norm limit.
8. The combined drug response prediction model based on multi-omics data and transfer learning according to claim 1, wherein, In step S4, the evaluation metrics of the prediction model are specifically the mean squared error MSE and the coefficient of determination R 2 , jointly evaluating the prediction performance of the prediction model; the mean squared error MSE is used to evaluate the error between the predicted binding affinity of the prediction model and the true value, and the coefficient of determination R 2 is used to evaluate the explanatory ability of the prediction model for the change in binding affinity; The specific calculation formula is: where y i is the true data in the experimental data; f(x i ) is the data predicted by the model; N is the number of samples; is the mean of the true values in the experimental data.
9. Application of a combined drug response prediction model based on multi-omics data and transfer learning, characterized in that, The combined - drug response prediction model based on multi - omics data and transfer learning constructed according to any one of claims 1 - 7 is used to predict the combined - drug response of cancer samples with multi - omics data, give the drug combination and dose with the best therapeutic effect based on the drug response, and conduct organoid drug response experiments and drug recommendations.
10. Use of the combined drug response prediction model based on multi-omics data and transfer learning according to claim 9, characterized in that, It is used to construct a combined - drug response prediction model, including a data integration module, a model prediction module, and a drug recommendation module. The data integration module is used to obtain the original data set and pre - process and extract features from these data. The model prediction module migrates the dimensionality reduction function from the input layer of multi-omics data of cancer samples or cell lines to the gene function module layer through single-drug response two-tower prediction into the combination drug response three-tower prediction model, so as to predict the drug response of different candidate drugs in combination. Based on the drug response data, the drug recommendation module studies the synergistic effect of drugs and makes drug recommendations and screening for combination drugs according to the predicted drug response.
Citation Information
Cited By
Method and system for detecting thickness of LED wafer with vertical structure and computer equipment
CN120947454A