Methods and devices for predicting response to anticancer drugs

CN116543830BActive Publication Date: 2026-08-18GENEIS TECH BEIJING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310516615.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-09
Publication Date
2026-08-18
Estimated Expiration
2043-05-09

AI Technical Summary

Technical Problem

例如,SRMF模型其优化问题为一个非凸函数,该模型采用交替极小化算法只搜索局部极小值而不搜索全局极小值;MC模型、RR模型及两者结合的MCRR模型预测精度有待进一步提高等

Benefits of technology

[0016] The present invention provides a method for predicting anticancer drug response, proposing an MCLRP model based on known data and gene expression profiles to predict drug sensitivity matrix, thereby improving interpretability and prediction accuracy. Preferably, the method further includes a model evaluation or validation step, mainly comprising: a comparison of prediction performance based on PCC and SCC, a new drug-cancer association prediction analysis based on the GDSC dataset (at the cellular level), and a reliability analysis of prediction results based on consistency identification (at the gene level). The evaluation of the model and its validation at both the cellular and gene levels are other major innovations of the present invention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116543830B_ABST
    Figure CN116543830B_ABST
Patent Text Reader

Abstract

The application discloses a method and device for predicting anticancer drug response. The method for predicting anticancer drug response comprises the following steps: obtaining drug response spectrum and gene expression spectrum data of anticancer drugs, thereby obtaining a response matrix and an expression matrix; pre-processing the response matrix and the expression matrix, thereby obtaining a standardized matrix and a processed expression matrix; constructing an MCLRP model based on the standardized matrix and the processed expression matrix, solving the MCLRP model by using an ADMM algorithm, and obtaining a prediction matrix; and using the prediction matrix to predict the anticancer drug response. In some embodiments, the model is further evaluated, including a comparison of prediction performance based on a Pearson correlation coefficient (PCC) and a Spearman correlation coefficient (SCC), a new drug-cancer correlation prediction analysis based on a GDSC dataset from a cell level, and a reliability analysis of a prediction result based on consistency identification from a gene level. Thus, the purposes of improving interpretability and prediction accuracy are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of anticancer drug development, and more specifically to a method and apparatus for predicting anticancer drug response. Background Technology

[0002] Predicting drug response and resistance in cancer cells plays a crucial role in cancer drug selection, yet it remains largely unknown. Large-scale experimental studies are helpful in selecting drugs to treat different types of cancer, but they are limited by experimental conditions, high costs, and a lack of data integration. Furthermore, patients with the same type of cancer may respond differently to specific drug treatments. The goal of precision cancer drugs is to decipher the causes of a specific patient's cancer at the molecular level and then tailor treatment plans to slow cancer progression. Therefore, identifying predictive biomarkers of individual drug sensitivity is key to advancing the development of precision cancer drugs.

[0003] Compared to human and animal models, human cancer cell lines offer greater manipulability for cancer biology and drug discovery through simpler experimental procedures. Several large-scale high-throughput screenings have categorized the genomic and pharmacological data of hundreds of human cancer cell lines. Current research indicates the coexistence of multiple mutated genes within tumors, increasingly limiting the development of effective treatments based solely on tumor characteristics. The development of computational methods that link the genomic profiles of cancer cell lines to drug responses could facilitate the development of precision cancer drugs. Identified genomic biomarkers can be used to predict anticancer drug responses, thus enabling the identification of clinically relevant biomarkers using large-scale drug sensitivity profiles within conventional cancer cell line models.

[0004] With the continuous development of computing technology, computational methods have been increasingly applied to the field of drug response prediction in order to further reduce drug development costs and shorten the development cycle, attracting considerable attention from scholars. Currently, computational models for predicting drug response are mainly divided into four categories: regression methods, matrix factorization methods, Bayesian inference methods, and matrix completion-based methods. However, existing computational methods have certain limitations. For example, some machine learning-based models classify cell line drugs into sensitive or resistant groups, predicting only binary results; some network-based or matrix factorization-based methods provide limited information on the relationship between genes / biomarkers and drug response, resulting in poor interpretability. Several computational models that currently perform well in predicting drug response also have limitations. For instance, the SRMF model's optimization problem is a non-convex function, and this model uses an alternating minimization algorithm to search only for local minima without searching for global minima; the prediction accuracy of the MC model, RR model, and the combined MCRR model needs further improvement. Summary of the Invention

[0005] To address at least some of the technical problems in the prior art, the present invention provides a method for predicting response to anticancer drugs. Specifically, the present invention includes the following.

[0006] A first aspect of the present invention provides a method for predicting response to anticancer drugs, comprising:

[0007] Obtain drug response profiles and gene expression profiles of anticancer drugs to obtain response matrices and expression matrices;

[0008] The response matrix and expression matrix are preprocessed to obtain the standardized matrix and the processed expression matrix;

[0009] An MCLRP model is constructed based on the standardized matrix and the processed expression matrix. The ADMM algorithm is then used to solve the MCLRP model to obtain the prediction matrix.

[0010] Prediction matrices are used to predict responses to anticancer drugs.

[0011] A second aspect of the present invention provides an apparatus for predicting response to anticancer drugs, comprising:

[0012] The acquisition module is used to acquire drug response profiles and gene expression profiles of anticancer drugs, thereby obtaining response matrices and expression matrices.

[0013] The preprocessing module is used to preprocess the response matrix and expression matrix to obtain the standardized matrix and the processed expression matrix.

[0014] The model building module is used to construct an MCLRP model based on the standardized matrix and the processed expression matrix, and to solve the MCLRP model using the ADMM algorithm to obtain the prediction matrix;

[0015] The prediction module is used to predict responses to anticancer drugs using a prediction matrix.

[0016] The present invention provides a method for predicting anticancer drug response, proposing an MCLRP model based on known data and gene expression profiles to predict drug sensitivity matrix, thereby improving interpretability and prediction accuracy. Preferably, the method further includes a model evaluation or validation step, mainly comprising: a comparison of prediction performance based on PCC and SCC, a new drug-cancer association prediction analysis based on the GDSC dataset (at the cellular level), and a reliability analysis of prediction results based on consistency identification (at the gene level). The evaluation of the model and its validation at both the cellular and gene levels are other major innovations of the present invention. Attached Figure Description

[0017] Figure 1 A flowchart of a method for predicting anticancer drug response as an exemplary embodiment;

[0018] Figure 2A A schematic diagram of the Pearson correlation coefficient between predicted and observed values ​​in CCLE, as disclosed in an exemplary embodiment;

[0019] Figure 2B A schematic diagram of the Spearma correlation coefficient between predicted and observed values ​​in CCLE, as an exemplary embodiment;

[0020] Figure 2C A schematic diagram of the Pearson correlation coefficient between predicted and observed values ​​in ERKAUC, as an exemplary embodiment;

[0021] Figure 2D A schematic diagram of the Pearson correlation coefficient between ERKIC50 predicted values ​​and observed values, as shown in an exemplary embodiment;

[0022] Figure 2E A schematic diagram of the Spearman correlation coefficient between predicted and observed values ​​in ERKAUC, as an exemplary embodiment;

[0023] Figure 2F A schematic diagram of the Spearman correlation coefficient between ERKIC50 predicted values ​​and observed values, as shown in an exemplary embodiment;

[0024] Figure 2G A scatter plot of observing and predicting active regions using MCLRP for an exemplary embodiment of the AZD6244;

[0025] Figure 2H A scatter plot illustrating the observation and prediction of active regions using MCLRP in Irinotecan as an exemplary embodiment;

[0026] Figure 2I A scatter plot of PD-0325901 using MCLRP to observe and predict active regions as an exemplary embodiment;

[0027] Figure 2J A scatter plot illustrating the observation and prediction of active regions using MCLRP in Topotecan as an exemplary embodiment;

[0028] Figure 3A A schematic diagram of a drug-cell line association network in the PI3K pathway where the observed and predicted P-values ​​for both mutant and wild types are less than 0.05, as an exemplary embodiment.

[0029] Figure 3B A schematic diagram of a drug-cell line association network in the ERK pathway where the observed and predicted values ​​for both mutant and wild-type cells are less than 0.05, as an exemplary embodiment.

[0030] Figure 3C A schematic diagram illustrating the consistency identification results of BRAF mutant and wild-type cell lines PD-0325901 in ERKAUC as an exemplary embodiment.

[0031] Figure 3D A schematic diagram illustrating the consistency identification results of RDEA119 in ERKAUC for predictive and observational data of BRAF mutant and wild-type cell lines, as an exemplary embodiment.

[0032] Figure 3E A schematic diagram illustrating the consistency identification results of predicted and observed data of BRAF mutant and wild-type cell lines SB590885 in ERKAUC as an exemplary embodiment.

[0033] Figure 3F A schematic diagram illustrating the results of ERKIC50 concordance identification between predictive and observational data of BRAF mutant and wild-type cell lines using AZ628, as an exemplary embodiment.

[0034] Figure 3G A schematic diagram illustrating the results of ERKIC50 concordance assessment between predictive and observational data of NRAS mutant and wild-type cell lines using AZ628, as an exemplary embodiment.

[0035] Figure 3H A schematic diagram illustrating the results of ERKIC50 concordance assessment between predictive and observational data of RB1 mutant and wild-type cell lines using AZ628, as an exemplary embodiment.

[0036] Figure 4A A Sankey schematic diagram illustrating the functional annotation of GO-MF (cellular components) in at least two of the 28 drugs of PI3KAUC, as an exemplary embodiment;

[0037] Figure 4B A Sankey schematic diagram illustrating the functional annotation of KEGG in at least two of the 28 drugs of PI3KAUC, as an exemplary embodiment;

[0038] Figure 4C A heatmap of the average response values ​​of tissues and cell lines using the drug in an exemplary embodiment of the PI3KAUC dataset;

[0039] Figure 5A This is a schematic diagram illustrating the number of cell lines sensitive to each drug in the new drug-cancer association prediction based on the GDSC dataset, using different thresholds of 35%, 40%, and 45% of the observations on PI3KAUC, as an exemplary embodiment.

[0040] Figure 5B This is a schematic diagram illustrating the number of cell lines sensitive to each drug in the new drug-cancer association prediction based on the GDSC dataset, using different thresholds of 35%, 40%, and 45% of the observations on PI3KIC50, as an exemplary embodiment.

[0041] Figure 5C This is a schematic diagram illustrating the number of cell lines sensitive to each drug in the new drug-cancer association prediction based on the GDSC dataset, using different thresholds of 35%, 40%, and 45% of the observations on ERKAUC as an exemplary embodiment.

[0042] Figure 5D This is a schematic diagram illustrating the number of cell lines sensitive to each drug in the new drug-cancer association prediction based on the GDSC dataset, using different thresholds of 35%, 40%, and 45% of the observations on ERKIC50, as an exemplary embodiment.

[0043] Figure 5E This is a schematic diagram illustrating the response values ​​of cell lines sensitive to each drug in new drug-cancer association prediction based on the GDSC dataset, using a threshold of 40% of the observations on PI3KAUC as an exemplary embodiment.

[0044] Figure 5F This is a schematic diagram illustrating the response values ​​of cell lines sensitive to each drug in new drug-cancer association prediction based on the GDSC dataset, using a threshold of 40% of the observations on ERKAUC as an exemplary embodiment. Detailed Implementation

[0045] Various exemplary embodiments of the present invention will now be described in detail. This detailed description should not be considered as a limitation of the present invention, but rather as a more detailed description of certain aspects, features, and embodiments of the present invention.

[0046] The calculation parameters are mathematical factors that participate in the calculation.

[0047] This embodiment combines gene expression data with drug sensitivity data for data augmentation. Based on the fundamental assumption that similar drugs with similar molecular pathways have similar drug sensitivities to the same cancer cell line, anticancer drug response prediction is treated as a noise matrix completion problem, and a low-rank regularization method is proposed to address this issue. During model construction, instead of strictly fitting known elements, constraints are used to allow deviations of the drug response matrix from known values. Based on the above analysis, the anticancer drug response prediction problem can be transformed into a minimum-rank matrix problem within the allowable noise range. Figure 1 As shown, the main construction and solution steps of this model are as follows.

[0048] First, drug response profiles and gene expression profiles of anticancer drugs are obtained, and response matrices and expression matrices are derived. The expression matrix can be represented as X = [X... ij ] m×n , where m is the number of cell lines and n is the number of genes.

[0049] Secondly, the two matrix data are preprocessed. All missing values ​​in the response matrix are replaced with 0, and then each column is standardized to obtain a standardized matrix, denoted as M′=[M ik ′] m×p , where m is the number of cell lines and p is the number of drugs; principal component analysis is used to reduce the dimensionality of the expression matrix to avoid overfitting.

[0050] Next, we construct the MCLRP model. In constructing this model, we assume that M can be linearly expressed through the columns of X. Since drug sensitivity data contains noise, the model should be able to effectively tolerate potential noise. Based on the above analysis, this problem can be transformed into a minimum rank problem within the allowable noise range. However, since the rank function is non-convex, it can be replaced by a kernel function, transforming it into a convex optimization problem. Furthermore, since the noise level is uncertain, selecting appropriate parameters for the model is challenging. To simplify the solution, we can relax the model into a regularized model and use methods such as variable substitution and augmented Lagrangian to perform equivalent transformations on the objective function, further simplifying the solution.

[0051] Therefore, the ADMM algorithm is used to solve the model to obtain the prediction matrix. The specific process of MCLRP model construction and ADMM algorithm solution will be described in detail in the implementation section.

[0052] Finally, the model was evaluated. To increase the robustness of the estimated parameters, cross-validation, such as ten-fold cross-validation, was used. To improve the reliability of the results, the cross-validation process was repeated ten times to obtain the average index. To evaluate the predictive performance of the model, the Person correlation coefficient (PCC) and Spearman correlation coefficient (SCC) for each drug were calculated. The PCC value represents the degree of correlation between the predicted and observed drug response curves, and the SCC represents the correlation between the predicted and observed values. The larger the two indices, the better the predictive performance.

[0053] Based on two publicly available datasets, CCLE and GDSC, the aforementioned model was used to predict anticancer drug responses, and two sets of model evaluation metrics, SCC and PCC, were obtained. The model-based evaluation metrics were compared with those of four other drug response prediction models: MC, RR, MCRR, and SRMF. Overall, the model disclosed in this embodiment performed better. Furthermore, a series of drug sensitivity effects were predicted, and top-level genes related to drug sensitivity or resistance were generated.

[0054] The specific process of constructing the MCLRP model and solving it using the ADMM algorithm is as follows:

[0055] The model construction process is based on the assumption that similar drugs with similar molecular pathways have similar drug sensitivities to the same cancer cell line. Furthermore, since drug sensitivity data contains noise, the model should be able to effectively tolerate potential noise. Based on the above analysis, the anticancer drug response prediction problem can be transformed into a minimum-rank matrix problem within the allowable noise range. Let M′=[M ik ′] m×p Represents the standardized drug response matrix, X = [X ij ] m×n Let X represent the gene expression matrix, where m is the number of cell lines, p is the number of drugs, n is the number of genes, and M′ is the normalized response matrix. In the model construction process, it is assumed that M can be linearly expressed through the columns of X. The estimation and optimization problem of the drug response matrix is ​​as follows:

[0056] Treating the problem of predicting anticancer drug sensitivity as a matrix completion problem—an algorithm for filling in missing terms in a partial observation matrix, equivalent to data imputation in statistics—further transforms the problem into a low-rank matrix completion problem based on the premise that similar anticancer drugs elicit similar responses in similar cancer cell lines.

[0057] min M rank(M),

[0058] stP Ω (M)=P Ω (M′),

[0059] Where rank(·) is the rank function, and Ω is the set of observations on M′. It is a projection operator, that is

[0060]

[0061] This embodiment assumes a linear relationship between the gene expression profile of cell lines and the response to anticancer drugs, and the problem is transformed into...

[0062] min Mrank(M),

[0063] stP Ω (M)=P Ω (M′),

[0064] M∈range(X),

[0065] Here, range(·) represents the range space.

[0066] When conducting experimental measurements of drug sensitivity data, due to the existence of systematic and random errors, the measured data results will inevitably have a certain degree of deviation. Therefore, the MCLRP model should be able to allow for a certain deviation from the known values. The model for predicting the sensitivity matrix of anticancer drugs is as follows:

[0067] min M rank(M)st||P Ω (MM′)|| F ≤ε, M∈range(X) (1),

[0068] Where ε is the allowable noise range between the final matrix M and the initial matrix M′ in the known values, and is P Ω The projection operator over Ω. Since the rank(.) function is a non-convex and discontinuous function, we use ||.|| * To approximate the rank(·) function,

[0069] rank(·) is the rank function, Ω is the set of observations on matrix M′, and F is the calculation parameter;

[0070] Since the rank function is a non-convex and discontinuous function, a kernel function is used to approximate it. The kernel function is expressed as:

[0071] min M ||M|| * st||P Ω (MM′)|| F ≤ε,M∈range(X) (2),

[0072] Among them, ||M|| * Let M represent the nuclear norm of M, which is the sum of all singular values ​​of M, and M represent the matrix to be recovered, i.e., the prediction matrix. The model is transformed into a convex optimization problem. Candes et al. proved that, under certain conditions, the solutions obtained by both methods are equivalent. Since the cell line-gene expression matrix X has a relatively high dimensionality and significant information redundancy, principal component analysis is used to reduce the dimensionality of X, which can be expressed as:

[0073]

[0074] Where U = PCA(X), PCA(·) refers to principal component analysis of gene expression profile X, β is the equilibrium parameter, and Λ is a diagonal matrix.

[0075] Since the noise level is uncertain, choosing appropriate parameters for this model is challenging. Furthermore, proposing an efficient solver for a constrained model is not straightforward. Therefore, to simplify the solution, model equation (3) is relaxed to a regularized model. The introduction of the regularization term not only allows for tolerance to unknown noise but also provides computational convenience. Therefore, the unconstrained optimization problem is as follows:

[0076]

[0077] Where G = {M:||P} Ω (MM′)|| F ≤ε, to facilitate the solution, the optimization function (4) is substituted with variables, and is expressed as:

[0078]

[0079] Therefore, applying the augmented Lagrange to formula (5) yields:

[0080]

[0081] Where Γ is the Lagrange multiplier, ρ>0 is the penalty parameter, and δ and G are computation parameters. Next, the optimization formula (6) is solved by using the alternating multiplication method (ADMM), and at the k-th iteration.

[0082]

[0083] in, K represents the number of iterations, ρ is the penalty parameter, and Y is the calculation parameter. To solve formula (7), first obtain the partial derivative of the objective function, and then obtain 0∈U. T (UΛ-M k ),U T UΛ∈U T M k Finally, we can obtain:

[0084] Λ k+1 =(U T U) -1 U T M k (11)

[0085] Solve for formula (8), find the partial derivatives of the objective function, and then begin calculating formula (8).

[0086]

[0087]

[0088] Finally, we can get

[0089]

[0090] Among them prox τ (·) is the singular value reduction operator. Equation (9) is solved using the projected gradient method. First, let let and The following formula can then be obtained from the projection gradient method.

[0091]

[0092] in

[0093] Based on the above process, the construction and solution of the MCLRP model are completed.

[0094] Specific implementation process and conclusions based on two databases, CCLE and GDSC:

[0095] First, data acquisition and preprocessing were performed. The CCLE dataset is publicly available at https: / / www.broadinstitute.org / ccle; the GDSC dataset is publicly available at https: / / www.cancerrxgene.org. For CCLE, there are drug response profiles (measured by active regions, where the active region is the cell line's response value to the drug, meaning a higher active region indicates a more sensitive cell line) and expression profiles of 18,900 genes for 24 anticancer drugs from 491 cancer cell lines. For GDSC, there are drug response profiles for 119 anticancer drugs from 655 cancer cell lines, including the PI3K and ERK pathways, and expression profiles of 12,072 available genes were selected. Specifically, for GDSC, the cell line-drug response information can be represented as a 491×24 matrix, where rows represent cell lines and columns represent drugs, with 424 (3.6%) missing values. Detailed information about the datasets is shown in Table 1. Missing values ​​in the drug response matrix were replaced with 0, and each column was standardized; the gene expression profile matrix was reduced in dimensionality using PCA.

[0096] Table 1 provides detailed information about the dataset.

[0097]

[0098] Next, the responses and expression matrices from the two datasets were input into the MCLRP model and solved. Following the process described in Part I, the predicted response matrices were obtained. Furthermore, the response and expression matrices were input into the MC model, RR model, MCRR model, and SRMF model to obtain prediction matrices based on these four models, which were then used for subsequent model comparison and evaluation.

[0099] Next, we will discuss model evaluation and result validation. This section mainly introduces three aspects: a comparison of the predictive performance of five models in CCLE and GDSC, a new drug-cancer association prediction analysis based on the GDSC dataset, and a reliability analysis of the prediction results based on consistency identification.

[0100] In comparing the predictive performance of the five models in CCLE and GDSC, Figure 2A For the PCC comparison of the five models based on CCLE data, using PCC as the evaluation index, it can be concluded that the PCC obtained by the model in this implementation is higher than 0.6 for 21 drugs, higher than 0.7 for 12 drugs, and higher than 0.8 for 3 drugs. 75% of the 24 drugs are better than the MCRR model, and 75% of the 24 drugs are better than the SRMF model. In particular, PD-0325901 has a PCC of 0.895. Figure 2B The SCC of the five models based on CCLE data is compared. Using SCC as the evaluation index, it can be concluded that the SCC of 17 drugs obtained by the model in this embodiment is higher than 0.6, 7 drugs are higher than 0.7, and 3 drugs are higher than 0.8. Among the 24 drugs, 83.3% are better than the MCRR model, and 75% of the 24 drugs are better than the SRMF model. In particular, AZD6244 has an SCC of 0.864. To further evaluate the performance of MCLRP, scatter plots of predicted and observed values ​​for four drugs are drawn, as shown below. Figures 2G to 2JAs shown. The results indicate that a good PCC is reasonable due to the small number of outliers. For the other four datasets in GDSC, the model in this embodiment also performs well in PCC and SCC. For PCC, the model in this embodiment outperforms the MCRR model in 57% of drugs in PI3KAUC, 69% of drugs in PI3KIC50, 73% of drugs in ERKAUC, and 69% of drugs in ERKIC50. The model in this embodiment outperforms the SRMF model in 100% of drugs in PI3KAUC, 100% of drugs in PI3KIC50, 97% of drugs in ERKAUC, and 100% of drugs in ERKIC50. For SCC, the model in this embodiment outperforms the MCRR model in terms of 68% of drugs in PI3KAUC, 69% of drugs in PI3KIC50, 70% of drugs in ERKAUC, and 66% of drugs in ERKIC50; the model in this embodiment outperforms the SRMF model in terms of 93% of drugs in PI3KAUC, 93% of drugs in PI3KIC50, 100% of drugs in ERKAUC, and 100% of drugs in ERKIC50. Figures 2A to 2F In the chart, each group of bars, from left to right, represents MC, RR, MCRR, SRMF, and MCLRP.

[0101] Based on four response matrices (PI3KAUC, PI3KIC50, ERKAUC, and ERKIC50), the number of cell lines sensitive to each drug was selected using different thresholds of observed values, and the results were visualized. Figures 5A to 5D ,in Figures 5A to 5DThese represent four matrices based on PI3KAUC, PI3KIC50, ERKAUC, and ERKIC50, with each column representing a different threshold: 35%, 40%, and 45%, respectively. For the PI3KAUC dataset, when filtering predicted values ​​at a threshold of 35%, Imatinib has the largest proportion among anticancer drugs; when the threshold is 40%, both Temsirolimus and Imatinib have the largest proportions; and when the threshold is 45%, Rapamycin has the largest proportion among anticancer drugs. For the PI3KIC50 dataset, when the split threshold is 35%, the top three drugs with the highest percentage of drugs sensitive to cancer cell lines are Imatinib, PF-02341066, and Lapatinib; when the split threshold is 40%, the drugs are Imatinib, PHA-665752, and PF-02341066; and when the split threshold is 45%, the drugs are Imatinib, PF-02341066, and Lapatinib. For the ERKAUC dataset, when filtering predicted values ​​with thresholds of 35%, 40%, and 45%, Imatinib consistently has the highest percentage among anticancer drugs. For the ERKIC50 dataset, when the split threshold is 35%, the top three drugs with the highest percentage of drugs sensitive to cancer cell lines are Imatinib, PF-02341066, and Lapatinib; when the split threshold is 40%, the drugs are Imatinib, PHA-665752, and PF-02341066; and when the split threshold is 45%, the drugs are Imatinib, PHA-665752, and PF-02341066. Figure 5A The legend is applicable Figures 5A to 5D .

[0102] Figures 5E to 5FTo screen for drug-sensitive cancer cell lines in the PI3K and ERK pathways, using AUC as a metric, a threshold of 40% of observed values ​​was established. To further validate the predictions, this section presents a case study of the Imatinib-sensitive M14 cancer cell line. The M14 cell line in the dataset represents a type of melanoma. Sims JT et al. conducted biological experiments on the M14 cancer cell line using the c-Abl / Arg inhibitor Imatinib. Their results showed that Imatinib induces G2 / M phase arrest and promotes apoptosis in M14 cancer cells expressing highly active c-Abl and Arg. Furthermore, Imatinib prevents intrinsic resistance by promoting doxorubicin-mediated NF-κB / p65 nuclear localization and inhibiting the NF-κB target in a STAT3-dependent manner, while also preventing the activation of novel STAT3 / HSP27 / p38 / Akt survival pathways. Additionally, Imatinib prevents acquired resistance by inhibiting the upregulation of the ABC drug transporter ABCB1, directly suppressing ABCB1 function and eliminating survival signaling.

[0103] Regarding the reliability analysis of prediction results based on consistency identification, to further test the reliability of the model in identifying missing values, an evaluation was conducted based on the classification of drug-sensitive or drug-resistant cell lines. If the missing data exhibited the same or consistent distribution pattern as the existing data, the estimation of the missing data was considered reliable. Based on this, we first focused on the predicted and observed responses of the BRAF gene in the ERKAUC dataset to three MEK inhibitors: PD-0325901, RDEA119, and SB590885. BRAF is a serine / threonine kinase that plays a crucial role in regulating the mitogen-activated protein kinase (MAPK) cascade. Under physiological conditions, MAPK regulates gene expression involved in cellular function. BRAF gene alterations have been found in a large proportion of melanoma, thyroid cancer, and histiocytic tumors, as well as a small proportion of lung and colorectal cancers. We used the BRAF mutation status of cell lines to study the relationship between gene mutations and responses to the three drugs. The predicted responses to the three drugs were divided into BRAF mutant and BRAF wild-type groups, and the observed drug responses were further grouped. The three sets of results are illustrated, demonstrating the consistency between the predicted sensitivity of BRAF-mutant cell lines to the three drugs and existing datasets. Furthermore, we can clearly see that BRAF-mutant cell lines are significantly more sensitive to MEK inhibitors. Figures 3A to 3H As shown in the figure. These predictions are consistent with those of cell lines with response data and with previously published studies. The above results indicate that our results are reliable.

[0104] Finally, genes associated with drug sensitivity or resistance were generated and functional enrichment analysis was performed. Based on the MCLRP model, the top 1000 genes (supplementary) for each drug in the five datasets were extracted according to importance using gene expression profiling. Notably, many of the selected genes are well-documented biomarkers of drug sensitivity. For example, erlotinib prevents HER1 / EGFR phosphorylation and related downstream signaling events by inhibiting tyrosine kinase activity, and blocks tumorigenesis mediated by inappropriate HER1 / EGFR signaling; lapatinib effectively inhibits ERBB2 tyrosine phosphorylation, causing growth arrest or apoptosis in ERBB2-dependent tumor cell lines; the anticancer compound ZD6474 indirectly inhibits two key pathways in tumor growth by inhibiting VEGF-dependent tumor angiogenesis and VEGF-mediated endothelial cell survival, and directly inhibits EGFR-mediated tumor proliferation and survival. Notably, in the CCLE dataset, KIT appeared in the top genes of 12 drugs, indicating that KIT expression is a therapeutic option for many anticancer drugs. In addition, functional enrichment analysis of the top 1000 genes was performed using DAVID, and Sankey diagrams were plotted for the enriched pathways containing at least two of the 28 drugs, resulting in a visualization of the GO enrichment analysis, as shown below. Figure 4A Visualization of KEGG enrichment analysis, such as Figure 4B , Figure 4A , Figure 4B , Figure 4C The horizontal and vertical axes represent gene names, pathway names, tissue names, drug names, etc., and their English translations are the common terms used in this field.

[0105] Based on the above analysis, this embodiment proposes an MCLRP model based on known data and gene expression profiles to predict drug sensitivity matrices. In the case study, based on drug response profiles and gene expression profiles from the CCLE and GDSC databases, the MCLRP model is used to predict the drug response matrix, yielding model evaluation indicators. To further evaluate the model, the model disclosed in this embodiment is compared with other models, demonstrating that the prediction of the model disclosed in this embodiment is superior to previous methods. Case studies further prove the reliability of the model's predictions. Furthermore, genes related to drug sensitivity or resistance are generated. Notably, many of the selected genes in this embodiment are validated drug sensitivity biomarkers.

[0106] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.

Claims

1. A method for predicting response to anticancer drugs, characterized in that, include: Obtain drug response profiles and gene expression profiles of anticancer drugs to obtain response matrices and expression matrices; The response matrix and expression matrix are preprocessed to obtain the standardized matrix and the processed expression matrix; An MCLRP model is constructed based on the standardized matrix and the processed expression matrix. The ADMM algorithm is then used to solve the MCLRP model to obtain the prediction matrix. Prediction matrices are used to predict responses to anticancer drugs. The MCLRP model is constructed based on the standardized matrix and the processed expression matrix. The ADMM algorithm is used to solve the MCLRP model to obtain the prediction matrix, including: The construction process assumes that the reaction matrix can be linearly expressed through the columns of the expression matrix. In constructing the MCLRP model, the problem of predicting anticancer drug response is transformed into a minimum rank problem of a matrix within the allowable noise range. Since the rank function of the minimum rank of a matrix is ​​a non-convex function, it is replaced by a kernel function, thus transforming the minimum rank problem into a convex optimization problem. The MCLRP model is relaxed into a regularized model, thereby performing an equivalent transformation on the objective function of the MCLRP model; The relaxation of the MCLRP model into a regularized model is described above. The unconstrained optimization problem with the introduction of regularization terms is as follows: , in The unconstrained optimization problem function, after variable substitution, is expressed as: Applying the augmented Lagrange method to the function after variable substitution yields: , Where Γ is the Lagrange multiplier. >0 is the penalty parameter. and To calculate the parameters, U = PCA(X), PCA( ) refers to principal component analysis of gene expression profile X, where β is the equilibrium parameter. It is a diagonal array; Let X represent the gene expression matrix, where m is the number of cell lines and n is the number of genes. In the MCLRP model construction process, it is assumed that the response matrix is ​​linearly expressed through the columns of X.

2. The method for predicting anticancer drug response according to claim 1, characterized in that, After obtaining the prediction matrix by solving the MCLRP using the ADMM algorithm, the following steps are also included: The prediction matrix is ​​evaluated, including: calculating the Pearson correlation coefficient for each drug to obtain the Pearson correlation coefficient value, and calculating the Spearman correlation coefficient for each drug to obtain the Spearman correlation coefficient value. The Pearson correlation coefficient value represents the degree of correlation between the predicted and observed drug response curves, and the Spearman correlation coefficient value represents the correlation between the predicted and observed values. The larger the Pearson correlation coefficient value and / or the Spearman correlation coefficient value, the better the prediction effect. It also includes new drug-cancer association prediction analysis based on the GDSC dataset, i.e., validating the model at the cellular level; It also includes reliability analysis of prediction results based on consistency identification, that is, validating the model at the gene level.

3. The method for predicting anticancer drug response according to claim 1, characterized in that, The response matrix and expression matrix are preprocessed to obtain a standardized matrix and a processed expression matrix, including: Replace all missing values ​​in the response matrix with 0, and then standardize each column of the response matrix to obtain the standardized matrix, represented as follows. Where m is the number of cell lines and p is the amount of drug; Principal component analysis was used to reduce the dimensionality of the expression matrix. It is the matrix after the reaction matrix has been standardized.

4. The method for predicting anticancer drug response according to claim 1, characterized in that, The equivalent transformation of the objective function of the MCLRP model includes: Variable substitution or augmenting Lagrange.

5. The method for predicting anticancer drug response according to claim 4, characterized in that, The transformation of the anticancer drug response prediction problem into a minimum-rank matrix problem within an allowable noise range includes: The problem of estimating and optimizing the drug response matrix is ​​as follows: , in, Given the values ​​of matrix M and matrix... The permissible noise range between It is the projection operator on Ω, where rank(·) is the rank function and Ω is the matrix. The set of observations on ′ For calculation parameters; Represents the range space; Since the rank function is a non-convex and discontinuous function, a kernel function is used to approximate it. The kernel function is expressed as: , Where M represents the matrix to be recovered, i.e., the prediction matrix.

6. The method for predicting anticancer drug response according to claim 5, characterized in that, The dimension reduction representation of the expression matrix using principal component analysis is as follows: 。 7. The method for predicting anticancer drug response according to claim 1, characterized in that, The MCLRP model is constructed based on the standardized matrix and the processed expression matrix. The ADMM algorithm is used to solve the MCLRP model to obtain the prediction matrix, including: in, , Indicates the number of iterations. For penalty parameters, These are the parameters for calculation.

8. A device for predicting response to anticancer drugs, characterized in that, include: The acquisition module is used to acquire drug response profiles and gene expression profiles of anticancer drugs, thereby obtaining response matrices and expression matrices. The preprocessing module is used to preprocess the response matrix and expression matrix to obtain the standardized matrix and the processed expression matrix. The model building module is used to construct an MCLRP model based on the standardized matrix and the processed expression matrix, and to solve the MCLRP model using the ADMM algorithm to obtain the prediction matrix; The prediction module is used to predict anticancer drug responses using a prediction matrix. The MCLRP model is constructed based on the standardized matrix and the processed expression matrix. The ADMM algorithm is used to solve the MCLRP model to obtain the prediction matrix, including: The construction process assumes that the reaction matrix can be linearly expressed through the columns of the expression matrix. In constructing the MCLRP model, the problem of predicting anticancer drug response is transformed into a minimum rank problem of a matrix within the allowable noise range. Since the rank function of the minimum rank of a matrix is ​​a non-convex function, it is replaced by a kernel function, thus transforming the minimum rank problem into a convex optimization problem. The MCLRP model is relaxed into a regularized model, thereby performing an equivalent transformation on the objective function of the MCLRP model; The relaxation of the MCLRP model into a regularized model is described above. The unconstrained optimization problem with the introduction of regularization terms is as follows: , in The unconstrained optimization problem function, after variable substitution, is expressed as: Applying the augmented Lagrange method to the function after variable substitution yields: , Where Γ is the Lagrange multiplier. >0 is the penalty parameter. and These are the parameters for calculation.