Anti-cancer drug response prediction method and device, computer equipment and storage medium

By extracting features and reducing dimensionality from genomic data of tumor samples, and combining this with a prediction network, the accuracy of predicting anticancer drug response has been improved, solving the problem of inaccurate predictions in existing models and promoting the implementation of personalized medicine.

CN118398086BActive Publication Date: 2026-01-06SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410483313.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-19
Publication Date
2026-01-06
Estimated Expiration
2044-04-19

AI Technical Summary

Technical Problem

Existing anticancer drug sensitivity prediction models have low accuracy in predicting anticancer drug responses, which increases the difficulty of implementing personalized medicine.

Method used

By acquiring genomic data of tumor samples to be predicted, feature extraction and dimensionality reduction of tumor cell line gene mutation data are performed using a gene expression encoder of the target tumor cell line. Combined with a target anticancer drug response prediction network, low-dimensional characterization of gene expression and mutation data is performed to predict the sensitivity values ​​of tumor samples to different drugs.

Benefits of technology

It improves the accuracy of anticancer drug response prediction models, makes full use of genomic data of tumor cell lines, maintains as much genetic feature information as possible, and promotes the realization of personalized medicine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118398086B_ABST
    Figure CN118398086B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of drug response prediction, and discloses an anticancer drug response prediction method and device, computer equipment and a storage medium, the method comprising: obtaining genomics data of a tumor sample to be predicted, the genomics data comprising gene expression data and gene mutation data; inputting the genomics data of the tumor sample to be predicted into a target anticancer drug response prediction model; using a target tumor cell line gene expression encoder to extract features from the gene expression data, to obtain a low-dimensional representation of the gene expression data; using a tumor cell line gene mutation data dimension reduction processing module to perform dimension reduction processing on the gene mutation data, to obtain a low-dimensional representation of the gene mutation data; inputting the low-dimensional representation of the gene expression data and the low-dimensional representation of the gene mutation data into a target anticancer drug response prediction network, to obtain a predicted value of the sensitivity numerical value of the tumor sample to be predicted to different drug responses, and the present application improves the prediction accuracy of anticancer drug responses.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of drug response prediction technology, specifically to methods, devices, computer equipment, and storage media for predicting anticancer drug responses. Background Technology

[0002] The medical treatment of malignant tumors is gradually developing towards personalization, precision, and efficiency. The varying clinical responses of cancer patients to the same drug offer promising prospects for personalized medicine. However, implementing precise treatment for all patients is not easy. Therefore, establishing suitable anticancer drug sensitivity prediction models is particularly important in overcoming the challenges of high costs and implementation difficulties in precision medicine.

[0003] However, due to insufficient use of gene data mining in cancer cell lines, the predictive accuracy of anticancer drug sensitivity prediction models for anticancer drug response is not high. Summary of the Invention

[0004] In view of this, the present invention provides a method, apparatus, computer equipment and storage medium for predicting anticancer drug response, in order to solve the problem of low accuracy in predicting anticancer drug response by anticancer drug sensitivity prediction models.

[0005] In a first aspect, the present invention provides a method for predicting anticancer drug response, the method comprising:

[0006] Obtain genomic data of the tumor sample to be predicted, wherein the genomic data includes gene expression data and gene mutation data;

[0007] The genomic data of the tumor sample to be predicted is input into the target anticancer drug response prediction model, which includes a target tumor cell line gene expression encoder, a tumor cell line gene mutation data dimensionality reduction processing module, and a target anticancer drug response prediction network.

[0008] The gene expression encoder of the target tumor cell line is used to extract features from the gene expression data of the tumor sample to be predicted, so as to obtain a low-dimensional representation of the gene expression data of the tumor sample to be predicted.

[0009] The gene mutation data of the tumor cell line is reduced in dimensionality using the tumor cell line gene mutation data dimensionality reduction module to obtain a low-dimensional representation of the gene mutation data of the tumor sample to be predicted.

[0010] The low-dimensional representations of gene expression data and gene mutation data of the tumor sample to be predicted are input into the target anticancer drug response prediction network to obtain the predicted values ​​of the sensitivity values ​​of the tumor sample to different drug responses.

[0011] The anticancer drug response prediction method provided in this embodiment extracts features from the gene expression data of the tumor sample to be predicted using a gene expression encoder of the target tumor cell line, obtaining a low-dimensional representation of the gene expression data of the tumor sample to be predicted. Then, it uses a tumor cell line gene mutation data dimensionality reduction module to reduce the dimensionality of the gene mutation data of the tumor sample to be predicted, obtaining a low-dimensional representation of the gene mutation data of the tumor sample to be predicted. The low-dimensional representations of the gene expression data and gene mutation data of the tumor sample to be predicted are input into the target anticancer drug response prediction network to obtain the predicted values ​​of the sensitivity values ​​of the tumor sample to different drug responses. This method makes full use of the tumor's genomic data, promotes the preservation of as much gene feature information as possible, and improves the prediction accuracy of the anticancer drug response prediction model.

[0012] In one optional implementation, the target anticancer drug response prediction model is obtained through the following steps:

[0013] Construct training samples, which include genomic data of different tumor samples and sensitivity values ​​of different tumor samples to different drugs;

[0014] A model for predicting anticancer drug response is constructed, comprising: a tumor cell line gene expression encoder, a tumor cell line gene mutation data dimensionality reduction module, and an anticancer drug response prediction network. The tumor cell line gene expression encoder is used to extract features from the gene expression data of tumor samples to obtain a low-dimensional representation of the gene expression data. The tumor cell line gene mutation data dimensionality reduction module is used to perform dimensionality reduction processing on the gene mutation data of tumor samples using a multidimensional scaling algorithm to obtain a low-dimensional representation of the gene mutation data. The anticancer drug response prediction network is used to predict the sensitivity values ​​of the tumor samples to different drug responses based on the low-dimensional representations of the gene expression data and gene mutation data of the tumor samples, thereby obtaining predicted values ​​of the sensitivity values ​​of the tumor samples to different drug responses.

[0015] The tumor cell line gene expression encoder was self-supervised learning of tumor cell line gene expression data using training samples to obtain the target tumor cell line gene expression encoder.

[0016] The anticancer drug response prediction model, which includes a target tumor cell line gene expression encoder and a tumor cell line gene mutation data dimensionality reduction module, was trained using training samples to optimize the parameters of the anticancer drug response prediction network, determine the target anticancer drug response prediction network, and obtain the target anticancer drug response prediction model.

[0017] The anticancer drug response prediction method provided in this embodiment is based on a target anticancer drug response prediction model trained using low-dimensional representations of tumor sample genomic data and numerical training on the sensitivity of tumor samples to different drug responses. The target tumor cell line gene expression encoder is obtained through self-supervised learning of tumor cell line gene expression data. By introducing a self-supervised approach based on tumor cell line gene expression data and a multi-dimensional scaling algorithm that maintains the distance between samples in the original gene mutation data space in a low-dimensional space, the method fully utilizes tumor cell line genomic data, promotes the preservation of as much gene feature information as possible, and improves the prediction accuracy of the target anticancer drug response prediction model.

[0018] In one optional implementation, the step of training the anticancer drug response prediction model, which includes a target tumor cell line gene expression encoder and a tumor cell line gene mutation data dimensionality reduction module, using training samples to optimize the parameters of the anticancer drug response prediction network, determine the target anticancer drug response prediction network, and obtain the target anticancer drug response prediction model includes:

[0019] Using genomic data of different tumor samples as input data for the anticancer drug response prediction model, and the sensitivity values ​​of different tumor samples to different drug responses as labels, the anticancer drug response prediction model, which includes a target tumor cell line gene expression encoder and a tumor cell line gene mutation data dimensionality reduction processing module, is trained to optimize the parameters of the anticancer drug response prediction network, determine the target anticancer drug response prediction network, and obtain the target anticancer drug response prediction model.

[0020] The anticancer drug response prediction method provided in this embodiment uses genomic data from different tumor samples as input data and sensitivity values ​​of different tumor samples to different drugs as labels. Through training, the model can learn the correlation between gene expression and gene mutation data and drug response, thereby predicting the response of unknown tumor samples to anticancer drugs. By fully utilizing tumor cell line genomic data, it helps to retain as much genetic feature information as possible, improving the accuracy of the target anticancer drug response prediction model.

[0021] In one optional implementation, constructing the training samples includes:

[0022] Gene expression and mutation data of different tumor samples were obtained from the Encyclopedia of Cancer Cell Lines;

[0023] The half-inhibition concentrations of different tumor samples in the Cancer Cell Line Encyclopedia were obtained from the Cancer Drug Sensitivity Genomics Project for responses to different drugs, wherein the half-inhibition concentrations were used to characterize sensitivity values.

[0024] Gene expression and mutation data of different tumor samples were obtained from the cancer genome atlas.

[0025] The training sample was constructed based on gene expression and mutation data of different tumor samples obtained from the Encyclopedia of Cancer Cell Lines, half-inhibitory concentrations of different tumor samples to different drugs obtained from the Cancer Drug Sensitivity Genomics Project, and gene expression and mutation data of different tumor samples obtained from the Cancer Genome Atlas.

[0026] The anticancer drug response prediction method provided in this embodiment makes full use of tumor cell line genomics data, promotes the preservation of as much genetic feature information as possible, and improves the prediction accuracy of the target anticancer drug response prediction model.

[0027] In one optional implementation, the step of using training samples to perform self-supervised learning of tumor cell line gene expression encoders on tumor cell line gene expression data to obtain target tumor cell line gene expression encoders includes:

[0028] We used gene expression and mutation data from different tumor samples obtained from the Cancer Genome Atlas to perform self-supervised learning on the tumor cell line gene expression encoder, in order to obtain the target tumor cell line gene expression encoder.

[0029] The anticancer drug response prediction method provided in this embodiment uses a target tumor cell line gene expression encoder obtained through self-supervised learning of tumor cell line gene expression data. By introducing a self-supervised approach based on tumor cell line gene expression data, the method fully utilizes tumor cell line genomics data, promotes the retention of as much gene feature information as possible, and improves the accuracy of the target anticancer drug response prediction model in predicting anticancer drug response.

[0030] In one optional implementation, the step of using genomic data of different tumor samples as input data for the anticancer drug response prediction model, and using the sensitivity values ​​of different tumor samples to different drug responses as labels to train the anticancer drug response prediction model, which includes a target tumor cell line gene expression encoder and a tumor cell line gene mutation data dimensionality reduction processing module, includes:

[0031] The anticancer drug response prediction model was trained using genomic data of different tumor samples obtained from the Encyclopedia of Cancer Cell Lines as input data and half-inhibition concentrations of different tumor samples for different drug responses as labels. The model included a target tumor cell line gene expression encoder and a tumor cell line gene mutation data dimensionality reduction module.

[0032] The anticancer drug response prediction method provided in this embodiment is based on a target anticancer drug response prediction model trained by low-dimensional representation of tumor sample genomic data and numerical sensitivity of tumor samples to different drug responses. This fully utilizes tumor cell line genomic data, promotes the retention of as much genetic feature information as possible, and improves the prediction accuracy of the target anticancer drug response prediction model.

[0033] In an optional implementation, before constructing the training samples based on gene expression and mutation data of different tumor samples obtained from the Encyclopedia of Cancer Cell Lines, half-inhibitory concentrations of different tumor samples to different drugs obtained from the Cancer Drug Sensitivity Genomics Project, and gene expression and mutation data of different tumor samples obtained from the Cancer Genome Atlas, the method further includes:

[0034] Gene expression and gene mutation data from different tumor samples obtained from the Encyclopedia of Cancer Cell Lines were standardized and missing values ​​were filled.

[0035] The half-inhibition concentrations (WICs) of different tumor samples from the Cancer Cell Line Encyclopedia obtained from the Cancer Drug Sensitivity Genomics Project were standardized and missing values ​​were filled.

[0036] Gene expression and mutation data from different tumor samples obtained from the Cancer Genome Atlas were standardized and missing values ​​were filled.

[0037] The anticancer drug response prediction method provided in this embodiment ensures the integrity, reliability, and accuracy of the training samples by standardizing and filling in missing values, thereby improving the performance and efficiency of data analysis and modeling.

[0038] In a second aspect, the present invention provides an anticancer drug response prediction device, the device comprising:

[0039] The genomics data acquisition module is used to acquire genomics data of the tumor sample to be predicted, wherein the genomics data includes gene expression data and gene mutation data;

[0040] The genomics data input module is used to input the genomics data of the tumor sample to be predicted into the target anticancer drug response prediction model. The target anticancer drug response prediction model includes a target tumor cell line gene expression encoder, a tumor cell line gene mutation data dimensionality reduction processing module, and a target anticancer drug response prediction network.

[0041] The low-dimensional characterization acquisition module for gene expression data is used to extract features from the gene expression data of the tumor sample to be predicted using the gene expression encoder of the target tumor cell line, and obtain a low-dimensional characterization of the gene expression data of the tumor sample to be predicted.

[0042] The low-dimensional representation acquisition module for gene mutation data is used to perform dimensionality reduction processing on the gene mutation data of the tumor cell line gene mutation data to be predicted using the tumor cell line gene mutation data dimensionality reduction processing module, so as to obtain a low-dimensional representation of the gene mutation data of the tumor sample to be predicted.

[0043] The anticancer drug response prediction value acquisition module is used to input the low-dimensional representation of the gene expression data and the low-dimensional representation of the gene mutation data of the tumor sample to be predicted into the target anticancer drug response prediction network to obtain the predicted value of the sensitivity value of the tumor sample to different drug responses.

[0044] Thirdly, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the anticancer drug response prediction method of the first aspect or any corresponding embodiment described above.

[0045] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the anticancer drug response prediction method of the first aspect or any corresponding embodiment thereof.

[0046] Fifthly, the present invention provides a computer program product, including computer instructions for causing a computer to execute the anticancer drug response prediction method of the first aspect or any corresponding embodiment described above. Attached Figure Description

[0047] To more clearly illustrate the technical solutions in the specific embodiments or related technologies of the present invention, the drawings used in the description of the specific embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0048] Figure 1 This is a flowchart illustrating a method for predicting anticancer drug response according to an embodiment of the present invention;

[0049] Figure 2 This is a flowchart illustrating another method for predicting anticancer drug response according to an embodiment of the present invention;

[0050] Figure 3 This is a schematic diagram of a target anticancer drug response prediction model according to an embodiment of the present invention;

[0051] Figure 4 This is a structural block diagram of an anticancer drug response prediction device according to an embodiment of the present invention;

[0052] Figure 5 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0054] The same anticancer drug can have different therapeutic effects on different patients with the same type of cancer. Pharmacogenomics aims to study how genomic alterations and transcriptomic processes determine drug response based on a patient's genotype, develop rational drug therapies, and optimize treatment to ensure maximum efficacy and minimal side effects. Due to tumor heterogeneity and intratumoral subclones, accurately predicting drug response and identifying new anticancer drugs remains a challenging task. With the advent of high-throughput technologies and the rapid development of high-throughput screening techniques, a wealth of information (transcriptional, metabolic, gene expression) can be measured at a reasonable cost. Several online public phenotypic screening datasets have been established to support research on anticancer drug response, such as data from The Cancer Genome Atlas (TCGA), Genomics of Drug Sensitivity in Cancer (GDSC), and the Cancer CellLine Encyclopedia (CCLE) project, which has performed molecular profiling analyses of cancer cell lines. The establishment of these genomic databases for cancer bioinformatics and systems biology provides a wealth of cell line sensitive data and cell line omics characteristics. This rich genomic data offers opportunities for systematic identification of biomarkers and prediction of tumor drug sensitivity.

[0055] With the emergence of the concept of personalized medicine and the continuous development of big data and modern molecular biology technologies, the medical treatment of malignant tumors is gradually moving towards personalization, precision, and efficiency. How to predict and analyze treatment side effects based on genomic characteristics and drug response patterns has become a major challenge for precision medicine.

[0056] Alterations in the cancer genome significantly influence clinical responses to treatment and, in many cases, serve as effective biomarkers for drug response. The varying clinical responses of cancer patients to the same drug offer promise for personalized medicine; however, providing precision treatment to all patients requires substantial financial and material resources, making it difficult to implement. Therefore, establishing suitable predictive models for anticancer drug sensitivity is crucial in overcoming the challenges of high costs and implementation difficulties in precision medicine.

[0057] In related technologies, the insufficient use of gene data mining in cancer cell lines has resulted in low accuracy in predicting anticancer drug response by anticancer drug sensitivity prediction models.

[0058] This invention provides a method for predicting anticancer drug response. It extracts features from gene expression data of tumor samples using a target tumor cell line gene expression encoder to obtain a low-dimensional representation of the gene expression data. Then, it uses a tumor cell line gene mutation data dimensionality reduction module to reduce the dimensionality of the gene mutation data of the tumor samples, obtaining a low-dimensional representation of the gene mutation data. These low-dimensional representations of gene expression data and gene mutation data are input into a target anticancer drug response prediction network to obtain predicted values ​​of the tumor sample's sensitivity to different drug responses. This improves the accuracy of the target anticancer drug response prediction model in predicting anticancer drug responses.

[0059] According to an embodiment of the present invention, an embodiment of a method for predicting anticancer drug response is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0060] This embodiment provides a method for predicting anticancer drug response, which can be used in mobile terminals, such as central processing units, servers, etc. Figure 1 This is a flowchart of an anticancer drug response prediction method according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps:

[0061] Step S101: Obtain the genomic data of the tumor sample to be predicted. The genomic data includes gene expression data and gene mutation data.

[0062] The genomic data of the tumor sample to be predicted can be the genomic data of the tumor cell line to be predicted. This genomic data can include gene expression data and gene mutation data.

[0063] Gene expression data reflects the activity levels of genes in cells or tissues. This data is typically obtained using high-throughput sequencing technologies (such as RNA sequencing) to measure the expression levels of different genes under various conditions. Gene mutation data describes the variations in the genome. In this embodiment, gene mutation data considers four types of non-synonymous mutations: missense mutations, nonsense mutations, frameshift insertions, and deletions. These variations may be associated with the occurrence and development of diseases; therefore, gene mutation data is crucial for studying the genetic basis of diseases and inter-individual genetic differences.

[0064] It should be noted that gene copy number, gene mutation data, and gene expression data are interconnected. An increase in gene copy number typically leads to an increase in gene expression levels, meaning more gene copies in the cell produce more corresponding proteins. This is because an increase in gene copy number provides more opportunities for transcription and translation, thereby increasing protein synthesis. In this embodiment, the gene expression data package contains changes in gene copy number.

[0065] Gene copy number variation refers to the change in the number of times a gene is copied repeatedly within the genome. Some gene copy number variations may be caused by gene mutations, such as gene duplication misplacement, insertion, or deletion. Such gene copy number variations can affect gene function, thereby influencing an individual's phenotype. In this embodiment, the gene mutation data package contains gene copy number variations.

[0066] Step S102: Input the genomic data of the tumor sample to be predicted into the target anticancer drug response prediction model. The target anticancer drug response prediction model includes a target tumor cell line gene expression encoder, a tumor cell line gene mutation data dimensionality reduction processing module, and a target anticancer drug response prediction network.

[0067] In this process, after obtaining the genomic data of the tumor sample to be predicted, the genomic data of the tumor sample to be predicted is input into the target anticancer drug response prediction model to obtain the predicted value of the sensitivity value of the tumor sample to different anticancer drugs.

[0068] The target anticancer drug response prediction model is an anticancer drug response prediction model that has already been trained.

[0069] Step S103: Use the gene expression encoder of the target tumor cell line to extract features from the gene expression data of the tumor sample to be predicted, and obtain a low-dimensional representation of the gene expression data of the tumor sample to be predicted.

[0070] In this process, after the genomic data of the tumor sample to be predicted is input into the target anticancer drug response prediction model, the target tumor cell line gene expression encoder in the target anticancer drug response prediction model extracts features from the gene expression data of the tumor sample to be predicted, and obtains a low-dimensional representation of the gene expression data of the tumor sample to be predicted.

[0071] The target tumor cell line gene expression encoder is a trained tumor cell line gene expression encoder.

[0072] Step S104: Use the tumor cell line gene mutation data dimensionality reduction processing module to perform dimensionality reduction processing on the gene mutation data of the tumor sample to be predicted, and obtain a low-dimensional representation of the gene mutation data of the tumor sample to be predicted.

[0073] Specifically, after the genomic data of the tumor sample to be predicted is input into the target anticancer drug response prediction model, the tumor cell line gene mutation data dimensionality reduction module in the target anticancer drug response prediction model performs dimensionality reduction processing on the gene mutation data of the tumor sample to be predicted, and obtains a low-dimensional representation of the gene mutation data of the tumor sample to be predicted.

[0074] It should be noted that the tumor cell line gene mutation data dimensionality reduction module uses the MDS (Multidimensional Scaling) algorithm to reduce the dimensionality of the gene mutation data of the tumor sample to be predicted.

[0075] Dimensionality reduction of gene mutation data for tumor samples to be predicted can also be described as feature scaling of the gene mutation data for tumor samples to be predicted. Obtaining a low-dimensional representation of the gene mutation data for tumor samples to be predicted can also be described as obtaining a compressed representation of the gene mutation data for tumor samples to be predicted.

[0076] The MDS algorithm is a statistical analysis method used to map high-dimensional data to a low-dimensional space. It aims to preserve the relative distances between data points to visualize and interpret the data in a lower-dimensional space. In this embodiment, the MDS algorithm is used to project gene mutation data into a low-dimensional space to preserve the distances between data points as much as possible, extract structural information from the gene mutation data, and retain some features of masked genes in the gene sequence. By mapping high-dimensional data to a low-dimensional space, the MDS algorithm can help understand the relationships between data, discover patterns and structures in the data, and visualize the data.

[0077] It is understood that steps S103 and S104 can be performed simultaneously, or steps S103 can be performed first and then steps S104, or steps S104 can be performed first and then steps S103. No specific restrictions are imposed here.

[0078] Step S105: Input the low-dimensional representation of gene expression data and gene mutation data of the tumor sample to be predicted into the target anticancer drug response prediction network to obtain the predicted value of the sensitivity value of the tumor sample to different drug responses.

[0079] Specifically, after obtaining low-dimensional representations of gene expression data and gene mutation data of the tumor sample to be predicted, these low-dimensional representations are input into the target anticancer drug response prediction network to obtain the predicted values ​​of the sensitivity values ​​of the tumor sample to different drug responses.

[0080] Specifically, the low-dimensional representations of gene expression data and gene mutation data of the tumor sample to be predicted are concatenated to obtain the concatenated result. An anticancer drug response prediction network is then used to predict the sensitivity values ​​of the tumor sample to different drug responses, outputting the predicted values.

[0081] The target anticancer drug response prediction network is a pre-trained anticancer drug response prediction network.

[0082] It should be noted that the target anticancer drug response prediction model outputs predicted values ​​of the sensitivity of the tumor sample to different drugs. These sensitivity values ​​are calculated using the half-inhibitory concentration (IC50). 50 Characterization was performed. The lower the predicted sensitivity value, the higher the antibody sensitivity, the stronger the specificity, and the better the tumor sample to be predicted would respond to the drug.

[0083] The anticancer drug response prediction method provided in this embodiment extracts features from the gene expression data of the tumor sample to be predicted using a gene expression encoder of the target tumor cell line, obtaining a low-dimensional representation of the gene expression data of the tumor sample to be predicted. Then, it uses a tumor cell line gene mutation data dimensionality reduction module to reduce the dimensionality of the gene mutation data of the tumor sample to be predicted, obtaining a low-dimensional representation of the gene mutation data of the tumor sample to be predicted. The low-dimensional representations of the gene expression data and gene mutation data of the tumor sample to be predicted are input into the target anticancer drug response prediction network to obtain the predicted values ​​of the sensitivity values ​​of the tumor sample to different drug responses. This method makes full use of the tumor's genomic data, promotes the preservation of as much gene feature information as possible, and improves the prediction accuracy of the anticancer drug response prediction model.

[0084] This embodiment provides a method for predicting anticancer drug response, which can be used in mobile terminals such as central processing units and servers. The method for predicting anticancer drug response includes the following steps:

[0085] Step 1: Obtain the genomic data of the tumor sample to be predicted. Genomic data includes gene expression data and gene mutation data. For details, please refer to [link to relevant documentation]. Figure 1 Step S101 of the illustrated embodiment will not be described again here.

[0086] The second step is to input the genomic data of the tumor sample to be predicted into the target anticancer drug response prediction model. The target anticancer drug response prediction model includes a target tumor cell line gene expression encoder, a tumor cell line gene mutation data dimensionality reduction processing module, and a target anticancer drug response prediction network.

[0087] The process of obtaining the target anticancer drug response prediction model in the second step includes the following steps:

[0088] Step a1: Construct training samples, which include genomic data of different tumor samples and sensitivity values ​​of different tumor samples to different drugs.

[0089] To obtain a predictive model for response to the target anticancer drug, training samples are first constructed. Specifically, step a1 above includes:

[0090] Step a11: Obtain gene expression and gene mutation data for different tumor samples from the Encyclopedia of Cancer Cell Lines.

[0091] This involved obtaining gene expression and mutation data for different cancer cell lines from the Encyclopedia of Cancer Cell Lines. It should be noted that the tumor samples can be cancer cell lines.

[0092] The CCLE (Cancer Cell Line Encyclopedia) contains molecular descriptors of over 1,000 human cancer cell lines from various tissue types and cancer subtypes, including molecular characteristics such as gene expression, protein expression, and protein phosphorylation.

[0093] The gene expression data for cancer cell lines in the Encyclopedia of Cancer Cell Lines is E. CCLE ,in, This refers to the number of transcriptions per million genes g in cancer cell line c.

[0094] The gene mutation data for cancer cell lines in the Encyclopedia of Cancer Cell Lines is M. CCLE ,in, The value indicates the gene mutation status, with 0 indicating no mutation and 1 indicating a mutation.

[0095] It should be noted that the gene expression and gene mutation data of different cancer cell lines obtained from the Encyclopedia of Cancer Cell Lines include the gene expression and gene mutation data of all cancer cell lines contained in the Encyclopedia of Cancer Cell Lines.

[0096] Step a12: Obtain the half-inhibition concentrations (WICs) of different tumor samples to different drugs from the Encyclopedia of Cancer Cell Lines project, where the WIC is used to characterize the sensitivity values.

[0097] Among them, the half-inhibitory concentrations of different cancer cell lines to different drugs were obtained from the Cancer Drug Sensitivity Genomics project.

[0098] The half-inhibition concentration is IC50. 50 IC 50 The value is in μM and is measured on a logarithmic scale. IC CCLE Let be the half-inhibitory concentration (WIC) of cancer cell line c to drug d. Then, obtain the WIC of different tumor samples from the Cancer Cell Line Encyclopedia to different drugs.

[0099] Step a13: Obtain gene expression data and gene mutation data of different tumor samples from the cancer genome atlas.

[0100] Specifically, gene expression and mutation data of pan-cancer samples in the cancer genome map are obtained from the open-source genomic data visualization and analysis platform UCSC Xena, which means obtaining gene expression and mutation data of different tumor samples in the cancer genome map.

[0101] TCGA (The Cancer Genome Atlas) includes information on various indicators such as gene expression, DNA methylation, gene mutation, protein expression, and miRNA expression. It aims to reveal the molecular mechanisms of cancer occurrence and development through comprehensive molecular characterization of various cancer samples, and to provide new targets and strategies for cancer treatment.

[0102] Among them, the gene expression data of different tumor samples in the Cancer Genome Atlas are E TCGA ,in,

[0103] Gene mutation data of different tumor samples in the Cancer Genome Atlas is M TCGA ,in,

[0104] It should be noted that the gene expression data and gene mutation data of different tumor samples obtained from the Cancer Genome Atlas are the gene expression data and gene mutation data of all tumor samples included in the Cancer Genome Atlas.

[0105] Step a14: Based on gene expression and gene mutation data of different tumor samples obtained from the Encyclopedia of Cancer Cell Lines, half-inhibitory concentrations of different tumor samples to different drugs obtained from the Encyclopedia of Cancer Cell Lines obtained from the Cancer Drug Sensitivity Genomics Project, and gene expression and gene mutation data of different tumor samples obtained from the Cancer Genome Atlas, a training sample is constructed.

[0106] Among them, after obtaining gene expression and gene mutation data of different tumor samples from the Encyclopedia of Cancer Cell Lines, half-inhibitory concentrations of different tumor samples to different drugs from the Encyclopedia of Cancer Cell Lines from the Cancer Drug Sensitivity Genomics Project, and gene expression and gene mutation data of different tumor samples from the Cancer Genome Atlas, these data were used to construct training samples.

[0107] Figure 2 This is a flowchart of an anticancer drug response prediction method according to an embodiment of the present invention. It should be noted that steps a11 to a14 above correspond to... Figure 2 The steps involved in obtaining cell line genome data.

[0108] Step a2: Construct an anticancer drug response prediction model. The anticancer drug response prediction model includes: a tumor cell line gene expression encoder, a tumor cell line gene mutation data dimensionality reduction module, and an anticancer drug response prediction network. The tumor cell line gene expression encoder is used to extract features from the gene expression data of tumor samples to obtain a low-dimensional representation of the gene expression data of tumor samples. The tumor cell line gene mutation data dimensionality reduction module is used to perform dimensionality reduction processing on the gene mutation data of tumor samples using a multidimensional scaling transformation algorithm to obtain a low-dimensional representation of the gene mutation data of tumor samples. The anticancer drug response prediction network is used to predict the sensitivity values ​​of tumor samples to different drug responses based on the low-dimensional representations of gene expression data and gene mutation data of tumor samples, so as to obtain the predicted values ​​of the sensitivity values ​​of tumor samples to different drug responses.

[0109] Step a3: Use training samples to perform self-supervised learning of tumor cell line gene expression encoders to obtain the target tumor cell line gene expression encoder.

[0110] After constructing the training samples, the tumor cell line gene expression encoder is used to perform self-supervised learning of the tumor cell line gene expression data to obtain the target tumor cell line gene expression encoder.

[0111] Step a4: Use training samples to train the anticancer drug response prediction model, which includes the target tumor cell line gene expression encoder and the tumor cell line gene mutation data dimensionality reduction module, in order to optimize the parameters of the anticancer drug response prediction network, determine the target anticancer drug response prediction network, and obtain the target anticancer drug response prediction model.

[0112] Specifically, after obtaining the gene expression encoder of the target tumor cell line, the anticancer drug response prediction model, which includes the gene expression encoder of the target tumor cell line and the dimensionality reduction processing module for tumor cell line gene mutation data, is trained using training samples to optimize the parameters of the anticancer drug response prediction network, determine the target anticancer drug response prediction network, and obtain the target anticancer drug response prediction model.

[0113] Among them, the anticancer drug response prediction model, which includes a target tumor cell line gene expression encoder and a tumor cell line gene mutation data dimensionality reduction processing module, can be an anticancer drug response prediction model with a fully connected network containing a target tumor cell line gene expression encoder and a tumor cell line gene mutation data dimensionality reduction processing module, and has an activation function.

[0114] like Figure 2 As shown, an autoencoder is used to reduce the dimensionality of high-dimensional features of gene expression. That is, after obtaining the gene expression encoder of the target tumor cell line, the gene expression encoder of the target tumor cell line is used to extract features from the gene expression data of the training sample to obtain the low-dimensional features of the gene expression data of the training sample.

[0115] The high-dimensional features of gene mutations in cell lines are compressed using the MDS compression method. Specifically, the MDS algorithm in the tumor cell line gene mutation data dimensionality reduction processing module is used to reduce the dimensionality of the gene mutation data of the training samples to obtain the low-dimensional features of the gene mutation data of the training samples.

[0116] The anticancer drug sensitivity prediction model, which includes a pre-trained expression editor, is trained using training samples. Specifically, the anticancer drug response prediction model, which includes a target tumor cell line gene expression encoder and a tumor cell line gene mutation data dimensionality reduction module, is trained using training samples to optimize the parameters of the anticancer drug response prediction network, determine the target anticancer drug response prediction network, and obtain the target anticancer drug response prediction model.

[0117] It is understandable that the target anticancer drug response prediction model is used to predict the sensitivity values ​​of the tumor sample to different drugs, that is, to predict the drug sensitivity values ​​using a trained anticancer drug sensitivity prediction model. The target anticancer drug response prediction model can be a trained anticancer drug sensitivity prediction model.

[0118] Step 3: Utilize the gene expression encoder of the target tumor cell line to extract features from the gene expression data of the tumor sample to be predicted, obtaining a low-dimensional representation of the gene expression data of the tumor sample to be predicted. For details, please refer to [link to details]. Figure 1 Step S103 of the illustrated embodiment will not be described again here.

[0119] Step 4: Utilize the tumor cell line gene mutation data dimensionality reduction module to reduce the dimensionality of the gene mutation data of the tumor sample to be predicted, obtaining a low-dimensional representation of the gene mutation data of the tumor sample to be predicted. For details, please refer to [link to relevant documentation]. Figure 1 Step S104 of the illustrated embodiment will not be described again here.

[0120] Step 5: Input the low-dimensional representations of gene expression data and gene mutation data of the tumor sample to be predicted into the target anticancer drug response prediction network to obtain the predicted sensitivity values ​​of the tumor sample to different drug responses. For details, please refer to... Figure 1 Step S105 of the illustrated embodiment will not be described again here.

[0121] The anticancer drug response prediction method provided in this embodiment is based on a target anticancer drug response prediction model trained using low-dimensional representations of tumor sample genomic data and numerical training on the sensitivity of tumor samples to different drug responses. The target tumor cell line gene expression encoder is obtained through self-supervised learning using tumor cell line gene expression data. By introducing a self-supervised approach based on tumor cell line gene expression data and a multi-dimensional scaling algorithm that maintains the distance between samples in the original gene mutation data space in a low-dimensional space, the method fully utilizes tumor cell line genomic data, promotes the preservation of as much gene feature information as possible, and makes the target anticancer drug response prediction model more effective, more generalizable, and more accurate, thereby improving the prediction accuracy of the target anticancer drug response prediction model for anticancer drug responses.

[0122] This embodiment uses gene expression data and gene mutation data to predict drug response (measured by logarithmic IC50 value). This method fully integrates the gene expression and mutation characteristics of drug-producing cell lines and uses a fully connected network to predict anticancer drug sensitivity. This addresses the shortcomings of related technologies that fail to adequately consider genomic characteristics and the relationship between mutation expression and drug efficacy in accurately predicting drug response based on genomic sequences and improving drug sensitivity prediction. By using self-supervised learning and the MDS algorithm to reduce the dimensionality of the original gene expression data and high-dimensional mutation data, a fully connected network is constructed using CCLE genomics data and drug response data to predict the genomic sensitivity of different individuals to anticancer drugs.

[0123] In some alternative implementations, step a3 above includes:

[0124] Step a31: Use gene expression data and gene mutation data of different tumor samples obtained from the Cancer Genome Atlas to perform self-supervised learning of tumor cell line gene expression encoder to obtain target tumor cell line gene expression encoder.

[0125] In this process, after obtaining the training samples, the gene expression data and gene mutation data of different tumor samples obtained from the cancer genome map are used to perform self-supervised learning of the tumor cell line gene expression encoder, so as to train the target tumor cell line gene expression encoder.

[0126] It should be noted that during the self-supervised learning phase of the tumor cell line gene expression encoder, the gene expression data and gene mutation data of different tumor samples obtained from the Cancer Genome Atlas can be used as the training set, and the gene expression data and gene mutation data of different tumor samples obtained from the Cancer Cell Line Encyclopedia can be used as the validation set to train the tumor cell line gene expression encoder.

[0127] Self-supervised learning is a machine learning paradigm in which algorithms learn representations from the data itself without the need for manually labeled data.

[0128] Autoencoders are a type of self-supervised learning. An autoencoder is a neural network architecture consisting of two parts: an encoder and a decoder. The encoder is responsible for mapping the input data to a low-dimensional representation, also known as encoding. This low-dimensional representation is usually considered as the latent features or representation of the input data. The encoder can be a neural network composed of multiple hidden layers, which maps the input data from the original space to the latent space through layer-by-layer transformations. The decoder is responsible for mapping the low-dimensional representation generated by the encoder back to the original data, also known as decoding. The goal of the decoder is to restore the original data as closely as possible to the input data. The structure of the decoder is usually symmetrical to that of the encoder and can be a neural network composed of multiple hidden layers. The decoder maps the low-dimensional representation back to the space of the original data through layer-by-layer transformations, restoring the dimension and structure of the data.

[0129] In this embodiment, the tumor cell line gene expression encoder is the encoder portion of an autoencoder. It is understood that this embodiment may also include a tumor cell line gene expression decoder, corresponding to... Figure 3 The expression decoder in the code can be used to verify whether the gene expression encoder of the tumor cell line has been trained successfully.

[0130] In some alternative implementations, step a4 above includes:

[0131] Step a41: Using genomic data of different tumor samples as input data for the anticancer drug response prediction model, and the sensitivity values ​​of different tumor samples to different drug responses as labels, the anticancer drug response prediction model containing the target tumor cell line gene expression encoder and the tumor cell line gene mutation data dimensionality reduction module is trained to optimize the parameters of the anticancer drug response prediction network, determine the target anticancer drug response prediction network, and obtain the target anticancer drug response prediction model.

[0132] Specifically, genomic data of different tumor samples obtained from the Encyclopedia of Cancer Cell Lines were used as input data for the anticancer drug response prediction model. The half-inhibition concentration of different tumor samples to different drug responses was used as a label to train the anticancer drug response prediction model, which includes a target tumor cell line gene expression encoder and a tumor cell line gene mutation data dimensionality reduction module, in order to optimize the parameters of the anticancer drug response prediction network, determine the target anticancer drug response prediction network, and obtain the target anticancer drug response prediction model.

[0133] Figure 3 This is a schematic diagram of a target anticancer drug response prediction model according to an embodiment of the present invention, as shown below. Figure 3 As shown, in this target anticancer drug response prediction model, the encoder E is expressed. encThe structure (number of network layers and number of neurons per layer) is fixed; that is, the structure of the tumor cell line gene expression encoder is fixed. Their synaptic parameters are initialized using parameters obtained from genomic data of different tumor samples in the pre-trained TCGA. P is a 5-layer feedforward neural network, i.e., an anticancer drug response prediction network. Its first layer is used to merge gene expression data and gene mutation data from CCLE tumor samples, and the last layer, with neuron d, generates the IC50 response to drug d. 50 The synaptic parameters are initialized using a uniform distribution and updated during training.

[0134] In this embodiment, genomic data from 80%, 10%, and 10% of tumor samples in CCLE are used as training, validation, and test datasets, respectively, to train the anticancer drug response prediction model. Of course, the proportions of data in the training, validation, and test sets can be randomly allocated. The training set is used to train the model, the validation set is used to prevent overfitting, and the test set is used to evaluate model performance (MSE(IC)). CCLE (:,C test ), IC CCLE (:,C test C test This represents the cell line data in the test set.

[0135] Using a trained anticancer drug response prediction model, we predict the drug response data of tumor samples in TCGA. For a tumor t, we analyze its gene mutation data and gene expression data {M}. TCGA (:t),E TCGA The data (:t) is fed into the model, where a low-dimensional representation of gene expression data is obtained through an expression encoder. Mutation feature MDS compression is performed on the gene mutation data to reduce its dimensionality, resulting in a low-dimensional representation. These low-dimensional representations of gene expression and mutation data are then input into a prediction network to obtain the predicted half-inhibitory concentrations (WICs) of the tumor response to different drugs. That is, the genomic sequence of cancer cell lines relative to the IC50 of different drugs. 50 Among them, the lower the predicted value, the higher the antibody sensitivity, the stronger the specificity, and the better the response to the drug.

[0136] In some optional embodiments, prior to step a14, the anticancer drug response prediction method further includes:

[0137] Step b1 involves standardizing and imputing missing values ​​for gene expression and gene mutation data from different tumor samples obtained from the Encyclopedia of Cancer Cell Lines.

[0138] This can be understood as follows: after obtaining gene expression and mutation data from different tumor samples from the Cancer Cell Line Encyclopedia, it is necessary to standardize and impute missing values ​​in this data. Data standardization refers to converting the data into a standardized form with a uniform scale and range. In data analysis and machine learning, standardization is typically used to eliminate dimensional differences in the data, enabling comparison and standardized processing of different features.

[0139] In this embodiment, the data can be standardized using the following method:

[0140] Z-Score standardization: Transforms data into a standard normal distribution with a mean of 0 and a standard deviation of 1. For each data point, calculate the difference between the mean and the standard deviation, and divide by the standard deviation.

[0141] Min-Max standardization: Linearly maps data to a specified interval between a minimum and a maximum value. For each data point, calculates the difference between it and the minimum value, and divides it by the difference between the maximum and the minimum value.

[0142] Decimal scaling standardization: Divide the data by a fixed base, usually the largest absolute value in the data, so that the data falls within the range of [-1, 1].

[0143] The appropriate standardization method can be selected based on actual needs. Standardization can improve the comparability and interpretability between different features, which helps to improve the performance of data analysis and machine learning models.

[0144] Imputing missing values ​​is the process of handling missing values ​​in data. When some observations or features are missing in a dataset, imputing missing values ​​can help us better utilize the data for analysis and modeling.

[0145] In this embodiment, missing values ​​can be filled in the data using the following method:

[0146] Removing missing values: The simplest method is to directly delete the observations or features containing missing values. However, this may result in data loss and is only suitable for situations with a small number of missing values.

[0147] Mean / Median / Mode Imputation: For numerical features, missing values ​​can be imputed using the mean, median, or mode. Mean imputation is suitable when the data distribution is approximately normal, median imputation is suitable when there are many outliers, and mode imputation is suitable when the data is discrete.

[0148] Regression imputation: This method uses other feature variables or the target variable to perform regression modeling, predict missing values, and then impute them. This approach requires considering the correlation between features and selecting an appropriate regression model.

[0149] Interpolation imputation: Using interpolation methods based on existing data, such as linear interpolation, polynomial interpolation, spline interpolation, etc., to estimate missing values.

[0150] Fill in missing values ​​using expertise: You can use expertise or domain knowledge in this field to fill in missing values.

[0151] It should be noted that the choice of an appropriate method for imputing missing values ​​depends on the characteristics of the data, the type of missing values, and their distribution. Furthermore, it is essential to perform validation and sensitivity analysis on the data after imputation to ensure that the imputed data retains its original characteristics and structure.

[0152] Step b2 involves standardizing and filling missing values ​​for the half-inhibition concentrations of different tumor samples in the Cancer Cell Line Encyclopedia obtained from the Cancer Drug Sensitivity Genomics Project for different drug responses.

[0153] It is understandable that after obtaining the half-inhibition concentrations (WICs) of different tumor samples to different drugs from the Cancer Cell Line Encyclopedia from the Cancer Drug Sensitivity Genomics Project, it is necessary to standardize and fill in missing values ​​for the WICs of different tumor samples to different drugs from the Cancer Cell Line Encyclopedia.

[0154] Step b3 involves standardizing and imputing missing values ​​for gene expression and gene mutation data from different tumor samples obtained from the cancer genome atlas.

[0155] It can be understood that after obtaining gene expression data and gene mutation data of different tumor samples from the cancer genome atlas, it is necessary to standardize and fill in missing values ​​for these gene expression data and gene mutation data.

[0156] Furthermore, the training samples were constructed by standardizing and imputing missing values ​​on gene expression and gene mutation data of different tumor samples obtained from the Encyclopedia of Cancer Cell Lines, half-inhibitory concentrations of different tumor samples to different drug responses obtained from the Encyclopedia of Cancer Cell Lines obtained from the Cancer Drug Sensitivity Genomics Project, and gene expression and gene mutation data of different tumor samples obtained from the Cancer Genome Atlas.

[0157] This embodiment provides an anticancer drug response prediction method that utilizes knowledge obtained from the Human Genome Project and employs deep learning methods to construct a target anticancer drug response prediction model to study how genetically differentially expressed genes influence tumor response to drugs. Specifically, a model for predicting tumor drug response is proposed based on self-supervised learning and the MDS algorithm. This model comprises two subnetworks: a gene expression encoder pre-trained using a large pan-cancer dataset to abstract the core representation of high-dimensional gene expression data, and a drug response prediction network that integrates the output of the expression encoder with the low-dimensional feature representation of gene mutation data obtained using the MDS algorithm. For a given tumor sample's genomic data, this model can predict the IC50 response of that sample to different drugs. 50 These values ​​provide a basis for personalized treatment and the development of new anticancer drugs.

[0158] This embodiment integrates additional genomic mutation information (such as copy number alterations) into the data matrix M. TCGA and M CCLE This enriches the complexity of tumor mutations used in model training and further reduces the training mean squared error (MSE). The targeted anticancer drug response prediction model utilizes a large number of tumor samples from TCGA to learn the genetic background from gene expression data. Then, it further trains the model using pharmacogenomics data developed in human cancer cell lines by the GDSC project, along with corresponding genomic and transcriptomic alterations. Finally, it is applied again to TCGA data to predict tumor drug responses. Predicting the effectiveness of a drug against a specific tumor based on genetic differences, helping to prevent adverse drug reactions, and discovering new anticancer drugs have significant reference value for cancer treatment. This model can be widely applied to the integration of other omics data and broader drug research.

[0159] This embodiment also provides an anticancer drug response prediction device, which is used to implement the above embodiments and preferred embodiments, and will not be repeated as already described. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0160] This embodiment provides an anticancer drug response prediction device, such as... Figure 4 As shown, it includes:

[0161] The genomics data acquisition module 401 is used to acquire genomics data of the tumor sample to be predicted. The genomics data includes gene expression data and gene mutation data.

[0162] The genomics data input module 402 is used to input the genomics data of the tumor sample to be predicted into the target anticancer drug response prediction model. The target anticancer drug response prediction model includes a target tumor cell line gene expression encoder, a tumor cell line gene mutation data dimensionality reduction processing module, and a target anticancer drug response prediction network.

[0163] The low-dimensional characterization acquisition module 403 for gene expression data is used to extract features from the gene expression data of the tumor sample to be predicted using the gene expression encoder of the target tumor cell line, thereby obtaining a low-dimensional characterization of the gene expression data of the tumor sample to be predicted.

[0164] The low-dimensional representation acquisition module 404 for gene mutation data is used to perform dimensionality reduction processing on the gene mutation data of the tumor cell line gene mutation data to be predicted using the tumor cell line gene mutation data dimensionality reduction processing module, so as to obtain a low-dimensional representation of the gene mutation data of the tumor sample to be predicted.

[0165] The anticancer drug response prediction value acquisition module 405 is used to input the low-dimensional representation of the gene expression data and the low-dimensional representation of the gene mutation data of the tumor sample to be predicted into the target anticancer drug response prediction network to obtain the predicted value of the sensitivity value of the tumor sample to different drug responses.

[0166] In some optional embodiments, the anticancer drug response prediction device further includes:

[0167] The training sample construction module is used to construct training samples, which include genomic data of different tumor samples and sensitivity values ​​of different tumor samples to different drugs.

[0168] The anticancer drug response prediction model construction module is used to construct an anticancer drug response prediction model. The anticancer drug response prediction model includes: a tumor cell line gene expression encoder, a tumor cell line gene mutation data dimensionality reduction module, and an anticancer drug response prediction network. Among them, the tumor cell line gene expression encoder is used to extract features from the gene expression data of tumor samples to obtain a low-dimensional representation of the gene expression data of tumor samples. The tumor cell line gene mutation data dimensionality reduction module is used to perform dimensionality reduction processing on the gene mutation data of tumor samples using a multidimensional scaling transformation algorithm to obtain a low-dimensional representation of the gene mutation data of tumor samples. The anticancer drug response prediction network is used to predict the sensitivity values ​​of tumor samples to different drug responses based on the low-dimensional representations of gene expression data and gene mutation data of tumor samples, so as to obtain the predicted values ​​of tumor sample sensitivity to different drug responses.

[0169] The target tumor cell line gene expression encoder acquisition module is used to perform self-supervised learning of tumor cell line gene expression data on the tumor cell line gene expression encoder using training samples, so as to obtain the target tumor cell line gene expression encoder.

[0170] The target anticancer drug response prediction model acquisition module is used to train the anticancer drug response prediction model, which includes a target tumor cell line gene expression encoder and a tumor cell line gene mutation data dimensionality reduction module, using training samples to optimize the parameters of the anticancer drug response prediction network, determine the target anticancer drug response prediction network, and obtain the target anticancer drug response prediction model.

[0171] In some optional implementations, the target anticancer drug response prediction model acquisition module includes:

[0172] The target anticancer drug response prediction model acquisition subunit is used to train the anticancer drug response prediction model, which includes a target tumor cell line gene expression encoder and a tumor cell line gene mutation data dimensionality reduction module, using genomic data of different tumor samples as input data and sensitivity values ​​of different tumor samples to different drug responses as labels. This optimizes the parameters of the anticancer drug response prediction network, determines the target anticancer drug response prediction network, and obtains the target anticancer drug response prediction model.

[0173] In some optional implementations, the training sample construction module includes:

[0174] The first data acquisition unit is used to obtain gene expression data and gene mutation data of different tumor samples from the Encyclopedia of Cancer Cell Lines.

[0175] The half-inhibition concentration acquisition unit is used to obtain the half-inhibition concentrations of different tumor samples in the Cancer Cell Line Encyclopedia for the response to different drugs from the Cancer Drug Sensitivity Genomics Project, where the half-inhibition concentration is used to characterize the sensitivity value.

[0176] The second data acquisition unit is used to acquire gene expression data and gene mutation data of different tumor samples from the cancer genome map.

[0177] The training sample construction subunit is used to construct training samples based on gene expression and gene mutation data of different tumor samples obtained from the Encyclopedia of Cancer Cell Lines, half-inhibitory concentrations of different tumor samples to different drugs obtained from the Encyclopedia of Cancer Cell Lines obtained from the Cancer Drug Sensitivity Genomics Project, and gene expression and gene mutation data of different tumor samples obtained from the Cancer Genome Atlas.

[0178] In some optional implementations, the target tumor cell line gene expression encoder acquisition module includes:

[0179] The target tumor cell line gene expression encoder acquisition subunit is used to perform self-supervised learning of the tumor cell line gene expression encoder using gene expression data and gene mutation data of different tumor samples obtained from the Cancer Genome Atlas, so as to obtain the target tumor cell line gene expression encoder.

[0180] In some optional implementations, the target anticancer drug response prediction model acquisition subunit includes:

[0181] The training unit is used to train an anticancer drug response prediction model by using genomic data of different tumor samples obtained from the Encyclopedia of Cancer Cell Lines as input data for the model, and the half-inhibition concentration of different tumor samples for different drug responses as labels. The model includes a gene expression encoder for the target tumor cell line and a dimensionality reduction module for tumor cell line gene mutation data.

[0182] In some optional implementations, before constructing training samples based on gene expression and mutation data of different tumor samples obtained from the Encyclopedia of Cancer Cell Lines, half-inhibitory concentrations of different tumor samples to different drug responses obtained from the Encyclopedia of Cancer Cell Lines obtained from the Cancer Drug Sensitivity Genomics Project, and gene expression and mutation data of different tumor samples obtained from the Cancer Genome Atlas, the anticancer drug response prediction device further includes:

[0183] The first data processing unit is used to standardize and fill missing values ​​in gene expression and gene mutation data of different tumor samples obtained from the Encyclopedia of Cancer Cell Lines.

[0184] The second data processing unit is used to standardize and fill missing values ​​for the half-inhibition concentrations of different tumor samples in the Cancer Cell Line Encyclopedia obtained from the Cancer Drug Sensitivity Genomics Project for responses to different drugs.

[0185] The third data processing unit is used to standardize and fill missing values ​​in the gene expression and gene mutation data of different tumor samples obtained from the cancer genome map.

[0186] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0187] In this embodiment, the anticancer drug response prediction device is presented in the form of a functional unit. Here, a unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0188] This invention also provides a computer device having the above-described features. Figure 4 The device shown is for predicting the response to anticancer drugs.

[0189] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 5 As shown, the computer device includes one or more processors 501, memory 502, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 5 Take a processor 501 as an example.

[0190] Processor 501 may be a central processing unit, a network processor, or a combination thereof. Processor 501 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.

[0191] The memory 502 stores instructions executable by at least one processor 501 to cause at least one processor 501 to perform the method shown in the above embodiments.

[0192] Memory 502 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, memory 502 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, memory 502 may optionally include memory remotely located relative to processor 501, and this remote memory may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0193] Memory 502 may include volatile memory, such as random access memory; memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; memory 502 may also include combinations of the above types of memory.

[0194] The computer device also includes a communication interface 503 for communicating with other devices or communication networks.

[0195] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.

[0196] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0197] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for predicting a response to an anticancer drug, characterized by, The method comprises: obtaining genomics data of a tumor sample to be predicted, the genomics data comprising gene expression data and gene mutation data; inputting the genomics data of the tumor sample to be predicted into a target anticancer drug response prediction model, the target anticancer drug response prediction model comprising a target tumor cell line gene expression encoder, a tumor cell line gene mutation data dimension reduction processing module, and a target anticancer drug response prediction network; extracting features from the gene expression data of the tumor sample to be predicted by using the target tumor cell line gene expression encoder to obtain a low-dimensional representation of the gene expression data of the tumor sample to be predicted; performing dimension reduction processing on the gene mutation data of the tumor sample to be predicted by using the tumor cell line gene mutation data dimension reduction processing module to obtain a low-dimensional representation of the gene mutation data of the tumor sample to be predicted; inputting the low-dimensional representation of the gene expression data of the tumor sample to be predicted and the low-dimensional representation of the gene mutation data of the tumor sample to be predicted into the target anticancer drug response prediction network to obtain a predicted value of the sensitivity of the tumor sample to be predicted to different drugs; obtaining the target anticancer drug response prediction model by the following steps, comprising: constructing training samples, the training samples comprising genomics data of different tumor samples and sensitivity values of different tumor samples to different drugs; constructing an anticancer drug response prediction model, the anticancer drug response prediction model comprising a tumor cell line gene expression encoder, a tumor cell line gene mutation data dimension reduction processing module, and an anticancer drug response prediction network, wherein the tumor cell line gene expression encoder is used to extract features from gene expression data of a tumor sample to obtain a low-dimensional representation of the gene expression data of the tumor sample, the tumor cell line gene mutation data dimension reduction processing module is used to perform dimension reduction processing on gene mutation data of a tumor sample by using a multidimensional scaling algorithm to obtain a low-dimensional representation of the gene mutation data of the tumor sample, and the anticancer drug response prediction network is used to predict sensitivity values of a tumor sample to different drugs according to the low-dimensional representation of the gene expression data of the tumor sample and the low-dimensional representation of the gene mutation data of the tumor sample to obtain predicted values of the sensitivity of the tumor sample to different drugs; performing self-supervised learning of tumor cell line gene expression data on the tumor cell line gene expression encoder by using the training samples to obtain a target tumor cell line gene expression encoder; training the anticancer drug response prediction model comprising the target tumor cell line gene expression encoder and the tumor cell line gene mutation data dimension reduction processing module by using the training samples to optimize parameters of the anticancer drug response prediction network, determine a target anticancer drug response prediction network, and obtain the target anticancer drug response prediction model.

2. The method of claim 1, wherein, The training sample is used to train the anticancer drug response prediction model comprising the target tumor cell line gene expression encoder and the tumor cell line gene mutation data dimension reduction processing module, to optimize the parameters of the anticancer drug response prediction network, determine the target anticancer drug response prediction network, and obtain the target anticancer drug response prediction model, comprising: The genomic data of different tumor samples is used as the input data of the anticancer drug response prediction model, and the sensitivity values of different tumor samples to different drugs are used as labels to train the anticancer drug response prediction model comprising the target tumor cell line gene expression encoder and the tumor cell line gene mutation data dimension reduction processing module, to optimize the parameters of the anticancer drug response prediction network, determine the target anticancer drug response prediction network, and obtain the target anticancer drug response prediction model.

3. The method of claim 2, wherein, The training sample is constructed, comprising: Obtaining gene expression data and gene mutation data of different tumor samples from a cancer cell line encyclopedia; Obtaining half-inhibitory concentrations of different tumor samples in the cancer cell line encyclopedia to different drugs from a cancer drug sensitivity genomics project, wherein the half-inhibitory concentrations are used to represent sensitivity values; Obtaining gene expression data and gene mutation data of different tumor samples from a cancer genome atlas; Based on the gene expression data and gene mutation data of different tumor samples obtained from the cancer cell line encyclopedia, the half-inhibitory concentrations of different tumor samples in the cancer cell line encyclopedia to different drugs obtained from the cancer drug sensitivity genomics project, and the gene expression data and gene mutation data of different tumor samples obtained from the cancer genome atlas, the training sample is constructed.

4. The method of claim 3, wherein, The training sample is used to train the tumor cell line gene expression encoder for self-supervised learning of tumor cell line gene expression data, to obtain the target tumor cell line gene expression encoder, comprising: The tumor cell line gene expression encoder is used for self-supervised learning of tumor cell line gene expression data based on the gene expression data and gene mutation data of different tumor samples obtained from the cancer genome atlas, to obtain the target tumor cell line gene expression encoder.

5. The method of claim 3, wherein, The genomic data of different tumor samples is used as the input data of the anticancer drug response prediction model, and the sensitivity values of different tumor samples to different drugs are used as labels to train the anticancer drug response prediction model comprising the target tumor cell line gene expression encoder and the tumor cell line gene mutation data dimension reduction processing module, to optimize the parameters of the anticancer drug response prediction network, determine the target anticancer drug response prediction network, and obtain the target anticancer drug response prediction model. The genomic data of different tumor samples obtained from the cancer cell line encyclopedia is used as the input data of the anticancer drug response prediction model, and the half-inhibitory concentrations of different tumor samples to different drugs are used as labels to train the anticancer drug response prediction model comprising the target tumor cell line gene expression encoder and the tumor cell line gene mutation data dimension reduction processing module.

6. The method of claim 3, wherein, Before constructing the training sample based on the gene expression data and gene mutation data of different tumor samples obtained from the Cancer Cell Line Encyclopedia, the semi-inhibitory concentrations of different tumor samples in the Cancer Cell Line Encyclopedia obtained from the Cancer Drug Sensitivity Genomics Project to different drugs, and the gene expression data and gene mutation data of different tumor samples obtained from The Cancer Genome Atlas, the method further comprises: standardizing and filling missing values of the gene expression data and gene mutation data of different tumor samples obtained from the Cancer Cell Line Encyclopedia; standardizing and filling missing values of the semi-inhibitory concentrations of different tumor samples in the Cancer Cell Line Encyclopedia obtained from the Cancer Drug Sensitivity Genomics Project to different drugs; standardizing and filling missing values of the gene expression data and gene mutation data of different tumor samples obtained from The Cancer Genome Atlas.

7. An anticancer drug response prediction device characterized by comprising: The device comprises: a genomics data acquisition module for acquiring genomics data of a tumor sample to be predicted, the genomics data comprising gene expression data and gene mutation data; a genomics data input module for inputting the genomics data of the tumor sample to be predicted into a target anticancer drug response prediction model, the target anticancer drug response prediction model comprising a target tumor cell line gene expression encoder, a tumor cell line gene mutation data dimensionality reduction processing module, and a target anticancer drug response prediction network; a low-dimensional representation acquisition module for gene expression data for extracting features from the gene expression data of the tumor sample to be predicted using the target tumor cell line gene expression encoder to obtain a low-dimensional representation of the gene expression data of the tumor sample to be predicted; a low-dimensional representation acquisition module for gene mutation data for performing dimensionality reduction processing on the gene mutation data of the tumor sample to be predicted using the tumor cell line gene mutation data dimensionality reduction processing module to obtain a low-dimensional representation of the gene mutation data of the tumor sample to be predicted; an anticancer drug response prediction value acquisition module for inputting the low-dimensional representation of the gene expression data of the tumor sample to be predicted and the low-dimensional representation of the gene mutation data of the tumor sample to be predicted into the target anticancer drug response prediction network to obtain a predicted value of the sensitivity numerical value of the tumor sample to be predicted to different drugs; The anticancer drug response prediction device further comprises: a training sample construction module for constructing a training sample, the training sample comprising genomics data of different tumor samples and sensitivity numerical values of different tumor samples to different drugs; The anti-cancer drug response prediction model construction module is used for constructing an anti-cancer drug response prediction model, and the anti-cancer drug response prediction model comprises a tumor cell line gene expression encoder, a tumor cell line gene mutation data dimension reduction processing module and an anti-cancer drug response prediction network, wherein the tumor cell line gene expression encoder is used for feature extraction on gene expression data of a tumor sample to obtain a low-dimensional representation of the gene expression data of the tumor sample, the tumor cell line gene mutation data dimension reduction processing module is used for dimension reduction processing on gene mutation data of the tumor sample by using a multidimensional scaling transformation algorithm to obtain a low-dimensional representation of the gene mutation data of the tumor sample, and the anti-cancer drug response prediction network is used for predicting a sensitivity value of the tumor sample to different drug responses according to the low-dimensional representation of the gene expression data of the tumor sample and the low-dimensional representation of the gene mutation data of the tumor sample to obtain a prediction value of the sensitivity value of the tumor sample to different drug responses. The target tumor cell line gene expression encoder acquisition module is used for self-supervised learning of tumor cell line gene expression data of the tumor cell line gene expression encoder by using training samples to obtain the target tumor cell line gene expression encoder. The target anti-cancer drug response prediction model acquisition module is used for training the anti-cancer drug response prediction model comprising the target tumor cell line gene expression encoder and the tumor cell line gene mutation data dimension reduction processing module by using training samples to optimize parameters of the anti-cancer drug response prediction network, determine the target anti-cancer drug response prediction network and obtain the target anti-cancer drug response prediction model.

8. A computer device, comprising: Comprise: A memory and a processor, which are mutually connected in communication, the memory stores computer instructions, and the processor executes the computer instructions to perform the anti-cancer drug response prediction method of any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for making a computer execute the anti-cancer drug response prediction method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Cancer cell line therapeutic drug prediction method based on multidimensional network

    CN110232978A

  • Neural network-based anticancer drug synergistic effect prediction method

    CN110277174A