A method for constructing a bladder cancer chemotherapy drug sensitivity prediction model and application thereof

By constructing a bladder cancer chemotherapy drug sensitivity prediction model based on whole transcriptome sequencing and machine learning, the problem of low accuracy of existing models has been solved. This model achieves high-precision, economical, and applicable prediction of bladder cancer chemotherapy drug sensitivity in the Chinese population, guiding personalized treatment.

CN119360975BActive Publication Date: 2026-03-31HARBIN MEDICAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-16
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing bladder cancer chemotherapy drug sensitivity prediction models have low accuracy and predict IC50 values, which cannot intuitively represent the drug effect. They also lack real clinical data specific to Chinese patients, resulting in a discrepancy between the predicted results and the actual effects.

Method used

The training set was constructed by acquiring whole transcriptome sequencing data, chemotherapy drug type data, and chemotherapy sensitivity data of tumor samples from bladder cancer patients. Multiple candidate models were obtained by training them using various machine learning algorithms, and prediction was performed using hard voting. Clinical data from the Chinese population were used for model training and prediction.

Benefits of technology

It achieves high-precision prediction of bladder cancer chemotherapy drug sensitivity, with an accuracy of over 80%. The prediction results are clear, applicable to the Chinese population, with a short cycle and low cost. Sequencing, analysis and prediction can be completed within 7 days, providing personalized treatment guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119360975B_ABST
    Figure CN119360975B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of drug sensitivity prediction, and discloses a construction method of a bladder cancer chemotherapy drug sensitivity prediction model and application thereof, which comprises the following steps: constructing a training set; according to the training set, a plurality of candidate models are trained through a plurality of machine learning algorithms respectively; and according to preset model indexes, the plurality of candidate models are screened to obtain at least two bladder cancer chemotherapy drug sensitivity prediction models. The application is based on the administered bladder cancer chemotherapy drug type data of a target group, the immune cell infiltration level data of a tumor sample tissue and PTC data, utilizes a machine learning algorithm, trains a plurality of high-precision bladder cancer chemotherapy drug sensitivity prediction models, and can be used to realize accurate prediction of the bladder cancer chemotherapy drug sensitivity in combination with a hard voting method. The model has high prediction accuracy, is more targeted for the applicable population, the prediction result is simple and clear, the use cycle is short, and the cost is low.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of drug sensitivity prediction technology, and in particular to a method for constructing a drug sensitivity prediction model for bladder cancer chemotherapy and its application. Background Technology

[0002] Bladder cancer is the most common malignant tumor of the urinary system, ranking first in incidence among urogenital tumors in my country. Due to the heterogeneity of bladder cancer and its insensitivity to chemotherapy drugs, treatment can be somewhat unpredictable. Chemotherapy drugs themselves are cytotoxic; during chemotherapy, the tumor may not be controlled, and adverse reactions may occur, causing unnecessary suffering for patients, affecting their quality of life, and even their prognosis.

[0003] With the rapid development of molecular biology, it has been discovered that gene expression in cancer is closely related to tumor occurrence, development, metastasis, and prognosis, and also has a certain correlation with the efficacy of chemotherapy. Furthermore, an increasing number of machine learning models are being used to predict drug sensitivity. How to utilize machine learning methods to accurately predict patient drug sensitivity and assist in guiding personalized treatment for cancer patients has significant clinical implications. Currently, most drug sensitivity prediction models use data primarily from public databases, lacking real clinical data specific to the Chinese population. This results in inconsistent prediction results compared to actual clinical outcomes, and the models' predictions are based on IC50 values, failing to directly reflect the drug's effect.

[0004] Therefore, there is an urgent need for a method to construct a predictive model for the sensitivity of bladder cancer chemotherapy drugs, so as to help predict patients' drug sensitivity intuitively and accurately, guide personalized treatment for cancer patients, and improve the cure rate. Summary of the Invention

[0005] This invention provides a method for constructing a bladder cancer chemotherapy drug sensitivity prediction model and its application, which addresses the shortcomings of existing drug sensitivity prediction models, such as low prediction accuracy and the fact that the prediction results are IC50 values, which fail to intuitively represent the drug effect.

[0006] This invention provides a method for constructing a training set, comprising:

[0007] The study aimed to acquire whole transcriptome sequencing data of tumor samples from the target population, data on the types of chemotherapy drugs administered for bladder cancer, and data on the sensitivity of bladder cancer chemotherapy drugs. The target population was bladder cancer patients, and the data on the sensitivity of bladder cancer chemotherapy drugs included data on bladder cancer chemotherapy drug response and data on bladder cancer chemotherapy drug resistance.

[0008] Gene expression data of tumor sample tissues from the target population were obtained based on whole transcriptome sequencing data of tumor sample tissues from the target population.

[0009] Based on gene expression data of tumor sample tissues from the target population, data on the level of immune cell infiltration in tumor sample tissues from the target population were obtained.

[0010] A training set was constructed using data on the types of bladder cancer chemotherapy drugs administered to the target population and data on the level of immune cell infiltration in tumor tissue samples as sample data, and data on the bladder cancer chemotherapy drug sensitivity of the target population as label data.

[0011] In one implementation, the data on the types of chemotherapy drugs administered for bladder cancer includes any one or any combination of the following: gemcitabine data, paclitaxel data, and docetaxel data.

[0012] In one implementation, obtaining gene expression data of tumor sample tissues from the target population based on whole transcriptome sequencing data of tumor sample tissues from the target population includes:

[0013] Obtain reference genome data;

[0014] Data quality control was performed on the whole transcriptome sequencing data of tumor samples from the target population.

[0015] The quality-controlled whole transcriptome sequencing data was compared with the reference genome data to obtain the bam file;

[0016] Gene expression levels are calculated based on the BAM file to obtain gene expression data of tumor sample tissues from the target population.

[0017] In one implementation, the data quality control of the whole transcriptome sequencing data of tumor samples from the target population includes:

[0018] According to preset data quality control conditions, the whole transcriptome sequencing data of tumor samples from the target population were subjected to data quality control. These preset data quality control conditions included:

[0019] Reads with a nitrogen content exceeding a preset threshold are filtered out.

[0020] Reads with sequencing quality values ​​below a preset quality threshold are filtered out.

[0021] In one implementation, the preset quality value threshold is 15.

[0022] In one implementation, obtaining immune cell infiltration level data of tumor sample tissues from the target population based on gene expression data of tumor sample tissues from the target population includes:

[0023] Based on gene expression data of tumor samples from the target population, multiple cell infiltration level assessment results were obtained through various cell infiltration level assessment methods, thereby obtaining immune cell infiltration level data of tumor samples from the target population. Among them, the multiple cell infiltration level assessment methods include any one of the following or any combination thereof: CIBERSORT, CIBERSORT-ABS, EPIC, MCPCOUNTER, QUANTISEQ, TIMER, XCELL.

[0024] In one embodiment, the present invention provides a method for constructing a training set, which further includes:

[0025] The training set is cleaned according to preset cleaning rules, which include any one of the following or any combination thereof:

[0026] For each immune cell infiltration level data in the tumor sample tissue of the target population, if the value of the immune cell infiltration level assessment result is >60%, then the immune cell infiltration level data is deleted.

[0027] For immune cell infiltration level data of tumor sample tissues of the target population obtained by multiple cell infiltration level assessment methods, if ≥80% of the values ​​in the immune cell infiltration level assessment results obtained by a certain cell infiltration level assessment method are all 0, the immune cell infiltration level assessment results obtained by that cell infiltration level assessment method shall be deleted.

[0028] When the data on the type of chemotherapy drug used for bladder cancer is represented as the same chemotherapy drug, if the ratio of bladder cancer chemotherapy drug response to bladder cancer chemotherapy drug resistance in the target population is outside the preset ratio range, then some bladder cancer chemotherapy drug sensitivity data will be randomly deleted until the ratio of bladder cancer chemotherapy drug response to bladder cancer chemotherapy drug resistance is within the preset range.

[0029] This invention provides a method for constructing a bladder cancer chemotherapy drug sensitivity prediction model, comprising:

[0030] Obtain the training set constructed according to any of the above methods;

[0031] Based on the training set, multiple candidate models were trained using various machine learning algorithms.

[0032] Based on the preset model indicators, multiple candidate models are screened to obtain at least two final candidate models, which will serve as at least two bladder cancer chemotherapy drug sensitivity prediction models.

[0033] In one implementation, at least two final candidate models include four final candidate models, wherein the four final candidate models are candidate models trained on the training set using the EasyEnsemble Classifier (an ensemble classifier based on random undersampling), BalancedRandomForest Classifier (a balanced random forest classifier), BalancedBagging Classifier (an ensemble classifier based on Bagging technology), and RUSBoost Classifier (a classifier based on random undersampling), respectively.

[0034] In one implementation, the preset model metrics include any one or any combination of the following: Area Under Curve (AUC), F1 score, accuracy, and Matthews Correlation Coefficient (MCC).

[0035] This invention provides a bladder cancer chemotherapy drug sensitivity prediction system, comprising:

[0036] The data receiving module is used to receive data on the type of chemotherapy drug to be administered to the subject of bladder cancer and data on the level of immune cell infiltration in tumor sample tissue sent by at least one client, wherein the subject of the test is a bladder cancer patient;

[0037] The prediction module is used to: obtain the bladder cancer chemotherapy drug sensitivity prediction result of the test subject by combining at least two bladder cancer chemotherapy drug sensitivity prediction models obtained by the construction method of the bladder cancer chemotherapy drug sensitivity prediction model described above with hard voting, based on the data of the type of bladder cancer chemotherapy drug to be administered to the test subject and the data of immune cell infiltration level of tumor sample tissue.

[0038] The data sending module is used to send the predicted results of the bladder cancer chemotherapy drug sensitivity of the test subject back to at least one client.

[0039] In one implementation, the prediction module includes:

[0040] The prediction submodule is used to: obtain at least two bladder cancer chemotherapy drug sensitivity prediction results for the test subject based on the data of the type of bladder cancer chemotherapy drug to be administered to the test subject and the data of immune cell infiltration level of tumor sample tissue, respectively, through at least two bladder cancer chemotherapy drug sensitivity prediction models obtained by the construction method of bladder cancer chemotherapy drug sensitivity prediction model described above.

[0041] The determination submodule is used to: obtain the final bladder cancer chemotherapy drug sensitivity prediction result for the test subject based on at least two bladder cancer chemotherapy drug sensitivity prediction results and a hard voting method. The hard voting method is as follows: when there is a unique bladder cancer chemotherapy drug sensitivity prediction result with the most votes among the at least two bladder cancer chemotherapy drug sensitivity prediction results for the test subject, that unique bladder cancer chemotherapy drug sensitivity prediction result with the most votes is used as the final bladder cancer chemotherapy drug sensitivity prediction result for the test subject; when there is no unique bladder cancer chemotherapy drug sensitivity prediction result with the most votes among the at least two bladder cancer chemotherapy drug sensitivity prediction results for the test subject, the bladder cancer chemotherapy drug sensitivity prediction result of the bladder cancer chemotherapy drug sensitivity prediction model with the highest accuracy corresponding to the bladder cancer chemotherapy drug type is used as the final bladder cancer chemotherapy drug sensitivity prediction result for the test subject.

[0042] The present invention also provides an electronic device, including a processor and a memory storing a computer program, wherein the processor executes the computer program to implement the training set construction method and the bladder cancer chemotherapy drug sensitivity prediction model construction method described above.

[0043] The present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method for constructing a training set or the method for constructing a bladder cancer chemotherapy drug sensitivity prediction model as described above.

[0044] The present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute any of the above-described methods for constructing a training set and for constructing a bladder cancer chemotherapy drug sensitivity prediction model.

[0045] This invention provides a method for constructing a bladder cancer chemotherapy drug sensitivity prediction model and its application. Based on data on the types of bladder cancer chemotherapy drugs administered to the target population, data on the level of immune cell infiltration in tumor sample tissues, and bladder cancer chemotherapy drug sensitivity data (PTC data), multiple high-precision bladder cancer chemotherapy drug sensitivity prediction models are trained using machine learning algorithms. These models can be used in conjunction with hard voting methods to achieve accurate prediction of bladder cancer chemotherapy drug sensitivity. The bladder cancer chemotherapy drug sensitivity prediction model constructed in this invention has high prediction accuracy, with an average accuracy greater than 80%; it is more targeted to the Chinese cancer patient population; the prediction results are simple and clear, divided into bladder cancer chemotherapy drug response and bladder cancer chemotherapy drug resistance; the usage cycle is short, and prediction results can be obtained within 7 days (including sequencing, analysis, and prediction) from obtaining tumor sample tissue from bladder cancer patients; the cost is low, with the main cost being the decreasing cost of second-generation sequencing year by year. Attached Figure Description

[0046] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0047] Figure 1 This is a flowchart illustrating a method for constructing a bladder cancer chemotherapy drug sensitivity prediction model provided by the present invention.

[0048] Figure 2 This is a schematic diagram of the structure of a bladder cancer chemotherapy drug sensitivity prediction system provided by the present invention.

[0049] Figure 3 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, embodiments of this invention, and should not be construed as limiting the invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention. In the description of this invention, it should be understood that the terminology used is for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0051] The following is combined with Figures 1-3This invention describes the construction method and application of a bladder cancer chemotherapy drug sensitivity prediction model. It should be noted that the execution entity for the construction method and application of the bladder cancer chemotherapy drug sensitivity prediction model provided by this invention can be any network-side device / terminal-side device that meets the technical requirements, such as a bladder cancer chemotherapy drug sensitivity prediction device.

[0052] Figure 1 This is a flowchart illustrating the method for constructing a bladder cancer chemotherapy drug sensitivity prediction model provided by the present invention.

[0053] Reference Figure 1 The present invention provides a method for constructing a bladder cancer chemotherapy drug sensitivity prediction model, which may include:

[0054] Step S110: Construct a training set, wherein the training set uses data on the types of bladder cancer chemotherapy drugs administered to the target population and data on the level of immune cell infiltration in tumor sample tissues as sample data, and data on the sensitivity of the target population to bladder cancer chemotherapy drugs as label data.

[0055] In this embodiment, step S110 may include:

[0056] S1101. Obtain whole transcriptome sequencing data, bladder cancer chemotherapy drug type data, and bladder cancer chemotherapy drug sensitivity data from tumor samples of the target population. The target population is a bladder cancer patient population. The bladder cancer chemotherapy drug sensitivity data includes bladder cancer chemotherapy drug response data and bladder cancer chemotherapy drug resistance data. The bladder cancer chemotherapy drug type data includes any one of the following or any combination thereof: gemcitabine data, paclitaxel data, and docetaxel data. The bladder cancer chemotherapy drug sensitivity data is obtained by detecting tumor samples of the target population using the microtumor drug sensitivity testing method (PTC).

[0057] S1102. Based on the whole transcriptome sequencing data of tumor sample tissues of the target population, obtain the gene expression data of tumor sample tissues of the target population. Specifically, this step can first perform data quality control on the whole transcriptome sequencing data of tumor sample tissues of the target population to filter out reads with the content of base N exceeding a preset threshold and reads with sequencing quality values ​​lower than a preset quality value threshold (e.g., 15). Then, compare the quality-controlled whole transcriptome sequencing data with the reference genome data to obtain a BAM file. Then, calculate the gene expression level (TPM) based on the BAM file to obtain the gene expression data of tumor sample tissues of the target population.

[0058] S1103. Based on the gene expression data of tumor sample tissues of the target population, multiple cell infiltration level assessment results are obtained through various cell infiltration level assessment methods, thereby obtaining immune cell infiltration level data of tumor sample tissues of the target population. Among them, the multiple cell infiltration level assessment methods include: CIBERSORT, CIBERSORT-ABS, EPIC, MCPCOUNTER, QUANTISEQ, TIMER, and XCELL.

[0059] In this embodiment, the immune cell infiltration level of tumor sample tissues of the target population was assessed using seven cell infiltration level assessment methods based on gene expression data (TPM). The seven immune cell infiltration level assessment results were obtained, and then the seven immune cell infiltration level assessment results were combined to obtain the final immune cell infiltration level assessment result of the tumor sample tissues of the target population, which is used as the immune cell infiltration level data of the tumor sample tissues of the target population.

[0060] S1104. Clean the training set according to preset cleaning rules, wherein the preset cleaning rules include any one of the following or any combination thereof:

[0061] For each immune cell infiltration level data in the tumor sample tissue of the target population, if more than 60% of the immune cell infiltration level assessment results are 0, then the immune cell infiltration level data is deleted (because an immune cell infiltration level data includes seven immune cell infiltration level assessment results, that is, if more than 60% of the seven immune cell infiltration level assessment results are 0, then the immune cell infiltration level data is deleted).

[0062] For immune cell infiltration level data of tumor sample tissues of the target population obtained through multiple cell infiltration level assessment methods, if ≥80% of the values ​​in the immune cell infiltration level assessment results obtained by a certain cell infiltration level assessment method are all 0, then the immune cell infiltration level assessment results obtained by that cell infiltration level assessment method are deleted (that is, for all immune cell infiltration level assessment results, if ≥80% of the values ​​in the immune cell infiltration level assessment results obtained by a certain cell infiltration level assessment method are all 0, then the immune cell infiltration level assessment results of tumor sample tissues obtained by that cell infiltration level assessment method are deleted).

[0063] When the data on the type of chemotherapy drug used for bladder cancer is represented as the same chemotherapy drug, if the ratio of bladder cancer chemotherapy drug response to bladder cancer chemotherapy drug resistance in the target population is outside the preset ratio range, then some bladder cancer chemotherapy drug sensitivity data will be randomly deleted until the ratio of bladder cancer chemotherapy drug response to bladder cancer chemotherapy drug resistance is within the preset range.

[0064] S1105. Using data on the types of bladder cancer chemotherapy drugs administered to the target population and data on the level of immune cell infiltration in tumor sample tissues as sample data, and using data on the bladder cancer chemotherapy drug sensitivity of the target population as label data, a training set is constructed.

[0065] Step S120: Based on the training set, train multiple candidate models using various machine learning algorithms. These algorithms may include those for classification tasks, such as decision trees, Naive Bayes, logistic regression, K-nearest neighbors, support vector machines, random forests, adaptive boosting, and gradient boosting tree algorithms.

[0066] Step S130: Based on the preset model indicators, screen multiple candidate models to obtain at least two final candidate models as at least two bladder cancer chemotherapy drug sensitivity prediction models. The preset model indicators include any one of the following or any combination thereof: area under the ROC curve (AUC), F1 score, accuracy, and Matthews correlation coefficient (MCC).

[0067] In this embodiment, four final candidate models are obtained, which are candidate models trained by the machine learning algorithms EasyEnsemble Classifier (an ensemble classifier based on random undersampling), BalancedRandomForest Classifier (a balanced random forest classifier), BalancedBagging Classifier (an ensemble classifier based on Bagging technology), and RUSBoostClassifier (a classifier based on random undersampling). This embodiment yields four bladder cancer chemotherapy drug sensitivity prediction models.

[0068] Among them, EasyEnsemble Classifier (an ensemble classifier based on random undersampling) performs multiple random samplings on the original dataset to obtain multiple subsets. For each subset, an algorithm (decision tree, support vector machine, etc.) is used to train a classifier. The prediction results of all classifiers are then evaluated using a voting method, with the one receiving the most votes being the classification result. This effectively reduces the problem of overfitting and improves the model's generalization ability.

[0069] The BalancedRandomForest Classifier works similarly to Random Forest, but differs in that it automatically considers the balance of subset samples during training. The final result is determined through voting or averaging, resulting in high accuracy and generalization performance for the overall model. It boasts high accuracy and generalization performance; it can handle high-dimensional data without dimensionality reduction; it can handle both discrete and continuous data without requiring dataset normalization; and it has fast training speed and can obtain variable importance rankings.

[0070] BalancedBagging Classifier (an ensemble classifier based on bagging technology), also known as the bagging algorithm, uses a sampling method with replacement to generate training data. It generates multiple training sets in parallel by randomly sampling the initial training set multiple times with replacement, corresponding to multiple base learners (with no strong dependencies between them). These base learners are then combined to construct a strong learner. BalancedBagging Classifier resamples each subset of the dataset before each estimator in the training set, otherwise similar to the bagging algorithm. It increases sample randomness to reduce variance and address the overfitting problem.

[0071] RUSBoost Classifier (based on random undersampling) is an ensemble learning method whose core idea is to construct a strong learning algorithm by combining weak learning algorithms. Boosting methods train multiple learning algorithms, with each model focusing on samples that previous models did not handle well. RUSBoost Classifier combines random undersampling and boosting methods in an ensemble specifically designed to handle imbalanced data. During the learning process, random undersampling of samples in each iteration alleviates the class imbalance problem. It improves the accuracy of weak classification algorithms; the algorithm system has a high detection rate and is less prone to overfitting.

[0072] After training and obtaining the bladder cancer chemotherapy drug sensitivity prediction model, it can be applied to the bladder cancer chemotherapy drug sensitivity prediction system. Based on the data of the type of bladder cancer chemotherapy drug to be administered to the test subject (there can be multiple types of bladder cancer chemotherapy drugs to be administered), the data of immune cell infiltration level of tumor sample tissue, and the hard voting method, the sensitivity prediction results of the test subject to different bladder cancer chemotherapy drugs can be obtained. This can help to predict the patient's drug sensitivity intuitively and accurately, guide the personalized treatment of cancer patients, and improve the patient's cure rate.

[0073] The following two examples will describe the construction and application of a bladder cancer chemotherapy drug sensitivity prediction model.

[0074] Example 1: Construction of a Predictive Model for Chemotherapy Drug Sensitivity in Bladder Cancer

[0075] Example 1 involved tumor tissue samples from 935 bladder cancer patients (target group) participating in the modeling process. The procedures included:

[0076] 1. LncRNA (whole transcriptome) sequencing was performed using the Illumina NovaSeq X Plus platform, with 30G of sequencing data per sample, to obtain whole transcriptome sequencing data of tumor tissue samples from the target population.

[0077] 2. Bioinformatics analysis of sequencing samples: Fastp (v 0.23.2) was used to perform data quality control on the whole transcriptome sequencing data of tumor samples from the target population, generating cleandata (whole transcriptome sequencing data after data quality control). STAR software (v 2.7.10a) was used to compare the cleandata data with the reference genome data (GRCh38) to obtain the bam file.

[0078] 3. Use RSEM (v 1.3.3) to calculate gene expression levels from the generated BAM file to obtain TPM values, which are the gene expression data of tumor sample tissues of the target population.

[0079] 4. The immune cell infiltration level of tumor sample tissues of the target population was assessed using gene expression data and seven cell infiltration level assessment methods (CIBERSORT, CIBERSORT-ABS, EPIC, MCPCOUNTER, QUANTISEQ, TIMER, XCELL) to obtain immune cell infiltration level data of tumor sample tissues of the target population.

[0080] 5. The classification of PTC drug sensitivity reactions in bladder cancer patients is divided into two categories: bladder cancer chemotherapy drug response (0 indicates) and bladder cancer chemotherapy drug resistance (1 indicates).

[0081] 6. The dataset is integrated, and the data columns include: bladder cancer patient ID, type ID of bladder cancer chemotherapy drug administered, immune cell infiltration level of tumor sample tissue assessed by different methods, and bladder cancer chemotherapy drug sensitivity classification.

[0082] 7. By dividing the dataset into a 70% training set and a 30% test set, and performing 10-fold cross-validation, four final candidate models were selected using AUC and F1 score metrics as four bladder cancer chemotherapy drug sensitivity prediction models. During the sampling process, the SMOTEENN method was used to address the imbalance in classification (response / resistance).

[0083] The optimal key parameters for the model set when training the four final candidate models in Example 1 are shown in Table 1.

[0084] Table 1 Optimal Key Parameter Settings for the Model

[0085]

[0086] The training results of the final candidate model in Example 1 are shown in Table 2.

[0087] Table 2 Model training results

[0088] Model Name AUC(%) F1 score Accuracy (%) MCC EasyEnsembleClassifier 87.7 0.81 90.3 0.81 BalancedRandomForestClassifier 85.9 0.76 88 0.74 BalancedBaggingClassifier 85 0.71 87 0.71 RUSBoostClassifier 84.2 0.73 86.5 0.71

[0089] After determining the optimal key parameters of the model, Example 1 uses a test set to verify the generalization ability of each final candidate model. The model test accuracy is shown in Table 3.

[0090] Table 3 Model Test Accuracy

[0091]

[0092]

[0093] Example 2: The bladder cancer chemotherapy drug sensitivity prediction model constructed in Example 1 was used to predict the drug sensitivity of 251 bladder cancer patients to bladder cancer chemotherapy drugs (gemcitabine, paclitaxel, docetaxel).

[0094] The specific steps include:

[0095] 1. LncRNA sequencing data (i.e., whole transcriptome sequencing data) were obtained from tumor tissue samples of 251 bladder cancer patients. The sequencing volume of the samples was guaranteed to be at least 30G. All 251 bladder cancer patients were on single-drug therapy. Specifically: 62 patients had clinical drug response data, and the drug response type was classified according to tag grading; 189 patients had drug sensitivity values ​​measured by PTC, and the drug response type was classified according to the drug sensitivity data.

[0096] 2. Perform bioinformatics analysis on the whole transcriptome sequencing data in step 1, including data quality control, alignment with reference genome data, and quantification of gene expression.

[0097] 3. Using gene expression data from tumor tissue samples of 251 bladder cancer patients, seven methods (CIBERSORT, CIBERSORT-ABS, EPIC, MCPCOUNTER, QUANTISEQ, TIMER, and XCELL) were used to assess the level of immune cell infiltration in the tumor tissue samples.

[0098] 4. Generate model input data d1, with column names including bladder cancer patient ID, type ID of bladder cancer chemotherapy drug to be administered, and immune cell infiltration level of tumor sample tissue assessed by different methods.

[0099] 5. Input the model input data d1 into the four bladder cancer chemotherapy drug sensitivity prediction models constructed in Example 1, and combine the hard voting method (combining the bladder cancer chemotherapy drug sensitivity prediction results of the four bladder cancer chemotherapy drug sensitivity prediction models, the voting result of the majority category is the final result, if there is a tie, the result of the model with the highest drug prediction accuracy in the test data is selected as the final result) to obtain the bladder cancer chemotherapy drug sensitivity prediction results of 251 bladder cancer patients.

[0100] 6. The consistency of the predicted bladder cancer chemotherapy drug sensitivity results and the actual drug response types of 251 bladder cancer patients was compared. The comparison results of the clinical sample and the PTC sample are shown in Table 4 and Table 5, respectively:

[0101] Table 4 Comparison of Consistency of Clinical Samples

[0102] Giscitabine Paclitaxel Dorsey answer(%) 86.7 84.9 85 Drug resistance (%) 86 86 84.1 Average accuracy (%) 86.35 85.45 84.55

[0103] Table 5. Comparison of PTC Sample Consistency

[0104] Giscitabine Paclitaxel Dorsey answer(%) 87.5 85 84 Drug resistance (%) 88.4 82.5 87 Average accuracy (%) 87.95 83.75 85.5

[0105] The accuracy of the bladder cancer chemotherapy drug sensitivity prediction model was tested using clinical data and PTC (partial clinical triglyceride) data from 251 cases. Comparative analysis showed that the model performed well in predicting bladder cancer chemotherapy drug sensitivity (gemcitabine, paclitaxel, and docetaxel). Specifically, in the clinical sample, the accuracy was 86.35% for gemcitabine, 85.45% for paclitaxel, and 84.55% for docetaxel. In patients who underwent PTC testing, the accuracy was 87.95% for gemcitabine, 83.75% for paclitaxel, and 85.5% for docetaxel.

[0106] The present invention provides a method for constructing a bladder cancer chemotherapy drug sensitivity prediction model and its application, which has at least the following beneficial effects:

[0107] (1) High accuracy;

[0108] The results in Tables 4 and 5 show that the accuracy of the bladder cancer chemotherapy drug sensitivity prediction model constructed in this invention is very high (>80%), whether compared with clinical drug use data or drug sensitivity data obtained from PTC.

[0109] (2) Shorter cycle, more economical and efficient;

[0110] From obtaining tumor samples from bladder cancer patients, predictive results can be obtained within 7 days (including sequencing, analysis, and prediction), while obtaining drug sensitivity data through PTC takes 2 weeks. As sequencing costs decrease year by year, the cost of drug sensitivity testing based on gene expression data for patients will also decrease year by year.

[0111] (3) The modeling data used is clinical data specifically for Chinese people, which has ethnic specificity and high consistency of clinical effects.

[0112] (4) The prediction results are simple and clear, divided into two types: response and drug resistance, which is convenient for use as an adjunct to medication guidance.

[0113] The following describes the bladder cancer chemotherapy drug sensitivity prediction system provided by the present invention. The bladder cancer chemotherapy drug sensitivity prediction system described below can be referred to in correspondence with the bladder cancer chemotherapy drug sensitivity prediction method described above.

[0114] Reference Figure 2 The present invention provides a bladder cancer chemotherapy drug sensitivity prediction system, which may include:

[0115] The data receiving module is used to receive data on the type of chemotherapy drug to be administered to the subject of bladder cancer and data on the level of immune cell infiltration in tumor sample tissue sent by at least one client, wherein the subject of the test is a bladder cancer patient;

[0116] The prediction module is used to: obtain the bladder cancer chemotherapy drug sensitivity prediction result of the test subject by combining at least two bladder cancer chemotherapy drug sensitivity prediction models obtained by the construction method of the bladder cancer chemotherapy drug sensitivity prediction model described above with hard voting, based on the data of the type of bladder cancer chemotherapy drug to be administered to the test subject and the data of immune cell infiltration level of tumor sample tissue.

[0117] The data sending module is used to send the predicted results of the bladder cancer chemotherapy drug sensitivity of the test subject back to at least one client.

[0118] In one implementation, the prediction module includes:

[0119] The prediction submodule is used to: obtain at least two bladder cancer chemotherapy drug sensitivity prediction results for the test subject based on the data of the type of bladder cancer chemotherapy drug to be administered to the test subject and the data of immune cell infiltration level of tumor sample tissue, respectively, through at least two bladder cancer chemotherapy drug sensitivity prediction models obtained by the construction method of bladder cancer chemotherapy drug sensitivity prediction model described above.

[0120] The determination submodule is used to: obtain the final bladder cancer chemotherapy drug sensitivity prediction result for the test subject based on at least two bladder cancer chemotherapy drug sensitivity prediction results and a hard voting method. The hard voting method is as follows: when there is a unique bladder cancer chemotherapy drug sensitivity prediction result with the most votes among the at least two bladder cancer chemotherapy drug sensitivity prediction results for the test subject, that unique bladder cancer chemotherapy drug sensitivity prediction result with the most votes is used as the final bladder cancer chemotherapy drug sensitivity prediction result for the test subject; when there is no unique bladder cancer chemotherapy drug sensitivity prediction result with the most votes among the at least two bladder cancer chemotherapy drug sensitivity prediction results for the test subject, the bladder cancer chemotherapy drug sensitivity prediction result of the bladder cancer chemotherapy drug sensitivity prediction model with the highest accuracy corresponding to the bladder cancer chemotherapy drug type is used as the final bladder cancer chemotherapy drug sensitivity prediction result for the test subject.

[0121] Figure 3 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 3 As shown, the electronic device may include a processor 810, a communications interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communications interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute the following steps:

[0122] Receive data from at least one client on the type of chemotherapy drug to be administered to the subject for bladder cancer and data on the level of immune cell infiltration in the tumor sample tissue, wherein the subject is a bladder cancer patient;

[0123] Based on the data of the type of bladder cancer chemotherapy drug to be administered to the test subject and the data of immune cell infiltration level of tumor sample tissue, at least two bladder cancer chemotherapy drug sensitivity prediction models obtained by the construction method of bladder cancer chemotherapy drug sensitivity prediction model described above are combined with the hard voting method to obtain the bladder cancer chemotherapy drug sensitivity prediction result of the test subject.

[0124] The predicted sensitivity of bladder cancer chemotherapy drugs for the test subjects is sent back to at least one client.

[0125] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0126] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program, the computer program being able to be stored on a non-transitory computer-readable storage medium, and the computer program being executed by a processor, enabling the computer to perform the following steps:

[0127] Receive data from at least one client on the type of chemotherapy drug to be administered to the subject for bladder cancer and data on the level of immune cell infiltration in the tumor sample tissue, wherein the subject is a bladder cancer patient;

[0128] Based on the data of the type of bladder cancer chemotherapy drug to be administered to the test subject and the data of immune cell infiltration level of tumor sample tissue, at least two bladder cancer chemotherapy drug sensitivity prediction models obtained by the construction method of bladder cancer chemotherapy drug sensitivity prediction model described above are combined with the hard voting method to obtain the bladder cancer chemotherapy drug sensitivity prediction result of the test subject.

[0129] The predicted sensitivity of bladder cancer chemotherapy drugs for the test subjects is sent back to at least one client.

[0130] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:

[0131] Receive data from at least one client on the type of chemotherapy drug to be administered to the subject for bladder cancer and data on the level of immune cell infiltration in the tumor sample tissue, wherein the subject is a bladder cancer patient;

[0132] Based on the data of the type of bladder cancer chemotherapy drug to be administered to the test subject and the data of immune cell infiltration level of tumor sample tissue, at least two bladder cancer chemotherapy drug sensitivity prediction models obtained by the construction method of bladder cancer chemotherapy drug sensitivity prediction model described above are combined with the hard voting method to obtain the bladder cancer chemotherapy drug sensitivity prediction result of the test subject.

[0133] The predicted sensitivity of bladder cancer chemotherapy drugs for the test subjects is sent back to at least one client.

[0134] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0135] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for constructing a training set, comprising: obtaining whole transcriptome sequencing data of tumor sample tissues of a target population, administered bladder cancer chemotherapy drug type data, and bladder cancer chemotherapy drug sensitivity data, wherein the target population is a bladder cancer patient population, and the bladder cancer chemotherapy drug sensitivity data comprises bladder cancer chemotherapy drug response data and bladder cancer chemotherapy drug resistance data; obtaining gene expression data of the tumor sample tissues of the target population from the whole transcriptome sequencing data of the tumor sample tissues of the target population; obtaining immune cell infiltration level data of the tumor sample tissues of the target population from the gene expression data of the tumor sample tissues of the target population; constructing a training set using the administered bladder cancer chemotherapy drug type data of the target population and the immune cell infiltration level data of the tumor sample tissues of the target population as sample data, and using the bladder cancer chemotherapy drug sensitivity data of the target population as label data.

2. The method of claim 1, wherein, The method further comprises: obtaining reference genome data; performing data quality control on the whole transcriptome sequencing data of the tumor sample tissues of the target population; aligning the data quality controlled whole transcriptome sequencing data with the reference genome data to obtain a bam file; calculating gene expression from the bam file to obtain the gene expression data of the tumor sample tissues of the target population.

3. The method of claim 2, wherein the training set is constructed by: The method further comprises: performing data quality control on the whole transcriptome sequencing data of the tumor sample tissues of the target population according to a preset data quality control condition, wherein the preset data quality control condition comprises: filtering out reads with a content of base N exceeding a preset content threshold; filtering out reads with a sequencing quality value lower than a preset quality value threshold.

4. The method of claim 3, wherein the training set is constructed by, The method further comprises: obtaining a plurality of cell infiltration level evaluation results from the gene expression data of the tumor sample tissues of the target population by a plurality of cell infiltration level evaluation methods, and then obtaining the immune cell infiltration level data of the tumor sample tissues of the target population, wherein the plurality of cell infiltration level evaluation methods comprise any one or any combination of the following: CIBERSORT, CIBERSORT-ABS, EPIC, MCPCOUNTER, QUANTISEQ, TIMER, and XCELL.

5. The method of claim 4, wherein the training set is constructed by, The method further comprises: performing data cleaning on the training set according to a preset cleaning rule, wherein the preset cleaning rule comprises any one or any combination of the following: for each piece of immune cell infiltration level data in the immune cell infiltration level data of the tumor sample tissues of the target population, if the value of > 60% of the immune cell infiltration level evaluation results is 0, then delete the piece of immune cell infiltration level data. When the immune cell infiltration level evaluation results obtained by a certain cell infiltration level evaluation method are all 0 in ≥80% of the values in the immune cell infiltration level evaluation results, the immune cell infiltration level evaluation results obtained by the certain cell infiltration level evaluation method are deleted; When the bladder cancer chemotherapy drug type data to be administered by the subject indicates the same bladder cancer chemotherapy drug, if the bladder cancer chemotherapy drug sensitivity data of the subject indicates that the ratio of bladder cancer chemotherapy drug response and bladder cancer chemotherapy drug resistance is outside the preset ratio range, then part of the bladder cancer chemotherapy drug sensitivity data is randomly deleted until the ratio of bladder cancer chemotherapy drug response and bladder cancer chemotherapy drug resistance is within the preset range.

6. A method for constructing a bladder cancer chemotherapeutic drug sensitivity prediction model, characterized by, The method comprises the following steps: obtaining a training set obtained by the construction method of the training set according to any one of claims 1-5; training a plurality of candidate models by using a plurality of machine learning algorithms respectively according to the training set; selecting the plurality of candidate models according to a preset model index to obtain at least two final candidate models as at least two bladder cancer chemotherapy drug sensitivity prediction models.

7. The method of claim 6, wherein the bladder cancer chemotherapy drug sensitivity prediction model is constructed by using the bladder cancer chemotherapy drug sensitivity prediction model construction program. The at least two final candidate models include four final candidate models, wherein the four final candidate models are candidate models trained by EasyEnsembleClassifier, BalancedRandomForestClassifier, BalancedBaggingClassifier, and RUSBoostClassifier learning algorithms respectively according to the training set.

8. A bladder cancer chemotherapeutic drug sensitivity prediction system characterized by, The method comprises the following steps: The data receiving module is configured to receive the bladder cancer chemotherapy drug type data to be administered by the subject and the immune cell infiltration level data of the tumor sample tissue sent by at least one client, wherein the subject is a bladder cancer patient; The prediction module is configured to obtain the bladder cancer chemotherapy drug sensitivity prediction result of the subject by using the at least two bladder cancer chemotherapy drug sensitivity prediction models obtained by the construction method of the bladder cancer chemotherapy drug sensitivity prediction model according to claim 6 or 7 and combining the hard voting method according to the bladder cancer chemotherapy drug type data to be administered by the subject and the immune cell infiltration level data of the tumor sample tissue of the subject. The data sending module is configured to send the bladder cancer chemotherapy drug sensitivity prediction result of the subject back to the at least one client.

9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor performs the following steps: The data receiving module is configured to receive the bladder cancer chemotherapy drug type data to be administered by the subject and the immune cell infiltration level data of the tumor sample tissue sent by at least one client, wherein the subject is a bladder cancer patient; The prediction module is configured to obtain the bladder cancer chemotherapy drug sensitivity prediction result of the subject by using the at least two bladder cancer chemotherapy drug sensitivity prediction models obtained by the construction method of the bladder cancer chemotherapy drug sensitivity prediction model according to claim 6 or 7 and combining the hard voting method according to the bladder cancer chemotherapy drug type data to be administered by the subject and the immune cell infiltration level data of the tumor sample tissue of the subject. The bladder cancer chemotherapy drug sensitivity prediction result of the to-be-tested person is sent back to the at least one client. 10.A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program performs the following steps by the processor: Receiving the type data of the bladder cancer chemotherapy drug to be administered to the to-be-tested person and the immune cell infiltration level data of the tumor sample tissue sent by the at least one client, wherein the to-be-tested person is a bladder cancer patient; According to the type data of the bladder cancer chemotherapy drug to be administered to the to-be-tested person and the immune cell infiltration level data of the tumor sample tissue, at least two bladder cancer chemotherapy drug sensitivity prediction models obtained by the construction method of the bladder cancer chemotherapy drug sensitivity prediction model of claim 6 or 7 are combined with the hard voting method to obtain the bladder cancer chemotherapy drug sensitivity prediction result of the to-be-tested person; The bladder cancer chemotherapy drug sensitivity prediction result of the to-be-tested person is sent back to the at least one client.

Citation Information

Patent Citations

  • Drug sensitivity prediction method, electronic equipment and computer readable storage medium

    CN112951327A

  • Tumor drug sensitivity prediction method, system and equipment and storage medium

    CN115966316A