An esophageal cancer personalized treatment decision-making method and system, and a storage medium containing the same
Patent Information
- Application Number
- CN202211542915.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-02
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2042-12-02
AI Technical Summary
[0054]本发明提供了疾病的个性化治疗决策方法及执行该决策方法的决策系统,所述决策系统包括:食管癌治疗预测模型的构建装置、模型比较模块和方案决策模块,通过结合机器算法构建治疗反应的预测模型,可向待治疗患者推荐适合其的个性化治疗方案,并在一定程度上实现个性化治疗。
Smart Images

Figure CN116434962B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of personalized treatment of human diseases, and relates to a predictive model and personalized treatment decision-making method for esophageal cancer based on proteome panel typing and constructed through machine learning algorithms, as well as a computer-readable storage medium and electronic device containing the same. Background Technology
[0002] Esophageal cancer (EC) is a common and deadly cancer with a poor prognosis and high mortality rate. It is the eighth most common cancer worldwide and the sixth leading cause of cancer-related death globally. Early detection rates for esophageal cancer are low, and late-stage cure rates are also low, with a 5-year survival rate of less than 19%. Surgery is the preferred treatment for esophageal cancer patients. For some patients who cannot undergo surgery, medication is the only option to improve their quality of life or even maintain survival. Although significant progress has been made in recent years in the iterative updates of first-line treatments, the development of second-line treatments, and the targeted immunosuppressant pembrolizumab (Keytruda), drug tolerance remains a problem, and the overall prognosis for esophageal cancer treatment remains poor. Clinically, there are significant individual differences in the effectiveness of cancer treatment, but there is insufficient evidence for personalized drug selection. There is an urgent clinical need for biomarkers to guide personalized precision medicine and alleviate the problem of drug tolerance.
[0003] The development of genomics and transcriptomics technologies has advanced the research and large-scale discovery of mutation patterns in tumorigenesis-related genes and important driver genes, driving the emergence and development of precision medicine. However, exploring tumor precision diagnosis and treatment solely at the gene level cannot meet clinical needs. Proteomics research elucidates the causes of specific biological phenomena at the protein level, revealing developmental patterns. Multidimensional omics, with the proteome at its core, is of great significance to the development of life science research and precision medicine.
[0004] According to reports, current research mainly involves identifying characteristic genes and signaling networks of tumors by integrating genomic datasets, and genome-based molecular classification systems have been proposed for different human populations. In January 2017, the Cancer Genome Atlas (TCGA) project mapped the genomes of 164 esophageal cancer patients and classified esophageal cancer into three subtypes: ① ESCC1; ② ESCC2; ③ ESCC3. In May of the same year, BGI Genomics and other institutions jointly published an article on esophageal cancer subtyping, classifying 360 cases of esophageal squamous cell carcinoma into four subtypes based on transcriptome data: ① well-differentiated ESC1; ② metastasis-associated ESCC2; ③ moderately modified ESCC3; ④ chromosomally unstable positive (CIN+) ESCC4. Different molecular subtypes have different molecular characteristics and are associated with different clinical outcomes. In the past few years, many similar studies have been conducted through large-scale cancer genome sequencing, showing that esophageal cancer is divided into molecularly genetically heterogeneous subgroups, rather than a single type of cancer. Therefore, to achieve personalized treatment for esophageal cancer, it is necessary to identify subtypes based on molecular genetic and pathological characteristics and discover and apply corresponding target genes. Furthermore, esophageal cancer research has reported results demonstrating the ability to classify the prognosis of esophageal cancer based on its subtypes. Several research patents have emerged that focus on early screening for esophageal cancer based on gene expression levels in the genome and transcriptome, such as esophageal cancer prognostic biomarkers and their applications (patent number CN106701992A) and esophageal cancer diagnostic and treatment biomarkers (patent number CN105886627B). However, in clinical practice, first-line treatment regimens for esophageal cancer patients, such as platinum-based chemotherapy plus paclitaxel and chemotherapy combined with radiotherapy (platinum-based chemotherapy plus paclitaxel combined with radiotherapy), do not benefit all patients. Cancer treatment tolerance remains a critical issue that needs to be addressed. Currently, there is still a lack of methods to screen patients for benefits from different treatment regimens, thus hindering personalized treatment for cancer patients.
[0005] As "executors of life," proteins determine phenotypes and may bridge the gap between research and clinical practice. Currently, at the protein level, molecular subtyping of diffuse gastric cancer, protein biomarkers for subtyping, and their screening methods and applications have been achieved (patent number CN 108445097 A), suggesting a strong correlation between protein-based molecules and patient survival and treatment response. However, personalized treatment plans for esophageal cancer and the identification of esophageal cancer patients who can benefit from different treatment options are still lacking. Summary of the Invention
[0006] To address the lack of predictive models for esophageal cancer treatment in existing technologies, thereby failing to effectively provide computer-aided personalized treatment decision-making solutions, this invention provides a personalized treatment decision-making method, system, and storage medium containing the same for esophageal cancer.
[0007] This invention utilizes proteomics technology combined with machine learning algorithms to construct predictive models for different treatment responses in esophageal cancer based on protein expression data from treated patients. These predictive models cover current chemotherapy and radiotherapy regimens for the digestive tract, helping clinicians to develop personalized drug treatment plans based on the degree of agreement between protein expression data from untreated patients and the predictive models, thereby maximizing patient benefits and laying the foundation for clinical drug selection.
[0008] To solve the above-mentioned technical problems, one of the technical solutions of the present invention is: to provide a method for constructing an esophageal cancer treatment prediction model, which includes the following steps:
[0009] (1) Divide the sample groups: Obtain protein expression data and clinical treatment response information of patients treated with different treatment regimens before treatment, and divide the samples into sensitive groups and non-sensitive groups; screen the differentially expressed proteins DEP in the sensitive group with a protein expression level difference of more than 2 times that of the non-sensitive group and p<0.05, and input them into the matrix;
[0010] (2) Feature protein selection: The Akaike Information Criterion (AIC) of DEP described in step (1) is calculated by using multivariate binary stepwise logistic regression. The group of DEP with the smallest AIC value is selected as the feature protein of the treatment regimen.
[0011] (3) Constructing a prediction model: Randomly select at least 50% of the clinical samples from each treatment plan, use the corresponding feature proteins as the training set and the feature proteins corresponding to the remaining clinical samples as the test set, and use multivariate binary classification stepwise logistic regression to cross-validate the training set. The cross-validation is 10-fold cross-validation; thus, the prediction model of the treatment plan is obtained.
[0012] In some preferred embodiments:
[0013] In step (2), the selection of the characteristic protein is performed using R Studio software; the pROC data package is used, and the family function is set to binomial in the glm function and the direction function is set to 'backward' in the step function; and / or,
[0014] In step (3), the at least 50% of the samples is at least 60%, at least 70%, or at least 80% of the samples, and the 10-fold cross-validation is repeated 10 times; and / or,
[0015] The clinical treatment response information includes treatment efficacy evaluated using conventional X-ray, CT scans, and / or MRI scans; and / or,
[0016] The patients were grouped based on RECIST criteria, wherein the sensitive group included complete remission and partial remission, and the non-sensitive group included stable disease and disease progression.
[0017] In some more preferred embodiments: the treatment regimen is selected from chemotherapy and radiotherapy regimens; preferably,
[0018] In step (1), based on the protein expression data and clinical treatment response information of pre-treatment samples from cancer patients treated with chemotherapy and radiotherapy regimens, the samples are divided into chemotherapy-sensitive group CTSG, chemotherapy-insensitive group CTNSG, radiotherapy-sensitive group CRTSG, and radiotherapy-insensitive group CRTNSG. Differentially expressed proteins DEP with a protein expression level difference of more than 2-fold and p<0.05 between the sensitive group and the insensitive group are screened and input into the matrix.
[0019] In step (2), the AIC value of DEP for different treatment regimens is calculated, and the group of DEPs with the smallest AIC value is selected as the characteristic protein of a certain treatment regimen. The multivariate binary classification stepwise logistic regression method is used to establish the prediction model for chemotherapy and radiotherapy regimens respectively.
[0020] In step (3), 80% of the clinical samples in each of the chemotherapy and radiotherapy regimens are randomly selected as the training set and the remaining clinical samples are used as the validation set. 10-fold cross-validation is performed and repeated 10 times to obtain the prediction model for the chemotherapy and radiotherapy regimens.
[0021] In some further preferred embodiments: obtaining the protein expression data includes the following steps:
[0022] (A) Sample preparation: The sample is prepared into formalin-fixed paraffin-embedded tissue sections, and / or, the esophageal cancer cells in the sample account for more than 80% of the total cell number; and / or, the lysis buffer used for the sample is 0.1M Tris-HCl, pH 8.0, with 0.1M DTT and 1mM PMSF added; and / or, the protein in the sample is enzymatically digested with 50mM NH4HCO3 containing trypsin at 37°C for 18-20 hours; and / or, the enzymatically digested peptides are collected by centrifugation; and / or,
[0023] (B) Mass spectrometry detection: The peptide was detected using a Q-Exactive HF-X hybrid quadrupole orbital trap mass spectrometer and a high-performance liquid chromatography system to obtain the corresponding mass spectrometry data; and / or, data acquisition was controlled using Xcalibur software; and / or,
[0024] (C) Data processing: The collected data were processed using the Firmiana database and MaxQuant software; and / or, the first search quality tolerance was 20 ppm and the major search peptide tolerance was 0.5 da; and / or, the calculation method was label-free iBAQ, and FOT was used to represent the normalized abundance of proteins in the sample; and / or, proteins with at least one exclusive peptide and an FDR of less than 1% were selected.
[0025] To solve the above-mentioned technical problems, a second technical solution of the present invention is: to provide a decision-making method for personalized esophageal cancer treatment, which includes the following steps:
[0026] (a) Obtaining prediction models for different treatment regimens using the method for constructing a prediction model for esophageal cancer treatment as described in one of the technical solutions of this invention;
[0027] (b) Extract protein expression data from clinical samples of patients to be treated, combine them with prediction models of different treatment options, use the predict function of R Studio software with parameter set type="prob" to calculate the prediction probability of each treatment option, and recommend the treatment option with the highest prediction probability to the patient.
[0028] To solve the above-mentioned technical problems, the third technical solution of the present invention is: to provide a device for constructing a predictive model for esophageal cancer treatment, comprising:
[0029] Sample grouping module: This module obtains protein expression data and clinical treatment response information of pre-treatment samples from patients treated with different regimens, and divides the samples into sensitive and non-sensitive groups. It then filters proteins whose expression level in the sensitive group is more than twice that of the non-sensitive group (p < 0.05), i.e., S protein; and proteins whose expression level in the non-sensitive group is more than twice that of the sensitive group (p < 0.05), i.e., NS protein, and inputs these values into a matrix; and,
[0030] Feature protein selection module: The feature protein selection module calculates the Akaike Information Criterion (AIC) of the DEPs described in step (1) using a multivariate binary stepwise logistic regression method, and obtains the group of DEPs with the smallest AIC value as the feature proteins of this treatment regimen; and,
[0031] Prediction model construction module: The prediction model construction module randomly selects at least 50% of the clinical samples from each treatment plan, uses the corresponding feature proteins as the training set, and uses the feature proteins corresponding to the remaining clinical samples as the test set. Multivariate binary classification stepwise logistic regression is used to cross-validate the training set. The cross-validation is 10-fold cross-validation and repeated 10 times; thus, the prediction model of the treatment plan is obtained.
[0032] In some more preferred embodiments,
[0033] In the prediction model building module, the AUC value is calculated using R Studio software; the pROC data package is used, and the family in the glm function is set to binomial, and the direction in the step function is set to 'backward'; and / or,
[0034] In the optimal prediction model training module, at least 50% of the samples are at least 60%, at least 70%, or at least 80% of the samples, and the 10x cross-validation is repeated 10 times.
[0035] Clinical treatment response information includes treatment efficacy evaluated using conventional X-ray, CT scans, and / or MRI scans; and / or,
[0036] The patients were grouped based on RECIST criteria, wherein the sensitive group included complete remission and partial remission, and the non-sensitive group included stable disease and disease progression.
[0037] In some preferred embodiments, the treatment regimen is selected from chemotherapy and radiotherapy regimens; preferably,
[0038] In the sample grouping module, based on the protein expression data and clinical treatment response information of pre-treatment samples from cancer patients treated with chemotherapy and radiotherapy regimens, the samples are divided into chemotherapy-sensitive group, chemotherapy-insensitive group, radiotherapy-sensitive group, and radiotherapy-insensitive group. Characteristic proteins of each group are screened to obtain the difference fold of more than 2 and p < 0.05, i.e., DEP, and input into the matrix.
[0039] In the feature protein selection module, the DEP obtained by the sample grouping module is processed using a multivariate binary classification stepwise logistic regression method to calculate the AIC value of DEP for different treatment regimens. The group of DEPs with the smallest AIC value is selected as the feature protein of a certain treatment regimen, thereby obtaining the feature proteins of chemotherapy and radiotherapy regimens.
[0040] In the prediction model building module, 80% of the clinical samples from each chemotherapy and radiotherapy regimen are randomly selected. The corresponding feature proteins are used as the training set, and the feature proteins corresponding to the remaining clinical samples are used as the test set. Multivariate binary classification stepwise logistic regression is applied to the training set for 10-fold cross-validation, and this is repeated 10 times to obtain the prediction model for chemotherapy and radiotherapy regimens.
[0041] In some further preferred embodiments, the construction apparatus further includes:
[0042] Sample preparation module: The sample preparation module prepares formalin-fixed paraffin-embedded tissue sections, wherein esophageal cancer cells account for more than 80% of the total cell count in the samples; the lysis buffer used for sample lysis is 0.1M Tris-HCl, pH 8.0, with 0.1M DTT and 1mM PMSF added; the proteins in the samples are enzymatically digested using 50mM NH4HCO3 containing trypsin, incubated at 37°C for 18-20 hours, and the digested peptides are collected by centrifugation; and / or,
[0043] Mass spectrometry detection module: The peptide is detected using a Q-Exactive HF-X hybrid quadrupole orbital trap mass spectrometer and a high-performance liquid chromatography system, and the corresponding mass spectrometry data is obtained; data acquisition is controlled using Xcalibur software; and / or,
[0044] Data processing module: The data collected by the data processing module is processed using the Firmiana database and MaxQuant software; the first search quality tolerance is 20 ppm, and the main search peptide tolerance is 0.5 da; the calculation method is label-free iBAQ, and FOT is used to represent the normalized abundance of proteins in the sample; proteins with at least one exclusive peptide and an FDR of less than 1% are selected.
[0045] To solve the above-mentioned technical problems, the fourth technical solution of the present invention is: to provide a personalized esophageal cancer treatment decision-making system, which includes:
[0046] The apparatus for constructing an esophageal cancer treatment prediction model as described in technical solution three of the present invention; and
[0047] Treatment Decision Module: The treatment decision module extracts protein expression data from clinical samples of patients to be treated (e.g., obtained through mass spectrometry detection), combines it with prediction models of different treatment options, and uses the predict function of Rstudio software with the parameter set to type="prob" to calculate the prediction probability sensitive to each treatment option, and recommends the treatment option with the highest prediction probability to the patient.
[0048] To solve the above-mentioned technical problems, the fifth technical solution of the present invention is: to provide an electronic device, which includes a memory and a processor; the memory includes a computer program stored therein that can run on the processor; wherein,
[0049] When the processor executes the computer program, it implements the method for constructing an esophageal cancer treatment prediction model as described in one of the technical solutions of the present invention, the decision-making method for personalized esophageal cancer treatment as described in another of the technical solutions of the present invention, or the method for constructing an esophageal cancer treatment prediction model as described in a third of the technical solutions of the present invention.
[0050] To solve the above-mentioned technical problems, the sixth technical solution of the present invention is: providing a computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, can implement the steps of the method for constructing an esophageal cancer treatment prediction model as described in the first technical solution of the present invention, the steps of the decision-making method for personalized esophageal cancer treatment as described in the second technical solution of the present invention, or the steps of the method for constructing an esophageal cancer treatment prediction model as described in the third technical solution of the present invention.
[0051] Based on common knowledge in the field, the above-mentioned preferred conditions can be combined arbitrarily to obtain various preferred embodiments of the present invention.
[0052] The reagents, raw materials, equipment and drugs used in this invention are all commercially available.
[0053] The positive and progressive effects of this invention are as follows:
[0054] This invention provides a personalized treatment decision-making method for diseases and a decision-making system for implementing the method. The decision-making system includes: a device for constructing a predictive model for esophageal cancer treatment, a model comparison module, and a treatment plan decision-making module. By combining machine algorithms to construct a predictive model of treatment response, it can recommend suitable personalized treatment plans to patients to be treated and achieve personalized treatment to a certain extent. Attached Figure Description
[0055] Figure 1 This is a flowchart of the personalized tumor treatment decision-making method of the present invention;
[0056] Figure 2 A schematic diagram of the modules of a personalized cancer treatment decision-making system;
[0057] Figure 3 This is a schematic diagram of the structure of the electronic device according to Embodiment 4 of the present invention;
[0058] Figure 4 Volcano plot of differentially expressed proteins CTSG and CTNSG in a chemotherapy cohort;
[0059] Figure 5 Volcano plot of differentially expressed CRTSG and CRTNSG proteins in a radiotherapy / chemotherapy cohort;
[0060] Figure 6 Process for building a predictive model for esophageal cancer treatment;
[0061] Figure 7 A flowchart for building a model based on proteomics panel analysis and recommending personalized esophageal cancer treatment plans;
[0062] Figure 8 Develop a predictive model for chemotherapy cohorts;
[0063] Figure 9 A predictive model for a radiotherapy and chemotherapy cohort was constructed. Detailed Implementation
[0064] The present invention is further illustrated below by way of embodiments, but the invention is not limited to the scope of the embodiments described herein. Experimental methods in the following embodiments that do not specify specific conditions were performed according to conventional methods and conditions, or as selected according to the product instructions.
[0065] Example 1: Method for establishing a predictive model for esophageal cancer treatment based on proteomic data
[0066] The method 201 for constructing a tumor treatment prediction model of the present invention includes the following steps (such as...) Figure 1 As shown):
[0067] Step 101: Divide the sample groups: Collect protein expression data and clinical treatment response information of samples from patients treated with different treatment regimens before treatment, divide the samples into sensitive groups and non-sensitive groups, screen for differentially expressed proteins DEP whose protein expression level in the sensitive group is more than 2 times that in the non-sensitive group, or whose protein expression level in the non-sensitive group is more than 2 times that in the sensitive group and p<0.05, and input them into the matrix;
[0068] Step 102, Feature protein selection: The DEPs described in Step 101 are calculated using a multivariate binary stepwise logistic regression method. The Akaike Information Criterion (AIC) of the DEPs is then obtained, and the group of DEPs with the smallest AIC value is selected as the feature proteins of this treatment regimen.
[0069] Step 103: Construct a prediction model: Randomly select at least 50% of the clinical samples from each treatment plan, use the corresponding feature proteins obtained in Step 102 as the training set, and the feature proteins corresponding to the remaining clinical samples as the test set, and use multivariate binary classification stepwise logistic regression to cross-validate the training set; thus, the prediction model of the treatment plan is obtained.
[0070] This embodiment uses esophageal cancer as an example. In step 101, obtaining the protein expression data includes the following steps:
[0071] I. Esophageal Cancer Sample Collection
[0072] The esophageal cancer samples required for this invention were all provided by the Department of Pathology, Zhongshan Hospital, Fudan University: 150 esophageal cancer patients, including a platinum-based combined with paclitaxel subcohort (75 patients treated with platinum-based combined with paclitaxel), a platinum-based plus paclitaxel combined with radiotherapy subcohort (55 patients treated with platinum-based plus paclitaxel combined with radiotherapy), and other subcohorts (7 patients treated with radiotherapy and 13 different chemotherapy regimens). All treatment regimens were administered at the standard dose for first-line treatment of esophageal cancer.
[0073] For evaluating the efficacy of treatment for target lesions, conventional X-ray, CT scan, or MRI scan are commonly used methods, and other clinically acceptable diagnostic methods can also be used. Based on the currently recognized response evaluation criteria in solid tumors (RECIST), the efficacy evaluation of target lesions can be divided into complete response (CR), partial response (PR), stable disease (SD), or progressive disease (PD). Objective response rate (ORR), as an indicator of tumor response efficacy, refers to the proportion of patients whose tumors shrink to a certain extent and remain so for a certain period, including cases of complete response (CR) and partial response (PR). Clinically, the Food and Drug Administration (FDA) defines the objective response rate (ORR) as the sum of PR and CR, which can directly measure the antitumor activity of the drug.
[0074] Based on RECIST criteria, esophageal cancer patients were divided into 56 sensitive patients (sum of CR and PR) and 94 non-sensitive patients (sum of SD and PD). In the platinum-based combined paclitaxel subcohort, patients were divided into 14 sensitive patients (CR, 4; PR, 10, designated CTSG) and 61 non-sensitive patients (SD, 60; PD, 1, designated CTNSG); in the platinum-based plus paclitaxel combined with radiotherapy subcohort, patients were divided into 35 sensitive patients (CR, 20; PR, 15, designated CRTSG) and 20 non-sensitive patients (SD, 20; PD, 0, designated CRTNTSG); in other subcohorts, patients were classified into 7 sensitive patients (CR, 2; PR, 5) and 13 non-sensitive patients (SD, 11; PD, 2). There was no bias in case selection. All cases were staged according to the American Joint Committee on Cancer (AJCC) 7th edition staging system. With the approval of the hospital's ethics committee (B2019-200R), all patients provided written informed consent.
[0075] II. Preparation of esophageal cancer protein samples:
[0076] The clinical samples were formalin-fixed paraffin-embedded (FFPE) tissues.
[0077] Sample pretreatment: 3-10 μm thick sections were taken from FFPE blocks for macroscopic dissection, dewaxing with xylene, washing with ethanol, and air-drying to obtain white slides. Hematoxylin and eosin (H&E) slides were used as microscopic references for tumor examination. Slides with a tumor cell content exceeding 80% were included in the study, as assessed by two esophageal pathologists.
[0078] Protein and peptide extraction from samples: Equal volumes of FFPE tissue were collected in EP tubes, and lysis buffer (0.1 M Tris-HCl pH 8.0, 0.1 M DTT, 1 mM PMSF) was added. The mixture was then ground for 3 minutes. Sodium dodecyl sulfate (SDS) was added to a final concentration of 4%, and the mixture was incubated at 99°C and 1800 rpm for 2-2.5 hours. The supernatant was collected by centrifugation at 12,000 g for 5 minutes and added to an EP tube. Four volumes of acetone were added, and the mixture was incubated at -20°C for 4 hours or overnight. The supernatant was discarded by centrifugation at 4°C for 1 minute, and the precipitate was washed three times with cold acetone. The protein precipitate was then air-dried in a clean bench. The protein precipitate was reconstituted with 8 M Urea and 50 mM NH4HCO3 and added to a FASP tube. Urea was removed by repeated centrifugation with 50 mM NH4HCO3. 50 μL of a solution containing 5.5 μg trypsin was added to the precipitate. Add 50 mM NH4HCO3 to a FASP tube and incubate at 37°C for 18-20 hours for enzymatic digestion; centrifuge at 12,800 g for 15 minutes to collect peptides. To improve peptide yield, wash twice with 200 μL MS water and collect the peptides together; vacuum dry at 60°C to obtain the peptides required for mass spectrometry detection.
[0079] III. Mass spectrometry detection of esophageal cancer protein samples:
[0080] The peptide sample was detected using a Q-Exactive HF-X hybrid quadrupole orbital trap mass spectrometer (Thermo Fisher Scientific, Rockford, IL, USA) and a high-performance liquid chromatography system (EASY nLC 1200, Thermo Fisher), and the corresponding mass spectrometric data were obtained. The dried peptide sample was redissolved in solvent A (0.1% formic acid in water) and loaded onto a trap column (100 μm × 2 cm; particle size, 3 μm; pore size, ...). The particles were then separated on an analytical column (150 μm × 12 cm, particle size 1.9 μm; pore size...). The elution gradient was 5-35% mobile phase B (80% acetonitrile and 0.1% formic acid) at a flow rate of 600 nL / min for a total elution time of 75 min. MS analysis of the QE-HFX was performed using a single full scan (300-1400 m / z, resolution = 12000), with the maximum allowed number of ions in the ion trap (automatic gain control target, AGC target) set at 3E+06 ions. This was followed by 60 data correlation MS / MS scans on an orbital (resolution = 7500) using high-energy collision-induced dissociation (isolation window 1.6 m / z, collision energy 27%, AGC target 5E+04 ions, maximum injection time 30 ms, dynamic exclusion set to 18 seconds). Data acquisition was performed using the Xcalibur software (Thermo Scientific) on the liquid chromatography-tandem mass spectrometry system.
[0081] IV. Data Processing:
[0082] The original files were retrieved from the Refseq protein database of the National Center for Biotechnology Information (NCBI), using the Firmiana database (https: / / phenomics.fudan.edu.cn / firmiana / gardener / ) and MaxQuant software. Firmiana is a workflow based on the Galaxy system, consisting of multiple functional modules including a user login interface, raw data, identification and quantification, data analysis, and knowledge mining. Trypsin was selected as the proteolytic enzyme, with a maximum allowable two missed cleavage sites. The enzyme was fixed with carbamidomethyl (C) and dynamically modified with protein acetyl (protein N-term) and oxidation (M). The first search quality tolerance was 20 ppm, and the primary search peptide tolerance was 0.5 da. The false discovery rate (FDR) for peptide profiling (PSMs) and proteins was less than 1%. This invention employs label-free quantification, using a label-free, intensity-based absolute quantification (iBAQ) method. This study primarily uses fraction of total (FOT) to represent the normalized abundance of specific proteins in the sample. FOT is defined as the iBAQ of a protein divided by the total iBAQ of the entire experimental sample. Proteins with at least one unique peptide and an FDR of less than 1% are selected for further analysis.
[0083] Specifically, based on the quantitative proteomic data of the chemosensitive group (CTSG) and the chemosensitive group (CTNSG), 556 gene products (GPs) of DEP (expression level more than 2-fold and non-parametric test Wilcox p < 0.05) of CTSG relative to CTNSG were screened. Figure 4 Table 1). Quantitative proteomic data were collected from the chemoradiosensitive group (CRTSG) and the chemoradiosensitive group (CRTNSG). 147 proteoglobulins (GPs) with a DEP (expression level more than 2-fold higher than CRTNSG and a non-parametric test Wilcox p < 0.05) for CRTSG were screened. Figure 5 (Table 2)
[0084] Table 1: DEP in chemotherapy subcohorts
[0085]
[0086]
[0087]
[0088]
[0089]
[0090]
[0091]
[0092]
[0093] Table 2: DEP in the chemoradiotherapy subcohort
[0094]
[0095]
[0096]
[0097]
[0098]
[0099]
[0100] Based on the protein expression data and clinical treatment response information of the cancer patients who received chemotherapy and radiotherapy respectively before treatment, the samples were divided into chemotherapy-sensitive group (CTSG) and chemotherapy-insensitive group (CTNSG), radiotherapy-sensitive group (CRTSG) and radiotherapy-insensitive group (CRTNSG). The characteristic proteins of each group were screened to obtain CTSG, CTNSG, CRTSG and CRTNSG proteins respectively, and the protein profile input matrix for chemotherapy and radiotherapy was constructed.
[0101] V. Model Construction
[0102] Specifically, we used R Studio software, loaded the pROC data analysis package, and set 80% of the randomly selected samples from the input cohorts (chemotherapy subcohort and radiotherapy / chemotherapy subcohort) as the training set and 20% as the test set. For DEP, we applied multivariate binary logistic regression analysis, setting `family=binomial` in the `glm` function and `direction='backward'` in the `step` function. Based on the AIC values, we established combined models of chemotherapy and radiotherapy / chemotherapy regimens based on multivariate binary logistic regression. Figure 6 In the chemotherapy cohort, the DEP group with the lowest AIC, as a model feature protein, showed a prediction accuracy of 1 (95% confidence interval [CI]: 0.952-1) on the 80% training set, with both sensitivity and specificity of 100%; and a prediction accuracy of 1 (95% confidence interval [CI]: 0.782-1) on the 20% test set, with both sensitivity and specificity of 100%. Figure 8 In the chemoradiotherapy cohort, the DEP group with the lowest AIC, as a model feature protein, showed a prediction accuracy of 1 (95% confidence interval [CI]: 0.9169-1) on the 80% training set, with both sensitivity and specificity of 100%; and a prediction accuracy of 1 (95% confidence interval [CI]: 0.7151-1) on the 20% test set, with both sensitivity and specificity of 100%. Figure 9 ).
[0103] Example 2: Decision-making method for personalized tumor treatment
[0104] For patients with esophageal cancer awaiting treatment, personalized treatment plans are recommended based on predictive models (see...). Figure 7 Specifically, it includes the following steps ( Figure 1 ).
[0105] Step 201: Construct prediction models for different treatment options based on Example 1.
[0106] Step 202:
[0107] (1) Extracting proteomic data from samples of patients to be treated
[0108] Specifically, five patients with esophageal cancer awaiting treatment were identified as #1, #2, #3, #4, and #5. Their clinical samples were obtained as described in Example 1, and their proteomic data were obtained by mass spectrometry.
[0109] (2) Calculate the predicted probability of patients for different treatment options.
[0110] Protein expression data from five esophageal cancer patients awaiting treatment were input into a prediction model for conventional chemotherapy and conventional chemotherapy combined with targeted therapy, as described in Example 1. The prediction probability of sensitivity to each treatment regimen was calculated using the predict function in R Studio software with the parameter set to type="prob".
[0111] (3) Recommendations for treatment plans
[0112] Based on the predicted probabilities in Table 3 below, select the one with the highest predicted probability that is sensitive to the treatment plan, and recommend that treatment plan to the patient.
[0113] Table 3: Predictive outcomes and treatment decisions for 5 patients with esophageal cancer
[0114]
[0115] Example 3: Device for constructing a predictive model for esophageal cancer treatment and a personalized esophageal cancer treatment decision-making system
[0116] I. Construction Device for Esophageal Cancer Treatment Prediction Model
[0117] This embodiment provides a device 51 for constructing a predictive model for esophageal cancer treatment, such as... Figure 2 As shown, it includes: a sample grouping module 41, a feature protein selection module 42, and a prediction model construction module 43.
[0118] The sample grouping module 41 is used to collect protein expression data and clinical treatment response information of samples from patients treated with different treatment regimens before treatment. The samples are divided into sensitive groups and non-sensitive groups. S protein in the sensitive group is at least twice as much as that in the non-sensitive group and p < 0.05, and NS protein in the non-sensitive group is at least twice as much as that in the sensitive group and p < 0.05. The results are then input into a matrix.
[0119] The feature protein selection module 42 compares the DEPs described in the sample grouping module with the Akaike Information Criterion (AIC) of the DEPs using a multivariate binary stepwise logistic regression method. The DEPs with the smallest AIC value are then identified as the feature proteins of the treatment regimen.
[0120] The prediction model building module 43 randomly selects at least 80% of the clinical samples from each treatment plan, uses the corresponding feature proteins as the training set, and uses the feature proteins corresponding to the remaining 20% of the clinical samples as the validation set. Multivariate binary classification stepwise logistic regression is used to cross-validate the training set; thus, the prediction model of the treatment plan is obtained.
[0121] II. Personalized Esophageal Cancer Treatment Decision System
[0122] The personalized esophageal cancer treatment decision system includes: a device for constructing an esophageal cancer treatment prediction model 51 and a treatment plan decision module 52.
[0123] Information about the device 51 for constructing the treatment prediction model is provided in Part I; and,
[0124] The treatment plan decision module 52 extracts protein expression data from clinical samples of patients to be treated, combines them with prediction models of different treatment plans, and uses the predict function in R Studio software with the parameter set to type="prob" to calculate the predicted probability of the patient's sensitivity to each treatment plan, and recommends the treatment plan with the highest predicted probability to the patient.
[0125] Example 4 Electronic device
[0126] This embodiment provides an electronic device, which can be represented in the form of a computing device (e.g., a server device), including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it can implement the prediction model construction method in Embodiment 1, the specific model construction method in Embodiment 2, and the decision-making method for personalized tumor treatment.
[0127] Figure 3 This embodiment shows a hardware structure diagram, and the electronic device 9 specifically includes:
[0128] At least one processor 91, at least one memory 92, and a bus 93 for connecting different system components (including processor 91 and memory 92), wherein:
[0129] Bus 93 includes a data bus, an address bus, and a control bus.
[0130] The memory 92 includes volatile memory, such as random access memory (RAM) 921 and / or cache memory 922, and may further include read-only memory (ROM) 923.
[0131] The memory 92 also includes a program / utility 925 having a set (at least one) of program modules 924, including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.
[0132] The processor 91 executes various functional applications and data processing by running computer programs stored in the memory 92, such as the predictive model and personalized tumor treatment decision-making method in Embodiment 1 of the present invention.
[0133] Electronic device 9 can further communicate with one or more external devices 94 (e.g., keyboard, pointing device, etc.). This communication can be performed via input / output (I / O) interface 95. Furthermore, electronic device 9 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public network, such as the Internet) via network adapter 96. Network adapter 96 communicates with other modules of electronic device 9 via bus 93. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 9, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID (disk array) systems, tape drives, and data backup storage systems, etc.
[0134] It should be noted that although several units / modules or sub-units / modules of the electronic device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.
[0135] Example 5: Computer-readable storage medium
[0136] This invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the predictive model construction method in Embodiment 1, the specific model construction method in Embodiment 2, and the decision-making method for personalized tumor treatment.
[0137] The readable storage medium may be more specifically adopted, including but not limited to: portable disk, hard disk, random access memory, read-only memory, erasable programmable read-only memory, optical storage device, magnetic storage device, or any suitable combination thereof.
[0138] In a possible implementation, the present invention can also be implemented as a program product comprising program code, wherein when the program product is run on a terminal device, the program code is used to cause the terminal device to perform the steps of the method for constructing the prediction model in Embodiment 1 of the present invention, the steps of the method for constructing a specific model in Embodiment 2, and the steps of the decision-making method for personalized tumor treatment.
[0139] The program code for executing the present invention can be written in any combination of one or more programming languages. The program code can be executed entirely on the user device, partially on the user device, as a standalone software package, partially on the user device and partially on a remote device, or entirely on a remote device.
[0140] While specific embodiments of the present invention have been described above, those skilled in the art should understand that these are merely illustrative examples, and the scope of protection of the present invention is defined by the appended claims. Those skilled in the art can make various changes or modifications to these embodiments without departing from the principles and essence of the present invention, but all such changes and modifications fall within the scope of protection of the present invention.
Claims
1. A combination of biomarkers for predicting chemotherapy sensitivity in esophageal cancer, characterized in that, The biomarker combination consists of the following biomarkers: ABCB8, ACADSB, ACE, ACOX1, ACP1, ADAM15, AEBP1, ALDH1B1, ALOX15B, ANGPTL2, ANK3, ANP32B, AP1B1, APCS, APOM, APPL1, ARHGDIB, BRCC3, C10orf76, C11orf68, C19orf70, C1S, C2orf54, C3orf58, C4BPA, C5orf51, C6orf132, C8G, CD2BP2, CD63, CDC26, CDCP1, CES1, CETN2, CLPB, CLU. CNPY4, COA3, COL10A1, CPA3, CPSF3L, CRMP1, CTHRC1, CTSB, CTSZ, DCAKD, DDX54, DHCR7, DHRS9, DPYSL3, DPYSL4, EEA1, EFEMP1, EVPL, F10, F2, FABP6, FB LN1, FBN1, FTH1, GBP3, GEMIN5, GINS2, GK, GNAI1, GOSR2, GSTM3, H1FX, HDAC1, HIF1AN, HOOK1, HPSE, IGF2BP1, IGF2BP2, IGFALS, IRF2BP2, ITGA5, ITGB2, KBTBD11, KHNYN, KNG1, KRAS, LAMC1, LPCAT2, LPP, LRRC47, LSM4, LUM, LXN, LYZ, MAP1B, MCM5, MED27, MMP1, MMP2, MMP8, MOB2, MPG, MPHOSPH10, MREG, MRP S18C, MTA3, MXRA7, NAA20, NACA2, NBN, NCOR1, NDUFA11, NNMT, NOL9, NUDT2, OSBPL2, PADI3, PAK4, PARP9, PARVA, PDIA5, PFN2, PHACTR4, PLG, POLR2D, PPM 1F, PPP2R1B, PPP2R5C, PRELP, PRPF3, QPCTL, RABL6, RBL1, RPAP3, SAMD9, SERPINA5, SERPINC1, SERPIND1, SERPINE2, SERPINF1, SHROOM3, SIRT3, SLC12A 7. SLC1A3, SLTM, SMAD2, SMARCD2, SPIN1, SRSF4, STARD7, TAGLN, TALDO1, TFCP2, TIMP3, TMSB4X, TNIP2, TNS1, TRIM24, TUBB1, UBA7, UBE2I, UBE2M, UGDH,UROD, VTN, WASL, WDFY1 and YES1.
2. The use of the biomarker combination of claim 1 or the reagent for detecting the biomarker combination of claim 1 in the preparation of a formulation for predicting the chemosensitivity of esophageal cancer.
Citation Information
Patent Citations
Biomarkers for the diagnosis and treatment of esophageal cancer
CN105886627B
Esophageal carcinoma prognosis marker and application thereof
CN106701992A
Molecular typing of diffuse gastric cancer and protein markers for typing and screening method and application thereof
CN108445097A