Microbial marker combination for predicting colorectal cancer neoadjuvant chemotherapy sensitivity and application thereof
By using microbial marker combinations of spirochemo, Staphylococcus, Streptococcus, Facultative Bitococcus and Fusobacterium, combined with metagenomic sequencing and machine learning models, the problem of predicting sensitivity of neoadjuvant chemotherapy in colorectal cancer is solved, and precise personalized treatment plans are achieved, improving treatment effect and safety.
Patent Information
- Application Number
- CN202510861758.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-25
AI Technical Summary
The lack of reliable methods for predicting neoadjuvant chemotherapy sensitivity in colorectal cancer leads to difficulties in formulating personalized treatment options, which may lead to increased risk of treatment and poor treatment effectiveness.
The genus Treponema, Staphylococcus, Streptococcus, Gemella and Fusobacterium were used as microbial markers to analyze the composition of intestinal flora through metagenomic sequencing to construct a machine learning model to predict chemotherapy sensitivity.
Accurate prediction of the sensitivity of neoadjuvant chemotherapy in patients with colorectal cancer is achieved, providing personalized treatment basis, reducing treatment risks, and improving treatment effect.
Smart Images

Figure CN120350126A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of biological detection technology, and specifically relates to a combination of microbial markers for predicting the sensitivity of colorectal cancer to neoadjuvant chemotherapy and its application. Background Art
[0002] Colorectal cancer (CRC) is a malignant tumor that originates in the colon (large intestine) or rectum. It is one of the most common cancers in the world. The treatments for colorectal cancer currently include surgery, chemotherapy, radiotherapy, targeted therapy, and immunotherapy. Surgical treatment is the main treatment for colorectal cancer. Since advanced colorectal cancer is common, postoperative recurrence and progression are common. For patients in the middle and late stages, preoperative neoadjuvant therapy is conducive to observing the efficacy, and can effectively reduce the stage and improve the efficacy of radical surgery. Due to factors such as tumor heterogeneity, the efficacy of adjuvant chemotherapy for colorectal cancer varies from person to person. Patients who are insensitive to neoadjuvant chemotherapy may show severe toxic side effects after neoadjuvant therapy, or affect the subsequent surgery. Some patients may also have a poor prognosis. Therefore, in view of the differences in the response of different patients to neoadjuvant therapy, neoadjuvant therapy cannot be used uniformly, and personalized treatment plans need to be formulated in a targeted manner.
[0003] However, there is currently no reliable prediction method for colorectal neoadjuvant chemotherapy to screen out patients who are sensitive to neoadjuvant chemotherapy, formulate targeted treatment plans, reduce treatment risks, and improve treatment effects. Summary of the invention
[0004] In view of this, the primary purpose of this application is to provide a group of microbial marker combinations for predicting the sensitivity of colorectal cancer neoadjuvant chemotherapy, the microbial marker combination includes: Treponema, Staphylococcus, Streptococcus, Gemella and Fusobacterium. It can accurately predict whether colorectal cancer patients are sensitive to neoadjuvant chemotherapy, with high accuracy and good sensitivity, and provide a basis for personalized treatment methods.
[0005] In order to achieve the above objectives, this application adopts the following technical solutions: One aspect of the present application provides a combination of microbial markers for predicting the sensitivity of colorectal cancer to neoadjuvant chemotherapy, wherein the combination of microbial markers includes Treponema, Staphylococcus, Streptococcus, Gemella and Fusobacterium.
[0006] Another aspect of the present application provides the use of a reagent for detecting the abundance of the above-mentioned microbial biomarker combination in a sample in the preparation of a product for predicting the sensitivity of neoadjuvant chemotherapy for colorectal cancer.
[0007] Another aspect of the present application provides a product for predicting the sensitivity of neoadjuvant chemotherapy for colorectal cancer, wherein the product includes a reagent for detecting the abundance of the above-mentioned microbial biomarker combination in a sample.
[0008] Another aspect of the present application provides a model for predicting the sensitivity of neoadjuvant chemotherapy for colorectal cancer, wherein the input variable of the model is the abundance of the above-mentioned microbial biomarker combination in a sample.
[0009] Another aspect of the present application provides a system for predicting the sensitivity of neoadjuvant chemotherapy for colorectal cancer, including: A data input module, which is used to input the abundance data of the above-mentioned microbial biomarker combination; A data storage module, which is used to store the abundance data of the above-mentioned microbial biomarker combination in the biological samples of a population, and the population includes colorectal cancer patients receiving neoadjuvant chemotherapy; An output prediction module, which is respectively connected to the data input module and the data storage module, constructs a prediction model by using the abundance data of the above-mentioned microbial biomarker combination in the biological samples of the population, and obtains and outputs a prediction result of the sensitivity of colorectal cancer patients to neoadjuvant chemotherapy based on the abundance data of the microbial biomarker combination input by the data input module.
[0010] Another aspect of the present application provides a computer device, including a memory for storing a computer program and a processor for executing the computer program, and when the computer program is executed, the steps of the following method are implemented: Obtain the abundance data of the above-mentioned microbial biomarker combination in a sample; Construct a prediction model by using the abundance data of the above-mentioned microbial biomarker combination in the samples of a population, and the population includes colorectal cancer patients receiving neoadjuvant chemotherapy; Input the obtained abundance data of the microbial biomarker combination into the prediction model, and obtain and output a prediction result of the sensitivity of colorectal cancer patients to neoadjuvant chemotherapy.
[0011] Another aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the following method are implemented: Obtain the abundance data of the above-mentioned microbial biomarker combination in a sample; Construct a prediction model using the abundance data of the microbial biomarker combination in the samples of the population, where the population includes colorectal cancer patients who receive neoadjuvant chemotherapy; Input the obtained abundance data of the microbial biomarker combination into the prediction model, and obtain and output the prediction result of the sensitivity of colorectal cancer patients to neoadjuvant chemotherapy.
[0012] Advantages of the present application: By collecting samples of tumor tissues of colorectal cancer patients and analyzing the composition information of the intestinal flora, the present application finds that Treponema, Staphylococcus, Streptococcus, Gemella, and Fusobacterium have significant statistical differences between the sensitive group and the insensitive group. Therefore, the microbial biomarker combination provided in the present application can predict whether colorectal cancer patients are sensitive to neoadjuvant chemotherapy.
[0013] The present application constructs a machine learning model based on the obtained abundance data of the microbial biomarker combination to predict the sensitivity of colorectal cancer patients to neoadjuvant chemotherapy. Through experimental verification, the model has high accuracy and good sensitivity in predicting sensitivity. Therefore, it can clinically judge whether a patient is suitable for receiving neoadjuvant chemotherapy, select the best treatment method according to individual differences, and achieve personalized precision medicine. Description of the drawings
[0014] Figure 1 Presents the five characteristic microorganisms obtained after screening and decontamination by XGBoost in Example 3 and their characteristic importance scores in the random forest model.
[0015] Figure 2 Shows the confusion matrix of the random forest prediction model in Example 3 on the test set cases.
[0016] Figure 3 Is the ROC curve of the training set and the test set of the random forest model constructed in Example 4, where Figure 3 a in is the ROC curve of the training set, Figure 3 b in is the ROC curve of the test set. Detailed implementation manners
[0017] The embodiments of the present application are described in detail below. The embodiments described below are exemplary and are only used to explain the present application, and should not be construed as a limitation to the present application.
[0018] The first aspect of the present application provides a combination of microbial markers for predicting the sensitivity of neoadjuvant chemotherapy for colorectal cancer. The combination of microbial markers includes: Treponema, Staphylococcus, Streptococcus, Gemella, and Fusobacterium.
[0019] In the present application, tumor tissue samples of colorectal cancer patients who received neoadjuvant chemotherapy were retrospectively collected, and the intestinal flora composition information was analyzed. It was found that Treponema, Staphylococcus, Streptococcus, Gemella, and Fusobacterium had significant statistical differences between the neoadjuvant chemotherapy sensitive group and the insensitive group. Based on this combination of microbial markers, it is possible to predict whether colorectal cancer patients are sensitive to neoadjuvant chemotherapy, so as to identify whether colorectal cancer patients can benefit from neoadjuvant chemotherapy, formulate precise individualized treatment, reduce the pain of patients, avoid over-medical treatment, and improve the treatment effect.
[0020] In the present application, the neoadjuvant chemotherapy refers to the standard first-line treatment regimen based on platinum (oxaliplatin) and fluorouracil (5-FU or capecitabine) recommended by the National Comprehensive Cancer Network (NCCN) in the United States and the Chinese Society of Clinical Oncology (CSCO) guidelines.
[0021] The second aspect of the present application discloses the use of a reagent for detecting the abundance of the above-mentioned combination of microbial markers in a sample in the preparation of a product for predicting the sensitivity of neoadjuvant chemotherapy for colorectal cancer.
[0022] The third aspect of the present application discloses a product for predicting the sensitivity of neoadjuvant chemotherapy for colorectal cancer, and the product includes a reagent for detecting the abundance of the above-mentioned combination of microbial markers in a sample.
[0023] In the present application, the abundance refers to the DNA abundance or RNA abundance of the microbial markers. The detection method of the abundance can be those well-known in the art, such as metagenomic sequencing, qPCR detection, or 16S sequencing. In some specific embodiments of the present application, metagenomic sequencing is preferably used. Through metagenomic sequencing, it is possible to more comprehensively obtain and analyze the intestinal flora composition information of colorectal cancer patients and obtain microbial markers with significant statistical differences between the neoadjuvant chemotherapy sensitive group and the insensitive group.
[0024] In this application, the sample is the tumor tissue of a colorectal cancer patient. For example, it can be a biopsy tissue of a colorectal cancer patient, or a paraffin-embedded block of the tumor tissue of a colorectal cancer patient. In some specific embodiments of this application, the sample is a paraffin-embedded block of the tumor tissue of a colorectal cancer patient. The paraffin-embedded tissue sample can be stored for a long time, which is convenient for tumor research and diagnosis. Moreover, paraffin can effectively prevent tissue from spoiling and can be stored at room temperature for many years without special treatment. The same paraffin-embedded tissue sample can be sectioned repeatedly, which is convenient for researchers to conduct multiple observations and analyses, ensuring the reliability of the research results. In addition, compared with other tissue preservation methods, paraffin-embedded samples also have the advantages of relatively low cost and being more economical and affordable.
[0025] In this application, the product refers to an in vitro diagnostic product, and the types of the in vitro diagnostic products include but are not limited to chips or reagent kits.
[0026] Specifically, the product includes reagents for detecting the DNA abundance or RNA abundance of microbial markers. These reagents can be primers or probes that can specifically amplify or specifically bind to these microbial markers, and can be specifically selected according to needs. In addition, it can be understood that the product also includes other reagents required based on different detection techniques, such as buffer solutions, digestion solutions, cleaning solutions, etc., which can also be configured according to needs here, so there is no special limitation.
[0027] The fourth aspect of this application discloses a model for predicting the sensitivity of colorectal cancer to neoadjuvant chemotherapy. The input variable of the model is the abundance of the above-mentioned combination of microbial markers in the sample.
[0028] Specifically, the abundance of the above-mentioned combination of microbial markers is input into the prediction model, and based on the abundance of the microbial markers, it can be judged and output whether a colorectal cancer patient is sensitive to neoadjuvant chemotherapy.
[0029] The fifth aspect of this application discloses a system for predicting the sensitivity of colorectal cancer to neoadjuvant chemotherapy, including: A data input module, which is used to input the abundance data of the above-mentioned combination of microbial markers; A data storage module, which is used to store the abundance data of the above-mentioned combination of microbial markers in the biological samples of a population, and the population includes colorectal cancer patients who receive neoadjuvant chemotherapy; An output prediction module, which is respectively connected to the data input module and the data storage module, constructs a prediction model using the abundance data of the above-mentioned combination of microbial markers in the biological samples of the population, and based on the abundance data of the combination of microbial markers input by the data input module, obtains and outputs the prediction result of the sensitivity of colorectal cancer patients to neoadjuvant chemotherapy.
[0030] A sixth aspect of the present application discloses a computer device, including a memory for storing a computer program and a processor for executing the computer program. When the computer program is executed, the following method steps are implemented: Obtain the abundance data of the microbial biomarker combination described above in the sample; Construct a prediction model using the abundance data of the microbial biomarker combination in the samples of the population, where the population includes colorectal cancer patients receiving neoadjuvant chemotherapy; Input the obtained abundance data of the microbial biomarker combination into the prediction model, and obtain and output the prediction result of the sensitivity of colorectal cancer patients to neoadjuvant chemotherapy.
[0031] A seventh aspect of the present application discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following method steps are implemented: Obtain the abundance data of the microbial biomarker combination described above in the sample; Construct a prediction model using the abundance data of the microbial biomarker combination in the samples of the population, where the population includes colorectal cancer patients receiving neoadjuvant chemotherapy; Input the obtained abundance data of the microbial biomarker combination into the prediction model, and obtain and output the prediction result of the sensitivity of colorectal cancer patients to neoadjuvant chemotherapy.
[0032] In the present application, the construction of the prediction model includes the following steps: Randomly divide the abundance data of the microbial biomarker combination in the biological samples of the population into two groups, one group is the training set and the other group is the test set; Use the training set data to construct a prediction model based on a machine learning algorithm and conduct multiple verifications; In the test set, verify the obtained prediction model.
[0033] Among them, the machine learning algorithm used can be one of the logistic regression algorithm, random forest algorithm, XGBoost, and support vector machine algorithm. In some specific embodiments of the present application, the machine learning algorithm used is the random forest algorithm.
[0034] Through the model constructed in the present application, the sensitivity of colorectal cancer patients to neoadjuvant chemotherapy can be accurately predicted, so as to provide a more precise treatment plan clinically, identify colorectal cancer patients who are insensitive to neoadjuvant chemotherapy, avoid over-medical treatment, and improve the treatment effect of patients.
[0035] The present application will be described below through specific embodiments. It should be noted that the following specific embodiments are only for the purpose of illustration and do not limit the scope of the present application in any way. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used in the description of the present application herein are only for the purpose of describing specific embodiments and are not intended to limit the present application. In addition, unless otherwise specified, the methods without specific conditions or steps recorded are all conventional methods, and the reagents and materials used can be obtained from commercial channels.
[0036] Example 1 In this example, a total of 44 paraffin-embedded tumor tissue samples from colorectal cancer patients after neoadjuvant chemotherapy were collected. These paraffin-embedded tumor tissue samples were all from the Gastrointestinal Oncology Center of Hefei Institutes of Physical Science, Chinese Academy of Sciences (colorectal tumor patients at the first diagnosis and initial treatment). After being approved by the Ethics Committee of Hefei Institutes of Physical Science, Chinese Academy of Sciences, informed consent forms of the patients were provided, and they agreed to participate in the study and collect biological specimens.
[0037] The specific information of these colorectal cancer patients is as follows:
[0038] All of the above patients received the standard first-line treatment regimen recommended by the National Comprehensive Cancer Network (NCCN) in the United States and the Chinese Society of Clinical Oncology (CSCO) guidelines, which is based on platinum (oxaliplatin) and fluorouracil (5-FU or capecitabine).
[0039] Among them, the determination of sensitivity was based on clinical criteria: Based on the size and scope of the colonoscopy before and after neoadjuvant chemotherapy of the tumor, and imaging data such as CT, magnetic resonance, and ultrasound, the dynamic change amplitude of the tumor in terms of size, depth of invasion, and involved scope was judged, and the TRG grading based on pathology was also considered.
[0040] 1. MR imaging evaluation criteria for the effect of neoadjuvant chemoradiotherapy for rectal cancer The axial small FOV high-resolution T2WI non-fat-suppressed sequence is the main sequence for evaluating TRG. Signal definition: The tumor shows medium signal higher than the rectal muscular layer but lower than the submucosa; mucus shows extremely high signal higher than the submucosa; fiber shows low signal or even lower signal similar to muscle.
[0041] 2. MRI diagnosis criteria for rectal cancer TRG according to the pathological Mandard diagnostic criteria.
[0042] 1. mrTRG1: No residual tumor.
[0043] 2. mrTRG2: A large amount of fibrous components and a small amount of residual tumor.
[0044] 3. mrTRG3: Fibrous / mucinous components and residual tumor each account for about 50%.
[0045] 4. mrTRG4: A small amount of fibrous / mucinous components, and most is residual tumor.
[0046] 5. mrTRG5: No obvious change in the tumor.
[0047] (1) The original tumor area is normal on digital rectal examination, and no tumor mass can be palpated; (2) There are no visible tumor signs under endoscopy, or only a small amount of superficial ulcers or scars; (3) The imaging assessment of the tumor lesion shows a complete response (CR) or a large amount of fibrous components, and only a small amount of residue meets the MRI diagnostic criteria of grades 1-2; (4) The pathological TRG grading is 0-1, or part of grade 2 (not meeting grade 1, but the proportion of fibroblastic components is close to 50%). Patients with the above two factors are judged to be sensitive to neoadjuvant chemotherapy, and the rest are judged to be insensitive to neoadjuvant chemotherapy.
[0048] According to the above clinical criteria, the sensitivity distribution of 44 patients to neoadjuvant chemotherapy was determined: 31 were insensitive and 13 were sensitive.
[0049] Example 2 The paraffin-embedded tissue blocks of the tumor tissues collected in Example 1 were respectively subjected to DNA sample and RNA sample extraction for subsequent metagenomic sequencing.
[0050] 1. DNA sample extraction It was carried out using the QIAamp DNA FFPE Tissue Kit (cas: 56404, QIAGEN) kit, and the specific steps are as follows: (1) Trim the excess paraffin on the sample block with a scalpel.
[0051] (2) Cut paraffin ribbons with a thickness of 5-20 µm. If the surface of the sample has been exposed to air, discard the first 2-3 segments.
[0052] (3) Immediately put the sections into a 1.5 or 2 ml microcentrifuge tube and close the lid.
[0053] (4) Add 160 μL or 320 μL of xylene solution, vortex vigorously for 10 s, and centrifuge briefly to make the sample gather at the bottom of the tube.
[0054] (5) Incubate at 56 °C for 3 min, and then cool at room temperature (15-25 °C).
[0055] (6) Add 150 μL or 240 μL of Buffer PKD and mix by vortexing.
[0056] (7) Centrifuge at 11,000 × g (10,000 rpm) for 1 min.
[0057] (8) Add 10 μL of Proteinase K to the lower clear liquid and mix gently by pipetting up and down.
[0058] (9) Incubate at 56 °C for 15 min, then at 80 °C for 15 min. If using a heating block without a vibration function, vortex briefly every 3 - 5 min for mixing. If using only one heating block, after incubation at 56 °C, place the sample at room temperature until the heating block reaches 80 °C.
[0059] (10) Transfer the lower clear liquid to a new 2 - ml microcentrifuge tube.
[0060] (11) Incubate on ice for 3 min, then centrifuge at 20,000 × g (13,500 rpm) for 15 min.
[0061] (12) Transfer the supernatant to a new microcentrifuge tube, being careful not to aspirate the particles.
[0062] (13) Add DNase Booster Buffer (16 μL or 25 μL) and 10 μL of DNase I stock solution, mix by inverting the centrifuge tube up and down, and briefly centrifuge to collect the residual liquid on the side of the tube. Note that mixing should only be done by gently inverting the centrifuge tube, without creating a vortex.
[0063] (14) Incubate at room temperature for 15 min.
[0064] (15) Add 320 μL or 500 μL of Buffer RBC to adjust the binding conditions and mix the lysate thoroughly.
[0065] (16) Add 720 μL or 1200 μL of ethanol (100%) to the sample, mix well by pipetting, do not centrifuge, and immediately proceed to step 17.
[0066] (17) Transfer 700 μL of the sample, including any possible precipitate, to the RNeasy MinElute spin column placed in a 2 - ml collection tube; cap the lid and centrifuge at ≥8000 × g (≥10,000 rpm) for 15 s; discard the filtrate.
[0067] (18) Repeat step (17) until the entire sample has passed through the RNeasy MinElute spin column. Reuse the collection tube from step (19).
[0068] (19) Add 500 μL of Buffer RPE to the RNeasy MinElute spin column. Gently close the lid and centrifuge at ≥8000 × g (≥10000 rpm) for 15 s. Discard the filtrate. Note: Buffer RPE is provided as a concentrate. Ensure that ethanol is added before use.
[0069] (20) Add 500 μL of Buffer RPE to the RNeasy MinElute spin column. Gently close the lid and centrifuge at ≥8000 × g (≥10000 rpm) for 2 minutes. Discard the collection tube along with the filtrate.
[0070] (21) Place the RNeasy MinElute spin column into a new 2 ml collection tube. Centrifuge at full speed for 5 minutes. Discard the collection tube and the filtrate; to avoid damage to the lid, place the spin column in the centrifuge with at least one empty space between columns. Adjust the direction of the lids so that they point in the opposite direction of the rotor rotation (e.g., if the rotor rotates clockwise, adjust the lids counterclockwise).
[0071] (22) Place the RNeasy MinElute spin column into a new 1.5 ml collection tube. Directly add 14 - 30 μL of nuclease-free water onto the spin column membrane. Gently close the lid and centrifuge at full speed for 1 minute to elute the RNA.
[0072] 2. RNA Sample Extraction Use the RNeasy FFPE Kit (cas:73504, QIAGEN) for the extraction, and the specific steps are as follows: (1) Trim the excess paraffin on the sample block with a scalpel.
[0073] (2) Cut wax ribbons 5 - 20 µm thick. If the sample surface has been exposed to air, discard the first 2 - 3 segments.
[0074] (3) Immediately place the sections into a 1.5 or 2 ml microcentrifuge tube and close the lid.
[0075] (4) Add 160 μL or 320 μL of xylene solution, vortex vigorously for 10 s, and briefly centrifuge to collect the sample at the bottom of the tube.
[0076] (5) Incubate at 56 °C for 3 min, then cool at room temperature (15 - 25 °C).
[0077] (6) Add 150 μL or 240 μL of Buffer PKD and mix by vortexing.
[0078] Centrifuge at 11000 × g (10000 rpm) for 1 min.
[0079] (8) Add 10 μL of Proteinase K to the lower clear liquid and gently mix by pipetting up and down.
[0080] (9) Incubate at 56 °C for 15 min, then at 80 °C for 15 min. If using a heating block without a vibration function, vortex briefly every 3 - 5 min for mixing. If using only one heating block, after incubation at 56 °C, let the sample sit at room temperature until the heating block reaches 80 °C.
[0081] (10) Transfer the lower clear liquid to a new 2 ml microcentrifuge tube.
[0082] (11) Incubate on ice for 3 min, then centrifuge at 20000 × g (13500 rpm) for 15 min.
[0083] (12) Transfer the supernatant to a new microcentrifuge tube, being careful not to aspirate any particles.
[0084] (13) Add 16 μL or 25 μL of DNase Booster Buffer and 10 μL of DNase I stock solution, mix by inverting the centrifuge tube, and briefly centrifuge to collect any residual liquid on the sides of the tube. Note that mixing should only be done by gently inverting the centrifuge tube and no vortexing should be generated.
[0085] (14) Incubate at room temperature for 15 min.
[0086] (15) Add 320 μL or 500 μL of Buffer RBC to adjust the binding conditions and mix the lysate thoroughly.
[0087] (16) Add 720 μL or 1200 μL of ethanol (100%) to the sample, mix well by pipetting, do not centrifuge, and immediately proceed to step (17).
[0088] (17) Transfer 700 μL of the sample, including any possible precipitate, to the RNeasy MinElute spin column placed in a 2 ml collection tube; cap the tube and centrifuge at ≥8000 × g (≥10000 rpm) for 15 s; discard the filtrate.
[0089] (18) Repeat step (17) until the entire sample has passed through the RNeasy MinElute spin column. Reuse the collection tube from step (19).
[0090] (19) Add 500 μL of Buffer RPE to the RNeasy MinElute spin column. Gently close the lid and centrifuge at ≥8000 × g (≥10000 rpm) for 15 s. Discard the filtrate. Note: Buffer RPE is provided as a concentrate. Ensure that ethanol is added before use.
[0091] (20) Add 500 μL of Buffer RPE to the RNeasy MinElute spin column. Gently close the lid and centrifuge at ≥8000 × g (≥10000 rpm) for 2 minutes. Discard the collection tube along with the filtrate.
[0092] (21) Place the RNeasy MinElute spin column into a new 2 ml collection tube. Centrifuge at full speed for 5 minutes. Discard the collection tube and the filtrate; to avoid damage to the lid, place the spin column in the centrifuge with at least one empty space between columns. Adjust the direction of the lids so that they point in the opposite direction of the rotor rotation (e.g., if the rotor rotates clockwise, adjust the lids counterclockwise).
[0093] (22) Place the RNeasy MinElute spin column into a new 1.5 ml collection tube. Add 14 - 30 μL of nuclease-free water directly onto the spin column membrane. Gently close the lid and centrifuge at full speed for 1 minute to elute the RNA.
[0094] Example 3 1. Using metagenomic sequencing technology for the DNA or RNA samples extracted in Example 2, analyze the composition of the gut tumor tissue microbiota of the patient and calculate its abundance value. The specific steps are as follows: (1) Microbial metagenomic sequencing: Obtain the genomic information of each species in the microbial community by performing metagenomic sequencing on the DNA or RNA samples extracted in Example 2.
[0095] (2)Sequence preprocessing: First, use Cutadapt to remove the adapter sequences introduced during sequencing. Subsequently, use quality control tools such as FastQC to comprehensively evaluate the quality of the original sequence data. On this basis, screen and remove sequences with poor quality, including low-quality bases and sequences shorter than the minimum length threshold. For samples containing host DNA, use alignment tools such as Bowtie2 or BWA to align the host DNA sequences with the reference genome (specifically, align the sequencing reads to the database GRCh38, then extract the reads that are not aligned to the database GRCh38, and align the reads that are not aligned to the database GRCh38 to the database GRCh38 again, then extract the reads that are not aligned to the database GRCh38, and further align the reads that are not aligned to the database GRCh38 to the database T2T-CHM13v2.0 to capture those sequences that cannot be aligned in the database GRCh38, so as to minimize data loss and improve the comprehensiveness of the analysis), and remove all matching host sequences to reduce host contamination and ensure the accuracy and purity of subsequent analysis. This step is crucial for improving data quality and reducing background noise.
[0096] (3)Sequence assembly: Use the assembly software SPAdes for the assembly of conventional genomic data, or use MetaSPAdes specifically for metagenomic data to assemble the preprocessed short reads into longer contigs or scaffolds, and construct a draft genome of the metagenome. Subsequently, use the assembly evaluation software QUAST to evaluate indicators such as the length, number, N50 value, and integrity of the assembled contigs to ensure the quality of the assembly results.
[0097] (4)Species annotation and classification: Use reference databases, such as the NCBI Nucleotide database or the Greengenes database, to align the assembled contigs or scaffolds, and perform species annotation and classification based on the alignment results. Identify contigs that are highly similar to the genomes of species in the database and perform species annotation. Count the number of contigs of different species, and estimate the relative abundance of species based on the proportion of the number of these contigs. For example, determine the abundance of different species in the sample according to the proportion of the number of contigs, so as to perform species classification.
[0098] 2. Machine learning for model construction (1) Calculate the average abundance difference of each microorganism between sensitive and insensitive: Randomly select 14 samples from the 44 total samples in Example 1 (among them, 10 sensitive samples and 4 insensitive samples). Extract the microbial abundance data of sensitive and insensitive respectively from the 14 samples, calculate the mean difference of each microorganism, and select the top 50 microorganisms with the largest absolute difference for subsequent PCA verification.
[0099] (2) According to PCA verification, observe whether the selected consistent features can separate the two types of samples: For the 50 screening features of the 14 samples in step (1), after standardizing the data, perform PCA dimensionality reduction analysis. The PCA results show that the samples in the sensitive group and the insensitive group are significantly separated in the principal component space, and can accurately distinguish the two types of samples. Therefore, it can be considered that the selected microbial features have strong discriminant ability in distinguishing sensitive and insensitive samples.
[0100] (3) Use Euclidean distance to detect outliers and perform sample rejection: In this stage, we use the data after PCA dimensionality reduction to avoid being affected by the original dimension of the data and ensure that the calculated outliers can reflect the anomalies in the main feature space. Use Euclidean distance to detect outliers, compare the distance of the samples with a threshold to identify outliers, use a 95% distance threshold, and establish an outlier index to find outlier samples. After outlier detection based on Euclidean distance, no outliers exceeding the set threshold (such as 95% distance threshold) are found in the 14 samples, and all samples are within the normal range, so there is no need to perform sample rejection operations. The feature matrix and labels remain unchanged and can be directly used for subsequent data analysis or modeling work.
[0101] (4) Read XGBoost model data: Import the microbiome feature data and the corresponding chemotherapy response labels from the standardized CSV file. After import, perform Z-score standardization on all features to eliminate the adverse effects of different dimensions on model training. Subsequently, divide the samples of 14 colorectal cancer patients into a training set and a test set according to a certain proportion: The training set contains 12 samples, among which 9 are sensitive to neoadjuvant chemotherapy and 3 are insensitive, and the test set contains 2 samples, 1 sensitive and 1 insensitive each, ensuring that the sensitive and insensitive each account for 50% in the test set to more accurately evaluate the classification performance of the XGBoost model.
[0102] (5) Initialize the XGBoost model and set the objective function: In this step, first initialize the XGBoost model, set the objective function to logarithmic loss, and add first-order and second-order regularization terms to the objective function to effectively suppress overfitting while improving the model fitting ability.
[0103] (6) Gradient and Hessian matrix calculation: Based on the prediction results of the current XGBoost model, the first-order derivative (i.e., the gradient) and the second-order derivative (i.e., the Hessian matrix) of the residuals are calculated for each sample in the training set, and these values are used as sample weights for subsequent evaluation of the split gain.
[0104] (7) Construction of the optimal split of the decision tree: XGBoost uses a greedy strategy to traverse the candidate split points of all features and selects the optimal split according to the gain value jointly determined by the gradient and the Hessian matrix. If the gain is lower than the preset threshold, the node is stopped from further splitting to reduce the risk of overfitting.
[0105] (8) Model iterative update and learning rate adjustment: In each round of iteration, XGBoost multiplies the newly constructed weak learner by the learning rate and accumulates it with the existing model to form a new overall prediction function; subsequently, the processes of gradient calculation, optimal split, weak learner construction, and model update are repeated until the preset number of iterations is reached or the early stopping condition is triggered.
[0106] (9) XGBoost feature importance scoring and screening: After training, the XGBoost model outputs the importance scores of each feature microbial genus for model prediction. These scores reflect the frequency of each feature appearing in the split nodes and its contribution to the information gain. According to the importance ranking of the features, the top 15 microbial genera were selected, and it was considered that these features had significant discriminative power in predicting chemotherapy response and might be potential biomarkers.
[0107] (10) Removal of interference from contaminated microbial genera: To further enhance the biological relevance of the analysis, microbial genera that might originate from environmental or technical contamination were excluded from these 15 microbial genera, and only the target microbial genera known to be related to human colorectal cancer were retained, including Treponema, Staphylococcus, Streptococcus, Gemella, and Fusobacterium. Based on this screening, the feature importance scores of these target microorganisms in the XGBoost model were re-obtained, and the results are as Figure 1 shown.
[0108] (11) The 5 selected microbial markers were used as features to evaluate the model performance. A prediction model was constructed using the random forest algorithm. The model was fitted and trained on the training set to learn the data features; subsequently, it was verified on the test set to test the prediction performance of the random forest algorithm, thereby comprehensively evaluating the performance of the model.
[0109] Example 4 Verification of Model Performance After removing the 14 samples used for feature screening in Example 3, the remaining 30 samples (21 sensitive samples and 9 insensitive samples) from the 44 total samples in Example 1 were used to independently validate the random forest model and verify the model's performance.
[0110] The predictive performance of the model was quantified by the receiver operating characteristic (ROC) curve. Specifically, first, the 30 samples were randomly divided into a training set (24 samples) and a test set (6 samples). Among them, 20% of the samples in the total dataset (i.e., 6 samples) were randomly selected as the test set. In the test set, there were 3 sensitive group samples (accounting for 50%) and 3 insensitive group samples (accounting for 50%), which were used to evaluate the model's ability to distinguish the sensitivity of patients with colorectal cancer to neoadjuvant chemotherapy. A random forest prediction model was constructed using the training set. Finally, the test set was input into the model to predict the sensitivity of patients with colorectal cancer to neoadjuvant chemotherapy, and the ROC curve was plotted and the area under the curve (AUC) was calculated. At the same time, a confusion matrix was generated to comprehensively evaluate the classification performance of the model. The specific steps are as follows: (1) Detection of relative abundances of gut microbial markers: First, the samples were subjected to metagenomic sequencing according to the method in Example 2 to determine the relative abundances of 5 specific bacterial species. These data will be used as key features and input into the constructed random forest model to provide a basis for subsequent modeling and analysis.
[0111] (2) Data reading: First, in the data reading stage, microbiome features and corresponding clinical response labels were imported from a standardized CSV file; then data preprocessing was performed, mainly including handling missing values, such as using median imputation or deleting features with severe missing values, and standardizing the features, such as Z - score standardization, to eliminate the influence of dimensions and improve the stability of the model.
[0112] (3) Data partitioning: Next, the preprocessed data was partitioned into a training set and a test set, allocated in an 8:2 ratio, so as to objectively measure its performance on unseen data during model evaluation.
[0113] (4) Random forest model training: In the model training stage, the random forest algorithm was introduced. By constructing a large number of decision trees and combining their output results for integrated prediction, the robustness and accuracy of the model were improved. To enhance the ability to identify class imbalance, the setting of class weights was introduced into the model to strengthen the ability to identify minority classes, such as patients with treatment response.
[0114] (5) Model hyperparameter tuning: During the model hyperparameter tuning process, the grid search method was adopted to systematically search for multiple hyperparameter combinations, such as the number of trees, maximum depth, minimum number of samples for splitting, etc., to determine the optimal parameter combination and avoid overfitting.
[0115] (6)Model performance evaluation: In terms of model performance evaluation, a variety of metrics such as the ROC curve, AUC value, confusion matrix, and classification report are combined to comprehensively measure the classification ability and generalization performance of the model. The ROC curve is used to visualize the balance between sensitivity and specificity at different thresholds, and the AUC value provides an overall performance score; the confusion matrix can reveal the numbers of TP, FP, TN, and FN in the prediction results, and the classification report further details metrics such as precision, recall, and F1 score.
[0116] The verification results are as follows: Figure 2 The shown confusion matrix indicates that the model has excellent classification accuracy and can effectively distinguish positive and negative class samples. For positive class samples, the model shows high sensitivity and can accurately identify most positive class samples; at the same time, the specificity is good, and there are very few cases of misclassifying negative class samples as positive class. When verifying 6 samples in the test set, the numbers of both false positives and false negatives are small. As Figure 3 shown, the AUC value of the ROC curve of the model on the training set is as high as 0.992 ( Figure 3 a in Figure 3 ), and the AUC value on the test set also reaches 0.889 (
[0117] b in
[0118] This indicates that the prediction performance of the model is extremely excellent and it has strong ability to distinguish positive and negative samples. In summary, the microbial biomarker combination provided in this application and the constructed prediction model have significant efficacy and can provide an accurate basis for predicting the sensitivity of neoadjuvant chemotherapy for colorectal cancer patients.
[0117] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not conflict, it should be considered as the scope described in this specification.
[0118] The above-described embodiments only represent several implementation manners of this application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several deformations and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of the patent of this application should be subject to the appended claims.
Claims
1. A microbial biomarker combination for predicting the sensitivity of neoadjuvant chemotherapy for colorectal cancer, characterized in that, The microbial biomarker combination includes: Treponema, Staphylococcus, Streptococcus, Gemella, and Fusobacterium.
2. Use of a reagent for detecting the abundance of the microbial biomarker combination according to claim 1 in a sample in the preparation of a product for predicting the sensitivity of neoadjuvant chemotherapy for colorectal cancer.
3. A product for predicting the sensitivity of neoadjuvant chemotherapy for colorectal cancer, characterized in that, The product includes a reagent for detecting the abundance of the microbial biomarker combination according to claim 1 in a sample.
4. The product according to claim 3, characterized in that, The sample is a tumor tissue of a colorectal cancer patient; and / or, the sample is a paraffin-embedded block of a tumor tissue of a colorectal cancer patient.
5. The product according to claim 3, wherein, The product is an in vitro detection product; and / or, the in vitro detection product is a kit.
6. A model for predicting the sensitivity of neoadjuvant chemotherapy for colorectal cancer, characterized in that, The input variable of the model is the abundance of the microbial biomarker combination according to claim 1 in the sample; and / or, the abundance of the microbial biomarker combination is obtained by metagenomic sequencing, qPCR detection, or 16S sequencing; and / or, the abundance of the microbial biomarker combination is obtained by metagenomic sequencing.
7. A system for predicting the sensitivity of neoadjuvant chemotherapy for colorectal cancer, characterized in that, Including: A data input module for inputting the abundance data of the microbial biomarker combination described in claim 1; A data storage module for storing the abundance data of the microbial biomarker combination in the biological samples of a population, the population including colorectal cancer patients receiving neoadjuvant chemotherapy; An output prediction module, which is respectively connected to the data input module and the data storage module, constructs a prediction model using the abundance data of the microbial biomarker combination in the biological samples of the population, and based on the abundance data of the microbial biomarker combination input by the data input module, obtains and outputs a prediction result of the sensitivity of colorectal cancer patients to neoadjuvant chemotherapy.
8. The system according to claim 7, wherein The construction of the prediction model includes the following steps: Randomly divide the abundance data of the microbial biomarker combination in the biological samples of the population into two groups, one group is a training set and the other group is a test set; Using the training set data, construct a prediction model based on a machine learning algorithm and perform multiple validations; In the test set, validate the obtained prediction model; and / or, the machine learning algorithm is a random forest algorithm.
9. A computer device, comprising a memory for storing a computer program and a processor for executing the computer program, characterized in that, When the computer program is executed, the system as described in claim 7 or 8 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the system as described in claim 7 or 8 is implemented.
Citation Information
Patent Citations
Microbial marker of colorectal cancer and application of marker
CN109943636A
Biomarker for identifying community acquired pneumonia and application thereof
CN115976198A
Gene markers for predicting or evaluating neoadjuvant chemotherapy response of breast cancer and application of gene markers
CN120005996A
Methods for comparative metagenomic analysis
WO2019232357A1
Cited By
Method and system for predicting juvenile depression based on intestinal flora
CN121306573A