Microbial marker combination and its application in predicting sensitivity of neoadjuvant chemotherapy in colorectal cancer

By analyzing the intestinal flora of colorectal cancer patients and using a combination of microbial markers such as Treponema to construct a machine learning model, the difficult problem of predicting the sensitivity of colorectal cancer to neoadjuvant chemotherapy was solved, accurate personalized treatment plans were achieved, and the treatment effect and safety were improved.

CN120350126BActive Publication Date: 2025-09-23HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510861758.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-09-23
Estimated Expiration
2045-06-25

Smart Images

  • Figure CN120350126B_ABST
    Figure CN120350126B_ABST
Patent Text Reader

Abstract

This application relates to the field of bioassay technology and specifically discloses a combination of microbial markers for predicting the sensitivity of colorectal cancer to neoadjuvant chemotherapy and its application. The combination of microbial markers includes Treponema, Staphylococcus, Streptococcus, Gemella, and Fusobacterium. This application analyzes the composition of the microbial community in the tumor microenvironment of colorectal cancer patients to identify microbial markers associated with sensitivity to neoadjuvant chemotherapy. Based on the microbial marker combination disclosed in this application, it is possible to accurately predict patients who are insensitive to neoadjuvant chemotherapy, reduce treatment risks, and reduce ineffective treatments, providing a reliable basis for developing precise and personalized treatments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of biological detection technology, and specifically relates to a combination of microbial markers for predicting the sensitivity of colorectal cancer to neoadjuvant chemotherapy and its application. Background Art

[0002] Colorectal cancer (CRC) is a malignant tumor that originates in the colon (large intestine) or rectum. It is one of the most common cancers worldwide. Current treatment options for CRC include surgery, chemotherapy, radiotherapy, targeted therapy, and immunotherapy. Surgery is the mainstay of treatment for CRC. However, due to the prevalence of advanced stage colorectal cancer, postoperative recurrence and progression are common. For patients with advanced stage disease, preoperative neoadjuvant therapy facilitates observation of efficacy, effectively downstaging the disease, and improving the efficacy of radical surgery. Due to factors such as tumor heterogeneity, the efficacy of adjuvant chemotherapy for CRC varies from patient to patient. Patients who are insensitive to neoadjuvant chemotherapy may experience severe side effects after treatment, which may affect subsequent surgery, and some patients may even have a poor prognosis. Therefore, considering the varying responses to neoadjuvant therapy, a uniform approach to neoadjuvant therapy is recommended; tailored treatment plans are required.

[0003] However, there is currently no reliable prediction method for colorectal neoadjuvant chemotherapy to screen patients who are sensitive to neoadjuvant chemotherapy, formulate targeted treatment plans, reduce treatment risks, and improve treatment effects. Summary of the Invention

[0004] In light of this, the primary objective of this application is to provide a combination of microbial markers for predicting the sensitivity of colorectal cancer patients to neoadjuvant chemotherapy. The combination includes Treponema, Staphylococcus, Streptococcus, Gemella, and Fusobacterium. This combination can accurately predict the sensitivity of colorectal cancer patients to neoadjuvant chemotherapy with high accuracy and sensitivity, providing a basis for personalized treatment approaches.

[0005] In order to achieve the above objectives, this application adopts the following technical solutions:

[0006] One aspect of the present application provides a combination of microbial markers for predicting the sensitivity of colorectal cancer to neoadjuvant chemotherapy, wherein the combination of microbial markers includes Treponema, Staphylococcus, Streptococcus, Gemella, and Fusobacterium.

[0007] Another aspect of the present application provides the use of a reagent for detecting the abundance of the above-mentioned combination of microbial markers in a sample in the preparation of a product for predicting the sensitivity of colorectal cancer to neoadjuvant chemotherapy.

[0008] Another aspect of the present application provides a product for predicting the sensitivity of colorectal cancer to neoadjuvant chemotherapy, wherein the product comprises a reagent for detecting the abundance of the combination of microbial markers described above in a sample.

[0009] Another aspect of the present application provides a model for predicting the sensitivity of colorectal cancer to neoadjuvant chemotherapy, wherein the input variable of the model is the abundance of the combination of microbial markers described above in the sample.

[0010] Another aspect of the present application provides a system for predicting the sensitivity of colorectal cancer to neoadjuvant chemotherapy, comprising:

[0011] A data input module, which is used to input the abundance data of the microbial marker combination described above;

[0012] a data storage module for storing abundance data of the microbial marker combination in biological samples of a population, the population comprising colorectal cancer patients receiving neoadjuvant chemotherapy;

[0013] An output prediction module is connected to the data input module and the data storage module respectively, and uses the abundance data of the microbial marker combination in the biological samples of the group to build a prediction model, and based on the abundance data of the microbial marker combination input by the data input module, obtains and outputs the prediction results of the sensitivity of colorectal cancer patients to neoadjuvant chemotherapy.

[0014] Another aspect of the present application provides a computer device, comprising a memory for storing a computer program and a processor for executing the computer program, wherein when the computer program is executed, the steps of the following method are implemented:

[0015] Obtaining abundance data of the above-mentioned microbial marker combinations in the sample;

[0016] constructing a predictive model using the abundance data of the microbial marker combination in samples from a population comprising colorectal cancer patients receiving neoadjuvant chemotherapy;

[0017] The obtained abundance data of the microbial marker combination is input into the prediction model to obtain and output the prediction results of the sensitivity of colorectal cancer patients to neoadjuvant chemotherapy.

[0018] Another aspect of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the following method are implemented:

[0019] Obtaining abundance data of the above-mentioned microbial marker combinations in the sample;

[0020] constructing a predictive model using the abundance data of the microbial marker combination in samples from a population comprising colorectal cancer patients receiving neoadjuvant chemotherapy;

[0021] The obtained abundance data of the microbial marker combination is input into the prediction model to obtain and output the prediction results of the sensitivity of colorectal cancer patients to neoadjuvant chemotherapy.

[0022] Beneficial effects of this application:

[0023] This application collects samples of tumor tissue from patients with colorectal cancer and analyzes the composition of the intestinal flora. It is found that Treponema, Staphylococcus, Streptococcus, Gemella and Fusobacterium have significant statistical differences between the sensitive group and the insensitive group. Therefore, the microbial marker combination provided in this application can predict whether colorectal cancer patients are sensitive to neoadjuvant chemotherapy.

[0024] This application constructs a machine learning model based on the abundance data of microbial marker combinations to predict the sensitivity of colorectal cancer patients to neoadjuvant chemotherapy. Experimental verification shows that this model predicts sensitivity with high accuracy and good sensitivity. Therefore, it can be used to clinically determine whether a patient is suitable for neoadjuvant chemotherapy and select the optimal treatment method based on individual differences, thus achieving personalized precision medicine. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 The five characteristic microorganisms obtained after XGBoost screening and decontamination in Example 3 and their feature importance scores in the random forest model are presented.

[0026] Figure 2 The confusion matrix of the random forest prediction model in Example 3 on the test set cases is shown.

[0027] Figure 3 is the ROC curve of the training set and test set of the random forest model constructed in Example 4, wherein, Figure 3 The a in is the ROC curve of the training set, Figure 3 The b in is the ROC curve of the test set. DETAILED DESCRIPTION

[0028] The embodiments of the present application are described in detail below. The embodiments described below are exemplary and are only used to explain the present application, and should not be understood as limiting the present application.

[0029] The first aspect of the present application provides a combination of microbial markers for predicting the sensitivity of colorectal cancer to neoadjuvant chemotherapy, wherein the combination of microbial markers includes: Treponema, Staphylococcus, Streptococcus, Gemella and Fusobacterium.

[0030] In this application, tumor tissue samples from colorectal cancer patients who received neoadjuvant chemotherapy were retrospectively collected, and the intestinal flora composition information was analyzed. It was found that Treponema, Staphylococcus, Streptococcus, Gemella and Fusobacterium had significant statistical differences in the neoadjuvant chemotherapy-sensitive and -insensitive groups. Based on these microbial marker combinations, it is possible to predict whether colorectal cancer patients are sensitive to neoadjuvant chemotherapy, thereby identifying whether colorectal cancer patients can benefit from neoadjuvant chemotherapy, so as to develop precise individualized treatment, reduce patient suffering, avoid excessive medical treatment, and improve treatment effects.

[0031] In this application, the neoadjuvant chemotherapy refers to the standard first-line treatment regimen based on platinum (oxaliplatin) and fluorouracil (5-FU or capecitabine) recommended by the National Comprehensive Cancer Network (NCCN) and the Chinese Society of Clinical Oncology (CSCO) guidelines.

[0032] The second aspect of the present application discloses the use of a reagent for detecting the abundance of the above-mentioned combination of microbial markers in a sample in the preparation of a product for predicting the sensitivity of colorectal cancer to neoadjuvant chemotherapy.

[0033] The third aspect of the present application discloses a product for predicting the sensitivity of colorectal cancer to neoadjuvant chemotherapy, wherein the product comprises a reagent for detecting the abundance of the combination of microbial markers described above in a sample.

[0034] In this application, the abundance refers to the DNA abundance or RNA abundance of a microbial marker. The abundance can be detected using methods well known in the art, such as metagenomic sequencing, qPCR detection, or 16S sequencing. In some specific embodiments of this application, metagenomic sequencing is preferably used. Metagenomic sequencing can more comprehensively obtain and analyze the intestinal flora composition of colorectal cancer patients, and obtain microbial markers with significant statistical differences between the neoadjuvant chemotherapy-sensitive and chemotherapy-insensitive groups.

[0035] In the present application, the sample is a tumor tissue of a colorectal cancer patient, for example, it can be a biopsy tissue of a colorectal cancer patient, or it can be a paraffin-embedded block of a tumor tissue of a colorectal cancer patient. In some specific embodiments of the present application, the sample is a paraffin-embedded block of a tumor tissue of a colorectal cancer patient. Paraffin-embedded tissue samples can be preserved for a long time, which is convenient for tumor research and diagnosis, and paraffin can effectively prevent tissue corruption and deterioration, and can be stored at room temperature for many years without special treatment. The same paraffin-embedded tissue sample can be repeatedly sectioned, which is convenient for researchers to observe and analyze multiple times and ensure the reliability of the research results. In addition, compared with other tissue preservation methods, paraffin-embedded samples also have the advantages of being relatively low in cost and more economical.

[0036] In this application, the product refers to an in vitro detection product, and the types of the in vitro detection product include but are not limited to chips or kits.

[0037] Specifically, the product includes reagents for detecting the abundance of DNA or RNA of microbial markers. These reagents can be primers or probes capable of specifically amplifying or specifically binding to these microbial markers, and the specific reagents can be selected based on needs. Furthermore, it is understood that the product may also include other reagents required by different detection technologies, such as buffers, digestion solutions, and cleaning solutions. These can also be configured as needed and are not particularly limited.

[0038] The fourth aspect of the present application discloses a model for predicting the sensitivity of colorectal cancer to neoadjuvant chemotherapy, wherein the input variable of the model is the abundance of the combination of microbial markers described above in the sample.

[0039] Specifically, the abundance of the microbial marker combination described above is input into the prediction model, and based on the abundance of the microbial marker, it can be judged and output whether the colorectal cancer patient is sensitive to neoadjuvant chemotherapy.

[0040] The fifth aspect of the present application discloses a system for predicting the sensitivity of colorectal cancer to neoadjuvant chemotherapy, comprising:

[0041] A data input module, which is used to input the abundance data of the microbial marker combination described above;

[0042] a data storage module for storing abundance data of the microbial marker combination in biological samples of a population, the population comprising colorectal cancer patients receiving neoadjuvant chemotherapy;

[0043] An output prediction module is connected to the data input module and the data storage module respectively, and uses the abundance data of the microbial marker combination in the biological samples of the group to build a prediction model, and based on the abundance data of the microbial marker combination input by the data input module, obtains and outputs the prediction results of the sensitivity of colorectal cancer patients to neoadjuvant chemotherapy.

[0044] A sixth aspect of the present application discloses a computer device, comprising a memory for storing a computer program and a processor for executing the computer program. When the computer program is executed, the following method steps are implemented:

[0045] Obtaining abundance data of the above-mentioned microbial marker combinations in the sample;

[0046] constructing a predictive model using the abundance data of the microbial marker combination in samples from a population comprising colorectal cancer patients receiving neoadjuvant chemotherapy;

[0047] The obtained abundance data of the microbial marker combination is input into the prediction model to obtain and output the prediction results of the sensitivity of colorectal cancer patients to neoadjuvant chemotherapy.

[0048] A seventh aspect of the present application discloses a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the following method are implemented:

[0049] Obtaining abundance data of the above-mentioned microbial marker combinations in the sample;

[0050] constructing a predictive model using the abundance data of the microbial marker combination in samples from a population comprising colorectal cancer patients receiving neoadjuvant chemotherapy;

[0051] The obtained abundance data of the microbial marker combination is input into the prediction model to obtain and output the prediction results of the sensitivity of colorectal cancer patients to neoadjuvant chemotherapy.

[0052] In this application, the construction of the prediction model includes the following steps:

[0053] The abundance data of the microbial marker combination in the biological samples of the group are randomly divided into two groups, one group is a training set, and the other group is a test set;

[0054] Using training set data, we build a prediction model based on machine learning algorithms and conduct multiple verifications;

[0055] The obtained prediction model is verified in the test set.

[0056] The machine learning algorithm used may be one of a logistic regression algorithm, a random forest algorithm, an XGBoost algorithm, and a support vector machine algorithm. In some specific embodiments of the present application, the machine learning algorithm used is a random forest algorithm.

[0057] The model constructed in this application can accurately predict the sensitivity of colorectal cancer patients to neoadjuvant chemotherapy, thereby providing more accurate treatment plans in clinical practice, identifying colorectal cancer patients who are insensitive to neoadjuvant chemotherapy, avoiding over-medicalization, and improving patient treatment outcomes.

[0058] The application is described below by specific examples, it should be noted that the following specific examples are merely for illustrative purposes, and do not limit the scope of the application in any way, unless otherwise defined, all technical and scientific terms used herein are identical with the meaning generally understood by those skilled in the art belonging to the technical field of the application. The terms used herein in the specification of the application are only for the purpose of describing specific embodiments, and are not intended to limit the application. In addition, unless otherwise specified, the method for not specifically recording conditions or steps is conventional method, and the reagents and materials adopted can be obtained from commercial sources.

[0059] Example 1

[0060] In this example, a total of 44 postoperative tumor tissue paraffin block samples from patients with colorectal cancer who received neoadjuvant chemotherapy were collected. These tumor tissue paraffin block samples were all from the Gastrointestinal Tumor Center of Hefei Cancer Hospital, Chinese Academy of Sciences (patients with newly diagnosed and treated colorectal cancer). With the approval of the Ethics Committee of Hefei Cancer Hospital, Chinese Academy of Sciences, the patients provided informed consent to participate in the study and collect biological specimens.

[0061] The specific information of colorectal cancer patients in these samples is as follows:

[0062]

[0063] All the above patients received the standard first-line treatment regimen based on platinum (oxaliplatin) and fluorouracil (5-FU or capecitabine) recommended by the National Comprehensive Cancer Network (NCCN) and the Chinese Society of Clinical Oncology (CSCO) guidelines.

[0064] Among them, the determination of sensitivity is based on clinical criteria:

[0065] Based on the size and range of the tumor before and after neoadjuvant chemotherapy, CT, MRI, ultrasound and other imaging data, the dynamic changes in tumor size, infiltration depth, and involvement range are judged, as well as the TRG grade based on pathology.

[0066] 1. MR imaging evaluation criteria for the efficacy of neoadjuvant chemoradiotherapy for rectal cancer

[0067] Axial, small FOV, high-resolution T2WI (non-fat-suppressed) sequences are the primary sequence for evaluating TRG. Signal definition: Tumors have an intermediate signal, above the rectal muscle layer but below the submucosa; mucus has an extremely high signal, above the submucosa; and fibers have a low signal, similar to muscle, or even lower.

[0068] 2. The MRI diagnostic criteria for rectal cancer TRG were derived based on the pathological Mandard diagnostic criteria.

[0069] 1.mrTRG1: no residual tumor.

[0070] 2.mrTRG2: a large amount of fibrous components and a small amount of residual tumor.

[0071] 3.mrTRG3: Fibrous / mucinous components and residual tumor each account for approximately 50%.

[0072] 4.mrTRG4: A small amount of fibrous / mucinous components, mostly residual tumor.

[0073] 5.mrTRG5: No obvious changes were observed in the tumor.

[0074] (1) Digital rectal examination shows that the original tumor area is normal, with no palpable tumor mass; (2) No visible tumor signs under endoscopy, or only a small amount of superficial ulcers or scars; (3) Complete remission (CR) or a large amount of fibrous components, with only a small amount of residual tumor lesions meeting the MRI diagnostic criteria of 1-2; (4) Pathological TRG grade 0-1, or partially grade 2 (not meeting grade 1, but the fibroblast component accounts for nearly 50%). Patients with the above two factors are judged to be sensitive to neoadjuvant chemotherapy, while the rest are judged to be insensitive to neoadjuvant chemotherapy.

[0075] The sensitivity of 44 patients to neoadjuvant chemotherapy was determined by the above clinical criteria: 31 patients were insensitive and 13 patients were sensitive.

[0076] Example 2

[0077] DNA samples and RNA samples were extracted from the tumor tissue wax block samples collected in Example 1 for subsequent metagenomic sequencing.

[0078] 1. DNA sample extraction

[0079] The QIAamp DNA FFPE Tissue Kit (cas: 56404, QIAGEN) was used for the analysis. The specific steps are as follows:

[0080] (1) Use a scalpel to trim excess paraffin from the sample block.

[0081] (2) Cut wax rolls with a thickness of 5 to 20 µm. If the sample surface has been exposed to air, discard the first 2 to 3 sections.

[0082] (3) Immediately place the slices into a 1.5 or 2 ml microcentrifuge tube and close the lid.

[0083] (4) Add 160 μL or 320 μL of xylene solution, vortex vigorously for 10 seconds, and centrifuge briefly to allow the sample to accumulate at the bottom of the tube.

[0084] (5) Incubate at 56°C for 3 min and then cool to room temperature (15-25°C).

[0085] (6) Add 150 μL or 240 μL Buffer PKD and mix by vortexing.

[0086] (7) Centrifuge at 11,000 × g (10,000 rpm) for 1 min.

[0087] (8) Add 10 μL of proteinase K to the lower clear liquid and mix gently by pipetting up and down.

[0088] (9) Incubate at 56°C for 15 min, then at 80°C for 15 min. If using a heating block without a vibrating function, vortex briefly every 3–5 min to mix. If using only one heating block, place the sample at room temperature after incubation at 56°C until the heating block reaches 80°C.

[0089] (10) Transfer the lower clear liquid to a new 2 ml microcentrifuge tube.

[0090] (11) Incubate on ice for 3 min, then centrifuge at 20,000 × g (13,500 rpm) for 15 min.

[0091] (12) Transfer the supernatant to a new microcentrifuge tube, taking care not to aspirate the pellet.

[0092] (13) Add DNase Booster Buffer (16 μL or 25 μL) and 10 μL DNase I stock solution, mix by inverting the tube upside down, and centrifuge briefly to remove any residual liquid from the side of the collection tube. Note that mixing should only be performed by gently turning the tube over, and do not create a vortex.

[0093] (14) Incubate at room temperature for 15 minutes.

[0094] (15) Add 320 μL or 500 μL Buffer RBC to adjust the binding conditions and mix the lysate thoroughly.

[0095] (16) Add 720 μL or 1200 μL of 100% ethanol to the sample and mix well by pipetting. Do not centrifuge and proceed immediately to step 17.

[0096] (17) Transfer 700 μL of the sample, including any precipitate that may have formed, to an RNeasy MinElute spin column in a 2 ml collection tube; cap the column and centrifuge at ≥8000 × g (≥10000 rpm) for 15 s; discard the filtrate.

[0097] (18) Repeat step (17) until the entire sample has passed through the RNeasy MinElute spin column. Reuse the collection tube from step (19).

[0098] (19) Add 500 μL of Buffer RPE to the RNeasy MinElute spin column. Gently close the cap and centrifuge at ≥8000 × g (≥10000 rpm) for 15 seconds. Discard the filtrate. Note: Buffer RPE is provided as a concentrate; be sure to add ethanol before use.

[0099] (20) Add 500 μL of Buffer RPE to the RNeasy MinElute spin column. Gently close the cap and centrifuge at ≥8000 × g (≥10000 rpm) for 2 minutes. Discard the collection tube and the filtrate.

[0100] (21) Place the RNeasy MinElute spin column in a new 2 ml collection tube. Centrifuge at full speed for 5 minutes and discard the collection tube and filtrate. To avoid damage to the caps, place the spin columns in the centrifuge with at least one empty space between the columns. Orient the caps so that they point in the opposite direction of the rotor's rotation (e.g., if the rotor rotates clockwise, orient the caps counterclockwise).

[0101] (22) Place the RNeasy MinElute spin column in a new 1.5 ml collection tube. Add 14–30 μL of nuclease-free water directly onto the spin column membrane. Gently close the cap and centrifuge at full speed for 1 minute to elute the RNA.

[0102] 2. RNA Sample Extraction

[0103] The RNeasy FFPE Kit (cas:73504, QIAGEN) was used for the analysis. The specific steps are as follows:

[0104] (1) Use a scalpel to trim excess paraffin from the sample block.

[0105] (2) Cut wax rolls with a thickness of 5 to 20 µm. If the sample surface has been exposed to air, discard the first 2 to 3 sections.

[0106] (3) Immediately place the slices into a 1.5 or 2 ml microcentrifuge tube and close the lid.

[0107] (4) Add 160 μL or 320 μL of xylene solution, vortex vigorously for 10 seconds, and centrifuge briefly to allow the sample to accumulate at the bottom of the tube.

[0108] (5) Incubate at 56°C for 3 min and then cool to room temperature (15-25°C).

[0109] (6) Add 150 μL or 240 μL Buffer PKD and mix by vortexing.

[0110] (7) Centrifuge at 11,000 × g (10,000 rpm) for 1 min.

[0111] (8) Add 10 μL of proteinase K to the lower clear liquid and mix gently by pipetting up and down.

[0112] (9) Incubate at 56°C for 15 min, then at 80°C for 15 min. If using a heating block without a vibrating function, vortex briefly every 3–5 min to mix. If using only one heating block, place the sample at room temperature after incubation at 56°C until the heating block reaches 80°C.

[0113] (10) Transfer the lower clear liquid to a new 2 ml microcentrifuge tube.

[0114] (11) Incubate on ice for 3 min, then centrifuge at 20,000 × g (13,500 rpm) for 15 min.

[0115] (12) Transfer the supernatant to a new microcentrifuge tube, taking care not to aspirate the pellet.

[0116] (13) Add DNase Booster Buffer (16 μL or 25 μL) and 10 μL DNase I stock solution, mix by inverting the tube upside down, and centrifuge briefly to remove any residual liquid from the side of the collection tube. Note that mixing should only be performed by gently turning the tube over, and do not create a vortex.

[0117] (14) Incubate at room temperature for 15 minutes.

[0118] (15) Add 320 μL or 500 μL Buffer RBC to adjust the binding conditions and mix the lysate thoroughly.

[0119] (16) Add 720 μL or 1200 μL of 100% ethanol to the sample and mix well by pipetting. Do not centrifuge and proceed immediately to step (17).

[0120] (17) Transfer 700 μL of the sample, including any precipitate that may have formed, to an RNeasy MinElute spin column in a 2 ml collection tube; cap the column and centrifuge at ≥8000 × g (≥10000 rpm) for 15 s; discard the filtrate.

[0121] (18) Repeat step (17) until the entire sample has passed through the RNeasy MinElute spin column. Reuse the collection tube from step (19).

[0122] (19) Add 500 μL of Buffer RPE to the RNeasy MinElute spin column. Gently close the cap and centrifuge at ≥8000 × g (≥10000 rpm) for 15 seconds. Discard the filtrate. Note: Buffer RPE is provided as a concentrate; ensure that ethanol is added before use.

[0123] (20) Add 500 μL of Buffer RPE to the RNeasy MinElute spin column. Gently close the cap and centrifuge at ≥8000 × g (≥10000 rpm) for 2 minutes. Discard the collection tube and the filtrate.

[0124] (21) Place the RNeasy MinElute spin column in a new 2 ml collection tube. Centrifuge at full speed for 5 minutes and discard the collection tube and filtrate. To avoid damage to the caps, place the spin columns in the centrifuge with at least one empty space between the columns. Orient the caps so that they point in the opposite direction of the rotor's rotation (e.g., if the rotor rotates clockwise, orient the caps counterclockwise).

[0125] (22) Place the RNeasy MinElute spin column in a new 1.5 ml collection tube. Add 14–30 μL of nuclease-free water directly onto the spin column membrane. Gently close the cap and centrifuge at full speed for 1 minute to elute the RNA.

[0126] Example 3

[0127] 1. Using metagenomic sequencing technology, analyze the bacterial composition of the patient's intestinal tumor tissue using the DNA or RNA sample extracted in Example 2 and calculate its abundance value. The specific steps are as follows:

[0128] (1) Microbial metagenomic sequencing: The DNA or RNA samples extracted in Example 2 are subjected to metagenomic sequencing to obtain the genome information of each species in the microbial community.

[0129] (2) Sequence preprocessing: First, the adapter sequences introduced during the sequencing process were removed using Cutadapt, and then the raw sequence data were comprehensively evaluated for quality using quality control tools such as FastQC. On this basis, poor quality sequences were screened and eliminated, including low-quality bases and sequences shorter than the minimum length threshold. For samples containing host DNA, alignment tools such as Bowtie2 or BWA were used to align the host DNA sequence with the reference genome (specifically, the sequencing reads were aligned to the database GRCh38, and then the reads that were not aligned to the database GRCh38 were extracted, and the reads that were not aligned to the database GRCh38 were aligned to the database GRCh38, and then the reads that were not aligned to the database GRCh38 were extracted, and the reads that were not aligned to the database GRCh38 were further aligned to the database T2T-CHM13v2.0 to capture those sequences that could not be aligned in the database GRCh38, thereby minimizing data loss and improving the comprehensiveness of the analysis), and all matching host sequences were removed to reduce host contamination and ensure the accuracy and purity of subsequent analysis. This step is crucial to improving data quality and reducing background noise.

[0130] (3) Sequence assembly: Use the assembly software SPAdes to assemble conventional genomic data, or use MetaSPAdes to specifically process metagenomic data. Splice the preprocessed short read sequences into longer contigs or scaffolds to construct a draft genome of the metagenomic genome. Subsequently, use the assembly evaluation software QUAST to evaluate the length, number, N50 value, and completeness of the assembled contigs to ensure the quality of the assembly results.

[0131] (4) Species annotation and classification: Use a reference database, such as NCBI's Nucleotide database or Greengenes database, to align the spliced ​​contigs or scaffolds and perform species annotation and classification based on the alignment results. Identify contigs that are highly similar to the genomes of species in the database and perform species annotation. Count the number of contigs of different species and estimate the relative abundance of species based on the ratio of the number of these contigs. For example, determine the abundance of different species in a sample based on the ratio of the number of contigs, and then perform species classification.

[0132] 2. Machine Learning Model Building

[0133] (1) Calculate the mean abundance difference between sensitive and insensitive microorganisms: Randomly select 14 samples from the total 44 samples in Example 1 (10 sensitive samples and 4 insensitive samples). Extract the abundance data of sensitive and insensitive microorganisms from each of the 14 samples, calculate the mean difference of each microorganism, and select the top 50 microorganisms with the largest absolute difference to prepare for subsequent PCA verification.

[0134] (2) Based on PCA verification, observe whether the selected consistent features can separate the two types of samples: For the 50 screening features of the 14 samples in step (1), the data was standardized and then PCA dimensionality reduction analysis was performed. The PCA results showed that the sensitive group and the insensitive group samples showed a clear separation in the principal component space, which can accurately distinguish the two types of samples. Therefore, it can be considered that the selected microbial features have a strong discriminant ability in distinguishing sensitive from insensitive samples.

[0135] (3) Use Euclidean distance to detect outliers and perform sample removal: At this stage, we use the data after PCA dimensionality reduction to avoid being affected by the original dimension of the data and ensure that the calculated outliers can reflect the anomalies in the main feature space. Use Euclidean distance to detect outliers, compare the distance of the sample with a threshold, identify outliers, use a 95% distance threshold, and establish an outlier index to find outlier samples. After outlier detection based on Euclidean distance, no outliers exceeding the set threshold (such as the 95% distance threshold) were found in the 14 samples. All samples were within the normal range and no sample removal operation characteristics were required. The matrix and label remained unchanged and could be directly used for subsequent data analysis or modeling.

[0136] (4) XGBoost model data reading: Import microbiome feature data and corresponding chemotherapy response labels from standardized CSV files. After importing, all features were Z-score standardized to eliminate the adverse effects of different dimensions on model training. Subsequently, the samples of 14 colorectal cancer patients were divided into training and test sets in proportion: the training set contained 12 samples, of which 9 were sensitive to neoadjuvant chemotherapy and 3 were insensitive; the test set contained 2 samples, 1 sensitive and 1 insensitive, ensuring that the sensitive and insensitive samples in the test set accounted for 50% each, so as to more accurately evaluate the classification performance of the XGBoost model.

[0137] (5) XGBoost model initialization and objective function setting: In this step, the XGBoost model is first initialized, the objective function is set to logarithmic loss, and first-order and second-order regularization terms are added to the objective function to effectively suppress overfitting while improving the model fitting ability.

[0138] (6) Gradient and Hessian matrix calculation: Based on the prediction results of the current XGBoost model, the first-order derivative of the residual, i.e., the gradient, and the second-order derivative, i.e., the Hessian matrix, are calculated for each sample in the training set. These values ​​are used as sample weights for the subsequent evaluation of the split gain.

[0139] (7) Optimal split construction of the decision tree: XGBoost uses a greedy strategy to traverse candidate split points for all features and select the optimal split according to the gain value determined by the gradient and the Hessian matrix. If the gain is lower than the preset threshold, the node is stopped from splitting to reduce the risk of overfitting.

[0140] (8) Model iterative update and learning rate adjustment: In each iteration, XGBoost multiplies the newly constructed weak learner by the learning rate and adds it to the existing model to form a new overall prediction function; then repeats the gradient calculation, optimal split, weak learner construction and model update processes until the preset number of iterations is reached or the early stopping condition is triggered.

[0141] (9) XGBoost feature importance scoring and screening:

[0142] After training, the XGBoost model outputs a score for the importance of each microbial genus to the model's predictions. These scores reflect the frequency of each feature at the split node and its contribution to information gain. The top 15 microbial genera were selected based on feature importance, as they were considered to have significant discriminatory power in predicting chemotherapy response and may serve as potential biomarkers.

[0143] (10) Removal of interference from contaminating microbial genera: To further enhance the biological relevance of the analysis, microorganisms that may be derived from environmental or technological contamination were excluded from these 15 microbial genera, and only target microbial genera known to be associated with human colorectal cancer were retained, including Treponema, Staphylococcus, Streptococcus, Gemella, and Fusobacterium. Based on this screening, the feature importance scores of these target microorganisms in the XGBoost model were re-obtained, and the results are as follows: Figure 1 shown.

[0144] (11) The five selected microbial markers were included as features in the model performance evaluation, and a prediction model was constructed using the random forest algorithm. The model was fitted and trained on the training set to enable it to learn the data features; it was then validated on the test set to test the predictive performance of the random forest algorithm and thus comprehensively evaluate the performance of the model.

[0145] Example 4 Verification of Model Effectiveness

[0146] After removing the 14 samples used for feature screening in Example 3, the remaining 30 samples (21 sensitive samples and 9 insensitive samples) of the total 44 samples in Example 1 were used to independently validate the random forest model to verify the effectiveness of the model.

[0147] The predictive performance of the model was quantified by the receiver operating characteristic (ROC) curve. Specifically, the 30 samples were first randomly divided into a training set (24 cases) and a test set (6 cases). Among them, 20% of the samples (i.e., 6 cases) in the total data set (30 samples) were randomly selected as the test set. In the test set, 3 cases were in the sensitive group (accounting for 50%) and 3 cases were in the insensitive group (accounting for 50%) to evaluate the model's ability to distinguish the sensitivity of colorectal cancer patients to neoadjuvant chemotherapy. The training set was used to construct a random forest prediction model, and finally the test set was input into the model to predict the sensitivity of colorectal cancer patients to neoadjuvant chemotherapy, draw the ROC curve, and calculate the area under the curve (AUC). At the same time, a confusion matrix was generated to comprehensively evaluate the classification performance of the model. The specific steps are as follows:

[0148] (1) Detection of relative abundance of intestinal microbial markers: First, the samples were subjected to metagenomic sequencing according to the method in Example 2 to determine the relative abundance of five specific bacterial species. These data will be input into the constructed random forest model as key features, providing a basis for subsequent modeling and analysis.

[0149] (2) Data reading: First, the data reading stage imports microbiome features and corresponding clinical response labels from standardized CSV files; then data preprocessing is performed, which mainly includes the processing of missing values, such as using median interpolation or deleting features with severe missing values, and standardization of features, such as Z-score standardization, to eliminate dimensionality effects and improve model stability.

[0150] (3) Data partitioning: Next, the preprocessed data is divided into a training set and a test set in a ratio of 8:2, so that the performance of the model on unseen data can be objectively measured during model evaluation.

[0151] (4) Random Forest Model Training: During the model training phase, the random forest algorithm is introduced to improve the robustness and accuracy of the model by constructing a large number of decision trees and combining their output results for integrated prediction. In order to improve the ability to identify class imbalance, the model introduces the setting of class weights to strengthen the ability to identify minority classes, such as patients who respond to treatment.

[0152] (5) Model parameter tuning: During the model parameter tuning process, a grid search method is used to systematically search for multiple hyperparameter combinations, such as the number of trees, maximum depth, minimum number of split samples, etc., to determine the optimal parameter combination and avoid overfitting.

[0153] (6) Model performance evaluation: In terms of model performance evaluation, a combination of multiple indicators such as ROC curve, AUC value, confusion matrix and classification report is used to comprehensively measure the classification ability and generalization performance of the model. The ROC curve is used to visualize the balance between sensitivity and specificity of the model at different thresholds, and the AUC value provides an overall performance score; the confusion matrix can reveal the number of TP, FP, TN, and FN in the prediction results, and the classification report is further refined into indicators such as precision, recall rate and F1 score.

[0154] The verification results are as follows:

[0155] Figure 2 The confusion matrix shown shows that the model has excellent classification accuracy and can effectively distinguish between positive and negative samples. For positive samples, the model shows high sensitivity and can accurately identify most positive samples; at the same time, the specificity is good, and it rarely misclassifies negative samples as positive. When verifying the 6 samples in the test set, the number of both false positives and false negatives is small. Figure 3 As shown, the ROC curve AUC value of the model on the training set is as high as 0.992 ( Figure 3 a in the figure), the AUC value on the test set also reached 0.889 ( Figure 3(b) This indicates that the model's predictive performance is excellent, with strong ability to distinguish between positive and negative samples. In summary, the microbial marker combination and the constructed predictive model provided in this application have significant efficacy and can provide an accurate basis for predicting the sensitivity of colorectal cancer patients to neoadjuvant chemotherapy.

[0156] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0157] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. Use of a reagent for detecting the abundance of a combination of microbial markers in a sample in the preparation of a product for predicting the sensitivity of colorectal cancer to neoadjuvant chemotherapy, characterized in that: The microbial marker combination is Treponema, Staphylococcus, Streptococcus, Gemella and Fusobacterium.

2. The use according to claim 1, characterized in that The sample is tumor tissue from a colorectal cancer patient.

3. The use according to claim 2, characterized in that The sample is a paraffin-embedded block of tumor tissue from a patient with colorectal cancer.

4. The use according to claim 1, wherein The product described is an in vitro detection product.

5. The use according to claim 4, characterized in that The in vitro detection product is a kit.

6. A system for predicting the sensitivity of colorectal cancer to neoadjuvant chemotherapy, characterized in that: include: a data input module for inputting abundance data of a microbial marker combination, wherein the microbial marker combination is Treponema, Staphylococcus, Streptococcus, Gemella, and Fusobacterium; a data storage module for storing abundance data of the microbial marker combination in biological samples of a population, the population comprising colorectal cancer patients receiving neoadjuvant chemotherapy; An output prediction module is connected to the data input module and the data storage module respectively, and uses the abundance data of the microbial marker combination in the biological samples of the group to build a prediction model, and based on the abundance data of the microbial marker combination input by the data input module, obtains and outputs the prediction results of the sensitivity of colorectal cancer patients to neoadjuvant chemotherapy.

7. The system according to claim 6, wherein: The abundance of the microbial marker combination is obtained by metagenomic sequencing, qPCR detection or 16S sequencing.

8. The system according to claim 7, wherein: The abundance of the microbial marker combination is obtained by metagenomic sequencing.

9. The system according to claim 6, wherein: The construction of the prediction model includes the following steps: The abundance data of the microbial marker combination in the biological samples of the group are randomly divided into two groups, one group is a training set, and the other group is a test set; Using training set data, we build a prediction model based on machine learning algorithms and conduct multiple verifications; The obtained prediction model is verified in the test set.

10. The system according to claim 9, wherein: The machine learning algorithm is a random forest algorithm.

11. A computer device comprising a memory for storing a computer program and a processor for executing the computer program, characterized in that: Also comprises a system as described in any one of claims 6-10.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that: Also comprises a system as described in any one of claims 6-10.

Citation Information

Patent Citations

  • Biomarker for identifying community acquired pneumonia and application thereof

    CN115976198A