A microbial marker for detecting tumor radiotherapy sensitivity and its application

CN119220674BActive Publication Date: 2025-09-09FUJIAN MEDICAL UNIV UNION HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411321368.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-23
Publication Date
2025-09-09
Estimated Expiration
2044-09-23

AI Technical Summary

Technical Problem

然而,TRG分级未能在治疗前对结直肠癌患者的放射反应进行敏感度分级

Benefits of technology

[0019] The present invention provides a microbial marker for detecting tumor radiotherapy sensitivity and its application. The present invention utilizes the microbial marker to effectively distinguish the tumor radiotherapy sensitivity of cancer patients, and can also be used for grading and identifying the tumor radiotherapy sensitivity of cancer patients, thereby improving the accuracy of judging the radiotherapy sensitivity of cancer patients and providing an important theoretical basis for the efficient use of radiotherapy to treat cancer patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119220674B_ABST
    Figure CN119220674B_ABST
Patent Text Reader

Abstract

The present invention provides a microbial marker for detecting tumor radiotherapy sensitivity and its application, belonging to the technical field of tumor radiotherapy sensitivity diagnosis. The present invention utilizes microbial markers such as s__Roseburia_inulinivorans, s__Prevotella_stercorea, s__Anaerostipes_unclassified, s__Clostridioides_difficile, s__Prevotella_nigrescens, s__Schaalia_odontolytica, and s__Alistipes_obesi to effectively differentiate the tumor radiotherapy sensitivity of cancer patients. The markers can also be used for grading and identifying the tumor radiotherapy sensitivity of cancer patients, thereby improving the accuracy of determining the degree of radiotherapy sensitivity of cancer patients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of tumor radiotherapy sensitivity diagnosis, and in particular relates to a microbial marker for detecting tumor radiotherapy sensitivity and an application thereof. Background Art

[0002] Colorectal cancer accounts for approximately 10% of all cancers diagnosed and cancer-related deaths worldwide each year. It is the second most common cancer in women and the third most common cancer in men. Radiotherapy, as an important component of comprehensive colorectal cancer treatment, and preoperative chemoradiotherapy remain the standard treatment for locally advanced rectal cancer (stages II and III). However, while radiotherapy offers therapeutic benefits, it also carries increased toxic side effects, resulting in a relatively small proportion of patients benefiting from it, and many challenges remain to be addressed.

[0003] In the field of radiotherapy, clinical efficacy varies significantly between individuals. Even with the same radiotherapy regimen and dose, approximately half of patients still experience no significant tumor regression or even progression. The resistance of colorectal cancer (CRC) cells to radiotherapy contributes significantly to the poor clinical prognosis of CRC patients. However, the identification of CRC patients with high radiosensitivity (HS), moderate radiosensitivity (MS), and low radiosensitivity (LS) remains a challenge. Tumor regression grading (TRG) is widely accepted for assessing radiotherapy response, and previous studies have shown that the AJCC-TRG grading accurately predicts prognosis after chemoradiotherapy for rectal cancer. However, the TRG grading fails to stratify the sensitivity of CRC patients to radiotherapy before treatment. Therefore, developing a method to screen tumor radiosensitivity markers for predicting radiosensitivity is of great significance. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a microbial marker for detecting tumor radiotherapy sensitivity and its application. The present invention utilizes the microbial marker to effectively detect the patient's sensitivity to tumor radiotherapy.

[0005] In order to achieve the above-mentioned object of the invention, the present invention provides the following technical solutions:

[0006] The present invention provides a microbial marker for detecting tumor radiotherapy sensitivity, wherein the microbial marker includes one or more of s__Schaalia_odontolytica, s__Roseburia_inulinivorans, s__Mogibacterium_diversum, s__Campylobacter_gracilis, s__Alistipes_obesi, s__bacterium_New, s__Prevotella_nigrescens, s__Clostridioides_difficile, s__Bacteroides_clarus, s__Enterobacter_cloacae, s__Anaerostipes_unclassified, and s__Prevotella_stercorea.

[0007] The present invention also provides a use of the above-mentioned microbial marker in the preparation of a product for diagnosing tumor radiotherapy sensitivity.

[0008] The present invention also provides an application of the above-mentioned microbial marker in the preparation of a product for diagnosing tumor radiotherapy sensitivity grading.

[0009] Preferably, the tumor radiotherapy sensitivity classification is tumor radiotherapy high sensitivity and tumor radiotherapy moderate sensitivity.

[0010] The present invention also provides a system for predicting tumor radiotherapy sensitivity using the above-mentioned microbial markers, comprising a nucleic acid sample separation unit for separating a fecal microbial nucleic acid sample from a test sample;

[0011] The data processing unit establishes a logistic regression model using the above-mentioned microbial markers to obtain a critical value;

[0012] The result determination unit is used to compare the critical value obtained by the data processing unit with the set diagnostic value.

[0013] Preferably, the logistic regression model is a first logistic regression model and a second logistic regression model;

[0014] The first logistic regression model is a logistic regression model constructed by comparing the stool sample to be tested with the stool sample of the colorectal cancer patient who is not sensitive to radiotherapy. When the critical value is less than 0.62, it indicates that the tumor patient is not sensitive to radiotherapy. When the critical value is greater than or equal to 0.62, it indicates that the tumor patient is moderately sensitive to radiotherapy or highly sensitive to radiotherapy.

[0015] The second logistic regression model is a logistic regression model constructed by comparing the stool samples to be tested with a critical value ≥0.62 with the stool samples of colorectal cancer patients with moderate sensitivity to radiotherapy. When the critical value is <0.48, it indicates that the tumor patient is moderately sensitive to radiotherapy, and when the critical value is ≥0.48, it indicates that the tumor patient is highly sensitive to radiotherapy.

[0016] Preferably, when the first logistic regression model is established using the above-mentioned microbial markers, the microbial markers are s__Schaalia_odontolytica, s__Roseburia_inulinivorans, s__Mogibacterium_diversum, s__Campylobacter_gracilis, s__Alistipes_obesi, s__Prevotella_nigrescens, s__Clostridioides_difficile, s__Bacteroides_clarus, s__Enterobacter_cloacae, s__Anaerostipes_unclassified, and s__Prevotella_stercorea.

[0017] Preferably, when the second logistic regression model is established using the above-mentioned microbial markers, the microbial markers are s__Roseburia_inulinivorans, s__Prevotella_stercorea, s__Anaerostipes_unclassified, s__Clostridioides_difficile, s__Prevotella_nigrescens, s__Schaalia_odontolytica and s__Alistipes_obesi.

[0018] Compared with the prior art, the present invention has the following beneficial effects:

[0019] The present invention provides a microbial marker for detecting tumor radiotherapy sensitivity and its application. The present invention utilizes the microbial marker to effectively distinguish the tumor radiotherapy sensitivity of cancer patients, and can also be used for grading and identifying the tumor radiotherapy sensitivity of cancer patients, thereby improving the accuracy of judging the radiotherapy sensitivity of cancer patients and providing an important theoretical basis for the efficient use of radiotherapy to treat cancer patients. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is the 16S rDNA region of bacteria;

[0021] Figure 2 This is the PacBio sequencing principle;

[0022] Figure 3 It is the PacBioCCS model;

[0023] Figure 4 This is the full-length 16S experimental process;

[0024] Figure 5 ROC curves were drawn based on the logistic regression model fitting method. A is the ROC curve for detecting stool samples to be tested versus stool samples from colorectal cancer patients who were not sensitive to radiotherapy; B is the ROC curve for stool samples to be tested with P ≥ 0.62 versus stool samples from colorectal cancer patients who were moderately sensitive to radiotherapy;

[0025] Figure 6 To draw an ROC curve based on the microbial community involved in the logistic regression model equation, A is the ROC curve drawn for the microbial community involved in the fecal sample to be tested vs. the fecal sample of the colorectal cancer patient who is not sensitive to radiotherapy using the microbial community marker detection in Example 2; B is the ROC curve drawn for the microbial community involved in the fecal sample to be tested with P ≥ 0.62 vs. the fecal sample of the colorectal cancer patient who is moderately sensitive to radiotherapy using the microbial community marker detection in Example 2. DETAILED DESCRIPTION

[0026] The present invention provides a microbial marker for detecting tumor radiotherapy sensitivity, wherein the microbial marker includes one or more of s__Schaalia_odontolytica, s__Roseburia_inulinivorans, s__Mogibacterium_diversum, s__Campylobacter_gracilis, s__Alistipes_obesi, s__bacterium_New, s__Prevotella_nigrescens, s__Clostridioides_difficile, s__Bacteroides_clarus, s__Enterobacter_cloacae, s__Anaerostipes_unclassified, and s__Prevotella_stercorea.

[0027] In the present invention, the microbial markers are preferably s__Schaalia_odontolytica, s__Roseburia_inulinivorans, s__Mogibacterium_diversum, s__Campylobacter_gracilis, s__Alistipes_obesi, s__Prevotella_nigrescens, s__Clostridioides_difficile, s__Bacteroides_clarus, s__Enterobacter_cloacae, s__Anaerostipes_unclassified, and s__Prevotella_stercorea. The present invention uses the above 11 microbial markers to effectively diagnose whether a patient is sensitive to tumor radiotherapy. Furthermore, after the above-mentioned 11 microbial markers are used to effectively diagnose patients who are sensitive to tumor radiotherapy, in order to determine the degree of sensitivity of the above-mentioned patients to tumor radiotherapy, the microbial markers used are preferably s__Roseburia_inulinivorans, s__Prevotella_stercorea, s__Anaerostipes_unclassified, s__Clostridioides_difficile, s__Prevotella_nigrescens, s__Schaalia_odontolytica and s__Alistipes_obesi. The present invention uses these 7 microbial markers to effectively diagnose whether the patient is highly sensitive to tumor radiotherapy or moderately sensitive to tumor radiotherapy.

[0028] Based on this, the present invention also provides an application of the above-mentioned microbial marker in the preparation of a product for diagnosing tumor radiotherapy sensitivity.

[0029] The present invention also provides an application of the above-mentioned microbial marker in the preparation of a product for diagnosing tumor radiotherapy sensitivity grading.

[0030] In the present invention, the tumor radiotherapy sensitivity is graded into highly sensitive tumor radiotherapy and moderately sensitive tumor radiotherapy. The present invention divides tumor radiotherapy sensitivity into highly sensitive tumor radiotherapy, moderately sensitive tumor radiotherapy, and low sensitivity to tumor radiotherapy (also known as insensitive tumor radiotherapy) based on the NAR (Neoadjuvant Rectal) score. When the NAR score is less than 8, it means that the patient is highly sensitive to tumor radiotherapy; when the NAR score is 8-16, it means that the patient is moderately sensitive to tumor radiotherapy; when the NAR score is greater than 16, it means that the patient is insensitive to tumor radiotherapy.

[0031] The present invention also provides a system for predicting tumor radiotherapy sensitivity using the above-mentioned microbial markers, comprising a nucleic acid sample separation unit for separating a fecal microbial nucleic acid sample from a test sample;

[0032] The data processing unit establishes a logistic regression model using the above-mentioned microbial markers to obtain a critical value;

[0033] The result determination unit is used to compare the critical value obtained by the data processing unit with the set diagnostic value.

[0034] In the present invention, the critical value is also referred to as the optimal threshold for predicting P. In the present invention, the logistic regression model is preferably a first logistic regression model or a second logistic regression model. When determining whether the stool sample is sensitive to tumor radiotherapy, the diagnostic value is 0.62. When determining that the stool sample is sensitive to tumor radiotherapy and further determining the degree of sensitivity to tumor radiotherapy, the diagnostic value is 0.48.

[0035] In the present invention, the first logistic regression model is a logistic regression model constructed by comparing the stool sample to be tested with the stool sample of the colorectal cancer patient who is not sensitive to radiotherapy. When the critical value is <0.62, it indicates that the tumor patient is not sensitive to radiotherapy. When the critical value is ≥0.62, it indicates that the tumor patient is moderately sensitive to radiotherapy or highly sensitive to radiotherapy.

[0036] The second logistic regression model is a logistic regression model constructed by comparing the stool samples to be tested with a critical value ≥0.62 with the stool samples of colorectal cancer patients with moderate sensitivity to radiotherapy. When the critical value is <0.48, it indicates that the tumor patient is moderately sensitive to radiotherapy, and when the critical value is ≥0.48, it indicates that the tumor patient is highly sensitive to radiotherapy.

[0037] In the present invention, when the first logistic regression model is established using the above-mentioned microbial markers, the microbial markers are preferably s__Schaalia_odontolytica, s__Roseburia_inulinivorans, s__Mogibacterium_diversum, s__Campylobacter_gracilis, s__Alistipes_obesi, s__Prevotella_nigrescens, s__Clostridioides_difficile, s__Bacteroides_clarus, s__Enterobacter_cloacae, s__Anaerostipes_unclassified, and s__Prevotella_stercorea.

[0038] In the present invention, when the second logistic regression model is established using the above-mentioned microbial markers, the microbial markers are preferably s__Roseburia_inulinivorans, s__Prevotella_stercorea, s__Anaerostipes_unclassified, s__Clostridioides_difficile, s__Prevotella_nigrescens, s__Schaalia_odontolytica and s__Alistipes_obesi.

[0039] The present invention utilizes the above-mentioned microbial flora markers to effectively diagnose whether a tumor patient is sensitive to tumor radiotherapy and the degree of sensitivity to tumor radiotherapy, and the diagnosis is highly accurate.

[0040] In the present invention, unless otherwise specified, all raw material components are commercially available products well known to those skilled in the art.

[0041] The technical solutions provided by the present invention are described in detail below with reference to the embodiments, but they should not be construed as limiting the scope of protection of the present invention.

[0042] Example 1

[0043] A method for screening and detecting bacterial biomarkers of tumor radiotherapy sensitivity, comprising the following steps:

[0044] 1.1 Fecal DNA extraction

[0045] (1) Add 0.25-0.5 g of sample to a 2 mL centrifuge tube. If the sample is liquid, transfer 200 μL to the centrifuge tube. Add 500 μL of buffer SA, 100 μL of buffer SC, and 0.25 g of grinding beads. (Fecal samples may contain residual RNA. If RNA needs to be removed, it is recommended to add 10 μL of RNase A (TIANGEN, RT405-02, self-prepared). Vortex and mix thoroughly or use a tissue homogenizer to mix thoroughly. Heat and lyse the sample at 70°C for 15 minutes to improve lysis efficiency. Centrifuge at 12,000 rpm for 1 minute. Transfer the first supernatant (about 500 μL) to a new 2 mL centrifuge tube.

[0046] (2) Add 200 μL of buffer SH to the first supernatant, mix well, vortex for 5 seconds, and place at 4°C for 10 minutes.

[0047] (3) Centrifuge at 12000 rpm for 3 min, transfer the second supernatant to a new 2 mL centrifuge tube, add 500 μL of buffer GFA (please check whether isopropanol has been added before use), and mix by inversion.

[0048] (4) Add 10 μL of magnetic bead suspension G and shake to mix for 5 minutes.

[0049] (5) Place the centrifuge tube on the magnetic rack and let it stand for 30 seconds. After the magnetic beads are completely adsorbed, carefully remove the liquid.

[0050] (6) Remove the centrifuge tube from the magnetic rack, add 700 μL of deproteinized solution RD (please check whether anhydrous ethanol has been added before use), and shake to mix for 5 minutes.

[0051] (7) Place the centrifuge tube on the magnetic rack and let it stand for 30 seconds. After the magnetic beads are completely adsorbed, carefully remove the liquid.

[0052] (8) Remove the centrifuge tube from the magnetic stand, add 700 μL of rinse solution PWD (please check whether anhydrous ethanol has been added before use), and shake to mix for 3 minutes.

[0053] (9) Place the centrifuge tube on the magnetic rack and let it stand for 30 seconds. After the magnetic beads are completely adsorbed, carefully remove the liquid.

[0054] (10) Repeat steps (8) and (9) once.

[0055] (11) Place the centrifuge tube on a magnetic rack and dry it at room temperature for 5 to 10 minutes.

[0056] (12) Remove the centrifuge tube from the magnetic rack, add 50-100 μL of elution buffer TB, shake and mix, place at 56°C, and incubate for 5 minutes. Shake and mix three times, 3-5 times each time.

[0057] (13) Place the centrifuge tube on a magnetic stand and let it stand for 2 minutes. After the magnetic beads are completely adsorbed, carefully transfer the DNA solution to a new centrifuge tube and store it under appropriate conditions to obtain fecal DNA.

[0058] 1.2 Introduction to full-length amplicon

[0059] 16S rDNA is a gene that encodes the ribosome 30S small subunit component (16S rRNA) in prokaryotes. It is about 1500bp in length and includes 9 variable regions (V1-V9) and 10 conserved regions. The conserved regions reflect the relationship between species, while the variable regions reflect the differences between species. The degree of variation is closely related to the bacterial phylogeny and is considered to be the most suitable indicator for bacterial phylogeny and classification identification. Figure 1 .

[0060] Unlike the second-generation amplicons that only target 1-2 hypervariable regions, the third-generation amplicons sequence the full length of 16S (the amplification primer pair is 27F: AGRGTTYGATYMTGGCTCAG (SEQ ID NO. 1) and 1492R: RGYTACCTTGTTACGACTT (SEQ ID NO. 2)), where R represents A / G, Y represents C / T, and M represents A / C, avoiding the influence of different hypervariable regions and PCR preferences. More hypervariable region information can significantly improve the resolution and accuracy of species annotation, and has been widely used in the study of microbial communities in different habitats.

[0061] 1.3 Introduction to Sequencing Principles

[0062] PacBio (Pacific BioSciences)'s Single Molecular RealTime (SMRT) technology has the advantages of long read length and high accuracy; the Sequel system is currently the mainstream third-generation sequencing platform. Its SMRT cell for sequencing reaction has a large number of nanoscale holes, called Zero-Mode Waveguides (ZMWs). Typically, each hole fixes a template DNA and a molecule of DNA polymerase. The sequence information is determined by detecting the fluorescent signal of dNTPs during chain extension. The principle of PacBio sequencing can be found in Figure 2 .

[0063] The Circular Consensus Sequencing (CCS) mode sequences each template molecule multiple times and obtains accurate sequence information through subread correction, namely CCS reads (also known as HiFi reads); it is currently mainly used for libraries with insert fragments less than 2kb, which is highly consistent with the requirements of 16S rDNA full-length sequencing. Figure 3 .

[0064] 2. Workflow

[0065] 2.1 Experimental Procedure

[0066] In order to ensure the accuracy and reliability of sequencing data from the source, there are rigorous and reliable sample testing and quality control processes from DNA extraction to sequencing. The sample quality is strictly controlled at every step to ensure the authenticity and credibility of the sequencing data. Figure 4 .

[0067] 2.2 Information Analysis Process

[0068] After the on-machine sequencing is completed, the original off-machine data Raw CCS will be quality controlled to obtain high-quality ValidCCS. In order to meet the needs of different customers, two different analysis modes can be provided in the future. One is based on Qiime2's DADA2 (Divisive Amplicon Denoising Algorithm), which obtains representative sequences ASVs (Amplicon Sequence Variants) with single-base accuracy through "dereplication"; the other is to use VSEARCH for OTUs (Operational Taxonomic Units) clustering (97% similarity). At present, the former is more widely recognized and recommended due to its advantages such as higher phylogenetic resolution and comparability between different data sets. The project uses the ASVs method by default. After obtaining the ASVs / OTUs table, further diversity analysis, species classification annotation and difference analysis are carried out.

[0069] 3. Standard Analysis

[0070] 3.1 Sequencing data quality control statistics

[0071] After sequencing, first obtain the Raw CCS (circular consensus sequencing), and then perform further data quality control to obtain the Valid CCS. The steps are as follows:

[0072] (1) SMRT Links identification to obtain Raw CCS, parameters: minPasses ≥ 3, minPredictedAccuracy ≥ 0.99;

[0073] (2) limma uses default parameters to identify barcode split samples;

[0074] (3) cutadapt identified and removed the primers and then performed length screening, retaining reads of 1200-1650 bp;

[0075] (4) The final Valid CCS was denoised and chimeras were filtered using the qiime dada2 denoise-single method to obtain the ASV (feature) feature sequence and abundance table, and singletons ASVs were removed to obtain the processed Valid CCS.

[0076] 3.2 Screening of markers

[0077] The processed Valid CCS is then used to screen markers.

[0078] The processed Valid CCS were classified into 6 categories according to the NAR (Neoadjuvant Rectal) score and before and after radiotherapy: prHS, prMS, prLS, poHS, poMS and poLS. Specifically, prHS stands for sensitive data before radiotherapy (highly sensitive bacterial community subset before radiotherapy), prMS stands for moderately sensitive data before radiotherapy (moderately sensitive bacterial community subset before radiotherapy), prLS stands for insensitive data before radiotherapy (insensitive bacterial community subset before radiotherapy), poHS stands for sensitive data after radiotherapy (highly sensitive bacterial community subset after radiotherapy), poMS stands for moderately sensitive data after radiotherapy (moderately sensitive bacterial community subset after radiotherapy), and poLS stands for insensitive data after radiotherapy (insensitive bacterial community subset after radiotherapy), which were used as the initial sample set.

[0079] Step S1: The initial sample set is processed using principal component analysis and principal coordinate analysis (PCoA) to obtain an initial differential flora sample set; the initial differential flora sample set is preliminarily screened using a rank sum test to obtain a second differential flora sample set.

[0080] First, establish the null hypothesis: the two groups of samples come from the same population, mix the two groups of data, and rank them uniformly from small to large. When encountering the same data, take the average rank and calculate the statistic. Set the significance level a = 0.05, and obtain the probability of the current statistic value under the condition that the null hypothesis is true. If p-value < 0.05, reject the null hypothesis, indicating that there is a significant difference between the two samples.

[0081] Step S2: Based on the logistic regression model and P-value established between the differential samples in the second differential flora sample set and the significant clinical indicators, the differential flora after eliminating the interference of clinical factors is obtained, that is, the third differential flora sample set is obtained.

[0082] The initial clinical indicators were screened using multiple comparisons and PerMANOVA multivariate analysis of variance, and significant clinical indicators were obtained, including:

[0083] a. Obtaining initial clinical indicators:

[0084] A total of 271 clinical samples were collected, and 16 clinical indicator information were extracted, including treatment response group, age, gender, BMI, smoking, drinking, T, N, GTV, cm from the anus (MR / CT), CRM (MR), EMVI (MR), NLR, concurrent chemotherapy regimen, hormone use, and antibiotic use.

[0085] b. Screening the initial clinical indicators using the multiple comparison method to determine the second clinical indicator:

[0086] The multiple comparison method was used to retain samples with a p-value < 0.05 in the statistical test results of the above 16 clinical indicators. The p-value was obtained by performing multiple comparisons on the clinical indicators among multiple experimental groups. First, the homogeneity was tested by Bartlett Test (corresponding to the R language program function bartlett.test) (taking a p-value < 0.05 as a non-null hypothesis), and then the p-value was calculated by the oneway.test function of the R language program package. If p < 0.05, it indicated that the clinical indicator was significant.

[0087] c. Use the PerMANOVA multivariate analysis of variance method to screen the second clinical indicators and determine significant clinical indicators.

[0088] The permaonva analysis method was used to further screen out clinical indicator information that had a significant relationship with the expression of the microbiome profile of human fecal samples (i.e., the second clinical indicator);

[0089] d. Use the logistic regression model to use the significant clinical indicator information obtained in step c as the response variable to perform a final correction on the second differential microbial population sample set obtained in step S3 above, and screen out the differential microbial populations whose p-value after regression is still less than 0.05, to obtain the differential microbial populations after eliminating the interference of clinical factors (the third differential microbial population sample set), specifically including:

[0090] The clinically significant indicator information obtained in step c was used as the response variable, and the second differential substance bacterial cluster obtained in step S3 was used as the independent variable. A logistic regression model was established to obtain the p-value. If the p-value of the differentially expressed bacterial cluster was less than 0.05, it indicated that these bacterial clusters were not affected by the clinical indicators, that is, the differential bacterial clusters after eliminating the interference of clinical factors.

[0091] e. The differential microbiome obtained by the four comparison strategies of prHS and prMS, prHS and prLS, prMS and prLS, and prHS+prMS and prLS after excluding the interference of clinical factors was merged to form the third differential microbiome sample set. Three different machine learning algorithms, random forest method, support vector machine method, and LASSO, were used for further screening and analysis. Based on the key candidate differential microbiome combinations finally screened out by these three different algorithms, a joint ROC analysis was carried out. Finally, the final differential microbiome combination panel was determined according to the obtained AUC value. Specifically,

[0092] First, the third differential bacterial community sample set was divided into a training set and a test set (the ratio of training set to test set was 8:2). Different algorithms were used to screen differential bacterial communities, as follows:

[0093] Random Forest Method: The R package randomForest was used to fit a random forest model to the training set. Parameter tuning ultimately determined the treeNum in the fitted model. The training set was fitted with the optimal parameters, and the model was tested on the test set to obtain a confusion matrix for the prediction results. Furthermore, a 10-fold cross-validation was performed using the rfcv function in the R package randomForest to determine the relationship between variable selection and the average classification error rate of the model. The optimal differential bacterial community was obtained using MeanDecreaseAccuracy, where a larger value indicates greater importance of the variable.

[0094] Support Vector Machine (SVM): The support vector machine is a two-class classification model whose basic model is a linear classifier that searches for a separating hyperplane with the largest margin in feature space. SVM is an algorithm for sequential post-selection based on the maximum margin principle. In the first iteration, the SVM model is optimized for all feature sets in the dataset. The score of each feature is then calculated and sorted in descending order. The feature set with the lowest score is recorded, and the feature with the lowest score is deleted. The process iterates again until only one feature remains. The training set is trained using the R language svm function to determine the relationship between variable selection and the average classification error rate of the model, and the feature set with the lowest error rate is selected.

[0095] LASSO: LASSO applies a regression penalty to all variables, reducing the coefficients of relatively unimportant variables to 0, thus excluding them from the model. It then selects independent variables that have a significant impact on the dependent variable and calculates the corresponding regression coefficients, ultimately resulting in a predictive model. The R language glmnet is used to train the model on the training set. The lambda value with the smallest mean square error is selected based on the mean square error under different parameters lambda. A vertical line is drawn through the corresponding penalty value on the trajectory of each independent variable coefficient. Each curve represents a variable. The variable that intersects with the penalty value is the variable ultimately included in the model. The vertical coordinate corresponding to the variable is the regression coefficient of the variable, which also represents the contribution of the variable.

[0096] Secondly, the differential bacterial communities were screened out through random forest, support vector machine and LASSO. The data of these differential bacterial communities were found in the normalized data. The differentially expressed bacterial communities were used as the response variables, and the groups were divided into 1 as the experimental group and 0 as the control group as the dependent variables to establish a logistic regression model.

[0097] Finally, the probability given by logistic regression was calculated to draw the joint ROC curve. The larger the AUC value, the better the classification effect of the model. The optimal differential bacterial community combination panel was determined based on the size of the AUC.

[0098] Step S3: Using the optimal differential flora combination panel, combined with the k-fold cross-validation method, the logistic regression model, the random forest model, the support vector machine and the LASSO model are trained respectively to obtain a trained logistic regression model, a trained random forest model, a trained support vector machine and a trained LASSO model.

[0099] First, the optimal differential microbial community combination panels were determined for the three different comparison strategies of prHS and prMS, prHS and prLS, and prMS and prLS. Then, the optimal differential microbial community combination panels were merged and randomly sampled into training and test sets, with a training set:test set ratio of 7:3.

[0100] Secondly, for the training sets of different comparison strategies, the caret package was used to perform 10-fold cross-validation on the logistic regression, LASSO, support vector machine, and random forest models, respectively, with 1000 repetitions. The cross-validation was mainly performed using the trainControl function in the caret package in the R language.

[0101] The steps involved in the 10-fold cross validation are:

[0102] (1) Randomly divide the dataset into 10 subsets, (2) For each selected subset of data points, use that subset as a validation set; use all remaining subsets for training purposes; train the model and evaluate it on the validation set or test set; calculate the prediction error, (3) repeat the above steps 10 times, (4) generate the overall prediction error by taking the average of the prediction errors.

[0103] Step S4: Using ROC analysis, according to the AUC value, the train function is used to train the logistic regression, LASSO, support vector machine, and random forest models respectively, and the prediction probability of the test set under the trained model is calculated. At the same time, the ROC curves of different models are drawn. The larger the AUC value, the better the model effect. The optimal model is determined. Finally, the microbial communities obtained by the logistic regression, LASSO, support vector machine, and random forest models are merged to obtain microbial community markers for screening and detecting tumor radiotherapy sensitivity.

[0104] Finally, the following 12 microbial combinations were obtained: s__Schaalia_odontolytica, s__Roseburia_inulinivorans, s__Mogibacterium_diversum, s__Campylobacter_gracilis, s__Alistipes_obesi, s__bacterium_New, s__Prevotella_nigrescens, s__Clostridioides_difficile, s__Bacteroides_clarus, s__Enterobacter_cloacae, The panel composed of s__Anaerostipes_unclassified and s__Prevotella_stercorea had the largest AUC (area under the curve) value. The AUC value of the 12 differential bacterial communities in prHS+prMSvsprLS reached 0.913, the AUC value in prHSvsprLS reached 0.890, the AUC value in prHSvsprMS reached 0.552, and the AUC value in prMSvsprLS reached 0.798, revealing that the 12 intestinal bacterial communities have potential predictive value for the radiotherapy sensitivity of patients with colorectal cancer, and are expected to be used as fixed markers for prediction by collecting patient feces and performing 16srRNA high-throughput sequencing analysis.

[0105] Example 2

[0106] In this example, tumor radiotherapy sensitivity is detected using the microbial markers screened in Example 1. The present invention predicts the radiotherapy sensitivity of colorectal cancer patients using the critical value of a logistic regression model constructed using the microbial markers described in Example 1 (s__Schaalia_odontolytica, s__Roseburia_inulinivorans, s__Mogibacterium_diversum, s__Campylobacter_gracilis, s__Alistipes_obesi, s__Prevotella_nigrescens, s__Clostridioides_difficile, s__Bacteroides_clarus, s__Enterobacter_cloacae, s__Anaerostipes_unclassified, s__Prevotella_stercorea).

[0107] Among them, the reagents are divided into stool samples to be tested, stool samples of colorectal cancer patients who are insensitive to radiotherapy, and stool samples of colorectal cancer patients who are moderately sensitive to radiotherapy.

[0108] Determination method: Metaboanalyst6.0 (https: / / www.metaboanalyst.ca / ) was used to construct a logistic regression model to predict the optimal threshold of P for stool samples of patients with colorectal cancer who were not sensitive to radiotherapy. When P < 0.62, it indicated that the tumor patient was not sensitive to radiotherapy. When P ≥ 0.62, it indicated that the tumor patient was moderately sensitive to radiotherapy or highly sensitive to radiotherapy.

[0109] To further determine whether cancer patients are moderately sensitive or highly sensitive to radiotherapy, a logistic regression model was constructed by comparing stool samples with P ≥ 0.62 with stool samples of colorectal cancer patients with moderate sensitivity to radiotherapy to predict the optimal threshold of P. When P < 0.48, it indicates that the cancer patient is moderately sensitive to radiotherapy, and when P ≥ 0.48, it indicates that the cancer patient is highly sensitive to radiotherapy.

[0110] The inclusion criteria for the study subjects (tumor patients) included: (1) patients who received neoadjuvant radiotherapy or chemoradiotherapy followed by surgical treatment; (2) patients whose radiotherapy regimen was long-term radiotherapy. The exclusion criteria included: (1) patients with any other type of malignant tumor or severe immune, neurological, digestive, or hematological system diseases; (2) patients whose plasma samples could not be obtained; (3) patients with colorectal adenocarcinoma that had metastasized; and (4) patients whose radiotherapy regimen was short-term radiotherapy.

[0111] In this example, 22 stool samples were tested, 31 stool samples from patients with colorectal cancer insensitive to radiotherapy, and 60 stool samples from patients with colorectal cancer moderately sensitive to radiotherapy. Among them, patients with colorectal cancer insensitive to radiotherapy refer to those with a NAR score > 16, and patients with colorectal cancer moderately sensitive to radiotherapy refer to those with a NAR score of 8 to 16. The NAR score is significantly correlated with overall survival (OS).

[0112] The logistic regression model for the stool samples to be tested vs. the stool samples of colorectal cancer patients who are not sensitive to radiotherapy is logit(P)=log(P / (1-P))=0.004+0.188

[0113] s__Roseburia_inulinivorans+81.813s__Prevotella_nigrescens-9.098s__Enterobacter_cloacae-9.461s__Bacteroides_clarus-0.753s__Anaerostipes_unclassified-0.492s__Cl ostridioides_difficile-63.111s__Campylobacter_gracilis-0.119s__Prevotella_stercorea-0.537s__Mogibacterium_diversum-0.095s__Schaalia_odontolytica-0.46s__Alistipes_obesi, the optimal threshold (or cutoff value) for predicting P is 0.62. When P < 0.62, it indicates that the tumor patient is lowly sensitive to radiotherapy. When P ≥ 0.62, it indicates that the tumor patient is highly sensitive to radiotherapy.

[0114] In order to further determine whether the tumor patient is moderately sensitive or highly sensitive to radiotherapy, the stool samples with P ≥ 0.62 were compared with stool samples of colorectal cancer patients with moderate sensitivity to radiotherapy, and a logistic regression model constructed using the microbial markers screened in Example 1: s__Roseburia_inulinivorans, s__Prevotella_stercorea, s__Anaerostipes_unclassified, s__Clostridioides_difficile, s__Prevotella_nigrescens, s__Schaalia_odontolytica and s__Alistipes_obesi was used to predict the optimal threshold of P.

[0115] The logistic regression model for the stool samples to be tested with P ≥ 0.62 vs. the stool samples of colorectal cancer patients with moderate sensitivity to radiotherapy is logit(P) = log(P / (1-P)) = -0.043+0.581s__Roseburia_inulinivorans+0.016s__Prevotella_stercorea+0.997s__Anaerostipes_unclassified+0.009s__Clostridioides_difficile-0.946s__Prevotella_nigrescens+0.283s__Schaalia_odontolytica-0.765s__Alistipes_obesi. The optimal threshold (or cutoff value) for predicting P is 0.48. When P < 0.48, it indicates that the tumor patient is moderately sensitive to radiotherapy. When P ≥ 0.48, it indicates that the tumor patient is highly sensitive to radiotherapy. Colorectal cancer patients with high radiosensitivity are those with a NAR score of less than 8.

[0116] According to the logistic regression model fitting method, the ROC curve is drawn. Figure 5 The results showed that the AUC value of the stool samples to be tested vs. the stool samples of colorectal cancer patients who were not sensitive to radiotherapy reached 0.913, and the AUC value of the stool samples to be tested with P≥0.62 vs. the stool samples of colorectal cancer patients who were moderately sensitive to radiotherapy reached 0.552.

[0117] The ROC curve is drawn according to the bacterial flora involved in the logistic regression model equation. Figure 6 The results showed that the AUC of the present invention using microbial markers to diagnose stool samples to be tested versus stool samples of colorectal cancer patients who are not sensitive to radiotherapy reached 0.907, and the AUC value of the stool samples to be tested with P≥0.62 versus stool samples of colorectal cancer patients who are moderately sensitive to radiotherapy reached 0.571.

[0118] The relative levels of the microbiota in the group of colorectal cancer patients with radiotherapy insensitivity in Example 1 were brought into the prHS+prMSvsprLS logistic regression model to verify their grouping, and the P values ​​of each sample were 0.482658179, 0.499993271, 0.500693147, 0.492165009, 0.500693147, 0.421449718, 0.411294497, 0.476149864, 0.154763929, and 0.478 716206, 0.486857377, 0.414507253, 0.488936299, 0.008361682, 6.79526E-05, 0.495352268, 0.500693147, 0.500693147, 0.487279704, 0.443473258, 0.000487226, and 0.495477066 all met the P < 0.62. The verification results also proved that tumor patients have low sensitivity to radiotherapy.

[0119] It can be seen that the microbial flora marker of the present invention has high accuracy in detecting tumor radiotherapy sensitivity.

[0120] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A system for predicting tumor radiotherapy sensitivity using microbial biomarkers, characterized in that: It includes a nucleic acid sample separation unit for separating a fecal flora nucleic acid sample from a test sample; The data processing unit uses the microbial markers to establish a logistic regression model and obtain the critical value; A result determination unit, used for comparing the critical value obtained by the data processing unit with the set diagnostic value; The logistic regression model is a first logistic regression model and a second logistic regression model; The first logistic regression model is a logistic regression model constructed by comparing the stool sample to be tested with the stool sample of the colorectal cancer patient who is not sensitive to radiotherapy. When the critical value is less than 0.62, it indicates that the tumor patient is not sensitive to radiotherapy. When the critical value is greater than or equal to 0.62, it indicates that the tumor patient is moderately sensitive to radiotherapy or highly sensitive to radiotherapy. When the first logistic regression model was established using bacterial flora markers, the bacterial flora markers were s__Schaalia_odontolytica, s__Roseburia_inulinivorans, s__Mogibacterium_diversum, s__Campylobacter_gracilis, s__Alistipes_obesi, s__Prevotella_nigrescens, s__Clostridioides_difficile, s__Bacteroides_clarus, s__Enterobacter_cloacae, s__Anaerostipes_unclassified, and s__Prevotella_stercorea; The first logistic regression model is logit(P)=log(P / (1-P))=0.004+0.188s__Roseburia_inulinivorans+81.813s__Prevotella_nigrescens-9.098s__Enterobacter_cloacae-9.461s__Bacteroides_clarus-0.753s__Anaerostipes_unclassified-0.492s__Clostridioides_difficile-63.111s__Campylobacter_gracilis-0.119s__Prevotella_stercorea-0.537s__Mogibacterium_diversum-0.095s__Schaalia_odontolytica-0.46s__Alistipes_obesi; The second logistic regression model is a logistic regression model constructed by comparing stool samples with a critical value ≥ 0.62 to stool samples of colorectal cancer patients with moderate sensitivity to radiotherapy. When the critical value is < 0.48, it indicates that the tumor patient is moderately sensitive to radiotherapy, and when the critical value is ≥ 0.48, it indicates that the tumor patient is highly sensitive to radiotherapy. When the second logistic regression model was established using bacterial flora markers, the bacterial flora markers were s__Roseburia_inulinivorans, s__Prevotella_stercorea, s__Anaerostipes_unclassified, s__Clostridioides_difficile, s__Prevotella_nigrescens, s__Schaalia_odontolytica, and s__Alistipes_obesi; The second logistic regression model is logit(P)=log(P / (1-P))=-0.043+0.581s__Roseburia_inulinivorans+0.016s__Prevotella_stercorea+0.997s__Anaerostipes_unclassified+0.009s__Clostridioides_difficile-0.946s__Prevotella_nigrescens+0.283s__Schaalia_odontolytica-0.765s__Alistipes_obesi; The tumor is colorectal cancer.

Citation Information

Patent Citations

  • Marking composition for colorectal cancer detection and application thereof

    CN116183889A

  • Fecal microbiome marker combination for early multistage diagnosis of colorectal cancer and colorectal adenoma

    CN118086499A