Early screening model for esophageal cancer based on oral microorganisms and method for constructing the same.
A non-invasive esophageal cancer screening model using oral microbiome markers effectively distinguishes esophageal cancer from periodontitis, achieving high diagnostic accuracy through logistic regression analysis, addressing the limitations of current invasive methods.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2026-04-06
AI Technical Summary
Current esophageal cancer screening methods are invasive, have low patient compliance, and are hindered by the inability to distinguish between oral microbiome changes caused by periodontitis and esophageal cancer, leading to inaccurate biomarker diagnostics.
An early screening model for esophageal cancer based on oral microorganisms, utilizing Prevotella, Granulicatella, Rothia, Prevotellamassilia, Blautia, Abiotrophia, Peptostreptococcus, Actinomyces, Burkholderia, Akkermansia, Fudania, Kineothrix, Bacillus, Duncanella, and Eisenbergiella as markers, with logistic regression analysis to differentiate between healthy, periodontitis, and esophageal cancer patients.
The model achieves high specificity (0.839) and sensitivity (0.911) in the discovery cohort and maintains effectiveness (specificity 0.768, sensitivity 0.820) in the validation cohort, enabling non-invasive, efficient early detection of esophageal cancer.
Smart Images

Figure 2026058987000007 
Figure 2026058987000008 
Figure 2026058987000009
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of biomedicine and analysis, and specifically relates to an early screening model for esophageal cancer based on oral microorganisms and a method for constructing the same.
Background Art
[0002] Esophageal cancer is a malignant tumor of the upper gastrointestinal tract that progresses from atypical growth to precancerous lesions, carcinoma in situ, and metastatic cancer. It ranks 9th among the most common cancers worldwide and 6th among cancer-related deaths. According to epidemiological studies, in 2020, there were 604,100 new cases of esophageal cancer worldwide (6.3 / 100,000) and 544,076 deaths (5.6 / 100,000), and these are projected to increase to 957,000 (6.3 / 100,000) and 880,000 (5.6 / 100,000) by 2040, respectively. However, the cause of such a high incidence of esophageal cancer is still unclear, resulting in a lack of primary prevention based on etiology. Therefore, secondary prevention through "screening" and "early diagnosis and early treatment" is the main method of preventing esophageal cancer. Early screening is one of the most effective ways to detect cancer and save lives. Studies have shown that if detected early, more than 90% of cancer patients can be successfully treated. Unfortunately, the development of early screening strategies for esophageal cancer has been limited over the past few decades, and the above-mentioned early screening method, primarily using gastrointestinal endoscopy, remains the only option. Upper gastrointestinal endoscopy involves inserting a thin, flexible tube (endoscope) with a small camera at its tip from the patient's throat into the esophagus. This is an invasive examination method, and because patient compliance is low and widespread adoption is difficult, it fails to detect esophageal cancer early. Furthermore, because there are no specific symptoms in the early stages, many patients are diagnosed in the mid- or late stages, after which the disease progresses rapidly, limiting treatment options and making a cure unlikely, which contributes to the high mortality rate. Given the current state of clinical treatment for esophageal cancer, and the lack of targets for primary prevention, there is an urgent need for non-invasive methods that facilitate early screening, early diagnosis, and early treatment to suppress the progression of esophageal cancer.
[0003] The oral microbiome, as the second largest bacterial reservoir in the human body, is considered a potential biomarker for predicting head and neck cancer, lung cancer, and colorectal cancer. Previous studies have shown abnormalities in the oral microbiome in esophageal cancer patients, and poor oral hygiene has been linked to an increased risk of esophageal cancer. However, due to small sample sizes and the lack of a separate control group (periodontitis patients) in addition to healthy controls, the predictive value of the oral microbiome for esophageal cancer remains unclear. Since periodontitis is a risk factor for esophageal cancer, it may cause abnormalities in the oral microbiome. Bacteria that cause periodontitis (such as Porphyromonas gingivalis and Fusobacterium nucleatum) not only exist within cancerous lesions but also promote cancer progression through multiple mechanisms, posing a significant challenge to the accuracy of diagnostic / screening models. Determining whether the oral microbiome of cancer patients differs from that of periodontitis patients is crucial. The fact that the oral microbiota of the two groups is essentially the same means that existing biomarkers may misdiagnose periodontitis patients as cancer patients, and that the influence of periodontitis on esophageal cancer screening cannot be effectively eliminated. This is a major limitation of previous studies and the most important factor affecting sensitivity and accuracy. [Overview of the project]
[0004] To address the limitations of current early screening methods for esophageal cancer patients, the low compliance among subjects, and the difficulty in widespread adoption, as well as the problem that biomarker-based esophageal cancer screening cannot eliminate interference from periodontitis, the present invention provides an early screening model for esophageal cancer based on oral microorganisms and a method for constructing the same.
[0005] The present invention is realized by the following technical solution: an early screening model for esophageal cancer based on oral microorganisms is provided, the model using the following oral microbiomes as markers: Prevotella, Granulicatella, Rothia, Prevotellamassilia, Blautia, Abiotrophia, Peptostreptococcus, Actinomyces, Burkholderia, Akkermansia, Fudania, Kineothrix, Bacillus, Duncanella, and Eisenbergiella.
[0006] Selectively, the relative abundance information of the oral microbiome is used as the independent variable of the model, and whether or not the patient has esophageal cancer is used as the dependent variable, and logistic regression is performed. The model was modeled using a Regression model, and the equation of the model constructed by logistic regression is y = 1 / (1 + e^(2.89285631363524 + -0.130555223035498 * Prevotella + -0.234624415130977 * Granulicatella + -0.123421428328489 * Rothia + 0.0305284360696491 * Prevotellamassilia + 2.72774911990433 * Blautia + -0.671714432121879 * Abiotrophia + -0.643868902165489 * Pe The formula is ptostreptococcus+1.81604325065481*Actinomyces+9.28516953416389*Burkholderia+-3.52739890519049*Akkermansia+9.46455532514789*Fudania+-7.93770666127306*Kineothrix+-127.099900821647*Bacillus+-93.3296419351167*Duncaniella+-37.9213542816937*Eisenbergiella), and the names of the bacterial communities in the formula represent information about the relative abundance of the bacterial community.
[0007] The method for constructing the above-mentioned early screening model for esophageal cancer based on oral microorganisms is as follows: Step (1) involves collecting oral swab samples from healthy individuals, oral swab samples from patients with periodontitis, and oral swab samples from patients with esophageal cancer as analytical samples, Step (2) involves extracting DNA from the analysis sample to construct a library, performing library quality control using the Qubit3.0 Fluorometer and FEMTO Pulse system, and sequencing the qualified library using the Pacbio Sequel platform. Difference analysis of oral microbiota composition and structure for qualified libraries: Step (3) performed using the Bray-Curtis sample distance calculation method, Construction of an early screening model (initial model): The cohort is divided into a discovery cohort and a validation cohort using the "Rand" function in Excel. First, in the discovery cohort, species that distinguish esophageal cancer patients are screened using the random forest (randomforest V4.6.12) method. The parameters are set as randomForest obs ~ ., data = data, importance = T. A model is constructed using the logistic regression algorithm. The parameters of the R-3.6.3 function glm are set as glm(obs ~ ., family = binomial(), data). The classification effect of the initial model is determined using the area under the curve (AUC), specificities, and sensitivity in the receiver operating curve (ROC) (step 4). Model simplification: The impact of excluding a single species on the evaluation indices of the esophageal cancer early screening model (initial model) is investigated using a "sequential exclusion method." Finally, an attempt is made to exclude four species that have low Mean Decrease Gini coefficients and do not significantly impact the evaluation indices of the model after exclusion. Then, the early screening model is reconstructed as the final simplified model using a logistic algorithm, and this final simplified model is validated in a validation cohort. Similarly, the area under the curve (AUC), specificities, and sensitivities in the receiver operating curve (ROC) are used as evaluation indices of the model (5).
[0008] This model exhibits excellent esophageal cancer screening efficacy. Specific parameters characterizing the screening effect include a specificity of 0.839, sensitivity of 0.911, area under the receiver operating curve of 0.932, and a 95% confidence interval of 0.900–0.965 in the detection cohort. After evaluation using the receiver operating curve, the specificity of the model in the validation cohort is 0.768, sensitivity of 0.820, area under the curve of 0.856, and a 95% confidence interval of 0.806–0.906.
[0009] The present invention enables esophageal cancer screening simply by collecting an oral swab (saliva) sample from a patient and having a specialist complete the subsequent sequencing, analysis, and diagnostic processes. The present invention has the advantages of being non-invasive, convenient, and easy to implement, and subjects can perform the sample collection process at home, enabling early detection and treatment of esophageal cancer.
[0010] This invention aims to develop an efficient, specific, and highly sensitive early screening method for esophageal cancer by clarifying the differences in oral microbiota between esophageal cancer patients, healthy individuals, and periodontitis patients, and by identifying specific microbiota characteristics in esophageal cancer patients. From Shanxi Province in China, which has a high incidence of esophageal cancer, 109 age-matched healthy volunteers, 138 periodontitis patients, and 224 esophageal cancer patients were recruited. Oral swab samples were collected from the subjects, and 16S-rDNA sequencing was performed to analyze the characteristics of the bacterial flora in the samples. Subsequently, an early screening model based on oral microbiota was constructed using the relative abundance of indicator microbiota, and the diagnostic effect of the indicator microbiota was evaluated using receiver operation characteristics (ROC) analysis. This ensured that the early screening model could effectively screen esophageal cancer patients from healthy individuals and periodontitis patients. This invention provides a non-invasive screening method for the screening and early diagnosis of esophageal cancer, thereby providing patients with the opportunity for early treatment. This invention aims to identify non-invasive biomarkers to enable the early detection of asymptomatic esophageal cancer patients, thereby improving prognosis and saving lives. [Brief explanation of the drawing]
[0011] [Figure 1] This is a quality control flowchart for sequencing data of oral swab bacterial flora. [Figure 2] This is a principal coordinate diagram (PCoA diagram) showing the differences in diversity of the composition and structure of oral microbiota in esophageal cancer patients, healthy individuals, and periodontitis patients. [Figure 3] This is a schematic diagram illustrating how the "Rand" function is used to randomly divide a research cohort into a discovery cohort and a validation cohort. [Figure 4] This figure shows 20 genera-level species that were screened using the random forest method in the discovery cohort and can be applied to constructing an early screening model for esophageal cancer. [Figure 5] This figure shows 19 genera-level species screened using multiple evaluation indices of the receiver operating curve (area under the curve AUC, sensitivity, specificity) as indicators for constructing an early screening model for esophageal cancer (initial model). [Figure 6] This is the receiver-operated characteristic curve of the esophageal cancer early screening model (initial model) in the discovery cohort, and the AUC shown is the area under the receiver-operated characteristic curve. [Figure 7] This schematic diagram shows how the "sequential exclusion method" is used to investigate the impact of excluding a single species on the evaluation indicators of an esophageal cancer early screening model (initial model). Species marked in red have a significant impact on the evaluation indicators of the model after exclusion, while species marked in blue have an impact on the evaluation indicators of the model after exclusion. Ultimately, it is decided to exclude Solobacterium, Catonella, Filifactor, and Bergeyella, and to model the model using the remaining 15 genera. [Figure 8] This is the receiver-operated characteristic curve used when constructing an early esophageal cancer screening model (final simplified model) using 15 genera-level species in the discovery cohort. The AUC shown is the area under the receiver-operated characteristic curve. [Figure 9] This is the receiver-operated characteristics curve when the esophageal cancer early screening model (final simplified model) is applied to the validation cohort. The AUC shown is the area under the receiver-operated characteristics curve. [Modes for carrying out the invention]
[0012] To further clarify the object, technical solution, and advantages of the embodiments of the present invention, the technical solution in the embodiments of the present invention will be described clearly and completely below. It will be obvious that the embodiments described are only some embodiments of the present invention, not all embodiments. Any other embodiments that can be obtained by a person skilled in the art without creative effort based on the embodiments of the present invention are all within the scope of protection of the present invention.
[0013] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those generally understood by those skilled in the art, and all disclosures and materials cited herein are incorporated by reference.
[0014] Those skilled in the art should understand that technical equivalents of the specific embodiments described, which can be obtained through ordinary experiments, are included in this application.
[0015] Unless otherwise specified, the experimental methods in the following examples are all ordinary methods. Unless otherwise specified, the equipment and devices used in the following examples are all ordinary experimental equipment and devices, and the experimental materials used in the following examples are all purchased from ordinary biochemical reagent stores unless otherwise specified.
[0016] The reagents and consumables used in this example are as follows.
[0017] [[ID=
[0022] Selection criteria for periodontitis patients: Patients (aged 50 and above) diagnosed with periodontitis at admission; Exclusion criteria: Patients with oral mucosal lesions, patients with bacterial or viral infections in the tonsils, salivary glands or throat within 1 month before sample collection, patients who received periodontal treatment within 3 months before sample collection, patients who received surgery, radiotherapy, chemotherapy within the past 6 months, patients with oral squamous cell carcinoma, head and neck squamous cell carcinoma, colorectal cancer, vulnerable groups such as pregnant women.
[0023] Selection criteria for esophageal cancer patients: Patients diagnosed with esophageal cancer at admission; Exclusion criteria: Patients with obvious inflammation in the oral cavity, patients with bacterial or viral infections in the tonsils, salivary glands or throat within 1 month before sample collection, patients who received periodontal treatment within 3 months before sample collection, minors who received general anesthesia surgery within the past 6 months, vulnerable groups such as pregnant women.
[0024] All subject information is shown in Table 3.
[0025] TIFF2026058987000003.tif175170
[0026] 2. Method for collecting oral swabs: Before sample collection, about 50 ml of purified water was drunk, and the mouth was rinsed well for about 10 s and then spit out. The swab was extended into the mouth, and the tip of the swab was sufficiently contacted with the inner side of the cheek and the upper and lower gingival mucosa, rubbed up and down with a toothbrushing-like force, and at the same time the swab was rotated, the tip of the swab was sufficiently contacted with the oral mucosa, and this operation was repeated for 1 min. The swab was returned to a sterilized collection tube and stored at low temperature.
[0027] 3. Extraction of DNA and construction of library A. DNA extraction (extracted using Hipure Universal DNA kits, specific experimental methods refer to the manual, and the following is a simple procedure): 1) Transfer the swab sample to a 2 ml centrifuge tube, immediately add 500 ul of buffer ATL (included in the kit) and 20 ul of protease K (included in the kit), and incubate at 56°C for 1 hour with shaking. 2) Add 500 ul of buffer AL (included in the kit) and 500 ul of anhydrous ethanol, vortex for 10 seconds, 3) Place 2 ml of Hipure DNA Mini Column I (included in the kit) into a collection tube, transfer the mixture (less than 750 ul) to the column, and centrifuge at 10000 g for 1 minute. 4) Discard the waste liquid at the bottom, return the column to the sampling tube, transfer the remaining mixture to the column, centrifuge at 10000g for 1 minute, discard the filtrate and sampling tube, 5) Discard the waste liquid at the bottom, return the column to a new sampling tube, add 500 ul of buffer GW1 (diluted with ethanol) to the column, and centrifuge at 10000 g for 1 minute. 6) Discard the waste liquid at the bottom, return the column to the sampling tube, add 650 ul of buffer GW2 (diluted with ethanol) to column 1, and centrifuge at 10000 g for 1 min. 9) Discard the waste liquid at the bottom, return the column to the sampling tube, centrifuge at 10000g for 3 mins, and centrifuge dry the column. 10) Place the column in a 1.5 ml centrifuge tube, add 50-100 μl of Buffer AE (included in the kit) preheated to 70°C to the center of the column membrane, leave at room temperature for 3 minutes, and centrifuge at 10000 g for 1 minute. 11) The DNA-binding column was discarded, and the DNA was either used for subsequent experiments or stored frozen at -20°C.
[0028] B. PCR amplification: Template: Extracted DNA, 16S V1-V9 primers (with barcode): 27F-AGRGTTYGATYMTGGCTCAG, 1492R-RGYTACCTTGTTACGACTT.
[0029] Amplification system: 50 μl of TransGen High-Fidelity PCR SupermMix mixture containing 0.2 μM upstream and downstream primers and 5 ng of DNA template. The amplification program is shown in Table 4.
[0030] TIFF2026058987000004.tif55170
[0031] C. Purification and recovery using AMPure® PB Beads: If the sample concentration or purity does not meet the standards, the amplification product can be purified and recovered using AMPure® PB Beads, and 1-2 ug of amplification product was collected in a new 1.5 ml centrifuge tube.
[0032] Remove the AMP PB Bead from 2-8°C 30 min prior to the start of the experiment, allow it to stabilize at room temperature, mix the AMP PB Bead thoroughly, aspirate an equal volume, add it to the sample of the amplified product after screening, mix slowly 15 times with a wide-bore pipette at room temperature without slapping the tube, incubate at room temperature for 5 minutes, immediately centrifuge in a handheld centrifuge after incubation, place the sample on a magnetic rack, and carefully remove the supernatant after the solution has become clear. The sample was kept in a magnetic rack, the magnetic beads were rinsed with 200 µl of freshly prepared 80% ethanol, incubated at room temperature for 30 s, the supernatant was removed, the previous step was repeated, the tube was removed, and immediate centrifugation was performed for 3 s in a handheld centrifuge, the sample was left in the magnetic rack for about 10 s, the residual ethanol was carefully removed, the sample was kept in the magnetic rack, the lid was opened and the magnetic beads were allowed to dry at room temperature for about 5-10 min, the sample was removed from the magnetic rack, 45.5 µl of Elution Buffer was added, gently mixed with a wide-bore tip, left at room temperature for 15 min, then placed in the magnetic rack, and waited until the solution became clear, and 25 µl of the supernatant was carefully aspirated and transferred to a new nuclease-free EP tube.
[0033] D. Damage repair: The reaction system is as follows: Amplification product 25.0 µl; DNA Damage Repair Mix v2, 2.0 µl; Total Volume 27.0 µl; incubated at 37°C for 30 min, cooled to 4°C and stored, then proceeded to the next step.
[0034] E, end repair and A addition: The reaction system was as follows: amplification product 27.0 ul, End Prep Mix 3.0 ul; Total Volume 30.0 ul; incubated at 20°C for 10 min, incubated at 65°C for 30 min, stored at 4°C, and proceeded to the next step.
[0035] F. Adapter connection: The reaction system is shown in Table 5, incubated at 20°C for 60 minutes, stored at 4°C, and then proceeded to the next step.
[0036] TIFF2026058987000005.tif64170
[0037] G. Purification using AMPure® PB Bead: Add 3 μl of Elution Buffer to the reaction system from the previous step, mix thoroughly, aspirate 45 μl and add to the ligation product, slowly pipette 15 times with a wide-bore tip at room temperature, incubate at room temperature for 5 minutes without flicking the tube, immediately centrifuge with a handheld centrifuge after incubation, place the sample on a magnetic rack, carefully remove the supernatant after the solution becomes clear, keep the sample on the magnetic rack, rinse the magnetic beads with 200 μl of freshly prepared 80% ethanol, incubate at room temperature for 30 seconds, carefully remove the supernatant, repeat the previous step, remove the tube, and immediately centrifuge with a handheld centrifuge for 3 seconds.
[0038] Library quality control was performed using the Qubit3.0 Fluorometer and FEMTO Pulse system, and qualified libraries were sequenced using the Pacbio Sequel platform.
[0039] 4. Analysis of differences in the composition and structure of oral microbiota in esophageal cancer patients, healthy individuals, and periodontitis patients: β-diversity is an indicator of the degree of difference in the diversity of species composition and structure between different ecosystems. In this embodiment, the Bray-Curtis sample distance calculation method ( Beta diversity statistics were performed using TIFF2026058987000006.tif12170, and finally, differences in microecological composition and structure between different cohorts were shown using PCOA diagrams (see Figure 2). The calculations and statistics were mainly performed using the Vegan package in the R language.
[0040] 5. Construction of an esophageal cancer early screening model (initial model): The cohort was divided into a discovery cohort and a validation cohort using the "Rand" function in Excel. First, in the discovery cohort, the random forest method (randomforest(V4.6.12), parameters: randomForest(obs ~ ., data = data, importance = T)) was used to screen for species that could distinguish esophageal cancer patients. Then, using the logistic regression algorithm (R-3.6.3 function glm parameters: glm(obs ~ ., family = binomial(), data)), an early screening model (initial model) was constructed based on the discovery cohort, and a corresponding model algorithm was generated. The model classification effect was determined by the receiver operating curve (ROC).
[0041] 6. Simplification of the initial model: The impact of excluding a single species on the evaluation indices of the esophageal cancer early screening model (initial model) was investigated using the "sequential exclusion method." Finally, after attempting to exclude four species that had low Mean Decrease Gini coefficients and did not significantly impact the evaluation indices of the model after exclusion, the early screening model was reconstructed as the final simplified model using a logistic algorithm. This final simplified model was then validated in a validation cohort. Similarly, the area under the curve (AUC), specificities, and sensitivity of the receiver operating curve (ROC) were used as evaluation indices for the model.
[0042] 2. Experimental results: 1. There are significant differences in the structure of the oral microbiota between esophageal cancer patients and healthy individuals and patients with periodontitis. Through analysis of the microbiological β-diversity of the subject groups, it was revealed that the oral microbiological structure of esophageal cancer patients is significantly different from that of healthy individuals and patients with periodontitis. Specifically, in PCOA analysis based on the Bray-Curtis distance, the sample distribution of esophageal cancer patients was significantly different from that of healthy individuals and patients with periodontitis, and statistical analysis showed P=0.001 (see Figure 2). This indicates that the oral microbiota of esophageal cancer patients differs in overall composition and structure from that of non-esophageal cancer patients (including healthy individuals and patients with periodontitis).
[0043] 2. Randomly divide into discovery and validation cohorts: Using the "RAND" function in an Excel worksheet, the cohorts of healthy individuals, periodontitis patients, and esophageal cancer patients are randomly divided into discovery and validation cohorts (see Figure 3). The discovery cohort is used to construct an early screening model, and the validation cohort is used to re-verify the diagnostic effectiveness of the constructed screening model.
[0044] 3. Screening for marker species (genera level) applicable to early screening and diagnosis of esophageal cancer using a random forest model: First, in the discovery cohort, healthy individuals and periodontitis patients are mixed as a non-esophageal cancer patient cohort. Then, using a random forest model, the optimal genera level species that can distinguish between esophageal cancer patients and non-esophageal cancer patients is screened, and the results are sorted in descending order of the Mean Decrease Gini coefficient obtained by random forest analysis. The top 20 genera of species with high Gini coefficients (Prevotella, Gardnerella, Granulicatella, Rothia, Prevotellamassilia, Blautia, Abiotrophia, Peptostreptococcus, Actinomyces, Filifactor, Solobacterium, Bergeyella, Burkholderia, Akkermansia, Catonella, Fudania, Kineothrix, Bacillus, Duncanella, and Eisenbergiella) were selected as marker species for constructing an early screening model (Figure 4).
[0045] 4. Using receiver operating curves, evaluate the diagnostic / screening / classification effects of the 20 genera-level marker species in esophageal cancer patients: Area under the curve (AUC) is an index that evaluates the effectiveness of a particular species in the diagnosis / screening / classification of esophageal cancer, generally ranging from 0.5 to 1.0, with values closer to 1.0 indicating better effectiveness of that species in the diagnosis of esophageal cancer. Furthermore, specificities and sensitivities can reflect the reliability and sensitivity of a particular species applied to the diagnosis / screening / classification of esophageal cancer, respectively; higher specificities result in a lower false positive rate, and higher sensitivity results in a lower false negative rate. The results in Figure 5 show that the AUC for Gardnerella alone is 0.305, while the AUCs for all other species are above 0.5, indicating that Gardnerella does not have a superior ability to differentiate esophageal cancer patients. Therefore, in the subsequent logistic modeling analysis, Gardnerella was excluded, and an early screening model for esophageal cancer (initial model) was constructed using the remaining 19 genera (Prevotella, Granulicatella, Rothia, Prevotellamassilia, Blautia, Abiotrophia, Peptostreptococcus, Actinomyces, Filifactor, Solobacterium, Bergeyella, Burkholderia, Akkermansia, Catonella, Fudania, Kineothrix, Bacillus, Duncanella, Eisenbergiella).
[0046] 5. Construction of an early screening model for esophageal cancer (initial model): First, the 19 species screened using a random forest model in the discovery cohort are modeled using a logistic regression model. The independent variable of the model is the relative abundance information of each bacterial community, and the dependent variable is the disease group (whether or not the individual has esophageal cancer). After logistic regression analysis and modeling, the predictive effect of the model is evaluated using a receiver operating curve. As a result, the early screening model (initial model) constructed based on these 19 species showed high specificity (0.927) and sensitivity (0.839) in the discovery cohort, with an area under the curve reaching 0.934 (Figure 6). This indicates that the early screening model constructed based on 19 species at the genera level has excellent ability for the diagnosis / screening / classification of esophageal cancer in the discovery cohort.
[0047] 6. Simplification of the esophageal cancer early screening model: The impact of excluding a single species on the evaluation indicators of the esophageal cancer early screening model (initial model) was investigated using the "sequential exclusion method." Species that had a significant impact on the evaluation indicators of the model after exclusion were marked in red, and species that had an impact on the evaluation indicators of the model after exclusion were marked in blue (Figure 7). Simultaneously, based on the Mean Decrease Gini coefficient, it was decided to ultimately exclude Solobacterium, Catonella, Filifactor, and Bergeyella, and construct the esophageal cancer early screening model (final simplified version) using the remaining 15 genera.
[0048] 7. Construction of an esophageal cancer early screening model (final simplified model): In the discovery cohort, the 15 species screened above were modeled using logistic regression analysis. The independent variable of the model was the relative abundance information of each bacterial flora, and the dependent variable was the disease group (whether or not the patient had esophageal cancer). After logistic regression analysis and modeling, the predictive effect of the model was evaluated using receiver operating curves. As a result, the early screening model (final simplified model) constructed based on the 15 species had high specificity (0.839) and sensitivity (0.911) in the discovery cohort, with an area under the curve of 0.932 (Figure 8). Compared to the initial model before simplification, the area under the curve was almost the same, sensitivity improved, and specificity decreased. Overall, the esophageal cancer early screening model (final simplified model) still maintains excellent esophageal cancer diagnostic / screening / classification capabilities.
[0049] The equation for the model constructed using logistic regression is y = 1 / (1 + e^(2.89285631363524 + -0.130555223035498 * Prevotella + -0.234624415130977 * Granulicatella + -0.123421428328489 * Rothia + 0.0305284360696491 * Prevotellamassilia + 2.72774911990433 * Blautia + -0.671714432121879 * Abiotrophia + -0.643868902165489 * Peptostreptoco ccus+1.81604325065481*Actinomyces+9.28516953416389*Burkholderia+-3.52739890519049*Akkermansia+9.46455532514789*Fudania+-7.93770666127306*Kineothrix+-127.099900821647*Bacillus+-93.3296419351167*Duncaniella+-37.9213542816937*Eisenbergiella)), and the names of the bacterial communities in the formula represent information about the relative abundance of the bacterial community.
[0050] 8. Evaluation of the classification effect of the esophageal cancer early screening model (final simplified model) in the validation cohort: The early screening model constructed above was further applied to the validation cohort and evaluated by receiver operating curve (Figure 9). The model showed a specificity of 0.768, a sensitivity of 0.820, and an area under the curve of 0.856, indicating that the model still has excellent effectiveness for the diagnostic screening of esophageal cancer in an independent validation cohort.
[0051] It should be noted that the above embodiments are for illustrating the technical solutions of the present invention and do not limit the present invention. Although the present invention has been described in detail based on the embodiments described above, those skilled in the art should understand that they can modify the technical solutions described in the embodiments described above, or replace some or all of the technical features with equivalent ones, and that such modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of each embodiment of the present invention.
Claims
1. An early screening model for esophageal cancer based on oral microorganisms, characterized in that the model uses the oral microbiome of Prevotella, Granulicatella, Rothia, Prevotellamassilia, Blautia, Abiotropia, Peptostreptococcus, Actinomyces, Burkholderia, Akkermansia, Fudania, Kineothrix, Bacillus, Duncaniella, and Eisenbergierla as markers.
2. The relative abundance information of the oral microbiome was used as the independent variable of the model, and whether or not the patient has esophageal cancer was used as the dependent variable. A logistic regression model was used for modeling, and the equation of the model constructed by logistic regression is y = 1 / (1 + e^(2.89285631363524 + -0.130555223035498 * Prevotella + -0.234624) 415130977*Granulicatella+-0.123421428328489*Rothia+0.0305284360696491*Prevotellama ssilia+2.72774911990433*Blautia+-0.671714432121879*Abiotrophia+-0.643868902165489* Peptostreptococcus+1.81604325065481*Actinomyces+9.28516953416389*Burkholderia+-3. 52739890519049*Akkermansia+9.46455532514789*Fudania+-7.93770666127306*Kineothrix+- The early screening model for esophageal cancer based on oral microorganisms according to claim 1, characterized in that the formula is 127.099900821647*Bacillus+-93.3296419351167*Duncanialla+-37.9213542816937*Eisenbergialla), and the name of the bacterial community in the formula represents information on the relative abundance of the bacterial community.
3. A method for constructing an early screening model for esophageal cancer based on oral microorganisms according to claim 1 or 2, Step (1) involves collecting oral swab samples from healthy individuals, oral swab samples from patients with periodontitis, and oral swab samples from patients with esophageal cancer as analytical samples, Step (2) involves extracting DNA from the analysis sample to construct a library, performing library quality control using Qubit 3.0 Fluorometer and FEMTO Pulse system, and sequencing the qualified library using Pacbio Sequel platform. Difference analysis of oral microbiota composition and structure for qualified libraries: Step (3) reflecting differences between samples and between communities using the Bray-Curtis sample distance calculation method, Construction of an early screening model (initial model): Using the "Rand" function in Excel, the cohort was divided into a discovery cohort and a validation cohort. First, in the discovery cohort, the species that distinguish esophageal cancer patients were screened using the random forest method (randomforest V4.6.12), and the parameters of the random forest were set to randomForest obs ~ . Step (4) involves setting data = data, importance = T, constructing a model using the logistic regression algorithm, setting the parameters of the R-3.6.3 function glm as glm(obs ~ ., family = binomial(), data), and determining the classification effect of the initial model using the area under the curve (AUC), specificities, and sensitivities in the receiver operating curve (ROC), Simplification of the model: A construction method characterized by including the step (5) of examining the impact of excluding a single species on the evaluation indicators of the esophageal cancer early screening model (initial model) using a "sequential exclusion method," and finally attempting to exclude four species that have a low Mean Decrease Gini coefficient and do not have a significant impact on the evaluation indicators of the model after exclusion, then reconstructing the early screening model as the final simplified model using a logistic algorithm, and validating the final simplified model in a validation cohort, and similarly using the area under the curve (AUC), specificities, and sensitivity in the receiver operating curve (ROC) as evaluation indicators of the model.
4. The construction method according to claim 3, characterized in that, regarding the effectiveness of the model for early screening of esophageal cancer, the specificity in the detection cohort is 0.839, the sensitivity is 0.911, the area under the receiver operating curve is 0.932, and the 95% confidence interval is 0.900 to 0.965, and after evaluation using the receiver operating curve, the specificity of the model in the validation cohort is 0.768, the sensitivity is 0.820, the area under the curve is 0.856, and the 95% confidence interval is 0.806 to 0.906.