Internal reference sequence, primer probe composition, kit and method for quantitatively detecting caries microbial marker of young children and application of internal reference sequence, primer probe composition, kit and method
By designing a 16S internal reference sequence and primer-probe combination, and combining it with ddPCR technology, the problems of low sensitivity and poor specificity in the detection of dental caries microbial markers in young children have been solved, achieving efficient and economical dental caries risk assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PEKING UNIV SCHOOL OF STOMATOLOGY
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies suffer from low sensitivity, poor specificity, poor repeatability, weak anti-interference ability, slow detection speed, and high cost when identifying microbial biomarkers of dental caries in young children, which limits the early diagnosis and effective intervention of ECC.
A 16S internal reference sequence and primer-probe combination were designed for the quantitative detection of caries microbial biomarkers in young children. Biomarkers such as Cardiobacterium hominis, Rothia aeria, and Streptococcus mutans were screened through a combination of machine screening and manual optimization. Simultaneous and absolute quantitative detection was performed using ddPCR technology, and the primer-probe sequence was optimized to improve specificity and stability.
It achieves high sensitivity (detection limit 0.01 copies/μL), high specificity (NTC amplification <10 copies/μL), high reproducibility (relative standard deviation ≤5%) and low cost multiplex detection, significantly improving the accuracy and efficiency of caries risk assessment.
Smart Images

Figure CN121896338A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of oral preventive medicine and oral public health technology, and in particular to internal reference sequences, primer-probe compositions, kits, methods and applications for the quantitative detection of caries microbial markers in young children. Background Technology
[0002] Early Childhood Caries (ECC), a chronic oral infectious disease, poses a serious challenge to the oral health of children worldwide due to its early onset, rapid progression, and high prevalence.
[0003] Current research and technologies lack a standardized system of microbial biomarkers for identifying ECC (extracorporeal caries). Three common methods exist: independent studies using high-throughput sequencing of 16S rRNA genes in small samples; meta-analysis modeling based on data from multiple studies; and qPCR (quantitative polymerase chain reaction) detection modeling based on several biomarkers suggested by previous studies. While independent studies using high-throughput sequencing of 16S rRNA genes in small samples can suggest biomarkers for caries risk, the small sample size results in weak statistical power and poor reproducibility, affecting the accurate identification and validation of microbial biomarkers. Meta-analysis modeling based on data from multiple studies only yields biomarker models, lacking validation and technical practice in new populations. Furthermore, the high-throughput sequencing-based analysis workflow cannot meet the needs of rapid clinical diagnosis, and its relatively high cost limits its efficiency in practical scenarios. Combinations of biomarkers suggested by previous studies are randomized, and qPCR lacks sensitivity when detecting low-abundance biomarkers, requiring standard curve quantification, which affects the ease of biomarker identification.
[0004] In summary, existing technologies face numerous challenges in researching and identifying ECC-related microbial biomarkers, including low sensitivity, poor specificity, poor reproducibility, weak anti-interference ability, slow detection speed, and high cost. These shortcomings collectively limit the early diagnosis and effective intervention of ECC. Summary of the Invention
[0005] Based on the above analysis, the present invention aims to provide an internal reference sequence, primer-probe composition, kit, method and application for quantitative detection of caries microbial biomarkers in young children, in order to solve at least one of the problems existing in the detection of caries microbial biomarkers in young children, such as low sensitivity, poor specificity, poor repeatability, poor anti-interference ability, slow detection speed and high cost.
[0006] The objective of this invention is achieved through the following technical solution: This invention provides a 16S internal reference sequence for quantitative detection of caries microbial biomarkers in young children, the internal reference sequence comprising an upstream primer sequence, a downstream primer sequence, and a probe sequence; The upstream primer sequence is SEQ ID NO.1; The downstream primer sequences are SEQ ID NO.2, SEQ ID NO.3, and SEQ ID NO.4; The probe sequences are SEQ ID NO.5, SEQ ID NO.6 and SEQ ID NO.7.
[0007] This invention provides a primer and probe composition for quantitative detection of microbial biomarkers of dental caries in young children. The primer and probe composition includes the 16S internal reference sequence, as well as upstream primers, downstream primers, and probes for detecting three bacteria: Cardiobacterium hominis, Rothia aeria, and Streptococcus mutans. The upstream and downstream primer sequences for detecting Human poxvirus are shown in SEQ ID NO.8 and SEQ ID NO.9, respectively, and the probe sequence is shown in SEQ ID NO.10. The upstream and downstream primer sequences for detecting *Rhodesia spp.* are shown in SEQ ID NO.11 and SEQ ID NO.12, respectively, and the probe sequence is shown in SEQ ID NO.13. The upstream and downstream primer sequences for detecting Streptococcus mutans are shown in SEQ ID NO.14 and SEQ ID NO.15, respectively, and the probe sequence is shown in SEQ ID NO.16.
[0008] Furthermore, in the primer-probe composition, the molar ratio between the upstream primers for detecting *Humanoidobacterium tumefaciens*, *Roseobacterium spaceis*, and *Streptococcus mutans* is 1:1:1, and the molar ratio between the upstream primer of the internal control and the upstream primer of each of the three bacteria is 2:1.
[0009] Furthermore, in the primer-probe composition, the molar ratio between the downstream primers for detecting *Humanoidobacterium tumefaciens*, *Roseobacterium spaceis*, and *Streptococcus mutans* is 1:1:1, and the molar ratio of each downstream primer of the internal control to the downstream primer of each of the three bacteria is 1:1.
[0010] Furthermore, in the primer-probe composition, the molar ratio between the probes for detecting *Humanoidobacterium tumefaciens*, *Roseobacter stenosis*, and *Streptococcus mutans* is 1:1:1, and the molar ratio of each probe of the internal control to the probe of each bacterium is 1:1.
[0011] Furthermore, in the primer-probe composition, the molar number of the upstream or downstream primer for each of the three bacteria is twice that of the probe for each bacteria; and / or, The molar ratio of the upstream primer, downstream primer, and probe of the internal control is 4:6:3.
[0012] This invention provides a kit for quantitative detection of caries microbial biomarkers in young children. The kit includes a ddPCR system comprising a primer-probe mixture, a ddPCR premix, a DNA template, and sterile water. The primer-probe mixture includes the internal reference sequence or the primer-probe composition.
[0013] This invention provides a method for quantitatively detecting caries microbial markers in young children, comprising: (1) Extract genomic DNA from dental plaque samples; (2) Using the genomic DNA as a template, droplets are generated in a droplet generator through a ddPCR system including the primer and probe composition or the kit described above, and PCR amplification is performed to obtain the amplification product; (3) The amplification product is subjected to fluorescence detection based on the droplet reader, and the copy number of the three bacteria, Cardiobacterium hominis, Rothia aeria, and Streptococcus mutans, and the copy number of the internal reference are determined respectively in the amplification product. (4) Divide the copy number of the three bacteria by the copy number of the internal reference to calculate the relative DNA content of each of the three bacteria in the dental plaque sample.
[0014] Furthermore, in step (4), the relative DNA content of the three bacteria is used in the ROC model to determine the risk of dental caries; among them, the cut-off value of the ROC model constructed by Streptococcus mutans, Human cardiomyxobolus, and Rossella squarrosa is 0.453; if only one of the markers is used for detection, the cut-off value of the Streptococcus mutans ROC model is 0.476, the cut-off value of the Human cardiomyxobolus ROC model is 0.387, and the cut-off value of the Rossella squarrosa ROC model is 0.378; and / or, The detection limits for the three bacteria are as low as 0.01 copies / μL; and / or, The standard deviation of the relative DNA content of each bacterium is ≤5%.
[0015] This invention provides the application of the internal reference sequence or the primer-probe composition described herein in the preparation of detection products for rapid quantitative detection of caries microbial biomarkers in young children.
[0016] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects: This invention designs the 16S internal reference sequence by combining machine screening with manual optimization, systematically solving at least one of the problems in existing caries microorganism detection: low sensitivity, poor specificity, poor repeatability, weak anti-interference ability, slow detection speed, and high cost.
[0017] (1) Machine screening lays the foundation for high sensitivity and broad spectrum. By using bioinformatics tools to perform initial screening on the conserved regions of the 16S rRNA gene or 16S rDNA V6 region, its efficient amplification potential for the vast majority of oral bacteria (coverage of 94.9%) is ensured, providing a guarantee for the detection of low-abundance microorganisms.
[0018] (2) Artificial optimization achieved breakthroughs in specificity and stability. Addressing the non-specific amplification of initial screening sequences in the human genome and complex samples, low-degeneracy and parallel allelic sequences were introduced to precisely match the polymorphisms of different strains, significantly reducing non-specific background signals to <10 copies / μL. Simultaneously, fine-tuning of Tm values and secondary structures ensured stable amplification efficiency across different batches (relative standard deviation ≤5%), significantly improving the specificity and reproducibility of the detection.
[0019] (3) This internal control sequence enables a rapid and economical multiplex detection scheme. Its design allows it to coexist stably with multiple caries biomarkers in the same ddPCR reaction. Combined with an optimized annealing procedure, it enables simultaneous, rapid, and absolute quantification of total bacterial count and specific pathogens. This not only avoids the cumbersome process of qPCR relying on standard curves and improves detection speed, but also effectively reduces overall costs by reducing repeated testing and reagent consumption.
[0020] The primer-probe composition provided by this invention screens core microbial biomarkers (Cardiobacterium hominis, Rothia aeria, Streptococcus mutans) from large-scale multi-cohort data by combining machine learning and LEfSe analysis, and further optimizes its primer-probe sequences through targeted artificial optimization. This systematically solves at least one of the problems in existing caries microbial detection, such as low sensitivity, poor specificity, poor repeatability, weak anti-interference ability, slow detection speed, and high cost. (4) The combination of biomarkers significantly improves the accuracy and reliability of diagnosis. Three biomarker combinations were selected from a large dataset of 559 samples across 9 studies using machine learning, increasing the sample size by approximately 6 times compared to traditional studies (which typically have an average sample size of 30-50). Following the PRISMA guidelines for systematic reviews, a unified data preprocessing workflow (QIIME 2 + DADA2 denoising) was employed to eliminate batch effects and significantly improve the comparability of data across studies. A random forest model with 10-fold cross-validation (n_estimators=50, max_depth=5) was used, ranked by feature importance scores (Ch=0.354, Ra=0.250). Seven potential biomarkers (such as Cardiobacterium hominis and Rothia aeria) were identified, ultimately validated as a combination of three core biomarkers (Ch, Ra, Sm), demonstrating significantly improved predictive performance. The combined biomarker achieved an AUC of 0.775 (the highest AUC for a single biomarker was 0.726). This multi-marker-based joint judgment model effectively overcomes the shortcomings of single markers, such as high randomness and easy missed diagnosis, and provides a more accurate and reliable basis for assessing caries risk.
[0021] (5) Manual optimization ensures high specificity and compatibility with multiplex detection. Based on machine screening, primers and probes for each bacterium were optimized manually. By precisely adjusting primer length, 3' end bases, and probe coverage sites, cross-reactivity with closely related bacteria and the risk of primer dimerization were effectively eliminated. For example, a high GC tail was added to *S. mutans* to increase the Tm value, and the 3' end of *R. aeria* was replaced to eliminate complementarity with the internal control. These optimizations ensured that the three biomarkers and the internal control could be amplified efficiently and specifically simultaneously in the same ddPCR system, achieving "one reaction, simultaneous quantification," and solving the problem of low efficiency caused by multiple detections required in traditional methods.
[0022] (6) Achieve high-sensitivity quantification and strong anti-interference capability The optimized primer-probe combination achieves a detection limit as low as 0.01 copies / μL, enabling precise capture of low-abundance pathogens in oral plaque. In actual sample validation, even with the inclusion of high concentrations of human genome, the quantification accuracy of the target bacteria was not significantly affected (deviation ≤5%), demonstrating excellent anti-interference capabilities and effectively addressing the pain points of insufficient sensitivity and susceptibility to inhibition in complex sample matrices for qPCR.
[0023] (7) In practical applications, this invention utilizes a multi-channel ddPCR system (such as FAM, ROX, CY5 fluorescence channels, etc.) combined with a 16S universal internal control to achieve simultaneous and absolute quantitative detection of three caries-related microbial biomarkers. This system has high detection sensitivity, with a detection limit of 0.01 copies / μL, and is particularly suitable for the accurate detection of low-abundance microorganisms. In a clinical validation cohort (e.g., n=209), the standard deviation of the relative abundance of biomarkers was significantly reduced (standard deviation ≤5%), the detection process was shortened by about 50% compared to traditional 16S rRNA gene sequencing, and reagent consumption was reduced, resulting in good economic benefits.
[0024] In terms of detection performance, the combined biomarker model used in this invention shows significant advantages over single biomarkers, with a 6.75% increase in the area under the ROC curve (AUC) (0.775 vs. 0.726), a specificity of 81.1%, strong interpretability, and high clinical applicability. The classification model built based on these validation results can be further embedded in portable diagnostic devices, laying a technological foundation for rapid and automated risk screening of dental caries in young children.
[0025] In this invention, the above-described technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of this invention will be set forth in the following description, and some advantages may become apparent from the description or be learned by practicing the invention. The objects and other advantages of this invention can be realized and obtained from what is particularly pointed out in the description and drawings. Attached Figure Description
[0026] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts. Figure 1 This is a schematic diagram of the initial screening results of the internal reference candidate region in an embodiment of the present invention; Figure 2 1D Amplitude scatter plots (Group 1) for droplet digital PCR detection of different bacteria, blank control and human genome. Figure 3 1D Amplitude scatter plots (Group 2) for droplet digital PCR detection of different bacteria, blank control and human genome. Figure 4 1D Amplitude scatter plots (Group 3) for droplet digital PCR detection of different bacteria, blank control and human genome. Figure 5 1DAmplitude scatter plots for droplet digital PCR detection of different bacteria, blank control and human genome (Group 4). Figure 6 1D Amplitude scatter plot of different primer-probe combinations for detecting internal control in droplet digital PCR; Figure 7 1D Amplitude scatter plot of different primer-probe combinations for detecting internal control in droplet digital PCR; Figure 8 1D Amplitude scatter plot of different primer-probe combinations for detecting internal control in droplet digital PCR; Figure 9 1D Amplitude scatter plot of different primer-probe combinations for detecting internal control in droplet digital PCR; Figure 10 1D Amplitude scatter plot of different primer-probe combinations for detecting internal control in droplet digital PCR; Figure 11 1D Amplitude scatter plot for the detection of internal control primer-probe combination by droplet digital PCR at different annealing temperatures; Figure 12 1D Amplitude scatter plot for the detection of internal control primer-probe combination by droplet digital PCR at different annealing temperatures; Figure 13 1D Amplitude scatter plot (sum) for droplet digital PCR detection of internal control primer-probe combination at different annealing temperatures; Figure 14 Scatter plot of 1D Amplitude for droplet digital PCR detection of internal control primer-probe combination at different annealing times; Figure 15 This is a 1D Amplitude scatter plot (blank control) obtained by combining the biomarker system and the internal control 16S system of this invention with droplet digital PCR for validation. Figure 16 The 1D Amplitude scatter plot (human genome) is obtained by combining and validating the biomarker system + internal control 16S system of this invention through droplet digital PCR. Figure 17 This is a two-dimensional amplitude scatter plot (for quantitative analysis of biomarkers) obtained by combining the biomarker system and the internal control 16S system of this invention with droplet digital PCR. Figure 18 This is a 2D Amplitude scatter plot (quantitative analysis of biomarkers) obtained by combining the biomarker system and the internal control 16S system of this invention with droplet digital PCR for validation. Figure 19This is a scatter plot of 1D Amplitude obtained by detecting different primer-probe combinations of the internal control 16S system in this embodiment of the invention using droplet digital PCR. Figure 20 The chart shows the ROC curves and area under the curve (AUC) of the five-phase primer-probe system (four biomarkers + internal reference 16S) used in Example 4 under single-phase and multi-phase detection conditions. Figure 21 The bar chart shows the results of the T-test for the relative abundance differences of the four microbial biomarkers in the non-carious (CF) group and the carious (CA) group in Example 4, displaying the mean (%) and statistical significance (P value) of each biomarker between the two groups. Detailed Implementation
[0027] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.
[0028] For experiments not specifically described in the examples, the procedures or conditions should be followed according to the conventional experimental procedures described in the literature in this field. Reagents, bacterial strains, or instruments whose manufacturers are not specified are all commercially available conventional reagent products.
[0029] In a first aspect, the present invention provides a 16S internal reference sequence for quantitative detection of caries microbial markers in young children, the internal reference sequence comprising an upstream primer sequence, a downstream primer sequence, and a probe sequence; The upstream primer sequence is SEQ ID NO.1; The downstream primer sequences are SEQ ID NO.2, SEQ ID NO.3, and SEQ ID NO.4; The probe sequences are SEQ ID NO.5, SEQ ID NO.6 and SEQ ID NO.7.
[0030] Secondly, the present invention provides a primer-probe composition for quantitative detection of microbial biomarkers for dental caries in young children, the primer-probe composition comprising the aforementioned 16S internal reference sequence and microbial biomarkers; the microbial biomarkers are respectively for detecting *Humanoidobacterium* (…). Cardiobacterium hominis ), Space Rose ( Rothia aeria ), Streptococcus mutans ( Streptococcus mutans The upstream primers, downstream primers, and probes for the three bacteria; The upstream and downstream primer sequences for detecting Human poxvirus are shown in SEQ ID NO.8 and SEQ ID NO.9, respectively, and the probe sequence is shown in SEQ ID NO.10. The upstream and downstream primer sequences for detecting *Rhodesia spp.* are shown in SEQ ID NO.11 and SEQ ID NO.12, respectively, and the probe sequence is shown in SEQ ID NO.13. The upstream and downstream primer sequences for detecting Streptococcus mutans are shown in SEQ ID NO.14 and SEQ ID NO.15, respectively, and the probe sequence is shown in SEQ ID NO.16.
[0031] It should be noted that the degenerate bases W (A / T) and R (A / G) in the upstream primer sequence (SEQ ID NO.1) are significant.
[0032] Table 1. Details of primer and probe compositions for internal controls and microbial biomarkers.
[0033] Example 1: Design and optimization of internal control sequence and primer-probe composition The primer and probe composition for quantitative detection of caries microbial markers in young children and the preferred amplification conditions described in this invention are obtained through the following steps: (1) Initial screening was conducted by searching "16S AND caries" on PubMed, yielding 455 records. Further screening was performed based on inclusion and exclusion criteria, and finally, 16S rRNA gene sequencing datasets from 9 studies were included for meta-analysis.
[0034] Specifically, the inclusion criteria were: (a) 16S rRNA gene sequencing using 454 or Illumina sequencing, (b) samples of supragingival plaque obtained from deciduous teeth of children aged 3–5 years, and (c) publicly available or shared sequences and related metadata. Specifically, the exclusion criteria were: (a) a history of systemic diseases, infectious diseases, and systemic medication use; (b) systemic antibiotic use within one month prior to sampling; (c) incomplete records or missing data in previous studies; and (d) use of topical fluoride within two weeks prior to sampling.
[0035] Specifically, the original sequence data and metadata of the nine studies ultimately included were obtained through three methods: downloading from NCBI Sequence ReadArchive, obtaining from supplemental materials of the original papers, or requesting them from the authors. The dataset includes both caries-free (CF) and caries-affected (CA) populations, and all studies focused on supragingival plaque samples from primary dentition. Detailed information on the final dataset, including authors, country of study, year, children's age, data source, sequencing platform, sequencing variable regions, and sample size, is shown in Table 2. Table 2. Characteristics of the nine studies included in the meta-analysis
[0036] (2) A method combining machine learning and linear discriminant analysis (LEfSe) was used to screen potential ECC microbial biomarkers. The specific technical solution is as follows: (2.1) In the machine learning process, a random forest model was constructed using the randomForest package in R language. Random forest models were constructed for each of the nine studies and analyzed. Receiver operating characteristic curves were then constructed based on the model using the pROC package. MeanDecreaseGini scores were ranked and plotted using ggplot2. The most important microbial features were extracted from the random forest model. Subsequently, 10-fold cross-validation was used to determine the optimal number of features in the random forest constructed based on CLR, and the results from multiple rounds were averaged and sorted (Table 3). This process was entirely based on R. First, the microbiome data for each study was imported using the phyloseq R package to extract species abundance tables and sample grouping information, converting the raw ASV counts into relative abundance. A random forest classification model was constructed using the randomForest R package, with the target variable being group labels (CF vs. CA) and the input feature being the species relative abundance matrix. The core model parameters were set as follows: decision tree ntree=50; minimum number of samples per terminal node nodesize=5 to prevent overfitting; and feature importance assessment was enabled to calculate the feature importance score based on Gini index decline (Mean Decrease Gini). Hierarchical 10-fold cross-validation was implemented using the caret R package. The dataset was divided into 10 equal parts, with one part selected as the validation set and the remaining 9 parts as the training set, iterating through all folds in turn. In each training round, the random forest model was trained using the same parameters as the main model to predict the caries probability in the validation set, and the area under the receiver operating characteristic (AUC) was calculated using the pROC R package. Finally, the average and standard deviation of the AUC over 10 rounds were used as the model performance evaluation metrics. Mean Decrease Gini scores for each species are extracted from the trained random forest model and sorted from highest to lowest. The top 20 species by importance score in each study are selected, and the frequencies across studies are summed to select the top 10 species by frequency.
[0037] (2.2) During the LEfSe analysis, LEfSe analysis was performed on each of the nine studies, and the LEfSe ranking results of each study were summarized and sorted by frequency (Table 3). For the microbiome data of each study, a standardized feature table (ASV table) was generated using QIIME 2 (version 2021.8), containing the microbial amplicon sequence variants (ASVs) of the samples and their relative abundance. Based on the extended human oral microbiome database (eHOMD, 16S rRNA RefSeq version 15.22), ASVs were annotated for species using a Naive Bayes classifier. The ASV tables were converted to a LEfSe-compatible input format and analyzed using the LEfSe analysis module of the Beijing Genomics Institute Bioinformatics Cloud Platform (https: / / www.bic.ac.cn / BIC). The Kruskal-Wallis rank-sum test (a non-parametric test) was used to screen for species with significant inter-group differences, with a significance threshold of p-value of 0.05, retaining only species with p < 0.05. Linear discriminant analysis (LDA) was used to assess the magnitude of the difference in species influence, with an LDA score threshold of 2.0, retaining only species with an LDA score of 2.0 or higher (i.e., only species with significant biological significance). The analysis covers all taxonomic levels from phylum to species, with the species level selected as the primary output by default. Nonparametric tests and LDA analysis were performed to obtain LEfSe analysis results at the species level for 9 studies. The top 20 species from each study were selected, and the frequencies across studies were summed to select the top 10 species by frequency.
[0038] (2.3) The intersection of the data obtained from machine learning and LEfSe analysis was used to finally obtain 7 species. These 7 feature markers have the highest discriminative potential in distinguishing between children with and without caries. Considering that low relative abundance may affect the accuracy and generalizability of species detection, species with an average relative abundance of less than 0.5% in the caries (CA) and caries-free (CF) subgroups of the 9 datasets were excluded. The finally identified potential ECC microbial markers include: Cardiobacterium hominis (Ch), Veillonella sp_HMT_780, Rothia aeria (Ra), and Streptococcus mutans (Sm) (Table 4).
[0039] Table 3. Frequency results of analyses using Random Forest and LEfSe
[0040] Table 4. Candidate ECC microbial biomarkers identified by machine learning and LEfSe analysis
[0041] (3) Design and optimization of the universal internal reference 16S primer-probe system, the specific technical solution is as follows: (3.1) The bioinformatics screening steps for internal control primers and probes are as follows: S1: Initial screening of candidate regions (region names and sites refer to the Escherichia coli_NR_024570 sequence, such as...) Figure 1 ) S1.1 Extract the reference sequences of 16S rRNA genes of all oral microorganisms from the eHOMD V3.1 database, i.e. eHOMD 16S rRNA RefSeq Version 15.23 (Starts from position 28); S1.2 uses MAFFT Alignment for multiple sequence alignment to filter segments that meet the following criteria: (a) All oral microbiota conservation >90% (only the primer-probe paired region) (b) Fragment length 80-250bp (suitable for digital PCR amplification system) S1.3 Based on the comparison results, the final selection was made in Figure 1 Primers and probes for internal controls were designed for the conserved regions (i.e., C5, C6, and C7 segments) at both ends of the variable region V6.
[0042] S2: Internal reference primer and probe design S2.1 candidate primers were designed using Primer Premier 5.0: The forward primer is used to locate the conserved region C5 to the left of region V6. The reverse primer is used to locate the conserved region C7 to the right of the V6 region; The probe is located in the conserved region C6 to the right of the V6 region. It is a reverse probe and is close to the reverse primer. Primer probe T m The value range is 56-65℃, and the probe T m Value ratio of primer T m The value is 1-3℃ higher.
[0043] S2.2 via IDT OligoAnalyzer TM Tool and MFE Primer 3.1 Verification of primer and probe structure: Neither the primers nor the probes exhibited strong secondary structures, including hairpin structures, self-dimers, or heterodimers. The primer dimer has no strong pairing at the 3' end.
[0044] S2.3 Specificity was verified using NCBI Primer-BLAST: Parameter settings: Product size (70-1000bp), maximum mismatch number (3), T mValue (55-65℃), Database (nt), Organization: select Homo sapiens (taxid:9606); Excluding non-target genomes: No cross-reactivity was observed when comparing with the human genome, meaning that the human genome could not be amplified.
[0045] S2.4 Coverage verification via Silva database: Parameter settings: SILVA database (ssu-138.2 version), sequence set (reference non-redundant set), maximum number of mismatches (2 mismatches), length of the 3' end mismatch-free region (3 bases); The upstream and downstream primers have a bacterial coverage rate of more than 85% in the taxonomic unit.
[0046] Table 5 Validation and Acceptance Criteria for Internal Reference Primers and Probes
[0047] (3.2) Optimization of internal control primers and probes after screening Following the screening in step 3.1, the present invention further optimized the selected internal control primers and probes. Specifically, while ensuring high coverage of oral microbiota, low-degenerate or parallel allele sequences were introduced into the conserved sites of the upstream and downstream primers and probes of the internal control to eliminate mismatches in some strains and reduce background positivity on NTC and human DNA. This allowed the 16S internal control system to stably coexist with the specific primers and probes of dental caries marker bacteria in the same ddPCR reaction. The results before and after optimization are shown in Table 6. Specifically, the optimization motivation, optimization strategy, optimization steps, and optimization effects are as follows: (3.2.1) Optimization Motivation (a) Reduce nonspecific amplification: Some candidate primers have potential mismatches with human sequences or non-target bacterial species, resulting in background signals.
[0048] (b) Improve cross-species coverage: The binding efficiency of the initial screening primers is low in a few bacterial populations, which may affect the universality of the general internal reference.
[0049] (c) Enhance signal stability: The Tm value of the original probe does not match the primer perfectly, which can easily cause unstable fluorescence signal or decreased amplification efficiency.
[0050] (3.2.2) Optimization Strategy Site selection: The conserved regions at both ends of the V6 region of the 16S rRNA gene are preferred as primer binding sites. In probe design, the C6 region adjacent to the reverse primer is selected to ensure that the amplified fragment length is appropriate (80–250 bp), which is conducive to droplet separation and stable amplification in ddPCR.
[0051] Sequence optimization: By introducing degenerate bases, the cross-species coverage is improved, and the Tm value of the probe is adjusted to be 1–3 °C higher than that of the primers to enhance pairing stability.
[0052] Structural screening: OligoAnalyzer and MFEprimer are used to exclude sequences that may form dimers, self-complementary structures, or hairpin structures, thus avoiding false positives at the source.
[0053] (3.2.3) Optimization steps 1) Upstream primer: U16S_898BF1 First, Primer Premier 5.0 automatically generated the upstream primer U16S_898BF1 for the conserved region (C5) on the left side of V6. After verification by MAFFT, NCBI Primer-BLAST and SILVA testprimer, this sequence was able to obtain effective matches in 94.9% of oral bacteria and had no amplicon with the human genome. Therefore, it was retained as the basic sequence and finally adopted.
[0054] Based on this, the inventors introduced 1–2 bases at the 5′ end to form two parallel versions, U16S_898F1.1 (TTCGGGTAGCGAACAGGATTAGAT) and U16S_898F1.2 (GTTCGGGTAGCGAACAGGATTAGAT), which were used to further optimize the annealing temperature and amplification stability in the ddPCR reaction system.
[0055] 2) Downstream primers: U16S_1106R1→R1-1 / R1-2 / R1-3 / R1-4 / R1-5 The downstream primer, U16S_1106R1, was initially designed automatically by software in the conserved region C7 to the right of V6, with the original sequence being AAGGGTTGCGCTCGTT. This sequence showed good conservation in various oral bacteria 16S rRNA gene templates. Based on this, the inventors designed allelic extensions and length variants for U16S_1106R1 according to the measured base polymorphism at this site in the eHOMD database. [Original software segment] AAGGGTTGCGCTCGTT; R1-1: This is the U16S_1106R1 automatically designed by the software; R1-2: Final sequence: AAGGGTTGCGCT A GTT; [Reserved software segment] AAGGGTTGCGCT C GTT; [Artificially modified section] Change the middle C to A; The only base difference between R1-1 and R1-2 is the middle C → A (…CGCT) C GTT vs…CGCT A GTT): This was discovered after aligning the eHOMD oral bacteria sequence and finding that the dominant polymorphism at this site was C / A. Therefore, by using two completely independent primers instead of adding degeneracy in one primer, the Tm of both primers can be locked at the same temperature, avoiding the situation where "the Tm of the half of the sequence becomes lower after degeneracy".
[0056] R1-3: Final sequence: GTGGGTCTCGCTCGTT [Reserved software segment] CGCTCGTT [Manually modified section] Replace the entire AAGGGTTG segment at the 5' end with GTGGGTCT. R1-3: The head is changed to GTGGGTCT… to accommodate another type of sequence variation at the 5′ end. In order to inherit another group of sequences starting with G in the conserved region on the right side of V6 and to avoid the risk of self-dimerization caused by excessive degeneracy superimposed on the same primer, the significantly different sequence U16S_1106R1-3 was designed as a parallel allelic version.
[0057] In addition, the inventors designed two length variants: one is a short version U16S_1106R1-4 with the 5′ end “AA” omitted, and its sequence is GGGTTGCGCTCGTT; The other is a longer version, U16S_1106R1-5, which is formed by adding a sequence to the 5' end of U16S_1106R1-1. Its sequence is: CGTAACA AAGGGTTGCGCTCGTT.
[0058] The short version U16S_1106R1-4 and the long version U16S_1106R1-5 were used to evaluate the effects of 5′ end length variation on amplification specificity, annealing window, and background signal, respectively.
[0059] In summary, the software-generated U16S_1106R1 primer was expanded into five parallel allelic versions, U16S_1106R1-1 to U16S_1106R1-5, based on the polymorphism of the oral bacteria 16S rRNA gene at this landing site and the 5′ length requirement. Each base substitution or 5′ sequence rearrangement was designed to accommodate a mainstream variant still existing in the conserved region, while maintaining high consistency in Tm, GC content, and secondary structure scores across primers. This avoids the risk of dimers and false positives caused by introducing excessive degeneracy into a single primer. In subsequent ddPCR combinatorial screening, U16S_1106R1-1 to R1-3 were the main candidates, while U16S_1106R1-4 and R1-5 were used to assist in validating the performance of different length versions.
[0060] 3) Probe changes: U16S_1064rP1 → U16S_1064rP1-1 / P1-2 / P1-3 [Original Software Segment] TCGTCAGCTCGTG The probe U16S_1064rP1 falls within the conserved C6 region of V6, corresponding to the initial software sequence TCGTCAGCTCGTG. The short sequences such as AGCTC / AGCTG (collectively referred to as the AGCTN framework) in the middle of the C6 conserved region are highly conserved in oral bacteria 16S rRNA genes, but some single-base polymorphisms still exist on either side of them. In the initial design, the core sequence of U16S_1064rP1 was TCGTCAGCTCGTG, which is the current sequence of U16S_1064rP1-1.
[0061] The part was artificially split into three sections •P1-1: TCGTC+common segment AGCTCG +TG •P1-2: TA+common segment CGAGCTGACG+GC •P1-3: CA+common segment CGAGCTGTCG+AC.
[0062] The initially designed U16S_1064rP1 probe was split into three short probes, U16S_1064rP1-1, P1-2, and P1-3, based on the microvariable sites in the C6 region. All three probes were modified with MGB, ensuring that while maintaining high Tm, each probe corresponds to a different clinical sample sequence branch, thus maintaining compact fluorescence clustering and cross-sample consistency in the multiplex ddPCR system. The initial probe U16S_1064rP1 targets the conserved C6 site, but oral microbiota still exhibits several high-frequency single-base variations at this site. Introducing degeneracy into a single probe would lead to uneven fluorescence intensity in ddPCR droplets. Therefore, the original probe was split into three sequencing versions (rP1-1 / -2 / -3) with approximately the same length and similar Tm. Each version differs from the original backbone only in the first and last 2–3 bases, matching different alleles to achieve higher community coverage and more compact fluorescence clustering without increasing probe degeneracy.
[0063] (3.2.4) Optimization effect (a) Enhanced sensitivity: The detection limit can be as low as 0.01 copies / μL, enabling stable detection even in low abundance samples.
[0064] (b) Enhanced specificity: False positive amplifications in NTC and human genome samples were significantly reduced (<10 copies / μL), effectively avoiding interference from false signals.
[0065] (c) Improved stability: The amplification efficiency remained consistent under different bacterial strains and different experimental batches, with a relative standard deviation of ≤5%, ensuring the absolute quantitative reliability of ddPCR.
[0066] Table 6. Information on internal control primer and probe sequences obtained through biological screening.
[0067] In summary, this invention does not simply use existing universal 16S primers. Instead, based on multiple sequence alignment of the eHOMD oral microbiome database, the software first generates initial primer-probe sequences located in the conserved regions at both ends of V6. Then, according to the actual polymorphism of this site in the oral microbiota, a single degenerate base is introduced into the upstream primer at the key polymorphic site, the downstream primer is split into 5 parallel allelic versions, and the probe is split into 3 short probes + MGB parallel versions. Subsequently, ddPCR is used to screen for combinations of blank control, human genome, and various common oral bacteria. With fine adjustment of annealing temperature and annealing time, background signals caused by mismatches are gradually eliminated. The final primer-probe combination has high coverage (bacteria coverage 94.9%), high specificity (human and NTC amplification <10 copies / μL), and multiplex detection compatibility. Thus, absolute quantitative detection of dental caries microbial markers and total bacterial count in young children can be performed simultaneously in the same ddPCR reaction system.
[0068] (3.3) Based on the artificially optimized sequence in (3.2), the internal reference 16S primer probe was further screened and optimized, and the steps are as follows: S1: Optimization of 16S primer and probe screening; By using permutation and combination methods, multiple degenerate variants of the upstream primers, multiple parallel versions of the downstream primers, and multiple parallel versions of the probes listed in Table 6 were combined to obtain different primer-probe combinations. The amplification effects of different combinations were verified using the bacterial strains listed below, as well as blank controls and human genomes. Some of the combination methods and results are shown in Tables 7 and 8. Figures 2-5 As shown; NTC: Blank control; Homo-G: Human genome; Ab: Acinetobacter baumannii; Ef: Enterococcus faecium; Kp: Klebsiella pneumoniae; Sm1: Streptococcus mutans, Streptococcus mutans-1; Sm2: Streptococcus mutans; Table 7 Internal control primer-probe combination and amplification results
[0069] Table 8. Amplification results of different internal control primer-probe combinations
[0070] The results showed that the U16S_1064rP1-1 probe had good amplification effect (with the first group of primers showing the best effect), and it could effectively amplify the virus in four different bacteria. However, the specificity was poor, with a certain amount of non-specific amplification in both NTC and human genomes, requiring further optimization of primers.
[0071] S2: Screening of 16S-specific primers for internal control; By combining multiple degenerate variants of the upstream primers, multiple parallel versions of the downstream primers, and U16S_1064rP1-1 determined in step S1, different primer-probe combinations were obtained. The amplification effects of these different combinations were verified using Streptococcus mutans, Streptococcus mutans (Sm), and the blank control NTC. The combination methods and results are shown in Tables 9 and 9. Figures 6-10 As shown.
[0072] Specificity screening results at an annealing temperature of 58℃ showed that the combination of the U16S_1064rP1-1 probe with U16S_898F1.1 and U16S_1106R1-4 exhibited the best amplification specificity, with a significant reduction in non-specific amplification. Compared with other combinations, this combination maintained a high amplification efficiency in Sm templates (>1000 copies / μl), while the amplification amount in NTC was significantly lower than 30 copies / μl, indicating that its specificity was superior to other combinations.
[0073] Table 9 Primer-probe combinations and amplification results
[0074] S3: Specific optimization of the internal reference 16S system - annealing temperature The upstream primers U16S_898BF1 and U16S_898F1.1, which are internal controls in Table 6, were selected. These two upstream primers were then combined with U16S_1064rP1-1 and U16S_1106R1-4, which were selected in step S2, to obtain two sets of primer-probe combinations. Amplification was then performed at different annealing temperatures of 56℃, 58℃, and 60℃. The combination methods and results are shown in Tables 10 and 11. Figures 11-14 As shown.
[0075] The comparison results at three annealing temperatures of 56℃, 58℃ and 60℃ showed that the primer and probe combination of U16S_898F1.1 had better amplification effect at annealing temperature of 60℃, less non-specific amplification of NTC, and did not affect the amplification efficiency of target gene.
[0076] Table 10 Primer-probe combinations and amplification results (at different annealing temperatures)
[0077] Table 11 Amplification results of internal control primer-probe combinations at different annealing temperatures
[0078] S4: Specificity Optimization of the Internal Reference 16S System - Amplification Procedure Two different amplification programs were used, both starting with a 5-minute pre-denaturation at 95°C, followed by a 20-second denaturation at 94°C, and then annealing at 60°C for 60 seconds or 20 seconds, respectively. The experiment used DW-DNA PCR Super Mix ddPCR premix as the reaction system. Templates included DNA from various bacteria such as NTC (blank control), HOMO_G (human genome), Ab (Acinetobacter baumannii), Ef (Enterococcus faecium), and Sm (Streptococcus mutans). The selected primer-probe combination was U16S_898F1.1 (upstream primer), U16S_1106R1-4 (downstream primer), and U16S_1064rP1-1 (probe). The number of positive droplets, negative droplets, and total droplets were detected using droplet digital PCR (ddPCR) technology, and the copies / μl were calculated. The experimental protocol and results are shown in Tables 12 and 13. Figure 14 As shown.
[0079] Table 12 Primer-probe combinations and amplification results (different annealing times)
[0080] Table 13 Amplification results of internal control primer-probe combinations at different annealing temperatures
[0081] Results: The primer-probe combination (U16S_898F1.1, U16S_1106R1-4, U16S_1064rP1-1) significantly reduced nonspecific amplification (<10 nonspecific sites) when the amplification annealing temperature Ta was 60℃ for 20s, and had little impact on the quantification of medium and high concentration templates.
[0082] Furthermore, the artificially optimized upstream primers (U16S_898F1.1, U16S_898F1.2, U16S_898BF3, U16S_898BF4) in Table 6 were further optimized, and finally determined to be U16S_898F (SEQ ID NO.1) in Table 1. The specific optimization ideas and processes are as follows: Within the conserved C5 region to the left of the V6 region of the 16S rRNA gene, a set of upstream primer candidates, namely U16S_898BF1, U16S_898F1.1, U16S_898F1.2, U16S_898BF3, and U16S_898BF4, were designed based on multiple sequence alignment results from the eHOMD database (see Table 6). The binding sites of these candidate primers partially overlap, their 3′ ends generally contain common motifs such as "AGGA", and their 5′ ends also show high conservation at tetranucleotides such as "AAAG". Based on this, the inventors re-aligned the representative 16S sequence within the conserved region of C5 and optimized the starting position of the binding site and primer length by combining the coverage of candidate primers, mismatch sites, and secondary structure prediction results. A 21 bp binding window was redefined on the conserved fragment that partially overlapped with the binding region of the primers in Table 6. At the same time, degenerate bases W (A / T) and R (A / G) were introduced at two sites with obvious nucleotide polymorphisms, respectively, to design the final degenerate upstream primer U16S_898F: AAACTCAAAGGAATWGRCGG (SEQ ID NO.1). The primer extends “AAACTCAA” at the 5′ end to balance Tm and GC content, and retains conserved motifs such as “AAAG” and “AGGA” shared with the original candidate primers at the 3′ end. It can cover the main sequence variations involved in the original BF1 / F1.1 / F1.2 / BF3 / BF4 with a single primer, while reducing the complexity of the primer pool and facilitating stable application in multiplex ddPCR systems.
[0083] Furthermore, the artificially optimized downstream primers (U16S_1106R1-1, U16S_1106R1-2, U16S_1106R1-3, U16S_1106R1-4, U16S_1106R1-5) in Table 6 were further optimized, and finally determined to be U16S_1106R1-1 (SEQ ID NO.2), U16S_1106R1-2 (SEQ ID NO.3), and U16S_1106R1-3 (SEQ ID NO.4) in Table 1. The specific optimization ideas and processes are as follows: Using the Escherichia coli 16S rRNA gene reference sequence NR_024570 as coordinates, within the conserved C7 region to the right of the V6 region, a basic downstream primer U16S_1106R1 was first obtained based on multiple sequence alignment results and automatic software design. Considering that there are still certain nucleotide polymorphisms in the C7 region among different oral bacteria, in order to take into account the matching of different sequence branches, this invention artificially adjusted the 5′ length and individual bases in the middle around the binding site of U16S_1106R1, forming multiple candidate downstream primers U16S_1106R1-1~U16S_1106R1-5 as shown in Table 6, wherein: U16S_1106R1-1 and U16S_1106R1-2 mainly match common polymorphic sites such as C / A and G / A in the C7 region by replacing individual bases near the 3′ end. U16S_1106R1-3 maintains the conservative motif at the 3′ end while moderately extending or adjusting the sequence at the 5′ end to accommodate another type of G-starting sequence branch; U16S_1106R1-4 is a version with a moderately shortened 5′ end, used to explore the performance of shorter primers in terms of specificity and amplification efficiency; U16S_1106R1-5 is a version with a further extension at the 5′ end, used to evaluate the impact of a combination of high Tm and high GC content on amplification stability.
[0084] Candidate primers U16S_1106R1-1 to U16S_1106R1-5 collectively cover the main variations of representative sequences of different oral bacteria within the C7 region. After obtaining multiple candidate downstream primers in Table 6, this invention first systematically evaluates the conservation of their binding sites in the C7 region. Multiple sequence alignment of the 16S rRNA gene sequences of dominant oral bacteria in databases such as eHOMD was performed, focusing on the consistency of the 8–10 bp at the 3′ end of each primer. The results showed that the 3′ ends of U16S_1106R1-1, R1-2, and R1-3 all fall within the highly conserved core region of the C7 region, with almost no mismatches in the representative sequences; while the 3′ end of U16S_1106R1-4 is located on a relatively poorly conserved fragment, with 1–2 nucleotide mismatches appearing in some strains, and these mismatches are close to the critical pairing region at the 3′ end. Theoretically, 3′ mismatches can lead to a significant decrease in amplification efficiency in some strains (risk of false negatives) and may also increase the probability of non-specific amplification in complex templates. Therefore, considering both conservation and mismatch risk, this invention deems the position of U16S_1106R1-4 poorly conserved and unsuitable as a final internal control downstream primer, and thus R1-4 is no longer retained in subsequent optimizations.
[0085] In further structural prediction and experimental verification, this invention used tools such as OligoAnalyzer to evaluate each candidate downstream primer, and the results showed: The Tm values of U16S_1106R1-1, R1-2, and R1-3 are concentrated in a range similar to those of the upstream primer U16S_898F and the probe at site 1064, with moderate GC content and low risk of self-dimerization and cross-dimerization. U16S_1106R1-5, due to its significantly extended 5′ end, has a significantly higher overall Tm, making it difficult to unify the annealing temperature in multiplex systems, and it also shows more potential dimer structures in secondary structure prediction. Based on the results of singlet and multiplex ddPCR preliminary experiments, while R1-5 can achieve some amplification under high annealing temperatures, it exhibits a certain inhibitory trend on the amplification of other channels in multiplex systems. Therefore, U16S_1106R1-5 was ultimately not retained as a final downstream primer in this invention. Based on the above analysis, this invention retains three downstream primers—U16S_1106R1-1, U16S_1106R1-2, and U16S_1106R1-3—from Table 6, and mixes them in a certain molar ratio to use as a downstream primer combination for the 16S internal control (see Table 1). These three primers retain the same conserved motif at the 3′ end and are differentially designed at the 5′ or mid-segment positions targeting polymorphic sites in different branches of the C7 region. Functionally, they are equivalent to an allelic combination of a "degenerate downstream primer," which can: cover more C7 region variation types of oral bacterial strains, improving the universality and robustness of the internal control for complex oral flora; maintain Tm matching with the upstream primer U16S_898F and the 1064 site probe, facilitating multiplex ddPCR amplification at a uniform annealing temperature; control the risk of primer self-complementarity and cross-dimerization, and no significant channel crosstalk or amplification inhibition was observed when used in conjunction with dental caries microbial biomarker primers and probes.
[0086] Therefore, after comprehensive evaluation and experimental verification of the conservation, mismatch risk, Tm and structural characteristics of the downstream primer candidates in Table 6, this invention finally determined the final downstream primer scheme with the downstream primer mix composed of U16S_1106R1-1, U16S_1106R1-2 and U16S_1106R1-3 as the 16S internal reference.
[0087] (4) The primers and probes for the three biomarkers selected above were artificially optimized. The specific ideas and steps are as follows: 1. Cardiobacterium hominis (Ch) Forward F1: CCAGGCCTTGACATCCTAG [Original Software Segment] CCAGGCCTTGACATC [Artificial Tailing] CTAG The first 12bp (CCAGGCCTTGAC) is the GC high segment retained after the initial screening by the software, which ensures that the Tm is not too low. The following ATCCTAG is retained manually because these 7 bases appear consecutively in the Ch reference sequence. Among the several "interfering oral bacteria" to be excluded, the middle 1-2 positions will change to G or C. Therefore, this complete ATCCTAG segment is retained to keep out "closely related bacteria that are only 1-2bp different from it".
[0088] Reverse R1: CATGCAACACCTGTCTCTAGGTT [Original Software Segment] CATGCAACACCTGTCT [Artificial Extension] CTAGGTT The middle "ACACCTGTCTC" is the "ACACCTGTC" provided by the software, extended by two bases (...TC). This is to lengthen the fully matched region at the 3' end, so that even if the annealing time is reduced from 60 s to 20 s in multiplexing, it can still attach first. The last three "GTT" are the strong 3' end, which facilitates hot-start enzyme binding and also reduces weak pairings on the human genome.
[0089] Probe P1: TTGGCAGAGATGCCTT [Original Software Segment] TTGGCAGAGATG [Manual Finishing] CCTT The middle AGAGATG site is stable in Ch, but this position changes in other oral bacteria. The CCTT at the end is to raise Tm, maintaining "probe elevation of 1–3°C" with F / R; The Ch primer probe consists of a series of small segments, from the forward ATCCTAG to the AGAGATG segment within the probe. Each segment is a unique fragment specific to Ch, so it cannot be degenerate and can only be ordered.
[0090] II. Rothia aeria (Ra) 1. Forward F1: GAAGAACCTTACCAAGGCTG [Original Software Segment] GAAGAACCTTACCAAG [Manual GC Tail Completion] GCTG The first part, GAAGAA, is a common prefix in the genus *Rothia* and is selected by the software for initial screening. The middle part, CCTTACCAAG, represents a conserved segment within the genus but a variant segment outside the genus. The addition of GCTG at the end is a terminal base added to match the uniform annealing temperature of the multiplex system.
[0091] 2. Reverse R1: CACGAGCTGACGACAACC [Original Software Segment] CACGAGCTGACG [Manual replacement / extension] ACAACC While inheriting the CACGAGCTGACG backbone, the terminal sequence is replaced with AACC. This maintains the overall GC content and Tm unchanged, and eliminates the potential complementarity with the 16S probe and its reverse primer, thereby reducing the risk of heterodimers in multiple reactions.
[0092] 3. Probe P1: ACATGCATTAGATCGCGTCAGA [Original Software Segment] ACATGCATTAGATC [Artificial incorporation of the second characteristic site + termination] GCGTCAGA The preceding ACATGCATTA is a stable and conserved sequence within the Ra lineage. The following GATCGCGTCA is used to distinguish it from other oral actinomycetes. A GATC tetrad is inserted in the middle to prevent the probe from forming internal palindromes, thus balancing specific recognition and structural stability.
[0093] 3. Streptococcus mutans (Sm) 1. Forward F1: TTGGAAACGATAGATAATACCGCATA (26bp) [Original Software Segment] TTGGAAACGATAGATAATA [Artificial High GC Tail] CCGCATA For the automatically generated Sm forward primer, while keeping the characteristic continuous sequence 'AGATAATA' of Sm unchanged, 'CCGCATA' is added to the 3' end to enhance Tm and form a high-fidelity 3' end recognition region, thereby distinguishing Sm from other oral streptococci in a multivariate system.
[0094] 2. Reverse R2: TAATACAACGCAGGTCCATCTACTA (25bp) [Original Software Segment] TAATACAACGCAGGTCC [Artificial Extension] ATCTACTA Extend the primers by 2-3 bases to the 3' end to ensure specific recognition of the Sm sequence at multiple annealing temperatures, while avoiding complementarity with the reverse primer of the internal reference 16S.
[0095] 3. Probe rP1: ATCTTTCAATCAATTATCATG (21bp) [Original Software Segment] ATCTTTCAATCAA [Artificial Extension] TTATCATG Considering the high AT content in the Sm target region, stable hybridization cannot be achieved at 60℃ if the same 13-14bp short probe as the internal control is used. Therefore, while keeping the Sm characteristic sequence CAATCAA unchanged, the probe length is extended to 21bp and terminated with TCATG at the 3' end to improve hybridization stability and form a compact amplification window with the Sm primer pair.
[0096] After determining the internal control 16S primer-probe combination, this invention aims to simultaneously detect three microbial biomarkers related to dental caries in young children—Cardiobacterium hominis, Rothia aeria, and Streptococcus mutans. Species-specific sequence fragments were screened from their 16S rRNA genes. Initial primers and probes were generated using software, and then manually compared with sequences of easily confused strains in the oral microbiota. The primer length, 3' end bases, GC content, and probe coverage sites were fine-tuned site-by-site: for the Sm sequence with high AT content, high-GC bases were added to the tail to increase Tm; for the Ra sequence, which has thermodynamic properties similar to the internal control, the 3' end bases were replaced while inheriting the common backbone to eliminate complementarity with the internal control; for the Ch sequence, its species-specific continuous fragment was maintained and the 3' end of the reverse primer was extended to ensure specific amplification capability in multiplex systems. Each addition, substitution, or deviation from degenerate form mentioned above is a targeted optimization made in response to deficiencies in specificity, Tm, or potential dimer risks discovered in the previous round of software or experimental comparisons. It has clear technical motivations and continuity.
[0097] (5) Validation of the combination of biomarker system and internal control 16S system Experimental steps 1) Primer and probe configuration: A five-component primer-probe system was used to target four species (Cardiobacterium hominis, Rothiaaeria, Veillonella sp_HMT_780, Streptococcus mutans) and a universal 16S internal reference primer-probe combination. The specific sequences of the primers and probes are shown in Table 14.
[0098] Table 14 Primer and probe sequences for the biomarker system + internal control 16S system
[0099] 2) ddPCR reaction system preparation: see Table 15 Table 15. ddPCR reaction system configuration for primer and probe combination validation
[0100] 3) Sample grouping Group 1: NTC (blank control) and human genome (Homo-G).
[0101] Group 2: NTC, human genome, single-species templates (Sm, Ra, V.780), and mixed templates (S.m+R.a+Ch), see Table 16. Among them, 293T cells are a derivative of a human embryonic kidney cell line commonly used in biomedical research; Sm is an abbreviation for Streptococcus mutans; Ra is an abbreviation for Rothiaaeria; V.780 is an abbreviation for Veillonella sp_HMT_780; and Ch is an abbreviation for Cardiobacterium hominis. Sample-1 is Sample 1 from the kindergarten validation samples.
[0102] Table 16 Sample Grouping Information
[0103] The above experimental groupings were designed to verify the system's specificity, sensitivity, quantitative accuracy, and anti-interference ability. The grouping design is as follows: Group 1: Used to validate background signals and nonspecific amplification. NTC (Negative Control): No template added; used to detect background signals. Human Genome (Homo-G): Human genomic DNA added; used to detect nonspecific amplification.
[0104] Group 2: Used to verify the specificity, sensitivity, quantitative accuracy, and anti-interference ability of the system. NTC (blank control): Same as Group 1. Human genome (Homo-G): Same as Group 1. Single species template: DNA of a single species (e.g., Sm, Ra, V.780) is added to verify the specific amplification of a single target sequence. Mixed template: DNA of a mixture of species (e.g., Sm + Ra + Ch) is added to verify the performance of the system in mixed samples of multiple species. Sample-1: Used to verify the specificity / sensitivity and quantitative accuracy of the combination of the biomarker system designed in the embodiments of this invention and the internal control 16S system (Table 14).
[0105] 4) Amplification Amplification was performed using the dPCR-D600 system. The amplification program was as follows: (a) Pre-denaturation: 95°C, 5 minutes; (b) 45 cycles: denaturation at 94°C for 20 seconds; annealing at 60°C for 60 seconds, repeating the denaturation-annealing process for a total of 45 cycles; Photography parameters: FAM, HEX, ROX, Cy5 channels, with exposure times of 1.5 seconds, 1.5 seconds, 1.5 seconds, and 1 second respectively; see Table 17; In the experiment, the five fluorescent channels of ddPCR were used to detect the presence and intensity of the following microbial biomarkers and internal controls, as shown in fluorescence intensity (AU): HEX channel (internal control 16S rRNA gene, i.e., the total 16S internal control sequences listed in Table 1 or Table 14 as a bacterial total control), FAM channel (Streptococcus mutans), ROX channel (Rothiaaeria), CY5 channel (Veillonella HMT 780), and ATTO-425 channel (Cardiobacterium hominis). It is understandable that Cy5 labeling was used in some early exploratory experiments to evaluate the feasibility of different fluorescence schemes. Due to the channel configuration limitations of the instrument used at that time, Cy5 and ATTO-425 could not be used simultaneously in the same operating protocol; therefore, the channel data for ATTO-425 and Cy5 are not shown simultaneously in Tables 18 and 19. However, the multiplex ddPCR dental caries biomarker detection system ultimately established in this invention does not use the Cy5 channel. Instead, it uniformly uses ATTO-425 as the reporter fluorescent dye for the corresponding target, and combines it with other channels such as FAM and HEX. The selection of fluorescent dyes only involves wavelength labeling of different channels and does not change the primer and probe sequences, amplification fragments, or the reaction system itself, nor does it affect the construction and detection performance of the final multiplex ddPCR system of this invention. The method of this invention can achieve stable multiplex detection on qPCR / ddPCR instruments equipped with corresponding excitation / emission channels by configuring non-interfering fluorescent labels such as FAM, HEX, and ATTO-425.
[0106] Table 17: Amplification Procedure and Imaging Parameter Settings
[0107] 5) Results and Analysis (a) HEX channel (total 16s internal reference sequence) specificity: from Table 18 and Figure 15 and Figure 16The specificity analysis (NTC, human genome) showed that NTC and human genome (293T) had constant non-specific occurrences (800-900 copies / reaction); this indicates that the system still produces a certain background signal in the absence of the target sequence. Although the background signal exists, it is constant, which shows that the internal control 16S system has good stability when dealing with non-specific signals.
[0108] (b) From Table 19 and Figure 17 As can be seen from the Biomarker quantitative analysis, all channels of the quadruple (four biomarkers) system can be amplified normally, with good positive and negative differentiation; each channel can detect the target sequence, indicating that the biomarker system of the present invention has good sensitivity and can detect low concentrations of target sequences.
[0109] (c) From Table 19 and Figure 17 It can be seen that when amplifying a single biomarker (3-g2, 4-g2, 5-g2), the HEX (total 16s) quantitative value minus the non-specific value is consistent with the corresponding biomarker specific quantitative value (deviation <5%). This indicates that it has high quantitative accuracy in the detection of single target sequences, can accurately distinguish and quantify single target sequences, and has good specificity. (d) From Table 19 and Figure 17 It can be seen that when amplifying the mixed templates (7-g2, 8-g2), the quantitative values of S.m+R.a+Ch are consistent with the quantitative values of HEX (total 16s) (deviation <3%), which indicates that the system has high quantitative accuracy in mixed samples of multiple species; and the addition of high concentration of human gene template (10000 copies) has no effect on biomarker amplification, and the system can still accurately detect and quantify the target biomarker, indicating that the system has good anti-interference ability.
[0110] (e) Sample 6-g2 was derived from real clinical dental plaque samples. The detection results showed that different levels of signal were detected in the FAM (Streptococcus mutans), HEX (16S internal reference sequence), ROX (Rothia aeria), and Cy5 (Veillonella HMT 780) channels. This indicates that the five-fold primer-probe combination system of this invention can be successfully applied to the simultaneous detection of complex clinical samples. In this specific sample, the total bacterial load (HEX channel, 41837 copies / reaction) was high, and three target microorganisms were simultaneously present: *Streptococcus mutans* (3200.5 copies / reaction), *Rothia aquaticus* (5313.5 copies / reaction), and *Pontius cordata* (5114.25 copies / reaction).
[0111] (f) The data crossed out in Table 19 are excluded because their corresponding signal values are extremely low (e.g., 1.25 copies / reaction). These data are non-specific background signals and are therefore excluded in the data analysis.
[0112] In summary, the results above show that the five-phase primer-probe system of the present invention (Table 14) performs well in terms of specificity, sensitivity, quantitative accuracy and anti-interference ability, and is suitable for detection and quantitative analysis of complex samples and low concentrations.
[0113] The meanings of the sample names in Tables 17 and 18 are as follows: Taking 4-g1(NTC) as an example, g1 corresponds to the first group in Table 16, 4 corresponds to the fourth column in Table 16, and 4-g1(NTC) corresponds to NTC (blank control) in the fourth column of the first group in Table 16; the other sample names are similar.
[0114] Table 18: ddPCR amplification results - specificity analysis
[0115] Table 19 ddPCR Amplification Results - Quantitative Analysis
[0116] Example 2: Verification of the synergistic relationship between primers and probes for the internal reference sequence (1) Experimental steps (a) ddPCR reaction system configuration: A 25 μL reaction system was prepared, with specific components as shown in Table 21; among which, Temple is bacterial DNA, which was obtained from a dental plaque sample from a child in a Beijing kindergarten: (2) Experimental grouping: Group A: Complete system, including all primers and probes for the internal control; Group BH: Each group contains a primer or probe that lacks the internal control, as detailed in Table 21. (3) Amplification procedure: Table 20 Table 20 Amplification procedures and imaging parameters for validating the 16s internal reference synergy.
[0117] (4) Results Analysis To verify the amplification advantage of the internal reference 16s in the five-primer-probe system in the simultaneous detection of multiple targets, this embodiment designed eight amplification reactions (A~H) with different primer-probe combinations containing the internal reference, and compared their amplification efficiency under the same amplification conditions / programs.
[0118] Table 22 and Figure 18The results showed that the amplification effect was optimal only when all target primers and probes of the internal control 16s were used (Group A), with significantly higher detection sensitivity than other combinations, achieving a quantification result of 1659.62 copies / μl. In contrast, any combination lacking one primer or probe (Groups B–H) resulted in a significant decrease in amplification efficiency. The amplification amounts of Groups E and F were 328.84 copies / μl and 193.24 copies / μl, respectively, while Group H had the worst amplification effect, at only 11.94 copies / μl.
[0119] This result demonstrates that the internal reference 16s in the five-phase primer-probe system (four biomarkers + internal reference 16s) designed in Example 1 exhibits a significant synergistic effect in simultaneous multi-target detection. The absence of any primer or probe for the internal reference disrupts the overall synergistic amplification of the system, leading to a substantial decrease in detection sensitivity. Therefore, the complete five-phase system, through its synergistic effect, offers significant advantages in high-sensitivity quantitative detection and complex sample analysis.
[0120] Table 21. ddPCR system and grouping for synergistic relationship of internal reference 16S.
[0121] Table 22 Test results for different primer-probe combinations for internal reference 16s.
[0122] Example 3: Application of the five-phase primer-probe system (four biomarkers + internal control 16s) in actual detection, and verification of the primer-probe synergistic relationship in the system. (1) Experimental steps (a) Sample collection and processing: Sample collection: After obtaining the permission and informed consent of the guardians, 209 supragingival plaque samples were collected from multiple kindergartens in Beijing, including 83 samples from the caries-free CF group and 126 samples from the caries-prone CA group.
[0123] The sample size of this invention is statistically significant. Specifically, this validation study is a nested case-control study design, aiming to compare the relative abundance of candidate bacterial genera in children with primary dental caries and those without caries. Based on previous research, four annotable candidate bacterial genera were obtained. The mean and standard deviation of the relative abundance of each genera in the caries-affected and caries-free groups were calculated. Under the conditions of significance level α=0.05 and test power 1–β=0.80 (β = 0.20), the required sample size for each bacterial genera was calculated using PASS statistical software according to the two-tailed test of the difference between the means of two independent samples.
[0124]
[0125] The results showed that the required total sample sizes for different genera were approximately: 21 children for Cardiobacterium hominis, 148 children for Veillonella sp. HMT 780, 121 children for Rothia aeria, and 26 children for Streptococcus mutans. Veillonella sp. HMT 780 had the largest sample size, with 148 children. To ensure a diagnostic power of at least 80% for all major candidate genera, this invention used this maximum value as the design basis, requiring 74 children each for the case group (carious group, CA) and the control group (non-carious group, CF), for a total sample size of 148 children (74 CF + 74 CA).
[0126] Sample processing: Subjects fasted and abstained from water for 2 hours prior to sample collection; dental plaque was collected from all surfaces of the children's primary teeth using a dental excavator by an experienced dentist; samples were transferred to labeled sterile centrifuge tubes containing 1 ml of sterile phosphate-buffered saline (PBS). Samples were transported to a laboratory freezer within 2 hours of sampling, placed on dry ice during transport, and then frozen at -80°C after transport.
[0127] (b) DNA extraction: Bacterial genomic DNA was extracted from dental plaque samples using a Tianlong GeneRotex96 according to the manufacturer’s instructions.
[0128] (c) ddPCR detection: Establish a ddPCR validation system and use the multiplex quantitative droplet digital PCR method to detect each sample and obtain the absolute quantity of the corresponding microbial markers.
[0129] The primer and probe sequences of the biomarker system + internal control 16S system obtained in Example 1 (see Table 14) and the reaction system in Table 15 were used for detection.
[0130] The relative abundance of the four microbial markers in each sample was obtained by dividing the absolute quantitative values of the four microbial markers (Cardiobacterium hominis, Rothia aeria, Veillonellasp_HMT_780, and Streptococcus mutans) by the absolute quantitative values of the total bacteria obtained by the universal internal control 16S primer probe in ddPCR.
[0131] (d) Standardize the relative abundance data of all samples obtained in step (c): Before training, standardize the data to have a mean of zero and a standard deviation of one.
[0132] (e) Model training and validation: Model selection: Use the Random Forest Classifier. Hyperparameter optimization: Set n_estimators to 50 and max_depth to 5. Cross-validation: Use 10-fold cross-validation to evaluate model performance. Model training: Train the Random Forest Classifier using standardized data.
[0133] (f) Model evaluation: The model performance was evaluated using a test set to obtain the probability that each sample belonged to the caries group (CA group); the true positive rate (TPR) and false positive rate (FPR) were calculated; and the ROC curve was plotted and the AUC value was calculated. Specifically, the steps for plotting the ROC curve are as follows: (f1) Data preparation: The relative abundance of the four microbial markers of all samples were compiled into a data table, which included the sample number, relative abundance value and the group to which the sample belonged (non-caries CF group or caries CA group).
[0134] (f2) Data standardization: Use Python's StandardScaler to standardize the data.
[0135] (f3) Training the model: The model was trained using RandomForestClassifier. The model performance was evaluated using 10-fold cross-validation.
[0136] (f4) Predicted probability: Use the trained model to predict the probability that each sample belongs to the caries group (CA group).
[0137] (f5) Calculate ROC curve data: Use the roc_curve function to calculate the true positive rate (TPR) and false positive rate (FPR). Use the auc function to calculate the AUC value.
[0138] (f6) Plotting ROC curves: Use Matplotlib to plot ROC curves.
[0139] (2) Sample grouping In this embodiment, to evaluate the impact of different primer-probe combinations on the detection effect, based on the above experimental step (1), some or all of the biomarkers listed in Table 14 were removed to study their impact on detection sensitivity, specificity, and anti-interference ability. Specific groupings and corresponding test results are shown in Table 23 and... Figure 20 .
[0140] Table 23 ROC area and cutoff value of random forest models with different combinations of biomarkers
[0141] (3) Results Analysis In this embodiment of the invention, there is a significant synergistic relationship between the multiple upstream and downstream primers and probes of the internal reference 16S. Specifically, the absence of any one of the internal reference primers or probes will negatively affect the final detection results.
[0142] Based on this, further verification through this embodiment shows, as can be seen from Table 23, that the four biomarkers also exhibit a significant synergistic effect. When all four biomarkers are present simultaneously, the detection performance reaches its optimal level, reflected in the highest area under the ROC curve and a moderate cut-off value; conversely, the absence of any one biomarker leads to a decrease in detection performance, specifically manifested in a reduction in the area under the ROC curve.
[0143] It is understandable that ROC curves (Receiver Operating Characteristic curves) and AUC values (area under the curve) are important tools for evaluating the performance of classification models. Cut-off values (thresholds) are used to determine the decision boundary of the classification model. These metrics are closely related to the sensitivity, specificity, repeatability, and robustness to interference of detection. According to common knowledge in the field, a high AUC value generally indicates that the model performs better in terms of sensitivity and specificity, and has higher repeatability and robustness to interference. A moderate cut-off value indicates that the model has achieved a good balance between sensitivity and specificity, because a high cut-off value can improve specificity but may reduce sensitivity.
[0144] Example 4: The five-phase primer-probe system (four biomarkers + 16s internal control) was further optimized and finally determined to be a four-phase primer-probe system (three biomarkers + 16s internal control). (1) Feature importance score The relative abundance data of the four microbial biomarkers obtained in Example 3 were analyzed using a random forest classifier with 10-fold cross-validation. The model was trained using the RandomForestClassifier from the sklearn library in Python 3.9. Before training, the data were standardized to a mean of zero and a standard deviation of one. Considering the small sample size, the hyperparameters of the random forest were optimized by setting n_estimators to 50 and max_depth to 5. The feature importance scores of each microbial biomarker were averaged to obtain the feature importance score for each fold, as shown in Table 24. Table 24 Characteristic Importance Scores of Microbial Biomarkers
[0145] (2) Statistical analysis (T test) The relative abundance data of the four microbial biomarkers obtained in Example 3 were analyzed using IBM SPSS Statistics software (v23.0) via a t-test. The mean and p-value of each biomarker were calculated in both the caries-free (CF) and caries (CA) groups. The results are shown in Table 25. Figure 21 If the p-value is ≤0.05, the difference is considered statistically significant; if the p-value is >0.05, the difference is considered not significant.
[0146] Table 25 Results of the T-test for the relative abundance of four microbial biomarkers
[0147] The results showed significant differences among Streptococcus mutans, Rothia aeria, and Cardiobacterium hominis. Veillonella sp_HMT_780 had a p-value > 0.05, showed no significant difference between the CF and CA groups, and had the lowest random forest feature importance score, thus it was excluded from the final microbial biomarkers. Ultimately, three combined microbial biomarkers—Streptococcus mutans, Rothia aeria, and Cardiobacterium hominis—were identified.
[0148] Specifically, the primer-probe composition for quantitative detection of caries microbial markers in young children has a fluorescent group labeled at the 5' end of the probe and a quencher group labeled at the 3' end of the probe; wherein the fluorescent group is selected from FAM, VIC, Cy5, ATTO series, ROX, or HEX; and the quencher group is selected from MGB, BHQ series, TAMRA, or Eclipse. Preferably, the fluorescent group is selected from FAM, HEX, or ATTO 425.
[0149] For example, the probe of Cardiobacterium hominis is labeled with a 5' ATTO 425 fluorescent group at the 5' end and a 3' MGB quencher group at the 3' end.
[0150] For example, the probe of the space Rothia aeria is labeled with a 5' ROX fluorescent group at the 5' end and a 3' BHQ2 quencher group at the 3' end.
[0151] For example, the probe of Streptococcus mutans is labeled with a 5'6-FAM fluorescent group at the 5' end and a 3'MGB quencher group at the 3' end.
[0152] Preferably, in the primer-probe composition, the molar ratio between the upstream primers for detecting *Humanoidobacterium stoloniferum*, *R. stoloniferum*, and *Streptococcus mutans* is 1:1:1, and the molar ratio of the upstream primer of the internal control to the upstream primer of each of the three bacteria is 2:1. Preferably, in the primer-probe composition, the molar ratio between the downstream primers for detecting *Humanoidobacterium stoloniferum*, *R. stoloniferum*, and *Streptococcus mutans* is 1:1:1, and the molar ratio of each downstream primer of the internal control to the downstream primer of each of the three bacteria is 1:1.
[0153] Preferably, in the primer-probe composition, the molar ratio between the probes for detecting *Humanoidobacterium tumefaciens*, *Roseobacter stenosis*, and *Streptococcus mutans* is 1:1:1, and the molar ratio of each probe of the internal control to the probe of each bacterium is 1:1.
[0154] Preferably, in the primer-probe composition, the molar number of the upstream or downstream primer for each of the three bacteria is twice that of the probe for each bacteria.
[0155] Preferably, the molar ratio of the upstream primer, downstream primer, and probe of the internal control is 4:6:3.
[0156] Thirdly, the present invention provides a kit for quantitative detection of caries microbial biomarkers in young children, the kit comprising a ddPCR system, which includes a primer-probe mixture, a ddPCR premix, a DNA template, and sterile water; the primer-probe mixture includes the internal reference sequence described in the first aspect or the primer-probe composition described in the second aspect.
[0157] It is understandable that ddPCR stands for Droplet Digital Polymerase Chain Reaction.
[0158] In one embodiment, the ddPCR system, in 25 μL, includes 12.5 μL of PCR premix, 0.125 μL each of the upstream and downstream primers for each of *Humanoidea*, *Roseola stolonifer*, and *Streptococcus mutans* at a final concentration of 500 nM, 0.625 μL of the probe at a final concentration of 250 nM, 0.25 μL of the upstream primer for the internal control at a final concentration of 1000 nM, and three downstream primers for the internal control. Each downstream primer (U16S_1106R1-1, U16S_1106R1-2, U16S_1106R1-3) was added in a final concentration of 500 nM. 0.625 μL of each of the three internal control probes (U16S_1064rP1-1, U16S_1064rP1-2, U16S_1064rP1-3) was added in a final concentration of 250 nM. 2.5 μL of DNA template was added, and the total volume was brought up to 25 μL with sterile water (double-distilled water, ddH2O).
[0159] It can be understood that the final concentration (nM) = [stock solution concentration (μM) × added volume (μL)] / total reaction volume (μL) × 1000. For example, referring to Table 15, taking a 100 μM upstream primer as an example: final concentration = (100 μM × 0.125 μL) / 25 μL = 0.5 μM = 500 nM.
[0160] Fourthly, the present invention provides a method for quantitatively detecting caries microbial markers in young children, comprising: (1) Extract genomic DNA from dental plaque samples; (2) Using the genomic DNA as a template, droplets are generated in a droplet generator through a ddPCR system including the primer and probe composition described in the second aspect or the kit described in the third aspect, and PCR amplification is performed to obtain the amplification product; (3) The amplification product is subjected to fluorescence detection based on the droplet reader, and the copy number of the three bacteria, Cardiobacterium hominis, Rothia aeria, and Streptococcus mutans, and the copy number of the internal reference are determined respectively in the amplification product. (4) Divide the copy number of the three bacteria by the copy number of the internal reference to calculate the relative DNA content of each of the three bacteria in the dental plaque sample.
[0161] Preferably, the PCR amplification program is as follows: 95℃ pre-denaturation for 5 min, 94℃ denaturation for 20 s, 60℃ annealing and extension for 1 min, and repeating the denaturation-annealing-extension cycle for a total of 45 cycles.
[0162] Specifically, in step (4), the relative DNA content of the three bacteria is used in the ROC model to determine the risk of dental caries. The cut-off value of the ROC model constructed by Streptococcus mutans, Human cardiomycosis, and Rossella squarrosa is 0.453. If only one of the markers is used for detection, the cut-off value of the Streptococcus mutans ROC model is 0.476, the cut-off value of the Human cardiomycosis ROC model is 0.387, and the cut-off value of the Rossella squarrosa ROC model is 0.378.
[0163] Specifically, the detection limit for the three bacteria is as low as 0.01 copies / μL.
[0164] Specifically, the standard deviation of the relative DNA content of each type of bacteria is ≤5%.
[0165] Fifthly, the present invention provides the application of the internal reference sequence described in the first aspect or the primer-probe composition described in the second aspect in the preparation of detection products for rapid quantitative detection of dental caries microbial markers in young children.
[0166] Comparative Example Currently, dental caries detection kits on the market can be roughly divided into three categories: (1) Chairside rapid detection products based on immunochromatography or ATP bioluminescence, such as GC's Saliva-Check Mutans rapid detection kit for Streptococcus mutans. These products are easy to operate and produce results quickly, but they are mostly qualitative or coarse semi-quantitative, and are difficult to reflect the structure and combination patterns of caries-related microbiota.
[0167] (2) Caries activity or bacterial counting kits based on selective culture media, such as Ivoclar Vivadent's CRTbacteria Caries Risk Test and the Dentocult SMStrip mutans kit for determining the load of Streptococcus mutans in saliva. These products require culture, have long cycles, and only cover culturable species, making the results susceptible to the influence of culture conditions.
[0168] (3) Caries-associated bacteria detection kits based on real-time quantitative PCR are mostly designed with a single target or 2–3 targets per tube, such as Bioneer's AccuPower® Streptococcus mutans Real-Time PCR Kit, Skygen's Rothia dentocariosa Probe qPCR Kit, and the dental caries actinomycete probe qPCR kit. The quantitative results of these products rely on a standard curve, and problems such as increased Ct values, decreased amplification efficiency, and false negatives in low-copy samples are prone to occur in the complex oral matrix. Moreover, they often lack unified internal controls and multi-species combination biomarkers for systematic screening.
[0169] Compared to existing commercial solutions, the multiplex ddPCR caries microbial biomarker detection system proposed in this invention has significant improvements in the following aspects: First, this invention uses large-scale 16S rRNA sequencing data and machine learning methods to screen for microbial combinations highly associated with caries in young children, such as Streptococcus mutans, Cardiobacterium hominis, and Rothia aeria. It employs a quadruple primer-probe system combining three microbial biomarkers with an internal 16S reference, obtaining absolute quantitative information for multiple targets in a single reaction, thus more comprehensively reflecting the caries-related microecological characteristics. Second, this invention utilizes ddPCR technology for counting, achieving high sensitivity and precision absolute quantification without relying on a standard curve, effectively reducing the impact of the complex oral matrix on amplification efficiency and significantly lowering the risk of false negatives in low-copy samples. Third, this invention designs a dedicated universal 16S internal reference primer-probe combination to normalize the differences in total bacterial count between different samples, improving the comparability and clinical interpretability of the results. In summary, this invention overcomes the shortcomings of existing dental caries detection kits, such as single target, reliance on standard curves, susceptibility to matrix interference, and lack of clinically usable multiple molecular diagnostic systems, and is more suitable for the needs of early screening and risk stratification assessment of dental caries in young children.
[0170] Compared to 16S sequencing methods, 16S high-throughput sequencing is suitable for community structure and diversity analysis, but it outputs relative abundance and does not always provide accurate values for low-abundance, highly targeted pathogen-associated bacteria. Furthermore, the sequencing cost and timeframe are not suitable for routine testing in outpatient clinics or large-scale epidemiological studies. Existing studies on oral microbiota in children with early dental caries (Tanner, 2011; Teng F et al., 2015) describe "which bacteria are more / less abundant" based on sequencing results, without providing copy numbers that can be directly used clinically. The multiplex ddPCR system of this invention uses universal 16S as an internal control, converting the absolute amounts of the three marker bacteria into a relative index of "target bacteria / total bacterial count." This achieves three layers of information—presence / absence + quantification + normalization with the microbiota—on the same oral sample, a feat difficult to accomplish in a single step with existing kits and sequencing methods.
[0171] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A 16S internal reference sequence for quantitative detection of caries microbial biomarkers in young children, characterized in that, The internal reference sequence includes the upstream primer sequence, the downstream primer sequence, and the probe sequence; The upstream primer sequence is SEQ ID NO.1; The downstream primer sequences are SEQ ID NO.2, SEQ ID NO.3, and SEQ ID NO.4; The probe sequences are SEQ ID NO.5, SEQ ID NO.6 and SEQ ID NO.
7.
2. A primer-probe composition for quantitative detection of caries microbial markers in young children, characterized in that, The primer-probe composition includes the 16S internal reference sequence as described in claim 1, as well as upstream primers, downstream primers, and probes for detecting three bacteria: Cardiobacterium hominis, Rothia aeria, and Streptococcus mutans. The upstream and downstream primer sequences for detecting Human poxvirus are shown in SEQ ID NO.8 and SEQ ID NO.9, respectively, and the probe sequence is shown in SEQ ID NO.
10. The upstream and downstream primer sequences for detecting *Rhodesia spp.* are shown in SEQ ID NO.11 and SEQ ID NO.12, respectively, and the probe sequence is shown in SEQ ID NO.
13. The upstream and downstream primer sequences for detecting Streptococcus mutans are shown in SEQ ID NO.14 and SEQ ID NO.15, respectively, and the probe sequence is shown in SEQ ID NO.
16.
3. The primer-probe composition according to claim 2, characterized in that, In the primer-probe composition, the molar ratio between the upstream primers for detecting *Humanoidobacterium tumefaciens*, *Roseobacter stenosis*, and *Streptococcus mutans* is 1:1:1, and the molar ratio between the upstream primer of the internal control and the upstream primer of each of the three bacteria is 2:
1.
4. The primer-probe composition according to claim 2, characterized in that, In the primer-probe composition, the molar ratio between the downstream primers for detecting *Humanoidobacterium tumefaciens*, *Roseobacter stenosis*, and *Streptococcus mutans* is 1:1:1, and the molar ratio of each downstream primer of the internal control to the downstream primer of each of the three bacteria is 1:
1.
5. The primer-probe composition according to claim 2, characterized in that, In the primer-probe composition, the molar ratio between the probes for detecting *Humanoidobacterium tumefaciens*, *Roseobacter stenosis*, and *Streptococcus mutans* is 1:1:1, and the molar ratio of each probe of the internal control to the probe of each bacterium is 1:
1.
6. The primer-probe composition according to claim 2, characterized in that, In the primer-probe composition, the molar number of the upstream or downstream primer for each of the three bacteria is twice that of the probe for each bacteria; and / or, The molar ratio of the upstream primer, downstream primer, and probe of the internal control is 4:6:
3.
7. A kit for quantitative detection of caries microbial markers in young children, characterized in that, The kit includes a ddPCR system comprising a primer-probe mixture, a ddPCR premix, a DNA template, and sterile water; the primer-probe mixture comprises the internal reference sequence of claim 1 or the primer-probe composition of any one of claims 2-6.
8. A method for quantitatively detecting caries microbial markers in young children, characterized in that, include: (1) Extract genomic DNA from dental plaque samples; (2) Using the genomic DNA as a template, droplets are generated in a droplet generator using a ddPCR system comprising the primer and probe composition of any one of claims 2-6 or the kit of claim 7, and PCR amplification is performed to obtain the amplification product; (3) The amplification product is subjected to fluorescence detection based on the droplet reader, and the copy number of the three bacteria, Cardiobacterium hominis, Rothia aeria, and Streptococcus mutans, and the copy number of the internal reference are determined respectively in the amplification product. (4) Divide the copy number of the three bacteria by the copy number of the internal reference to calculate the relative DNA content of each of the three bacteria in the dental plaque sample.
9. According to the method of claim 8, in step (4), the relative DNA content of the three bacteria can be used to determine the risk of dental caries by introducing it into the ROC model; wherein, The cut-off value of the ROC model constructed using *Streptococcus mutans*, *Humanoidomyces cerevisiae*, and *R. cerevisiae* was 0.453; if only one of the markers was used for detection, the cut-off value of the *Streptococcus mutans* ROC model was 0.476, the *Humanoidomyces cerevisiae* ROC model was 0.387, and the *R. cerevisiae* ROC model was 0.378; and / or, The detection limits for the three bacteria are as low as 0.01 copies / μL; and / or, The standard deviation of the relative DNA content of each bacterium is ≤5%.
10. The use of the internal reference sequence of claim 1 or the primer-probe composition of any one of claims 2-6 in the preparation of a detection product for rapid quantitative detection of dental caries microbial markers in young children.