A sequencing data calibration system and method based on microbial amplicon sequencing

By constructing a sequencing data correction system for droplet-based digital PCR, the quantitative limitations and biases of microbial amplicon sequencing technology were resolved, achieving absolute quantification and cross-platform data comparability, and improving the accuracy and reliability of the test results.

CN122344622APending Publication Date: 2026-07-07NATIONAL INSTITUTE OF METROLOGY CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NATIONAL INSTITUTE OF METROLOGY CHINA
Filing Date
2026-03-11
Publication Date
2026-07-07

AI Technical Summary

Technical Problem

Existing microbial amplicon sequencing technology has limitations in quantification, difficulty in correcting technical biases, a lack of standard materials, and insufficient reliability of measurement values. This results in the inability to convert relative quantitative results into absolute quantitative results, and a lack of comparability and credibility of cross-platform and cross-batch data.

Method used

A sequencing data correction system based on droplet digital PCR was constructed, including a standard substance preparation module, a ddPCR value determination module, and a data correction module. Through primer and probe design, reaction condition optimization, and rigorous verification, multi-target correction substances were prepared, and a full-process deviation correction mechanism and a traceability system for measurement values ​​were established.

Benefits of technology

It achieves absolute quantification of microbial amplicon sequencing data, eliminates systematic errors, improves the accuracy and cross-laboratory comparability of test results, and meets the stringent requirements of fields such as medical diagnosis and environmental monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The application discloses a sequencing data calibration system and method based on microbial amplicon sequencing, and belongs to the field of microbial detection and bioinformatics analysis. The application takes ddPCR technology as the core, first screens common human intestinal microorganisms, designs and obtains 14 reconstructed sequences containing bacterial 16S and fungal 18S / ITS, prepares high-purity DNA standard substances through chemical synthesis and molecular cloning, completes standardization and uniformity verification, then designs high-specificity primers and probes, optimizes ddPCR reaction conditions, establishes a calibration method, completes uncertainty evaluation through joint calibration of 9 laboratories, and prepares multi-target amplicon abundance standard substances. The application realizes the transformation of amplicon sequencing data from relative quantification to absolute quantification, corrects the system error of the whole sequencing process, improves the comparability and reliability of cross-platform and cross-batch data, and provides technical support for the standardization of microbiome research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of microbial detection and bioinformatics analysis technology, specifically relating to a sequencing data correction system and method based on microbial amplicon sequencing. Background Technology

[0002] Microbial amplicon sequencing, as a core technology for elucidating the composition, structure, and diversity of microbial communities, has been widely applied in various research and application fields such as medical testing, environmental ecology, agricultural breeding, and food microbial safety due to its advantages of low detection cost, high sequencing efficiency, and strong host contamination resistance. It has become a mainstream technique in microbiome research (Caporaso JG, Kuczynski J, Stombaugh J, et al. QIIME allows analysis of high-throughput community sequencing data [J]. Nat Methods. 2010,7(5):335-336.). This technology targets microbial-specific genes (such as bacterial 16S rRNA genes, fungal 18S rRNA genes, and ITS regions) for amplification. Through PCR amplification, high-throughput sequencing, and subsequent bioinformatics analysis, it achieves qualitative and relative quantitative analysis of microbial species in samples. This technology has advantages such as high throughput, low cost, and small sample requirements, and can quickly achieve classification and identification and relative abundance assessment at the genus and species level of microorganisms, providing important technical support for research on diseases related to dysbiosis and environmental microbial monitoring.

[0003] However, existing amplicon sequencing technology has significant technical bottlenecks, which seriously restrict its quantitative accuracy and data comparability: (1) Prominent quantitative limitations: Amplicon sequencing can only provide relative abundance information of microbial communities. Affected by the "zero-sum game" effect, changes in the proportion of a certain group will lead to passive shifts in the relative readings of other groups, making it impossible to distinguish between the real changes in the absolute number of microorganisms and the structural fluctuations in the relative proportions. This defect is particularly critical in scenarios such as clinical diagnostic threshold determination and probiotic intervention effect evaluation, and there is an urgent need to transform relative quantification into absolute quantification. (2) Difficulty in correcting technical deviations: There are various systematic errors in the sequencing process, such as GC bias and primer binding bias in the PCR amplification process, read length limitations and base recognition differences in different sequencing platforms (Illumina, third-generation sequencers, etc.), and fluctuations in sample processing and nucleic acid extraction efficiency, resulting in poor consistency of detection results for the same sample in different laboratories and different platforms, and a lack of data comparability. (3) Scarcity and single function of standard materials: At present, multi-target, multi-abundance, cross-platform quantitative standard materials for microbial amplicon sequencing are extremely scarce internationally. Existing standard materials mostly cover only a single bacterial genus or a limited number of targets, lack simulation of complex community composition, and have not established a complete absolute quantitative traceability system, making it impossible to simultaneously achieve multiple functions such as instrument calibration, method validation, and process quality control. (4) Insufficient reliability of measurement values: The calibration of existing sequencing data has not formed a standardized uncertainty assessment system, and has not fully considered the uniformity and stability of calibration materials, as well as the influencing factors such as inter-laboratory differences and instrument system errors in the calibration process, resulting in the calibration results lacking metrological credibility and failing to meet the strict requirements for data accuracy in fields such as clinical diagnosis and public health monitoring.

[0004] Therefore, it is urgent to construct a complete sequencing data correction system and method based on microbial amplicon sequencing. By designing multi-target correction substances that simulate the characteristics of real communities, a deviation correction mechanism and quantitative traceability system covering the entire process can be established to achieve systematic correction of amplicon sequencing data, convert relative quantitative results into absolute quantitative data, improve the comparability and reliability of cross-platform and cross-batch data, and provide core technical support for the standardization and precision development of microbiome research. Summary of the Invention

[0005] To address the problems existing in current technologies, this invention establishes a complete sequencing data correction system and method based on droplet digital PCR (ddPCR) for microbial amplicon sequencing. This system covers the entire process from primer and probe design and reaction condition optimization to specificity verification. Through systematic methodological comparison and condition optimization, it ensures that optimal quantitative analysis results can be obtained for each gene sequence. The entire study adopts a phased, step-by-step optimization strategy. First, primer and probe design at the bioinformatics level is completed, followed by the establishment and optimization of experimental methods, and finally, the reliability of the method is confirmed through rigorous validation experiments.

[0006] On one hand, the present invention provides a sequencing data correction system for microbial amplicon sequencing, including a standard substance preparation module, a ddPCR value determination module, and a data correction module; the standard substance preparation module is used to prepare microbial multi-target amplicon abundance standard substances; the ddPCR value determination module is used to perform absolute quantification of the standard substances; and the data correction module is used to correct the amplicon sequencing data based on the values ​​of the standard substances.

[0007] Specifically, the standard substance preparation module includes a sequence design unit, a synthesis and cloning unit, and a dispensing and verification unit; the sequence design unit screens human gut microorganisms and obtains reconstructed sequences through sequence alignment; the synthesis and cloning unit obtains high-purity DNA fragments through chemical synthesis and molecular cloning; and the dispensing and verification unit completes the standardized dispensing and uniformity detection of the standard substance.

[0008] Specifically, the GC content of the reconstructed sequence is less than 65%, preferably less than 60%, and the reconstructed sequence includes Bacillus, Enterococcus, Salmonella, Fusobacterium, Broutella, Klebsiella, Megamonas, Penicillium, Candida, Escherichia, Saccharomyces, Bifidobacterium, Cladosporium, and Aspergillus.

[0009] Specifically, the reconstructed sequence covers the full 16s length, the 18s region, and the ITS region. Preferably, the reconstructed sequence comprises 14 DNA fragments, the nucleotide sequences of which are shown in SEQ ID NO.1-SEQ ID NO.14.

[0010] Specifically, the synthetic cloning unit uses a high-fidelity enzyme, and the reaction system includes 5× high-fidelity enzyme reaction buffer, dNTPs, primers, template and enzyme. The annealing temperature is 50℃-65℃, preferably 55℃-60℃. Preferably, the primers include universal primers 27F, 1492R, 515F, 1119R, ITS1F, ITS3F and ITS4R, and some fragments use specific primers Bif164-F / Pbi R2, whose nucleotide sequences are shown in SEQ ID NO.15-SEQ ID NO.23.

[0011] Specifically, the ddPCR determination module includes a primer and probe design unit, a reaction condition optimization unit, and a combined determination unit; the primer and probe design unit designs primers with a length of 18-24 bp, a Tm value of 58-60℃, and probes with a length of 13-25 bp, with the first base at the 5' end being non-G.

[0012] Specifically, the joint value setting unit is jointly customized by no fewer than nine laboratories with testing capabilities.

[0013] On the other hand, the present invention provides a sequencing data correction method based on microbial amplicon sequencing, comprising the following steps: S1, screening human gut microbes, designing and obtaining 14 reconstructed sequences; S2, designing primers and probes and optimizing ddPCR reaction conditions, and establishing a quantification method; S3, preparing DNA standard substances for the reconstructed sequences and completing aliquoting and homogeneity verification; S4, determining the values ​​of the standard substances through joint quantification by 9 laboratories and assessing the uncertainty; S5, performing absolute quantitative correction of the microbial amplicon sequencing data based on the values ​​of the standard substances.

[0014] Specifically, the method for obtaining the reconstructed sequence is as follows: if the microbial sequence contains a complementary sequence to the universal primers, it is directly synthesized; if it is missing, bases are added; if it contains both 18S and ITS sequences, sequence splicing is performed. The universal primers include 27F / 1492R for bacterial 16S, 515F / 1119R for fungal 18S, and ITS1F / ITS3F / ITS4R for ITS. Preferably, after the joint determination, the data are statistically analyzed using the Cochrane test, Grubbs test, and Shapiro-Wilke test. The arithmetic mean of the results from nine laboratories is used as the standard value, and the coverage factor k=2 is taken. The expansion uncertainty is evaluated at a 95% confidence level. Based on this standard value, the relative quantitative data of amplicon sequencing is converted into absolute quantitative data, and the sequencing data correction is completed.

[0015] On the other hand, the present invention utilizes the sequencing data correction system or the method described herein for applications in correcting amplicon sequencing technology deviations, in the calibration of environmental microbial detection instruments, or in the performance evaluation of sequencing kits.

[0016] Compared with existing technologies, the beneficial technical effects of this invention are as follows: This invention constructs a complete microbial amplicon sequencing data correction system. The prepared multi-target standard materials cover key gene sequences of bacteria and fungi, simulating real community characteristics and solving the problem of single standard materials in existing systems. The ddPCR-based quantification method can achieve absolute quantification, traceable to natural unit 1, breaking through the limitations of relative quantification in amplicon sequencing and eliminating the influence of the "zero-sum game" effect. This invention effectively corrects systematic errors such as PCR bias and sequencing platform differences. Verified jointly by multiple laboratories, the standard materials exhibit good homogeneity and high stability, and the uncertainty assessment system is comprehensive, significantly improving the accuracy and cross-laboratory comparability of detection results. It can meet the stringent requirements for sequencing data in fields such as medical diagnostics and environmental monitoring, promoting the standardization and precision development of microbiome research. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of 16S sequence universal primer synthesis and amplification.

[0018] Figure 2 This is a diagram of 16S universal primer addition sequence synthesis.

[0019] Figure 3 This is a schematic diagram of the 18S and ITS splicing sequence.

[0020] Figure 4 This is a flowchart of the preparation process for standard substance candidates.

[0021] Figure 5 This is a graph showing the plasmid identification results.

[0022] Figure 6 This is a graph showing the results of PCR product purification and identification.

[0023] Figure 7 This is a microarray electrophoresis result of the sample shown in SEQ ID NO.1.

[0024] Figure 8 This is a Sanger sequencing alignment result of the sample shown in SEQ ID NO.13.

[0025] Figure 9 These are droplet diagrams for ddPCR under different primer and probe concentration systems. F05-G05 represent primer and probe concentrations of 700nM-125nM, H05-A06 of 700nM-250nM, C06-D06 of 700nM-500nM, F06-G06 of 900nM-125nM, A07-B07 of 900nM-250nM, and D07-E07 of 900nM-500nM.

[0026] Figure 10These are the results of digital PCR amplification at different annealing temperatures, where A01-F01 represent different channels, B01-C01 represent annealing temperatures of 55℃, and D01-F01 represent annealing temperatures of 60℃.

[0027] Figure 11 These are the results of the crossover experiment for gene sequence 1. A04, B04, C04, E04, and F04 represent channels 1-5. Channel 1 is the blank control, channels 2 and 3 are the test samples, and channels 4 and 5 are the mixture of other samples besides the test samples. Detailed Implementation

[0028] The present invention will be further described below with reference to specific embodiments, and the advantages and features of the present invention will become clearer as a result of the description. However, these embodiments are merely illustrative and do not constitute any limitation on the scope of protection defined by the claims of the present invention.

[0029] It should be understood that the terminology used in this invention is merely for describing particular embodiments and is not intended to limit the invention. Furthermore, with respect to numerical ranges in this invention, it should be understood that the upper and lower limits of the range and each intermediate value between them are specifically disclosed. Any stated value or intermediate value within a stated range, as well as each smaller range between any other stated value or intermediate value within said range, are also included in this invention. The upper and lower limits of these smaller ranges may be independently included or excluded from the range.

[0030] Unless otherwise stated, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. While only preferred methods and materials have been described herein, any methods and materials similar or equivalent to those described herein may be used in the implementation or testing of this invention. All references to this specification are incorporated by way of citation to disclose and describe methods and / or materials associated with those references. In the event of any conflict with any incorporated reference, the content of this specification shall prevail.

[0031] Example 1: Design of DNA Standard Materials

[0032] 1. Strains selection

[0033] Microorganisms were screened based on the principle of being common and not homologous to deep-sea organisms, with priority given to screening from the human gut microbiome database hGMB and NCBI. Gut microorganisms can be divided into two main categories: fungi and bacteria. In the study of fungi and bacteria, special sequences are often studied, namely 16S for bacteria and 18S and ITS for fungi. Therefore, our sequence selection also focused on these special gene sequences. The diversity of human fungal communities is relatively low. At the genus level, Candida and yeast are the main species (NASH AK, AUCHTUNG TA, WONG MC, et al. The gut mycobiome of the HumanMicrobiome Project healthy cohort [J]. Microbiome, 2017, 5(1): 153.). Bacteria are mainly Bifidobacterium. Therefore, in the sequence screening process, common phyla and genera with a large proportion in the gut microbiome were selected as strain sequences. The classification information of the strains can be obtained from hGMB, and the sequence information can be downloaded from NCBI.

[0034] 2. Reference sequence and universal primer selection

[0035] Literature review shows that amplicon sequences are generally the most studied for microorganisms. For bacteria, the main focus is on the 16S sequence, while for fungi, the focus is on some variable regions on the 18S and ITS sequences. By detecting these variable regions, the types of microorganisms can be quickly and accurately identified. Therefore, the selection of standard material sequences will cover these sequences. Commonly used universal primers for amplifying the full-length 16S sequence in bacteria are 27F (SEQ ID NO. 15: AGAGTTTGATCMTGGCTCAG) and 1492R (SEQ ID NO. 16: GGTTACCTTGTTACGACT). For 18S, 515F (SEQ ID NO. 17: GTGCCAGCMGCCGCGGTAA) and 1119R (SEQ ID NO. 18: GGTGCCCTTCCGTCA) can amplify the V4-V5 variable region of the 18S sequence. Common ITS sequences are ITS1F (SEQ ID NO. 19: CTTGGTCATTTAGAGGAAGTAA), ITS3F (SEQ ID NO. 20: GCATCGATGAAGAACGCAGC), and ITS4R (SEQ ID NO. 21: TCCTCCGCTTATTGATATGC). In addition to the universal primers for bacteria, we designed non-universal primers accounting for 10%. Table 1 lists the universal primers for sequence 12 (Bifidobacterium) (SEQ ID NO. 15: AGAGTTTGATCMTGGCTCAG). SEQ ID NO. 22: Bif164-F GGGTGGTAATGCCGGATG; SEQ ID NO. 23: Pbi R2: GACCATGCACCACCTGTGAA) differs from other genera in that it can only amplify the sequences of Bifidobacterium. After the universal primers are determined, the sequence information of the selected strains is compared with the sequences complementary to the universal primers using SnapGene. If the sequence downloaded from NCBI itself contains a sequence complementary to the universal primers, it is not processed. Figure 1 Sequence synthesis can be performed directly; if upstream primers or complementary sequences to downstream primers are lacking, bases of universal primer length, such as... Figure 2 This is done to obtain the final reconstructed sequence; if the same strain has both 18S and ITS sequences, and its universal primers are inside the sequence, then the 18S and ITS sequences are spliced ​​together, such as... Figure 3 As shown, the final reconstructed sequences are obtained (as shown in SEQ ID NO.1-SEQ ID NO.14). The reconstructed sequence information is shown in Table 1, which is the sequence information obtained by adding primers or splicing the corresponding NCBI accession number sequence.

[0036] Table 1. Reconstructed Sequence Information Table

[0037]

[0038] Example 2: Preparation and Dispensing of Standard Substances

[0039] The preparation of this reference material strictly follows the pre-designed sequence information. The core preparation strategy is as follows: First, high-precision DNA samples of 14 sequences are obtained and verified through chemical synthesis and molecular cloning techniques. Then, the purified fragments are mixed according to the concentration of the synthesized samples and the set fractions (the final added volumes to each aliquot are 0.004 μL, 0.055 μL, 0.172 μL, 0.199 μL, 2.293 μL, 0.318 μL, 0.059 μL, 0.747 μL, 1.001 μL, 0.974 μL, 1.582 μL, 14.039 μL, 7.032 μL, and 3.576 μL, respectively) to prepare the final reference material.

[0040] 1. Preparation of standard substances

[0041] Fourteen sequences were synthesized using sequence fragment synthesis. Strains containing these sequences were constructed using vectors and then subjected to Sanger sequencing. Strains with correct sequencing results were streaked twice to form single clones, which were then preserved. Simultaneously, the single clones were enriched by shaking for plasmid extraction. The target fragments were then amplified using the plasmids as templates. The products were subjected to agarose gel electrophoresis (1.5%), and the target fragments were recovered by gel extraction and purified. Finally, the purified products were identified using Sanger sequencing. Fourteen DNA samples with accurate sequence information were obtained. The flowchart for the preparation of standard material candidates is shown below. Figure 4 As shown. The specific steps are as follows:

[0042] (1) Synthesis of gene fragments: Shenzhen BGI Genomics Co., Ltd. was commissioned to complete the chemical synthesis of all target DNA fragments based on the designed sequence information. The synthesis service covers the entire process from sequence design, oligonucleotide synthesis, splicing to complete fragments.

[0043] (2) Cloning and identification of the fragment: In order to obtain a template that can be stably amplified and has a precise sequence, the synthesized fragment was cloned and identified:

[0044] i) Cloning vector construction: Using homologous recombination technology, each target fragment was ligated into a cloning vector and transformed into competent E. coli cells; the ligation reaction system of the target fragment and the vector is shown in Table 2:

[0045] Table 2. Reaction System

[0046]

[0047] Transformation method: Take 1-3 μL of plasmid with a concentration of about 100 ng / μL and add it to about 100 μL of competent cells. Gently shake and rotate to mix. Place on ice for 3 minutes, then incubate in a 42°C water bath for 90 seconds without shaking. Finally, place in an ice bath for about 3 minutes.

[0048] ii) Screening for positive clones: The transformed bacterial culture was plated on LB agar plates containing the appropriate antibiotics. The plates were incubated at 37°C for 15 minutes, then inverted and incubated at 37°C for 12-16 hours until colonies appeared. Colonies were picked from the plates and shaken at 37°C and 250 rpm for 14 hours. Positive clones were screened by PCR. The PCR reaction used a 20 μL system: 2 μL bacterial culture, 0.5 μL polymerase buffer, 3 μL buffer, and 14 μL ddH2O. The cycling parameters were: 96°C pre-denaturation for 3 min; 95°C for 15 s, 58°C for 15 s, 72°C for 20 s, 23 cycles, and a final extension at 72°C for 1 min.

[0049] iii) Sequence Validation and Plasmid Preparation: Positive clones from the initial screening were validated using Sanger sequencing. Clones with 100% sequence identity were amplified and high-purity plasmids were extracted using a plasmid extraction kit. The extracted plasmids were identified by agarose gel electrophoresis (1%) (band size equals vector plus sequence). Some results are shown below. Figure 5 As shown. The correct strains, verified by sequencing, were purified by streak purification twice to form single clones, which were then prepared into a glycerol strain library and stored at -80℃ for long-term preservation.

[0050] (3) Amplification and purification of the target fragment: High-purity and high-accuracy target DNA fragments were prepared using the verified plasmid as a template to provide high-quality raw materials for the subsequent construction of standard substances. The entire process underwent rigorous optimization and multiple quality controls to ensure that the 14 fragments obtained in the end met the requirements of standard substances in terms of sequence accuracy and physicochemical properties.

[0051] i) PCR amplification system optimization and high-fidelity enzyme selection: In the PCR amplification stage, the focus was on screening and optimizing high-fidelity DNA polymerases. This study compared two commonly used high-fidelity enzymes: NEB Q5 High-Fidelity DNA Polymerase and BGI's self-produced high-fidelity PFU enzyme. Through systematic comparison, the most suitable amplification system for this project was determined.

[0052] High-fidelity enzyme comparative experimental design: The same plasmid template and primer combination were used. A uniform amplification program was employed: 98℃ pre-denaturation for 30 seconds; 35 cycles (98℃ 10 seconds, 60℃ 30 seconds, 72℃ 30 seconds); final extension at 72℃ for 5 minutes. A comprehensive evaluation of the yield, accuracy, and fragment integrity of the amplified products was performed.

[0053] Technical parameter comparison and analysis: NEB Q5 High-Fidelity DNA Polymerase exhibits superior overall performance, with a fidelity approximately 200 times that of ordinary Taq enzymes and an error rate of 2.8 × 10⁻⁶. -7 In comparison, while PFU enzymes offered comparable fidelity, their amplification efficiency was slightly lower than that of Q5 enzymes. Sanger sequencing of the amplified products confirmed that the sequence amplified using Q5 enzymes was 100% identical to the reference sequence, with no base errors detected.

[0054] Final optimized reaction system: Based on the above comparison results, a 25 μL standard reaction system was selected, and its specific composition is shown in Table 3.

[0055] Table 3. Optimized reaction system

[0056]

[0057] The system was verified by three independent repeated experiments, showing good reproducibility.

[0058] ii) Product purification and identification: The PCR products were separated by agarose gel electrophoresis (1.5%). Target bands of the same size as the target fragment were excised and purified using an agarose gel DNA recovery kit. A small amount of the purified product was verified by agarose gel electrophoresis (1.5%). Some results are shown below. Figure 6 The fragment shown is the correct size and has good purity.

[0059] iii) Final quality control: The purified final DNA product is subjected to Sanger sequencing again to ensure that no mutations were introduced during the amplification and purification process.

[0060] Through the aforementioned rigorously optimized processes and quality control measures, 14 DNA samples with completely accurate sequences and meeting the required purity were successfully prepared. These samples met the stringent requirements for standard substances in terms of concentration, purity, and sequence accuracy.

[0061] (4) Determination of DNA fragment purity and sequence information

[0062] The 14 DNA sequences prepared above underwent systematic quality testing, which included two key parts: fragment purity analysis and Sanger sequencing verification. Multiple quality controls ensured that each DNA sample met the stringent requirements for the preparation of standard substances.

[0063] i) Purity detection of DNA fragments

[0064] Sample pretreatment: The 14 purified PCR products obtained in step (3) were diluted with TE buffer (10 mM Tris-HCl, 1 mM EDTA) at pH 8.0. Each sample was vortexed for 30 seconds, briefly centrifuged, and then a standardized DNA suspension of the corresponding volume was prepared for use.

[0065] Chip testing process

[0066] Chip preparation: Using an Agilent high-sensitivity DNA chip, add 9 μL of gel-dye mixture to each well. Sample loading: Accurately transfer 1 μL from each standardized sample into the designated sample well. Instrument analysis: Microfluidic electrophoresis analysis was performed using an Agilent 2100 Bioanalyzer.

[0067] Quality control standards and results

[0068] Purity determination: A qualified sample must show a single main peak at the expected size in the electrophoresis pattern, without primer dimers (<100bp) or other non-specific amplification products or other contaminating peaks; Integrity confirmation: The size of the fragment corresponding to the main peak must be consistent with the designed sample length, with fluctuations controlled within ±5%. Results are as follows: Figure 7 As shown.

[0069] The test results showed that all 14 samples met the above quality control standards, proving that each DNA sample had high purity and good integrity, and met the requirements for subsequent preparation of standard substances.

[0070] ii) Detection of fragment DNA Sanger sequencing sequence information

[0071] To ensure sequence accuracy to the greatest extent possible, this study employs a multiple validation strategy:

[0072] ① Tripartite verification: Each sample was sent to two authoritative sequencing institutions, Sangon Biotech and Riboxin Biotech, for independent verification and compared with the monoclonal plasmid sequencing results provided by BGI.

[0073] ② Bidirectional sequencing: Each sample undergoes both forward and reverse bidirectional sequencing to ensure full-length sequence coverage;

[0074] ③ Repeated testing: Two independently prepared samples were provided for sequencing for each institution.

[0075] Sequence analysis quality control:

[0076] ① Sequence alignment: Use BioEdit software to align the sequencing peak diagram with the reference sequence;

[0077] ② Quality assessment: High clarity of sequencing peaks, low background noise, and Q values ​​greater than 30 are required;

[0078] ③ Consistency standard: The sequencing results of the three institutions are required to have 100% consistency with the reference sequence, and the forward and reverse sequencing results must be completely consistent.

[0079] Fourteen samples underwent purity testing and multiplexing verification using the aforementioned system. The Sanger sequencing alignment results for sample SEQ ID NO. 13 are as follows: Figure 8 As shown, this ensures that the 14 sequences meet the stringent requirements of standard reference materials in terms of concentration, purity, and sequence accuracy, laying a solid foundation for the construction of high-quality standard reference materials.

[0080] 2. Dispensing and Storage

[0081] To ensure the excellent stability and consistency of the reference materials during storage and use, this application establishes a standardized dispensing process and strict storage conditions. All operations are performed in a clean environment that meets the requirements of molecular biology experiments, minimizing exogenous contamination and nucleic acid degradation.

[0082] (1) Pre-packaging and concentration standardization

[0083] Initial concentration determination

[0084] Absolute quantification of 14 DNA candidates was performed using droplet digital PCR technology, and their copy number concentrations were accurately determined as shown in Table 4; three technical replicates were set up for each sample.

[0085] Dilution calculation and preparation

[0086] Based on the ddPCR results, the required dilution volume for each sample was determined using the standard calculation formula:

[0087] The target was set at 1050 ng to meet the sample introduction requirements of current sequencing platforms.

[0088] Table 4. Preliminary Values ​​for Copy Number Concentration

[0089]

[0090] (2) Preservation and packaging

[0091] TE buffer (10 mM Tris-HCl, 1 mM EDTA) at pH 8.0 was used as the standard dilution and storage medium. All dispensing operations were performed in a biosafety cabinet to ensure a sterile environment. The specific procedure is as follows:

[0092] i) Sample pretreatment

[0093] The DNA stock solution stored at -80℃ was placed on ice and thawed slowly. After it was completely thawed, it was vortexed at 3000 rpm for 10 seconds to ensure that the solution was homogeneous. The droplets on the tube wall were collected by short-term centrifugation.

[0094] ii) Dilution and Mixing

[0095] Using a metered and calibrated pipette, accurately transfer the calculated volume of DNA sample and mix it with TE preservation solution in a sterile container; place the mixture in a shaker at 4°C and oscillate at 150 rpm for 60 minutes to ensure thorough mixing.

[0096] iii) Packaging

[0097] Using a pipette, accurately aliquot the well-mixed DNA solution into pre-chilled 0.5 mL sterile cryovials. Each aliquot should contain 32 μL.

[0098] iv) Storage and Management

[0099] The dispensed standard substances are numbered according to batch, and 100 tubes are sealed in a special cryopreservation box; they are immediately transferred to an ultra-low temperature freezer at -80℃ for storage, with temperature fluctuations controlled within ±5℃.

[0100] The standardized dispensing and storage process described above ensures that this reference material maintains good stability and metrological consistency during storage, laying a solid foundation for subsequent homogeneity testing, stability monitoring, and metrological evaluation.

[0101] Table 5. Specific Volume Required for Each Sequence

[0102]

[0103] 3. Preliminary concentration determination of standard substances

[0104] To confirm the quality of the dispensing process and ensure good consistency among the dispensing units, this study conducted a preliminary homogeneity assessment of the dispensed standard substances.

[0105] (1) Experimental procedure

[0106] Eleven tubes of samples drawn from the same batch were used for concentration determination using a high-sensitivity fluorescence quantitative method. Qubit was used. TM 4.0 fluorometer and Qubit TM Absolute quantification was performed using the dsDNA HS Assay Kit. The kit's operating procedure was strictly followed: 199 μL of working solution was mixed with 1 μL of standard sample, incubated at room temperature in the dark for 2 minutes, and then analyzed. Two technical replicates were performed for each sample, and the average value was taken as the concentration measured in that tube.

[0107] (2) Results and Analysis

[0108] The concentration of the 11 standard reference tubes was measured to be 33.39 ng / μL, with an RSD of 4.0%. This initially indicates that the concentration variation between each dispensing unit is extremely small. The excellent concentration consistency among the dispensing units preliminarily demonstrates the stability of the dispensing process and the uniformity of the dispensing results for this standard reference.

[0109] Table 6. Qubit detection results of standard substances

[0110]

[0111] Example 3: Establishment and Optimization of the Fixed Value Method

[0112] To ensure accurate and reliable absolute quantification of 14 sequences, this study established a complete quantification method system based on droplet digital PCR (ddPCR) technology. This system covers the entire process from primer and probe design and reaction condition optimization to specificity verification. Through systematic methodological comparison and condition optimization, optimal quantitative analysis results were obtained for each gene sequence. The entire study adopted a phased and step-by-step optimization strategy. First, primer and probe design at the bioinformatics level was completed, followed by the establishment and optimization of experimental methods, and finally, the reliability of the methods was confirmed through rigorous verification experiments. Through bioinformatics pre-design and rigorous verification, 14 sets of highly specific primers were obtained. (2) Through systematic optimization, the primer and probe concentration and annealing temperature for each sequence were determined as the optimal reaction conditions. Finally, a quantification method for 14 sequences was established.

[0113] 1. Primer design

[0114] Fourteen synthesized gene sequences (SEQ ID NO.1-SEQ ID NO.14) were obtained from the company. Primers and probes were designed using NCBI, Premier 3 Plus, and DNAMAN software, and were designed according to primer and probe design principles. Primer design strictly adhered to the following principles: primer length 18-24 bp, with 20 bp being the optimal choice; Tm value 58-60℃, with the Tm value difference between the upstream and downstream primers not exceeding 2℃; GC content 30-80%. Probe design principles included: length 13-25 bp; GC content 30-80%.

[0115] Table 7. Specific Primer Table

[0116]

[0117] 2. Establishment and optimization of digital PCR determination method

[0118] (1) Optimization of amplification conditions

[0119] Primer concentration optimization: Primer concentrations were set at 700 and 900 nmol / L, and probe concentrations at 125 nmol / L, 250 nmol / L, and 500 nmol / L, respectively. Primer concentrations were screened based on the distinctness of the distinction between positive and negative droplet clusters in the digital PCR one-dimensional plot. Taking sequence 5 (SEQ ID NO. 5) as an example, the digital PCR scatter plot was analyzed, and the results are as follows... Figure 9 As shown, the horizontal axis, Event Number (number of events / droplet number), represents the cumulative number of detected droplets. The vertical axis (Y-axis): Ch1 ​​Amplitude (fluorescence amplitude / fluorescence intensity) represents the relative fluorescence signal intensity detected in each droplet, reflecting the amplification of target DNA / RNA in that droplet. The higher the value, the more target molecules are contained in the droplet or the higher the amplification efficiency. When the primer-probe combination concentration is 700-500 nmol / L, the separation of positive and negative droplet clusters is the clearest and the signal-to-noise ratio is the highest, so this is determined to be the optimal primer-probe concentration.

[0120] (2) Annealing temperature optimization: Digital PCR experiments were conducted at annealing temperatures of 55℃ and 60℃. The separation degree of positive and negative droplet clusters was compared to select the optimal annealing temperature. For example, the fragment amplification efficiency was highest and the droplet cluster separation effect was best at annealing at 55.0℃; therefore, 55.0℃ was chosen as the annealing temperature. Experimental results are as follows: Figure 10 As shown.

[0121] The optimal PCR reaction conditions for 14 sequences were obtained by optimizing primer and probe concentrations and annealing temperatures, as shown in Table 8.

[0122] Table 8. Summary of Primer Concentration-Probe Concentration-Annealing Temperature

[0123]

[0124] (3) Primer and probe homogenization

[0125] Due to the varying primer and probe concentrations, the primers and probes were homogenized to facilitate subsequent detection. The synthesized primer and probe powders were uniformly dissolved in ddH2O to a concentration of 10 nM. For primers with a concentration of 700 nM, 7 volumes of 10 nM primer were added to 2 volumes of ddH2O for a 7 / 9 dilution. No dilution was required for 900 nM primers. For probes with a concentration of 125 nM, a 10 nM solution was diluted 0.25 times. For probes with a concentration of 250 nM, a 10 nM solution was diluted 0.5 times. No dilution was required for 500 nM probes. The primers and probes for each sequence were mixed at a ratio of 1.8 μL each of upstream and downstream primers and 1 μL of probe, vortexed for 30 seconds, and then briefly separated.

[0126] Therefore, the ddPCR reaction system for the microbial multitarget amplicon abundance standard material was determined as follows: 10 μL of 2×ddPCRProbes Supermix, 1.8 μL each of forward and reverse primers, 1 μL of probe, 4 μL of DNA template, and ddH2O added to bring the volume to 20 μL. The ddPCR reaction system and PCR procedure for the microbial multitarget amplicon abundance standard material are shown in the table below:

[0127] Table 9. ddPCR Reaction System

[0128]

[0129] Table 10. PCR Reaction Procedure

[0130]

[0131] 3. Methods and Specificity Validation

[0132] To verify the specificity of the established probe-based detection method and the amplification system, a crossover experiment was designed among 14 samples, and a blank control was set up.

[0133] First, a certain volume of each individual sample is aspirated. Then, the remaining 13 samples (excluding the individual sample) are mixed together and simultaneously detected using the designed primers. Each sample is tested twice. The specificity of the established ddPCR method means that each specific primer can only amplify the target sequence and cannot non-specifically amplify other sequences. Figure 11 As shown, when using the primers for the sample shown in SEQ ID NO.1, only the PCR reaction result for the template shown in SEQ ID NO.1 was positive, while the PCR reaction results for the templates shown in SEQ ID NO.2-SEQ ID NO.14 were all negative. This indicates that the established ddPCR detection method for the sample shown in SEQ ID NO.1 is highly targeted and has good specificity for detecting this DNA template.

[0134] The same crossover experimental method was used to detect 14 gene sequences, and the ddPCR detection results were the same as described above, showing good specificity. The specific results are shown in the figure below. The results indicate that this ddPCR detection method can be used for subsequent detection of standard substances and determination of characteristic values.

[0135] 4. Establishment of the constant value method

[0136] (1) Fourteen sets of highly specific primers were obtained through bioinformatics pre-design and rigorous validation. (2) Through system optimization, the optimal reaction conditions for each sequence were determined, including primer-probe concentration and annealing temperature. Finally, a method for determining the specificity of the 14 sequences was established.

[0137] The customized reference materials were subjected to homogeneity analysis and stability assessment according to the technical requirements for homogeneity assessment of reference materials in the national metrological technical specification JJF 1343-2022 "Assignment of Values ​​and Homogeneity and Stability Assessment of Reference Materials". The copy number concentration of the reference materials was detected using digital PCR. The values ​​of 14 fragments of the microbial multi-target amplicon abundance reference material all met the requirements of F. <F 0.05 (v1,v2) indicates that the differences between and within groups meet the actual usage requirements, and the uncertainty caused by non-uniformity is taken into account in the uncertainty assessment section; the values ​​of the microbial multi-target amplicon abundance standard material are within 1 year. The slope is not significant, therefore no instability was observed.

[0138] Example 4: Standard Reference Material Determination

[0139] 1. Screening and verification by joint value-setting laboratories

[0140] A multi-laboratory joint assay was used to determine the values ​​of 14 genes. The main principles for selecting laboratories were: first, whether they had the relevant measurement capabilities; second, whether they were representative of the industry or field; and third, consideration of different digital PCR platforms. We selected the following 8 laboratories and 4 models of testing instruments from several fields, including agriculture, third-party testing centers, testing companies, and metrology institutions, as shown in Table 11.

[0141] Table 11. Joint Value Setting Laboratory Numbers and Instrument Models

[0142]

[0143] The organizing unit, the National Institute of Metrology (NIM), distributed blind samples (human (male) genomic DNA quantification standard, GBW09856) to the nine participating institutions. Upon receiving the blind samples, the participating institutions conducted quantification experiments according to the organizing unit's quantification scheme and submitted their measurement results. The organizing unit compiled the measurement results and compared them with the blind sample results. The results of this blind sample test showed that the measurement results of all nine participating laboratories did not exceed the quantification range of the certified reference material and had acceptable repeatability. These laboratories possess basic and reliable detection capabilities in DNA quantification, meeting the initial requirements for participating in this multi-gene joint quantification study.

[0144] 2. Standard reference value determination scheme

[0145] The digital PCR reaction system established using primers designed for 14 sequences exhibits good specificity and can be used for standard substance determination studies. Based on the preliminary determination results of the stock solution (Table 4), the 14 sample sequences were analyzed using a customized protocol: Three randomly selected tubes of microbial multi-target amplicon abundance standard substances were tested for concentration using digital PCR methods on different platforms, with two replicates per tube. The standard substances were diluted gravimetrically according to the requirements of different digital PCR instruments in nine laboratories (different laboratories may have different measurement platforms), ensuring the standard substances were diluted to the detection range applicable to each digital PCR instrument. Then, the droplet digital PCR system of primers, probes, and mix was prepared according to the actual instructions of each digital PCR instrument. The prepared PCR premix was vortexed and centrifuged. The diluted standard substances were then mixed with the PCR reaction solution. PCR amplification conditions were set according to the instructions of each instrument, with the annealing temperature set to the optimized temperature. The corresponding readout channels (FAM and HEX) for the probes were selected. After naming the experiment, the experiment was initiated to generate and read droplets. The Dixon criterion was used to check for any suspicious values ​​in the measurements. Test the mean and standard deviation of the measured data for statistical significance. If there is no statistical significance, calculate the overall mean and standard deviation.

[0146] 3. Standard reference value determination

[0147] In accordance with the requirements of JJF 1343-2022 "Assessment and Evaluation of Homogeneity and Stability of Standard Reference Materials", digital PCR was used to assess the values ​​of the standard reference materials. Nine laboratories participated in the collaborative assessment. Each laboratory independently measured three sample units under repeatability conditions, with each sample unit measured twice, resulting in a total of nine sets (six values ​​per set) of raw measurements.

[0148] During the data statistical processing stage, the following steps are performed:

[0149] Cochrane Test: First, the Cochrane test was used to test the homogeneity of variances for the nine groups of data. The test showed that the variances of each group were within the significance range. There was no significant difference at 0.05. This indicates that the measurement precision of the nine laboratories is consistent, and the data can be further merged for analysis.

[0150] Grubbs' test: After confirming consistent precision within the group, the results of each laboratory unit were calculated, yielding 54 independent data points. The Grubbs' test was used to test for outliers in this group. The test rule is as follows: If a measured value x... i There are residuals ,when At that time, It should be removed. It depends on the number of measurements and the given significance level. The relevant values. No outliers were found, and the means from all laboratories have been retained.

[0151] Shapiro-Wilk test: The Shapiro-Wilk method was used to test the normality of the nine laboratory data points retained after the above test. The test results show that the data follow a normal distribution.

[0152] Based on the above systematic verification, all data involved in the value determination were of reliable quality and met the requirements of parameter statistics, as shown in Table 12. Therefore, the arithmetic mean of the results from the nine laboratories was used as the standard value of this reference material.

[0153] Table 12, Values ​​of Sequence 1 (shown as SEQ ID NO.1)

[0154]

[0155] Example 5: Uncertainty Assessment

[0156] The uncertainty in the copy number concentration determination of this standard reference mainly comes from the following three aspects: uncertainty introduced by the determination process (uchar): including the repeatability of the measurement method (Type A assessment) and systematic errors such as balance weighing and droplet volume (Type B assessment). Through assessments of the uncertainty introduced by method repeatability, balance, droplet volume deviation, homogeneity, stability, combined standard uncertainty, expanded uncertainty, and reference sequence abundance uncertainty, Table 13 (uncertainty components of the microbial multitarget amplicon abundance standard reference) and Table 14 (calculation table of reference sequence abundance uncertainty) are obtained.

[0157] Table 13. Uncertainty Components of Microbial Multitarget Amplicon Abundance Standard Material

[0158]

[0159] Taking a coverage factor k equal to 2, at a 95% confidence level, according to the formula... The expanded uncertainty (U) of the standard material can be obtained, and the specific results are shown in the table below. Analysis shows that the relative expanded uncertainty corresponding to the gene copy number concentration characteristic value of the microbial multitarget amplicon abundance standard material is % (k=2), where Ci represents the standard value of the i-th sequence, and Ui is the standard uncertainty of the i-th sequence.

[0160] Table 14. Calculation of Abundance Uncertainty of Reference Sequence

[0161]

[0162] Example 6, Results Analysis

[0163] The prepared microbial multitarget amplicon abundance standard material was characterized and analyzed using digital PCR. The standard values ​​of the copy number concentration of each gene in the microbial multitarget amplicon abundance standard material and the expanded uncertainty are shown in Tables 15 and 16.

[0164] Table 15. Standard values ​​and reference sequence information for each sequence of microbial multitarget amplicon abundance standard materials

[0165]

[0166] Table 16. Relative abundance and expansion uncertainty among target sequences in microbial multitarget amplicon abundance standard materials

[0167]

[0168] Based on the principle of digital PCR: the target gene is encapsulated in tens of thousands of nanoliter-sized droplets, and the signal is amplified by PCR amplification. By counting the droplets with fluorescent signals, the absolute quantification of the initial target gene copy number is achieved through Poisson distribution. Therefore, this method can trace back to the natural unit 1 through molecular counting. This traceability approach has been approved by the Nucleic Acid Analysis Working Group (NAWG) of the International Committee on the Quality of Materials (CCQM). Furthermore, the balances used for sample dilution during the determination process have undergone metrological verification, further ensuring the traceability of the measurement results.

[0169] This application successfully developed a microbial multi-target amplicon abundance standard containing 14 sequences. Its nucleotide sequences are shown in SEQ ID NO.1~SEQ ID NO.14, with the abundance increasing sequentially. The GC content of each sequence is below 60%, and the lowest concentration standard value is (4.23±1.23)×10⁻⁶. 6 The highest number of copies / mL was (3.80±1.51)×10. 9 copies / mL (k=2). The relative uncertainty of synthesis for each sequence ranged from 8.87% to 19.26%; the abundance ranged from 0.03% to 28.0%, with corresponding uncertainties of 0.005% to 2.7% (k=2). Sequence information for all sequences was confirmed and is provided as nominal characteristics.

[0170] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A sequencing data correction system for microbial amplicon sequencing, characterized in that, It includes a standard substance preparation module, a ddPCR quantification module, and a data correction module; the standard substance preparation module is used to prepare microbial multi-target amplicon abundance standard substances; the ddPCR quantification module is used to perform absolute quantification of the standard substances; and the data correction module is used to correct amplicon sequencing data based on the values ​​of the standard substances.

2. The sequencing data correction system according to claim 1, characterized in that, The standard substance preparation module includes a sequence design unit, a synthesis and cloning unit, and a dispensing and verification unit. The sequence design unit screens human gut microbiota and obtains reconstructed sequences through sequence alignment. The synthesis and cloning unit obtains high-purity DNA fragments through chemical synthesis and molecular cloning. The dispensing and verification unit completes the standardized dispensing and uniformity testing of the standard substance.

3. The sequencing data correction system according to claim 2, characterized in that, The GC content of the reconstructed sequence is less than 65%, preferably less than 60%, and the reconstructed sequence includes Bacillus, Enterococcus, Salmonella, Fusobacterium, Broutella, Klebsiella, Megamonas, Penicillium, Candida, Escherichia, Saccharomyces, Bifidobacterium, Cladosporium, and Aspergillus.

4. The sequencing data correction system according to claim 3, characterized in that, The reconstructed sequence covers the full 16s region, the 18s region, and the ITS region. Preferably, the reconstructed sequence comprises 14 DNA fragments, the nucleotide sequences of which are shown in SEQ ID NO.1-SEQ ID NO.

14.

5. The sequencing data correction system according to claim 2, characterized in that, The synthetic cloning unit uses a high-fidelity enzyme. The reaction system includes 5× high-fidelity enzyme reaction buffer, dNTPs, primers, template, and enzyme. The annealing temperature is 50℃-65℃, preferably 55℃-60℃. Preferably, the primers include universal primers 27F, 1492R, 515F, 1119R, ITS1F, ITS3F, and ITS4R. Some fragments use specific primers Bif164-F / Pbi R2, and their nucleotide sequences are shown in SEQ ID NO.15-SEQ ID NO.

23.

6. The sequencing data correction system according to claim 1, characterized in that, The ddPCR determination module includes a primer and probe design unit, a reaction condition optimization unit, and a combined determination unit. The primer and probe design unit designs primers with a length of 18-24 bp, a Tm value of 58-60℃, and probes with a length of 13-25 bp, with the first base at the 5' end being non-G.

7. The sequencing data correction system according to claim 6, characterized in that, The joint value setting unit is jointly customized by no fewer than nine laboratories with testing capabilities.

8. A sequencing data correction method based on microbial amplicon sequencing, characterized in that, Includes the following steps: S1. Screening human gut microbiota and designing and obtaining 14 reconstructed sequences; S2. Designing primers and probes and optimizing ddPCR reaction conditions, and establishing a quantification method; S3. Preparing DNA standard materials for the reconstructed sequences and completing aliquoting and homogeneity verification; S4. Determining the values ​​of the standard materials and assessing uncertainty through joint quantification by 9 laboratories; S5. Performing absolute quantitative correction on microbial amplicon sequencing data based on the values ​​of the standard materials.

9. The correction method according to claim 8, characterized in that, The method for obtaining the reconstructed sequence is as follows: if the microbial sequence contains a complementary sequence to the universal primers, it is directly synthesized; if it is missing, bases are added; if it contains both 18S and ITS sequences, sequence splicing is performed. The universal primers include 27F / 1492R for bacterial 16S, 515F / 1119R for fungal 18S, and ITS1F / ITS3F / ITS4R for ITS. Preferably, after the joint determination, the data are statistically analyzed using the Cochrane test, Grubbs test, and Shapiro-Wilke test. The arithmetic mean of the results from nine laboratories is used as the standard value, and the coverage factor k=2 is taken. The expansion uncertainty is evaluated at a 95% confidence level. Based on this standard value, the relative quantitative data of amplicon sequencing is converted into absolute quantitative data, and the sequencing data correction is completed.

10. The application of the sequencing data correction system according to any one of claims 1-7 or the method according to any one of claims 8-9 in correcting amplicon sequencing technology deviations, in calibrating environmental microbial detection instruments, or in evaluating the performance of sequencing kits.