Internal reference composition, kit and method for 16s full-length absolute quantitative detection based on third-generation sequencing
By designing 12 quantitative internal reference sequences and PacBio kinnex 16S rRNA tandem technology, the difficulty of controlling internal reference sequences in the absolute quantification of microbial communities was solved, and efficient and accurate detection of the absolute abundance of microbial communities was achieved. It is suitable for a variety of sample types and reduces costs.
Patent Information
- Application Number
- CN202510556919.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-09-12
AI Technical Summary
In the existing technology for quantifying the absolute abundance of microbial communities, the amount of added internal reference sequences is difficult to control, resulting in dilution of sequencing data or insufficient data volume, making it impossible to accurately quantify the absolute abundance of microbial communities.
A third-generation sequencing-based internal reference composition for full-length absolute quantitative detection of 16S rRNA was designed, including 12 quantitative internal reference sequences P1-P12. By inserting primer binding sites for the 16S-V1V9 fragment into the internal reference sequence, and using PacBio Kinnex 16S rRNA tandem technology for amplification and sequencing, combined with gel electrophoresis analysis and quality control steps, the internal reference sequence can be distinguished from the sample, simplifying the experimental process.
It achieves accurate quantification of the copy number of the bacterial 16S rRNA gene V1V9 fragment, reduces random errors, improves experimental repeatability and stability, reduces costs, is applicable to a variety of sample types, and provides a more in-depth tool for microbial community structure and function analysis.
Smart Images

Figure CN120624685A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of biotechnology and sequencing genomics analysis, and specifically relates to an internal reference composition, a kit and a method for 16s full-length absolute quantitative detection based on third-generation sequencing. Background Art
[0002] The absolute abundance of a microbial community refers to the actual number of microorganisms in a specific environment. This data is crucial for assessing the role and impact of microorganisms in an ecosystem. However, traditional microbial community research mainly relies on the determination of relative abundance. This method cannot directly reflect the actual number of microorganisms, nor can it accurately compare the differences in microbial communities between different samples or under different conditions.
[0003] High-throughput sequencing technology randomly fragments DNA, adds adapters, and prepares sequencing libraries. It then performs extension reactions on tens of thousands of clones within the library, detecting the corresponding signals and ultimately obtaining sequence information. High-throughput sequencing technology can generate large amounts of sequencing data in a short period of time. Compared with traditional first-generation sequencing technologies, high-throughput sequencing technology offers higher throughput and lower costs, providing a powerful tool for in-depth research on microbial communities.
[0004] With the development of high-throughput sequencing (NGS) technology, particularly the application of the Illumina platform, absolute quantification of microbial abundance has become possible. Currently, absolute quantification techniques are mostly limited to second-generation sequencing (NGS). By extracting DNA from a sample, amplifying the 16S region through PCR, and constructing libraries, high-throughput sequencing based on the Illumina platform can generate extensive microbial community data. Compared to second-generation sequencing, which only sequences one or two hypervariable regions, third-generation sequencing (NGS) can obtain full-length 16S sequences. This increased information from more variable regions improves the resolution of species identification. Third-generation sequencing technologies, such as the PacBio platform, can simultaneously obtain sequence information from nine hypervariable regions, enabling taxonomic identification annotated to species. This not only improves species identification resolution but also allows for accurate species annotation using full-length sequences, thereby more realistically reconstructing the microbial community structure within a sample.
[0005] The advantage of third-generation sequencing technology lies in its long read length capability. PacBio sequencing read length can reach tens or even hundreds of kb, which can easily span the full-length sequence of the 16S rRNA gene.
[0006] In practical applications, selecting appropriate internal reference sequences is crucial. They need to be sufficiently different from the target sequence to avoid cross-reactions, while also ensuring similar biochemical properties during amplification and sequencing. This ensures that the amplification efficiency and sequencing coverage of the internal reference sequence can serve as a reliable reference for the target sequence.
[0007] Currently, adding an internal reference sequence with a known copy number to a sample is a common quantitative method. However, adding an internal reference sequence and amplifying and sequencing it along with the sample will cause the internal reference sequence to occupy a portion of the sequencing data. This means that the sequencing data of the actual microorganisms in the sample will be "swallowed" by the internal reference sequence. If the proportion of the internal reference sequence cannot be controlled, it will be impossible to ensure that the sample has sufficient valid data. If the amount of internal reference added is too low, the internal reference sequence will be too low, and the sample will not be sufficiently representative. When the internal reference is added too little, low concentrations of the internal reference cannot be detected, and the standard curve cannot be fitted based on different concentration gradients of the internal reference, resulting in failure of the quantitative experiment. If the amount of internal reference added is too high, the sequencing data of the actual microorganisms in the sample will be diluted by the internal reference sequence, reducing the amount of sequencing data for the actual microorganisms in the sample, and thus failing to meet the expected experimental requirements. Therefore, the internal reference sequence and its addition amount are crucial to the success of the experiment. Summary of the Invention
[0008] In order to solve at least one of the above problems, the present invention provides an internal reference composition, a kit and a method for absolute quantitative detection of 16s full length based on third-generation sequencing.
[0009] In order to achieve the above object, the present invention adopts the following technical means: The first aspect of the present invention provides an internal reference composition for 16s full-length absolute quantitative detection based on third-generation sequencing, wherein the internal reference composition comprises 12 quantitative internal references P1-P12, and the nucleotide sequences of the 12 quantitative internal references are shown in SEQ ID NO: 1-SEQ ID NO: 12.
[0010] In some embodiments of the present invention, the quantitative internal reference P1-P12 contains a pair of primer binding sites for amplifying a selected region of the 16S-V1V9 fragment, and the primer pair is a forward primer 27F and a reverse primer 1492R for amplifying the 16S-V1V9 fragment, and the sequences are SEQ ID NO: 13 and SEQ ID NO: 14, respectively.
[0011] In some embodiments of the present invention, the sequence composition of the quantitative internal reference P1-P12 is: primer 27F + random sequence 1 + primer 1492R + random sequence 2; the random sequence 1 in the quantitative internal reference P1 is shown as SEQ ID NO: 15, and the composition of the random sequence 1 in the quantitative internal reference P2-P12 is: recognition sequence A + SEQ ID NO: 17 + recognition sequence A + SEQ ID NO: 18 + recognition sequence B + SEQ ID NO: 19 + recognition sequence A + SEQ ID NO: 20 + recognition sequence A + SEQ ID NO: 21 + recognition sequence B + SEQ ID NO: 22 + recognition sequence B; the recognition sequence A and the recognition sequence B in the quantitative internal reference P2-P12 are different; the random sequence 2 is shown as SEQ ID NO: 16.
[0012] The second aspect of the present invention provides a kit for absolute quantitative detection of 16s full-length based on third-generation sequencing, wherein the kit comprises the internal reference composition described in the first aspect.
[0013] In some embodiments of the present invention, the kit comprises a quantitative standard S1 which is a mixture of an internal reference composition according to a mass ratio of P1:P2:P3:P4:P5:P6:P7:P8:P9:P10:P11:P12=1:1:1:1:10:10:10:10:100:100:1000:1000.
[0014] In some embodiments of the present invention, quantitative standards S2-S6 obtained by serial dilution of the quantitative standard S1 by 10 times, 100 times, 1000 times, 10,000 times, and 100,000 times are also included.
[0015] In some embodiments of the present invention, the total final concentration of the quantitative reference sequence in the quantitative standards S1-S6 is known to be 2.52×10 9 copies / μL, 2.52×10 8 copies / μL, 2.52×10 7 copies / μL, 2.52×10 6 copies / μL, 2.52×10 5 copies / μL, 2.52×10 4 copies / μL.
[0016] The third aspect of the present invention provides a method for absolute quantitative detection of 16S full-length based on third-generation sequencing, comprising the following steps: Step 1: Provide the sample to be tested, weigh and extract the total DNA of the sample to be tested; Step 2: The quantitative standard S1 in the kit described in the second aspect is diluted in a gradient of 10 to obtain quantitative standards S2-S6 of different concentrations.
[0017] Step 3: Mixing the test sample DNA obtained in step 1 with one of the quantitative standards S1-S6 in step 2 to obtain a test sample mixture containing the quantitative standard; Step 4: Amplify the 16S-V1V9 fragment of the sample mixture to be tested in step 3 to obtain a 16S amplification product containing a quantitative standard product; Step 5: Run the amplified product obtained in step 4 on gel electrophoresis and analyze the bands. Estimate the proportion of quantitative internal reference reads based on the band concentration. If the proportion of quantitative internal reference reads is within the range of 8% to 50%, send the sample for testing. Step 6: Construct a library for the amplicon sample screened in step 5 and perform sequencing. Sequence the quantitative standard based on the offline data, count the sequence numbers of the quantitative internal reference sequences P1-P12, and fit the standard curve based on the copy numbers of P1-P12. Step 7: Filter the internal reference sequences P1-P12 from the quality control-qualified data to obtain high-quality non-internal reference valid sequences. Cluster the valid sequences according to consistency to obtain the quantitative results of each bacterial community in the sample.
[0018] In some embodiments of the present invention, step 5 further includes a gel recovery step: for samples in which the quantitative internal reference sequence is estimated to account for 8% to 50%, gel recovery is performed by cutting out the strips at the positions of the quantitative internal reference and the sample.
[0019] In some embodiments of the present invention, samples with an estimated quantitative internal reference sequence content in the range of 8% to 50% can be mixed according to the concentration obtained by gel electrophoresis analysis and then subjected to magnetic bead recovery.
[0020] In some embodiments of the present invention, in step 5, if the proportion of the quantitative internal reference is not within the range of 8% to 50%, the concentration of the quantitative standard mixed with the sample is adjusted and re-amplification is performed.
[0021] In some embodiments of the present invention, in step 5, since the number of base pairs of the internal reference sequence is less than that of the sample, after amplification and running the test gel, the quantitative internal reference and sample can be distinguished based on the position of the bands. The concentration of each band can be obtained by gel electrophoresis analysis using GeneTools analysis software, and the sequence proportion of the quantitative internal reference can be estimated based on the band concentration.
[0022] In some embodiments of the present invention, step 7 further includes a copy number correction step: the copy number of the sample bacterial community is further corrected according to the constructed standard curve, and the absolute copy number of each bacterial community in the sample is calculated.
[0023] In some embodiments of the present invention, step 7 also includes a step of performing quality control on the offline data, including removing adapter sequences, low-quality and repetitive sequences, to obtain high-quality sequencing data.
[0024] In some embodiments of the present invention, 16S-V1V9 fragment amplification is based on PacBio kinnex 16S rRNA tandem technology, the sequence of the forward primer 27F for 16S-V1V9 fragment amplification is shown in SEQ ID NO: 13, and the sequence of the reverse primer 1492R is shown in SEQ ID NO: 14.
[0025] In some embodiments of the present invention, the sample type is soil, water, distiller's grains, fermented grains, cellar mud, feces, etc.; DNA extraction from the sample adopts an extraction method known in the industry, and an appropriate DNA extraction method is selected according to the different types of samples to increase the DNA concentration and total amount; In some specific implementation cases, high-quality non-reference valid sequences are used for OTU abundance analysis, OTU cluster analysis, species annotation analysis, and statistical analysis, etc. Clustered into one OTU with at least 97% consistency.
[0026] Beneficial effects of the present invention Compared with the prior art, the present invention has the following beneficial effects: The present application designs 12 internal reference sequences as quantitative standards by inserting universal primers in the 16S-V1V9 region into the internal reference sequence, thereby achieving accurate quantification of the copy number of the bacterial 16S rRNA gene V1V9 fragment. On the one hand, 12 internal references usually cover a wider concentration gradient and can more accurately calibrate target molecules with different expression levels in the sample. On the other hand, 12 internal references can provide more data points and reduce the impact of random errors on the calibration model. There is no need to indicate the internal reference in the present application. When the internal reference is designed, the length of the internal reference is designed to differ from the sample length by 117bp. During gel electrophoresis, the internal reference sample can be distinguished from the internal reference sample, and there is no need to indicate the internal reference to estimate the internal reference ratio. There is no need to indicate the internal reference in the internal reference composition. When premixing the quantitative standard, there is no need to premix the indicated internal reference, which simplifies the experimental process and saves costs.
[0027] The 12 plasmids designed in this application are divided into four concentration gradients. The first two concentration gradients each have four plasmid spots (p1-p4, p5-p8), and the last two concentration gradients each have two plasmid spots (p9-p10, p11-p12). This effectively prevents low-concentration gradients from being "swallowed" by high-concentration gradient plasmids due to insufficient data volume, and enhances signal stability through redundant design. The multi-plasmid spot design also offsets experimental batch differences, significantly improving cross-laboratory reproducibility and providing technical support for large-scale applications.
[0028] In addition, the amplification, library construction, and sequencing steps in the experimental protocol are all based on PacBio's kinnex 16S rRNA tandem technology. Tandem primers concatenate samples into a long SMRTbell® structure. While maintaining species-level resolution, the MAS-Seq method improves the throughput of PacBio's long-read sequencer, enabling multiplex sequencing and bacterial population quantification of up to 384 samples. This significantly reduces the cost of sequencing and absolute quantification of individual samples, greatly improving the cost-effectiveness and efficiency of the experiment. This provides higher resolution and higher-quality full-length 16S analysis, resulting in a more cost-effective approach.
[0029] The advantages of this method include accurate results, rapid testing, and low cost. It is applicable to a variety of sample types, including clinical samples, environmental samples, and food samples. This method can provide a deeper understanding of the structure and function of microbial communities, providing an effective tool for microbial diversity research and related applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 The standard curve obtained by fitting the data obtained after adding AS4 plasmid in Example 2 and using the Kinnex 16S rRNA kit to build the library is shown; Figure 2 The standard curve obtained by fitting the data obtained after adding AS5 plasmid in Example 2 and using the Kinnex 16S rRNA kit to build the library is shown; Figure 3 The standard curve obtained by fitting the data obtained after adding BS4 plasmid in Example 2 and using the Kinnex 16S rRNA kit to build the library is shown; Figure 4 The standard curve obtained by fitting the data obtained after adding AS4 plasmid in Example 2 and using the Thermo Fisher Scientific IonAmpliSeq™ for PacBio® kit to build a library is shown; Figure 5The standard curve obtained by fitting the data obtained after adding AS5 plasmid in Example 2 and using the Thermo Fisher Scientific IonAmpliSeq™ for PacBio® kit to build a library is shown; Figure 6 The standard curve obtained by fitting the data obtained after adding BS4 plasmid in Example 2 and using the Thermo Fisher Scientific IonAmpliSeq™ for PacBio® kit to build a library is shown; Figure 7 The standard curve obtained by fitting the sample Y10 using the AS4 quantitative standard in Example 3 is shown; Figure 8 The standard curve obtained by fitting the sample Y10 using the AS5 quantitative standard in Example 3 is shown; Figure 9 The figure shows the accumulation diagram of the bacterial community structure in sample Y10 in Example 3. DETAILED DESCRIPTION
[0031] The following examples are provided to illustrate preferred embodiments of the present invention. Those skilled in the art will appreciate that the techniques disclosed in the following examples represent techniques discovered by the inventors that can be used to practice the present invention and, therefore, can be considered preferred embodiments of the present invention. However, those skilled in the art will appreciate from this disclosure that many modifications may be made to the specific embodiments disclosed herein while still achieving the same or similar results without departing from the spirit or scope of the present invention.
[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one skilled in the art to which this invention belongs, and the disclosures herein and the materials they cite are hereby incorporated by reference. Those skilled in the art will recognize or be able to ascertain, through routine experimentation, many technical equivalents to the specific embodiments of the invention described herein. Such equivalents are intended to be encompassed by the claims.
[0033] The technical solution of the present application will be further described in detail below in conjunction with specific implementation methods.
[0034] Example 1 Plasmid design and synthesis Plasmid design: Insert universal primers for the bacterial 16S-V1V9 fragment into a random sequence. The random sequence is an artificially edited sequence that needs to contain a universal prokaryotic primer binding site, an optimized synthetic filler sequence with the same length and GC content as the in vivo target, and an easily accessible and processable cloning vector. The designed sequence has been verified by NCBI and has no match with known natural species sequences in nature.
[0035] Plasmid synthesis: 12 plasmid sequences were sent to Shanghai Sangon Biotechnology Co., Ltd. for synthesis. For specific plasmid sequences, please refer to the sequence listing.
[0036] To facilitate separation of the plasmids during decontamination, each plasmid, except P1, has two recognition sequences located at the 3' end of the universal primer. The recognition sequences for each plasmid are shown in Table 1 below.
[0037] Table 1 Plasmid recognition sequences
[0038] After the synthesized plasmids are expanded and cultured, they are extracted and the qubit concentration of each plasmid is measured and then mixed in the following two ways: Method 1: Mix according to the ratio of P1:P2:P3:P4:P5:P6:P7:P8:P9:P10:P11:P12=1:1:1:1:10:10:10:10:100:100:1000:1000; Method 2: Mix in the ratio of P1:P2:P3:P4:P5:P6:P7:P8:P9:P10:P11:P12 = 1:1:1:1:5:5:5:5:25:25:125:125.
[0039] Use a Qubit to accurately measure plasmid concentrations. Mix each quantitative reference sequence (plasmid) according to the internal control dosage listed in Table 2, add TE (10mM Tris-Cl, 1mM EDTA, pH 8.0) to 1000μL, label them AS1 and BS1, and vortex to mix thoroughly. Use a Qubit to accurately measure the S1 concentration. Perform three independent replicates. The concentrations of AS1 and BS1 should be approximately 11.22 ng / μL and 1.62 ng / μL, corresponding to a total molecular count of approximately 2.52×10^9 copies / μL and 3.64E×10^8 copies / μL. If the actual measured concentration differs by more than 15%, reconstitution is necessary.
[0040] Table 2 Premixed amount of internal reference plasmid DNA
[0041] The quantitative standards AS1 and BS1 were serially diluted by 10-fold, 100-fold, 1000-fold, 10,000-fold, and 100,000-fold to produce quantitative standards AS2-AS6 and BS2-BS6.
[0042] The copy number of each plasmid in different concentration gradients was calculated according to the formula provided by Kazuyoshi Koike:
[0043] Remark: ①6.02×10^23 is the approximate value of Avogadro's constant, which represents the number of particles (atoms, molecules, ions, etc.) contained in 1 mol of a substance; ②660 (Daltons) is the average molecular weight of a DNA base pair (sodium salt); ③10 ^9 is the conversion factor between g and ng.
[0044] The calculated copy numbers of each plasmid at different concentration gradients of AS1-AS6 are shown in Table 3 below.
[0045] Table 3. Plasmid copy numbers (copies / μL) at different concentration gradients of AS1-AS6
[0046] The calculated copy numbers of each plasmid at different concentration gradients of BS1-BS6 are shown in Table 4 below.
[0047] Table 4. Plasmid copy numbers at different concentration gradients from BS1 to BS6 (copies / μL)
[0048] Example 2 Artificial simulation community experiment Two pure bacterial samples were selected, and the concentrations of the two strains were measured using Qubit, and the purity of each strain was measured using Nanodrop. The specific information of the two pure bacterial samples is shown in Table 5.
[0049] Table 5 Information of two pure bacterial samples constituting the artificial simulated community
[0050] Two pure bacterial samples were mixed at different ratios of 1:1 to create an artificial simulated colony sample. The simulated colony samples were then mixed with different concentrations of the internal reference plasmids AS1-AS6 and BS1-BS6 from Example 1 for amplification using primers 16s-v1v9. The amplification system consisted of 25 μL of amplification, 25 ng of sample, 1 μL of the internal reference plasmid, and water to the top-up. Three replicates were used for each amplicon. The amplification system and procedure are shown in Table 6 and Table 7, respectively.
[0051] Table 6 Amplification system
[0052] Table 7 Amplification procedure
[0053] Forward primer 27F and reverse primer 1492R amplified the 16S-V1V9 region, corresponding to SEQ ID NO: 13 and SEQ ID NO: 14, respectively. The amplified target fragment was approximately 1465 bp in size. The internal reference sequence was 1348 bp in length, and the sample target fragment differed from the internal reference sequence by 117 bp.
[0054] After amplification, the sample was run on a 1% agarose gel and the gel image was saved. Due to the size difference between the sample and internal reference fragments, they could be distinguished during the run. The saved gel image was then analyzed by gel electrophoresis using GeneTools analysis software (Version 4.03.05.0 SynGene). After comparing the concentrations of the PCR products, the internal reference percentage was estimated based on the band concentrations. The estimated internal reference percentages for the two sets of internal reference plasmids in Table 2 of Example 1 are shown in Table 8.
[0055] Table 8 Results of internal reference estimation ratio
[0056] The results showed that according to method 1, the proportions of the two concentration gradient plasmids AS4 and AS5 were within the appropriate range (8%~50%), while according to method 2, only the proportion of the BS4 concentration gradient plasmid was within the range. In comparison, the premixed plasmid scheme of method 1 has a wider range of applicability.
[0057] Based on the PCR product band concentrations obtained using GeneTools analysis software (Version 4.03.05.0 SynGene), the required sample volume for each of the three samples was calculated based on the principle of equal volume. The PCR products were mixed and quantified by Qubit assay. The PCR product mixture was recovered using the HiPure Gel Pure DNA Mini Kit, and the target DNA fragment was eluted with TE buffer. The recovered DNA was then concentrated.
[0058] The mixed samples were then used to construct libraries using the Kinnex 16S rRNA kit. The library construction process includes specific PCR amplification, Kinnex array formation, DNA repair, and nuclease treatment. Through Kinnex array formation, 16S rRNA amplicons were ligated into approximately 19 kb SMRTbells, enabling multiplex sequencing of up to 384 samples, significantly reducing the sequencing cost per sample. After library construction, specific SMRTbell cleanup beads were used for multiple purifications to ensure library purity and quality.
[0059] In order to compare the library construction effect of the Kinnex 16S rRNA kit, Thermo Fisher Scientific Ion AmpliSeq™ for PacBio® was also used for library construction and sequencing.
[0060] Index codes for distinguishing samples and universal sequences required by the sequencing platform were added to both ends of the library, and the constructed amplicon library was sequenced using the PacBio Sequel lle platform.
[0061] After sequencing is completed, the samples are split according to the sample barcode encoding, and then the data is quality controlled, low-quality amplicon sequences are removed, and similar sequences are clustered. Generally, sequences with a similarity greater than or equal to 97% are divided into an OTU, which is then divided into a smaller number of taxonomic units, and species annotation is performed based on these taxonomic units.
[0062] The standard curve was fitted based on the internal reference concentration and the number of sequences. A linear regression equation was established between the number of plasmid DNA copies (log copies) and the number of sequences (log reads) in the sequencing results based on the plasmid copy number information and the number of internal reference DNA reads in the sequencing results. Where x is the logarithm of the copy number and y is the logarithm of the reads. The fitted standard curve, equation and R 2 The actual proportion of internal reference (the proportion of added internal reference sequence to the total number of sequencing reads) was counted, as shown in Table 9. The standard curves obtained by adding different internal reference gradients to each library construction kit are shown in Table 9. Figures 1 to 6 shown.
[0063] Table 9 Actual proportion of internal reference and standard curve
[0064] The results showed that the estimation of the proportion of internal references by gel electrophoresis analysis using GeneTools analysis software can more accurately reflect the proportion of internal references in the actual sequencing data.
[0065] Regardless of the library construction kit used, the standard curve R fitted by the premixed plasmids (AS4, AS5) in method 1 2 The values are all higher than the standard curve R fitted by the second method (BS4) premixed plasmid. 2 The value is high, so the plasmid premixing method of method 1 is used to premix the plasmid for subsequent experiments.
[0066] The standard curve obtained by premixing plasmids in method 1 and sequencing libraries with the Kinnex 16S rRNA kit was used to further calibrate the copy number of the sample flora using the fitted standard curve. The absolute copy number of each flora in the sample was calculated, and the corresponding reads of the sample were substituted into the linear regression formula to calculate the log10 (copy number) of the sample, which is recorded as c:
[0067] Get the number of copies of the sample nucleic acid d (copies):
[0068] The number of nucleic acid copies (copies / ng) is obtained based on the total amount of template DNA added to the amplification (unit: ng) The number of copies of nucleic acid: the unit is (copies / ng nucleic acid):
[0069] The sequence numbers of the samples and the absolute copy numbers after correction by the standard curve are shown in Table 10.
[0070] Table 10 Sequence numbers of samples and absolute copy numbers after calibration with the standard curve
[0071] Comparison of species copy numbers measured using internal references at different concentration gradients added to artificially simulated communities revealed that the copy numbers of species within the same artificially simulated community were essentially consistent, indicating that adding different concentrations of internal references and varying sample inputs did not significantly affect the absolute quantification of bacteria in the sample. This indicates that the internal reference method exhibits high reproducibility and stability, accurately reflecting the absolute abundance of each species in the sample. Furthermore, AS1 and AS2 internal reference concentration gradients are not suitable for co-amplification with samples, as otherwise the proportion of internal reference sequences would be significantly higher. Amplification and run of the AS6 internal reference concentration gradient with samples revealed a low proportion of quantitative internal references, and even the absence of target bands in the internal reference. This is particularly true when the strains in the artificially simulated community are of high quality, where the proportion of internal references is often zero.
[0072] Example 3 Real sample experiment In order to evaluate the accuracy and applicability of the method of the present invention, a cotton swab sample Y10 was collected.
[0073] The sample DNA was extracted using an extraction kit (Ark Bio). The sample to be tested was provided, weighed (g), and the total DNA of the sample to be tested was extracted, and the total amount of DNA extracted was recorded (ng). After the DNA concentration was accurately quantified and standardized using Qubit, the DNA of the sample to be tested was mixed with the three quantitative standards AS3, AS4, and AS5 in the quantitative standard kit to obtain the sample to be tested containing the quantitative standards.
[0074] Perform 16S full-length amplification on the sample to be tested in the previous step to obtain an amplified product of the 16S-V1V9 fragment containing a quantitative standard product. Run the amplified product on a gel and analyze the run bands by gel electrophoresis. Estimate the proportion of quantitative internal reference reads based on the band concentration. If the proportion of quantitative internal reference reads is within the range of 8% to 50%, send the sample for testing.
[0075] Samples that meet the standards were constructed using the Kinnex 16S rRNA library construction kit, which includes specific PCR amplification, Kinnex array formation, DNA repair, and nuclease treatment. Through Kinnex array formation, 16S rRNA amplicons were ligated into approximately 19 kb SMRTbells, making this structure suitable for PacBio's long-read sequencing. Library purity and quality were ensured through multiple purifications using specific SMRTbell cleanup beads.
[0076] Sequencing was performed using the PacBio Sequel 1e sequencing platform, and the data were processed according to the method of Example 2. The actual proportion of internal reference and the fitted standard curve are shown in Table 11. The copy number of the sample was converted based on the standard curve fitted by the internal reference sequence number and the copy number. The conversion unit is copies / ng. The absolute copy number of the top 10 at the genus level was selected for display. The corrected copy number results are shown in Table 12, and the standard curve is shown in Table 13. Figure 7 and 8 As shown, the bacterial community structure is accumulated Figure 9 shown.
[0077] Table 11 Estimated and actual proportions of internal reference and standard curve data
[0078] Table 12 Sequence numbers of samples and absolute copy numbers after calibration with the standard curve
[0079] Comparing the species copy numbers measured by adding internal references of different concentration gradients to real samples, the results were consistent with the results of artificial simulated communities. The copy numbers of species in the same sample were basically consistent, indicating that adding internal references of different concentrations had no significant difference in the absolute quantification of bacteria in the sample.
[0080] The method of the present invention can effectively detect the absolute copy number of 16S at all taxonomic levels in cotton swab samples, with high reproducibility between replicates. This means that the internal reference method has good reproducibility and stability, and can truly reflect the absolute abundance of each species in the sample.
[0081] All documents mentioned in this application are incorporated herein by reference, just as if each document were incorporated herein by reference individually. It should also be understood that after reading the above teachings of the present invention, those skilled in the art may make various changes or modifications to the present invention, and that such equivalents also fall within the scope of the present application.
Claims
1. An internal reference composition for 16s full-length absolute quantitative detection based on third-generation sequencing, characterized by: The internal reference composition includes 12 quantitative internal references P1-P12, and the nucleotide sequences of the 12 quantitative internal references are shown as SEQ ID NO: 1-SEQ ID NO:
12.
2. The internal reference composition according to claim 1, wherein: The quantitative internal references P1-P12 contain primer pair binding sites for amplifying the selected region of the 16S-V1V9 fragment. The primer pair is the forward primer 27F and the reverse primer 1492R for amplifying the 16S-V1V9 fragment, and the sequences are SEQ ID NO: 13 and SEQ ID NO: 14, respectively.
3. The internal reference composition according to claim 2, wherein: The sequence composition of the quantitative internal reference P1-P12 is: primer 27F+random sequence 1+primer 1492R+random sequence 2; the random sequence 1 in the quantitative internal reference P1 is shown as SEQ ID NO: 15, and the composition of the random sequence 1 in the quantitative internal reference P2-P12 is: recognition sequence A+SEQ ID NO: 17+recognition sequence A+SEQ ID NO: 18+recognition sequence B+SEQ ID NO: 19+recognition sequence A+SEQ ID NO: 20+recognition sequence A+SEQ ID NO: 21+recognition sequence B+SEQ ID NO: 22+recognition sequence B; the recognition sequence A and the recognition sequence B in the quantitative internal reference P2-P12 are different; the random sequence 2 is shown as SEQ ID NO:
16.
4. A kit for absolute quantitative detection of 16S full-length based on third-generation sequencing, characterized by: The kit comprises the internal reference composition according to any one of claims 1 to 3.
5. The kit according to claim 4, wherein: The kit includes a quantitative standard S1 which is a mixture of an internal reference composition according to a mass ratio of P1:P2:P3:P4:P5:P6:P7:P8:P9:P10:P11:P12=1:1:1:1:10:10:10:10:100:100:1000:1000.
6. A method for absolute quantitative detection of 16S full length based on third generation sequencing, characterized in that: The steps include: Step 1: Provide the sample to be tested, weigh and extract the total DNA of the sample to be tested; Step 2: The quantitative standard S1 in the kit according to claim 5 is diluted in a gradient of 10 to obtain quantitative standards S2-S6 of different concentrations. Step 3: Mixing the test sample DNA obtained in step 1 with one of the quantitative standards S1-S6 in step 2 to obtain a test sample mixture containing the quantitative standard; Step 4: Amplify the 16S-V1V9 fragment of the sample mixture to be tested in step 3 to obtain a 16S amplification product containing a quantitative standard product; Step 5: Run the amplified product obtained in step 4 on gel electrophoresis and analyze the bands. Estimate the proportion of quantitative internal reference reads based on the band concentration. If the proportion of quantitative internal reference reads is within the range of 8% to 50%, send the sample for testing. Step 6: construct a library for the amplicon sample screened in step 5 and perform sequencing. Sequence the quantitative standard according to the offline data, count the sequence numbers of the quantitative internal reference sequences P1-P12, and fit the standard curve according to the copy numbers of P1-P12. Step 7: Filter the internal reference sequences P1-P12 from the quality control-qualified data to obtain high-quality non-internal reference valid sequences. Cluster the valid sequences according to consistency to obtain the quantitative results of each bacterial community in the sample.
7. The method according to claim 6, characterized in that: In step 5, magnetic bead purification is used to purify and recover samples with an estimated quantitative internal reference sequence ratio of 8% to 50%.
8. The method according to claim 7, wherein: Step 7 also includes a copy number correction step: based on the constructed standard curve, the copy number of the sample flora is further corrected to calculate the absolute copy number of each flora in the sample.
9. The method according to claim 7, wherein: Step 7 also includes a step of quality control of the data, including removing adapter sequences, low-quality and repeated sequences, to obtain high-quality sequencing data.
10. The method according to claim 7, wherein: In step 4, the 16S-V1V9 fragment amplification is based on the PacBiokinnex 16S rRNA tandem technology. The sequence of the forward primer 27F for 16S-V1V9 fragment amplification is shown in SEQ ID NO: 13, and the sequence of the reverse primer 1492R is shown in SEQ ID NO: 14.
Citation Information
Patent Citations
Internal reference composition, kit and method for quantifying absolute abundance of microbial population
CN117844939A
Internal reference composition, kit and method for detecting absolute content of microorganisms based on high-throughput sequencing
CN119753183A
Plant endophytic bacterium absolute quantification method based on gradient internal reference
CN119842944A
Internal standard nucleic acid fragment for microbial analysis and use thereof
JP2021132618A