Method and Application for Absolute Quantification of Metagenome Based on Gradient Internal Parameters

By constructing a gradient internal reference system and rolling ring amplification technology, the absolute quantitative problem in traditional metagenomic sequencing is solved, and the accurate quantification of bacterial flora and genes in the metasample is achieved, which improves the accuracy and universality of the experiment.

CN118389641BActive Publication Date: 2025-05-30SHANGHAI HAOWEITAI BIOTECHNOLOGY CO LTD

Patent Information

Application Number
CN202410610502.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-16
Publication Date
2025-05-30
Estimated Expiration
2044-05-16

AI Technical Summary

Technical Problem

Traditional shotgun metagenomic sequencing technology is difficult to achieve absolute quantification of macrosamples, and the existing methods have problems such as insufficient universality, experimental deviations and large differences in copy numbers.

Method used

Using a gradient internal reference method, a standard internal reference system and sequencing library was constructed, combined with rolling loop amplification and NGS sequencing, an RCA standard curve was constructed to calculate the absolute copy number of the sample target.

Benefits of technology

Absolute quantification of bacterial abundance and gene copy number in macrosamples is achieved, which improves the accuracy and universality of the experiment, and overcomes the problems of deviation and copy number differences in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118389641B_ABST
    Figure CN118389641B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and application for absolute quantification of metagenome based on gradient internal reference. The method includes: constructing a standard internal reference system: using the plasmids amplified from the internal reference genes shown in SEQ ID NO.1 - SEQ ID NO.9 as RCA_BSI1, RCA_BSI2, RCA_BSI3, RCA_BSI4, RCA_BSI5, RCA_BSI6, RCA_BSI7, RCA_BSI8, RCA_BSI9, and forming a standard system with gradient concentrations in a buffer; constructing a sequencing library: extracting all DNA from a sample, after determining its concentration consistency and integrity, adding an appropriate amount of the standard internal reference system, breaking the DNA to an appropriate length range, at least through end repair and filling, adding A, ligating adapters, fragment length sorting, and then PCR amplification to obtain a sequencing library, and sequencing the sequencing library to obtain PE150 sequencing data; constructing an RCA standard curve: comparing the sequencing data with the plasmid sequence to extract data, drawing a standard curve, and obtaining a linear equation; calculating the absolute copy number corresponding to the sample target according to the linear equation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to biotechnology, and particularly to a method and application for absolute quantification of metagenome based on gradient internal reference. Background Art

[0002] With the development of NGS sequencing, shotgun metagenomic sequencing has gradually become a powerful tool for studying microbial community structure. Different from traditional 16S and ITS amplicon sequencing, shotgun metagenomic sequencing sequences, analyzes, and studies all DNA sequences in a metasample, enabling higher-resolution annotation of species; in addition to community structure research, through de novo assembly of sequences, prediction of gene structures, and annotation of gene functions, it can further achieve research on the gene functions in a metasample. The research on metagenomics has evolved from "what microorganisms are in the sample" in the amplicon era to "what these microorganisms in the sample are doing". Furthermore, with the development of bioinformatics technologies such as MGS and BINNING analysis, it has become possible to assemble complete microbial genomes from metasample data, and its importance for microorganisms that cannot be isolated and cultured is self-evident. However, traditional shotgun metagenomic sequencing technology only focuses on the species and functional composition of the research object and is a relative quantitative experiment that only focuses on proportions; as the research progresses, the scientific questions of concern are becoming increasingly complex, and the research on absolute quantification of metasamples has gradually become the industry consensus. In amplicon sequencing, absolute quantification techniques based on internal reference sequencing methods and qPCR methods have gradually become popular. After synthesizing a large number of research results, Nature Biotechnology strongly recommends that "absolute quantification" be used for the microbiome, pointing out that it is "unfortunate" that the absolute quantification detection of the microbiome has not been popularized in current research. For the research on shotgun metagenomic absolute quantification, there are also methods reported. By adding 1-2 DNA fragments with known copy numbers (CN110804655A) or bacteria with known cell numbers (DOI: 10.2144 / btn-2018-0089) to the sample to be tested, they participate in all links of DNA extraction, library construction, and NGS sequencing. Finally, using the data obtained from the sequencing results, with the known copy number / cell number added as the anchor, the number of bacteria per unit mass of the sample to be tested is calculated. Such methods are simple and easy to implement, but there are 3 problems that cannot be ignored: 1) If the added microorganisms exist under natural conditions, although the types of microorganisms that may exist in the research object will be avoided during selection, it also limits the universality of this method; in addition, due to the uncertainty of scientific research and the inability to correctly assemble and annotate microbial sequences due to homology and other problems, the experiment may fail; 2) Deviations inevitably occur during the library construction process of NGS experiments. The library cannot truly restore the original species or gene composition in the sample, and adding only 1-2 bacteria or DNA fragments cannot monitor this deviation, and the risk of quantitative result distortion caused by experimental deviation cannot be avoided. 3) In metasamples (especially environmental soil samples), the copy number differences among microorganisms are extremely large, and the amount of known bacteria or known copy DNA added to the sample to be tested cannot be determined, and it is difficult to simultaneously take into account species or genes with low and high copy numbers.

[0003] The information disclosed in this background section is only intended to enhance the general understanding of the background of the present invention and should not be regarded as an admission or any form of implication that this information constitutes prior art already known to those of ordinary skill in the art. Summary of the Invention

[0004] The object of the present invention is to provide a method and application for absolute quantification of metagenome based on gradient internal reference.

[0005] To achieve the above object, an embodiment of the present invention provides a method for absolute quantification of metagenome based on gradient internal reference, including:

[0006] Constructing a standard internal reference system: Using the plasmids amplified from the internal reference genes shown in SEQ ID NO.1 - SEQ ID NO.9 as RCA_BSI1, RCA_BSI2, RCA_BSI3, RCA_BSI4, RCA_BSI5, RCA_BSI6, RCA_BSI7, RCA_BSI8, RCA_BSI9, and forming a standard system with gradient concentrations in a buffer. The standard system is as follows: Grouping the plasmids RCA_BSI1 - RCA_BSI9, with the number of plasmids in each group being at least 1, the dosages between groups being inconsistent, and the dosage multiples of a single plasmid in each group satisfying 10 n , where n is a positive integer (preferably, n is not greater than 4);

[0007] Constructing a sequencing library: Extracting all DNA from the sample, after determining its concentration consistency and integrity, adding an appropriate amount of the above standard internal reference system, breaking the DNA to an appropriate length range, at least through end repair and filling, adding A, ligating adapters, fragment length sorting, and then PCR amplification to obtain a sequencing library, and sequencing the sequencing library to obtain PE150 sequencing data (preferably, the sequencing data is selected from PE30 sequencing data, PE100 sequencing data, PE150 sequencing data, PE250 sequencing data, etc.);

[0008] Constructing an RCA standard curve: Comparing the sequencing data with the plasmid sequence to extract data, plotting a standard curve, and obtaining a linear equation;

[0009] Calculating the absolute copy number corresponding to the sample target according to the linear equation.

[0010] In one or more embodiments of the present invention, the addition amounts of the internal reference genes in the standard internal reference system satisfy that, based on a total volume of 1350 μL: the amounts of RCA_BSI1 and RCA_BSI2 are both 10000 ng, the amounts of RCA_BSI3 and RCA_BSI4 are both 1000 ng, the amounts of RCA_BSI5 and RCA_BSI6 are both 100 ng, and the amounts of RCA_BSI7 and RCA_BSI9 are both 10 ng.

[0011] In one or more embodiments of the present invention, amplification is carried out by inserting the reference gene into a vector plasmid, then transfecting competent cells, and then selecting monoclonal cells for culture. After culture, the plasmid is extracted and subjected to rolling circle amplification.

[0012] In one or more embodiments of the present invention, the competent cells are selected from: Escherichia coli, Bacillus subtilis, Saccharomyces cerevisiae.

[0013] In one or more embodiments of the present invention, the culture medium used for culture is selected from: LB medium, NB medium, YPD medium.

[0014] In one or more embodiments of the present invention, the samples are selected from soil, feces, and sewage.

[0015] In one or more embodiments of the present invention, the construction of the linear equation is as follows: The sequencing data is aligned with the sequence of the plasmid, and the sequence-matched data of each plasmid is extracted from the original data, and at the same time, the number of Reads of each sequence is obtained; the number of reads after normalization by the reference length is used as the abscissa, and the absolute copy number is used as the ordinate. After taking the logarithm respectively, a standard curve is plotted. The slope of the standard curve is a and the intercept is b.

[0016] In one or more embodiments of the present invention, the process of calculating the absolute copy number corresponding to the sample target according to the linear equation is as follows: The total copy number C of a specific taxonomic unit or gene in a unit sample, C = 10 ^ (a * Log10(R / EffectiveLength) + b) * N / 200I; R is the number of reads remaining after removing the reads from RCA source, and the remaining reads are assembled de novo, and then the microbial community annotation and gene annotation are carried out to obtain the number of reads of the source bacteria and genes; EffectiveLength is the length of the contig obtained by assembling the obtained reads de novo; N is the total amount of DNA extracted from the sample; I is the amount of the sample used.

[0017] In one or more embodiments of the present invention, there is an application of the method for absolute quantification of metagenome based on gradient reference in environmental sample detection as described above.

[0018] In one or more embodiments of the present invention, the environmental sample is a bacteria-containing substance collected from the environment.

[0019] Compared with the prior art, according to the method and application for absolute quantification of metagenome based on gradient reference in the embodiments of the present invention, by adding a series of reference sequences with a 10-fold gradient and jointly constructing a library with the metagenomic DNA, the abundance of the microbial community and the gene copy number in the metagenomic sample can be absolutely quantified. Description of the Drawings

[0020] Figure 1 Agarose gel electrophoresis images of the original Gs_BSI1, RCA_BSI1 after rolling circle amplification, zymo standard (ZymoBIOMICS Microbial Community DNA Standard, ZYMO, D6305), and human DNA in Example 1 according to an embodiment of the present invention;

[0021] Figure 2 Fragment distribution diagram in Example 1 according to an embodiment of the present invention;

[0022] Figure 3 Standard curve graph in Example 2 according to an embodiment of the present invention;

[0023] Figure 4 Standard curves of samples with different amounts of RCA9 added in Example 3 according to an embodiment of the present invention;

[0024] Figure 5 Quantification result graphs when adding 2 ng, 0.2 ng, and 0.02 ng of RCA9 in Example 3 according to an embodiment of the present invention;

[0025] Figure 6 Quantification result graphs when adding 2 ng, 0.2 ng, and 0.02 ng of RCA9 in Example 3 according to an embodiment of the present invention;

[0026] Figure 7 nifH gene detection result graph of soil samples in Example 4 according to an embodiment of the present invention. Detailed implementation manners

[0027] The following describes the detailed implementation manners of the present invention in detail, but it should be understood that the protection scope of the present invention is not limited by the detailed implementation manners.

[0028] Unless otherwise clearly stated, throughout the specification and claims, the term "comprising" or its variations such as "including" or "having" etc. will be understood to include the stated elements or components, without excluding other elements or other components.

[0029] (1) Synthesize and prepare "RCA amplification internal reference sequence mixture"

[0030] 1. Synthesize 9 reference sequences and insert them into the vector pUC57. The reference sequences are Gs_BSI1, Gs_BSI2, Gs_BSI3, Gs_BSI4, Gs_BSI5, Gs_BSI6, Gs_BSI7, Gs_BSI8, and Gs_BSI9 of Patent CN 109943654B; the above plasmids are transferred into Escherichia coli competent cells by chemical transformation, and single colonies are picked and cultured in LB medium for expansion;

[0031] 2. Extract plasmids from the bacterial culture and use Qiagen Single Cell Kit to perform rolling circle amplification on the 9 reference plasmids to obtain RCA_BSI1, RCA_BSI2, RCA_BSI3, RCA_BSI4, RCA_BSI5, RCA_BSI6, RCA_BSI7, RCA_BSI8, and RCA_BSI9:

[0032] 3. Mix each reference sequence (plasmid) according to the ratio in Table 1, add TE buffer (10 mM Tris-Cl, 1 mM EDTA, pH 8.0) to 1350 μL, and label it as RCA9.

[0033]

[0034] 4. Use Qubit to accurately measure the concentration of RCA9, perform 3 independent replicates. The concentration of RCA9 should be approximately 16.4 ng / μL, corresponding to a molecular number of approximately 4.4×10^9 copies / μL. If the difference in the actual measured concentration exceeds 10%, it needs to be re-prepared;

[0035] 5. Dilute RCA9 to 2 ng / μL and add 1 μL (i.e., 2 ng) during the experiment; the corresponding molecular copy numbers are shown in Table 2:

[0036]

[0037] (2) Add RCA9 to the sample and construct a sequencing library

[0038] 1. Extract total DNA from the sample by any method. The amounts of each sample used during extraction should be equivalent, and the amounts of the samples used (such as grams of soil, grams of feces, liters of sewage, etc.) should be strictly recorded, denoted as I, and the total amount of DNA extracted, N (in ng);

[0039] 2. Use the Qubit fluorescence quantitative system to accurately measure the DNA concentration and standardize the sample to 220 ng / 22 μL using 1×TE;

[0040] 3. Take 2 μL of the standardized DNA for electrophoresis detection to judge the consistency of the concentration among samples and simultaneously judge the integrity of the DNA; samples with obvious concentration deviation need to be re-standardized, and degraded samples are regarded as quality control failure samples and not subjected to subsequent experiments;

[0041] 4. Add 0.2 ng or 2 ng of RCA9 to the remaining 200 ng of DNA above (the corresponding RCA addition ratios are 0.1% and 1%). Use the ultrasonic fragmenter Covaris to break the DNA into an appropriate length range, followed by end repair, filling, adding A, ligating adapters, fragment length sorting and other steps, and finally obtain the NGS sequencing library by PCR.

[0042] 5. The NGS sequencing library is measured for concentration by the Qubit fluorescence quantification system, the molar concentration is detected by qPCR, and the inserted fragment length is detected by Agilent 2100. Finally, NGS sequencing is performed on a compatible Illumina or BGI sequencer (such as Novaseq 6000, DNBSEQ-T7) to obtain PE150 data.

[0043] (3) According to the RCA standard curve, perform the detection of the absolute copy number of colonies and gene abundance

[0044] 1. First, align the sequencing data with the RCA plasmid sequence and extract it from the original data, and simultaneously obtain the number of Reads of each RCA sequence; use the number of reads normalized by the internal reference length as the abscissa and the absolute copy number as the ordinate, take the logarithm respectively and then draw the standard curve, and obtain the linear equation;

[0045] 2. After removing the reads from the RCA source, the remaining reads are de novo assembled, and then the microbial community annotation and gene annotation are carried out to obtain the number of reads of the source bacteria and genes, denoted as R;

[0046] 3. According to the linear equation, substitute the number of reads annotated as a specific bacterium or a specific gene into the linear equation to calculate the absolute copy number of this bacterium or this gene;

[0047] From this, the total copy number C of a specific taxonomic unit or gene in the unit sample is deduced backwards. The formula is C = 10^(a * Log10(R / EffectiveLength) + b) * N / 200 (a and b are the slope and intercept of the standard curve).

[0048] Example 1: Fragment length distribution after ultrasonic fragmentation of the internal reference sequence before and after rolling circle amplification

[0049] Synthesize the following sequence and insert it into the pUC57 plasmid;

[0050] SEQ ID NO.1>Gs_BSI1

[0051] ACTGAGATACGGCCCAGACTCCTACGGGAGGCAGCAGTGGGGAATATTGCGCAATGGCCGAAAGGCTGACCCTGAACTCTTGGGATGCGACGTTGAGGGCTGCTGGCATTAGCGAACCGCAATCCCGTACTTGGTAGACAACGAACCCAACACTCCGGAGAGACTGCCTACGCACCGGCTAACTCCGTGCCAGCAGCCGCGGTAATACGGAGGGTGCGAGCGTTGTCCGGAATCACTGGGCGTAAAGGGCGCGTAGGCGGTCCGATACGAAACTTCTTGCACAGGCATGAGGCACGCGTGCGTACCAGACGGCCTCGGAATACACCGGAAACCTTTGAGGCCGCTCCCAGGTGTACGCGAGGCACCAAAGCGGTATTCCATGGTAGAGAACACTGCTGCGGATTACCGACATTTAGCCTCGCGTATAGCACCCTGCCGTTGCGTGGGGAGCAAACAGGATTAGATACCCTGGTAGTCCACGCCGTAAACGATGGACAACCACGCGGTTACACGCGGAGATCCCGCCAGATGAGAGCCCAGGCATCACAGCGATCAGGCACTTGACATACCGCAAGGTTGAAACTCAAAGGAATTGACGGGGGCCCGCACAAGCGGTGGAGCATGTGGTTTAATTCGACGCAACGCGAAGAACCTTACCCAGGCT

[0052] SEQ ID NO.2>Gs_BSI2

[0053] ACTGAGACACGGTCCAGACTCCTACGGGAGGCAGCAGTGGGGAATATTGCACAATGGGCGCAAGCCTGATCATCTGGTCAGGTCTCCTCGACCTACCTACATGTGTGGCACGTTCGAGAGCGTGCTTTCGACCAGTATTGGCGCTCGTCCATAGCTAAGCGCCTAACTGCATAGCTACCTCACCGGCTAACTCCGTGCCAGCAGCCGCGGTAATACGGAGGGTGCAAGCGTTAATCGGAATTACTGGGCGTAAAGCGCACGCAGGCGGTTTGGTTAGGCACACATGGCACCGGTCTGCGTTGTGAGCGTACCTTTCACCTCCTATCGCCGTCGTCGGCGTTATAGACAGCACTCCTTCGTTGCTACTACCTGGGTAGGTTAGCACCACGAATCATCCCGACGAAATGACACCGGATCGGGCTGATAGGTGACTCGGAAAGCTGAATCGAGGTGGGGAGCAAACAGGATTAGATACCCTGGTAGTCCACGCCGTAAACGATGCTTTCCTATGGACTGGACAAGTGCCTCTCCAGCTAACTGCAACAGTCGTGGATCAGGTAATGGTAGAGCTCAGGCCGCAAGGTTAAAACTCAAATGAATTGACGGGGGCCGCACAAGCGGTGGAGCATGTGGTTTAATTCGATGCAACGCGAAGAACCTTACCTGGTCT

[0054] SEQ ID NO.3>Gs_BSI3

[0055] ACTGAGACACGGCCCAGACTCCTACGGGAGGCAGCAGTGGGGAATATTGCACAATGGGGGAAACCCTGATATTCCGTCAGAGCCGCACATAAGGCCAGCAGGGATGACTAGATATTCCCGCACGCCGACACTGACACTGTCAACGGGTGCAGGTACCACGGCTAACTACGTGCCAGCAGCCGCGGTAATACGTAGGTGGCAAGCGTTGTCCGGATTTACTGGGCGTAAAGGGAGCGTAGGTGGATTGCCTAGCCTAGCGGTTGGAGCGGGAATACTAAGGTGCAGTGTTTCCCATGGCCGGATTTAGGCGCTTCACGGGATCGTCGAGTTCGTCTGGGCATAGGTCGCGTCACGCTTTCATTCGAAGCCGATCTCGGACCGCCACCGATAAGGCAGTAGTCGCCTGTGGCATGTGCCTCGCTGAGGTGGGGAGCAAACAGGATTAGATACCCTGGTAGTCCACGCCGTAAACGATTGTAAGGTAGCCCGTGTAAGGGTACCTCGCAGATTCGCCAAGAACGAGCAAGGGCTAGTGTGTTGACGGTTACTCTCGCAAGATTAAAACTCAAAGGAATTGACGGGGGCCCGCACAAGCAGCGGAGCATGTGGTTTAATTCGAAGCAACGCGAAGAACCTTACCTAGACT

[0056] SEQ ID NO.4>Gs_BSI4

[0057] ACTGAGACACGGTCCAGACTCCTACGGGAGGCAGCAGTGGGGAATATTGCACAATGGGCGCAAGCCTGATTTCGGACCTGTGAGCTATGCTGCGCTATGTGCTTAGGGTCGCCTGGACGTTCACAGGATTCCGTCGTGATGCCCATCTTCGAGGTGTGGTAGCGGACGGACGCCATCCGTCACCGGCTAACTCCGTGCCAGCAGCCGCGGTAATACGGAGGGTGCAAGCGTTAATCGGAATTACTGGGCGTAAAGCGCACGCAGGCGGTTGCCGGGTATTGTAGAGACGTCGCACTTACTTGCTCCACCCGACTCGACCCTGTTGGGTTACCGCGGAAAGTTTGGTCGTCCTCTCTGCCATCAGGCGGGATATAGGGAGCTCCGGCAAACGTGGTGCATCCGCAGAGCGGGATGCTGTGGTCAGCTTGGACTTGACGACTGTACCTCAGCGTGGGGAGCAAACAGGATTAGATACCCTGGTAGTCCACGCCGTAAACGATACACGGGCTAATCTGTAATCTCCGGTTCCTTGACGCTGCCATGGCGTTTACGACGTTACGGACCACTTGCCAGAACCGCAAGGTTAAAACTCAAATGAATTGACGGGGGCCGCACAAGCGGTGGAGCATGTGGTTTAATTCGATGCAACGCGAAGAACCTTACCTGGTCT

[0058] SEQ ID NO.5>Gs_BSI5

[0059] ACTGAGACACGGTCCAGACTCCTACGGGAGGCAGCAGTGGGGAATATTGCACAATGGGCGCAAGCCTGATAATGCGACGCACGTTAGCAGGCCCTAGTTATTAGCCCGTAGCTTGAAGCACTAGATTCTACGCGGGTTCATCAGCCCAGACCCAACAATGAGGGTCCAATCCATGGCTAGCACCGGCTAACTCCGTGCCAGCAGCCGCGGTAATACGGAGGGTGCAAGCGTTAATCGGAATTACTGGGCGTAAAGCGCACGCAGGCGGTTGTAAGGCGACTTCTCTTATGACCAAAGTGGGCGTCCATGGCTTAGACTCGTGTGGCTCGAACCGAAGTCTTGACGTGATCTCGGGAGGGATGGTCGAGCTACTACCACACTCTCGGCTCAATTACCGTGTGACATCGGATACTCCAACATGGCACGGCGACTGTATTACACGATCCTGGTGTGGGGAGCAAACAGGATTAGATACCCTGGTAGTCCACGCCGTAAACGATGCGTTCATGATAGGTTCTCGGCAGCTAAAGGACTGCTATCTACTGGGAATAGCTGCCTTGTGACACTGTTCCTTGCCGCAAGGTTAAAACTCAAATGAATTGACGGGGGCCGCACAAGCGGTGGAGCATGTGGTTTAATTCGATGCAACGCGAAGAACCTTACCTGGTCT

[0060] SEQ ID NO.6> Gs_BSI6

[0061] TGGGACTGAGATACGGCCCAGACTCCTACGGGAGGCAGCAGCTAAGAATATTCCGCAATGGACGAAAGTCTGACGTAAATCGCCACCTGGAATCAGGGTGCGCTGTCGTGTGCGGATCGCATGACCGCCAATTCCGTGTAGCAGGGATAGCCTCCCACCTTCGATGATGGCTCGGCATCTGCTCCAACGGCTAATTACGTGCCAGCAGCCGCGGTAACACGTAAGTTGCGAGCGTTGTTCGGAATTATTGAGCGTAAAGGGCATGTAGGCGGTTGTTTGCTCGACATGGTTCGAGCTGGTAGAGATGCGGCCGTCCTAAGAGGAGAGTAGTGCGTCGCAAAGCACCCGGGTCAAGAGCCGGAGTTGACAGCACACCTTGACTCACTCGTATGGCAATAGCAGGACCACATTCGGGTTCGCGCATTACGATTACCACGTGGTCCTTGCTCGACCCGCGGGGAGCAAACAGGATTAGATACCCTGGTAGTCCGCACAGTCAACTATAGGGTGGCTGATGGAGGTAGAGACGACGGACTGCGAGGTGTGGTAGTTTCCTTCCGAGGGTGCACGGTTAGCCCGCAAGGGTGAAACTCAAAGGAATTGACGGGGGCCCGCACAAGCGGTGGAGCATGTGGTTTAATTCGATGGTACGCGAGGAACCTTACCTGGGTT

[0062] SEQ ID NO.7>Gs_BSI7

[0063] ACTGAGACACGGTCCAGACTCCTACGGGAGGCAGCAGTGGGGAATATTGCACAATGGGCGCAAGCCTGATTACTAGCTTCGTTTCCCACCAGGATAGTTAGGAGTGCCGACCCGTTATAGAAGTGCAGTGTCCTTTCTCTGCACTCGAGTTAAGTCGACAAGTCCTCTTACGCTAGGACTCACCGGCTAACTCCGTGCCAGCAGCCGCGGTAATACGGAGGGTGCAAGCGTTAATCGGAATTACTGGGCGTAAAGCGCACGCAGGCGGTTCATCGCGAGGCTTTATACGAGGCACCAAATAAGCACCGTAATAAGTGAGTCCCGCGGGCTTATTGTGCTGCAGTATAGCTACTATAGCGTAGGGATCGATATCAGCTATACCTAGATGAGAGCCCATTTCCGCTCGATATACCTAGGGACACGTAGATGTACTATTTCGGCGACTTGGATGTGGGGAGCAAACAGGATTAGATACCCTGGTAGTCCACGCCGTAAACGATCTACCACATCAGGCACTTGGCTATGAAGACTGCGTAAGCCATTTAGAGTTCGGGCTCCTTCTAAGGCTTAGCAGGCCGCAAGGTTAAAACTCAAATGAATTGACGGGGGCCGCACAAGCGGTGGAGCATGTGGTTTAATTCGATGCAACGCGAAGAACCTTACCTGGTCT

[0064] SEQ ID NO.8>Gs_BSI8

[0065] ACTGAGACACGGTCCAGACTCCTACGGGAGGCAGCAGTGGGGAATATTGCACAATGGGCGCAAGCCTGATTGCGCTCCGAGTATCGACTCAAGCCTCACTAGGAACAGCGGGCGTTTGCACGAACCCTGTTCCGCCGGCTTTGAGTTACTGTCGTGCTGACGCCGTTGAGGCGAGGGTTGCACCGGCTAACTCCGTGCCAGCAGCCGCGGTAATACGGAGGGTGCAAGCGTTAATCGGAATTACTGGGCGTAAAGCGCACGCAGGCGGTTGAAGGCGTTGCGTTCTTCTCGCTGGACCAGACTCTAATGGCATGTCTCATTCTGTCGTGGCCTGTATTCGCGGAGTGTGTGAGTCTTTGGGCAATGTCAGCAGTGAGCTACCCTGGGAGCAGCCGGACCTCTCTTGAGCGAATGCAGACGGGTTACCGAATCACACCCACCCTCACAGCGGTGGGGAGCAAACAGGATTAGATACCCTGGTAGTCCACGCCGTAAACGATAGGCGTTCTGGTGCCCACAATCCTGATCCTGGCATGGGACTACTGTTGGGCGTTATCGAACACACGGAGAGACGTCCGCAAGGTTAAAACTCAAATGAATTGACGGGGGCCGCACAAGCGGTGGAGCATGTGGTTTAATTCGATGCAACGCGAAGAACCTTACCTGGTCT

[0066] SEQ ID NO.9>Gs_BSI9

[0067] ACTGAGACACGGTCCAAACTCCTACGGGAGGCAGCAGTGAGGAATATTGGTCAATGGGCGAGAGCCTGAAGCTATGCCACAGGTTCGGACGGCTGTTAGTGCGGACTGCGGTCCTTAAAGCCGTCAGCCATCCGTACCGTTAGCTCAGCGTCCCTAGCCTTCGCATCGAACGCGACACCGGCTAACTCCGTGCCAGCAGCCGCGGTAATACGGAGGGTGCAAGCGTTAATCGGAATTACTGGGCGTAAAGCGCACGCAGGCGGTTTGCTGGCTTGCTATGGAGTTGGATCTCACAGTGTGTGTAACATGGCAGCGCTCCCGATTTCAACAGAGGCCGTTGACGCGCCGATTGAGCCTCCCATGGCATTGGTCGACCATCAACACGAGCTCGTTGGCTGTGCTGAATAGGGTGCGAGCAGCATCTAGGCGGAATTTCCATCGTGGTGTGGGTATCAAACAGGATTAGATACCCTGGTAGTCCACACGGTAAACGATGCTTAGATGACGGGTGCAGATCCTCTAGATTCCACCGCGAATAGGTCCCGTGAACTCTGCTCCCGGTTGGTACGGCAACGGTGAAACTCAAAGGAATTGACGGGGGCCGGCACAAGCGGAGGAACATGTGGTTTAATTCGATGATACGCGAGGAACCTTACCCGGGCT

[0068] Use Qiagen Single Cell Kit kit to perform rolling circle amplification on the above plasmid. First, denature the template DNA, and the system and conditions are as follows;

[0069] 0.5 μL DLB Reconstitution Buffer 2 μL Deionized Water 2.5 μL Template DNA

[0070] The reaction condition is to incubate at room temperature for 3 min.

[0071] After the denaturation reaction, add 5 μL of the following system (configured using Qiagen REPLI- Cell Kit kit) to terminate the reaction:

[0072] 0.75 μL Stop Solution 4.25 μL Deionized Water

[0073] Then add the following system for rolling circle amplification experiment (the reagents are from Qiagen SingleCell Kit):

[0074] 9 μL <![CDATA[H 2 O sc]]> 29 μL REPLI-g sc Reaction Buffer 2 μL REPLI-g sc DNA Polymerase

[0075] Reaction conditions: 30°C, culture for 12 h.

[0076] After the original plasmid was opened with BamHI restriction endonuclease, the concentration was measured using the Qubit fluorescence quantification system and mixed in the following ratio, denoted as BSI9;

[0077]

[0078]

[0079] After the rolling circle amplification product was purified by magnetic beads, the concentration was measured using the Qubit fluorescence quantification system and mixed in proportion, denoted as RCA9;

[0080] Take two portions of 200 ng of E. coli genomic DNA and add 0.2 ng of BSI and RCA respectively;

[0081] Use Covaris ME220 to fragment the above DNA;

[0082] Use a conventional library construction kit (Tianhao Gene, S1006L) for library construction to fill in the ends and add A in the following system

[0083]

[0084] Reaction conditions: 20°C for 30 minutes, 70°C for 30 minutes

[0085] After filling in the ends and adding A, perform adapter ligation. First, dilute the Basic-AD adapter to 2 μM with deionized water and prepare the following system:

[0086] 30 μL End-Filling + A-Tailing Fragment 26 μL Ligation Buffer 3 μL T4 DNA Ligase 1 μL 2 μM Basic-AD Adapter

[0087] Reaction conditions: 20°C for 30 minutes;

[0088] Use MagicPure Size Selection DNA Beads from TransGen Biotech to purify the ligation product and finally elute with 15 μL of deionized water;

[0089] Use NEBNext μLtra II Q5 Master Mix from NEB for PCR amplification, with 8 thermal cycles to enrich the library;

[0090] After purification with MagicPure Size Selection DNA Beads of TransGen Biotech for Library 1.0X, the library was sequenced on the Novaseq 6000 platform to obtain PE150 data, with a data volume of approximately 5G;

[0091] For data analysis, after removing adapters and low-quality sequences from the downloaded fastq files, the bowtie2 software was used to align to the reference sequences of the internal reference sequence and the E. coli genome respectively. The picard CollectInsertSizeMetrics was used for inter-group analysis of the library insert fragment lengths. Finally, an insert fragment length distribution map was drawn based on the average fragment length per 100bp bin;

[0092] The results are shown in Figure 2 , it can be seen that the consistency of the insert fragment length distribution between the RCA9 internal reference and the gene after full amplification is better than that of the original BSI9 internal reference; Simulating the length-screened library according to the length interval of 200bp, and calculating the proportion of the internal reference and the genome in each interval, it can also be seen that the RCA9 library is better than the BSI9 internal reference. The insert fragment distribution of the RCA internal reference is equivalent to that of the genome in the range of 101-800bp (the ratio is close to 1), and only drops significantly until it exceeds 800bp; while the BSI9 drops significantly at 600bp, and the ratio fluctuates more than that of the RCA9. The above results intuitively show that after rolling circle amplification, the properties of the RCA internal reference are similar to those of the genome, and it can be more accurately used for calculating the absolute values of species or genes in metagenomic absolute quantification.

[0093] Example 2 RCA Library Gradient

[0094] Synthesize the internal reference sequence according to Example 1, and perform rolling circle amplification. Purify it with MagicPure Size Selection DNA Beads of TransGen Biotech at 1.8X, and quantify it with the Qubit fluorescence quantification system;

[0095] Mix 9 RCA samples into RCA9 according to the following ratio;

[0096] Take 200ng of RCA9, fragment it according to Example 1, fill in the ends, add A, and ligate the adapters;

[0097] Use MagicPure Size Selection DNA Beads to perform length sorting on the ligation product at a ratio of 0.82X on the left and 0.67X on the right,

[0098] After purification with MagicPure Size Selection DNA Beads of TransGen Biotech for Library 1.0X, the library was sequenced on the Novaseq 6000 platform to obtain PE150 data, with a data volume of approximately 2G;

[0099] For data analysis, after removing adapters and low-quality sequences from the fastq files downloaded from the sequencer, the bowtie2 software was used to align to the reference sequences of the internal reference sequences respectively, and the number of reads aligned to each internal reference sequence was extracted. Combining with its corresponding molecular copy number, a standard curve was plotted;

[0100] The results are as Figure 3 shown. It can be seen that after library construction and sequencing, the sequencing data can truly reproduce the gradient among the copy numbers of each internal reference sequence, and is linear with good repeatability.

[0101] Example 3: Add internal reference sequences to the ZYMO standard sample to obtain the absolute content of the standard bacteria;

[0102] Synthesize the internal reference sequences according to Example 1, and perform rolling circle amplification. Purify using 1.8X MagicPure SizeSelection DNA Beads from TransGen Biotech, and quantify using the Qubit fluorescence quantification system;

[0103] Mix 9 portions of RCA samples into RCA9 according to the following ratio;

[0104] Take 200 ng of ZYMO standard DNA (ZymoBIOMICS Microbial Community DNA Standard, ZYMO, D6305), and add 2 ng / 0.2 ng / 0.02 ng of RCA9 respectively. Fragment according to Example 1, fill in the ends, add A, and ligate the adapters;

[0105] Use MagicPure Size Selection DNA Beads to perform length sorting on the ligation products at a ratio of 0.82X on the left and 0.67X on the right,

[0106] After purifying the library using 1.0X MagicPure Size Selection DNA Beads from TransGen Biotech, load the library onto the Novaseq 6000 platform to obtain PE150 data, with a data volume of approximately 10 G;

[0107] For data analysis, after removing adapters and low-quality sequences from the fastq files downloaded from the sequencer, the bowtie2 software was used to align to the reference sequences of the internal reference sequences respectively, and the number of reads aligned to each internal reference sequence was extracted. Combining with its corresponding molecular copy number, a standard curve was plotted;

[0108] The results are as Figure 4As shown, after library construction and sequencing, for the ZYMO samples with 0.2 ng and 2 ng of RCA9 added, the sequencing data can truly reproduce the gradient among the copy numbers of each internal reference sequence, and is linear with good repeatability; while for the addition amount of 0.02 ng, the linearity is poor and it cannot truly reflect the gradient among the copy numbers of each internal reference sequence added. This indicates that the addition amounts of 0.2 ng and 2 ng of RCA9 are more appropriate in this experiment.

[0109] As Figure 5 shown, after de novo assembly of the sequencing data, predicting and quantifying annotated genes, it can be seen that for the ZYMO samples with 0.2 ng and 2 ng of RCA9 added, the quantification results have a very strong correlation; while for the ZYMO samples with 0.02 ng and 2 ng added, the correlation of the quantification results is relatively poor. This also indicates that the addition amounts of 0.2 ng and 2 ng of RCA9 are more appropriate in this experiment.

[0110] As Figure 6 shown, after quantifying the annotated species, for the ZYMO samples with 0.2 ng and 2 ng of RCA9 added, the correlation of the species quantification results is also better than that between 0.02 ng and 2 ng; further verifying that the addition amounts of 0.2 ng and 2 ng of RCA9 are more appropriate in this experiment.

[0111] Example 4 Add internal reference sequences to ordinary field soil samples to obtain the absolute copy numbers of species and genes. Synthesize internal reference sequences according to Example 1, perform rolling circle amplification, purify using 1.8X MagicPure Size Selection DNABeads of TransGen Biotech, and quantify with the Qubit fluorescence quantification system;

[0112] Mix 9 RCA samples in proportion to form RCA9;

[0113] Take 12 soil samples (randomly sampled from Pudong, Shanghai). After DNA extraction, take 200 ng of qualified DNA, add 0.2 ng of RCA9, fragment according to Example 1, fill in the ends, add A, and ligate adapters;

[0114] Use MagicPure Size Selection DNA Beads to perform length sorting on the ligation products at a ratio of 0.82X on the left and 0.67X on the right,

[0115] After purifying the library with 1.0X MagicPure Size Selection DNA Beads of TransGen Biotech, load the library onto the Novaseq 6000 platform to obtain PE150 data, and the data volume is about 10G;

[0116] For data analysis, for the fastq files after sequencing, after removing adapters and low-quality sequences, the bowtie2 software was used to align to the reference sequences of the internal reference sequences respectively, and the number of reads aligned to each internal reference sequence was extracted. Combining with its corresponding molecular copy number, a standard curve was plotted;

[0117] The standard curve equation and R2 are shown in the following table: It can be seen that the gradient between internal references can be well reproduced in the sequencing data;

[0118] Sample Standard Curve r2 Sample 1 y = 1.1019x + 5.0592 0.99755 Sample 2 y = 1.0864x + 5.1751 0.99794 Sample 3 y = 1.0891x + 5.1861 0.99745 Sample 4 y = 1.063x + 5.2936 0.99733 Sample 5 y = 1.0497x + 5.19 0.99664 Sample 6 y = 1.0688x + 5.2438 0.99886 Sample 7 y = 1.0777x + 5.293 0.99827 Sample 8 y = 1.0724x + 5.2341 0.99639 Sample 9 y = 1.0895x + 5.1612 0.99726 Sample 10 y = 1.076x + 5.3537 0.99732 Sample 11 y = 1.0599x + 5.2653 0.9955 Sample 12 y = 1.0333x + 5.2921 0.99599

[0119] The sequencing data was de novo assembled to predict and annotate genes, and the absolute copy number was calculated according to the above standard curve and converted to the copy number per gram of soil;

[0120] For qPCR detection, the nifH gene was selected for qPCR absolute copy number detection. The qPCR primers used were as follows: polF: TGCGAYCCSAARGCBGACTC; polR: ATSGCCATCATYTCRCCGGA. In NCBI, a reference gene sequence matching the primers was selected to synthesize a plasmid to prepare a standard for qPCR detection. The synthesized plasmid was digested with BamHI restriction endonuclease and then accurately quantified using the Qubit fluorescence quantification system, denoted as the nifH standard stock solution. The theoretical copy number of the nifH standard stock solution was calculated based on the concentration measured by the Qubit fluorescence quantification system.

[0121] The nifH standard stock solution was serially diluted 10-fold to obtain a series of gradient standards; the series of standards and the soil DNA samples were simultaneously subjected to absolute copy number detection on a qPCR plate, and each gradient standard and sample was set with 3 replicates. TaKaRa TB Green TM Premix Ex Taq TM II (TaKaRa) kit was used. The qPCR detection system and reaction conditions are as follows:

[0122] 5 μL 2×TB Green II 0.2 μL 50×Rox 0.2 μL F / R Prime (10 μM) 3.6 μL Sterile Water 1 μL Template

[0123] Reaction conditions:

[0124] step1 95℃30S step2 (40 Cycles) 95℃5S 55℃30S 72℃30S* step3 (Dissociation Stage) 95℃15S 60 °C 1 min 99℃15S 60℃15S

[0125] * The step is the fluorescence signal collection step.

[0126] After the qPCR reaction, the average of the 3 replicates of the Ct values of the series of gradient standards was taken, and a standard curve was made with the logarithm of the average value and the theoretical copy number of each gradient. Substituting the Ct value detected for each sample into the standard curve equation, the absolute copy number information of the nifH gene of each sample can be obtained.

[0127] The results are visible Figure 7 , and there is a good correlation between the absolute copy number of nifH obtained by the conventional qPCR absolute quantification method and the absolute quantification metagenomic detection method. It shows that the method of the present invention for detecting the absolute copy number of genes in macro samples is operable. At the same time, the metagenomic absolute quantification experiment quantifies after assembling all genes, and the detection throughput is significantly better than the conventional qPCR method.

[0128] Figure 1 In Example 1, the agarose gel electrophoresis patterns of the length of the original Gs_BSI1 and RCA_BSI1 after rolling circle amplification, compared with the zymo standard (ZymoBIOMICS Microbial Community DNA Standard, ZYMO, D6305), show that the fragment length of RCA_BSI1 after rolling circle amplification has increased significantly and is comparable to the genome.

[0129] Figure 2 In Example 1, BSIS9 and RCA9 were respectively added with E. coli genomic DNA, fragmented by Covaris under the same conditions, library constructed and sequenced. After the sequencing reads obtained the inserted fragment length by alignment, it can be seen that the fragment length distribution of the original BSIS9 is smaller than that of the fully amplified RCA9, and also smaller than the fragment length distribution of the ZYMO standard; while the fragment distribution lengths of RCA9 and the ZYMO standard are comparable.

[0130] Figure 3 In Example 2, the sequencing data of the pure RCA internal reference library was aligned to the reference sequence of the internal reference; after the number of reads was normalized by the internal reference length, a standard curve was plotted with the logarithm of the copy number, showing excellent linearity.

[0131] Figure 4 In Example 3, after adding RCA9 to the ZYMO standard and extracting it from the sequencing data, a standard curve was plotted with the logarithm of the copy number after normalizing the number of reads aligned to the reference sequence of the internal reference by the internal reference length. For the samples adding 2 ng and 0.2 ng of RCA9, the standard curve shows excellent linearity. Figure 5 In Example 3, when 2 ng, 0.2 ng and 0.02 ng of RCA9 were added to the ZYMO sample, the gene quantification results in the sample were shown. It shows that the correlation of the gene quantification results in the sample is excellent when adding 2 ng and 0.2 ng; while when adding 2 ng and 0.02 ng, the gene quantification results are relatively worse.

[0132] Figure 6In Example 3, the quantitative results of species in the ZYMO sample when 2 ng, 0.2 ng, and 0.02 ng of RCA9 were added. It shows that the quantitative results of species in the sample are highly linear when 2 ng and 0.2 ng are added, and the slope of the linear equation is closer to 1. The correlation results for the addition amounts of 2 ng and 0.02 ng are relatively poor.

[0133] Figure 7 In Example 4, the correlation results between the qPCR quantification of the nifH gene and the metagenomic absolute quantification (MGS-AQ) detection results in soil samples.

[0134] The foregoing description of specific exemplary embodiments of the invention has been presented for purposes of illustration and example. These descriptions are not intended to limit the invention to the precise forms disclosed, and it is apparent that, according to the above teachings, many modifications and variations are possible. The purpose of selecting and describing the exemplary embodiments was to explain the specific principles of the invention and its practical application so that those skilled in the art can implement and utilize the various different exemplary embodiments of the invention, as well as various different selections and modifications. The scope of the invention is intended to be defined by the claims and their equivalents.

Claims

1. A method for absolute quantification of a metagenome based on a gradient internal reference for detecting functional genes in a microbial community, comprising: Construct a standard internal reference system: using the sequence SEQ ID NO.1-SEQ ID The plasmids of the internal reference genes shown in NO.9 after amplification are marked as RCA_BSI1, RCA_BSI2, RCA_BSI3, RCA_BSI4, RCA_BSI5, RCA_BSI6, RCA_BSI7, RCA_BSI8, and RCA_BSI9, and a standard system with gradient concentrations is formed in the buffer. The amount of the internal reference genes added in the standard system meets the following requirements, based on a total volume of 1350 μL: the amounts of RCA_BSI1 and RCA_BSI2 are both 10000 ng, the amounts of RCA_BSI3 and RCA_BSI4 are both 1000 ng, the amounts of RCA_BSI5 and RCA_BSI6 are both 100 ng, and the amounts of RCA_BSI7 and RCA_BSI9 are both 10 ng. The amplification is to insert the internal reference genes into the vector plasmid pUC57 and then transfect the competent cells, then select the monoclonal cell Escherichia coli for culture in LB medium, and extract the plasmid after culture for rolling circle amplification; Constructing a sequencing library: extracting all DNA from the sample, and after confirming that it has concentration consistency and integrity, adding an appropriate amount of the standard internal reference system, breaking the DNA to a length of no more than 800 bp, at least performing end repair and padding, adding A, connecting adapters, and fragment length sorting, and then PCR amplification to obtain a sequencing library, and sequencing the sequencing library to obtain sequencing data; Constructing an RCA standard curve: aligning the sequencing data with the plasmid sequence to extract data, drawing a standard curve, and obtaining a linear equation; According to the linear equation, the absolute copy number corresponding to the sample target is calculated. The linear equation is constructed as follows: the sequencing data is aligned with the plasmid sequence, the sequence matching data of each plasmid is extracted from the original data, and the number of reads of each sequence is obtained at the same time; the reads number is normalized by the internal reference length as the horizontal coordinate, and the absolute copy number is the vertical coordinate. The logarithm is taken to draw the standard curve. The slope of the standard curve is a and the intercept is b. According to the linear equation, the process of calculating the absolute copy number corresponding to the sample target is: the total copy number C of a specific taxonomic unit or gene in a unit sample, C=10^(a*Log10(R / EffectiveLength)+b)*N / 200I; R is the number of reads after removing the RCA source reads, and the remaining reads are assembled from scratch, and then the bacterial community annotation and gene annotation are performed to obtain the number of reads of the source bacteria and genes; EffectiveLength is the contig length obtained by de novo assembly of the obtained reads; N is the total amount of DNA extracted from the sample; I is the amount of sample used.

2. The method for absolute quantification of metagenomes based on gradient internal references for detecting functional genes in microbial communities according to claim 1, characterized in that: The sample is selected from soil, feces, sewage.

3. Application of the method for absolute quantification of metagenomics based on gradient internal references for detecting functional genes in microbial communities as described in any one of claims 1-2 in environmental sample detection.

4. The use according to claim 3, characterized in that The environmental sample is bacteria-containing material collected from the environment.

Citation Information

Patent Citations

  • Methods for detecting bacterial community composition and absolute content based on internal reference sequences

    CN109943654B

  • Metagenome absolute quantitation method

    CN110804655A

  • Method for detecting bacterial flora composition and absolute content based on internal reference sequences

    CN109943654A

  • Internal standard molecule, kit and method for quantitative detection of pathogenic microorganism metagenome

    CN116240297A

Cited By

  • Nucleic acid quantitative access judgment method and system based on sequencing response characteristic auditing

    CN122201445A

  • Accurate quantitative sequencing method based on multi-level atomic structure interior label

    CN122357702A