Microbial community analysis method using reference community-based calibration model

The reference community-based correction model addresses sequencing biases in microbial community analysis, enhancing accuracy and precision in microbial community structure prediction.

WO2026029644A1PCT designated stage Publication Date: 2026-02-05KOREA RES INST OF STANDARDS & SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/095383
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-30
Filing Date
2025-06-09
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing microbial community analysis methods, particularly those using next-generation sequencing (NGS), suffer from biases due to imbalances in GC content, varying protocols, and experimental workflows such as sample transport and storage conditions, leading to inaccuracies in microbial community profiling.

Method used

A reference community-based correction model is developed to correct sequencing biases, involving steps of quantifying genomic DNA, constructing a mock community, confirming gene sequences, and calculating microbial community composition using a specific mathematical formula to enhance accuracy.

Benefits of technology

The model increases the accuracy of microbial community structure analysis by correcting sequencing biases, enabling precise prediction of community structures, even for unknown microorganisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025095383_05022026_PF_FP_ABST
    Figure KR2025095383_05022026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a microbial community analysis method using a reference community-based calibration model.
Need to check novelty before this filing date? Find Prior Art

Description

Microbial community analysis method using a reference cluster-based correction model

[0001] The present invention relates to a method for analyzing a microbial community using a reference community-based correction model.

[0002] The microbiome is known as a community of microorganisms with physical and chemical characteristics that exist in the habitats where humans live. The number of microorganisms in the human body is 10 14 ~10 15 The number of human cells is two to three times greater than the number of human cells. Furthermore, the number of microbial genes in the human body is estimated to be over 100 times greater than the number of human genes. Among the bacteria in the human body, gut microbes have the greatest ability to influence our health and susceptibility to disease. They strengthen the immune system, support homeostasis and metabolic processes, protect against infection, and influence most bodily processes within the gastrointestinal tract.

[0003] With the inception of the Human Microbiome Project, elucidating the relationship between gene distribution and the role of microbes within the human body has become a crucial research topic. Consequently, the importance of the microbiome is increasingly understood, and extensive research is currently underway to study it. Recently, complex microbial communities, such as the human gut microbiome, have been studied using targeted amplicon-based metagenomic sequencing using the 16S rRNA gene. Despite the advantages of next-generation sequencing (NGS), such as metagenomics studies or whole-genome sequencing, biases in experimental workflows, such as unknown sample transport and storage conditions, DNA extraction methods, and amplification, can arise. Furthermore, strong sequencing biases can be introduced during NGS due to imbalances in GC content and various protocols. Library preparation kits, processes for various NGS platforms, and various bioinformatics tools can also bias sequencing profiling.

[0004] In this study, eight probiotic bacteria were selected from 19 probiotics selected by the Ministry of Food and Drug Safety and mixed into a mock population containing various ratios. Using the MOSTUR pipeline, NGS platforms such as MiSeq (Illumina, USA), Sequel II (Pacific Biosciences, USA), IonTorrent (Thermofish Scientific, USA), MGIseq-2000 (BGI, China), and MinION (Oxford Nanopore, UK) were compared to determine amplicon bias. Furthermore, the biases of four amplicon regions on the MiSeq or IonTorrent platforms were compared. The results revealed that specific strains or platforms exhibited biases, and the development of reference materials (RMs) can correct these biases in microbiome analysis.

[0005] The present invention was derived from the above-mentioned needs, and the inventors developed a reference community-based correction model to correct bias in microbial community analysis, and confirmed a microbial community analysis method using the same, thereby completing the present invention.

[0006] In order to solve the above problem, the present invention provides a method for analyzing a microbial community, including: 1) a step of quantifying genomic DNA of a microbial community to be analyzed; 2) a step of constructing a mock community based on the ratio of the quantified genomic DNA; 3) a step of confirming the gene sequence of the mock community using next generation sequencing; 4) a step of confirming the composition of the mock community from 3); and 5) a step of calculating the composition of the microbial community to be analyzed by correcting the composition of the mock community of 4) using the following [Mathematical Formula 2].

[0007] [Equation 2]

[0008]

[0009] N t : t: Number of sequences after clustering amplification cycles, N0: Initial number of sequences before amplification, t: Number of cycles, E: Amplification efficiency.

[0010] In one example of the present invention, the microbial community to be analyzed may be all kinds of microorganisms including bacteria and viruses, preferably probiotic bacteria, and more preferably Lactobacillus acidophilus, Lacticaseibacillus casei, Lactobacillus delbrueckiisubsp.bulgaricus, Limosilactobacillus fermentum, Lactobacillus gasseri, Lactobacillus helveticus, Lacticaseibacillus paracaseisubsp. paracasei, Lactiplantibacillus plantarum, Lacticaseibacillus rhamnosus, Limosilactobacillus reuteris subsp. reuteri, Ligilactobacillus salivarius, Bifidobacterium animalis subsp. lactis, Bifidobacterium bifidum, Bifidobacterium breve, Bifidobacterium longum subsp. longum, Streptococcus thermophilus, Lactococcus lactis subsp. lactis.lactis), Enterococcus faecalis and Enterococcus faecium, but is not limited thereto.

[0011] In another example of the present invention, the step of quantifying the DNA can be performed by any means for quantifying DNA, and preferably, it can be performed by one or more selected from the group consisting of real-time PCR, droplet digital PCR, or recombinase polymerase amplification, but is not limited thereto.

[0012] In another example of the present invention, the next generation sequencing may include any means for analyzing the sequence of a full-length or specific gene, and preferably, but not limited to, one or more selected from the group consisting of MiSeq, Sequel II, IonTorrent, MGIseq-2000, and MinION.

[0013] In another example of the present invention, the genes of the simulated community may include all genes included in the microorganism, and preferably, at least one selected from the group consisting of 16S rRNA, 18S rRNA, ITS, rpoB, polC, gyrB, mcrA, and nifH, but are not limited thereto.

[0014] The present invention provides a computer program recorded on a computer-readable recording medium for executing a microbial community analysis method, including the steps of: 1) quantifying genomic DNA of a microbial community to be analyzed; 2) composing a mock community based on the ratio of the quantified genomic DNA; 3) confirming the 16S rRNA sequence of the mock community by next generation sequencing; 4) confirming the composition of the mock community from the step 3); and 5) calculating the composition of the microbial community to be analyzed by correcting the composition of the mock community of the step 4) using the following [Mathematical Formula 2].

[0015] [Equation 2]

[0016]

[0017] N t : t: Number of sequences after clustering amplification cycles, N0: Initial number of sequences before amplification, t: Number of cycles, E: Amplification efficiency.

[0018] In the present invention, a computer program can configure a processing device to operate as desired or can independently or collectively command the processing device. The computer program can be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium, or device to be interpreted by the processing device or to provide instructions or data to the processing device. The software can be distributed over networked computer systems and stored or executed in a distributed manner. The computer program can be stored on one or more computer-readable recording media.

[0019] In addition, the present invention provides a microbial community analysis system including at least one processor implemented to execute computer-readable instructions, wherein the at least one processor 1) quantifies genomic DNA of a microbial community to be analyzed; 2) constructs a mock community based on the ratio of the quantified genomic DNA; 3) confirms the 16S rRNA sequence of the mock community by next generation sequencing; 4) confirms the mock community composition from the above 3); and 5) calculates the composition of the microbial community to be analyzed by correcting the mock community composition of the above 4) using the following [Mathematical Formula 2].

[0020] [Equation 2]

[0021]

[0022] N t : t: Number of sequences after clustering amplification cycles, N0: Initial number of sequences before amplification, t: Number of cycles, E: Amplification efficiency.

[0023] The present invention provides a method for identifying microorganisms, comprising: 1) a step of quantifying genomic DNA of a microbial community to be analyzed; 2) a step of constructing a mock community based on the ratio of the quantified genomic DNA; 3) a step of confirming the 16S rRNA sequence of the mock community using next generation sequencing; 4) a step of confirming the composition of the mock community from 3); and 5) a step of calculating the composition of the microbial community to be analyzed by correcting the composition of the mock community of 4) using the following [Mathematical Formula 2].

[0024] [Equation 2]

[0025]

[0026] N t: t: Number of sequences after clustering amplification cycles, N0: Initial number of sequences before amplification, t: Number of cycles, E: Amplification efficiency.

[0027] It was shown that the microbial community analysis method using the reference community-based correction model of the present invention can contribute to increasing the accuracy in the analysis of microbial community structure within a predictable range.

[0028] Furthermore, when analyzing the community structures of known and unknown microorganisms, utilizing a reference model containing the known microbial composition can accurately predict the community structure of unknown microorganisms. This reference-based correction method could serve as a model for addressing sequencing bias in metagenomic research.

[0029] Figure 1 illustrates examples of variable regions of the 16S rRNA gene and target primer pairs used in simulated cluster sequencing according to one embodiment of the present invention. Variable regions are shown in blue, while conserved regions are shown in green. The amplified regions of each sequencing platform are shown in grayscale.

[0030] Figure 2 illustrates the composition of eight mock communities according to one embodiment of the present invention. The X-axis and Y-axis represent the names of the mock communities and the relative abundance of each strain. Mock A and E: Lactobacillus group, Mock B and F: Bifidobacterium group, Mock C: Lactobacillus and Lactococcus group, Mock F: Bifidobacterium and Lactococcus group, Mock G and H: Lactobacillus, Bifidobacterium and Lactococcus groups.

[0031] Figure 3 illustrates the composition of microbial profiles in eight simulated clusters based on sequence reads performed on a MiSeq platform, an IonTorrent platform, and an NGS platform according to one embodiment of the present invention. V1-V2, V3, V4, and V1-V3 represent variable regions of amplified amplicons. AH represents a specific simulated cluster. The Y-axis represents the relative abundance of sequencing profiling for the simulated clusters.

[0032] Figure 4 illustrates a heatmap analysis of MiSeq sequencing profiling on eight simulated clusters according to one implementation example of the present invention. Original, Mis: Miseq. The numbers following the letters indicate the copy number of each amplicon.

[0033] Figure 5 illustrates principal component analysis (PCA) of raw and sequenced simulated clusters according to one implementation example of the present invention. Symbols represent ●, original; □, Miseq; ▲, Ion Torrent; X, MGIseq-2000; ◆, Sequel II; and ◇, MinION. Orange, green, blue, violet, and red colors represent V1-V2, V3, V4, V1-V3, and V1-V9 regions, respectively.

[0034] Figure 6 illustrates a heatmap analysis of IonTorrent sequencing profiling on eight simulated clusters according to one implementation example of the present invention. Original, Mis: Miseq. The numbers following the letters indicate the copy number of each amplicon.

[0035] Figure 7 illustrates a heatmap analysis of eight simulated cluster phase sequencing profilings according to one implementation example of the present invention. Original, Mis: Miseq. The numbers following the letters indicate the copy number of each amplicon.

[0036] FIG. 8 shows a correction ratio of a sequencing result by a reference-based bias correction model according to an implementation example of the present invention.

[0037] Hereinafter, preferred embodiments of the present invention will be described in detail. Furthermore, the following description includes numerous specific details, such as specific components, but these are provided to facilitate a more comprehensive understanding of the present invention. It will be apparent to those skilled in the art that the present invention can be practiced without these specific details. Furthermore, in describing the present invention, detailed descriptions of known functions or configurations will be omitted if they are deemed to unnecessarily obscure the gist of the present invention.

[0038]

[0039] <Example 1> Microbial culture (1)

[0040] The selected probiotic strain is one of 19 strains designated by the Ministry of Food and Drug Safety and obtained from the Korea Center for Biological Resources (KCTC). The eight strains are listed as follows (Tables 1 and 2): Lactobacillus acidophilus KCTC 3164 T ,Limosilactobacillus fermentumKCTC 3112 T ,Lactobacillus gasseriKCTC 3163 T ,Lacticaseibacillus paracaseisubsp.paracaseiKCTC 3510 T ,Limosilactobacillus reuteriKCTC 3594 T , Bifidobacterium animalis subsp. lactis KCTC 5854 T ,BifidobacteriumbreveKCTC 3220 T , andLactococcuslactissubsp.lactisKCTC 3769 T .

[0041] Lactobacillus, Limosilactobacillus, and Lactacaseibacillus strains were cultured at 37°C under aerobic conditions using MRS medium (Difco, USA). L. lactis was grown at 30°C under aerobic conditions and in MRS medium. Bifidobacterium strains were grown under anaerobic conditions using MRS medium at 37°C with H2 (4%), CO2 (10%), and N2 (balance). Eight strains were cultured for 48 h in 2 mL of MRS medium according to each culture guideline. 2 mL of the medium culture was centrifuged at 13,000 rpm for 5 minutes to remove cells.

[0042]

[0043] TaxonomyKCTCnumberNumberofcontigsAccessionnumberGC (%)16S rRNAcopynumberGenomesize (Mb)Lactobacillusacidophilus(L. ap)KCTC3164 T 1GCA_003047065.134.742.00997Limosilactobacillusfermentum(L. fe)KCTC3112 T 1GCA_000010145.151.552.09868Lactobacillus gasseri(L. ga)KCTC3163 T 1GCA_000014425.135.361.89436Lacticaseibacillus paracasei(L. pa)KCTC3510 T 3GCA_000829035.146.653.01780Limosilactobacillus reuteri(L. rt)KCTC3594 T 1GCA_000010005.138.962.03941Bifidobacteriumanimalis(B. al)KCTC5854 T1GCA_000022965.160.541.93269Bifidobacteriumbreve(B. br)KCTC3220 T 1GCA_900637145.158.932.26941Lactococcuslactissubsp.lactis(L. ll)KCTC3769 T 3ATCC 19435*35.462.62782

[0044] T Type strain * ATCC genome portal

[0045]

[0046] InformationStrainsV1-V2V3V4V1-V3V1-V9GC(%)Length(bp)GC(%)Length (bp)GC(%)Length (bp)GC(%)Length (bp)GC (%)Length (bp)Lactobacillusacidophilus54.833246.316052.224752.5512541528Limosilactobacillusfermentum52.534346.916050.624751.2523531540Lactobacillusgasseri50.733946.316051.424749.9519521537Lacticaseibacillusparacasei51.933551.916051.824752.4515531534Limosilactobacillusreuteri51.333951.916053.824752523541540Bifidobacteriumanimalis61.331056.615259.924760482601502Bifidobacteriumbreve61.430853.114558.724759473591495Lactococcuslactissubsp.lactis49.431649.716151.224650.1497521526

[0047]

[0048] <실시예 2> 게놈 DNA 추출

[0049] Genomic DNA (gDNA) of eight probiotics was purified using GenElute according to the manufacturer's instructions. TM Bacterial genomic DNA was extracted using a kit (Sigma Aldrich). 2 ml of bacterial culture medium was centrifuged at 13,000 rpm for 5 min. After discarding the supernatant, the pellet was mixed with 20 μl of lysozyme solution (45 mg / ml) and RNase A at 37°C for 30 min. The mixture was then added to 20 μl of proteinase K and 200 μl of lysis solution C and incubated at 55°C for 10 min. The lysis mixture was gently mixed with 200 μl of 100% ethanol by vortexing for 10 s. The contents of the tube were transferred to a binding column and then centrifuged at 13,000 rpm for 1 min. 500 μl of wash solution I was added to the column and centrifuged at 13,000 rpm for 1 min. The column was then filled with 500 μl of concentrated wash solution and centrifuged at 13,000 rpm for 1 minute. This procedure was repeated after discarding the collection tube. 200 μl of elution solution was immediately pipetted into the middle of the column and incubated at room temperature for 5 minutes. To elute the DNA, the sample was centrifuged at 13,000 rpm for 1 minute. The extracted DNA was then stored at -20°C.

[0050]

[0051] <Example 3> Amplification of 16S rRNA gene

[0052] The bacterial 16S rRNA gene was amplified by PCR using universal primers 27F (AGAGTTTGATCMTGCTCAG) and 1492R (TACGGYTACCTT). Amplification conditions consisted of initial denaturation at 95°C for 1 min, followed by 30 cycles of denaturation at 95°C for 15 s, primer annealing at 55°C for 15 s, primer extension at 72°C for 90 s, and a final extension at 72°C for 7 min. PCR products were purified using the QIAQuick PCR Purification Kit (QIAGEN).

[0053] The 16S rRNA gene amplicons of the probiotic strains were sequenced at Macrogen using sequencing primers 518F (CCAGCCGCGGTAATACG) and 800R (TACCAGGGTATCTAATCC). The raw sequence reads were merged into a single contig using MOTHUR (version 1.41.1). A single contig of bacterial sequences was identified using EzBioCloud.

[0054]

[0055] <Example 4> Quantification of genomic DNA using a fluorometer and ddPCR

[0056] The initial concentration of genomic DNA (gDNA) or amplicons was determined using a Quantus™ Fluorometer (Promega, USA) equipped with the QuantiFluor dsDNA System. In a 0.5 ml PCR tube, 1 μl of sample was added to 199 μl of QuantiFluor Dye working solution. The sample was placed in the instrument holder. The lid was then closed, and the sample concentration was calculated using the instrument.

[0057] The measured samples were serially diluted for accurate detection over the dynamic range, and their copy numbers were quantified using droplet digital PCR (ddPCR; BioRad, Hercules, CA, USA). A 20 μl mixture consisted of 10 μl of 2X QX200 ddPCR EvaGreen Supermix (BioRad, Hercules, CA, USA), 0.5 μl of 337F (GACTCCTACGGGAGGCWGCAG) forward primer, 0.2 μl of 518R (CGTATTACCGCGGCTGCTGG) reverse primer, and 8.3 μl of distilled water. After dispensing the PCR premix into the center line of the cartridge, 70 μl of droplet-generating oil was dispensed into the bottom line of the cartridge to bind the gasket to the cartridge. 45 μl of the droplet product was transferred to a 96-well plate, which was amplified by conventional PCR. PCR amplification consisted of initial denaturation at 95°C for 5 min, 40 cycles of 95°C for 30 s, primer heating and extension at 60°C for 1 min, 4°C for 5 min, and 1 cycle of signal stabilization at 90°C for 5 min. All thermal cycling was performed using a Veriti Thermal Cycler (Thermofisher, USA) using a slow ramp rate of 2°C / s.

[0058]

[0059] <Example 5> Mock community composition

[0060] Using the gDNA of eight quantified strains, mock communities were generated based on specific ratios. The mock communities consisted of eight groups (A, B, C, D, E, F, G, and H). Each mock community was generated gravimetrically and consisted of the following components: Mock A contained Limosilactobacillus, Lactobacillus, and Lacticaseibacillus strains in similar ratios. Mock B contained two Bifidobacterium strains in similar ratios. Mock C contained Limosilactobacillus, Lacticaseibacillus, and Lactococcus lactis subsp. lactis strains in similar ratios. Mock D contained two Bifidobacterium strains and Lactococcus lactis subsp. lactis strains in similar ratios. Mock E contained Limosilactobacillus, Lactobacillus, and Lacticaseibacillus strains in different proportions. Mock F contained two Bifidobacterium stains in different proportions. Mocks G and H contained eight probiotic strains in different and similar proportions.

[0061]

[0062] <Example 6> MiSeq metagenome sequencing

[0063] In the simulated community, the V1-V2, V3, V4, and V1-V3 regions of the 16S rRNA gene were amplified by a mixture of miseq27F and miseq337R (V1-V2), miseq337F and miseq518R (V3), miseq518F and miseq800R (V4), and miseq27F and miseq518R (V1-V3). The miseq27F primer mixture had a ratio of miseq27F to miseqbif27F of 8:2 (Fig. 1 and Table 3). The amplification consisted of an initial denaturation at 95°C for 1 min, followed by 30 cycles of primer annealing at 55°C for 15 s, extension at 72°C for 90 s, and a final extension at 72°C for 7 min. Secondary amplification was then performed using the I5 forward primer and the I7 reverse primer, attaching the Illumina NexTera barcode. The conditions for the second amplification were the same as above, except that the amplification cycle was changed to 8 cycles. PCR results were verified and identified using a 1% agarose gel and the Gel Doc system (BioRad, Hercules, CA, USA). The amplified products were purified using CleanPCR (CleanNA, Netherlands). Equal amounts of purified products were pooled, and short fragments (non-target products) were removed using CleanPCR. The quality and size of the amplicons were assessed using a DNA 7500 chip and a Bioanalyzer 2100 (Agilent, USA). All amplicons were pooled and sequenced using an Illumina MiSeq sequencing system at Chunlab, Inc. (Korea) according to the manufacturer's instructions.

[0064]

[0065] <Example 7> InoTorrent metagenome sequencing

[0066] In the simulated community, the V1-V2, V3, V4, and V1-V3 regions of the 16S rRNA gene were amplified with a mixture of 27F and 337R (V1-V2), 337F and 518R (V3), 518F and 800R (V4), and 27F and 518R (V1-V3). The primer mixture of the 27F primer had a 27F to bif27F ratio of 8:2 (Fig. 1 and Table 3). Sequencing DNA libraries were constructed using the Ion Plus Fragment Library Kit (Thermofisher Scientific, USA) together with the Ion Code Barcode Adapter Kit (Thermofisher Scientific, USA) according to the manufacturer's library preparation technique for the amplicons. qPCR was used to quantify the barcoded libraries using the Ion Universal Library Quantitation Kit (Thermofisher Scientific, USA). Each library was diluted to a final concentration of 100 pmol, and the reaction was performed according to the manufacturer's instructions. To ensure that each barcode library was equally represented in the sequencing run, the barcode libraries from each amplicon were individually quantified and pooled at equimolar total concentrations. Templates were prepared using Ion 520 according to the manufacturer's instructions. TM Ion 530 TM The Ion Chef System (Thermofisher Scientific, USA) was used to prepare the ExT Kit-Chef and Sequencing Kit (Thermofisher Scientific, USA) and the Ion 530 Chip (Thermofisher Scientific, USA). Twenty-seven barcode libraries were sequenced simultaneously using a single chip.

[0067]

[0068] <Example 8> MGIseq-2000 Metagenome Sequencing

[0069] The V3 region of the 16S rRNA gene in the simulated community was amplified using primers 337F and 518R (V3) (Fig. 1 and Table 3). 1 μg of genomic DNA was fragmented using an S220 sonicator (Covaris, USA). Sequencing libraries were prepared using the MGIEasy DNA library prep kit (MGI, China) according to the manufacturer's instructions. Fragmented genomic DNA was size-selected using AMPure XP magnetic beads (Beckman Coulter Genomics, USA), and then end-repaired and a-tailed at 37°C for 30 min and 65°C for 15 min. The ends of the DNA fragments were ligated to indexing adapters at 23°C for 60 min. PCR was performed to enrich DNA fragments with adapter molecules under the following PCR conditions: 95°C for 3 min; 7 cycles of 98°C for 20 s, 60°C for 15 s, and 72°C for 30 s; and a final extension at 72°C for 10 min. The concentration of the indexed library was quantified by the QuantiFluor ONE dsDNA System (Promega, USA). The library was first incubated at 37°C for 60 min, followed by 30 min of digestion and circularization product washing. The library was treated with DNB enzyme at 30°C for 25 min to generate DNA nanoballs (DNBs). The final concentration of the library was then measured by the QuantiFluor ssDNA System (Promega, USA). The prepared DNBs were sequenced using 150-bp paired-end reads on the MGIseq-2000 system (MGI, China).

[0070]

[0071] <Example 9> Sequel metagenome sequencing

[0072] The V1-V9 region of the 16S rRNA gene in the simulated community was amplified using a mixture of primers containing 27F and 1492R (V1-V9) (Fig. 1 and Table 3). The 27F primer mixture had an 8:2 ratio of 27F to bif27F. The following PCR conditions were used for amplification: initial denaturation at 95°C for 3 min, followed by 20 cycles of primer annealing at 55°C for 30 s, extension at 72°C for 30 s, and a final extension at 72°C for 5 min. PCR products were analyzed using 1% agarose gel electrophoresis and visualized with a Gel Document system (BioRad, Heracle, CA, USA). AMpure PB beads (Beckman Coolter Genomics, USA) were used to clean PCR products (Pacific Biosciences, USA). After pooling equal amounts of purified products, short fragments (off-target products) were removed using AMpure PB beads (Beckman Coolter Genomics, USA). The quality and size of the library were verified by Bioanalyzer 2100 (Agilent, USA) using a DNA 7500 chip (Agilent, USA). The SMRTbell library was constructed and sequenced using the PacBio sequel system (Pacific Biosciences, USA) according to the manufacturer's instructions.

[0073]

[0074] <Example 10> MinION metagenome sequencing

[0075] In the simulated community, the V1-V2, V3, V4, V1-V3, and V1-V9 regions of the 16S rRNA gene were amplified using primers such as a mixture of 27F and 1492R (V1-V9) (Fig. 1 and Table 3). The 27F primer mixture was an 8:2 ratio of 27F to bif27F. The following PCR conditions were used for amplification: initial denaturation at 95°C for 3 min, followed by 20 cycles of primer annealing at 55°C for 30 s, extension at 72°C for 30 s, and a final extension at 72°C for 5 min. The PCR products were purified using the QIAquick PCR Purification Kit (Qiagen, USA). 45 μl of PCR product was used for the end prep process using dA-tailing by adding 3 μl of Ultra II End-Prep enzyme mixture and 7 μl of Ultra II End-Prep reaction buffer. A total of 50 μl of the barcode mixture was composed of 22.5 μl of the end prep fragment after cleaning up the end prep DNA using AMPure XP beads (Beckman Coolter Genomics, USA), 25 μl of Blunt / TA Ligase Master Mix, and 2.5 μl of each barcode (NB01-NB08) from the Native Barcode Kit (SQK-LSK108). After purifying the barcode-ligated library using AMPure XP beads (Beckman Coolter Genomics, USA), equimolar amounts of pooled barcode amplicons were added to 100 μl of a mixture consisting of 50 μl of adapter ligation, 20 μl of Native Barcode Adapter Mix (BAM), 20 μl of NEBNext Quick Ligation Reaction Buffer (5X), and 10 μl of Quick T4 DNA Ligase. Amplification was performed to enrich for DNA fragments containing adapter molecules. A 75 μl mixture was prepared for sample loading by adding 35 μl of RBF, 25.5 μl of LLB, 12 μl of prepared DNA library, and 2.5 μl of nuclease-free water.This library mixture was loaded onto a 1D flow cell and run for 24 h. Raw data were collected and base calling was performed using MinKNOW software (version 19.10.1).

[0076]

[0077] PlatformsPrimersSequence (5' - 3')NoteMiSeqmiseq27FTCGTCGGCAGCGTC-AGATGTGTATAAGAGACAG-AGAGTTTGATCMTGGCTCAGGenomic DNAampliconmiseqbif27FTCGTCGGCAGCGTC-AGATGTGTATAAGAGACAG-GGGTTCGATTCTGGCTCAGmiseq337RGTCTCGTGGGCTCGG-AGATGTGTATAAGAGACAG-CTGCWGCCTCCCGTAGGAGTCmiseq337FTCGTCGGCAGCGTC-AGATGTGTATAAGAGACAG-GACTCCTACGGGAGGCWGCAGmiseq518RGTCTCGTGGGCTCGG-AGATGTGTATAAGAGACAG-CGTATTACCGCGGCTGCTGGmiseq518FTCGTCGGCAGCGTC-AGATGTGTATAAGAGACAG-CCAGCAGCCGCGGTAATACGmiseq800RGTCTCGTGGGCTCGG-AGATGTGTATAAGAGACAG-TACCAGGGTATCTAATCCI5 forward primerAATGATACGGCGACCACCGAGATCTACAC-XXXXXXXX-TCGTCGGCAGCGTCNextera barcodeI7 reverse primerCAAGCAGAAGACGGCATACGAGAT-XXXXXXXX-GTCTCGTGGGCTCGGIonTorrentMGI-2000SequelMinION27FAGAGTTTGATCMTGGCTCAGGenomic DNAampliconbif27FGGGTTCGATTCTGGCTCAG337FGACTCCTACGGGAGGCWGCAG337RCTGCWGCCTCCCGTAGGAGTC518FCCAGCAGCCGCGGTAATACG518RCGTATTACCGCGGCTGCTGG800RTACCAGGGTATCTAATCC1492RTACGGYTACCTTGTTACGACTT

[0078] <실시예 11> MOTHUR 파이프라인에 기반한 NGS 데이타 분석

[0079] Raw data from five platforms were analyzed using the MOTHUR pipeline according to the MOTHUR SOP (Fig. 2). From the results from the five platforms, all datasets were subsampled to 10,000 reads for comparative analysis. The raw reads underwent a quality control step, which was simplified using the "unique.seqs" command. These reads were then aligned to the EzBioCloud 16S database for MOTHUR using the "align.seqs" command. This database contains 66,303 manually curated rRNA gene sequences at the species / type and subspecies levels and serves as a de facto standard for taxonomic studies of prokaryotes. The aligned reads were trimmed using the "screen.seqs" command, and duplicate reads were removed using the "unique.seqs" and "filter.seqs" commands. Chimeric reads were also removed during the analysis using the "chimera.vseq" command. The "classify.seqs" command assigned taxonomically filtered reads using the Eztaxon-e database as a reference. The relative abundance for each phylotype was determined using the "make.shared" and "classify.otu" commands, which calculated taxonomic ranks for each phylotype. The "rarefaction.single" command was used to generate rarefaction curves. The sequencing error rate was calculated using the "seq.error" command.

[0080]

[0081] <Example 12> Microbiome analysis

[0082] Heatmaps of the simulated clusters were created to visualize taxonomic clusters using the Shiny Heatmap Online (http: / / shinyheatmap.com) with Euclidean distances. The R Shiny web server was used to run this program. Additionally, microbial profiles were visualized for principal component analysis (PCA) using Perseus (version 1.6.14.0). The proportion of specific strains in the original clusters and the sequencing results were used to define the relative difference value for each strain. To compare the bias between the simulated clusters, a bias index (BI) value was developed. The BI value is defined as the normalized Euclidean distance between the sequencing results and the original simulated clusters, based on the relative difference in the proportion of each component. Positive values ​​of relative difference in the sequencing results indicated overrepresentation, whereas negative values ​​of relative difference indicated underrepresentation. A specific strain was considered significantly over- or under-represented in the sequencing results if its absolute value was 0.5 or greater. If the BI converges to 0, the sequenced mock populations are comparable to the original mock populations. Higher BI values ​​may indicate greater bias in microbial profiling.

[0083]

[0084]

[0085] N: total number of strains in one simulated colony, xi: original proportion of each strain, bar(x)i: proportion of profile strains.

[0086]

[0087] <Example 13> Reference-based bias correction model

[0088] To correct for sequencing bias, a reference-based bias correction model was developed. This model considers the simulated population H as a reference and calculates the PCR efficiency using the sequence reads from the simulated H, the PCR cycles (t) used for amplification, and the virtual PCR efficiency. The resulting PCR efficiency was applied to simulated A to G to correct for bias. The relationship between the raw sequence reads and the resulting sequence reads can be defined as follows. The concentration of each strain (N0) is known in the reference material. The reference material is amplified for a specific cycle number (t) and subsequently sequenced (N t ) can be calculated based on the amplification efficiency. Since the sequence reads originate from a specific portion of the total DNA in the reference material, the source of the entire sequence reads can be estimated by assuming a uniform amplification efficiency across the entire DNA. In the examples below, the overall amplification efficiency is assumed to be 0.9. The table also presents the ratio of each strain in the reference material, so that the amplification efficiency for each strain can be calculated.

[0089]

[0090]

[0091] N t : t: Number of sequences after clustering amplification cycles, N0: Initial number of sequences before amplification, t: Number of cycles, E: Amplification efficiency.

[0092]

[0093] <Test Example 1> Genomic DNA Quantification

[0094] Genomic DNA from eight strains was quantified in terms of concentration and copy number using the Evagreen assay using a Quantus fluorometer (Promega, USA) and ddPCR (BioRad, Heracle, CA, USA). Templates were serially diluted 10-fold, and one template was used across the dynamic range to determine copy number.

[0095] Lactobacillus acidophilus, Limosilactobacillus fermentum, Lactobacillus gasseri, Lacticaseibacillus paracaseisubsp.paracasei, Limosilactobacillus reuteri, Bifidobacterium animalissubsp.lactis, Bifidobacterium breve, and Lactococcus lactissubsp.lactis were 1.2, 1.3, 1.5, 1.4, 1.4, 1.5, 1.6, and 1.9 per 1 ng / μl, respectively. In addition, the copy numbers of the strains were 6.7 x 10 per 1 μl, respectively. 6 , 6.5 x 10 6 , 2.3 x 10 7 , 5.2 x 10 6 , 3.8 x 10 6 , 6.9 x 10 6 , 4.9 x 10 6 , and 2.8 x 10 6 (Table 4).

[0096]

[0097] StrainReplication of copy number (1μL)Average copynumber (1μL)SD ±1st2nd3rdLactobacillus acidophilus814645550670133.7Limosilactobacillus fermentum805608536650139.1Lactobacillus gasseri141311658181132298.7Lacticaseibacillus paracaseisubsp.paracasei689497384523154.4Limosilactobacillus reuteri581243317380177.6Bifidobacterium animalissubsp. lactis846628585686139.8Bifidobacterium breve652391419487143.3Lactococcus lactissubsp. lactis762469457563172.7

[0098] <시험예 2> 모의 군집의 조성

[0099] The mock communities were composed of the proportion of each quantified gDNA based on copy number (Fig. 2 and Table 5). The eight mock communities were composed as follows: In mock A, 18.6% L. acidophilus, 18.45% L. fermentum, 17.48% L. gasseri, 20.18% L. paracaseisubsp. paracasei, and 25.3% L. reuteri were present. In mock B, 46.55% B. animalissubsp. lactis and 53.45% B. breve were present. In mock C, 16.62% L. acidophilus, 16.79% L. fermentum, 15.98% L. gasseri, 18.16% L. paracaseisubsp. paracasei, 15.32% L. reuteri, and 17.13% L. lactis were present. In simulation D, 30.73% B. animalissubsp. lactis, 35.14% B. breve, and 34.13% L. lactis were present. In simulation E, 33.41% L. acidophilus, 26.73% L. fermentum, 19.12% L. gasseri, 14.57% L. paracaseisubsp. paracasei, and 6.16% L. reuteri were present. Simulation F contained 56.64% B. animalissubsp. lactis and 43.36% B. breve. In simulation G, 8.11% L. acidophilus, 8.08% L. fermentum, 7.78% L. gasseri, 8.88% L. paracaseisubsp. paracasei, 7.53% L. reuteri, 16.99% B. animalissubsp. lactis, 19.43% B. breve, and 23.21% L. lactis were present. In the simulated H, 12.56% L. acidophilus, 12.51% L. fermentum, 11.96% L. gasseri, 13.63% L. paracaseisubsp.paracasei, 11.58% L. reuteri, 11.55% B. animalis subsp. lactis, 13.3% B. breve, and 12.91% L. lactis were present.

[0100]

[0101] MockCommunityStrainTotal copynumber(copies / μL)Ratio(%)ALactobacillus acidophilus1.40E+0818.60Limosilactobacillus fermentum1.39E+0818.45Lactobacillus gasseri1.31E+0817.48Lacticaseibacillus paracasei1.52E+0820.18Limosilactobacillus reuteri1.90E+0825.30BBifidobacterium animalissubsp.2.58E+0846.55Bifidobacterium breve2.96E+0853.45CLactobacillus acidophilus1.39E+0816.62Limosilactobacillus fermentum1.40E+0816.79Lactobacillus gasseri1.34E+0815.98Lacticaseibacillus paracasei1.52E+0818.16Limosilactobacillus reuteri1.28E+0815.32Lactococcus lactissubsp. lactis1.43E+0817.13DBifidobacterium animalissubsp.1.29E+0830.73Bifidobacterium breve1.48E+0835.14Lactococcus lactissubsp. lactis1.44E+0834.13ELactobacillus acidophilus6.96E+0833.41Limosilactobacillus fermentum5.57E+0826.73Lactobacillus gasseri3.98E+0819.12Lacticaseibacillus paracasei3.03E+0814.57Limosilactobacillus reuteri1.28E+086.16FBifidobacterium animalissubsp.3.87E+0856.64Bifidobacterium breve2.96E+0843.36GLactobacillus acidophilus3.83E+088.11Limosilactobacillus fermentum3.81E+088.08Lactobacillus gasseri3.67E+087.78Lacticaseibacillus paracasei4.19E+088.88Limosilactobacillus reuteri3.55E+087.53Bifidobacterium animalissubsp.8.02E+0816.99Bifidobacterium breve9.17E+0819.43Lactococcus lactissubsp. lactis1.10E+0823.21HLactobacillus acidophilus1.40E+0812.56Limosilactobacillus fermentum1.39E+0812.51Lactobacillus gasseri1.33E+0811.96Lacticaseibacillus paracasei1.52E+0813.63Limosilactobacillus reuteri1.29E+0811.58Bifidobacterium animalissubsp.1.28E+0811.55Bifidobacterium breve1.48E+0813.30Lactococcus lactissubsp. lactis1.44E+0812.91.

[0102] <시험예 3> 시퀀스 읽기의 통계(Statistics of sequence reads)

[0103] A subsample of 10,000 reads was analyzed using the MOTHUR pipeline. MiSeq sequencing yielded 70,876 (mean) ± 19,543 (SD), 75,451 ± 18,013, 75,805 ± 18,028, and 56,793 ± 19,029 reads per sample for the V1-V2, V3, V4, and V1-V3 regions, respectively, and 69,731 ± 19,939 reads per sample for the full dataset. IonTorrent sequencing yielded 130,742 ± 27,042, 213,401 ± 45,200, 140,429 ± 16,299, and 126,217 ± 13,493 reads per sample for the V1-V2, V3, V4, and V1-V3 regions, respectively, and 152,697 ± 45,254 reads per sample for the full dataset. MGISeq-2000 sequencing yielded 1,379,653 ± 484,553 reads per sample for the V3 region. Sequel and MinION sequencing generated 20,556 ± 5,034 and 81,000 ± 32,726.14 reads per sample for the V1-V9 region, respectively. The average accuracies of MiSeq, IonTorrent, MGIseq-2000, Sequel, and MinION were 99.720%, 99.825%, 99.996%, 99.708%, and 99.453%, respectively (Tables 6 to 10).

[0104]

[0105] MiSeqSamplesAnalyzedreadsSubsampleLowqualityreadsChimerareadsNonbacteria16SrRNAreadsError rate(%)A156,37110,0001,7522,32702.36E-03A249,53897472401.02E-02A360,8991,13696001.37E-04A434,0887,86555912.47E-04B147,2941,4832,60503.04E-06B256,54178340705.74E-03B359,92495957808.73E-05B441,4327,49452201.22E-02C1107,3701,15299004.99E-04C2101,82966031009.45E-03C3104,99388089108.37E-05C478,8827,88339016.96E-06D188,3068741,03100.00E+00D289,69366413005.68E-03D397,60371152404.01E-05D490,7677,98313602.63E-05E160,8941,22384103.99E-04E276,0331,03431601.00E-02E368,8461,00444505.34E-05E451,9147,79430115.32E-05F171,4709791,41701.06E-03F285,11966112805.54E-03F376,54371215508.37E-05F456,22610,0007,16131908.72E-03G165,6961,47096301.06E-05G270,60875939107.94E-03G369,1331,08978708.16E-05G450,0637,52934314.71E-05H169,6051,2401,02703.59E-06H274,24282036708.63E-03H368,4961,10478206.72E-05H450,9747,76041716.61E-05

[0106] IonTorrentSamplesAnalyzedreadsSubsampleLow qualityreadsChimerareadsNon bacteria16SrRNAreadsError rate(%)A1137,46310,00089521902.30E-03A2209,73741537603.77E-04A3142,43363011803.91E-03A4121,78672690709.00E-05B1126,87995010301.20E-04B2230,74636320606.81E-04B3125,7636828506.85E-03B4114,10980256504.66E-04C1138,31189518802.70E-03C2162,32142433403.46E-04C3144,49964013104.87E-03C4120,84170548205.38E-04D1137,3886599802.20E-04D2274,62037019523.87E-04D3133,3956389106.01E-03D4142,43356529005.43E-04E1134,53288812009.93E-04E2203,61641421103.91E-04E3148,65262911704.10E-03E4129,25069659305.69E-05F1169,05490810109.74E-05F2196,30738320316.08E-04F3128,64766011606.47E-03F4119,20010,00077258703.73E-04G1108,68174128201.15E-04G2217,77040022214.11E-04G3154,07965912505.80E-03G4136,49365749305.20E-04H193,62882318501.28E-04H2212,09138618503.83E-04H3145,9686519404.74E-03H4125,62070160605.26E-04

[0107] MGIseq-2000SamplesAnalyzedreadsSubsampleLow qualityreadsChimerareadsNonbacteria16SrRNAreadsError rate(%)A11,416,16510,0001,15247105.14E-05A21,399,0921,13342303.29E-05A31,343,2831,18550806.49E-05A41,339,3021,14431503.31E-05B11,415,7111,24333703.31E-05B21,477,5901,25526904.85E-05B31,212,5921,16428302.44E-05H41,433,4921,08626103.42E-05

[0108] SequelSamplesAnalyzedreadsSubsampleLow qualityreadsChimerareadsNon bacteria16SrRNA readsError rate(%)A125,33010,00048555403.12E-03A216,60442937603.28E-03A319,27650375502.68E-03A425,37328124702.70E-03B117,72650567203.23E-03B220,27242121403.14E-03B319,63728760502.46E-03H420,22627482902.73E-03

[0109] MinIONSamplesAnalyzedreadsSubsampleLow qualityreadsChimerareadsNon bacteria16SrRNA readsError rate(%)A192,00010,000961015.42E-02A2116,000802005.80E-02A312,0001,027005.44E-02A4104,000860005.45E-02B152,0001,023005.43E-02B2112,000652005.85E-02B372,000957005.47E-02H488,000945005.52E-02

[0110] <Experimental Example 4> Microbial profiling of the short-read sequencing

[0111] In the MiSeq sequencer, the four regions (V1-V2, V3, V4, and V1-V3) of the simulated community were sequenced in triplicate (Fig. 1). In the results of simulation A, the proportions of L. acidophilus in the four regions (8.83%, 8.64%, 7.76%, and 6.59%) were significantly underrepresented compared to the original proportion of 18.6% for this strain. In contrast, the proportions of L. gasseri in the four regions (26.43%, 22.87%, 29.29%, and 27.71%) were significantly overrepresented compared to the original proportion of 17.48% for this strain. In this experimental community, L. reuteri is most similar to the original composition. In the results of simulation B, B. The four regions of B. breve, 46.45%, 44.85%, 25.9%, and 15.79%, were significantly underrepresented compared to the original 53.45% ratio of this strain. In contrast, the four regions of B. animalissubsp. lactis, 53.55%, 55.15%, 74.1%, and 84.21%, were significantly overrepresented compared to the original 46.55% ratio of this strain. In the results of simulated C, the four regions of L. acidophilus, 5.54%, 5.29%, 4.8%, and 2.26%, were significantly underrepresented compared to the original 16.62% ratio of this strain. In contrast, L. lactis in the four regions at 27.46%, 27.5%, 31.61%, and 40.88% were significantly overrepresented compared to the original 17.13% proportion of this strain. In this experimental community, L. routeri is most similar to the original composition. In the results of simulation D, B. breve in the four regions at 29.35%, 20.88%, 15.57%, and 9.16% were significantly underrepresented compared to the original 35.14% proportion of this strain. In contrast, L. lactis in the four regions at 35.61%, 49.88%, 53.05%, and 52.74% were significantly overrepresented compared to the original 34.13% proportion of this strain. In the results of simulation E, L. The ratios of 15.66%, 15.46%, 14.85% and 7.32% in the four regions of acidophilus were 33%, 14.85% and 7.32%, respectively, compared to the original 33% of this strain.was significantly underrepresented compared to the original 41% ratio of this strain. In contrast, the ratios of 9.14%, 10.11%, 9.93%, and 7.11% in the four regions of L. routeri were significantly overrepresented compared to the original 6.16% ratio of this strain. Except for the V1-V3 region, L. gasseri was most similar to the original composition in all regions. In the results of simulation F, the ratios of 38.26%, 36.88%, 24.09%, and 14.77% in the four regions of B. breve were significantly underrepresented compared to the original 43.36% ratio of this strain. In contrast, the ratios of 61.74%, 63.12%, 75.91%, and 85.23% in the four regions of B. animalissubsp. lactis were significantly overrepresented compared to the original 56.64% ratio of this strain. In the results of simulation G, L. acidophilus in the four regions at 1.89%, 2.74%, 2.87%, and 1.14% were significantly underrepresented compared to the original 8.11% proportion of this strain. In contrast, L. lactis in the four regions at 23.59%, 32.13%, 34.56%, and 35.35% were significantly overrepresented compared to the original 23.21% proportion of this strain. In this experimental community, L. fermentum is most similar to the original composition. In the results of simulated H, L. acidophilus in the four regions at 3.08%, 3.96%, 4.06%, and 1.33% were significantly underrepresented compared to the original 12.56% proportion of this strain. In contrast, L. The proportions of 15.12%, 21.96%, 23.81%, and 27.46% in the four regions of L. lactis were significantly overrepresented compared to the original 12.91% proportion of this strain. L. routeri was relatively similar to the original composition in these simulated communities (Fig. 3 and Table 11). In general, L. acidophilus was underrepresented, and L. lactis was overrepresented in all simulated communities from the microbial profiles of the MiSeq platform. In all simulated communities, the BI values ​​of the V1-V3 regions were relatively greater than those of the other regions (Table 12).This was supported by the results of PCA and heatmap analysis, which revealed that the V1-V3 region exhibited the greatest bias. While the profiles from the original region and other regions formed a single phylogenetic group in the phylogenetic tree, the profiles from the V1-V3 region formed a unique group in the heatmap analysis (Fig. 4). This clustering was also observed in the PCA results (Fig. 5).

[0112]

[0113] NGSPlatformMiSeqIonTorrentMGIseq-2000SequelMinIONMockcommunityStrainV1-V2(%)V3(%)V4(%)V1-V3(%)V1-V2(%)V3(%)V4(%)V1-V3(%)V3(%)V1-V9(%)AL. acidophilus8.88.67.86.67.17.46.812.88.313.813.3L. fermentum27.024.719.337.417.227.719.03.424.74.03.3L. gasseri26.422.929.327.721.220.228.540.722.948.950.3L. paracase15.218.819.88.120.418.521.434.316.831.330.6L. reuteri22.625.023.920.234.226.224.38.727.32.02.6BB.animalissubsp. Lactis53.655.174.184.244.445.357.145.358.052.960.8B.breve46.444.925.915.855.654.742.954.742.047.139.2CL. acidophilus5.55.34.82.34.64.44.16.710.36.55.2L. fermentum19.718.115.728.013.419.812.62.028.41.41.5L. gasseri16.816.518.210.115.314.318.122.620.125.022.2L. paracasei10.813.212.54.013.812.213.917.713.816.113.8L. reuteri19.719.417.214.826.718.616.15.512.80.41.5Lactococcus lactis27.527.531.640.926.330.635.245.614.650.655.8DB.animalissubsp. lactis35.029.231.438.112.218.820.511.252.225.720.5B.breve29.420.915.69.214.520.513.912.826.619.313.2Lactococcus lactis35.649.953.052.773.360.765.675.921.155.066.4EL. acidophilus15.715.514.97.315.913.613.125.46.822.523.2L. fermentum38.835.931.061.930.139.828.15.526.35.14.7L. gasseri24.723.527.718.925.523.333.043.120.649.250.1L. paracasei11.715.116.54.815.513.416.623.117.922.321.0L. reuteri9.110.19.97.113.09.99.12.828.40.80.9FB.animalissubsp. lactis61.763.175.985.255.357.366.856.562.763.770.8B.breve38.336.924.114.844.742.733.243.537.336.329.2GL. acidophilus1.92.72.91.12.42.32.23.12.62.62.0L. fermentum7.89.48.38.37.110.17.81.110.30.60.4L. gasseri6.68.310.94.97.87.410.111.17.79.26.2L. paracasei3.16.77.12.56.16.06.57.86.96.23.7L. reuteri8.010.49.75.514.710.08.83.310.60.20.4B.animalissubsp. lactis27.117.717.732.813.111.711.59.415.812.017.5B.breve21.912.79.09.514.212.58.510.910.710.111.1Lactococcus lactis23.632.134.635.434.739.944.653.235.359.058.7HL. acidophilus3.14.04.11.34.03.43.55.34.04.83.8L. fermentum11.913.611.316.911.215.712.91.714.41.40.9L. gasseri9.612.414.76.711.811.015.317.212.418.714.5L. paracasei5.89.910.42.99.69.411.713.49.912.49.4L. reuteri12.114.713.68.623.416.314.04.915.41.30.8B.animalissubsp. lactis23.913.814.729.29.89.18.89.813.910.116.1B.breve18.59.77.46.811.29.86.411.18.28.311.6Lactococcus lactis15.122.023.827.519.125.327.436.621.943.043.0.

[0114] MiSeqIonTorrentMGIseq-2000SequelMinIONMockV1-V2V3V4V1-V3V1-V2V3V4V1-V3V3V1-V9A0.2270.1770.2240.3720.1860.2010.2240.4650.1850.5630.579B0.0500.0610.1960.2680.0150.0090.0750.0090.0810.0450.102C0.2610.2480.2910.4910.3020.2910.3370.5290.2230.6270.681D0.0550.1540.1970.2370.3560.2420.2880.3800.2080.1950.295E0.2300.2330.2410.4190.3200.2520.2680.4260.9250.5160.519F0.0370.0470.1400.2080.0100.0050.0740.0010.0440.0510.103G0.2970.2420.2780.4230.3440.3090.3520.4800.2780.5630.575H0.3770.2770.3150.5940.3430.3440.3750.5610.2900.7020.705

[0115] In the IonTorrent sequencer, the four regions (V1-V2, V3, V4, and V1-V3) of the simulated community were sequenced in triplicate (Fig. 1). In the results of simulation A, the proportions of L. acidophilus in the four regions were significantly underrepresented at 7.12%, 7.38%, 6.82%, and 12.78%, respectively, compared to the original 18.6% proportion of this strain. In contrast, the proportions of L. gasseri in the four regions were significantly overrepresented at 21.18%, 20.21%, 28.5%, and 40.74%, respectively, compared to the original 17.48% proportion of this strain. In this experimental community, L. fermentum and L. reuteri were also underrepresented compared to the original composition. In the results of simulation B, the V4 region showed B. breve in this region was underrepresented at 42.94% compared to the original 53.45% ratio of this strain. In contrast, the 57.06% ratio of B. animalissubsp. lactis in the four regions was overrepresented compared to the original 46.55% ratio of this strain. In addition, the strain ratios in regions V1-V2, V3, and V4 showed relatively similarity compared to the original composition. In the results of simulated C, the ratios of 4.57%, 4.42%, 4.1%, and 6.72% in the four regions of L. acidophilus were significantly underrepresented compared to the original 16.62% ratio of this strain. In contrast, the ratios of 26.32%, 30.64%, 35.2%, and 45.57% in the four regions of L. lactis were significantly overrepresented compared to the original 17.13% ratio of this strain. In this experimental community, L. gasseri is the most similar in all sequencing results. In the results of simulation D, the ratios of 12.25%, 18.8%, 20.46%, and 11.23% in the four regions of B. animalissubsp. lactis were significantly underrepresented compared to the original 30.73% ratio of this strain. In contrast, the ratios of 73.26%, 60.74%, 65.64%, and 75.93% in the four regions of L. lactis were significantly overrepresented compared to the original 34.13% ratio of this strain. In the results of simulation E, L.acidophilus, with proportions of 15.88%, 13.63%, 13.13%, and 25.43% in the four regions, were significantly underrepresented compared to the original 33.41% proportion of this strain. In contrast, L. gasseri, with proportions of 25.51%, 23.28%, 33.03%, and 43.12% in the four regions, were significantly overrepresented compared to the original 19.12% proportion of this strain. In the results of the simulated F, B. breve, with proportions of 44.73%, 42.68%, 33.19%, and 43.49% in the four regions, were significantly underrepresented compared to the original 43.36% proportion of this strain. In contrast, B. The proportions of 55.27%, 57.32%, 66.81%, and 56.51% in the four regions of animalissubsp. lactis were significantly overrepresented compared to the original 56.64% proportion of this strain. In the results of simulation G, the proportions of 2.36%, 2.31%, 2.2%, and 3.14% in the four regions of L. acidophilus were significantly underrepresented compared to the original 8.11% proportion of this strain. In contrast, the proportions of 34.66%, 39.92%, 44.59%, and 53.23% in the four regions of L. lactis were significantly overrepresented compared to the original 23.21% proportion of this strain. In the results of simulation H, L. acidophilus in the four regions at 3.96%, 3.45%, 3.55%, and 5.32% were significantly underrepresented compared to the original 12.56% proportion of this strain. In contrast, L. lactis in the four regions at 19.13%, 25.31%, 27.43%, and 36.62% were significantly overrepresented compared to the original 12.91% proportion of this strain (Fig. 3 and Table 11). In general, L. acidophilus was underrepresented, whereas L. lactis was overrepresented in all mock communities from the microbial profiles of the IonTorrent platform. However, mocks B and D containing only Bifidobacterium species were relatively well represented compared to the MiSeq platform. In all mock communities, the BI values ​​of the V1-V3 regions were relatively larger than those of the other regions (Table 11).This indicates that the V1-V3 region exhibited the greatest bias, a finding supported by the PCA and heatmap results. While the profiles from the original region and other regions formed a single phylogenetic group in the phylogenetic tree, the profiles from the V1-V3 region formed a unique group in the heatmap analysis (Figure 6). This clustering was also observed in the PCA results (Figure 5).

[0116] In the MGIseq-2000 sequencer, only the V3 region of the simulated population was sequenced three times (Fig. 1). In the simulated A, C, E, G, and H results, the proportions of L. fermentum were significantly underrepresented by 3.96%, 1.41%, 5.06%, 0.62%, and 1.4%, respectively, compared to the original proportions of this strain (18.45%, 16.79%, 26.73%, 8.08%, and 12.51%). In contrast, the simulated C, D, G, and H proportions of L. lactis were significantly overrepresented by 50.57%, 55.03%, 59.05%, and 43%, respectively, compared to the original proportions of this strain (17.13%, 34.13%, 23.21%, and 12.91%). In the PCA results, points A, C, D, E, G, and H were separated from the original and other sample points. In all simulated clusters, the BI value of simulated A was relatively larger than that of other regions (Table 12). This indicates that simulated A exhibited the greatest bias, which was also supported by the PCA and heatmap results. While the profiles from the original and other regions formed a single phylogenetic group in the phylogenetic tree, the profiles from simulated E formed a unique group in the heatmap analysis (Figure 7). This clustering was also observed in the PCA results (Figure 5).

[0117]

[0118] <Experimental Example 5> Microbial profiling of the long-read sequencing

[0119] Only the V1-V9 region of the simulated cluster was sequenced in triplicate using Sequel sequencer (Fig. 1). In the results of simulated A, C, E, G, and H, the proportions of L. fermentum (3.96%, 1.41%, 5.06%, 0.62%, and 1.4%) were significantly underrepresented compared to the original proportions of this strain (18.45%, 16.79%, 26.73%, 8.08%, and 12.51%). In contrast, the proportions of L. lactis (50.57%, 55.03%, 59.05%, and 43% in simulated C, D, G, and H) were significantly overrepresented compared to the original proportions of this strain (17.13%, 34.13%, 23.21%, and 12.91%). In the PCA results, points A, C, D, E, G, and H were separated from the original and other sample points. In all simulated clusters, the BI value of simulated A was relatively larger than that of other regions (Table 11). This indicates that simulated A exhibited the greatest bias, which was also supported by the PCA and heatmap results. While the profiles from the original and other regions formed a single clade in the phylogenetic tree, the profiles from all samples in the Sequel platform formed a unique group in the heatmap analysis (Figure 7). This clustering was also observed in the PCA results (Figure 5).

[0120] Only the V1-V9 region of the simulated cluster was sequenced in triplicate using the MinION sequencer (Fig. 1). In the results of simulated A, C, E, G, and H, the proportions of L. fermentum were significantly underrepresented (3.27%, 1.51%, 4.71%, 0.37%, and 0.9%) compared to the original proportions of this strain (18.45%, 16.79%, 26.73%, 8.08%, and 12.51%). In contrast, the proportions of L. lactis in simulated C, D, G, and H were significantly overrepresented (55.8%, 66.35%, 58.75%, and 42.97%) compared to the original proportions of this strain (17.13%, 34.13%, 23.21%, and 12.91%). In the PCA results, points A, C, E, G, and H were separated from the original and other sample points. In all simulated clusters, the BI value of simulated A was relatively larger than that of other regions (Table 12). This indicates that simulated A exhibited the greatest bias, which was also supported by the PCA and heatmap results. While the profiles from the original and other regions formed a single clade in the phylogenetic tree, the profiles from all samples in the MinION platform formed a unique group in the heatmap analysis (Figure 7). This clustering was also observed in the PCA results (Figure 5).

[0121]

[0122] <Experimental Example 6> Comparison of Profiling Across All NGS Platforms

[0123] Five NGS platforms were compared using microbial profiling in the V3 region (Fig. 3). The short-read sequencer-based MiSeq, IonTorrent, and MGIseq-2000 platforms outperformed the long-read sequencer-based Sequel and MinION platforms in simulations A, E, G, and H. MiSeq and Sequel platforms were more accurate in simulation B. MGIseq-2000 and IonTorrent platforms were more accurate in simulations C, D, and F, respectively. As a result, paired-end sequencing technologies such as the MiSeq and MGIseq-2000 platforms were less biased than single-read sequencing platforms. In contrast, microbial profiling on the Sequel II and MinION platforms was significantly biased in simulations A, C, E, G, and H. In general, L. acidophilus was significantly underrepresented, whereas L. lactis was often overrepresented in most amplicon regions and across NGS platforms. In addition, L. fermentum and L. routeri were significantly underrepresented compared to the original simulated ratios on the Sequel and MinION platforms.

[0124]

[0125] <Experimental Example 7> Reference-based bias correction

[0126] The microbial profiling ratio was a corrected reference-based bias-corrected model. From the MiSeq and IonTorrent sequencing results, relatively high bias was observed in the reads in the V1-V3 region. However, the corrected ratio was comparable to the original and other region ratios (Figure 8 and Table 13). Furthermore, reads from Sequel and MinION showed relatively high bias in NGS platform-dependent results. Nevertheless, the corrected ratio was similar to the original and other region ratios (Figure 8 and Table 13). This indicates that the developed model can correct for biased sequencing results, enabling the estimation of the initial microbial community composition. The corrected ratio was utilized for BI calculation (Table 14). The BI of the corrected ratio converged to 0, showing similarity to the original BI value. Moreover, the BI value of the corrected ratio was dramatically reduced in the long-read sequencer.

[0127]

[0128] NGS PlatformMiSeqIonTorrentMGIseq-2000SequelMinIONMockcommunityStrainV1-V2(%)V3(%)V4(%)V1-V3(%)V1-V2(%)V3(%)V4(%)V1-V3(%)V3(%)V1-V9(%)AL. acidophilus23.924.323.723.023.924.323.723.022.623.320.5L. fermentum17.918.218.919.617.918.218.919.618.924.321.4L. gasseri20.819.819.719.320.819.819.719.320.019.919.6L. paracase23.422.922.221.523.422.922.221.521.122.121.0L. reuteri14.014.715.516.614.014.715.516.617.410.317.5BB.animalissubsp. Lactis43.8643.1842.9042.7443.4944.4344.7844.7742.0944.2649.36B.breve56.1456.8257.1057.2656.5155.5755.2255.2357.9155.7450.64CL. acidophilus17.1417.0516.9916.7616.0816.7416.6217.1128.5221.8315.06L. fermentum15.8416.2416.6416.9016.2016.0516.3116.3821.6412.6318.35L. gasseri15.8116.1416.5016.7015.9916.3016.7616.9517.0519.4416.06L. paracasei18.7318.3417.8017.7319.4618.8218.7018.2817.0821.4317.56L. reuteri14.7515.0515.1815.2313.5713.8914.3014.588.565.7618.31Lactococcus lactis17.7317.1716.8916.6818.7118.1917.3216.707.1518.9114.65DB.animalissubsp. lactis25.3226.3427.3728.6917.6823.0625.8429.8244.4737.7929.63B.breve31.1032.9334.0535.5621.7926.2830.3133.2143.3140.0030.29Lactococcus lactis43.5840.7438.5835.7560.5350.6643.8536.9612.2322.2240.08EL. acidophilus37.036.436.235.736.836.835.535.419.435.233.6L. fermentum22.623.424.225.024.724.124.424.021.428.828.7L. gasseri17.917.918.018.218.218.719.620.217.818.518.3L. paracasei16.916.415.815.315.615.415.114.822.914.513.5L. reuteri5.55.85.95.94.65.05.45.618.53.05.9FB.animalissubsp. lactis52.3951.5551.0650.5753.6453.4054.2156.5647.9755.3260.42B.breve47.6148.4548.9449.4346.3646.6045.7943.4452.0344.6839.58GL. acidophilus10.6111.5011.6111.468.38.48.38.48.29.19.0L. fermentum10.9511.5811.4311.448.98.28.28.39.77.87.1L. gasseri10.3811.1811.2411.168.28.28.58.98.07.87.3L. paracasei10.8311.9812.1912.278.89.09.29.410.49.07.5L. reuteri10.1611.0310.8310.797.87.77.77.78.91.99.0B.animalissubsp. lactis18.1419.4619.4019.4915.815.815.716.414.518.417.6B.breve21.3123.2723.3023.3917.719.318.918.518.722.217.8Lactococcus lactis27.3328.5627.6826.6424.623.523.622.421.623.824.7.

[0129] MiSeqIonTorrentMGIseq-2000SequelMinIONMockV1-V2V3V4V1-V3V1-V2V3V4V1-V3V3V1-V9A0.1460.110.10.1890.1460.1390.1250.110.1020.1850.096B0.050.0630.0680.0710.0570.040.0330.0330.0830.0430.053C0.0180.0090.010.0120.0340.0250.0190.0180.2130.160.056D0.1740.1240.0870.0410.480.30.1770.0520.4080.220.113E0.0690.0540.0420.0310.0720.060.0420.040.5340.1290.03F0.0980.1170.1290.140.0690.0750.0560.0020.20.030.087G0.1060.1430.1420.1390.0250.0150.0190.0250.0510.1120.046

Claims

1) A step of quantifying the genomic DNA of the microbial community to be analyzed; 2) Step of constructing a mock community based on the ratio of quantified genomic DNA; 3) Step of confirming the genetic sequence of the simulated population using next generation sequencing; 4) A step of confirming the simulated cluster composition from the above 3); and 5) A method for analyzing a microbial community, comprising: a step of calculating the composition of a microbial community to be analyzed by correcting the simulated community composition of the above 4) using [Mathematical Formula 2]; [Equation 2] N t : t: Number of sequences after clustering amplification cycles, N0: Initial number of sequences before amplification, t: Number of cycles, E: Amplification efficiency. In claim 1, the microbial community to be analyzed is Lactobacillus acidophilus, Lacticaseibacillus casei, Lactobacillus delbrueckiisubsp.bulgaricus, Limosilactobacillus fermentum, Lactobacillus gasseri, Lactobacillus helveticus, Lacticaseibacillus paracaseisubsp. paracasei, Lactiplantibacillus plantarum, Lacticaseibacillus rhamnosus, At least one selected from the group consisting of Limosilactobacillus reuterisubsp.reuteri, Ligilactobacillus salivarius, Bifidobacterium animalissubsp.lactis, Bifidobacterium bifidum, Bifidobacterium breve, Bifidobacterium longumsubsp.longum, Streptococcus thermophilus, Lactococcus lactissubsp.lactis, Enterococcus faecalis and Enterococcus faecium. Microbial community analysis method characterized by In the first paragraph, the step of quantifying the DNA is characterized in that it is performed by at least one selected from the group consisting of real-time PCR, droplet digital PCR, or recombinase polymerase amplification. In the first paragraph, the next generation base sequence analysis method is characterized in that at least one is selected from the group consisting of MiSeq, Sequel II, IonTorrent, MGIseq-2000, and MinION. In the first paragraph, a method for analyzing a microbial community, characterized in that the genes of the simulated community are at least one selected from the group consisting of 16S rRNA, 18S rRNA, ITS, rpoB, polC, gyrB, mcrA, and nifH. 1) A step of quantifying the genomic DNA of the microbial community to be analyzed; 2) Step of constructing a mock community based on the ratio of quantified genomic DNA; 3) Step of confirming the 16S rRNA sequence of the simulated community using next generation sequencing; 4) A step of confirming the simulated cluster composition from the above 3); and 5) A computer program recorded on a computer-readable recording medium for executing a microbial community analysis method, including a step of calculating the composition of the microbial community to be analyzed by correcting the simulated community composition of the above 4) using [Mathematical Formula 2]; [Equation 2] N t : t: Number of sequences after clustering amplification cycles, N0: Initial number of sequences before amplification, t: Number of cycles, E: Amplification efficiency. A microbial community analysis system comprising at least one processor configured to execute computer-readable instructions, At least one processor, 1) Quantify the genomic DNA of the microbial community to be analyzed; 2) Create a mock community based on the proportion of quantified genomic DNA; 3) Confirm the 16S rRNA sequence of the simulated community using next generation sequencing; 4) Confirm the simulated cluster composition from the above 3); and 5) A microbial community analysis system that calculates the composition of the microbial community to be analyzed by correcting the simulated community composition of 4) above using [Mathematical Formula 2]. [Equation 2] N t : t: Number of sequences after clustering amplification cycles, N0: Initial number of sequences before amplification, t: Number of cycles, E: Amplification efficiency. 1) A step of quantifying the genomic DNA of the microbial community to be analyzed; 2) Step of constructing a mock community based on the ratio of quantified genomic DNA; 3) Step of confirming the 16S rRNA sequence of the simulated community using next generation sequencing; 4) A step of confirming the simulated cluster composition from the above 3); and 5) A method for identifying microorganisms, including a step of calculating the composition of the target microbial community by correcting the simulated community composition of the above 4) using [Mathematical Formula 2]; [Equation 2] N t : t: Number of sequences after clustering amplification cycles, N0: Initial number of sequences before amplification, t: Number of cycles, E: Amplification efficiency.

Citation Information

Patent Citations

  • Method for microbiome analysis

    EP2985350A1

  • Vegetable storage container

    KR1020210052897A