Methods of microbial profiling
Long-read sequencing and computational combining of 16s rRNA genes using LNA primers and optional chloroplast removal enhance microbial profiling accuracy and harmonization across diverse niches, addressing biases and ambiguities in existing methods.
Patent Information
- Application Number
- PCT/IL2025/050563
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-01
- Filing Date
- 2025-07-01
- Publication Date
- 2026-01-08
AI Technical Summary
Existing microbial profiling methods using short-read sequencing and operational taxonomic units (OTUs) face biases due to primer-dependent amplification and bioinformatic decisions, leading to inconsistent taxonomies and limited harmonization across studies, especially in non-human niches, and current database-free approaches fail to harmonize results when different variable regions are amplified.
A method involving long-read sequencing of full-length 16s rRNA genes using locked nucleic acid (LNA) primers, followed by computational combining of short-read amplicons, to generate a comprehensive database for high-resolution microbial profiling, with optional chloroplast removal using a blocking oligo or CRISPR-gRNA complex.
Provides accurate, unbiased, and harmonized microbial profiling across diverse niches, including non-human environments, with improved phylogenetic resolution and detection of rare bacteria, overcoming alignment ambiguities and database limitations.
Smart Images

Figure IL2025050563_08012026_PF_FP_ABST
Abstract
Description
METHODS OF MICROBIAL PROFILINGCROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority of Israeli Patent Application No. 314064, filed July 1, 2024, the contents of which are all incorporated herein by reference in their entirety.REFERENCE TO AN ELECTRONIC SEQUENCE LISTING
[0002] The contents ofthe electronic sequence listing (YEDA-OPEN-P-051-PCT.xml; Size: 40,031 bytes; and Date of Creation: June 24, 2025) is herein incorporated by reference in its entirety.FIELD OF INVENTION
[0003] The present invention is in the field of microbial profiling.BACKGROUND OF THE INVENTION
[0004] During the last decade, in the advent of massively parallel sequencing, the study and exploration of microbiomes has become an integrative science that spans multiple disciplines and research fields. Medicine, environmental studies, agriculture, and even behavioral studies have associated specific phenotypes with bacteria. A majority of these studies rely on sequencing predefined variable regions of the 16s rRNA gene using short reads . However, it has been widely shown that results of such studies depend on the amplified region due to biases introduced by the specific primers applied, and on specific bioinformatic decisions and practices, hence limiting the ability to harmonize findings across studies.
[0005] Early microbiome studies used operational taxonomic units (OTUs) either in “closed reference” or “open reference” forms. The former aligns each short read to a database of full-length 16s rRNA gene (e.g., SILVA or Greengenes), which contain millions of small ribosomal subunits (SSUs). Such alignment implicitly “cleans” sequencing errors and allows harmonizing results from different experiments, using both the 16s rRNA gene sequence itself and its assigned taxonomy, which can also provide ecological analysis at highertaxonomic levels. However, there are two inherent caveats to “closed reference” OTU. First, it implicitly assumes that the SSU database is comprehensive, an assumption that probably holds only for human-associated niches. Second, reads are often aligned to multiple 16s rRNA gene that share the same short amplicon, and such ambiguity may result in inconsistent taxonomies. To allow database-free assignments researchers then used “openreference” OTUs, i.e., experimental reads were taken as is, which then led to the development of denoising methods, e.g., Deblur and DADA2. The output of such methods, referred to as amplicon sequence variants (ASVs), use error models of Illumina platforms to “denoise” reads, i.e., create a single ASV from different reads that harbor different sequencing errors. It was shown that ASVs can be compared across studies even when using different denoising algorithms, allowing a database-free approach. However, such ASV- based cross-study harmonization is not possible when studies amplify different variable regions, thus results still depend on their identity. In addition, the issue of ambiguous assignment to SSUs remains unsolved.
[0006] To circumvent the necessity to commit to a specific primer pair, the Short MUlti- Region Framework (SMURF) approach was created (Fuks, et al., “Combining 16s rRNA gene variable regions enables high-resolution microbial community profiling”, Microbiome. 2018 Jan 26;6(1), the contents of which are hereby incorporated by reference herein in their entirety). In SMURF each sample is amplified by any number of primer pairs, and reads from different regions are computationally combined to provide coherent high-resolution microbial profiling. To perform such integration, SMURF relies on 16s rRNA gene databases, making the approach suboptimal for less explored niches. A new method that can be used on novel niches that are not associated with humans is greatly needed.SUMMARY OF THE INVENTION
[0007] The present invention provides methods of profiling a microbial community within a biological niche comprising: receiving DNA extracted from samples from the niche, amplifying 16s rRNA genes from the samples to produce long amplicons, sequencing the long amplicons with long -read sequencing, generating a 16s rRNA database, amplifying and sequencing 16s rRNA genes from a sample using deep sequencing to produce reads and computationally combining the reads based on the database to produce full-length 16s rRNA sequences are provided. ENA primer pairs and kits comprising those LNA primer pairs are also provided.
[0008] According to a first aspect, there is provided a method of profiling a microbial community within a biological niche, the method comprising: a. receiving a pool of DNA comprising DNA extracted from a plurality of samples from the biological niche; b. amplifying 16s rRNA genes within the pool of DNA to produce long amplicons wherein each long amplicon comprises at least 80% of the sequence of a 16s rRNA gene; c. sequencing the long amplicons with a long-read sequencing technology to receive full-length 16s rRNA sequences comprising a contiguous at least 80% of a 16s rRNA gene: d. generating a 16s rRNA database comprising the received full-length 16s rRNA sequences; e. amplifying and sequencing 16s rRNA sequences from a sample from the biological niche using short deep-sequencing technology with a plurality of primer pairs wherein each primer pair produces a short amplicon of not greater than 500 nucleotides and amplicon reads of not longer than 300 nucleotides; and f. computationally combine the amplicon reads to produce full-length 16s rRNA sequences based on the generated 16s rRNA database, thereby identifying unique 16s rRNA genes within the sample from the biological niche; thereby profiling a microbial community within a biological niche.
[0009] According to another aspect, there is provided a method of profiling a microbial community within a plant niche, the method comprising: a. receiving a pool of DNA comprising DNA extracted from a plurality of samples from the plant niche; b. amplifying 16s rRNA genes within the pool of DNA to produce long amplicons wherein each long amplicon comprises at least 80% of the sequence of a 16s rRNA gene;c. sequencing the long amplicons with a long-read sequencing technology to receive full-length 16s rRNA sequences comprising a contiguous at least 80% of a 16s rRNA gene: d. generating a 16s rRNA database comprising the received full-length 16s rRNA sequences; e. amplifying and sequencing 16s rRNA sequences from a sample from the plant niche using short deep-sequencing technology with a plurality of primer pairs wherein each primer pair produces a short amplicon of not greater than 500 nucleotides and amplicon reads of not longer than 300 nucleotides; and f. computationally combine the amplicon reads to produce full-length 16s rRNA sequences based on the generated 16s rRNA database, thereby identifying unique 16s rRNA genes within the sample from the plant niche; and wherein, and wherein the amplifying or step (b), step (e) or both is performed: i. in the presence of a blocking oligo comprising a sequence that perfectly hybridizes with a 16s rRNA sequence of a chloroplast and does not perfectly hybridize with a 16s rRNA sequence of a microbe of the community; or ii. after cutting a 16s rRNA sequence of a chloroplast by contact with a CRISPR-guide RNA (gRNA) complex wherein the CRISPR- gRNA recognizes and cuts a 16s rRNA sequence of a chloroplast and not a 16s rRNA sequence of a microbe of the community; thereby profiling a microbial community within a plant niche.
[0010] According to some embodiments, the plurality of samples comprises at least 10 samples from the biological niche.[Oi l] According to some embodiments, the sample from (e) is a sample from the plurality of samples.
[0012] According to some embodiments, the method comprises receiving a plurality of samples from the biological niche, extracting DNA from each sample of the plurality of samples, and pooling the extracted DNA to produce the pool of DNA.
[0013] According to some embodiments, each long amplicon comprises at least 85% of the sequence of a 16s rRNA gene and wherein full-length 16s rRNA sequences comprise a contiguous at least 85% of a 16s rRNA gene.
[0014] According to some embodiments, the long amplicons each comprise at least 1000 nucleotides.
[0015] According to some embodiments, the long amplicons each comprise at least 1300 nucleotides.
[0016] According to some embodiments, the amplifying of step (b) is with at least two different sets of primers wherein each set of primers produces a long amplicon.
[0017] According to some embodiments, the amplifying of step (b) is with locked nucleic acid (LN A) primer pairs.
[0018] According to some embodiments, the LNA primers are 10-13 nucleotides long.
[0019] According to some embodiments, the LNA primer pairs are selected from the group consisting of: SEQ ID NO: 10-11, SEQ ID NO: 12-13 and SEQ ID NO: 14-15.
[0020] According to some embodiments, the amplifying of step (b) is with a cocktail of at least three different LNA primer pairs and wherein the three pairs are SEQ ID NO: 10-11, SEQ ID NO: 12-13 and SEQ ID NO: 14-15.
[0021] According to some embodiments, the long read sequencing technology is selected from the group consisting of: Single-molecule real-time sequencing, loop sequencing and long read nanopore sequencing.
[0022] According to some embodiments, the database comprises known full-length 16s rRNA sequences combined with the received full-length 16s rRNA sequences.
[0023] According to some embodiments, sequencing of step (e) comprises next generation sequencing, deep sequencing or massively parallel sequencing.
[0024] According to some embodiments, the plurality of primer pairs of step (e) amplifies at least 80% of a 16s rRNA gene.
[0025] According to some embodiments, the plurality of primer pairs of step (e) comprises SEQ ID NO: 1-9, SEQ ID NO: 29-40 or both.
[0026] According to some embodiments, all primers of the plurality of primer pairs of step (e) comprise: a. a sequencing adapter; b. a unique molecular identifier (UMI) or c. both.
[0027] According to some embodiments, the biological niche is on a plant.
[0028] According to some embodiments, the amplifying of step (e) is performed in the presence of a blocking oligo comprising a sequence that perfectly hybridizes with a 16s rRNA sequence of a chloroplast.
[0029] According to some embodiments, the blocking oligo is an LNA or PNA blocking oligo.
[0030] According to some embodiments, the amplifying of step (e) is performed after cutting a 16s rRNA sequence of a chloroplast by contact with a Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-guide RNA (gRNA) complex wherein the gRNA comprises a sequence that perfectly hybridizes with a 16s rRNA sequence of a chloroplast.
[0031] According to some embodiments, the CRISPR is a CAS9 protein.
[0032] According to some embodiments, the sequence that perfectly hybridizes with a 16s rRNA sequence of a chloroplast comprises or consists of the sequence CGTCTGTAGGTGG (SEQ ID NO: 28) or is 100% complementary to SEQ ID NO: 28.
[0033] According to some embodiments, method further comprises selecting a subset of primer pairs for step (e) that provide sufficient resolution for the 16s rRNA genes within the plurality of sample from the biological niche.
[0034] According to some embodiments, the computationally combining comprises performing Short MUlti-Region Framework (SMURF) analysis.
[0035] According to some embodiments, step (f) comprises grouping produced 16s rRNA sequences with substantially the same sequence to produce a list of unique 16s rRNA genes within the sample.
[0036] According to another aspect, there is provided a locked nucleic acid (LNA) primer pair comprising:a. an LNA forward primer comprising SEQ ID NO: 10 and an LNA reverse primer comprising SEQ ID NO: 11; b. an LNA forward primer comprising SEQ ID NO: 12 and an LNA reverse primer comprising SEQ ID NO: 13; or c. an LNA forward primer comprising SEQ ID NO: 14 and an LNA reverse primer comprising SEQ ID NO: 15.
[0037] According to some embodiments, each LNA primer is not longer than 13 nucleotides.
[0038] According to some embodiments, the LNA primer pair comprises a. an LNA forward primer consisting of SEQ ID NO: 10 and an LNA reverse primer consisting of SEQ ID NO: 11; b. an LNA forward primer consisting of SEQ ID NO: 12 and an LNA reverse primer consisting of SEQ ID NO: 13; or c. an LNA forward primer consisting of SEQ ID NO: 14 and an LNA reverse primer consisting of SEQ ID NO: 15.
[0039] According to another aspect, there is provided a kit comprising a first LNA primer pair of the invention, a second LNA primer pair of the invention and a third LNA primer pair of the invention, wherein the first, second and third LNA primer pairs each comprise different primers.
[0040] According to another aspect, there is provided a kit comprising an LNA primer pair of the invention and a blocking oligo comprising or consisting of SEQ ID NO: 28.
[0041] According to some embodiments, the blocking oligo is an LNA blocking oligo or a PNA blocking oligo.
[0042] According to another aspect, there is provided a kit comprising an LNA primer pair of the invention and a guide RNA (gRNA) comprising a sequence comprising or consisting of SEQ ID NO: 28.
[0043] According to some embodiments, the kit further comprises a CRISPR protein or a nucleic acid molecule encoding a CRISPR protein.
[0044] According to some embodiments, the CRISPR protein is a CAS9 protein.
[0045] According to some embodiments, the plurality of samples comprises at least 10 samples from the plant niche.
[0046] According to some embodiments, the method comprises receiving a plurality of samples from the plant niche, extracting DNA from each sample of the plurality of samples, and pooling the extracted DNA to produce the pool of DNA.
[0047] According to some embodiments, the amplifying of step (b) is performed in the presence of a blocking oligo comprising a sequence that perfectly hybridizes with a 16s rRNA sequence of a chloroplast and does not perfectly hybridize with a 16s rRNA sequence of a microbe of the community.
[0048] According to some embodiments, the amplifying of step (e) is performed in the presence of a blocking oligo comprising a sequence that perfectly hybridizes with a 16s rRNA sequence of a chloroplast and does not perfectly hybridize with a 16s rRNA sequence of a microbe of the community.
[0049] According to some embodiments, the blocking oligo is an LNA or PNA blocking oligo.
[0050] According to some embodiments, the sequence that perfectly hybridizes with a 16s rRNA sequence of a chloroplast and does not perfectly hybridize with a 16s rRNA sequence of a microbe of the community comprises or consists of the sequence CGTCTGTAGGTGG (SEQ ID NO: 28).
[0051] According to some embodiments, the method further comprises before the amplifying of step (e) cutting a 16s rRNA sequence of a chloroplast by contact with a CRISPR-guide RNA (gRNA) complex wherein the CRISPR-gRNA complex recognizes and cuts a 16s rRNA sequence of a chloroplast and not a 16s rRNA sequence of a microbe of the community.
[0052] According to some embodiments, the CRISPR-gRNA complex recognizes and cuts CGTCTGT (SEQ ID NO: 41) or SEQ ID NO: 28.
[0053] According to some embodiments, the gRNA comprises the sequence TTGGGCGTAAAGCGTCTGT (SEQ ID NO: 42)
[0054] According to some embodiments, further comprises selecting a subset of primer pairs for step (e) that provide sufficient resolution for the 16s rRNA genes within the plurality of sample from the plant niche.
[0055] According to some embodiments, the computationally combining comprises performing Short MUlti-Region Framework (SMURF) analysis.
[0056] According to some embodiments, step (f) comprises grouping produced 16s rRNA sequences with substantially the same sequence to produce a list of unique 16s rRNA genes within the sample.
[0057] According to another aspect, there is provided a locked nucleic acid (LNA) primer pair comprising: a. an LNA forward primer comprising SEQ ID NO: 10 and an LNA reverse primer comprising SEQ ID NO: 11; b. an LNA forward primer comprising SEQ ID NO: 12 and an LNA reverse primer comprising SEQ ID NO: 13; or c. an LNA forward primer comprising SEQ ID NO: 14 and an LNA reverse primer comprising SEQ ID NO: 15.
[0058] According to some embodiments, each LNA primer is not longer than 13 nucleotides.
[0059] According to some embodiments, the LNA primer pair comprises: a. an LNA forward primer consisting of SEQ ID NO: 10 and an LNA reverse primer consisting of SEQ ID NO: 11; b. an LNA forward primer consisting of SEQ ID NO: 12 and an LNA reverse primer consisting of SEQ ID NO: 13; or c. an LNA forward primer consisting of SEQ ID NO: 14 and an LNA reverse primer consisting of SEQ ID NO: 15.
[0060] According to another aspect, there is provided a kit comprising a first LNA primer pair of the invention, a second LNA primer pair of the invention and a third LNA primer pair of the invention, wherein the first, second and third LNA primer pairs each comprise different primers.
[0061] According to another aspect, there is provided a kit comprising an LNA primer pair of the invention and a. a blocking oligo comprising or consisting of SEQ ID NO: 28;b. a gRNA that binds to and cuts SEQ ID NO: 41 or SEQ ID NO: 28; or c. both.
[0062] According to some embodiments, the blocking oligo is an LNA blocking oligo or a PNA blocking oligo.
[0063] According to some embodiments, the gRNA comprises SEQ ID NO: 42.
[0064] According to some embodiments, the kit comprises the gRNA and further comprising a CAS9 enzyme or a nucleic acid molecule encoding the CAS9 enzyme.
[0065] Further embodiments and the full scope of applicability of the present invention will become apparent from the detailed description given hereinafter. However, it should be understood that the detailed description and specific examples, while indicating preferred embodiments of the invention, are given by way of illustration only, since various changes and modifications within the spirit and scope of the invention will become apparent to those skilled in the art from this detailed description.BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figures 1A-1F: Microbial profiling of Aster altaicus samples depends on the applied amplicon. Leaf, soil, and root samples of Aster altaicus were profiled by the Qiagen kit that provides six amplicons along the 16s rRNA gene and subsequently analyzes their reads using the CLC Genomics platform. (1A) The six amplicons of the Qiagen kit, each amplifying two or three adjacent variable regions. Taken together, these amplicons cover all variable regions. (IB) The number of OTUs detected for soil, leaf, and root samples of Aster altaicus for each of the six amplicons. (1C) Distribution of Jaccard distances of among pairs of amplicons. In each pair, e.g., V1V2 vs. V3V4, a Jaccard distance is calculated between the lists of taxons detected by CLC Genomics in each region. (1D-1F) Decay graph of the number of unique taxons identified by the amplicons analyzed. Amplicons are ordered by the number of unique taxons a region contributes to the former set of regions in (ID) Leaf, (IE) Soil, and (IF) Root.
[0067] Figures 2A-2F: Evaluating the performance of sets of primer pairs across niches. 24 samples corresponding to leaf, soil, and root of eight plants were profiled by SMURF, using all 63 combinations of the six amplicons. (2A) Ambiguity, i.e., the number of 16s rRNA gene per group as detected by SMURF, as a function of the number of amplicons. Each dot corresponds to the average ambiguity in one of the 63 distinct cases.The blue line and gray background correspond to the average and standard deviation of the ambiguity across combinations of a specific number of regions. (2B) We consider SMURF’s solution based on all six regions as the gold standard and display a fraction of species identified by subsets of the six regions. Shown is the percent of species whose taxonomy matches the six-region solution. (2C) Pearson correlation between profiling based on all six regions and each individual set of regions. (2D-2F) A comparison of amplicon sets for (2D) soil, (2E) root and (2F) leaf samples. Each dot corresponds to a specific region combination averaged across samples (soil or root), where the marker size scales with the ambiguity and axes correspond to Pearson correlation (horizontal) and mutual species (vertical). The marker’s color indicates the number of amplicons used for SMURF analysis. Hence, a single marker corresponds to 6 regions (yellow), and 6 green dots correspond to the different groups of 5 regions. The marker “x” corresponds to a specific set of 3 regions comprising V1V2, V2V3, and V4V5.
[0068] Figures 3A-3C: Enriching the database with novel full-length 16s rRNA genes. The full-length 16s rRNA gene was amplified using designed LNA primers and the resulting amplicons were sequenced by PacBio long-read sequencing. Four samples were processed, each containing a pool of samples from a different habitat, i.e., plants from the Arava desert in Israel, characterized by a climate type BWh, according to Koppen climate classification; plants from the lowland region in Israel, characterized by a climate type Csa, and pools of samples from the marine sponge Spongia officinalis collected in the Mediterranean Sea along the coast of either France or Israel. (3A) All non-redundant 16s rRNA genes (Plants BWh - 14852, Plants Csa - 14589, France - 6507, Israel - 5468) from the PacBio sequencing were aligned to the SILVA database using the CD-HIT algorithm. The percentage of 16s rRNA genes that did not appear in SILVA is shown as a function of the identity threshold. The vertical dashed line corresponds to SILVA’s standard threshold for calling novel SSUs (SILVA NR99). Colors indicate different biomes sampled. (3B) Histogram of read counts per 16s rRNA gene for the France sponge pool of samples, for novel 16s rRNA genes (upper panel), and for known SILVA SSUs (lower panel), using a 97% identity threshold. (3C) Rarefaction curves presenting the number of novel 16s rRNA genes identified (y-axis) based on the percentage of reads sampled (x-axis), for each of the sequenced pools.
[0069] Figure 4: The CoSMIC suggested flowchart. Gray (top) - Perform standard DNA extraction from environmental samples. Orange (right) - Establish a relevant 16s rRNA GENE database by LNA-based amplification and long -read sequencing. Orange-purple (middle) - Test multiple primer sets to find an optimal primer pair combination. Purple (Left)- Perform high throughput profiling in a large-scale project. Samples are amplified using the optimal primer combination and analyzed by SMURF based on the augmented database.
[0070] Figures 5A-5E: Figure 5. Experimental evaluation of standard methods vs. CoSMIC. Ten samples were sequenced - leaf, soil, and root samples of Thymelaea hirsuta, Solcinum lycopersicum, and Triticum aestivum, and a mock mixture of 12 known species (see Materials and Methods). Each sample underwent standard shotgun metagenomic sequencing, from which we detected its full-length 16s rRNA gene to be considered groundtruth. In addition, as part of CoSMIC we first pooled extracted DNA from all 10 samples, amplified by LNA primers, and sequenced via synthetic long reads (Loop genomics) to create 16s rRNA genes that were added to the SILVA database. Standard amplicon sequencing was performed by a commercially available kit by Swift, which uses primers spanning the 16s rRNA gene to amplify several variable regions (see Materials and Methods). Analysis was then performed by either standard Swift pipeline using 10 primer pairs (Swift all), by the V3V4 region (Swift V3V4), and by CoSMIC. (5A) Alignment identity, i.e., the average alignment score between detected 16s rRNA genes and 16s rRNA genes detected by shotgun metagenomics. Boxplots present the distribution of alignment identities for each tissue and for each analysis method (CoSMIC (C), Swift all (S) and Swift V3V4 (V)). Leaf, soil, and root panels correspond to an average of over 3 samples (one from each plant). A Mann- Whitney U test was performed for all comparisons, which resulted in very low p-values, with the highest being 0.0064. (5B) Distribution of the number of taxa matching an identification by each analysis type. (5C) Total frequency of ground truth detected by each method as a function identity cutoff. The insets present the total frequency achieved per sample at a cutoff of 99%, while the dotted horizontal line is their average. (CoSMIC-(C); Swift all-(S); Swift V3V4-(V)). (5D) Percentage of erroneous 16s rRNA genes, i.e., those that were identified by an analysis method yet were not detected by shotgun metagenomics. Each marker represents a specific plant, while each facet corresponds to a tissue. The dotted line marks the mean for each tissue and method. (5E) Radar plots combining all samples over several criteria (normalized on a 0-1 scale, where large values correspond to better performance): FP, mean false positive rate as described in 5D; Amb, mean ambiguity as described in 5B; Cov-99, mean total frequency ground truth coverage at 99% identity; Cov-80, mean ground truth coverage at 80% identity cutoff as described in 5C; Id, mean identity for each analysis method as described in 5A.
[0071] Figures 6A-6E: (6A) The number of OTUs detected for soil, leaf, and root samples of Brcmdisici hancei for each of the six amplicons. (6B) Distribution of Jaccard distances ofamong pairs of amplicons. In each pair, e.g., V 1 V2 vs. V3V4, a Jaccard distance is calculated between the lists of taxons detected by CLC Genomics in each region. (6C-6E) Decay graph of the number of unique taxons identified by the amplicons analyzed. Amplicons are ordered by the number of unique taxons a region contributes to the former set of regions in (6C) Leaf, (6D) Soil, and (6E) Root.
[0072] Figures 7A-7E: (7A) The number of OTUs detected for soil, leaf, and root samples of Lepidium apetcilum for each of the six amplicons. (7B) Distribution of Jaccard distances of among pairs of amplicons. In each pair, e.g., V1V2 vs. V3V4, a Jaccard distance is calculated between the lists of taxons detected by CLC Genomics in each region. (7C-7E) Decay graph of the number of unique taxons identified by the amplicons analyzed. Amplicons are ordered by the number of unique taxons a region contributes to the former set of regions in (7C) Leaf, (7D) Soil, and (7E) Root.
[0073] Figures 8A-8E: (8A) The number of OTUs detected for soil, leaf, and root samples of Lathyrus sativus for each of the six amplicons. (8B) Distribution of Jaccard distances of among pairs of amplicons. In each pair, e.g., V 1 V2 vs. V3V4, a Jaccard distance is calculated between the lists of taxons detected by CLC Genomics in each region. (8C-8E) Decay graph of the number of unique taxons identified by the amplicons analyzed. Amplicons are ordered by the number of unique taxons a region contributes to the former set of regions in (8C) Leaf, (8D) Soil, and (8E) Root.
[0074] Figures 9A-9E: (9 A) The number of OTUs detected for soil, leaf, and root samples of Angelica hirsutiflora for each of the six amplicons. (9B) Distribution of Jaccard distances of among pairs of amplicons. In each pair, e.g., V1V2 vs. V3V4, a Jaccard distance is calculated between the lists of taxons detected by CLC Genomics in each region. (9C-9E) Decay graph of the number of unique taxons identified by the amplicons analyzed. Amplicons are ordered by the number of unique taxons a region contributes to the former set of regions in (9C) Leaf, (9D) Soil, and (9E) Root.
[0075] Figures 10A-10E: (10A) The number of OTUs detected for soil, leaf, and root samples of Tetraena mongolica for each of the six amplicons. (10B) Distribution of Jaccard distances of among pairs of amplicons. In each pair, e.g., V 1 V2 vs. V3V4, a Jaccard distance is calculated between the lists of taxons detected by CLC Genomics in each region. (10C- 10E) Decay graph of the number of unique taxons identified by the amplicons analyzed. Amplicons are ordered by the number of unique taxons a region contributes to the former set of regions in (IOC) Leaf, (10D) Soil, and (10E) Root.
[0076] Figures 11A-11C: Ranking of Amplicons used in Qiagen analysis. Analysis of (11A) soil, (11B) leaf, and (11C) root samples involved six distinct amplicons. Each amplicon is ranked according to the amount of information it contributed to the analysis. There is no evident ranking, thus indicating that the information contribution for each measurement originates from varying amplicons.
[0077] Figure 12: Number of SSU derived from PacBio long read sequencing. In order to reduce redundancy of amplicons derived from the PacBio long read sequencing, each sample was clustered using cd-hit with an identity threshold of 99%. In black is the number of sequences derived for each sample prior to clustering; in gray is the number of seeds obtained after clustering.
[0078] Figures 13A-13C: (13A) Map of the E. coli and wheat chloroplast 16s rRNA sequence. The unique blocker sequence in V3-V4 (small pink box) is present in the wheat sequence but not the E. coli sequence. (13B) qPCR results demonstrating inhibition efficiency in synthetic plasmids containing Arabidopsis chloroplast 16S or E. coli 16S following LNA blocker treatment. (13C) qPCR demonstrating CRISPR-mediated reduction of chloroplast DNA amplification.
[0079] Figures 14A-14G: Bar plots represent the relative abundance of (14A) the standard bacterial community produced by the standard library preparation, with LNA blocker or with CRISPR; (14B, 14D) chloroplast sequences shown in green and standard bacterial community sequences in gray when a (14B) 50:50 mixture or (14D) 90: 10 mixture is prepared without LNA (No LNA-control), with addition of 10 pM LNA or 100 LNA pM used in one PCR or in both PCR step. 10 100 refers to 10 pM used in the first PCR and 100 pM used in the second PCR. CRISPR-mediated removal of chloroplast DNA before library preparation is shown to the right of the dashed line; (14C, 14E) the bacterial community (gray area zoomed in) analyzed in each condition for both (14C) the 50:50 mixture and the (14E) 90: 10 mixture. Pseudomonas (orange and with arrow) are not well detected without the LNA blocker or with only 10 pM added to one PCR in the 50:50 mixture and become essentially undetectable in the 90: 10 mixture; and (14F- 14G) chloroplast sequences shown in green and bacteria sequences shown in color in leaf samples processed (14F) with and without CRISPR or (14G) with and without the LNA blocker. Bacteria sequences are only detectable with CRISPR or blocker treatment.
[0080] Figures 15A-15D: Comparison of bacterial composition in wheat root samples between two analysis pipelines: UMI and no UMI. (15A) Bar plot showing the relativeabundance of the 30 most abundant bacterial orders across five wheat root samples. (+) indicates analysis using the UMI-based pipeline, and (-) indicates the standard No UMI pipeline. Bray-Curtis distances shown above each pair of bars represent the compositional similarity between pipelines, with values closer to zero indicating higher similarity. (15B) Bar plot of Bray-Curtis dissimilarity between the two pipelines calculated for 39 wheat samples across multiple taxonomic resolutions (Order, Family, Genus, and Species). (15C) Bar graph of species richness (total number of taxa detected) evaluated across all samples with and without using UMIs in the analysis. (15D) Bar graph of differences in species richness evaluated across all samples with and without using UMIs in the analysis. A positive A Richness indicates more species identified by the UMI pipeline, whereas a negative A Richness indicates more species detected without UMIs. Horizontal lines within each box represent median values.DETAILED DESCRIPTION OF THE INVENTION
[0081] The present invention, in some embodiments, provides methods of profiling a microbial community. The present invention further concerns locked nucleic acid (LNA) primer pairs. Kits comprising the LNA primer pairs are also provided.
[0082] The invention is based, at least in part, on the generation of the Comprehensive 16s rRNA geneSSU Mapping and Identification of Communities (CoSMIC), which builds on SMURF to provide a framework for high-resolution profiling of novel niches. Prior to applying SMURF in a new niche, amplifying a pool of samples using locked nucleic acid (LNA) primers is done, which provides an almost universal amplification of full-length 16s rRNA sequences. Sequencing these amplicons by long-read technologies (e.g., PacBio) enriches 16s rRNA gene databases with relevant sequences of 16s rRNA, thus allowing the use of SMURF to provide efficient and high-resolution profiling of these niches. The advantages of CoSMIC over standard methods are demonstrated herein by sequencing a mock mixture and over a variety of environmental samples as well as marine sponges.
[0083] The invention is further based, at least in part, on the surprising finding of a 13 bp sequence (SEQ ID NO: 28) that is common to chloroplast 16s rRNA, but is unique with respect to known microbes’ 16s rRNA. That is, this sequence can be used to specifically target chloroplast 16s rRNA genes and remove them from the sample. Both a blocking oligo used during amplification and CRISPR cutting targeted with a guide RNA (gRNA) madeuse of this unique sequence to deplete chloroplast 16s rRNA and thus enhance the analysis of bacterial 16s rRNA.
[0084] By a first aspect, there is provided a method of profiling a microbial community, the method comprising: a. receiving DNA from a biological niche; b. amplifying 16s ribosomal RNA (rRNA) genes within the received DNA; c. sequencing long amplicons to receive long 16s rRNA sequences; d. generating a 16s rRNA database comprising the long 16s rRNA sequences; e. amplifying and sequencing 16s rRNA genes from the biological niche to receive 16s rRNA short reads; and f. combining the short reads sequences based on the database to produce long 16s rRNA sequences; thereby profiling a microbial community.
[0085] In some embodiments, the method is an in vitro method. In some embodiments, the method is an ex vivo method. In some embodiments, the method is a method of discovering 16s rRNA sequences. In some embodiments, the method is a method of microbiome analysis. In some embodiments, the method is a method of identifying organisms in a microbial community. In some embodiments, the method is a method of identifying species of microbes in a microbial community. In some embodiments, the microbial community is in a biological niche. In some embodiments, the method is a method of identifying species of microbes in a biological niche. In some embodiments, the method is a method of identifying organisms in a biological niche. In some embodiments, identifying is discovering.
[0086] It is well known in the art that different species of bacteria can be identified by their different 16s rRNAs. Thus, microbiome analysis by 16s rRNA sequencing is well known. However, there are many drawbacks related to the current approaches. These approaches are highly variable depending on the region of the 16s probed. This results in lack of consistency between results making harmonization of findings difficult and unreliable. The current approaches also produce a large amount of bias based both on the primers selected, the region investigated, and the bioinformatics analysis used. In particular, by looking only at single small part of the gene, low phylogenetic resolution is provided that leads to bias and aninaccurate representation of the original sample’s microbial diversity. Previous approaches also had an unacceptably high level of ambiguity, as reads are often aligned to multiple 16s rRNA genes. This leads to inconsistent taxonomies and inaccurate results. Finally, the comparative databases used in the past are limited. These databases are always highly biased toward human-associated niches and due to lack of universal primers not all bacteria are represented in the database. These results in the missing of new bacteria during detection. The instant invention, however, provides an improved method of such analysis. The universal primers provided herein provide full coverage of the 16s gene, with very low bias and the ability to detect rare / new bacteria. By using long read technology the issues of regional bias, low phylogenetic resolution and ambiguity of alignment are abrogated. The database produced by the instant analysis and the analysis performed remove bias and enable an accurate representation of microbial diversity. The instant method thus solves these problems of the old methos and provides a universal method that can be used generally for microbiome profiling.
[0087] As used herein, the term “biological niche” refers to a specialized biological location with a specific microbial repertoire. In some embodiments, the niche is a human associated niche. In some embodiments, the niche is in a human body. In some embodiments, the niche is in a biological fluid. In some embodiments, the fluid is a human fluid. In some embodiments, the fluid is stool. In some embodiments, the niche is not associated with a human. In some embodiments, the niche is an ecological niche. In some embodiments, the niche is a location in nature. In some embodiments, the niche is soil. In some embodiments, the niche is a body of water. In some embodiments, a niche is a biome. In some embodiments, the niche is a new, previously unexamined niche. In some embodiments, the niche is a natural habitat. In some embodiments, the niche does not comprise or is devoid of humans or human cells. In some embodiments, the niche is a plant niche. In some embodiments, the niche is a plant or a part thereof. In some embodiments, the niche is a part of a plant. In some embodiments, the part is selected from leaves and roots. In some embodiments, the niche is a plant associated niche. In some embodiments, the plant associated niche is selected from a part of a plant and soil around the plant. In some embodiments, the biological niche is on or in a plant. In some embodiments, the biological niche is on or in a leaf. In some embodiments, the biological niche is on or in a root.
[0088] In some embodiments, the method comprises receiving a sample from a niche. In some embodiments, the sample is not from a human. In some embodiments, the sample is a non-human sample. In some embodiments, a sample is a plurality of samples. In someembodiments, a plurality is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 75, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 750, 800, 900 and 1000 samples. Each possibility represents a separate embodiment of the invention. In some embodiments, a plurality is at least 10 samples. In some embodiments, a plurality is at least 25 samples. In some embodiments, a plurality is at least 50 samples.
[0089] In some embodiments, DNA is extracted from the sample. In some embodiments, the method comprises extracting DNA from the sample. In some embodiments, the method comprises receiving DNA extracted from a sample. In some embodiments, receiving DNA is receiving a pool of DNA. In some embodiments, the DNA is pooled from the samples. In some embodiments, the method comprises receiving a plurality of samples from the niche, extracting DNA from each sample and pooling the extracted DNA to produce a pool of DNA. Methods of extracting DNA and kits for doing same are well known in the art and any such method may be used.
[0090] In some embodiments, 16s rRNA genes are amplified within the sample. In some embodiments, 16s rRNA genes are amplified within the pool. In some embodiments, the amplifying produces long amplicons. In some embodiments, a long is at least 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500 or 1600 nucleotides in length. Each possibility represents a separate embodiment of the invention. In some embodiments, long is at least 800 nucleotides in length. In some embodiments, long is at least 900 nucleotides in length. In some embodiments, long is at least 1000 nucleotides in length. In some embodiments, long is at least 1100 nucleotides in length. In some embodiments, long is at least 1200 nucleotides in length. In some embodiments, long is at least 1300 nucleotides in length. In some embodiments, long is at least 1400 nucleotides in length.
[0091] In some embodiments, the long amplicon comprises a full-length 16s rRNA sequence. In some embodiments, a 16s rRNA sequence is a 16s rRNA gene. In some embodiments, the long amplicon comprises a contiguous 16s rRNA sequence. In some embodiments, the long amplicon comprises at least 70% of a sequence of a 16s rRNA gene. In some embodiments, each long amplicon comprises at least 70% of a sequence of a 16s rRNA gene. It will be understood by a skilled artisan that primers for amplifying a 16s rRNA gene can be placed essentially anywhere along the gene body. However, in this method it is important to produce long amplicons that contain large amounts of a 16s genes on a single strand. Thus, the amplicon (a single strand) must contain at least 70% of an entire 16s rRNA gene. In some embodiments, the long amplicon comprises at least 75% of a sequence of a 16s rRNA gene. In some embodiments, the long amplicon comprises at least 80% of asequence of a 16s rRNA gene. In some embodiments, the long amplicon comprises at least 85% of a sequence of a 16s rRNA gene. In some embodiments, the long amplicon comprises at least 87% of a sequence of a 16s rRNA gene. In some embodiments, the long amplicon comprises at least 90% of a sequence of a 16s rRNA gene. In some embodiments, the long amplicon comprises at least 92% of a sequence of a 16s rRNA gene. In some embodiments, the long amplicon comprises at least 95% of a sequence of a 16s rRNA gene.
[0092] In some embodiments, the amplifying is with a pair of primers. In some embodiments, the amplifying is with a plurality of pairs of primers. In some embodiments, the pair or primers amplifies a long amplicon. In some embodiments, each pair or primers amplifies a long amplicon. In some embodiments, amplifies is produces. In some embodiments, a plurality is at least two. In some embodiments, each pair amplifies a different long amplicon. In some embodiments, combining the different long amplicons produces a sequence that is at least 70, 75, 80, 85, 87, 90, 92, 95 or 97% of a 16s rRNA gene. In some embodiments, a plurality of pairs of primers is at least 2 ,3, 4, or 5 pairs of primers. Each possibility represents a separate embodiment of the invention. In some embodiments, a plurality of pairs of primers is at least 3 pairs. In some embodiments, a plurality of pairs of primers is 3 pairs of primers. In some embodiments, a plurality of pairs of primers comprises at least 2 forward primers and at least 2 reverse primers. In some embodiments, a plurality of pairs of primers comprises at least 3 forward primers and at least 3 reverse primers.
[0093] It will be understood by a skilled artisan that the initial long amplification is meant to capture as many different 16s rRNA genes as possible. Thus, the long amplicon is not at least 70% of any specific 16s gene but rather is generally at least 70% of any 16s gene. The designing of primers for amplification is well known in the art. For the long amplification primers should be selected to be as short as possible so that they will hybridize to as many 16s genes as possible. They should target regions that are highly conserved between different 16s genes. Further, multiple primer pairs will be used to increase the chances of a forward and reverse primer binding to any given 16s gene. Since all the forward primers will be at the 5 ’ end of the gene and all the reverse primers at the 3 ’ end, each forward primer can potentially amplify with each reverse primer, increasing the chance that a 16s gene will be amplified.
[0094] In some embodiments, the primers are at most 10, 11, 12, 13, 14, 15, or 16 nucleotides long. Each possibility represents a separate embodiment of the invention. In some embodiments, the primers are at most 10 nucleotides long. In some embodiments, the primers are at most 13 nucleotides long. In some embodiments, the primers are at most 15nucleotides long. In some embodiments, the primers are at least 5, 6, 7, 8, 9, or 10 nucleotides long. Each possibility represents a separate embodiment of the invention. In some embodiments, the primers are at least 8 nucleotides long. Each possibility represents a separate embodiment of the invention. In some embodiments, the primers are at least 10 nucleotides long. In some embodiments, the primers are 5-8, 5-9, 5-10, 5-11, 5-12, 5-13, 5- 15, 7-8, 7-9, 7-10, 7-11, 7-12, 7-13, 7-15, 8-9, 8-10, 8-11, 8-12, 8-13, 8-15, 9-10, 9-11, 9- 12, 9-13, 9-15, 10-11, 10-12, 10-13, or 10-15 nucleotides long. Each possibility represents a separate embodiment of the invention. In some embodiments, the primers are 9-13 nucleotides long. In some embodiments, the primers are 10-13 nucleotides long.
[0095] In some embodiments, each primer comprises a melting temperature (Tm) of at least 45, 48, 50, 52, 53, 54, 55, 56, 57, 58, 59 or 60 degrees Celsius. Each possibility represents a separate embodiment of the invention. In some embodiments, each primer pair comprises a melting temperature (Tm) of at least 45, 48, 50, 52, 53, 54, 55, 56, 57, 58, 59 or 60 degrees Celsius. Each possibility represents a separate embodiment of the invention. In some embodiments, each primer comprises a Tm of at least 45. In some embodiments, each primer comprises a Tm of at least 50. In some embodiments, each primer comprises a Tm of at least 55. In some embodiments, each primer pair comprises a Tm of at least 45. In some embodiments, each primer pair comprises a Tm of at least 50. In some embodiments, each primer pair comprises a Tm of at least 55.
[0096] In order to use very short primers but retain a proper Tm and specificity non-natural nucleic acids (nucleosides) can be incorporated into the primers. In some embodiments, the primers comprise a non-natural nucleic acid. In some embodiments, a non-natural nucleic acid is a nucleic acid analog. In some embodiments, a non-natural nucleic acid is a nucleic acid mimic. In some embodiments, the non-natural nucleic acid is a locked nucleic acid (LNA). In some embodiments, the primer is an LNA primer. In some embodiments, the primer pair is an LNA primer pair. In some embodiments, the non-natural nucleic acid is a peptide nucleic acid (PNA). In some embodiments, the primer is a PNA primer. In some embodiments, the primer is a PNA-DNA primer. In some embodiments, a PNA-DNA primer is a PNA-DNA chimera. In some embodiments, the primer pair is a PNA primer pair. It will be understood that the primers that hybridize to the 16s gene and amplify it are themselves incorporated with LNA / PNA bases. This is distinct for 16s profiling methods that use an LNA / PNA clamp, that is using an LNA / PNA blocking oligonucleotide. The primers of the instant invention are not blocking molecules but rather are short highly specific primers thatbind to the target 16s sequence and initiate transcription by polymerase. In some embodiments, LNA primers are used only for long amplicon production.
[0097] Examples of LNA primers for performance of methods of the invention are provided hereinbelow. In some embodiments, the primer sequence is selected from SEQ ID NO: 10- 15. In some embodiments, the primer sequence comprises a sequence selected from SEQ ID NO: 10-15. In some embodiments, a forward primer comprises SEQ ID NO: 10. In some embodiments, the sequence of a forward primer is SEQ ID NO: 10. In some embodiments, the sequence of a forward primer consists of SEQ ID NO: 10. In some embodiments, a reverse primer comprises SEQ ID NO: 11. In some embodiments, the sequence of a reverse primer is SEQ ID NO: 11. In some embodiments, the sequence of a reverse primer consists of SEQ ID NO: 11. In some embodiments, the primer pair is SEQ ID NO: 10-11. In some embodiments, the primer pair comprises a forward primer comprising a sequence consisting of SEQ ID NO: 10 and a reverse primer comprising a sequence consisting of SEQ ID NO: 11. In some embodiments, a forward primer comprises SEQ ID NO: 12. In some embodiments, the sequence of a forward primer is SEQ ID NO: 12. In some embodiments, the sequence of a forward primer consists of SEQ ID NO: 12. In some embodiments, a reverse primer comprises SEQ ID NO: 13. In some embodiments, the sequence of a reverse primer is SEQ ID NO: 13. In some embodiments, the sequence of a reverse primer consists of SEQ ID NO: 13. In some embodiments, the primer pair is SEQ ID NO: 12-13. In some embodiments, the primer pair comprises a forward primer comprising a sequence consisting of SEQ ID NO: 12 and a reverse primer comprising a sequence consisting of SEQ ID NO: 13. In some embodiments, a forward primer comprises SEQ ID NO: 14. In some embodiments, the sequence of a forward primer is SEQ ID NO: 14. In some embodiments, the sequence of a forward primer consists of SEQ ID NO: 14. In some embodiments, a reverse primer comprises SEQ ID NO: 15. In some embodiments, the sequence of a reverse primer is SEQ ID NO: 15. In some embodiments, the sequence of a reverse primer consists of SEQ ID NO: 15. In some embodiments, the primer pair is SEQ ID NO: 14-15. In some embodiments, the primer pair comprises a forward primer comprising a sequence consisting of SEQ ID NO: 14 and a reverse primer comprising a sequence consisting of SEQ ID NO: 15.
[0098] In some embodiments, the long amplifying is with a cocktail of primer pairs. In some embodiments, a cocktail is a plurality. In some embodiments, all the primer pairs are LNA primer pairs. In some embodiments, the cocktail comprises at least 3 primer pairs. In some embodiments, the at least 3 primer pairs are SEQ ID NO: 10-11, SEQ ID NO: 12-13 andSEQ ID NO: 14-15. In some embodiments, the sequences of the at least 3 primer pairs are SEQ ID NO: 10-11, SEQ ID NO: 12-13 and SEQ ID NO: 14-15.
[0099] It will be understood that not every base in an LNA primer needs to be an LNA base. Rather LNA bases can be added so as to produce the desired Tm. Indeed, in a cocktail of primers, the incorporation of LNA allows bringing primers of different lengths and different GC contents to the same Tm by the incorporation of different numbers of LNA bases. The addition of LNA bases to produce a specific Tm is routine in the art and a skilled artisan would be able to incorporate LNA bases as needed. Further, LNA primers can be ordered from commercial vendors who are skilled at inserting LNA bases. For example, the LNA primers used herein were ordered from Qiagen (catalog #339406). LNA primers can be designed and ordered from geneglobe.qiagen.com / us / product-groups / custom-lna- oligonucleotides. Other companies that provide LNA primers include, but are not limited to, IDT, Metabion, Eurofins Genomics LLC, LGC Biosearch Technologies and Biomers.net. PNA primers are similarly available from these companies and PNA incorporations can be used in a similar manner to LNA incorporation.
[0100] In some embodiments, the primer comprises at least 1 non-natural nucleic acid. In some embodiments, the primer comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 non-natural nucleic acids. Each possibility represents a separate embodiment of the invention. In some embodiments, the primer comprises at least 5 non-natural nucleic acids. In some embodiments, the primer comprises at least 8 non-natural nucleic acids. In some embodiments, the primer comprises at least 9 non-natural nucleic acids. In some embodiments, the primer comprises at least 10 non-natural nucleic acids. In some embodiments, the primer comprises at least 1 LNA base. In some embodiments, the primer comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 LNA bases. Each possibility represents a separate embodiment of the invention. In some embodiments, the primer comprises at least 5 LNA bases. In some embodiments, the primer comprises at least 8 LNA bases. In some embodiments, the primer comprises at least 9 LNA bases. In some embodiments, the primer comprises at least 10 LNA bases.
[0101] In some embodiments, the amplification is performed in the presence of a blocking oligo. In some embodiments, the long read amplification is performed in the presence of a blocking oligo. In some embodiments, a blocking oligo is a blocker. In some embodiments, the amplifying of step (b) is performed in the presence of a blocking oligo. In some embodiments, the method further comprises adding a blocking oligo to the amplification. In some embodiments, the method further comprises adding a blocking oligo to the long readamplification. In some embodiments, the blocking oligo is an LNA oligo. In some embodiments, the blocking oligo is a PNA oligo. In some embodiments, the blocking oligo blocks amplification of 16s rRNA from a non-microbial source. In some embodiments, the blocking oligo does not block amplification of 16s rRNA from a microbial source. In some embodiments, a microbial source is a microbe of the community of microbes. In some embodiments, the non-microbial source is a plant. In some embodiments, the non-microbial source is a chloroplast. In some embodiments, the non-microbial source is an organism that is the biological niche. In some embodiments, the non-microbial source is an animal. In some embodiments, the blocking oligo blocks amplification of 16s rRNA from chloroplasts. In some embodiments, the blocking oligo blocks amplification of 16s rRNA from chloroplasts and does not block amplification of a 16s rRNA from a microbe of the community. In some embodiments, the biological niche is a plant and the blocking oligo blocks amplification of 16s rRNA from chloroplasts. In some embodiments, the blocking oligo hybridizes to 16s rRNA from the non-microbial source. In some embodiments, the blocking oligo does not hybridizes to 16s rRNA from a microbial source. In some embodiments, the blocking oligo does not hybridizes to 16s rRNA from a microbe of the community. In some embodiments, the blocking oligo hybridizes to 16s rRNA from chloroplasts. In some embodiments, the blocking oligo hybridizes to 16s rRNA from chloroplasts and does not hybridize to 16s rRNA from a microbe of the community. In some embodiments, a microbe of the community is all microbes of the community. In some embodiments, hybridizes is perfectly hybridizes. In some embodiments, hybridizes is 100% hybridizes. In some embodiments, perfectly hybridizes is without any mismatches. In some embodiments, the blocking oligo does not hybridize with microbial 16s rRNA. In some embodiments, the blocking oligo does not hybridize with most microbial 16s rRNA. In some embodiments, most is at least 50, 55, 60, 65, 70, 75, 80, 85, 90, 92, 95, 97 or 99% of all microbial 16s rRNAs. Each possibility represents a separate embodiment of the invention. In some embodiments, microbial 16s rRNA is of microbes of the niche . In some embodiments, microbial 16s rRNA is of microbes of the community. In some embodiments, the blocking oligo comprises the sequence CGTCTGTAGGTGG (SEQ ID NO: 28). In some embodiments, the blocking oligo consists of the sequence of SEQ ID NO: 28. In some embodiments, the blocking oligo hybridizes to the V3-V4 region of the non-microbial 16s rRNA. In some embodiments, the blocking oligo hybridizes to the V3-V4 region of chloroplast 16s rRNA. In some embodiments, hybridizes to is is reverse complementary to.
[0102] In some embodiments, the method comprises sequencing the long amplicons. In some embodiments, the method comprises sequencing the results of the amplification. In some embodiments, the long amplicons of 16s rRNA sequences are sequenced. In some embodiments, the sequencing produces long 16s rRNA sequences. In some embodiments, the sequencing produces full length 16s rRNA gene sequences. In some embodiments, the sequencing produces full length 16s rRNA sequences. In some embodiments, the full-length sequence comprises a contiguous at least 70, 75, 80, 85, 87, 90, 92, 95, 97, or 99% of a 16s rRNA gene. Each possibility represents a separate embodiment of the invention. In some embodiments, the sequence is with a long-read sequencing technology.
[0103] As used herein, the term “long -read sequencing technology” refers to any sequencing method that provides contiguous sequence reads that are longer than 500 nucleotides. To avoid confusion, long-read sequencing does not refer to the production of synthetic long reads in which short reads and combined to produce a long read, but rather refers to a continuous sequence read from a single template that is long (i.e., longer than 500 nucleotides). In some embodiments, long -read sequencing is third-generation sequencing. In some embodiments, long -read sequencing produces contiguous reads of at least 500, 600, 700, 800, 900, 1000, 1100, 1200 or 1300 nucleotides. Each possibility represents a separate embodiment of the invention. In some embodiments, long-read sequencing produces contiguous reads of at least 700 nucleotides. In some embodiments, long -read sequencing produces contiguous reads of at least 900 nucleotides. In some embodiments, long -read sequencing produces contiguous reads of at least 1000 nucleotides. In some embodiments, long -read sequencing produces contiguous reads of at least 1200 nucleotides. In some embodiments, the long read sequencing is PacBio sequencing. In some embodiments, the long read sequencing is single-molecule real-time (SMRT) sequencing. In some embodiments, the long read sequencing is Loop Genomics sequencing. In some embodiments, the long read sequencing is loop sequencing (LoopSeq). In some embodiments, the long read sequencing is Oxford Nanopore sequencing. In some embodiments, the long read sequencing is nanopore sequencing. In some embodiments, nanopore sequencing is long -read nanopore sequencing. Long -read sequencing is provided by several commercial providers including PacBio, Loop Genomics, Oxford Nanopore Technologies, Illumina, PathQuest, CD Genomics, and Twist Biosciences to name but a few.
[0104] In some embodiments, using the long -read sequencing technology produces full- length 16s rRNA sequences. In some embodiments, the long-read sequencing technology is used to receive full-length 16s rRNA sequences. In some embodiments, the full-length 16srRNA sequences comprises at least 70, 75, 80, 85, 87, 90, 92, 95, 97 or 99% of a 16s rRNA gene. Each possibility represents a separate embodiment of the invention.
[0105] In some embodiments, the method comprises generating a database. In some embodiments, the database is a database of 16s rRNA sequences. In some embodiments, the database is a 16s rRNA database. In some embodiments, the database comprises long 16s rRNA sequences. In some embodiments, the database comprises full-length 16s rRNA sequences. In some embodiments, the database comprises 16s rRNA genes. In some embodiments, the database comprises the long 16s rRNA sequences produced by the long- read sequencing. In some embodiments, the database comprises the long 16s rRNA sequences produced by the long-read sequencing. In some embodiments, the produced long 16s rRNA sequences are combined into a database. In some embodiments, the database comprises known full-length 16s rRNA sequences. In some embodiments, the known sequences are combined with the received / produced sequences to generate the database. In some embodiments, the database is a computerized database. In some embodiments, the method is a computerized method. In some embodiments, the method is a computer implemented method.
[0106] In some embodiments, the method further comprises receiving a sample from the biological niche. In some embodiments, the sample is a new sample. In some embodiments, the sample is a sample of the plurality of samples. In some embodiments, the sample is a sample whose DNA was used for generating the database. In some embodiments, the sample is a sample whose DNA was not used for generating the database. In some embodiments, the sample is a soil sample. In some embodiments, the sample is a plant sample. In some embodiments, the sample is a biological fluid sample. In some embodiments, the sample is a water sample. In some embodiments, the sample is from an organism. In some embodiments, the organism is a non-human organism. In some embodiments, the organism is a human.
[0107] In some embodiments, the method further comprises extracting DNA from the received sample. In some embodiments, DNA from a sample from the niche is received. In some embodiments, the sample is isolated DNA. In some embodiments, the sample comprises bacteria. In some embodiments, the DNA is bacterial DNA.
[0108] In some embodiments, the method comprises amplifying 16s rRNA genes. In some embodiments, the method comprises amplifying 16s rRNA sequences. In some embodiments, the amplifying produces short amplicons. In some short is not more than 750,700, 650, 600, 550, 500, 450, 400, 350, 300, 250, 200, 195, 190, 185, 180, 175, 170, 165, 160, 155 or 150 nucleotides. Each possibility represents a separate embodiment of the invention. In some short is not more than 500 nucleotides. In some short is not more than 300 nucleotides. In some embodiments, a short amplicon is not more than 35, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, or 85% of a 16s rRNA gene. Each possibility represents a separate embodiment of the invention.
[0109] In some embodiments, the amplifying is with a primer pair. In some embodiments, the amplifying is with a plurality of primer pairs. In some embodiments, the amplifying is with a cocktail of primer pairs. In some embodiments, each primer pair produces a short amplicon. In some embodiments, the short amplicons produced by the primer pairs combine to cover at least 70, 75, 80, 86, 87, 90, 92, 95, 97 or 99% of a 16s rRNA gene. Each possibility represents a separate embodiment of the invention. In some embodiments, the primer pairs amplify at least at least 70, 75, 80, 86, 87, 90, 92, 95, 97 or 99% of a 16s rRNA gene. Each possibility represents a separate embodiment of the invention. In some embodiments, the primer pairs combine to amplify at least at least 70, 75, 80, 86, 87, 90, 92, 95, 97 or 99% of a 16s rRNA gene. Each possibility represents a separate embodiment of the invention. In some embodiments, the amplifying produces amplicons that cover at least 80% of a 16s rRNA gene. In some embodiments, the amplifying produces amplicons that cover at least 90% of a 16s rRNA gene.
[0110] In some embodiments, the primers for short amplification are selected from SEQ ID NO: 1-9. In some embodiments, a forward primer comprises or consists of SEQ ID NO: 1.In some embodiments, a forward primer comprises or consists of SEQ ID NO: 3. In some embodiments, a forward primer comprises or consists of SEQ ID NO: 5. In some embodiments, a forward primer comprises or consists of SEQ ID NO: 7. In some embodiments, a forward primer comprises or consists of SEQ ID NO: 9. In some embodiments, a reverse primer comprises or consists of SEQ ID NO: 2. In some embodiments, a reverse primer comprises or consists of SEQ ID NO: 4. In some embodiments, a reverse primer comprises or consists of SEQ ID NO: 6. In some embodiments, a reverse primer comprises or consists of SEQ ID NO: 8. In some embodiments, the primer pair is selected from the group consisting of: SEQ ID NO: 1-2, 1- 4, 1-6, 1-8, 3-2, 3-4, 3-6, 3-8, 5-2, 5-4, 5-6, 5-8, 7-6, 7-8, 9-6 and 9-8. In some embodiments, the plurality of primer pairs comprises SEQ ID NO: 1-9.
[0111] In some embodiments, the primers for short amplification are selected from SEQ ID NO: 29-40. In some embodiments, a forward primer comprises or consists of SEQ ID NO:29. In some embodiments, a forward primer comprises or consists of SEQ ID NO: 31. In some embodiments, a forward primer comprises or consists of SEQ ID NO: 33. In some embodiments, a forward primer comprises or consists of SEQ ID NO: 35. In some embodiments, a forward primer comprises or consists of SEQ ID NO: 37. In some embodiments, a forward primer comprises or consists of SEQ ID NO: 39. In some embodiments, a reverse primer comprises or consists of SEQ ID NO: 30. In some embodiments, a reverse primer comprises or consists of SEQ ID NO: 32. In some embodiments, a reverse primer comprises or consists of SEQ ID NO: 34. In some embodiments, a reverse primer comprises or consists of SEQ ID NO: 36. In some embodiments, a reverse primer comprises or consists of SEQ ID NO: 38. In some embodiments, a reverse primer comprises or consists of SEQ ID NO: 40. In some embodiments, the primer pair is selected from the group consisting of: SEQ ID NO: 29-30, 31-32, 33-34, 35-36, 37-38, and 39-40. In some embodiments, the primer pair is selected from the group consisting of: SEQ ID NO: 29-30, 29-32, 29-34, 29-36, 29-38, 29-40, 31-32, 31-34, 31-36, 31-38, 31-40, 33-34, 33-36, 33-38, 33-40, 35-36, 35-38, 35-40, 37-38, 37-40, and 39-40. In some embodiments, the plurality of primer pairs comprises SEQ ID NO: 29- 40. In some embodiments, the plurality of primer pairs comprises SEQ ID NO: 1-9 and 29- 40.
[0112] In some embodiments, the primers for short amplification further comprise a sequence of a sequencing adapter. In some embodiments, the sequencing adapter is an Illumina sequencing adapter. In some embodiments, a forward primer comprises a forward sequencing adapter. In some embodiments, a reverse primer comprises a reverse sequencing adapter. In some embodiments, the forward sequencing adapter comprises TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG (SEQ ID NO: 16). In some embodiments, the forward sequencing adapter consists of SEQ ID NO: 16. In some embodiments, the reverse sequencing adapter comprisesGTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG (SEQ ID NO: 17). In some embodiments, the reverse sequencing adapter consists of SEQ ID NO: 17.
[0113] In some embodiments, the primers for short amplification further comprise a unique molecular identifier (UMI). In some embodiments, an UMI is a sequence of an UMI. In some embodiments, an UMI is a random barcode sequence. In some embodiments, an UMI is a random sequence. The technology for including random sequence when synthesizing nucleic acid molecules is well known, and thus a set of forward and reverse primers, wherein each contains a different unique sequence, can be produced by a skilled artisan or purchasedcommercially. In some embodiments, the UMI comprises at least 5, 6, 7, 8, 9 or 10 random nucleotides. Each possibility represents a separate embodiment of the invention. In some embodiments, the UMI comprises at least 5 random nucleotides. In some embodiments, the UMI comprises at least 6 random nucleotides. In some embodiments, the UMI comprises 5- 10 random nucleotides. In some embodiments, the UMI comprises 6-10 random nucleotides. In some embodiments, the UMI is 5’ to the primer. In some embodiments, in some embodiments, the UMI is the 5 ’ end of the primer. In some embodiments, the UMI is 5 ’ to the sequence of the primer that hybridizes to the 16s rRNA. In some embodiments, the UMI is between the primer and the sequencing adapter. In some embodiments, the UMI is between the sequence that hybridizes to the 16s rRNA and the sequencing adapter. In some embodiments, hybridizes to is reverse complementary to.
[0114] In some embodiments, the short read amplification is performed in the presence of a blocking oligo. In some embodiments, the long read amplification and the short read amplification are performed in the presence of a blocking oligo. In some embodiments, a blocking oligo is a blocker. In some embodiments, the amplifying of step (e) is performed in the presence of a blocking oligo. In some embodiments, the amplifying of step (b) and the amplification of step (e) are performed in the presence of a blocking oligo. In some embodiments, the method further comprises adding a blocking oligo to the short read amplification. In some embodiments, the method further comprises adding a blocking oligo to the long read amplification and the short read amplification. In some embodiments, the blocking oligo is an ENA oligo. In some embodiments, the blocking oligo is a PNA oligo. In some embodiments, the blocking oligo blocks amplification of 16s rRNA from a nonmi crobial source. In some embodiments, the blocking oligo does not block amplification of 16s rRNA from a microbial source. In some embodiments, a microbial source is a microbe of the community of microbes. In some embodiments, the non-microbial source is a plant. In some embodiments, the non-microbial source is a chloroplast. In some embodiments, the non-microbial source is an organism that is the biological niche. In some embodiments, the non-microbial source is an animal. In some embodiments, the blocking oligo blocks amplification of 16s rRNA from chloroplasts. In some embodiments, the blocking oligo blocks amplification of 16s rRNA from chloroplasts and does not block amplification of a 16s rRNA from a microbe of the community. In some embodiments, the biological niche is a plant and the blocking oligo blocks amplification of 16s rRNA from chloroplasts. In some embodiments, the blocking oligo hybridizes to 16s rRNA from the non-microbial source. In some embodiments, the blocking oligo does not hybridizes to 16s rRNA from a microbialsource. In some embodiments, the blocking oligo does not hybridizes to 16s rRNA from a microbe of the community. In some embodiments, the blocking oligo hybridizes to 16s rRNA from chloroplasts. In some embodiments, the blocking oligo hybridizes to 16s rRNA from chloroplasts and does not hybridize to 16s rRNA from a microbe of the community. In some embodiments, a microbe of the community is all microbes of the community. In some embodiments, hybridizes is perfectly hybridizes. In some embodiments, hybridizes is 100% hybridizes. In some embodiments, perfectly hybridizes is without any mismatches. In some embodiments, the blocking oligo does not hybridize with microbial 16s rRNA. In some embodiments, the blocking oligo does not hybridize with most microbial 16s rRNA. In some embodiments, most is at least 50, 55, 60, 65, 70, 75, 80, 85, 90, 92, 95, 97 or 99% of all microbial 16s rRNAs. Each possibility represents a separate embodiment of the invention. In some embodiments, microbial 16s rRNA is of microbes of the niche. In some embodiments, microbial 16s rRNA is of microbes of the community. In some embodiments, the blocking oligo comprises the sequence CGTCTGTAGGTGG (SEQ ID NO: 28). In some embodiments, the blocking oligo consists of the sequence of SEQ ID NO: 28. In some embodiments, the blocking oligo hybridizes to the V3-V4 region of the non-microbial 16s rRNA. In some embodiments, the blocking oligo hybridizes to the V3-V4 region of chloroplast 16s rRNA. In some embodiments, hybridizes to is reverse complementary to.
[0115] In some embodiments, the blocking oligo is at least 8, 9, 10, 11, 12 or 13 base pairs. Each possibility represents a separate embodiment of the invention. In some embodiments, the blocking oligo is at least 10 base pairs. In some embodiments, the blocking oligo is at least 12 base pairs. In some embodiments, the blocking oligo is at least 13 base pairs. In some embodiments, the blocking oligo comprises at least 8, 9, 10, 11, 12 or 13 base pairs that hybridize with the 16s rRNA from a non-microbial source. Each possibility represents a separate embodiment of the invention. In some embodiments, the blocking oligo is at most 13, 14, 15, 16, 17, 18, 19 or 20 base pairs. Each possibility represents a separate embodiment of the invention. In some embodiments, the blocking oligo is at most 15 base pairs.
[0116] In some embodiments, the amplification is performed after cutting a 16s rRNA sequence from anon-microbial source. In some embodiments, the long read amplification is performed after cutting a 16s rRNA sequence from a non-microbial source. In some embodiments, the short read amplification is performed after cutting a 16s rRNA sequence from a non-microbial source. In some embodiments, the long read amplification and the short read amplification are performed after cutting a 16s rRNA sequence from a non- microbial source. In some embodiments, the amplifying of step (b) is performed after cuttinga 16s rRNA sequence from a non-microbial source. In some embodiments, the amplifying of step (e) is performed after cutting a 16s rRNA sequence from a non-microbial source. In some embodiments, the amplifying of step (b) and the amplifying of step (e) are performed after cutting a 16s rRNA sequence from a non-microbial source. In some embodiments, the method further comprises cutting a 16s rRNA sequence from a non-microbial source. In some embodiments, the cutting is performed before the amplifying. In some embodiments, the cutting is performed before all amplifying steps. In some embodiments, cutting is cleaving. In some embodiments, the cutting is by Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR). In some embodiments, by CRISPR is by a CRISPR-guide RNA (gRNA) complex. In some embodiments, the CRISPR is a CRISPR protein. In some embodiments, a CRISPR protein is a CRISPR associated protein (CAS). In some embodiments, the CAS protein is CAS9.
[0117] In some embodiments, the gRNA comprises a sequence that hybridizes with a 16s rRNA sequence of a non-microbial source. In some embodiments, the gRNA does not comprises a sequence that hybridizes with a 16s rRNA sequence of a microbial source. In some embodiments, a microbial source is a microbe of the niche. In some embodiments, a microbial source is a microbe of the community. In some embodiments, the gRNA does not comprises a sequence that hybridizes with a 16s rRNA sequence of a microbe of the community. In some embodiments, a microbe of the community is all microbes of the community. In some embodiments, hybridizes is perfectly hybridizes. In some embodiments, the gRNA comprises a sequence that hybridizes with a 16s rRNA sequence of a chloroplast. In some embodiments, the gRNA comprises a sequence that hybridizes with a 16s rRNA sequence of a chloroplast and does not hybridize with a 16s rRNA sequence of a microbe of the community. In some embodiments, the biological niche is a plant and the gRNA comprises a sequence that hybridizes with a 16s rRNA sequence of a chloroplast.
[0118] In some embodiments, the CRISPR-gRNA complex recognizes a 16s rRNA sequence from a non-microbial source. In some embodiments, the CRISPR-gRNA complex recognizes a 16s rRNA sequence from a chloroplast. In some embodiments, the CRISPR- gRNA complex does not recognize a 16s rRNA sequence from microbial source. In some embodiments, the CRISPR-gRNA complex does not recognize a 16s rRNA sequence from a microbe of the community. In some embodiments, recognizes is hybridizes to. In some embodiments, recognizes is binds to. In some embodiments, the CRISPR-gRNA complex cuts a 16s rRNA sequence from a non-microbial source. In some embodiments, the CRISPR- gRNA complex cuts a 16s rRNA sequence from a chloroplast. In some embodiments, theCRISPR-gRNA complex does not cut a 16s rRNA sequence from microbial source. In some embodiments, the CRISPR-gRNA complex does not cut a 16s rRNA sequence from a microbe of the community. In some embodiments, a microbe of the community is all microbes of the community.
[0119] In some embodiments, the sequence that hybridizes with a 16s rRNA sequence of a chloroplast is SEQ ID NO: 28. In some embodiments, the gRNA comprises SEQ ID NO: 28. In some embodiments, the gRNA consists of SEQ ID NO: 28. In some embodiments, the seed region of the gRNA comprises or consists of SEQ ID NO: 28. In some embodiments, the sequence that hybridizes with a 16s rRNA sequence of a chloroplast is complementary to SEQ ID NO: 28. In some embodiments, the gRNA comprises or consists of a sequence that is complementary to SEQ ID NO: 28. In some embodiments, complementary is reverse complementary. In some embodiments, the CRISPR-gRNA complex recognizes SEQ ID NO: 28. In some embodiments, the CRISPR-gRNA complex cuts SEQ ID NO: 28. In some embodiments, the sequence is complementary to SEQ ID NO: 28. In some embodiments, complementary is 100% complementary.
[0120] It will be understood that SEQ ID NO: 28 comprises two possible PAMs for CAS9 as the 3’ end of SEQ ID NO: 28 is AGGTGG. Both AGG and TGG are possible PAMs. The gRNA itself does not include the PAM or a sequence complementary to the PAM but rather comprises the sequence just 5’ to the PAM (or a complement thereof). This sequence in the unique chloroplast 16s rRNA sequence is CGTCTGT (SEQ ID NO: 41) for the first PAM and CGTCTGTAGG (SEQ ID NO: 43) for the second PAM. In some embodiments, the CRISPR-gRNA complex recognizes SEQ ID NO: 41. In some embodiments, the sequence that hybridizes with a 16s rRNA sequence of a chloroplast is SEQ ID NO: 41. In some embodiments, the gRNA comprises SEQ ID NO: 41. In some embodiments, the seed region of the gRNA comprises or consists of SEQ ID NO: 41. In some embodiments, the sequence that hybridizes with a 16s rRNA sequence of a chloroplast is complementary to SEQ ID NO: 41. In some embodiments, the gRNA comprises or consists of a sequence that is complementary to SEQ ID NO: 41. In some embodiments, complementary is reverse complementary. In some embodiments, the CRISPR-gRNA complex recognizes SEQ ID NO: 43. In some embodiments, the sequence that hybridizes with a 16s rRNA sequence of a chloroplast is SEQ ID NO: 43. In some embodiments, the gRNA comprises SEQ ID NO: 43. In some embodiments, the seed region of the gRNA comprises or consists of SEQ ID NO: 43. In some embodiments, the sequence that hybridizes with a 16s rRNA sequence of a chloroplast is complementary to SEQ ID NO: 43. In some embodiments, the gRNAcomprises or consists of a sequence that is complementary to SEQ ID NO: 43. In some embodiments, complementary is reverse complementary. It is further known that CAS cleavage occurs between the 3rdand 4thnucleotides upstream of the PAM which will fall in the middle of SEQ ID NO: 41 or 43, which is within SEQ ID NO: 28. In some embodiments, the CRISPR-gRNA complex cuts SEQ ID NO: 41. In some embodiments, the CRISPR- gRNA complex cuts SEQ ID NO: 43.
[0121] As the 3’ end of the gRNA provides the specificity, the inclusion of SEQ ID NO: 41 or 43 ensures that cutting will only occur in chloroplast 16s rRNA and not in microbial 16s rRNA. However, a gRNA is usually at least 17 nucleotides long, so additional sequence complementary to the chloroplast 16s rRNA was added at the 5’ end. In some embodiments, the gRNA comprises the sequence TTGGGCGTAAAGCGTCTGT (SEQ ID NO: 42). In some embodiments, the gRNA consists of SEQ ID NO: 42. In some embodiments, the gRNA comprises the sequence GGCGTAAAGCGTCTGTAGG (SEQ ID NO: 44). In some embodiments, the gRNA consists of SEQ ID NO: 44. In some embodiments, the gRNA comprises the sequence TTGGGCGTAAAGCGTCTGTAGG (SEQ ID NO: 45). In some embodiments, the gRNA consists of SEQ ID NO: 45.
[0122] CRISPR genome editing is a well known technology and any CRISPR enzyme that produces sequence specific cutting may be used. In some embodiments, the CRISPR enzyme is CAS9 or a variant thereof. CAS9 variants are well known the art and include, for example, eSpCas9, SpCas9-HFl, HypaCas9, evoCas9, xCas9 3.7, Sniper-Cas9, SuperFi-Cas9. The CAS protein and the gRNA form a complex that binds to and cuts a target sequence that is complementary to the gRNA. Thus, the gRNA can be used to specifically cut unwanted, non-microbial 16s rRNAs, such as chloroplast rRNAs. In some embodiments, the cr-RNA portion of the gRNA comprises the sequence that hybridizes to the non-microbial 16s rRNA.In some embodiments, the non-microbial 16s rRNA is a chloroplast 16s rRNA. In some embodiments, the cr-RNA portion of the gRNA comprises SEQ ID NO: 28. In some embodiments, the cr-RNA portion of the gRNA comprises SEQ ID NO: 41. In some embodiments, the cr-RNA portion of the gRNA comprises SEQ ID NO: 42. In some embodiments, the cr-RNA portion of the gRNA comprises SEQ ID NO: 43. In some embodiments, the cr-RNA portion of the gRNA comprises SEQ ID NO: 44. In some embodiments, the cr-RNA portion of the gRNA comprises SEQ ID NO: 45. In some embodiments, the cr-RNA portion of the gRNA comprises a complement of SEQ ID NO: 28. In some embodiments, the cr-RNA portion of the gRNA comprises a complement of SEQ ID NO: 41. In some embodiments, the cr-RNA portion of the gRNA comprises acomplement of SEQ ID NO: 42. In some embodiments, the cr-RNA portion of the gRNA comprises a complement of SEQ ID NO: 43. In some embodiments, the cr-RNA portion of the gRNA comprises a complement of SEQ ID NO: 44. In some embodiments, the cr-RNA portion of the gRNA comprises a complement of SEQ ID NO: 45. In some embodiments, the gRNA is the crRNA. In some embodiments, the gRNA is a single guide RNA (sgRNA). In some embodiments, the seed of the crRNA comprises or consists of SEQ ID NO: 28. In some embodiments, the seed of the crRNA comprises or consists of SEQ ID NO: 41. In some embodiments, the seed of the crRNA comprises or consists of SEQ ID NO: 42. In some embodiments, the seed of the crRNA comprises or consists of SEQ ID NO: 43. In some embodiments, the seed of the crRNA comprises or consists of SEQ ID NO: 44. In some embodiments, the seed of the crRNA comprises or consists of SEQ ID NO: 45. In some embodiments, the sequence in the gRNA that is SEQ ID NO: 28 comprises the seed region. In some embodiments, the sequence in the gRNA that is SEQ ID NO: 41 comprises the seed region. In some embodiments, the sequence in the gRNA that is SEQ ID NO: 42 comprises the seed region. In some embodiments, the sequence in the gRNA that is SEQ ID NO: 43 comprises the seed region. In some embodiments, the sequence in the gRNA that is SEQ ID NO: 44 comprises the seed region. In some embodiments, the sequence in the gRNA that is SEQ ID NO: 45 comprises the seed region. In some embodiments, SEQ ID NO: 28 is the 3’ end of the gRNA. In some embodiments, SEQ ID NO: 28 is the 3’ end of the crRNA. In some embodiments, SEQ ID NO: 41 is the 3’ end ofthe gRNA. In some embodiments, SEQ ID NO: 41 is the 3’ end of the crRNA. In some embodiments, SEQ ID NO: 42 is the 3’ end of the gRNA. In some embodiments, SEQ ID NO: 42 is the 3’ end of the crRNA. In some embodiments, SEQ ID NO: 43 is the 3’ end of the gRNA. In some embodiments, SEQ ID NO: 43 is the 3’ end of the crRNA. In some embodiments, SEQ ID NO: 44 is the 3’ end of the gRNA. In some embodiments, SEQ ID NO: 44 is the 3’ end of the crRNA. In some embodiments, SEQ ID NO: 45 is the 3’ end of the gRNA. In some embodiments, SEQ ID NO: 45 is the 3’ end of the crRNA. It will be understood that while the sequences provided herein are provided with “T” bases, the RNA will actually contain “U” bases. In some embodiments, the method further comprises inactivating the CRISPR enzyme after the cutting.
[0123] In some embodiments, the sequencing is using short read sequencing technology. In some embodiments, short read sequencing technology is second generation sequencing. In some embodiments, short read sequencing technology is deep sequencing technology. In some embodiments, the sequencing is next generation sequencing. In some embodiments,short read sequencing technology is massively parallel sequencing. Short next generation sequencing is well known in the art and can be done through many commercially available methods and vendors. Illumina, SWIFT sequencing and 454 sequencing can be used for short read generation.
[0124] In some embodiments, short-read sequencing comprises contiguous sequence reads ofnot longer than 500, 450, 400, 350, 300, 250, 200, 195, 190, 185, 180, 175, 170, 165, 160, 155 or 150 nucleotides. Each possibility represents a separate embodiment of the invention. In some embodiments, short-read sequencing comprises contiguous sequence reads of not longer than 200 nucleotides. In some embodiments, short-read sequencing comprises contiguous sequence reads of not longer than 300 nucleotides. In some embodiments, the short-read sequences are amplicon reads.
[0125] In some embodiments, the method comprises combining the short reads into long 16s rRNA sequences. In some embodiments, the combining is computational combining. In some embodiments, the combining is by a computer. In some embodiments, the combining produces synthetic long reads. In some embodiments, the combining is based on the database. In some embodiments, the combining is calculated using the database as a list of templates for the combination. In some embodiments, a produced combination is a 16s rRNA gene. In some embodiments, the produced combination is a unique 16s rRNA gene. In some embodiments, the 16s rRNA gene is a gene within the sample. In some embodiments, the 16s rRNA gene is a gene within the niche.
[0126] In some embodiments, combining comprises performing Short Multi-Region Framework (SMURF) analysis. SMURF is a tool for combining sequencing results from different PCR-amplified regions to provide one coherent profding. The methodology is provided in Fuks et al., “Combining 16S rRNA gene variable regions enables high-resolution microbial community profiling”, Microbiome, 2018 Jan 26;6(1): 17, which is hereby incorporated by reference in its entirety. In some embodiments, the combining is to produce a sequence that starts at the most 5 ’ long amplification primer used and that ends at the most 3’ long amplification primer used. In some embodiments, the combining is to produce a sequence that starts at the most 5 ’ ENA primer used and that ends at the most 3 ’ LNA primer used. In some embodiments, the SMURF analysis is based on the sequence from the most 5’ to the most 3’ primers used for the long amplicon amplification.
[0127] In some embodiments, combining comprises grouping produced 16s rRNA sequences. In some embodiments, the grouping is bases on sequence similarity. In someembodiments, 16s rRNA sequences with the same sequence are grouped. In some embodiments, the grouping produces a list of unique 16s rRNA genes. In some embodiments, the list is the unique 16s rRNA genes in the niche. In some embodiments, the list is the unique 16s rRNA genes within the samples. In some embodiments, the list is a list of the unique organisms within the niche. In some embodiments, the list is a list of unique organisms within the samples.
[0128] In some embodiments, a produced 16s rRNA gene identifies a specific organism. In some embodiments, the method further comprises identifying an organism as being in the niche based on the presence of the organism’s 16s rRNA gene being present in the sample. In some embodiments, an organism is a microbe. In some embodiments, an organism is a bacterium. In some embodiments, a plurality of different 16s rRNA genes are produced and each different gene indicates the presence of a different organism within the niche. Thus, by identifying different 16s rRNA genes within the niche a skilled artisan is actually identifying different organisms within the niche. In some embodiments, the method further comprises correlating each identified 16s rRNA gene to an organism. In some embodiments, the method thereby identifies unique 16s rRNA genes within the niche. In some embodiments, the method thereby profiles a community of organisms within the niche. In some embodiments, the community is a microbial community.
[0129] In some embodiments, the method further comprises selecting the primer pairs to provide a sufficient resolution for the 16s rRNA genes within the niche. In some embodiments, the method further comprises selecting the primer pairs to provide a sufficient resolution for the 16s rRNA genes within the samples. In some embodiments, selecting a primer pair is selecting a subset of primer pairs from the plurality of primer pairs. In some embodiments, the subset provides sufficient resolution for the 16s rRNA genes within the niche. In some embodiments, the subset provides sufficient resolution for the 16s rRNA genes within the samples. In some embodiments, the subset is selected to reduce ambiguity between 16s rRNA genes. In some embodiments, the subset is selected to optimize the determination of the identity of the 16s rRNA genes. In some embodiments, the subset is selected to optimize the correlation between 16s rRNA genes. In some embodiments, the subset is selected to optimize a parameter selected from ambiguity, correlation and identity. In some embodiments, the subset is selected to optimize at least two parameters selected from ambiguity, correlation and identity. In some embodiments, the subset is selected to optimize ambiguity, correlation and identity.
[0130] By another aspect, there is provided a LNA primer pair comprising an LNA forward primer comprising SEQ ID NO: 10 and an LNA reverse primer comprising SEQ ID NO: 11.
[0131] By another aspect, there is provided a LNA primer pair comprising an LNA forward primer comprising SEQ ID NO: 12 and an LNA reverse primer comprising SEQ ID NO: 13.
[0132] By another aspect, there is provided a LNA primer pair comprising an LNA forward primer comprising SEQ ID NO: 14 and an LNA reverse primer comprising SEQ ID NO: 15.
[0133] In some embodiments, the primers are primers as described hereinabove. In some embodiments, the forward primer comprises a sequence comprising SEQ ID NO: 10. In some embodiments, the forward primer comprises a sequence consisting of SEQ ID NO: 10. In some embodiments, the forward primer consists of SEQ ID NO: 10. In some embodiments, the reverse primer comprises a sequence comprising SEQ ID NO: 11. In some embodiments, the reverse primer comprises a sequence consisting of SEQ ID NO: 11. In some embodiments, the reverse primer consists of SEQ ID NO: 11.
[0134] In some embodiments, the forward primer comprises a sequence comprising SEQ ID NO: 12. In some embodiments, the forward primer comprises a sequence consisting of SEQ ID NO: 12. In some embodiments, the forward primer consists of SEQ ID NO: 12. In some embodiments, the reverse primer comprises a sequence comprising SEQ ID NO: 13. In some embodiments, the reverse primer comprises a sequence consisting of SEQ ID NO: 13. In some embodiments, the reverse primer consists of SEQ ID NO: 13.
[0135] In some embodiments, the forward primer comprises a sequence comprising SEQ ID NO: 14. In some embodiments, the forward primer comprises a sequence consisting of SEQ ID NO: 14. In some embodiments, the forward primer consists of SEQ ID NO: 14. In some embodiments, the reverse primer comprises a sequence comprising SEQ ID NO: 15. In some embodiments, the reverse primer comprises a sequence consisting of SEQ ID NO: 15. In some embodiments, the reverse primer consists of SEQ ID NO: 15.
[0136] By another aspect, there is provided a kit comprises a plurality of LNA primer pairs of the invention.
[0137] In some embodiments, the kit comprises a first primer pair comprising an LNA forward primer comprising SEQ ID NO: 10 and an LNA reverse primer comprising SEQ ID NO: 11 and a second primer pair comprising an LNA forward primer comprising SEQ ID NO: 12 and an LNA reverse primer comprising SEQ ID NO: 13. In some embodiments, the kit comprises a first primer pair comprising an LNA forward primer comprising SEQ IDNO: 10 and an LNA reverse primer comprising SEQ ID NO: 11 and a second primer pair comprising an LNA forward primer comprising SEQ ID NO: 14 and an LNA reverse primer comprising SEQ ID NO: 15. In some embodiments, the kit comprises a first primer pair comprising an LNA forward primer comprising SEQ ID NO: 14 and an LNA reverse primer comprising SEQ ID NO: 15 and a second primer pair comprising an LNA forward primer comprising SEQ ID NO: 12 and an LNA reverse primer comprising SEQ ID NO: 13. In some embodiments, the kit comprises a first primer pair comprising an LNA forward primer comprising SEQ ID NO: 10 and an LNA reverse primer comprising SEQ ID NO: 11, a second primer pair comprising an LNA forward primer comprising SEQ ID NO: 12 and an LNA reverse primer comprising SEQ ID NO: 13 and a third primer pair comprising an LNA forward primer comprising SEQ ID NO: 14 and an LNA reverse primer comprising SEQ ID NO: 15.
[0138] In some embodiments, the kit further comprises reagents for performing amplification. In some embodiments, the kit further comprises reagents for performing sequencing. In some embodiments, the reagent is a polymerase. In some embodiments, the reagent is oligonucleotides. In some embodiments, the reagent is an amplification buffer in which the polymerase is active. In some embodiments, the reagent is a sequencing adapter.
[0139] In some embodiments, the kit further comprises a blocking oligo. In some embodiments, the blocking oligo is a chloroplast 16s rRNA blocking oligo. In some embodiments, the blocking oligo is a blocking oligo as described hereinabove. In some embodiments, the blocking oligo is an LNA blocking oligo. In some embodiments, the blocking oligo is a PNA blocking oligo. In some embodiments, the blocking oligo comprises a sequence that hybridizes to a 16s rRNA of a chloroplast.
[0140] In some embodiments, the kit further comprises a gRNA. In some embodiments, the gRNA is a chloroplast gRNA. In some embodiments, the gRNA is as described hereinabove. In some embodiments, the gRNA comprises a sequence that hybridizes to a 16s rRNA of a chloroplast. In some embodiments, the kit further comprises a CRISPR protein. In some embodiments, the kit further comprises a nucleic acid molecule encoding a CRISPR protein. In some embodiments, the nucleic acid molecule is a plasmid. In some embodiments, the plasmid is an expression plasmid. In some embodiments, the nucleic acid molecule comprises an open reading frame. In some embodiments, the open reading frame encodes the CRISPR protein. In some embodiments, the open reading frame is operably linked to a promoter. In some embodiments, the CRISPR protein is CAS9. The term “operably linked” is intended to mean that the nucleotide sequence of interest is linked to the regulatoryelement or elements in a manner that allows for expression of the nucleotide sequence (e.g. in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell).
[0141] In some embodiments, hybridizes to is perfectly hybridizes to. In some embodiments, hybridizes to is complementary to. In some embodiments, complementary to is perfectly complementary to. In some embodiments, complementary to is reverse complementary to. In some embodiments, the sequence that hybridizes to a 16s rRNA sequence of a chloroplast comprises SEQ ID NO: 28. In some embodiments, the sequence that hybridizes to a 16s rRNA sequence of a chloroplast consists of SEQ ID NO: 28.
[0142] By another aspect, there is provided a method of depleting chloroplast 16s rRNA from a solution, the method comprising contacting the solution with a CRISPR-guide RNA (gRNA) complex wherein the CRISPR-gRNA complex recognizes SEQ ID NO: 41, SEQ ID NO: 43 or SEQ ID NO: 28, thereby depleting chloroplast 16s rRNA from a solution.
[0143] By another aspect, there is provided a method of depleting chloroplast 16s rRNA from a solution, the method comprising amplifying 16s rRNA genes within the solution in the presence of a blocking oligo, wherein the blocking oligo hybridizes to or comprises SEQ ID NO: 28, thereby depleting chloroplast 16s rRNA from a solution.
[0144] In some embodiments, depleting is removing. In some embodiments, the method is a method of depleting chloroplast contamination from a solution. In some embodiments, depleting contamination is depleting chloroplast 16s rRNA contamination. In some embodiments, the solution is a sample. In some embodiments, the solution is a solution comprising RNA. In some embodiments, the RNA is rRNA. In some embodiments, the RNA is 16s rRNA. In some embodiments, the method further comprises inactivating the CRISPR enzyme after the cutting. In some embodiments, the solution comprises RNA isolated from a biological niche. In some embodiments, the biological niche is a plant niche. In some embodiments, the solution comprises RNA isolated from a plant. It will be understood that while it is the sequence from the rRNA that is cut, the method would also apply to cDNA reverse-transcribed from the rRNA. Thus, the solution may also comprises cDNA from the rRNA. In some embodiments, the solution comprises cDNA from the RNA and the method is a method of depleting cDNA from chloroplast 16s rRNA.
[0145] In some embodiments, the CRISPR-gRNA complex recognizes SEQ ID NO: 41. In some embodiments, the CRISPR-gRNA complex recognizes SEQ ID NO: 43. In some embodiments, the CRISPR-gRNA complex recognizes SEQ ID NO: 28. In someembodiments, the CRISPR-gRNA complex cuts SEQ ID NO: 41, SEQ ID NO: 43 or SEQ ID NO: 28. In some embodiments, the CRISPR-gRNA complex cuts SEQ ID NO: 41. In some embodiments, the CRISPR-gRNA complex cuts SEQ ID NO: 43. In some embodiments, the CRISPR-gRNA complex cuts SEQ ID NO: 28. In some embodiments, the gRNA comprises SEQ ID NO: 42. In some embodiments, the gRNA comprises SEQ ID NO: 44. In some embodiments, the gRNA comprises SEQ ID NO: 45. In some embodiments, the gRNA is a gRNA such as is described hereinabove.
[0146] In some embodiments, the blocking oligo comprises SEQ ID NO: 28. In some embodiments, the blocking oligo hybridizes to SEQ ID NO: 28. In some embodiments, the blocking oligo is complementary to SEQ ID NO: 28. In some embodiments, hybridizes is perfectly hybridizes. In some embodiments, complementary is 100% complementary. In some embodiments, the blocking oligo consists of SEQ ID NO: 28. In some embodiments, the blocking oligo consists of a 100% complement of SEQ ID NO: 28. It will be understood that the blocking oligo can hybridizes to either stand and thereby block amplification. In some embodiments, the blocking oligo is an LNA blocking oligo. In some embodiments, the blocking oligo is a PNA blocking oligo. In some embodiments, the blocking oligo is a blocking oligo as is described hereinabove.
[0147] In some embodiments, the amplifying is short amplicon amplifying. In some embodiments, the amplifying is long amplicon amplifying. In some embodiments, the amplifying comprises PCR. In some embodiments, the amplifying is amplifying as is described hereinabove. In some embodiments, the primers used for the amplifying are primers as are described hereinabove.
[0148] As used herein, the term "about" when combined with a value refers to plus and minus 10% of the reference value. For example, a length of about 1000 nanometers (nm) refers to a length of 1000 nm+- 100 nm.
[0149] It is noted that as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a polynucleotide" includes a plurality of such polynucleotides and reference to "the polypeptide" includes reference to one or more polypeptides and equivalents thereof known to those skilled in the art, and so forth. It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as "solely," "only" and the like in connection with the recitation of claim elements, or use of a "negative" limitation.
[0150] In those instances where a convention analogous to "at least one of A, B, and C, etc." is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (e.g., "a system having at least one of A, B, and C" would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). It will be further understood by those within the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms. For example, the phrase "A or B" will be understood to include the possibilities of "A" or "B" or "A and B."
[0151] It is appreciated that certain features of the invention, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable sub-combination. All combinations of the embodiments pertaining to the invention are specifically embraced by the present invention and are disclosed herein just as if each and every combination was individually and explicitly disclosed. In addition, all subcombinations of the various embodiments and elements thereof are also specifically embraced by the present invention and are disclosed herein just as if each and every such sub-combination was individually and explicitly disclosed herein.
[0152] As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents, unless the context clearly dictates otherwise. The terms “a” (or “an”) as well as the terms “one or more” and “at least one” can be used interchangeably.
[0153] Furthermore, “and / or” is to be taken as specific disclosure of each of the two specified features or components with or without the other. Thus, the term “and / or” as used in a phrase such as “A and / or B” is intended to include A and B, A or B, A (alone), and B (alone). Likewise, the term “and / or” as used in a phrase such as “A, B, and / or C” is intended to include A, B, and C; A, B, or C; A or B; A or C; B or C; A and B; A and C; B and C; A (alone); B (alone); and C (alone).
[0154] Wherever embodiments are described with the language “comprising,” otherwise analogous embodiments described in terms of “consisting of’ and / or “consisting essentially of’ are included.
[0155] Additional objects, advantages, and novel features of the present invention will become apparent to one ordinarily skilled in the art upon examination of the following examples, which are not intended to be limiting. Additionally, each of the various embodiments and aspects of the present invention as delineated hereinabove and as claimed in the claims section below finds experimental support in the following examples.
[0156] V arious embodiments and aspects of the pre sent invention as delineated hereinabove and as claimed in the claims section below find experimental support in the following examples.EXAMPLES
[0157] Generally, the nomenclature used herein and the laboratory procedures utilized in the present invention include molecular, biochemical, microbiological and recombinant DNA techniques. Such techniques are thoroughly explained in the literature. See, for example, "Molecular Cloning: A laboratory Manual" Sambrook et al., (1989); "Current Protocols in Molecular Biology" Volumes I-III Ausubel, R. M., ed. (1994); Ausubel et al., "Current Protocols in Molecular Biology", John Wiley and Sons, Baltimore, Maryland (1989); Perbal, "A Practical Guide to Molecular Cloning", John Wiley & Sons, New York (1988); Watson et al., "Recombinant DNA", Scientific American Books, New York; Birren et al. (eds) "Genome Analysis: A Laboratory Manual Series", Vols. 1-4, Cold Spring Harbor Laboratory Press, New York (1998); methodologies as set forth in U.S. Pat. Nos. 4,666,828; 4,683,202; 4,801,531; 5,192,659 and 5,272,057; "Cell Biology: A Laboratory Handbook", Volumes I- III Cellis, J. E., ed. (1994); "Culture of Animal Cells - A Manual of Basic Technique" by Freshney, Wiley-Liss, N. Y. (1994), Third Edition; "Current Protocols in Immunology" Volumes I-III Coligan J. E., ed. (1994); Stites et al. (eds), "Basic and Clinical Immunology" (8th Edition), Appleton & Lange, Norwalk, CT (1994); Mishell and Shiigi (eds), "Strategies for Protein Purification and Characterization - A Laboratory Course Manual" CSHL Press (1996); all of which are incorporated by reference. Other general references are provided throughout this document.Materials and Methods
[0158] Samples: Overall, 33 individual samples of plants and their surroundings, one mock community sample, two sample pools derived from multiple plants and their surroundings (collected in distinct locations in Israel), two sample pools derived from sponges werecollected in Israel and France, and another pool composed of all 33 individual samples used in the study. The following subsection describes the samples in each experiment.
[0159] Qiagen-processed samples: Leaf, soil and root samples of the following plants were collected: Aster poliothamnus, Brandisici hancei, Tetrcienci mongolica, Lathyrus sativus and Angelica hirsutiflora. In addition, two Lepidium apetalum plants were collected. In total, 24 samples were processed.
[0160] Samples for pooled LNA amplification and PacBio long read sequencing: Four pools were created, each by combining extracted DNA from relevant samples (extraction procedure appears below):
[0161] Spongia officinalis France pool: Samples from seven different individuals were collected along the coast of Marseille, France.
[0162] Spongia officinalis Israel pool: Samples from three different individuals were collected along the coast of Haifa, Israel.
[0163] Arava desert Israel pool: Leaf, root and surrounding soil samples from the following Arava desert species were pooled - Moringa oleifera, Pulicaria incisa, Ochradenus baccatus, Anabasis articulata, Salsola Vermiculata, Commiphora gileadensis, Ochradenus baccatus, Fagonia mollis, Zilla spinosa, and Astersicus gravelonis.
[0164] Temperate lowlands Israel pool: Leaf, root and surrounding soil samples from the following lowlands region species were collected near Rehovot, Israel - Retama raetam, Ferula communis, Vicia villosa, Heterotheca subaxillaris, Sarcopoterium spinosum, Prasium majus and Withania somnifera.
[0165] Swift-processed samples: Leaf, soil and root samples of the following plants were collected: Hymelaea hirsuta, Solanum lycopersicum and Triticum aestivum.
[0166] Swift-processed samples - mock community: A mock community of 12 species derived from environmental isolates - Acidovorax citrulli, Bacillus subtilis, Paenibacillus dendritiformis, Clavibacter michiganensis, Escherichia coli, Zymomonas mobilis, Proteus mirabilis, Pseudomonas syringae, Xanthomonas campestris, Agrobacterium tumefaciens, Sphingomonas yanoikuyae, and Bacillus pumilus was created. Species identification was performed by Sanger sequencing using E8f-939r as primers.
[0167] Samples for pooled LNA amplification and Loop Genomics synthetic long reads sequencing: A pool was created for LNA amplification and Loop Genomics synthetic long read sequencing by combining 11 different libraries: A mixture of the 24 samples that wereprepared using the Qiagen QIAseq kit; one library per each of the 9 samples that were prepared using the Swift 16s+ITS PANEL kit; and one library for the mock community.
[0168] Sample Collection and Preservation: Plant samples were collected in the field and kept in a nucleic acid protection buffer. Sponge samples were collected and immediately preserved in absolute ethanol. Upon arrival at the lab, samples were kept at -80°C until DNA extraction.
[0169] DNA Extraction: Genomic DNA was extracted from lOOmg of the root and leaf samples, 250mg soil samples, and an approximately 1cm x 0.5cm size piece from the sponges using ZymoBIOMICS DNA Kit (Cat. D4300) based on the manufacturer's protocol.
[0170] For the mock community, DNA was extracted from each isolate using GenElute Bacterial Genomic DNA Kits (Cat. NA2110-1KT), and 20ng from each species was used to create a mock community. Pools were created by adding 50ng from each amplicon post LNA amplification and cleaning up to generate a 500ng pool.
[0171] Library Preparation and Sequencing: The following sections describe sample preparation and sequencing procedures used for each method.
[0172] Qiagen-Processed Samples: A total of 24 samples were processed. Libraries for deep sequencing were prepared using the QIAseq 16s / ITS Region Panel protocol (Cat. 333842), which amplifies the V1V2, V2V3, V3V4, V4V5, V5V7, and V7V9 regions (primer sequences may be available upon request from Qiagen). Sequencing was performed on an Illumina MiSeq using 250|8|8|250 cycles, and the resulting fastq files were analyzed using Qiagen Genomics Workbench software with the CLC Microbial Genomics module.PacBio Sequencing
[0173] 16s rRNA full-length amplicons, obtained from the LNA amplification of the Spongia officinalis microbial community (France and Israel) and of microbial communities in plants from Israel’s lowland and Arava, were used as a starting template for the SMRTbell Express template Prep Kit 2.0 kit (Product PN. 100-938-900). Sequencing was done on the PacBio Sequel sequencer. An initial analysis step of demultiplexing, filtering, and removal of chimeric reads was performed using the SMRT analysis software.
[0174] Loop Genomics Sequencing: Samples used for validation of CoSMIC method (Fig. 5) were sequenced using Loop Genomics. 16s rRNA full-length amplicons were used as a template for the LoopSeq PCR Amplicon Kit. Sequencing was done on the Illumina NovaSeq sequencer using the following cycles 151|8|8|151, resulting in 2.5-17 Mbp readsper sample. Analysis was performed on the company’s remote server, yielding fasta fdes containing assembled long reads.
[0175] Shotgun metagenomics sequencing: Extracted DNA from each sample was used with the Nextera DNA Flex Library Prep (Cat. 20018704) to prepare the libraries for metagenomic sequencing. Libraries were sequenced on the Illumina NovaSeq sequencer using the following cycles 151|8|8|151 resulting in 20-80 Mbp reads per sample.
[0176] SWIFT-Processed Samples: Libraries for deep sequencing were prepared with the SWIFT AMPLICON 16s+ITS PANEL (Cat. AL-51696), using 5 and 4 forward and reverse primers, respectively, pooled in a single reaction (see Table 1). Sequencing was performed on the Illumina NovaSeq sequencer with the following cycles: 151|8|8| 151 yielding 2.5-15 Mbp reads per sample. Base calling and demultiplexing were performed using Illumina’s bcl2fastq command line interface.
[0177] Table 1. Primers used in the SWIFT AMPLICON 16s+ITS PANEL. Primers applied to amplify the variable regions along the 16s gene as described in the kit’s protocol.
[0178] Enriching the database with novel 16s rRNA genes.
[0179] Designing LNA primers for full-length 16s rRNA gene: Using a refined version on the Greengenes database containing 1.4M 16s rRNA genes, short primer pairs (of length 10- I3nt) were sought that maximize the number of amplified 16s rRNA genes. Considering only perfect matches while ignoring thermodynamic considerations, three candidate pairs were identified that potentially amplify 1.2M out the 1.4M 16s rRNA genes (Table 2). Candidate pairs were provided to Qiagen and LNA primers were produced by standard protocol.
[0180] Table 2. LNA primers. A. Three primer pairs (pairs 1-3) and their location along the E. coli 16s rRNA. Primers are available via the Qiagen catalog number. B. The number ofamplified 16s rRNA genes by each pair and by their intersection. About 360K are amplified by all three pairs; 570K are amplified by two pairs; and 250K 16s rRNA genes are amplified by just a single pair.A.B.
[0181] Sample preparation: To augment the database with novel 16s rRNA genes the LNA primer set was applied. The three forward primers and the three reverse primers were combined into two separate mixes at a stock concentration of 20pM.
[0182] Samples were individually amplified in an input concentration of 10-5 Ong. Each sample was amplified using the following volumes : 25 pl Kapa HiFi HotStart Ready Mix (cat. KK2602), 1.25 pl forward primer pool, and 1.25 pl reverse primer pool (each to a final concentration of 0.5pM), 50ng template, and DDW up to 50pl. PCR plan used: 95°C - 3 minutes, 2 cycles: 98°C - 20 seconds, 55°C - 15 seconds, 72°C - 1 minute; 24 cycles: 98°C - 20 seconds, 58°C - 15 seconds, 72°C - 1 minute; 72°C - 5 minutes, hold at 4°C. Following amplification, the products were cleaned using Ampure XP beads (Cat. A63881) at 0.6X concentration. The relevant samples from each niche were pooled at 130ng each.
[0183] Combinatorial estimation of sets of variable regions: SMURF was independently applied to the 63 combinations of the 6 amplified regions produced by the QIAseq 16s / ITS Region Panel. The performance of each combination was evaluated by: (a) ambiguity - the number of 16s rRNA genes in each SMURF-detected group, i.e., in each set of 16s rRNA genes having the same sequence over the amplified regions; (b) mutual species - the proportion of correctly identified species, i.e., the number of headers that appear in both groups (six regions-fiill, and region combination-partial) this parameter would be 1 if the partial combination has all the headers and 0 if they share none with the full six-region analysis; and (c) Pearson correlation between the correctly identified mutual species. A Jupyter Notebook tutorial is available (GitHub rezenman / CoSMIC).
[0184] Defining novel 16s rRNA genes following the sequencing of LNA amplicons: To evaluate the similarity to the SILVA database, candidate 16s rRNA genes were clustered within each experiment using CD-HIT at 99% identity to remove redundant 16s rRNA genes. Non-redundant sequences were compared to the SILVA database using the cd-hit- est-2d function to determine the fraction of 16s rRNA genes that are unmatched to SILVA at different identity levels (Fig. 3A). Script for this analysis appears in the GitHub repository “rezenman / CoSMIC”.
[0185] SWIFT reads analysis: Standard analysis: Microbial profiling was performed by the standard Swift analysis scripts, which are based on the Qiime2 package.
[0186] CoSMIC analysis: Each sample was demultiplexed to its variable regions using cutadapt. In principle, since 5 forward primers were mixed with 4 reverse primers, SMURF can be applied to up to 20 specific regions. 10 potential amplicons were selected (Table 3) while ignoring long amplicons (e.g., the pair Vl_f and V9_r was omitted) since the number of such reads was negligible.
[0187] Reads that matched the relevant primer pairs (i.e., their R1 and R2 prefixes matched these primer pairs) were filtered and analyzed using R’s DADA2 package to account for sequencing inaccuracies and error-corrected reads were submitted to SMURF. Scripts and usage instructions appear in the GitHub repository rezenman / CoSMIC.
[0188] Table 3. SWIFT primer pairs considered in SMURF. Multiple amplicon types are generated in a multiplex amplification of the SWIFT 9 primer pairs. SMURF analysis considered 10 forward-reverse pairs of adjacent regions. For example, two amplicon types starting with Vl_f were considered, i.e., those that end with V4_r or V5_r.
[0189] Generating ground truth for benchmarking CoSMIC to other methods: Metagenomic reads were analyzed using PhyloFlash together with the Loop-genomics long reads that were used as trusted contigs. The resulting community composition served as the ground truth for comparing existing analysis pipelines (Swift all and Swift V3V4) to the suggested CoSMIC pipeline. For each sample, reads aligned to the SILVA 16s rRNA gene database (after removing LSU sequences) were kept. Those reads were aligned to one of the following: (a) 16s rRNA genes extracted from the metaSPAdes assembly of each sample, (b) Full-length 16s rRNA gene sequences obtained by Loop Genomics sequencing of the LNA amplicons; or (c) SILVA database of 16s rRNA genes, for the reads that were unassigned to the two previous sources. The complete list of 16s rRNA genes served as the ground truth for CoSMIC evaluation purposes.
[0190] Calculating ambiguity, alignment identity, explained frequency & erroneous 16s rRNA genes: The output of SMURF is a set of groups, where each group contains 16s rRNA genes whose sequence is identical over the amplified regions yet vary over the rest of the gene. An analogous definition of a group was created for standard analysis based on the SILVA database, in which all relevant 16s rRNA genes that share the same taxonomy as decided by the Qiime2 pipeline were extracted.This definition is equivalent to SMURF’s groups as all 16s rRNA genes sharing the same taxonomy are indistinguishable based on the reads provided. These groups were used for the comparisons below:
[0191] Ambiguity: To calculate ambiguity (Fig. 5B), the number of indistinguishable (equally likely to be the relevant species) matches to the SILVA database were counted for each identification. For example, if a hit to the SILVA database was to the genus level, ambiguity was defined as the number of SILVA SSUs for that genus. To resolve ambiguity for later comparison to the gtl6s, for each identification, all possible hits were clustered to the database using cd-hit and the most frequent seed was used as a representative.
[0192] Alignment identity: To evaluate the method's performance against the gtl6s, all seeds produced were aligned by each method to its corresponding gtl6s list using blast, and its score serves as the alignment identity.
[0193] Explained frequency: The relative proportion originating from the correctly aligned 16s rRNA genes for a given alignment score.
[0194] Erroneous 16s rRNA genes: Percent of declared 16s rRNA genes not present in the gtl6s.Example 1: The outcome of microbial profiling highly depends on amplified regions
[0195] To demonstrate the effect of preselecting a single variable region, 16s rRNA profiling was performed of leaves, roots, and their surrounding soil, from 7 different plants (see Materials and Methods). Profiling was performed using a commercial Qiagen kit (QIAseq 16s / ITS Region Panel) that amplifies six regions along the gene (Fig. 1A; see Materials and Methods) resulting in 200K-600K reads per sample. Qiagen’s accompanying CLC Genomics Microbiome package was then applied to analyze the results, thus providing a set of Operational Taxonomic Units (OTUs) and their taxonomy for each region and sample.
[0196] Figure IB illustrates the number of OTUs detected in leaf, soil, and root samples of Aster poliothamnus using each of the six amplified regions. For example, the same root sample results in 600-1100 OTUs, depending on the amplified region (Fig. IB, right). Moreover, there is a low overlap in assigned taxonomies detected by different regions, as evident by the beta diversity among all pairs of amplified regions (Fig. 1C). Diversity was calculated by a Jaccard distance and is higher than zero for any pair of amplified regions (median Jaccard distances of 0.27, 0.30 and 0.28 for leaf, root, and soil, respectively).
[0197] To evaluate the contribution of each amplicon to the analysis of a given sample type, the amplicon was identified as having the largest number of unique OTUs and regions were sequentially added according to their additional unique OTUs with respect to regions selected thus far (Fig. 1D-1F). The “top” region varied across samples, e.g., the most “informative” regions for leaf, soil, and root in the case of Aster poliothamnus were V3V4, V1V2, and V4V5, respectively. In addition, it required four regions to exhaust the number of unique OTUs, yet the identity of these regions varied among leaf, soil, and root samples, thus highlighting the inherent biases in region selection. Comparable observations were made when profiling other plants - Brandisici hancei (Fig. 6A-6E), Lepidium apetcilum (Fig. 7A-7E), Lathyrus sativus (Fig. 8A-8E), Angelica hirsutiflora (Fig. 9A-9E), and Tetraena mongolica (Fig. 10A-10E), while the most "informative" amplicon was inconsistent both across plants for a given niche (e.g., leaves across all plants), and over niches for each plant (Fig. 11A-11C).Example 2: Combining results from several variable regions partially overcomes problems in profiling
[0198] There was previously suggested the use of the Short MUlti -Region Framework (SMURF) approach to avoid issues related to pre-selecting a single region. In SMURF, any combination of regions is amplified, and their sequencing results are computationally combined, resulting in a single coherent profiling solution. SMURF relies on a database of full-length 16s rRNA gene, i.e., SSUs, to solve an optimization problem, seeking the 16s rRNA gene and their relative frequencies that have given rise to the set of reads from the relevant regions. The output of SMURF is a set of “groups” and their relative frequencies, where each group contains 16s rRNA genes that have an identical sequence over the amplified regions. For example, if V1V2 and V7V9 regions are being amplified, 16s rRNA gene that have the same sequence over these two regions are, therefore, indistinguishable and, if detected by SMURF, would be part of the same group. Apart from circumventing the issue of selecting a single region, with its associated inherent complications, SMURF provides higher profiling resolution. The larger the number of regions, the smaller the size of a characteristic group, typically allowing to identify a single 16s rRNA gene sequence that unambiguously appears in the sample.
[0199] SMURF was applied to analyze all eight plants across leaf, soil, and root samples. Every sample underwent analysis through all 63 region combinations of the amplified sixregions. For example, an analysis was performed using solely V1V2, followed by V1V2 combined with V3V4, followed by V1V2 combined with V3V4 and V7V9, and so forth.
[0200] The assessment of each region's contribution was based on three parameters: (a) ambiguity, i.e., the average group size detected by SMURF, the lower the ambiguity, the higher the resolution; (b) detection, representing the percentage of correctly identified 16s rRNA genes when the full six-region combination is used as a benchmark; (c) Pearson correlation between the frequencies of 16s rRNA gene detected by a set of regions vs. those detected by the full six-region combination.
[0201] The average performance across all three parameters improves when more regions are incorporated into the SMURF analysis. The average ambiguity across all 24 samples was 20,000 16s rRNA genes when using a single region, yet by adding only two more regions to the analysis, ambiguity was reduced to less than 2,000 (Fig. 2A). Secondly, the capacity to accurately detect a substantial proportion of the population increased linearly (Fig. 2B). Each additional amplicon contributed to the correct identifications, eliminating unique SSUs that were inaccurately declared by single regions, then subsequently dismissed with the inclusion of more comprehensive information. Thirdly, as more amplicons are added, the correlation improves significantly, with many samples achieving a correlation coefficient (r) higher than 0.9 with the addition of just two or three regions (Fig. 2C). Figures 2A-2C display a significant variance in performance for a given number of regions, depending on their identity. Assuming that costs scale with the number of amplified regions, the optimal selection of primers may be crucial. Such selection can be tailored to suit the specific aims and budget. One can identify an optimal selection by visualizing mutual taxonomies and Pearson correlation along the horizontal and vertical axes, respectively, and ambiguity designated by the marker size (Fig. 2D-2E). For example, a combination of three primer pairs, V1V2, V2V3, and V4V5, yields better results in soil samples than some sets of four primers (Fig. 2D, red cross vs. all other pink circles). However, the same combination performs poorly for root (Fig. 2D-2E, red cross) and leaf (Fig. 2F, red cross) samples.Example 3: Enriching the database using long-read sequencing
[0202] Although SMURF provides improved profiling, its results depend on the comprehensiveness of the available databases of 16s rRNA gene. However, most databases are highly enriched for, e.g., human-associated bacteria, while other habitats may be severely underrepresented.
[0203] To bridge this gap and harness the advantages of SMURF across habitats, the inventors propose a simple preceding step in a microbiome sampling project designated to provide niche-specific full-length 16s rRNA genes. The procedure entails a PCR step, utilizing a multiplex set of locked nucleic acid (LNA) primers aimed at effectively amplifying the whole 16s rRNA gene. LNA primers are shorter than regular primers (10-13 nt instead of ~20nt), thus allowing an increased universality compared to regular primers while maintaining primer specificity and high Tm (see Materials and Methods). There was therefore proposed pooling samples from the relevant niche, applying PCR amplification using the designed LNA primers, and sequencing to create full-length 16s rRNA genes that potentially enrich existing databases with novel sequences.
[0204] To illustrate this idea and estimate the number of novel 16s rRNA genes that can be detected, this approach was used over pools of samples, each from a different niche: (a) plants from the Arava desert in Israel; (b) plants from the temperate lowland in Israel; (c) sponges of the species Spongia officinalis collected from the Mediterranean coast of France; and (d) Spongia officinalis collected from the Mediterranean coast of Israel. These pools were sequenced using PacBio long reads amplicon sequencing, resulting in -200K high- quality 16s rRNA genes (Fig. 12) of average length 1400 bp.
[0205] Figure 3A presents the percentage of novel 16s rRNA genes as a function of the identity threshold to SILVA 16s rRNA genes. In general, sponges seem to have significantly higher percentages of novel 16s rRNA genes, irrespective of threshold. Using the SILVA definition of a novel 16s rRNA gene, i.e., a threshold of 99% identity (SILVA NR99 shown as a vertical dashed line in Fig. 3 A), more than 95% of detected 16s rRNA genes were novel for samples of Spongia officinalis and about 85% for samples taken from plants.
[0206] To remove redundancy in the initial pools, the SILVA database standard threshold for their NR99 database we used, and samples with 99% identity were clustered. This resulted in a table where each cluster is represented by a single 16s rRNA gene. Figure 3B shows the histograms of cluster sizes, i.e., the number of PacBio reads matching this 16s rRNA gene for one pool of samples derived from .S', officinalis. Upper and lower panels correspond to SILVA alignments lower or higher than 97%, respectively. Overall, the distribution of read counts is very similar between known 16s rRNA genes and novel 16s rRNA genes. The tail of the distribution shows highly abundant 16s rRNA genes (e.g., having more than 1000 reads) that are not represented in SILVA, suggesting that the underrepresentation problem is not limited to rare species. Furthermore, 96% of the novel 16s rRNA genes have fewer than 10 reads (3758 of 3880), indicating more could be detectedwith deeper sequencing. The same effect was observed for other pools, as is evident from rarefaction curves in Figure 3C.Example 4: The CoSMIC approach
[0207] When considering the problems and solutions presented in the former sections, the Inventors propose guidelines for large-scale microbiome studies that would provide high- resolution profiling in a new or known niche while accommodating for different budget tiers (Fig. 4). It is recommended to treat each research project as a new ecological niche (orange, right-hand side in the flowchart) and enrich current 16s rRNA sequence databases with relevant sequences, i.e., pool multiple samples of different origins of the ecological niche, apply LNA amplification and perform long reads sequencing. Next, using the newly enriched database, sequence a representative set of samples using the largest set of primers available (e.g., all 10 Swift pairs or all 6 Qiagen pairs), apply SMURF, and select the subset of primer pairs that provide sufficient resolution per niche (or select all primer pairs in case costs do not scale considerably with the number of primers). Once these steps are finalized, the project transitions to a “familiar niche” case (purple, left-hand side), which performs routine amplification of the optimal primer set followed by short reads sequencing, subsequently analyzed by augmented-database SMURF.Example 5: Experimental evaluation of CoSMIC
[0208] To evaluate the advantages of CoSMIC, Hymelaeci hirsuta, Solcinum lycopersicum, and Triticum aestivum were sampled at their leaves, roots, and surrounding soil. In addition, amock community of 12 known species was created. To construct the ground truth microbial communities, standard shotgun metagenomic sequencing of each of the ten samples was performed and 16s rRNA genes and their relative frequencies were extracted in each sample (see Materials and Methods). These 16s rRNA genes are referred to as ground truth 16s rRNA genes (gtl 6s), and their relative frequencies are called ground truth frequencies.
[0209] As a first step in CoSMIC, all ten samples were pooled and the ENA primers were used to amplify full-length 16s rRNA genes, the genes were sequenced using synthetic long reads technology (Eoop genomics) to achieve full-length 16s rRNA genes. Novel 16s rRNA genes (using the SIEVA NR99 standard) were added to the SIEVA database. This database is referred to as the “augmented SIEVA”.
[0210] Each sample was then sequenced by Illumina short reads using the Swift pipeline, which amplifies several regions along the 16s rRNA gene (see Materials and Methodssection and detailed information about obtained reads in supplementary for Loop, SWIFT and metagenomics S 13- 14). To apply SMURF, 10 amplicons spanning the 16s gene were considered, based on the abovementioned primers, and the augmented SILVA database was applied. For comparison, the standard analysis provided by Swift was provided, based on either all primers or on the V3V4 amplicon.
[0211] The comparison was based on four criteria: (a) Alignment identity - the average alignment score between detected 16s rRNA genes and gtl6s; (b) Ambiguity - the number of 16s rRNA genes in the augmented SILVA database that match detected 16s rRNA genes by each method; (c) Explained frequency - the sum of ground truth relative frequencies for gtl6s that were also detected by each method; (d) Erroneous 16s rRNA genes - the number of detected 16s rRNA genes that were not present in gtl6s (for details regarding a-d see Materials and Methods).
[0212] CoSMIC outperforms traditional methods on all parameters (Fig. 5).
[0213] Alignment identity: Mean and overall alignment identities are higher for CoSMIC across all samples (Fig. 5A). The distribution varies according to the tissue sampled, with mock, root, and soil consistently producing better scores compared to leaf tissue. This is possibly related to the low number of bacteria detected due to chloroplasts occupying most of the reads. While the difference between analysis methods is evident across all tissues, the most striking example is the mock community scores. Swift-based analysis, based on all regions or on V3V4, had a median alignment identity of about 90% even for this simple, well-defined mixture, compared to an almost 100% alignment identity for CoSMIC.
[0214] Ambiguity - Figure 5B presents a histogram of the ambiguity across all 16s rRNA genes detected in the 10 samples. In most cases each detected OTU by Swift corresponds to hundreds or thousands of the database’s 16s rRNA genes, while CoSMIC detects a single 16s rRNA gene, yielding a more precise, actionable result.
[0215] Explained frequency - Significant improvements are also apparent in the accuracy of community composition. Figure 5C presents the total explained frequency as a function of the identity threshold averaged per sample type across plants. Using CoSMIC, more species are accurately reported, accounting for a higher proportion of the ground-truth sample. The difference between the methods is particularly evident when only high-quality alignments are considered, resulting in almost zero frequency explained by traditional methods.
[0216] Erroneous 16s rRNA gene - The percentage of erroneous 16s rRNA genes, i.e., those that are reported yet do not exist in the ground truth sample, is significantly lower for CoSMIC compared to the other methods (Fig. 5D).
[0217] Figure 5E shows all criteria using a radar plot for each method, clearly showing the improved performance by CoSMIC, especially when addressing ambiguity (Amb) and explained frequency at high alignment identity scores (Cov-99).Example 6: Chloroplast reduction-based improvements to CoSMIC
[0218] During the above described experiments, it was observed that short read amplification of these regions from host tissue often resulted in reduced sequencing efficiency due to host contamination. For example, when sequencing the V4 region of the 16S rRNA gene from plant tissue, host-derived plastid and mitochondrial sequences were found to account for as much as 95% of all sequenced reads from some samples.
[0219] A search of the SILVA database was performed in order to find a common sequence in chloroplast 16S rRNA genes. A unique 13 base pair (bp) sequence (CGTCTGTAGGTGG; SEQ ID NO: 28) was found that is present in most chloroplast sequences but not commonly found in other bacterial 16S rRNA genes. Figure 13A illustrates the location of the unique sequence (block 1 V3-V4 in pink) on the wheat chloroplast sequence, which cannot be found in the E. coli 16S rRNA sequence. The sequence was depleted in 5271 out of 6214 amplified sequences in the V3-V4 region. To use this unique sequence to prevent PCR amplification of the chloroplast 16s rRNA during amplicon sequencing library preparation two methods were tested.
[0220] In the first, chloroplast contamination was decreased during PCR by the use of an oligo blocker. This oligo blocker was an LNA molecule, but a PNA blocker would also have worked as both have been shown to inhibit transcription. LNA oligos were synthesized because it enhances the pairing with complementary nucleotide strands and increases the stability of the resulting duplex. For the PCR reaction v3-v4 primers and the oligo were added to synthetic plasmids containing either Arabi dopsis chloroplast 16S DNA or synthetic plasmids with E. coli 16S DNA. Since the oligo annealing temperature is at 74°C, while the annealing temperature for the primers is 50°C, the DNA polymerase is prevented from amplifying the v3-v4 region in chloroplasts only. Chloroplast contamination dropped sharply with the addition of increasing concentrations of the blocking oligo (X1-X10, indicating the ratio of blocker to primer) while E. coli DNA was unaffected (Fig. 13B).
[0221] In the second, in vitro cleavage of chloroplast 16s rRNA was carried out by CRISPR. A crRNA (CRISPR guide RNA) was designed to target the unique 13 bp sequence and was ordered from IDT. The 13 bp sequence contains two possible PAM sequences (NGG) both near the 3’ end of the sequence (SEQ ID NO: 28 ends in AGGTGG). AGG was selected as the PAM (though TGG could have been selected) and the gRNA was designed to hybridize directly before the PAM (the PAM is on the opposite strand). Thus, the gRNA contained the sequence CGTCTGT (SEQ ID NO: 41) which corresponds to 7 bases at the 5’ end of the unique sequence. The full gRNA sequence targeting the chloroplast 16s rRNA sequence was TTGGGCGTAAAGCGTCTGT (SEQ ID NO: 42). This renders the CRISPR complex specific to chloroplast 16s rRNA. The cleavage was carried out in vitro before the library preparation. After cleavage, DNA amplicon library preparation was carried out as normal. The CRISPR effectively depleted the sample of intact chloroplast 16s rRNA (Fig. 13C).
[0222] The second PAM (TGG) is also tested. The additional gRNAs contain the sequence CGTCTGT AGG (SEQ ID NO: 43) which corresponds to the 10 bases at the 5’ end of the unique sequence. The full gRNA sequence targeting the chloroplast 16s rRNA sequence is GGCGTAAAGCGTCTGTAGG (SEQ ID NO: 44) or TTGGGCGTAAAGCGTCTGTAGG (SEQ ID NO: 45). This renders the CRISPR complex specific to chloroplast 16s rRNA. Both SEQ ID NO: 44 and SEQ ID NO: 45 are tested in the same manner that SEQ ID NO: 42 is tested and are found to also be effective in cleaving chloroplast 16s rRNA without significant cleavage of microbial 16s rRNA.
[0223] Having established the general effectiveness of the blocker and CRISPR approaches, they were both investigated in greater depth. To this end a commercial Community Standard (CS) composition of a mix of bacteria was used. This composition was tested using the standard CoSMIC method as well as the method with the blocker or CRISPR. Importantly, bacterial frequency was not significantly changed by either modification to the protocol (Fig. 14A).
[0224] Samples were prepared with different ratios of plasmid containing the 16s rRNA sequence of chloroplasts and the Community Standard. Two ratios were tested 50:50 chloroplast to CS and 90:10 chloroplast to CS. During the preparation of the library two different concentrations of LNA blocker were used, 10 mM and 100 mM, and a control without blocker was included. The three blocker concentrations (10, 100 and 0) were tested in both the first PCR (V3V4 amplification) and the second PCR (indices addition). CRISPR cutting reaction was followed by an inactivation step and was performed before bead cleanup.
[0225] When the 50:50 mix was tested without an LNA blocker or CRISPR an inaccurate ratio was produced which greatly overrepresented (-90%) the chloroplast contribution (Fig. 14B). Addition of blocker to only one PCR step produced a mild improvement, which was enhanced when blocker was added to both PCRs. In particular the addition of lOOmM blocker to both the first and second PCR produced a 59:41 ratio which was much closer to the actual 50:50 ratio (Fig. 14B). The CRISPR method produced a very similar 59:41 ratio (Fig. 14B, right of the dashed line). When the bacterial community was investigated, it was notable that without the blocker or when lOnM blocker was included in only 1 PCR the Pseudomonas population that was only about 4% of the total could not be detected due to the high levels of chloroplast (Fig. 14C). Higher amounts of blocker, especially when included in both PCRs, and the CRISPR step resulted in detection of this small population, exemplifying the improvement produced by both methods.
[0226] When the 90: 10 mixture, with very high chloroplast contamination, was tested the results were even more striking. With the standard protocol only chloroplast sequence was detected (Fig. 14D) making bacterial analysis impossible (Fig. 14E). The blocker and CRISPR were able to produce the expected 90: 10 ratio (Fig. 14D), but with this level of chloroplast contamination the bacterial analysis was suboptimal with the smaller populations (Pseudomonas and Salmonella) not always detected (Fig. 14E).
[0227] So far, the experiments involved a synthetic standard community composed of 10 bacterial species and purified chloroplast genes in plasmids. Next, performance in environmental samples was tested. Leaves were collected from 15 different plant species and the relative abundance of bacteria on the leaves was determined using the CoSMIC method with and without CRISPR cutting of chloroplast 16s rRNA. As expected, CRISPR treatment greatly reduced the amount of chloroplast rRNA detected and allowed for the detection of less abundant bacteria (Fig. 14F). In all 15 leaf samples the most abundant bacteria was Rickettsiales. When standard Cosmic was performed this abundant bacterium was the only one detected on the leaves in every sample, but when CRISPR was applied at least one of three other bacteria (Pseudomonadaies, Bacillales and Enterobacterales) were detected in nearly every sample. The LNA blocker produced similar, though inferior results (Fig. 14G).
[0228] Having evaluated both methodologies across a phylogenetically diverse sample of plant species the universal nature of the target sequence (SEQ ID NO: 28) was firmly established. Both CRISPR and LNA treatments significantly reduced chloroplast abundance. Reduced chloroplast abundance was associated with increased detection of previouslyundetected low-abundance bacteria. Both methods thus result in reduced chloroplast reads and greatly improving the CoSMIC protocol.Example 7: UMI based improvement to CoSMIC
[0229] Another cause of noise and bias in the PCR amplification stems from copy number variation of the 16s rRNA gene and intra-individual SSU variation. To overcome this problem and allow accurate estimates of relative frequencies, unique molecular identifiers (UMIs) are used as part of the CoSMIC sample preparations, specifically during analysis using the SMURF pipeline. A random barcode of 5-10 nucleotides (“N”) was added to the short primers used for short read amplification. Examples of such primers for V3V4 amplification are provided in Table 4. After the sequencing, duplicates of the barcodes are removed thereby reducing PCR bias and allowing for more accurate determination of relative frequencies of the various bacteria present.
[0230] Table 4: V3-V4 primers with UMIs and Illumina sequencing adapters
[0231] Additionally, these new UMI-incorporated primers make use of a degenerate sequence in the 3’ primer that hybridizes with the 16s rRNA. This allows for broader binding to the various sequences present in the sample. Degenerate primers for the short amplifications are provided in Table 5. It will be understood that as provided in Table 4 for the V3-V4 amplification, the primers of Table 5 can also include UMIs and sequencing adapters.
[0232] Table 5. Degenerate primers for use in the SWIFT AMPLICON 16s+ITS PANEL.Primers applied to amplify the variable regions along the 16s gene.
[0233] A comparative analysis was conducted on 39 wheat root samples to evaluate microbial diversity using either the standard SMURF pipeline (No UMI) or a pipeline that incorporates UMIs to enhance accuracy in distinguishing true biological sequences from PCR duplicates and sequencing errors. To quantify differences between the pipelines, microbial biodiversity (alpha diversity), richness, and Bray-Curtis dissimilarity (beta diversity) were calculated at multiple taxonomic levels, e.g., Order, Family, Genus, andSpecies. The two pipelines indeed showed quite a bit of dissimilarity, with some bacteria increasing in abundance and others decreasing. Five representative wheat root samples are shown in Figure 15A. These differences were present regardless of whether bacteria were examined across Order, Family, Genus or Species, though differences were greater when looking at the Genus or Species level (Fig. 15B). Importantly, when overall microbial diversity (richness) was examined across the samples, it was found that inclusion of the UMI increased richness as compared to when no UMI was used (Fig. 15C). When the richness difference was examined, where a positive A Richness indicates more species were identified by the UMI pipeline, and a negative A Richness indicates more species were detected without UMIs, only a small number of samples showed greater richness without UMIs, indicating the superiority of their inclusion in the pipeline (Fig. 15D).
[0234] Although the invention has been described in conjunction with specific embodiments thereof, it is evident that many alternatives, modifications and variations will be apparent to those skilled in the art. Accordingly, it is intended to embrace all such alternatives, modifications and variations that fall within the spirit and broad scope of the appended claims.
Claims
CLAIMS:
1. A method of profiling a microbial community within a plant niche, the method comprising: a. receiving a pool of DNA comprising DNA extracted from a plurality of samples from said plant niche; b. amplifying 16s rRNA genes within said pool of DNA to produce long amplicons wherein each long amplicon comprises at least 80% of the sequence of a 16s rRNA gene; c. sequencing said long amplicons with a long-read sequencing technology to receive full-length 16s rRNA sequences comprising a contiguous at least 80% of a 16s rRNA gene: d. generating a 16s rRNA database comprising said received full-length 16s rRNA sequences; e. amplifying and sequencing 16s rRNA sequences from a sample from said plant niche using short deep-sequencing technology with a plurality of primer pairs wherein each primer pair produces a short amplicon of not greater than 500 nucleotides and amplicon reads of not longer than 300 nucleotides; and f. computationally combine said amplicon reads to produce full-length 16s rRNA sequences based on said generated 16s rRNA database, thereby identifying unique 16s rRNA genes within said sample from said plant niche; and wherein said amplifying or step (b), step (e) or both is performed: i. in the presence of a blocking oligo comprising a sequence that perfectly hybridizes with a 16s rRNA sequence of a chloroplast and does not perfectly hybridize with a 16s rRNA sequence of a microbe of said community; or ii. after cutting a 16s rRNA sequence of a chloroplast by contact with a CRISPR-guide RNA (gRNA) complex wherein said CRISPR-gRNA complex recognizes and cuts a 16s rRNA sequence of a chloroplast and not a 16s rRNA sequence of a microbe of said community;thereby profiling a microbial community within a plant niche.
2. The method of claim 1 , wherein said plurality of samples comprises at least 10 samples from said plant niche.
3. The method of claim 1 or 2, wherein said sample from (e) is a sample from said plurality of samples.
4. The method of any one of claims 1 to 3, comprising receiving a plurality of samples from said plant niche, extracting DNA from each sample of said plurality of samples, and pooling said extracted DNA to produce said pool of DNA.
5. The method of any one of claims 1 to 4, wherein each long amplicon comprises at least 85% ofthe sequence of a 16s rRNA gene and wherein full-length 16s rRNA sequences comprise a contiguous at least 85% of a 16s rRNA gene.
6. The method of any one of claims 1 to 5, wherein said long amplicons each comprise at least 1000 nucleotides.
7. The method of claim 6, wherein said long amplicons each comprise at least 1300 nucleotides.
8. The method of any one of claims 1 to 7, wherein said amplifying of step (b) is with at least two different sets of primers wherein each set of primers produces a long amplicon.
9. The method of any one of claims 1 to 8, wherein said amplifying of step (b) is with locked nucleic acid (LNA) primer pairs.
10. The method of claim 9, wherein said LNA primers are 10-13 nucleotides long.
11. The method of claim 9 or 10, wherein said LNA primer pairs are selected from the group consisting of: SEQ ID NO: 10-11, SEQ ID NO: 12-13 and SEQ ID NO: 14-15.
12. The method of claim 11, wherein said amplifying of step (b) is with a cocktail of at least three different LNA primer pairs and wherein said three pairs are SEQ ID NO: 10-11, SEQ ID NO: 12-13 and SEQ ID NO: 14-15.
13. The method of any one of claims 1 to 12, wherein said long read sequencing technology is selected from the group consisting of: Single-molecule real-time sequencing, loop sequencing and long read nanopore sequencing.
14. The method of any one of claims 1 to 13, wherein said database comprises known full- length 16s rRNA sequences combined with said received full-length 16s rRNA sequences.
15. The method of any one of claims 1 to 14, wherein sequencing of step (e) comprises next generation sequencing, deep sequencing or massively parallel sequencing.
16. The method of any one of claims 1 to 15, wherein said plurality of primer pairs of step (e) amplifies at least 80% of a 16s rRNA gene.
17. The method of any one of claims 1 to 16, wherein said plurality of primer pairs of step (e) comprises SEQ ID NO: 1-9, SEQ ID NO: 29-40 or both.
18. The method of any one of claims 1 to 17, wherein all primers of said plurality of primer pairs of step (e) comprise: a. a sequencing adapter; b. a unique molecular identifier (UMI) or c. both.
19. The method of any one of claims 1 to 18, wherein said amplifying of step (b) is performed in the presence of a blocking oligo comprising a sequence that perfectly hybridizes with a 16s rRNA sequence of a chloroplast and does not perfectly hybridize with a 16s rRNA sequence of a microbe of said community.
20. The method of any one of claims 1 to 19, wherein said amplifying of step (e) is performed in the presence of a blocking oligo comprising a sequence that perfectly hybridizes with a 16s rRNA sequence of a chloroplast and does not perfectly hybridize with a 16s rRNA sequence of a microbe of said community.
21. The method of any one of claims 1 to 20, wherein said blocking oligo is an LNA or PNA blocking oligo.
22. The method of any one of claims 1 to 21, wherein said sequence that perfectly hybridizes with a 16s rRNA sequence of a chloroplast and does not perfectly hybridize with a 16s rRNA sequence of a microbe of said community comprises or consists of the sequence CGTCTGTAGGTGG (SEQ ID NO: 28) or is 100% complementary to SEQ ID NO: 28.
23. The method of any one of claims 1 to 22, wherein said method comprises before said amplifying of step (e) cutting a 16s rRNA sequence of a chloroplast by contact with a CRISPR-guide RNA (gRNA) complex wherein said CRISPR-gRNA complex recognizes and cuts a 16s rRNA sequence of a chloroplast and not a 16s rRNA sequence of a microbe of said community.
24. The method of any one of claims 1 to 23, wherein said CRISPR-gRNA complex recognizes and cuts CGTCTGT (SEQ ID NO: 41) or SEQ ID NO: 28.
25. The method of claim 24, wherein said gRNA comprises the sequence TTGGGCGTAAAGCGTCTGT (SEQ ID NO: 42).
26. The method of any one of claims 1 to 25, further comprising selecting a subset of primer pairs for step (e) that provide sufficient resolution for the 16s rRNA genes within said plurality of sample from said plant niche.
27. The method of any one of claims 1 to 26, wherein said computationally combining comprises performing Short MUlti-Region Framework (SMURF) analysis.
28. The method of any one of claims 1 to 27, wherein step (f) comprises grouping produced 16s rRNA sequences with substantially the same sequence to produce a list of unique 16s rRNA genes within said sample.
29. A locked nucleic acid (LNA) primer pair comprising: a. an LNA forward primer comprising SEQ ID NO: 10 and an LNA reverse primer comprising SEQ ID NO: 11; b. an LNA forward primer comprising SEQ ID NO: 12 and an LNA reverse primer comprising SEQ ID NO: 13; or c. an LNA forward primer comprising SEQ ID NO: 14 and an LNA reverse primer comprising SEQ ID NO: 15.
30. The LNA primer pair of claim 29, wherein each LNA primer is not longer than 13 nucleotides.
31. The LNA primer pair of claim 29 or 30, comprising: a. an LNA forward primer consisting of SEQ ID NO: 10 and an LNA reverse primer consisting of SEQ ID NO: 11; b. an LNA forward primer consisting of SEQ ID NO: 12 and an LNA reverse primer consisting of SEQ ID NO: 13; or c. an LNA forward primer consisting of SEQ ID NO: 14 and an LNA reverse primer consisting of SEQ ID NO: 15.
32. A kit comprising a first LNA primer pair of any one of claims 29 to 31, a second LNA primer pair of any one of claims 29 to 31 and a third LNA primer pair of any one of claims 29 to 31, wherein said first, second and third LNA primer pairs each comprise different primers.
33. A kit comprising an LNA primer pair of any one of claims 29 to 31 and a. a blocking oligo comprising or consisting of SEQ ID NO: 28; b. a gRNA that binds to and cuts SEQ ID NO: 41 or SEQ ID NO: 28; or c. both.
34. The kit of claim 33, wherein said blocking oligo is an LNA blocking oligo or a PNA blocking oligo.
35. The kit of claim 33 or 34, wherein said gRNA comprises SEQ ID NO: 42.
36. The kit of any one of claims 33 to 35, comprising said gRNA and further comprising a CAS9 enzyme or a nucleic acid molecule encoding said CAS9 enzyme.
37. A method of profiling a microbial community within a biological niche, the method comprising: a. receiving a pool of DNA comprising DNA extracted from a plurality of samples from said biological niche;b. amplifying 16s rRNA genes within said pool of DNA to produce long amplicons wherein each long amplicon comprises at least 80% of the sequence of a 16s rRNA gene; c. sequencing said long amplicons with a long-read sequencing technology to receive full-length 16s rRNA sequences comprising a contiguous at least 80% of a 16s rRNA gene: d. generating a 16s rRNA database comprising said received full-length 16s rRNA sequences; e. amplifying and sequencing 16s rRNA sequences from a sample from said plant niche using short deep-sequencing technology with a plurality of primer pairs wherein each primer pair produces a short amplicon of not greater than 500 nucleotides and amplicon reads of not longer than 300 nucleotides; and f. computationally combine said amplicon reads to produce full-length 16s rRNA sequences based on said generated 16s rRNA database, thereby identifying unique 16s rRNA genes within said sample from said biological niche; thereby profding a microbial community within a biological niche.
Citation Information
Patent Citations
Antibacterial spittoon with incinerable container
FR6507E
Test for Huntington's disease
US4666828A
Process for amplifying nucleic acid sequences
US4683202A
Apo AI / CIII genomic polymorphisms predictive of atherosclerosis
US4801531A
Intron sequence analysis method for detection of adjacent and remote locus alleles as haplotypes
US5192659A