Methods, compositions, and kits for target enrichement and metagenomic next generation sequencing

teNGS addresses the challenge of high human sequence interference in mNGS by combining enriched and unenriched libraries for enhanced microbial detection, achieving high sensitivity and efficiency in clinical samples.

WO2026050333A1PCT designated stage Publication Date: 2026-03-05ABBOTT LAB INC
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/043654
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-28
Filing Date
2025-08-27
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Metagenomic next generation sequencing (mNGS) in clinical samples is hindered by high levels of human sequences, which obscure the detection of low abundance nucleic acids from viruses, necessitating improved methods for microbial detection.

Method used

A method combining pools of enriched and unenriched metagenomic libraries in a predetermined ratio for a single sequencing run, referred to as target enrichment and metagenomic next generation sequencing (teNGS), achieving 100-10,000X increases in depth and greater than 50% genomic coverage with reduced read requirements.

Benefits of technology

teNGS significantly enhances microbial detection sensitivity and efficiency, allowing simultaneous detection of all microbe types with 3-4% of the usual read requirement, saving time and resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025043654_05032026_PF_FP_ABST
    Figure US2025043654_05032026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to methods for detecting microbes in samples, for example, for characterizing the metagenomes of biological samples. Compositions and kits for implementing the use in the methods are also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS, COMPOSITIONS, AND KITS FOR TARGET ENRICHEMENT AND METAGENOMIC NEXT GENERATION SEQUENCINGRELATED APPLICATION INFORMATION

[0001] This application claims priority to U.S. Application No. 63 / 687,812, filed on August 28, 2024, the contents of which are herein incorporated by reference.SEQUENCE LISTING STATEMENT

[0002] The contents of the electronic sequence listing titled ABBTL_43653_601_SequenceListing.xml (Size: 4,096 bytes; and Date of Creation: August 27, 2025) are herein incorporated by reference in their entirety.TECHNICAL FIELD

[0003] The present disclosure relates to methods for detecting microbes in samples, for example, for characterizing the metagenomes of biological samples. Compositions and kits for implementing the use in the methods are also disclosed.BACKGROUND

[0004] While metagenomic next generation sequencing (mNGS) enables detection of known and novel microbes in clinical samples, unbiased priming leads to high levels of human sequences which hinder detection of low abundance nucleic acids, such as those from viruses. Probe-based target enrichment represents a highly sensitive and cost-effective approach for overcoming host background and detecting microbes, such as viruses, however demonstrated performance in clinical samples has been limited. These challenges underscore the need for improved methods of detecting microbes in clinical samples.SUMMARY

[0005] Aspects of the present disclosure relate to methods, compositions and kits for detecting microbes in samples, for example, to determine whether a subject has a microbial infection (e.g., a viral infection). The methods, compositions and kits can also be used for characterizing metagenomes of samples. In various aspects, the approach described herein combines pools of enriched metagenomic libraries that are enriched for microbes of interest with pools of unenriched metagenomic libraries in a predetermined ratio of enriched metagenomic libraries to unenriched metagenomic libraries for next generation sequencing on a single sequencing run. Surprisingly and unexpectedly, the approach disclosed herein, referred to as target enrichment and metagenomic next generation sequencing (“teamNGS”), can achieve 100-10, 000X increases in depth and greater than50% genomic coverage for samples with titers greater than or equal to 1000 cp / ml, while only requiring 3-4% the number of reads, and can be performed on a single sequencing run to detect all microbe types simultaneously, significantly saving time and resources.

[0006] Accordingly, in an aspect, the disclosure provides a method of characterizing metagenomes of at least one sample obtained from at least one subject suspected of having a microbial infection, the method comprising (a) preparing at least four total unenriched libraries for metagenomic next generation sequencing from total nucleic acids extracted from at least one sample obtained from at least one subject suspected of having a microbial infection, wherein each unenriched library comprises: (i) a first control unenriched nucleic acid library comprising nucleic acids of known microbial origin; (ii) a second control unenriched nucleic acid library comprising nucleic acids of human origin obtained from at least one human subject who is not suspected of having the microbial infection; (iii) a third control unenriched library that is substantially free of nucleic acids; or (iv) at least one test unenriched nucleic acid library comprising nucleic acids of human origin and nucleic acids of unknown microbial origin prepared from the at least one sample; wherein each unenriched nucleic acid library comprises fragmented double-stranded DNA that is barcoded for indexing; (b) partitioning the unenriched libraries, on an equimolar basis, into a first set of pools of unenriched libraries and a second set of pools of unenriched libraries; (c) enriching each of the unenriched libraries in the first set of pools for nucleic acids of microbial origin to produce a first set of pools of enriched libraries that are enriched for nucleic acids of microbial origin; (d) combining a volume from each pool in the first set of pools of enriched libraries, in a predetermined volume ratio, with a volume from each corresponding pool in the second set of pools of unenriched libraries into a single flow cell for next generation sequencing; (e) sequencing the libraries in the combined volumes of each corresponding pool in the single flow cell using a next generation sequencing platform; and (f) characterizing the metagenomes of the at least one sample obtained from the at least one subject suspected of having the microbial infection based on the sequencing results.

[0007] The methods of the disclosure contemplate preparing any number of libraries in step (a) that is suitable for automation. In some embodiments, step (a) comprises preparing at least five, at least six, at least eight, at least 12, at least 24, at least 48, at least 96, at least 384, or at least 1536 unenriched libraries.

[0008] Any number of samples can be used to prepare the libraries in step (a). In some embodiments, the at least one sample from which the libraries in step (a) are prepared comprises at least 2 samples, at least 3 samples, at least 5 samples, at least 6 samples, at least 9 samples, at least 12 samples, at least 18 samples, at least 21 samples, at least 24 samples, at least 30 samples, at least 36, at least 42 samples, at least 45 samples, at least 72 samples, at least 78 samples, at least 84 samples, at least 90 samples, at least 312 samples, at least 336 samples, at least 348 samples, at least 354 samples, at least 366 samples, at least 1152 samples, at least 1248 samples, at least 1344 samples, at least 1392 samples, or at least 1440 samples obtained from subjects suspected of having the microbial infection.

[0009] The samples can be partitioned in step (b) into sets comprising any number of pools of unenriched libraries. In some embodiments, each of the first set of pools and the second set of pools comprises at least one pool, at least two pools, at least three pools, at least four pools, at least six pools, at least eight pools, at least 12 pools, at least 16 pools, at least 24 pools, at least 32 pools, at least 48 pools, at least 64 pools, at least 96 pools, or at least 128 pools of unenriched libraries. In an embodiment, each set of the at least one pool comprises at least 4, at least 5, at least 6, at least 12, at least 24, or at least 48 unenriched libraries. In another embodiment, each set of the at least two pools comprises at least 12, at least 24, at least 48, or at least 96 unenriched libraries. In yet another embodiment, each set of the at least three pools comprises at least 12, at least 24, at least 48, or at least 96 unenriched libraries. In a further embodiment, each set of the at least three pools comprises at least 12 nucleic acid libraries. In still a further embodiment, each set of the at least four pools comprises at least 24, at least 48, or at least 96 unenriched libraries. In certain embodiments, each set of the at least six pools comprises at least 24, at least 48, at least 96, or at least 384 unenriched libraries. In some embodiments, each set of the at least eight pools comprises at least 24, at least 48, at least 96, or at least 384 unenriched libraries. In other embodiments, each set of the at least 12 pools comprises at least 48, at least 96, or at least 384 unenriched libraries. In yet other embodiments, each set of the at least 16 pools comprises at least 384 unenriched libraries. In still yet other embodiments, each set of the at least 24 pools comprises at least 384 unenriched libraries. In one particular embodiment, each set of the at least 32 pools comprises at least 1536 unenriched libraries. In another particular embodiment, each set of the at least 48 pools comprises at least 1536 unenriched libraries. In a further particular embodiment, each set of the at least 64 pools comprises at least 1536 unenriched libraries. In still a further particular embodiment, each set of the at least 96 pools comprises at least 1536 unenriched libraries. In still yet another embodiment, each set of the at least 128 pools comprises at least 1536 unenriched libraries.

[0010] The number of controls and test libraries in any particular set of pools can vary. In some embodiments, each set of pools comprises: (i) a first control unenriched nucleic acid library comprising nucleic acids of known microbial origin; (ii) a second control unenriched nucleic acid library comprising nucleic acids of human origin obtained from at least one human subject who is not suspected of having the microbial infection; (iii) a third control unenriched library that is substantially free of nucleic acids; and (iv) at least X test unenriched nucleic acid libraries prepared from at least a Y samples obtained from at least Y subjects suspected of having the microbial infection; wherein X is equal to a variable number Z that is three less than the total number of unenriched libraries prepared divided by the number of pools in each set, wherein Y is equal to the number of pools in each set times Z, wherein each test unenriched nucleic acid library comprises nucleic acids of human origin and nucleic acids of unknown microbial, and wherein each unenriched nucleic acid library comprises fragmented double-stranded DNA that is barcoded for indexing.

[0011] The methods can also include, prior to step (a), (i) extracting total nucleic acids from at least one sample of at least one microbe to prepare the first control unenriched nucleic acid library; (ii) extracting total nucleic acids from at least one sample obtained from at least one human subject who is not suspected of having the microbial infection to prepare the second control unenriched nucleic acid library; and / or (iii) extracting total nucleic acids from the at least one sample obtained from the at least one subject suspected of having the microbial infection to prepare the at least one test unenriched nucleic acid library.

[0012] In some embodiments, where step (a) comprises: preparing at least eight total unenriched libraries for metagenomic next generation sequencing from total nucleic acids extracted from at least five samples obtained from at least five subjects suspected of having a microbial infection, wherein the at least eight total unenriched libraries comprise: (i) a first control unenriched nucleic acid library comprising nucleic acids of known microbial origin; (ii) a second control unenriched nucleic acid library comprising nucleic acids of human origin obtained from at least one human subject who is not suspected of having the microbial infection; (iii) a third control unenriched library that is substantially free of nucleic acids; and(iv) at least five test unenriched nucleic acid libraries each comprising nucleic acids of human origin and nucleic acids of unknown microbial origin prepared from the at least five samples, wherein each unenriched nucleic acid library comprises fragmented double- stranded DNA that is barcoded for indexing. In those embodiments, step (b) can comprise: partitioning the at least eight total unenriched libraries, on an equimolar basis, into a first set of one pool of unenriched libraries and a second set of one pool of unenriched libraries, wherein the one pool of unenriched libraries in each of the first and second sets comprises an aliquot taken from:(i) the first control unenriched nucleic acid library; (ii) the second control unenriched nucleic acid library; (iii) the third control unenriched library; and (iv) the at least five test unenriched nucleic acid libraries.

[0013] In other embodiments, step (a) comprises: preparing at least 48 total unenriched libraries for metagenomic next generation sequencing from total nucleic acids extracted from at least 42 samples obtained from at least 42 subjects suspected of having a microbial infection, wherein the at least 48 total unenriched libraries comprise: (i) at least two first control unenriched nucleic acid libraries comprising nucleic acids of known microbial origin; (ii) at least two second control unenriched nucleic acid libraries comprising nucleic acids of human origin obtained from at least one human subject who is not suspected of having the microbial infection; (iii) at least two third control unenriched libraries that are substantially free of nucleic acids; and (iv) at least 42 test unenriched nucleic acid libraries each comprising nucleic acids of human origin and nucleic acids of unknown microbial origin prepared from the at least 42 samples, wherein each unenriched nucleic acid library comprises fragmented double-stranded DNA that is barcoded for indexing. In those embodiments, step (b) can comprise partitioning the 48 total unenriched libraries, on an equimolar basis, into a first set of two pools of unenriched libraries and a second set of two pools of unenriched libraries, wherein each pool in the first and second sets of pools of unenriched libraries comprises an aliquot taken from:(i) a different one of the at least two first control nucleic acid libraries; (ii) a different one of the at least two second control libraries; (iii) a different one of the at least two third control libraries; and (iv) at least 21 different test libraries from the at least 42 test libraries.

[0014] In an exemplary embodiment, step (a) comprises: preparing at least 96 total unenriched libraries for metagenomic next generation sequencing from total nucleic acids extracted from at least 84 samples obtained from at least 84 subjects suspected of having a microbial infection, wherein the at least 96 total unenriched libraries comprise: (i) at least four first control unenriched nucleic acid libraries prepared from total nucleic acids extracted from nucleic acids of known microbial origin; (ii) at least four second control unenriched nucleic acid libraries comprising nucleic acids of human origin obtained from at least one human subject who is not suspected of having the microbial infection; (iii) at least four third control unenriched libraries that are substantially free of nucleic acids; and (iv) at least 84 test unenriched nucleic acid libraries each comprising nucleic acids of human origin and nucleic acids of unknown microbial origin prepared from the at least 84 samples, wherein each unenriched nucleic acid library comprises fragmented double-stranded DNA that is barcoded for indexing. In such embodiment, step (b) can comprise partitioning the 96 total unenriched libraries, on an equimolar basis, into a first set of four pools of unenriched libraries and a second set of four pools of unenriched libraries, wherein each pool in the first and second sets of four pools of unenriched libraries comprises an aliquot taken from: (i) a different one of the at least four first control nucleic acid libraries; (ii) a different one of the at least four second control libraries; (iii) a different one of the at least four third control libraries; and (iv) at least 21 different test libraries from the at least 84 test libraries.

[0015] In some embodiments, enriching each of the unenriched libraries comprises enriching nucleic acids of unknown microbial origin from at least one microbe of interest in the libraries in each pool of the first set of pools.

[0016] In some embodiments, the methods further comprise diluting each pool of enriched nucleic acid libraries. Each pool can be diluted to between about .20 nM to about 20 nm, about 0.30 nM to about 19 nM, about 0.4 nM to about 18 nM, about 0.5 nm to about 17 nM, about 0.6 nm to about 16 nm, about 0.7 to about 15 nM, about 0.8 to about 14 nM, about 0.9 to about 13 nM, about 1 nM to about 12 nM, about 1.1 nM to about 11 nM, about 1.2 nM to about 10 nM, about 1.3 nM to about 9 nM, about 1.4 nM to about 8 nM, about 1.5 nM to about 7 nM, about 1.6 nM to about 6 nM, about 1.7 nM to about 5 nM, about 1.8 nM to about 4 nM, about 1.9 nM to about 3 nM, or about 2 nM. In an exemplary embodiment, each pool of enriched nucleic acid libraries is diluted to about 2 nM.

[0017] In some embodiments, each aliquot taken from each unenriched library in the second set of pools is diluted. Each aliquot taken from each unenriched library in the second set of pools can be diluted to between about .20 nM to about 20 nm, about 0.30 nM to about 19 nM, about 0.4 nM to about 18 nM, about 0.5 nm to about 17 nM, about 0.6 nm to about 16 nm, about 0.7 to about 15 nM, about 0.8 to about 14 nM, about 0.9 to about 13 nM, about 1 nM to about 12 nM, about 1.1 nM toabout 11 nM, about 1.2 nM to about 10 nM, about 1.3 nM to about 9 nM, about 1.4 nM to about 8 nM, about 1.5 nM to about 7 nM, about 1.6 nM to about 6 nM, about 1.7 nM to about 5 nM, about 1.8 nM to about 4 nM, about 1.9 nM to about 3 nM, or about 2 nM. In an exemplary embodiment, each aliquot taken from each unenriched library in the second set of pools can be diluted to about 2 nM.

[0018] The predetermined volume ratio can vary. Exemplary predetermined volume ratios in microliters suitable for use herein include, without limitation, 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 2:8, 3:7, 4:6, 9:1, 8:1, 7:1, 6:1, 5:1, 4:1, 3:1, 2:1, 8:2, 7:3, or 6:4. In an exemplary embodiment, the volume ratio in microliters is 1:9. In another exemplary embodiment, the volume ratio in microliters is 2:8. In some embodiments, between about 0.10 microliters and about 10 microliters from each pool in the first set of pools is combined with between about 0.10 microliters and about 10 microliters from each corresponding pool in the second set of pools. In other embodiments, about 1 microliter or about 2 microliters from each pool in the first set of pools is combined with about 8 or about 9 microliters from each corresponding pool in the second set of pools. In yet another exemplary embodiment, about 1 microliter from each pool in the first set of pools is combined with about 9 microliters from each corresponding pool in the second set of pools. In still another exemplary embodiment, about 2 microliters from each pool in the first set of pools is combined with about 8 microliters from each corresponding pool in the second set of pools.

[0019] The at least one sample obtained from the at least one subject can be obtained from subjects suspected of having any type of infection. In some embodiments, the at least one subject is suspected of having a viral infection. In some embodiments, the at least one subject has or is suspected of having acute febrile illness (AFI). In some embodiments, the at least one subject has or is suspected of having a severe acute respiratory illness. In some embodiments, the nucleic acids of known microbial origin are nucleic acids of known viral origin. In some embodiments, the nucleic acids of human origin are obtained from at least one human subject who does not have or is not believed to have the viral infection. In some embodiments, the nucleic acids of unknown microbial origin comprise nucleic acids of viral origin. In some embodiments, the viral infection or viral origin is a virus selected from the group consisting of a virus from an order listed in Table 1, a virus family listed in Table 2, a virus from genus listed in Table 3, a virus from a species fisted in Table 4, and a virus strain listed in Table 5.

[0020] The methods of the disclosure can be performed using a variety of samples. In some embodiments, the at least one sample is a biological sample. Each aliquot taken from each unenriched library in the second set of pools can be diluted to between about the at least one sample is selected from the group consisting of bronchoalveolar lavage, cerebral spinal fluid, plasma, serum, sputum, urine, whole blood, stool, and a swab.

[0021] In some embodiments, the method further comprises detecting a viral infection in the at least one subject based on the next generation sequencing results. In some embodiments, the method further comprises treating the at least one subject for the viral infection or a symptom, condition,disease or disorder caused by the viral infection. In some embodiments, the method further comprises monitoring the treatment of the subject having the viral infection.BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The patent or application file contains drawings executed in color. Copies of this patent or patent application publication with color drawings will be provided by the Office upon request and payment of the necessary fee.

[0023] Having thus described the presently disclosed subject matter in general terms, reference will now be made to the accompanying Figures, which are not necessarily drawn to scale, and wherein:

[0024] FIG. 1 A illustrates batches of ninety-six clinical specimens pre-treated with an exemplary nuclease (e.g., benzonase) and spiked with an internal control before extraction of total nucleic acid on an exemplary automation system (e.g., ABBOTT m2000sp or KingFisher instrument). An automated liquid handler (e.g., epMotion) in a pre-amplification room) was used to perform cDNA and second-strand synthesis steps. Following bead clean-up, an exemplary library preparation kit (e.g., Nextera XT) 'tagments’ double stranded cDNA with barcoded i5 / i7 adapters. Amplified libraries were purified by bead clean up (in a post-amplification room), quantified by an exemplary fluorometer (e.g., Qubit) and an exemplary biochemical analyzer (e.g., BioAnalyzer), diluted and pooled for mNGS sequencing on a sequencing platform (e.g., NextSeq). FIG. IB illustrates the process of virus target enrichment, equal volumes of a library preparation kit (e.g., Nextera XT) are pooled and dried down in four sets of twenty-four (blue=virus positive; brown=NC; teal=PC). DNA pellets were reconstituted with CVRP probes and hybridized overnight at 65 °C, followed by the addition of magnetic streptavidin beads for thirty minutes. FIG. 1C illustrates after a series of stringent washes, captured viral reads are amplified via PCR for sixteen cycles using primers annealing to Illumina adapters. Library peaks are quantified, pools are diluted, and then all are combined onto a sequencing platform (e.g., MiSeq or NextSeq) for Pl runs.

[0025] FIG. 2 A shows a bar graph indicating percent genome coverage of purified stocks of HIV (blue; first bar on the left in each group of 3), SARS-CoV-2 (orange; second bar in the middle in each group of 3), and Zika virus (light grey; last bar on the right in each group of 3) when resuspended in different variations of healthy donor plasma. FIG. 2B shows bar graphs indicating dilutions of each virus stock of EMCV at 10 TCID50 / ml (dark grey) and 1 TCID50 / ml (light grey) (left), HIV (center left), SARS-CoV-2 (center right) and Zika (right) at 5,000 cp / ml (dark grey) and 1,000 cp / ml (light grey) were captured in pools of 8, 16, 24, and 32 libraries to determine percent genome coverage. FIG. 2C shows bar graphs indicating dilutions of each virus stock of EMCV at 10 TCID50 / ml (dark grey) and 1 TCID50 / ml (light grey) (left), HIV (center left), SARS-CoV-2 (center right) and Zika (right) at 5,000 cp / ml (dark grey) and 1,000 cp / ml (light grey) were captured in pools 8, 16, 24, and32 libraries to determine on-target read percentages. FIG. 2D shows a bar graph displaying EMCV reads, indicating total reads (black), EMCV reads in positive (green) and negative (light grey) libraries. FIG. 2E shows a bar graph displaying percentage of genome coverage, indicating EMCV reads in positive (green) and negative (grey) libraries.

[0026] FIG 3A shows NGS Run 1 (top), with two dilutions of EMCV, HIV, Zika, and SARS- CoV-2 contrived samples were either captured alone (8-plex) or in the presence of 12 virus-negative, 6 low titer virus-positives, and 6 high titer virus positives (32-plex). The model virus libraries were tagged with different barcodes, notated in red or orange, respectively. Mapping results are shown for the added low and high titers viruses included within the 32-plex capture. (Virus-negative; grey, Low titer virus-positives; light green, high titter virus-positives; dark green). FIG 3B shows NGS Run 2 (bottom), with two dilutions of EMCV, HIV, Zika, and SARS-CoV-2 contrived samples were either captured in the presence of 4 virus-negative, 2 low titer virus-positives, and 2 high titer virus positives (16-plex) or in the presence of 8 virus-negative, 4 low titer virus-positives, and 4 high titer virus positives (24-plex). The model virus libraries were tagged with different barcodes, notated in red or orange, respectively. (Virus-negative; grey, Low titer virus-positives; light green, high titter virus- positives; dark green).

[0027] FIG 4A shows violin plots for each capture of 24 libraries in normal human plasma. The BEACH control contains BK polyomavirus, EMCV, human Adenovirus E, Chlamydia trachomatis, and HIV-1. The PARVA control contains Parechovirus A3, Adeno-associated virus-2, Rotavirus A, Varicella-Zoster virus, and human Adenovirus C. Reads per million (RPM) results for each species are reported for eleven independent experiments. FIG 4B shows violin plots for each capture of 24 libraries in normal human plasma. The BEACH control contains BK polyomavirus, EMCV, human Adenovirus E, Chlamydia trachomatis, and HIV-1. The PARVA control contains Parechovirus A3, Adeno-associated virus-2, Rotavirus A, Varicella-Zoster virus, and human Adenovirus C. Percent genome coverage for each species was reported for ten independent experiments.

[0028] FIG. 5A shows target coverage plots for Chikungunya (left) and Zika (right) virus infections indicating an embodiment of sequencing performed on CVRP libraries (e.g., target enriched NGS [teNGS]) (blue) and mNGS (orange). FIG. 5B shows CVRP coverage plots for HIV-1 in whole blood (left), JC Polyomavirus (JCV) in urine (center), and EBV in nasal swabs (right). FIG. 5C shows AFI viral infections detected in patients from Thailand, sequenced by both teNGS (orange) and mNGS (blue), with results plotted by reads per million (RPM). FIG. 5D shows AFI viral infections detected in patients from Thailand, sequenced by both teNGS (orange) and mNGS (blue), results were plotted by percent genome coverage. FIG. 5E shows a summary of sequencing method output relating total reads collected to reference coverage. Symbols for categorization of genome completion are denoted in the legend.

[0029] FIG 6A shows graphs of teNGS libraries that were prepared from the same individual using plasma (blue) and whole blood (orange). Examples are shown for HIV-1 CRF01 AE (left) andDENV4 (right) strains. FIG 6B shows bar graphs of Pan-viral (PV) teNGS libraries for a positive control containing BK polyomavirus (top), EMCV (middle), and adenovirus E (bottom); Nextera libraries (brown; top bar) and XLT | QuantaBio libraries (green; bottom bar). Percent genome coverage (left) and reads per million (right) are reported for each virus. FIG 6C shows mNGS (blue) and teNGS (orange) libraries were prepared using the Quanta Bio cDNA synthesis and NGS library prep reagents. Examples are shown for Hepatitis E virus (HEV; right).

[0030] FIG. 7 A shows serial dilutions of Wanowrie virus that were processed (complete L segment; upper left panel; average in bold black line) and genome coverages for mNGS and CVRP were plotted for 1:100 (dilution 2; red) and 1:10,000 (dilution 4; blue) dilutions. Regions of enrichment are indicated by yellow shading. Individual pairwise identities are indicated by gray lines. FIG. 7B shows serial dilutions of Wanowrie virus that were processed (complete M segment; upper right panel; average in bold black line) and genome coverages for mNGS and CVRP were plotted for 1:100 (dilution 2; red) and 1:10,000 (dilution 4; blue) dilutions. Regions of enrichment are indicated by yellow shading. Individual pairwise identities are indicated by gray lines. FIG. 7C shows serial dilutions of Wanowrie virus that were processed (complete S segment; lower left panel; average in bold black line) and genome coverages for mNGS and CVRP were plotted for 1 : 100 (dilution 2; red) and 1 : 10,000 (dilution 4; blue) dilutions. Individual pairwise identities are indicated by gray lines. FIG. 7D shows serial dilutions of Mt. Elgon Bat virus were processed (full genome; lower right panel; average in bold black line) and genome coverages for mNGS and CVRP were plotted for 1 : 100 (dilution 2; red) and 1:10,000 (dilution 4; blue) dilutions. Regions of enrichment are indicated by yellow shading. Individual pairwise identities are indicated by gray lines. 0031] FIG 8A shows a coverage plot for the RdRp encoding segments from both the mNGS (blue) and teNGS (orange) libraries. FIG 8B shows a coverage plot for the capsid encoding segments from both the mNGS (blue) and teNGS (orange) libraries. FIG 8C shows a maximum-likelihood tree of RdRp amino acid sequences for the novel virus and other related fungal and invertebrate bunyaviruses wherein the novel virus bears only 75.6% AA identity to Pen-rose-RDRP_AYP71800 and 62.7% identity to Fusarium-poae-neg-str-vir-RDRP_YP_00927.

[0032] FIG 9A shows a panel of samples (n=7) from Japan with AFI following hematopoietic stem cell transplant, evaluated by both mNGS and teNGS, with confirmation attempted by PCR for JC Polyomavirus-2 (JC PyV-2). Positive controls or positive samples are highlighted in yellow. FIG 9B shows a panel of samples (n=7) from Japan with AFI following hematopoietic stem cell transplant, evaluated by both mNGS and teNGS, with confirmation attempted by PCR for Herpesvirus 6B (HHV-6B). Positive controls or positive samples are highlighted in yellow. FIG 9C shows a panel of samples (n=7) from Japan with AFI following hematopoietic stem cell transplant, evaluated by both mNGS and teNGS, with confirmation attempted by PCR for Picomavirus (PBV). Positive controls or positive samples are highlighted in yellow. FIG 9D shows a genome coverage plot of a Canineprotoparvovirus / feline panleukopenia. FIG 9E shows a genome coverage plot of a Canine bocaparvovirus.

[0033] FIG 10 illustrates a representative image of target enrichment that was applied to 599 specimens sourced from around the world exhibiting symptoms of respiratory (Argentina, Uganda) or acute febrile (Honduras, Bolivia, India, Japan) illness. Viral species and number of positives detected in each country are shown in boxes.

[0034] FIG 11 A illustrates a workflow describing the combination of target enriched and mNGS libraries into one sequencing run to obtain the desired throughput according to sample multiplexing. FIG 1 IB shows nine infections from Thailand that were evaluated for total reads (red; top row in each infection shown), RPM (blue; second row in each infection shown), and reference coverage (green; third row in each infection shown) in 5 separate sequencing experiments: pure mNGS (100%), teamNGS with 10% CVRP on a P2, teamNGS with 20% CVRP on a P2, teamNGS with 10% CVRP on a Pl, and pure CVRP (100%; teNGS). The nine infections from Thailand evaluated were (starting from the top left and moving right): influenza A (FLUA), HIV-1, human rhinovirus B (HRV-B), rabies virus (RABV), dengue virus stereotype 3 (DENV-3), dengue virus stereotype 4 (DENV-4), avian influenza virus (AIV), human adenovirus C (HAdV-C) and C. trachomatis. FIG 11C shows nine infections from Senegal that were evaluated for total reads (red; first row in each infection shown), RPM (blue; second row in each infection shown), and reference coverage (green; third row in each infection shown) in 5 separate sequencing experiments: pure mNGS (100%), teamNGS with 10% CVRP on a P2, teamNGS with 20% CVRP on a P2, teamNGS with 10% CVRP on a Pl, and pure CVRP (100%; teNGS). The nine infections from Senegal evaluated were (starting from the top left and moving right): Plasmodium falciparum (P. falciparum) , HIV-1, Plasmodium falciparum (P. falciparum), Saffold virus (SAFV), HIV-1, hepatitis B virus (HBV), Borrelia crocidurae (B. crocidurae), rotavirus A (ROTAV-A) and Varicella-zoster virus (VZV).

[0035] Figure 12 shows teamNGS detection of various pathogen classes across a diverse panel of specimen types derived from a cohort of hospitalized patients experiencing influenza-like illness (ILI). Forty patients were recruited into the study, with each patient having between one and four specimen types / bodily fluids collected for analysis. The results are reported in terms of relative reads per million (RPM-r), in which the number of NGS reads (per million collected) identified for each pathogen in each specimen is normalized against the number of reads (per million collected) identified for the same pathogens in any of the negative controls. Only microbes with known potential to cause human disease are shown. Only RPM-r values > 10 are considered significant (RPM-r values in grey are less than 10; purple - from 10 to 100; green from 100 to 1,000 and yellow, greater than 1 ,000). The viruses, bacteria, parasites and fungi tested included: 1. Viruses - Alphpolyomavirus quintihominis; Betacoronavirus 1, Betapolyomavirus hominis, Betapolyomavirus secuhominis, Brisavirus, Cytomegalovirus humanbeta 5, Hepacivirus hominis, Human betaherpesvirus 6, Human coronoavirus HKU1, Human immunodeficiency virus 1, Human respiratory circular DNA virus,Human respirovirus 3, Lymphocryptovirus humangamma4, Orthoflavivirus dengue, Orthopicobimavirus hominis, Primate erythroparvovirus, Roseolovirus humanbeta6b, Roseolovirus humanbeta7, Rotavirus A, Simplexvirus human alpha 1, and Vientovirus; 2. Bacteria - Campylobacter hominis, Chlamydia trachomatis, Corynebacterium pseudodiphtheriticum, Comebacterim striatum, Enterococcus faecalis, Enterococcus faecium, Escherichia coli, Gordonia bronchialis, Haemophilus influenzae, Klebsiella oxytoca, Kiebsiella pneumoniae, Klebsiella varicola, Moraxella catarrhalis, Neisseria gonorrhoeae, Neisseria meningitidis, Salmonella enterica, Shigella dysenteriae, Shigella flexneri, Sphingomonas paucimobilis, Stenotrophomonas maltophilia, Streptococcus gordonii, Streptococcus gordonii, Streptococcus infantis, Streptococcus pnuemoniae, and Streptococcus pyogenes; 3. Fungi - Altemaria altemata, Aspergillus sydowii, Aureobasidium pallulans, and Penicillin brevicompactum; and 4. Parasites - Toxoplasma gondii. Abbreviations: BAL - bronchoalveolar lavage fluid, NAS - nasal swab, NPH - nasopharyngeal swab, PLA - blood plasma, URN - urine, SER - serum, SPT - sputum, and STL - stool.DETAILED DESCRIPTION

[0036] Aspects of the present disclosure relate to methods, compositions and kits for detecting microbes in samples, for example, to determine whether a subject has a microbial infection (e.g., a viral infection). The methods, compositions and kits can also be used for characterizing metagenomes of samples. In various aspects, the approach described herein combines pools of enriched metagenomic libraries that are enriched for microbes of interest with pools of unenriched metagenomic libraries in a predetermined ratio of enriched metagenomic libraries to unenriched metagenomic libraries for next generation sequencing on a single sequencing run. Surprisingly and unexpectedly, the approach disclosed herein, referred to as target enrichment and metagenomic next generation sequencing (“teamNGS”), can achieve 100-10, 000X increases in depth and greater than 50% genomic coverage for samples with titers greater than or equal to 1000 cp / ml, while only requiring 3-4% the number of reads, and can be performed on a single sequencing run to detect all microbe types simultaneously, significantly saving time and resources.

[0037] Generally, the methods disclosed herein involve: (a) preparing unenriched libraries for metagenomic next generation sequencing from total nucleic acids extracted from a sample obtained from a subject suspected of having a microbial infection; (b) partitioning the unenriched libraries, on an equimolar basis, into a first set of pools of unenriched libraries and a second set of pools of unenriched libraries; (c) enriching each of the unenriched libraries in the first set of pools for nucleic acids of microbial origin to produce a first set of pools of enriched libraries that are enriched for nucleic acids of microbial origin; (d) combining a volume of each pool in the first set of pools of enriched libraries and a volume of each corresponding pool in the second set of pools of unenriched libraries in a predetermined volume ratio into a single flow cell for next generation sequencing; (e)sequencing the libraries in the combined volumes of each corresponding pool in a single flow cell using a next generation sequencing platform; and (f) characterizing the metagenomes of the at least one sample obtained from the at least one subject suspected of having the microbial infection based on the next generation sequencing results.

[0038] In a particular aspect, the methods involve (a) preparing a total of at least four unenriched libraries for metagenomic next generation sequencing from the total nucleic acids extracted from at least one sample obtained from at least one subject suspected of having a microbial infection, wherein each unenriched library comprises: (i) a first control unenriched nucleic acid library comprising nucleic acids of known microbial origin; (ii) a second control unenriched nucleic acid library comprising nucleic acids of human origin obtained from at least one human subject who is not suspected of having the microbial infection; (iii) a third control unenriched library that is substantially free of nucleic acids; or (iv) at least one test unenriched nucleic acid library comprising nucleic acids of human origin and nucleic acids of unknown microbial origin prepared from the at least one sample; wherein each unenriched nucleic acid library comprises fragmented double-stranded DNA that is barcoded for indexing; (b) partitioning the unenriched libraries, on an equimolar basis, into a first set of pools of unenriched libraries and a second set of pools of unenriched libraries; (c) enriching each of the unenriched libraries in the first set of pools for nucleic acids of microbial origin to produce a first set of pools of enriched libraries that are enriched for nucleic acids of microbial origin; (d) combining a volume from each pool in the first set of pools of enriched libraries, in a predetermined volume ratio, with a volume of each corresponding pool in the second set of pools of unenriched libraries into a single flow cell for next generation sequencing; (e) sequencing the libraries in the combined volumes of each corresponding pool in the single flow cell using a next generation sequencing platform; and (f) characterizing the metagenomes of the at least one sample obtained from the at least one subject suspected of having the microbial infection based on the next generation sequencing results. In other aspects, the disclosure provides compositions used in the methods and kits for implementing the methods.(0039] Section headings as used in this section and the entire disclosure herein are merely for organizational pinposes and are not intended to be limiting.1. Definitions

[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. In case of conflict, the present document, including definitions, will control. Preferred methods and materials are described below, although methods and materials similar or equivalent to those described herein can be used in practice or testing of the present disclosure. All publications, patent applications, patents and other referencesmentioned herein are incorporated by reference in their entirety. The materials, methods, and examples disclosed herein are illustrative only and not intended to be limiting.

[0041] The terms “comprise(s),” “include(s),” “having,” “has,” “can,” “contain(s),” and variants thereof, as used herein, are intended to be open-ended transitional phrases, terms, or words that do not preclude the possibility of additional acts or structures. The singular forms “a,” “an” and “the” include plural references unless the context clearly dictates otherwise. The present disclosure also contemplates other embodiments “comprising,” “consisting of’ and “consisting essentially of,” the embodiments or elements presented herein, whether explicitly set forth or not.[01)42] For the recitation of numeric ranges herein, each intervening number there between with the same degree of precision is explicitly contemplated. For example, for the range of 6-9, the numbers 7 and 8 are contemplated in addition to 6 and 9, and for the range 6.0-7.0, the number 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, and 7.0 are explicitly contemplated.

[0043] As used herein, the term “adapter” refers to a sequence that permits universal amplification. A key feature of the adapter is to enable the unique amplification of only a hybridized nucleic acid without the need to remove an existing template nucleic acid or purify the hybridized nucleic acid. This feature enables an “add only” reaction with fewer steps and ease of automation. The adapter is attached to the 5' and 3' end of the hybridized nucleic acid. The adapter may be Y-shaped, U-shaped, hairpin-shaped, or a combination thereof. In a specific embodiment, the adaptor is Y- shaped. In some embodiments, the adapter may be an Illumina adapter for Illumina sequencing. An index sequence may also be attached to each nucleic acid fragment. The addition of an index sequence allows pooling of multiple samples into a single sequencing run. This greatly increases experimental scalability, while maintaining extremely low error rates and conserving read length. The index sequence maybe about 5 to about 10 nucleotides. Accordingly, the index sequence maybe 5, 6, 7, 8, 9 or 10 or more nucleotides. In an embodiment, the index sequence is about 6 nucleotides. An “amount” as used herein refers to a quantity specified (e.g., high or low) or a number e.g., where the number is a level, such as a position on a real or imaginary scale of amount or quantity, or a concentration, such as, for example, a relative amount of a given substance contained within a solution or in a particular volume of space, e.g., the amount of solute per unit volume of solution.

[0044] As used herein, the terms “amplifying” or “amplification” in the context of nucleic acids refers to the production of multiple copies of a polynucleotide, or a portion of the polynucleotide, typically starting from a small amount of the polynucleotide (e.g., a single polynucleotide molecule), where the amplification products or amplicons are generally detectable. Amplification of polynucleotides encompasses a variety of chemical and enzymatic processes. Generation of multiple DNA copies from one or a few copies of a target or template DNA molecule during a polymerase chain reaction (PCR), isothermal reaction, or a ligase chain reaction (LCR) are forms of amplification. Amplification is not limited to the strict duplication of the starting molecule. For example, the generation of multiple cDNA molecules from a limited amount of RNA in a sample using reversetranscription (RT)-PCR is a form of amplification. Furthermore, the generation of multiple RNA molecules from a single DNA molecule during the process of transcription is also a form of amplification.

[0045] As used herein, the terms "characterization", "characterize", "characterizing" refer to describing or categorizing by sequence information in some cases.

[0046] The terms “bead,” “microparticle,” and “particle” used herein interchangeably herein, refer to a substantially spherical solid support. The bead may be a magnetic bead or a magnetic particle. Microparticles that can be used herein can be any type known in the art. For example, the bead or microparticle can be a magnetic bead or magnetic particle. Magnetic beads / particles may be ferromagnetic, ferrimagnetic, paramagnetic, superparamagnetic or ferrofluidic. Examples of ferromagnetic materials include Fe, Co, Ni, Gd, Dy, CrO2, MnAs, MnBi, EuO, and NiO / Fe.Examples of ferrimagnetic materials include NiFe2O4, CoFe2O4, Fe3O4 (or FeO.Fe2O3). Beads can have a solid core portion that is magnetic and is surrounded by one or more non-magnetic layers. Alternately, the magnetic portion can be a layer around a non-magnetic core. The size of the microparticles used in the methods, compositions and kits described herein can vary. In some aspects, the microparticles have a substantially uniform diameter of less than about 0.10 pm, less than about 0.20 pm, less than about 0.30 pm, less than about 0.40 pm, less than about 0.50 pm, less than about 0.60 pm, less than about 0.70 pm, less than about 0.80 pm, less than about 0.9 pm, less than about 1 pm, less than about 2 pm, less than about 3 pm, less than about 4 pm, or less than about 5 pm.

[0047] As used herein, "coding sequence" or a sequence "encoding" an expression product, such as a RNA, polypeptide, protein, or enzyme, is used herein to refer to a nucleotide sequence which, when expressed, results in the production of that RNA, polypeptide, protein, or enzyme, i.e., the nucleotide sequence encodes an amino acid sequence for that polypeptide, protein or enzyme. A coding sequence for a protein may include a start codon (usually ATG) and a stop codon.

[0048] The terms “complementary" or "complementarity" are used herein in reference to "polynucleotides" and "oligonucleotides" (which are interchangeable terms that refer to a sequence of nucleotides) related by the base-pairing rules. These terms may also include mimics of or artificial bases that may not faithfully adhere to the base-pairing rules. For example, the sequence "C-A-G-T," is complementary to the sequence "G-T-C-A." Complementarity can be "partial" or "total." "Partial" complementarity is where one or more nucleic acid bases are not matched according to the base pairing rules. "Total" or "complete" complementarity between nucleic acids is where each and every nucleic acid base is matched with another base under the base pairing rules. The degree of complementarity between nucleic acid strands has significant effects on the efficiency and strength of hybridization between nucleic acid strands. This is of particular importance in amplification reactions, as well as detection methods which depend upon binding between nucleic acids.

[0049] As used herein, the term “fragment" refers to a portion of a nucleotide sequence. Fragments may range in size from 5 nucleotide residues to the entire nucleotide sequence minus one nucleic acid residue.

[0050] As used herein, the term “gene" refers to deoxyribonucleotide or ribonucleotide sequences comprising the coding region of a structural gene and including sequences located adjacent to the coding region on both the 5' and 3' ends for a distance of about 1 kb on either end such that the gene corresponds to the length of the full-length mRNA. The sequences which are located 5' of the coding region and which are present on the mRNA are referred to as 5' non-translated sequences. The sequences which are located 3' or downstream of the coding region and which are present on the mRNA are referred to as 3' non-translated sequences. The term "gene" encompasses both amplified and genomic forms of a gene. A genomic form or clone of a gene contains the coding region interrupted with non-coding sequences termed "introns" or "intervening regions" or "intervening sequences." Introns are segments of a gene which are transcribed into heterogeneous nuclear RNA (hnRNA); introns may contain regulatory elements such as enhancers. Introns are removed or "spliced out" from the nuclear or primary transcript; introns therefore are absent in the messenger RNA (mRNA) transcript. The mRNA functions during translation to specify the sequence or order of amino acids in a nascent polypeptide.

[0051] As used herein, the term “genome" refers to the entirety of an organism's hereditary information that is encoded in its primary DNA or RNA or nucleotide sequence (DNA or RNA as applicable). The genome includes both the genes and the non-coding sequences. For example, the genome may represent a microbial genome, such as an archaeal genome, a fungal genome (e.g., yeast), a protozoan genome, or a viral genome.

[0052] As used herein the phrase "hybridization product" refers to a complex formed between two nucleic acid sequences by virtue of the formation of hydrogen bounds between complementary G and C bases and between complementary A and T bases; these hydrogen bonds may be further stabilized by base stacking interactions. The two complementary nucleic acid sequences hydrogen bond in an antiparallel configuration. A hybridization product may be formed in solution or between one nucleic acid sequence present in solution and another nucleic acid sequence immobilized to a solid support (e.g., hybridization of a biotinylated probe used in target enrichment step (c) to a nucleic acid of interest and magnetic streptavidin beads).

[0053] As used herein, the terms "identification", "identify", and "identifying" refer to recognizing a specific microbe or microbes in a sample from a subject, for example, archaea, bacteria, fungi (e.g., yeast), protozoa, and / or viruses in sample from a subject.

[0054] As used herein, the term "identifier" refers to any unique, non-naturally occurring, nucleic acid sequence that may be used to identify the originating genome of a nucleic acid fragment. The identifier function can sometimes be combined with other functionalities such as adapters or primers and can be located at any convenient position.

[0955] As used herein, the term "isolated" means that the referenced material, such as, for example, biological material (e.g., DNA, cDNA, RNA, mRNA, a nucleic acid, a protein, or a polypeptide), is free of components found in the natural environment in which the material is normally found. In particular, isolated biological material is free of cellular components. In the case of nucleic acid molecules, an isolated nucleic acid includes a PCR product, an isolated mRNA, a cDNA, an isolated genomic DNA, or a restriction fragment. In some embodiments, an isolated nucleic acid is preferably excised from the chromosome in which it may be found. Isolated nucleic acid molecules can be inserted into plasmids, cosmids, artificial chromosomes, and the like. Thus, in some embodiments, a recombinant nucleic acid is an isolated nucleic acid. An isolated protein may be associated with other proteins or nucleic acids, or both, with which it associates in the cell, or with cellular membranes if it is a membrane-associated protein. In some embodiments, an isolated material can be purified.

[9056] As used herein, the term "library” refers to a collection or compilation of nucleic acids compatible with a nucleic acid sequencing device, such as, for example, a next-generation high throughput sequencing device.[0057 As used herein, the phrase "next-generation sequencing " refers to any nucleic acid sequencing device or methodology that utilizes massively parallel technology. For example, such a platform may include, but is not limited to, ILLUMINA sequencing platforms.

[0058] As used herein, the phrases "nucleic acid", and "polynucleotide" and "nucleic acid sequence" and "nucleotide sequence" include a nucleic acid, an oligonucleotide, a nucleotide, a polynucleotide, and any fragment, variant, or derivative thereof. The nucleic acid or polynucleotide may be double- stranded, single- stranded, or triple-stranded DNA or RNA (including cDNA), or a DNA-RNA hybrid of genetic or synthetic origin, wherein the nucleic acid contains any combination of deoxyribonucleotides and ribonucleotides and any combination of bases, including, but not limited to, adenine, thymine, cytosine, guanine, uracil, inosine, and xanthine hypoxanthine. As further used herein, the term "cDNA" refers to an isolated DNA polynucleotide or nucleic acid molecule, or any fragment, derivative, or complement thereof. It may be double -stranded, single- stranded, or triple- stranded, it may have originated recombinantly or synthetically, and it may represent coding and / or noncoding 5' and / or 3' sequences.

[0059] As used herein, the phrase, “nucleic acids of microbial origin” refers to nucleic acid molecules, including DNA and RNA, that are derived from microorganisms such as bacteria, viruses, fungi, archaea, and other microbial entities present in a biological sample obtained from a human subject. This term specifically excludes nucleic acids of human origin and encompasses all forms of microbial genetic material, including but not limited to, genomic DNA, plasmid DNA, viral RNA, ribosomal RNA (rRNA), messenger RNA (mRNA), and non-coding RNAs of microbial species. These nucleic acids are characterized by sequences that are distinct from the human genome andtranscriptome, thereby enabling their identification, quantification, and analysis separate from human nucleic acids within the same sample.

[0060] As used herein, "nucleic acid hybridization" or "hybridization" refer to anti-parallel hydrogen bonding between two single-stranded nucleic acids, in which A pairs with T (or U if an RNA nucleic acid) and C pairs with G. Nucleic acid molecules are "hybridizable" to each other when at least one strand of one nucleic acid molecule can form hydrogen bonds with the complementary bases of another nucleic acid molecule under defined stringency conditions. Stringency of hybridization is determined, e.g., by (i) the temperature at which hybridization and / or washing is performed, and (ii) the ionic strength and (iii) concentration of denaturants such as formamide of the hybridization and washing solutions, as well as other parameters. Hybridization requires that the two strands contain substantially complementary sequences. Depending on the stringency of hybridization, however, some degree of mismatches may be tolerated. Under "low stringency" conditions, a greater percentage of mismatches are tolerable (i.e., will not prevent formation of an anti-parallel hybrid).[0061 J As used herein, the term "oligonucleotide" refers to a nucleic acid, generally of at least 10, at least 15, at least 20 nucleotides, at least 30 nucleotides, at least 60 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 110 nucleotides, at least 120 nucleotides, at least 130 nucleotides, at least 140 nucleotides, at least 150 nucleotides, at least 160 nucleotides, or at least 180 nucleotides, that is hybridizable to a genomic DNA molecule, a cDNA molecule, or an mRNA molecule encoding a gene, mRNA, cDNA, or other nucleic acid of interest. The nucleic acids that comprise the oligonucleotides include but are not limited to DNA, RNA, linked nucleic acids (LNA), bridged nucleic acids (BNA) and peptide nucleic acids (PNA). Oligonucleotides can be labeled, e.g., with 32 P-nucleotides or nucleotides to which a label, such as biotin, has been covalently conjugated.

[0062] As used herein, the phrases "percent (%) sequence similarity", "percent (%) sequence identity", refer to the degree of identity or correspondence between different nucleotide sequences of nucleic acid molecules or amino acid sequences of proteins that may or may not share a common evolutionary origin. Sequence identity can be determined using any publicly available sequence comparison algorithms, such as, for example, BLAST, FASTA, DNA Strider, and GCG (Genetics Computer Group, Program Manual for the GCG Package, Version 7, Madison, Wisconsin).

[0063] To determine the percent identity between two amino acid sequences or two nucleic acid molecules, the sequences are aligned for optimal comparison purposes. The percent identity between the two sequences is a function of the number of identical positions shared by the sequences (i.e., percent identity = number of identical positions / total number of positions (e.g., overlapping positions) x 100). In one embodiment, the two sequences are, or are about, of the same length. The percent identity between two sequences can be determined using techniques similar to those described below,with or without allowing gaps. In calculating percent sequence identity, typically exact matches are counted.

[0064] As used herein, the phrase “Polymerase chain reaction" ("PCR") refers to the method disclosed in U.S. Patent Nos. 4,683,195 and 4,683,202, herein incorporated by reference, which describe a method for increasing the concentration of a segment of a target sequence in a mixture of genomic DNA without cloning or purification. The length of the amplified segment of the desired target sequence is determined by the relative positions of two oligonucleotide primers with respect to each other, and therefore, this length is a controllable parameter. By virtue of the repeating aspect of the process, the method is referred to as the "polymerase chain reaction" (hereinafter "PCR").Because the desired amplified segments of the target sequence become the predominant sequences (in terms of concentration) in the mixture, they are said to be "PCR amplified". With PCR, it is possible to amplify a single copy of a specific target sequence in genomic DNA to a level detectable by several different methodologies (e.g., hybridization with a labeled probe; incorporation of biotinylated primers followed by avidin-enzyme conjugate detection; incorporation of 32P-labeled deoxynucleotide triphosphates, such as dCTP or dATP, into the amplified segment). In addition to genomic DNA, any oligonucleotide sequence can be amplified with the appropriate set of primer molecules. In particular, the amplified segments created by the PCR process itself are, themselves, efficient templates for subsequent PCR amplifications. With PCR, it is also possible to amplify a complex mixture (library) of linear DNA molecules, provided they carry suitable universal sequences on either end such that universal PCR primers bind outside of the DNA molecules that are to be amplified.

[6065] As used herein, the term “primer” refers to an oligonucleotide, whether occurring naturally as in a purified restriction digest or produced synthetically, that is capable of acting as a point of initiation of synthesis when placed under conditions in which synthesis of a primer extension product that is complementary to a nucleic acid strand is induced (e.g., in the presence of nucleotides and an inducing agent such as a biocatalyst (e.g., a DNA polymerase or the like) and at a suitable temperature and pH). The primer is typically single stranded for maximum efficiency in amplification but may alternatively be double stranded. If double stranded, the primer is generally first treated to separate its strands before being used to prepare extension products. In some embodiments, the primer is an oligodeoxyribonucleotide. The primer is sufficiently long to prime the synthesis of extension products in the presence of the inducing agent. The exact lengths of the primers will depend on many factors, including temperature, source of primer and the use of the method.

[0066] As used herein, the term "sequencing" refers to any methods for determining the order of the nucleotide bases, adenine, guanine, cytosine, and thymine / uracil, in a molecule of DNA or RNA.

[0067] As used herein, the phrase “statistically significant” refers to the likelihood that a relationship between two or more variables is caused by something other than random chance. Statistical hypothesis testing is used to determine whether the result of a data set is statisticallysignificant. In statistical hypothesis testing, a statistically significant result is attained whenever the observed p-value of a test statistic is less than the significance level defined of study. The p-value is the probability of obtaining results at least as extreme as those observed, given that the null hypothesis is true. Examples of statistical hypothesis analysis include Wilcoxon signed-rank test, t-test, Chi- Square or Fisher’s exact test. As used herein, the term “significant” refers to a change that has not been determined to be statistically significant (e.g., it may not have been subject to statistical hypothesis testing).

[0068] As used herein, the term "stringency" is used in reference to the conditions of temperature, ionic strength, and the presence of other compounds such as organic solvents, under which nucleic acid hybridizations are conducted. "Stringency" typically occurs in a range from about Tm to about 20°C to 25°C below Tm. A "stringent hybridization" can be used to identify or detect identical polynucleotide sequences or to identify or detect similar or related polynucleotide sequences. For example, when fragments are employed in hybridization reactions under stringent conditions the hybridization of fragments which contain unique sequences (i.e., regions which are either non- homologous to or which contain less than about 50% homology or complementarity) are favored. Alternatively, when conditions of "weak" or "low" stringency are used hybridization may occur with nucleic acids that are derived from organisms that are genetically diverse (i.e., for example, the frequency of complementary sequences is usually low between such organisms).

[0069] The terms “subject” and “patient” as used herein interchangeably refers to any vertebrate, including, but not limited to, a mammal (e.g., cow, pig, camel, llama, horse, goat, rabbit, sheep, hamsters, guinea pig, cat, dog, rat, and mouse, a non-human primate (for example, a monkey, such as a cynomolgus or rhesus monkey, chimpanzee, etc.) and a human). In some embodiments, the subject may be a human or a non-human. In some embodiments, the subject is a human. The subject or patient may be undergoing other forms of treatment. The subject or patient may be undergoing monitoring of the course of a microbial infection and / or treatment of a microbial infection. The subject or patient may have or be suspected of having a viral infection and / or a symptom, disease, condition, or disorder associated with the viral infection. The viral infection can be from a virus of any order listed in Table 1, any family listed in Table 2, any genus listed in Table 3, any species listed in Table 4, or any strain listed in Table 5.

[0070] As used herein, “substantially free of nucleic acids” in the context of a control library means that the library contains only trace or negligible amounts of nucleic acids (both DNA and RNA), such that these nucleic acids do not significantly interfere with or contribute to the library preparation process or subsequent sequencing results. In this context, "substantially free" means that the concentration of nucleic acids is below the detectable limit of standard analytical methods used in the relevant field, or is present at levels insufficient to affect the accuracy, sensitivity, or specificity of next generation sequencing. This ensures that any results obtained are attributed solely to the nucleicacids of interest (e.g., archaeal, bacterial, fungal, protozoan, viral nucleic acids, etc.) from a test library, without contamination from extraneous nucleic acid sources.

[0071] As used herein, "target enrichment and metagenomic sequencing" and "teamNGS" as used interchangeably herein, to refer to the novel capture sequencing method of the disclosure that allows the simultaneous detection, identification and / or characterization of all microbes (e.g., viruses) known or suspected to infect vertebrates in any single sample in a next generation sequencing run on a single flow cell. The phrase denotes the platform in every form, including but not limited to the collection of synthetic oligonucleotides representing the coding sequences of at least one virus from every viral taxa known to infect vertebrates (i.e., "probe library"), either in solution or attached to a solid support, a database comprising information on the teamNGS platform including at least the length, nucleotide sequence, melting temperature, and microbial (e.g., viral) origin of each oligonucleotide probe, and computer-readable storage mediums with program code comprising information on the teamNGS platform including at least the length, nucleotide sequence, melting temperature, and microbial (e.g., viral) origin of each oligonucleotide probe.

[0072] As used herein, “single flow cell” refers to an individual, integrated microfluidic device used in next generation sequencing platforms, which contains a network of channels, wells, or lanes designed to accommodate the simultaneous sequencing of multiple nucleic acid samples. A single flow cell serves as the physical substrate on which sequencing reactions occur, allowing for the capture, amplification, and detection of nucleic acid fragments. It is configured to process samples in parallel, typically through the use of sequencing-by-synthesis or other sequencing chemistries and enables the collection of sequencing data from a defined number of spatially separated and independently addressable sites within the same unit, without the need for multiple devices or additional flow cells.

[0073] As used herein, “melting temperature” or "Tm" as used interchangeably herein, refers to the temperature at which a population of double-stranded nucleic acid molecules becomes half dissociated into single strands. As indicated by standard references, a simple estimate of the Tm value may be calculated by the equation: Tm=81.5+0.41 (% G+C), when a nucleic acid is in aqueous solution at IM NaCl. Anderson et al, "Quantitative Filter Hybridization" In: Nucleic Acid Hybridization (1985). More sophisticated computations take structural, as well as sequence characteristics, into account for the calculation of Tm.

[0074] As used herein, “total nucleic acids” refers to the complete set of nucleic acid molecules, including both DNA (deoxyribonucleic acid) and RNA (ribonucleic acid), extracted from a biological sample. This encompasses all forms of DNA and RNA present, such as genomic DNA, mitochondrial DNA, plasmid DNA, messenger RNA (mRNA), ribosomal RNA (rRNA), transfer RNA (tRNA), microRNA (miRNA), small interfering RNA (siRNA), long non-coding RNA (IncRNA), and other non-coding RNAs. The term "total nucleic acids" specifically includes all nucleic acid sequences, regardless of their length, structure, or modifications, and is intended to represent the completegenetic and transcriptomic material that can be used for downstream library preparation and sequencing applications.

[0075] Unless otherwise defined herein, scientific and technical terms used in connection with the present disclosure shall have the meanings that are commonly understood by those of ordinary skill in the art. For example, any nomenclatures used in connection with, and techniques of, cell and tissue culture, molecular biology, immunology, microbiology, genetics and protein and nucleic acid chemistry and hybridization described herein are those that are well known and commonly used in the art. The meaning and scope of the terms should be clear; in the event, however of any latent ambiguity, definitions provided herein take precedent over any dictionary or extrinsic definition. Further, unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular.2. Methods of Detecting Microbes and / or Characterizing Metagenomes in Samples

[0076] In an aspect, the disclosure provides a method of detecting one or more microbes in a sample. In another aspect, the disclosure provides a method of characterizing the metagenomes of samples. In yet another aspect, the disclosure provides a method of detecting a microbial infection in a subject. In a further aspect, the disclosure provides a method of aiding in the diagnosis of a disease, condition and / or disorder associated with a microbial infection in a subject. In an additional aspect, the disclosure provides a method of identifying a pathogen in a sample. In another aspect, the disclosure provides a method of determining whether an outbreak of a microbial infection or associated disease, condition, and / or disorder is pandemic, epidemic, or endemic. In yet another additional aspect, the disclosure provides a method of monitoring the spread of a microbial infection or associated disease, condition, and / or disorder during a pandemic, epidemic, and / or endemic.

[0077] Generally, the methods of the disclosure involve the steps of (a) preparing unenriched libraries for metagenomic next generation sequencing from total nucleic acids extracted from one or more samples, for example, one or more samples obtained from subjects suspected of having a microbial infection; (b) partitioning the unenriched libraries, on an equimolar basis, into a first set of pools of unenriched libraries and a second set of pools of unenriched libraries; (c) enriching each of the unenriched libraries in the first set of pools for nucleic acids of microbial origin to produce a first set of pools of enriched libraries that are enriched for nucleic acids of microbial origin; (d) combining a volume from each pool in the first set of pools of enriched libraries, in a predetermined volume ratio, with a volume from each corresponding pool in the second set of pools of unenriched libraries into a single flow cell for next generation sequencing; (e) sequencing the libraries in the combined volumes of each corresponding pool in a single flow cell using a next generation sequencing platform; and (f) analyzing the next generation sequencing results in accordance with the intended use of the method being performed. For example, the analysis could involve detecting one or more microbes in the sample(s), characterizing the metagenomes of the samples, detecting microbial infections in asubjects), aiding in the diagnosis of a disease, condition and / or disorder associated with a microbial infection in a subject, identifying a pathogen in a sample, determining whether an outbreak of a microbial infection or associated disease, condition, and / or disorder is pandemic, epidemic, or endemic, and / or monitoring the spread of a microbial infection or associated disease, condition, and / or disorder during a pandemic, epidemic, and / or endemic.

[0078] In one embodiment of the disclosed methods, samples, such as one or more samples obtained from a subject having or suspected of having a microbial infection comprising nucleic acids of human origin and nucleic acids of unknown microbial origin, a first control sample comprising nucleic acids of known microbial origin, a second control sample comprising nucleic acids of human origin obtained from at least one human subject who is not suspected of having the microbial infection, and a third control sample that is substantially free of nucleic acids are treated with a nuclease and total nucleic acids are extracted and fragmented, such as, for example, by mechanical or enzymatic shearing to form a library of fragments from each sample. Universal adapters can be ligated to fragmented sample nucleic acids to form adapter-ligated sample unenriched nucleic acid libraries. The libraries can then be amplified using a barcoded primer library to generate barcoded adapter-sample unenriched nucleic acid libraries. The unenriched libraries can be partitioned, such as on an equimolar basis, into a first set of pools of unenriched libraries and a second set of pools of unenriched libraries. The libraries may then be subjected to target enrichment, such as, for example, by hybridization with a plurality of oligonucleotide probes designed to hybridize to the nucleic acids of microbial origin, along with blocking oligonucleotides that prevent hybridization between probes and adapters. Capture of the sample nucleic acids-probe hybridization pairs and removal of probes allows for isolation / enrichment of nucleic acids of unknown microbial origin, which are then amplified and optionally sequenced. Various combinations of universal adapters and barcoded primers may be used. In some embodiments, barcoded primers comprise at least one barcode. In some embodiments, different types of barcodes are added to the sample nucleic acid using adapters or barcodes, or both. For example, a universal adapter can comprise an index barcode, and after ligation can be amplified with a barcoded primer comprising an additional index barcode. In some embodiments, a universal adapter comprises a unique molecular identifier barcode, and after ligation can be amplified with a barcoded primer comprising an index barcode. a. Step (a) library preparation

[0079] Aspects of the disclosure involve preparing unenriched libraries for sequencing, particularly, next generation sequencing (e.g., metagenomic next generation sequencing). Methods of preparing sequencing libraries are well known in the art. The disclosure contemplates using any method of library preparation for step (a) of the methods described herein.

[0080] The library preparation in step (a) can comprise contacting at least one sample of the at least one microbe, at least one sample from the at least one human subject, and / or at least one sampleobtained from the at least one subject suspected of having the microbial infection with a nuclease (e.g., BENZONASE®). Subsequent to, or simultaneously with contacting the samples with the nuclease, the method includes (i) extracting total nucleic acids from at least one sample from at least one microbe to prepare the first control unenriched nucleic acid library; (ii) extracting total nucleic acids from at least one sample obtained from at least one human subject who is not suspected of having the microbial infection to prepare the second control unenriched nucleic acid library; and / or (iii) extracting total nucleic acids from the at least one sample obtained from the at least one subject suspected of having the microbial infection to prepare the at least one test unenriched nucleic acid library.

[0081] Each of the at least one samples can also be contacted with a reverse transcriptase to synthesize single-stranded cDNA from total RNA extracted from the samples. The single-stranded cDNA can be contacted with a DNA polymerase to produce double-stranded DNA. The double- stranded DNA can then be fragmented, for example, by enzymatic digestion, sonication, nebulization, and / or hydrodynamic shearing to produce double-stranded DNA fragments.

[0082] The double-stranded DNA fragments can comprise various lengths. In some embodiments, the double-stranded DNA fragments comprise nucleic acids having a length of between about 100 nt and 300 nt, between about 105 nt and about 275 nt, between about 110 nt and about 250 nt, between about 115 nt and about 225 nt, between about 120 nt and about 200 nt, between about 125 nt and about 175 nt, between about 130 and 170 nt, between about 135 and 165 nt, between about 140 nt and about 160 nt, or about 145 nt about 146 nt, about 147 nt, about 148 nt, about 149 nt, about 150 nt, about 151 nt, about 152 nt, about 153 nt, about 154 nt, or about 155 nt.

[0083] Subsequent to, or simultaneously with fragmentation of the double-stranded DNA, step (a) comprises ligating the fragmented double-stranded DNA with barcode adaptor sequences for indexing. In some embodiments, fragmenting and / or ligating the double-stranded DNA with barcode adaptor sequences for indexing comprises subjecting the double-stranded DNA in the samples to transposome-mediated tagmentation.

[0084] The methods are not limited to any particular type of barcode adaptors. In an embodiment, the barcode adaptors comprise unique dual-index barcode adaptors.

[0085] The barcoded double-stranded DNA in the samples can also be purified and / or quantified using routine methods. The double- stranded DNA in the samples can comprise cDNA and genomic DNA.

[0086] The methods also contemplate the preparation of unenriched libraries from a large number of samples, for example, samples obtained from at least one subject suspected of having a microbial infection. In some embodiments, at least one sample comprises at least 2 samples, at least 3 samples, at least 5 samples, at least 6 samples, at least 9 samples, at least 12 samples, at least 18 samples, at least 21 samples, at least 24 samples, at least 30 samples, at least 36, at least 42 samples, at least 45 samples, at least 72 samples, at least 78 samples, at least 84 samples, at least 90 samples, at least 312samples, at least 336 samples, at least 348 samples, at least 354 samples, at least 366 samples, at least 1152 samples, at least 1248 samples, at least 1344 samples, at least 1392 samples, or at least 1440 samples obtained from subjects suspected of having the microbial infection.

[0087] The methods contemplate preparation of a large number of unenriched libraries during step (a) from the at least one sample. In some embodiments, at least five, at least six, at least eight, at least 12, at least 24, at least 48, at least 96, at least 384, or at least 1536 unenriched libraries are prepared during library preparation step (a) from the sample(s).

[0088] In one embodiment, the library preparation of step (a) comprises preparing at least four total unenriched libraries for metagenomic next generation sequencing from total nucleic acids extracted from at least one sample obtained from at least one subject suspected of having a microbial infection, where each unenriched library comprises: (i) a first control unenriched nucleic acid library comprising nucleic acids of known microbial origin; (ii) a second control unenriched nucleic acid library comprising nucleic acids of human origin obtained from at least one human subject who is not suspected of having the microbial infection; (iii) a third control unenriched library that is substantially free of nucleic acids; or (iv) at least one test unenriched nucleic acid library comprising nucleic acids of human origin and nucleic acids of unknown microbial origin prepared from the at least one sample; where each unenriched nucleic acid library comprises fragmented double-stranded DNA that is barcoded for indexing. b. Step (b) partitioning unenriched libraries into a first and second set of pools

[0089] Aspects of the disclosure involve partitioning the unenriched libraries produced in library preparation step (a), on an equimolar basis, into a first set of pools of unenriched libraries and a second set of pools of unenriched libraries. The methods are not limited to any particular number of pools in each set. The first set of pools and the second set of pools can comprise at least one pool, at least two pools, at least three pools, at least four pools, at least six pools, at least eight pools, at least 12 pools, at least 16 pools, at least 24 pools, at least 32 pools, at least 48 pools, at least 64 pools, at least 96 pools, or at least 128 pools of unenriched libraries.

[0090] In an embodiment, each set of the at least one pool comprises at least 4, at least 5, at least 6, at least 12, at least 24, or at least 48 unenriched libraries. In another embodiment, each set of the at least two pools comprises at least 12, at least 24, at least 48, or at least 96 unenriched libraries. In another embodiment, each set of the at least three pools comprises at least 12, at least 24, at least 48, or at least 96 unenriched libraries. In another embodiment, each set of the at least three pools comprises at least 12 nucleic acid libraries. In another embodiment, each set of the at least four pools comprises at least 24, at least 48, or at least 96 unenriched libraries. In another embodiment, each set of the at least six pools comprises at least 24, at least 48, at least 96, or at least 384 unenriched libraries. In another embodiment, each set of the at least eight pools comprises at least 24, at least 48,at least 96, or at least 384 unenriched libraries. In another embodiment, each set of the at least 12 pools comprises at least 48, at least 96, or at least 384 unenriched libraries. In another embodiment, each set of the at least 16 pools comprises at least 384 unenriched libraries. In another embodiment, each set of the at least 24 pools comprises at least 384 unenriched libraries. In another embodiment, each set of the at least 32 pools comprises at least 1536 unenriched libraries. In another embodiment, each set of the at least 48 pools comprises at least 1536 unenriched libraries. In another embodiment, each set of the at least 64 pools comprises at least 1536 unenriched libraries. In another embodiment, each set of the at least 96 pools comprises at least 1536 unenriched libraries. In another embodiment, each set of the at least 128 pools comprises at least 1536 unenriched libraries.

[0091] In one embodiment, each set of pools comprises: (i) a first control unenriched nucleic acid library comprising nucleic acids of known microbial origin; (ii) a second control unenriched nucleic acid library comprising nucleic acids of human origin obtained from at least one human subject who is not suspected of having the microbial infection; (iii) a third control unenriched library that is substantially free of nucleic acids; and (iv) at least X test unenriched nucleic acid libraries prepared from at least a Y samples obtained from at least Y subjects suspected of having the microbial infection; where X is equal to a variable number Z that is three less than the total number of unenriched libraries prepared divided by the number of pools in each set, where Y is equal to the number of pools in each set times Z, where each test unenriched nucleic acid library comprises nucleic acids of human origin and nucleic acids of unknown microbial, and where each unenriched nucleic acid library comprises fragmented double-stranded DNA that is barcoded for indexing.

[0092] Following partitioning step (b), and prior to target enrichment step (c), the method can comprise drying each pool of unenriched libraries in the first set of pools to produce pellets comprising the double-stranded DNA in each pool of unenriched libraries.

[0093] In one embodiment, step (a) comprises: preparing at least eight total unenriched libraries for metagenomic next generation sequencing from total nucleic acids extracted from at least five samples obtained from at least five subjects suspected of having a microbial infection, where the at least eight total unenriched libraries comprise: (i) a first control unenriched nucleic acid library comprising nucleic acids of known microbial origin; (ii) a second control unenriched nucleic acid library comprising nucleic acids of human origin obtained from at least one human subject who is not suspected of having the microbial infection; (iii) a third control unenriched library that is substantially free of nucleic acids; and (iv) at least five test unenriched nucleic acid libraries each comprising nucleic acids of human origin and nucleic acids of unknown microbial origin prepared from the at least five samples, where each unenriched nucleic acid library comprises fragmented double-stranded DNA that is barcoded for indexing.

[0094] In one embodiment, step (b) comprises: partitioning the at least eight total unenriched libraries, on an equimolar basis, into a first set of one pool of unenriched libraries and a second set of one pool of unenriched libraries, where the one pool of unenriched libraries in each of the first andsecond sets comprises an aliquot taken from: (i) the first control unenriched nucleic acid library; (ii) the second control unenriched nucleic acid library; (iii) the third control unenriched library; and (iv) the at least five test unenriched nucleic acid libraries.

[0095] In another embodiment, where step (a) comprises: preparing at least 48 total unenriched libraries for metagenomic next generation sequencing from total nucleic acids extracted from at least 42 samples obtained from at least 42 subjects suspected of having a microbial infection, the at least 48 total unenriched libraries comprise: (i) at least two first control unenriched nucleic acid libraries comprising nucleic acids of known microbial origin; (ii) at least two second control unenriched nucleic acid libraries comprising nucleic acids of human origin obtained from at least one human subject who is not suspected of having the microbial infection; (iii) at least two third control unenriched libraries that are substantially free of nucleic acids; and (iv) at least 42 test unenriched nucleic acid libraries each comprising nucleic acids of human origin and nucleic acids of unknown microbial origin prepared from the at least 42 samples, where each unenriched nucleic acid library comprises fragmented double-stranded DNA that is barcoded for indexing.

[0096] In yet another embodiment, step (b) comprises: partitioning the 48 total unenriched libraries, on an equimolar basis, into a first set of two pools of unenriched libraries and a second set of two pools of unenriched libraries, where each pool in the first and second sets of four pools of unenriched libraries comprises an aliquot taken from: (i) a different one of the at least two first control nucleic acid libraries; (ii) a different one of the at least two second control libraries; (iii) a different one of the at least two third control libraries; and (iv) at least 21 different test libraries from the at least 42 test libraries.

[0097] In yet another embodiment, where step (a) comprises: preparing at least 96 total unenriched libraries for metagenomic next generation sequencing from total nucleic acids extracted from at least 84 samples obtained from at least 84 subjects suspected of having a microbial infection, the at least 96 total unenriched libraries comprise: (i) at least four first control unenriched nucleic acid libraries prepared from total nucleic acids extracted from nucleic acids of known microbial origin; (ii) at least four second control unenriched nucleic acid libraries comprising nucleic acids of human origin obtained from at least one human subject who is not suspected of having the microbial infection; (iii) at least four third control unenriched libraries that are substantially free of nucleic acids; and (iv) at least 84 test unenriched nucleic acid libraries each comprising nucleic acids of human origin and nucleic acids of unknown microbial origin prepared from the at least 84 samples, where each unenriched nucleic acid library comprises fragmented double-stranded DNA that is barcoded for indexing.

[0098] In such an embodiment, step (b) comprises partitioning the 96 total unenriched libraries, on an equimolar basis, into a first set of four pools of unenriched libraries and a second set of four pools of unenriched libraries, where each pool in the first and second sets of four pools of unenriched libraries comprises an aliquot taken from: (i) a different one of the at least four first control nucleicacid libraries; (ii) a different one of the at least four second control libraries; (iii) a different one of the at least four third control libraries; and (iv) at least 21 different test libraries from the at least 84 test libraries. c. Step (c) enriching target nucleic acids of microbial origin

[0099] Aspects of the disclosure involve target enrichment of nucleic acids of microbial origin to produce enriched libraries that are enriched for the nucleic acids of microbial origin. The methods of the disclosure can be used to enrich target nucleic acids from any microbe, for example, to enrich for archaeal nucleic acids, bacterial nucleic acids, fungal nucleic acids (e.g., yeast), protozoan nucleic acids, and / or viral nucleic acids.

[0100] In some embodiments, step (c) comprises enriching each of the unenriched libraries in the first set of pools for nucleic acids of microbial origin to produce a first set of pools of enriched libraries that are enriched for nucleic acids of microbial origin. In yet other embodiments, step (c) comprises enriching each of the unenriched libraries in the first set of pools for nucleic acids of archaeal origin to produce a first set of pools of enriched libraries that are enriched for nucleic acids of archaeal origin. In still yet other embodiments, step (c) comprises enriching each of the unenriched libraries in the first set of pools for nucleic acids of bacterial origin to produce a first set of pools of enriched libraries that are enriched for nucleic acids of bacterial origin. In still yet other embodiments, step (c) comprises enriching each of the unenriched libraries in the first set of pools for nucleic acids of fungal origin to produce a first set of pools of enriched libraries that are enriched for nucleic acids of fungal origin (e.g., yeast). In still further embodiments, step (c) comprises enriching each of the unenriched libraries in the first set of pools for nucleic acids of protozoan origin to produce a first set of pools of enriched libraries that are enriched for nucleic acids of protozoan origin. In still further embodiments, step (c) comprises enriching each of the unenriched libraries in the first set of pools for nucleic acids of viral origin to produce a first set of pools of enriched libraries that are enriched for nucleic acids of viral origin.

[0101] The target enrichment step can include enriching nucleic acids of unknown microbial origin from at least one microbe of interest in the libraries in each pool of the first set of pools. In some embodiments, the target enrichment step comprises enriching nucleic acids of unknown microbial origin from at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 microbes of interest.

[0102] In some embodiments, the target enrichment step comprises enriching nucleic acids of unknown archaeal origin from at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 archaea of interest.[(1103] In some embodiments, the target enrichment step comprises enriching nucleic acids of unknown bacterial origin from at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 bacteria of interest.

[0104] In some embodiments, the target enrichment step comprises enriching nucleic acids of unknown fungal (e.g., yeast) origin from at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 microbes of interest.

[0105] In some embodiments, the target enrichment step comprises enriching nucleic acids of unknown protozoa origin from at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 protozoa of interest.

[0106] In some embodiments, the target enrichment step comprises enriching nucleic acids of unknown viral origin from at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 viruses of interest.[(1107] In some embodiments, the target enrichment step comprises enriching a combination of nucleic acids of unknown bacterial and viral origin from at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 bacteria and viruses of interest. In other similar embodiments, the target enrichment step comprises enriching nucleic acids of unknown microbial origin from any combination of at least two, at least three, at least four, or all microbes selected from the group consisting of archaea, bacteria, fungi (e.g., yeast), protozoa, and viruses.

[0108] The target enrichment step (c) can be achieved by contacting each pool in the first set of pools with a plurality of oligonucleotide probes, for example as described in section 3c herein, that are designed to and are capable of hybridizing with nucleic acids of the at least one microbe of interest. In some embodiments, the target enrichment step (c) comprises contacting each pool in the first set of pools with at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 250,000, 300,000, 350,000, 400,000, 450,000, 500,000, 550,000, 600,000, 650,000, 700,000, 750,000, 800,000, 850,000, 900,000, 950,000, 1,000,000 or more than 1,000,000 probes designed to and capable of hybridizing with nucleic acids of at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 microbes of interest (e.g., archaea, bacteria, fungi (e.g., yeast), protozoa, and / or viruses).

[0109] In some embodiments, the target enrichment step (c) comprises contacting each pool in the first set of pools with at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 250,000, 300,000, 350,000, 400,000, 450,000, 500,000, 550,000, 600,000, 650,000, 700,000, 750,000, 800,000, 850,000, 900,000, 950,000, 1,000,000, or more than 1,000,000 probes designed to and capable of hybridizing with nucleic acids of at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 bacteria of interest. In some embodiments, the target enrichment step (c) comprises contacting each pool in the first set of pools with at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 250,000, 300,000, 350,000, 400,000, 450,000, 500,000, 550,000, 600,000, 650,000, 700,000, 750,000, 800,000, 850,000, 900,000, 950,000, 1,000,000, or more than 1,000,000 probes designed to hybridize to and enrich nucleic acids of at least one, at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000, at least 3500, at least 4000, at least 5000, at least 6000, at least 7000, at least 7500, at least 8000, at least 8500, at least 9000, at least 10000, at least 11,000, at least 11,500, at least 12,000, at least 12,500, at least 13,000, at least 13,500, at least 14,000, at least 14,500, or at least 15,000 different viral strains of interest.

[0110] Target enrichment step (c) can comprise resuspending the dried pellets comprising the DNA in a solution comprising blocking oligonucleotides to prevent nonspecific binding of capture probes and reduce off-target capture during library enrichment. The solution can be contacted with a plurality of probes described herein that are complementary to nucleic acids of the least one microbe of interest and allowing the probes to hybridize with the nucleic acids of the at least one microbe of interest. Hybridization of the probes to the nucleic acid may be done via methods known in the art. For example, the nucleic acid may first be denatured such that it is single stranded and then the panel of probes and nucleic acid may be incubated at elevated temperature for about 1 to about 72 hours. More specifically, the nucleic acid may be denatured at >95° C. for about 20 minutes and then thepanel of probes and nucleic acid may be incubated at about 47° C. for about 64 hours to about 72 hours.

[0111] The disclosure contemplates the use of any blocking oligonucleotides available to the skilled artisan. In some embodiments, the blocking oligonucleotides comprise human Cot-1 DNA and / or universal blocking oligonucleotides.

[0112] The probes that specifically hybridize to microbial (e.g., viral) nucleic acid sequences within the sample can be isolated. Any methods for isolating probes known in the art can be used. In one embodiment, bead purification may be used to isolate the probes that specifically hybridize to viral nucleic acid sequences within the sample. For example, streptavidin beads may be used. In some embodiments, target enrichment step (c) can also comprise contacting the solution with magnetic streptavidin beads. In yet further embodiments, target enrichment step (c) can also comprise separating hybridized sequences from non-hybridizes sequences by affinity interaction on the streptavidin beads. The streptavidin beads may be incubated with the hybridized sample at about 47° C. for about 45 minutes.

[0113] In some embodiments, target enrichment step (c) can also comprise removing nucleic acids other than nucleic acids from the unknown microbial origin. The disclosure contemplates any method of removing the undesired nucleic acids from the libraries known in the art. In some embodiments, removing the undesired nucleic acids comprises performing one or more stringent washes. The sample(s) may then be washed to remove unbound beads. In some embodiments, a solid support is washed one or more times with buffer, preferably about 2 and 5 times to remove unbound polynucleotides before an elution buffer is added to release the enriched, adapter-tagged nucleic acid fragments from the solid support.

[0114] Alternative variables such as incubation times, temperatures, reaction volumes / concentrations, number of washes, or other variables consistent with the specification can also be employed in the method.

[0115] The isolated viral nucleic acid sequences may then be amplified. In some embodiments, target enrichment step (c) comprises amplifying the hybridized nucleic acid sequences to enrich for the at least one microbe of interest. In general, amplification is carried out using polymerase chain reaction (PCR). A PCR reaction may comprise isolated viral nucleic acid, primers, polymerase, water, buffer, and deoxynucleotide triphosphates (dNTPs). PCR may be performed according to standard methods in the art. By way of non-limiting example, the PCR reaction may comprise denaturation, followed by about 10-20 cycles of denaturation, annealing and extension, followed by a final extension. In one embodiment, the PCR reaction comprises denaturation at about 98° C. for about 30 seconds, followed by about 10 to about 20 cycles of (about 98° C. for about 10 seconds, about 60-72° C. for about 30 seconds, about 72° C. for about 30 seconds), followed by a final extension at about 72° C. for about 5 minutes. Optionally, the amplified viral nucleic acid is then purified, for example, via column purification.

[0116] In some embodiments, target enrichment step (c) comprises repurifying the amplified nucleic acids using magnetic PCR beads.

[0117] Subsequent to target enrichment step (c) and prior to combining step (d), the methods comprise diluting each pool of enriched nucleic acid libraries. Each pool of enriched nucleic acid libraries can be diluted to between about 0.20 nM to about 20 nm, about 0.30 nM to about 19 nM, about 0.4 nM to about 18 nM, about 0.5 nm to about 17 nM, about 0.6 nm to about 16 nm, about 0.7 to about 15 nM, about 0.8 to about 14 nM, about 0.9 to about 13 nM, about 1 nM to about 12 nM, about 1.1 nM to about 11 nM, about 1.2 nM to about 10 nM, about 1.3 nM to about 9 nM, about 1.4 nM to about 8 nM, about 1.5 nM to about 7 nM, about 1.6 nM to about 6 nM, about 1.7 nM to about 5 nM, about 1.8 nM to about 4 nM, about 1.9 nM to about 3 nM, or about 2 nM. In one embodiment, each pool of enriched nucleic acid library is diluted to about 2 nM.

[0118] Each aliquot taken from each unenriched library in the second set of pools can also be diluted. Each such aliquot can be diluted to between about 0.20 nM to about 20 nm, about 0.30 nM to about 19 nM, about 0.4 nM to about 18 nM, about 0.5 nm to about 17 nM, about 0.6 nm to about 16 nm, about 0.7 to about 15 nM, about 0.8 to about 14 nM, about 0.9 to about 13 nM, about 1 nM to about 12 nM, about 1.1 nM to about 11 nM, about 1.2 nM to about 10 nM, about 1.3 nM to about 9 nM, about 1.4 nM to about 8 nM, about 1 .5 nM to about 7 nM, about 1 .6 nM to about 6 nM, about 1.7 nM to about 5 nM, about 1.8 nM to about 4 nM, about 1.9 nM to about 3 nM, or about 2 nM. In one embodiment, each pool of unenriched nucleic acid library in the second set of pools is diluted to about 2 nM.

[0119] In some embodiments, the methods further comprise quantifying each of the pools of unenriched libraries. d. Step (d) combining pools of enriched and unenriched libraries for NGS

[0120] Aspects of the disclosure involve combining pools of enriched and unenriched libraries into a single flow cell for sequencing, specifically, next generation sequencing. In one embodiment, step (d) comprises combining a volume from each pool in the first set of pools of enriched libraries, in a predetermined volume ratio, with a volume from each corresponding pool in the second set of pools of unenriched libraries into a single flow cell for next generation sequencing.

[0121] The predetermined volume ratio can be expressed in any units of volume, for example, as a microliter volume ratio where the ratio of a volume (X microliters) from each pool in the first set of pools (Sample A) to a volume (Y microliters) from each pool in the second set of pools (Sample B) can be expressed in the following equation:

[0122] Examples of predetermined microliter volume ratios include, without limitation, 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 2:8, 3:7, 4:6, 9:1, 8:1, 7:1, 6:1, 5:1, 4:1, 3:1, 2:1, 8:2, 7:3, or 6:4. In some embodiments, the predetermined microliter volume ratio is 1 :9. In other embodiments, thepredetermined microliter volume ratio is 2:8. For example, a predetermined microliter volume ratio of 1 : 1 means that the volume from each pool in the first set is equal to the volume from each corresponding pool in the second set, a predetermined microliter volume ratio of 1:9 means that for every 1 microliter from each pool in the first set, there are 9 microliters from each corresponding pool in the second set.

[0123] In some embodiments, between about 0.10 microliters and about 10 microliters from each pool in the first set of pools is combined with between about 0.10 microliters and about 10 microliters from each corresponding pool in the second set of pools. In some embodiments, about 1 microliter or about 2 microliters from each pool in the first set of pools is combined with about 8 or about 9 microliters from each corresponding pool in the second set of pools. In some embodiments, about 1 microliter from each pool in the first set of pools is combined with about 9 microliters from each corresponding pool in the second set of pools. In some embodiments, about 2 microliters from each pool in the first set of pools is combined with about 8 microliters from each corresponding pool in the second set of pools.

[0124] The predetermined volume ratio can also be expressed as volume to volume %v / v of a volume from each pool in the first set to a volume from each corresponding pool in the second set. Examples of predetermined microliter volume ratios expressed in this way include, without limitation, 5% / 95%, 10% / 90%, 15% / 85%, 20% / 80%, 25% / 75%, 30% / 70%, 35% / 65%, 40% / 60%, 45% / 55%, 50% / 50%, 55% / 45%, 60% / 40%, 65% / 35%, 70% / 30%, 75% / 25%, 80% / 20%, 85% / 15%, 90% / 10%, or 95% / 5%. In some embodiments, the predetermined volume ratio is 20% / 80%. In some embodiments, the predetermined volume ratio is 10% / 90%.

[0125] Libraries in combined volumes from each corresponding pool can be sequenced on a single flow cell. e. Step (e) sequencing in a single flow cell

[0126] Aspects of the disclosure involve sequencing combined enriched and unenriched libraries. In one embodiment, step (e) comprises sequencing libraries in the single flow cell using a next generation sequencing platform. Downstream applications of nucleic acid libraries may include next generation sequencing. Sequencing can be performed on a massively parallel sequencing platform, many of which are commercially available including, but not limited to Illumina, Roche / 454, Ion Torrent, Oxford Nanopore Technologies and PacBio. In one embodiment, Illumina sequencing is used. Target enrichment of nucleic acid sequences with a controlled stoichiometry plurality of probes can result in more efficient sequencing. The performance of a nucleic acid libraries for capturing or hybridizing to targets may be defined by a number of different metrics describing efficiency, accuracy, and precision. For example, Picard metrics comprise variables such as HS library size (the number of unique molecules in the library that correspond to target regions, calculated from read pairs), mean target coverage (the percentage of bases reaching a specific coverage level), depth ofcoverage (number of reads including a given nucleotide) fold enrichment (sequence reads mapping uniquely to the target / reads mapping to the total sample, multiplied by the total sample length / target length), percent off-bait bases (percent of bases not corresponding to bases of the probes / baits), percent off-target (percent of bases not corresponding to bases of interest), usable bases on target, AT or GC dropout rate, fold 80 base penalty (fold over-coverage needed to raise 80 percent of non-zero targets to the mean coverage level), percent zero coverage targets, PF reads (the number of reads passing a quality filter), percent selected bases (the sum of on-bait bases and near-bait bases divided by the total aligned bases), percent duplication, or other variable apparent to the skilled artisan.

[0127] Read depth (sequencing depth, or sampling) represents the total number of times a sequenced nucleic acid fragment (a “read”) is obtained for a sequence. Theoretical read depth is defined as the expected number of times the same nucleotide is read, assuming reads are perfectly distributed throughout an idealized genome. Read depth is expressed as a function of % coverage (or coverage breadth). For example, 10 million reads of a 1 million base genome, perfectly distributed, theoretically results in 10x read depth of 100% of the sequences. In practice, a greater number of reads (higher theoretical read depth, or oversampling) may be needed to obtain the desired read depth for a percentage of the target sequences. Enrichment of target sequences with a controlled stoichiometry probe library increases the efficiency of downstream sequencing, as fewer total reads will be required to obtain an outcome with an acceptable number of reads over a desired % of target sequences. For example, in some embodiments 55x theoretical read depth of target sequences results in at least 30* coverage of at least 90% of the sequences. In some embodiments, no more than 55x theoretical read depth of target sequences results in at least 30* read depth of at least 80% of the sequences. In some embodiments no more than 55 x theoretical read depth of target sequences results in at least 30x read depth of at least 95% of the sequences. In some embodiments no more than 55x theoretical read depth of target sequences results in at least 1 Ox read depth of at least 98% of the sequences. In some embodiments, 55x theoretical read depth of target sequences results in at least 20x read depth of at least 98% of the sequences. In some embodiments, no more than 55x theoretical read depth of target sequences results in at least 5x read depth of at least 98% of the sequences. Increasing the concentration of probes during hybridization with targets can lead to an increase in read depth. In some embodiments, the concentration of probes is increased by at least 1.5x, 2. Ox, 2.5x, 3x, 3.5x, 4x, 5x, or more than 5x. In some embodiments, increasing the probe concentration results in at least a 1000% increase, or a 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 750%, 1000%, or more than a 1000% increase in read depth. In some embodiments, increasing the probe concentration by 3x results in a 1000% increase in read depth. In some embodiments, sequencing is performed to achieve a theoretical read depth of at least 30x, 50x, 100x, 150x, 200x, 250x, 300x, 500x, at least 1000x, at least 1000Ox, or at least 100,000x. In some embodiments, sequencing is performed to achieve a theoretical read depth of about 30x, 50x, 100x, 150x, 200x, 250x, 300x, 500x, or about 1000x, about 1000Ox, or about 100,000x. In someembodiments, sequencing is performed to achieve a theoretical read depth of no more than 30x, 50*, 100x, 150x, 200x, 250x, 300x, 500x, or no more than 1000x. In some embodiments, sequencing is performed to achieve an actual read depth of at least 30x, 50x, 100x, 150x, 200x, 250x, 300x, 500x, at least 1000x, at least 1000Ox, or at least 100,000x. In some embodiments, sequencing is performed to achieve an actual read depth of no more than 30x, 50x, 100x, 150x, 200x, 250x, 300x, 500x, or no more than 1000x. In some embodiments, sequencing is performed to achieve an actual read depth of about 30x, 50x, 100x, 150x, 200x, 250x, 300x, 500x, about 1000x, about 1000Ox, or about 100,000x.

[0128] On-target rate represents the percentage of sequencing reads that correspond with the desired target sequences. In some embodiments, a controlled stoichiometry polynucleotide probe library results in an on-target rate of at least 30%, or at least 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, or at least 90%. Increasing the concentration of plurality of probes during contact with target nucleic acids leads to an increase in the on-target rate. In some embodiments, the concentration of probes is increased by at least 1.5x, 2.0x, 2.5x, 3x, 3.5x, 4x, 5x, or more than 5x. In some embodiments, increasing the probe concentration results in at least a 20% increase, or a 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, or at least a 500% increase in on- target binding. In some embodiments, increasing the probe concentration by 3x results in a 20% increase in on-target rate.

[0129] Coverage uniformity is in some cases calculated as the read depth as a function of the target sequence identity. Higher coverage uniformity results in a lower number of sequencing reads needed to obtain the desired read depth. For example, a property of the target sequence may affect the read depth, for example, high or low GC or AT content, repeating sequences, trailing adenines, secondary structure, affinity for target sequence binding (for amplification, enrichment, or detection), stability, melting temperature, biological activity, ability to assemble into larger fragments, sequences containing modified nucleotides or nucleotide analogues, or any other property of polynucleotides. Enrichment of target sequences with controlled stoichiometry polynucleotide probe libraries results in higher coverage uniformity after sequencing. In some embodiments, 95% of the sequences have a read depth that is within lx ofthe mean library read depth, or about 0.05, 0.1, 0.2, 0.5, 0.7, 1, 1.2, 1.5, 1.7 or about within 2x the mean library read depth. In some embodiments, 80%, 85%, 90%, 95%, 97%, or 99% of the sequences have a read depth that is within lx of the mean.

[0130] In any of the embodiments, the detection or quantification analysis of the oligonucleotides can be accomplished by sequencing. The subunits or entire synthesized oligonucleotides can be detected via full sequencing of all oligonucleotides by any suitable methods known in the art, e.g., Illumina sequencing by synthesis, PacBio nanopore sequencing, or BGI / MGI nanoball sequencing, including the sequencing methods described herein. Sequencing can be accomplished through classic Sanger sequencing methods which are well known in the art. Sequencing can also be accomplished using high-throughput systems some of which allow detection of a sequenced nucleotide immediatelyafter or upon its incorporation into a growing strand, i.e., detection of sequence in red time or substantially real time. In some cases, high throughput sequencing generates at least 1,000, at least 5,000, at least 10,000, at least 20,000, at least 30,000, at least 40,000, at least 50,000, at least 100,000 or at least 500,000 sequence reads per hour, with each read being at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120 or at least 150 bases per read.

[0131] In some embodiments, high-throughput sequencing involves the use of technology available by Illumina's Genome Analyzer IIX, MiSeq personal sequencer, or HiSeq systems, such as those using HiSeq 2500, HiSeq 1500, HiSeq 2000, HiSeq 1000, iSeq 100, Mini Seq, MiSeq, NextSeq 550, NextSeq 2000, NextSeq 550, or NovaSeq 6000. These machines use reversible terminator-based sequencing by synthesis chemistry. These machines can generate 6000 Gb or more reads in 13-44 hours. Smaller systems may be utilized for runs within 3, 2, 1 days or less time. Short synthesis cycles may be used to minimize the time it takes to obtain sequencing results.

[0132] In some embodiments, high-throughput sequencing involves the use of technology available by ABI Solid System. This genetic analysis platform enables massively parallel sequencing of clonally-amplified DNA fragments linked to beads. The sequencing methodology is based on sequential ligation with dye-labeled oligonucleotides.

[0133] The next generation sequencing can comprise ion semiconductor sequencing (e.g., using technology from Life Technologies (Ion Torrent)). Ion semiconductor sequencing can take advantage of the fact that when a nucleotide is incorporated into a strand of DNA, an ion can be released. To perform ion semiconductor sequencing, a high-density array of micromachined wells can be formed. Each well can hold a single DNA template. Beneath the well can be an ion sensitive layer, and beneath the ion sensitive layer can be an ion sensor. When a nucleotide is added to a DNA, H+ can be released, which can be measured as a change in pH. The H+ ion can be converted to voltage and recorded by the semiconductor sensor. An array chip can be sequentially flooded with one nucleotide after another. No scanning, light, or cameras can be required. In some cases, an IONPROTON™ Sequencer is used to sequence nucleic acid. In some cases, an IONPGM™ Sequencer is used. The Ion Torrent Personal Genome Machine (PGM) can do 10 million reads in two hours.

[0134] In some embodiments, high-throughput sequencing involves the use of technology available by Helicos BioSciences Corporation (Cambridge, Mass.) such as the Single Molecule Sequencing by Synthesis (SMSS) method. SMSS is unique because it allows for sequencing the entire human genome in up to 24 hours. Finally, SMSS is powerful because, like the 1 \4 W technology, it does not require a pre amplification step prior to hybridization. In fact, SMSS does not require any amplification.

[0135] In some embodiments, high-throughput sequencing involves the use of technology available by 454 Lifesciences, Inc. (Branford, Conn.) such as the Pico Titer Plate device which includes a fiber optic plate that transmits chemiluminescent signal generated by the sequencingreaction to be recorded by a CCD camera in the instrument. This use of fiber optics allows for the detection of a minimum of 20 million base pairs in 4.5 hours.

[0136] Methods for using bead amplification followed by fiber optics detection are described in Marguiles, M., et al. “Genome sequencing in microfabricated high-density picolitre reactors”, Nature, doi: 10.1038 / nature03959.

[0137] In some embodiments, high-throughput sequencing is performed using Clonal Single Molecule Array (Solexa, Inc.) or sequencing-by-synthesis (SBS) utilizing reversible terminator chemistry. Constans, A., The Scientist, 2003, 17(13):36. High-throughput sequencing of oligonucleotides can be achieved using any suitable sequencing method known in the art, such as those commercialized by Pacific Biosciences, Complete Genomics, Genia Technologies, Halcyon Molecular, Oxford Nanopore Technologies and the like. Overall, such systems involve sequencing a target oligonucleotide molecule having a plurality of bases by the temporal addition of bases via a polymerization reaction that is measured on a molecule of oligonucleotide, i.e., the activity of a nucleic acid polymerizing enzyme on the template oligonucleotide molecule to be sequenced is followed in real time. Sequence can then be deduced by identifying which base is being incorporated into the growing complementary strand of the target oligonucleotide by the catalytic activity of the nucleic acid polymerizing enzyme at each step in the sequence of base additions. A polymerase on the target oligonucleotide molecule complex is provided in a position suitable to move along the target oligonucleotide molecule and extend the oligonucleotide primer at an active site. A plurality of labeled types of nucleotide analogs are provided proximate to the active site, with each distinguishably type of nucleotide analog being complementary to a different nucleotide in the target oligonucleotide sequence. The growing oligonucleotide strand is extended by using the polymerase to add a nucleotide analog to the oligonucleotide strand at the active site, where the nucleotide analog being added is complementary to the nucleotide of the target oligonucleotide at the active site. The nucleotide analog added to the oligonucleotide primer as a result of the polymerizing step is identified. The steps of providing labeled nucleotide analogs, polymerizing the growing oligonucleotide strand, and identifying the added nucleotide analog are repeated so that the oligonucleotide strand is further extended and the sequence of the target oligonucleotide is determined.

[0138] The next generation sequencing technique can comprise real-time (SMRT™) technology by Pacific Biosciences. In SMRT, each of four DNA bases can be attached to one of four different fluorescent dyes. In some aspects, these dyes can be phospho linked. A single DNA polymerase can be immobilized with a single molecule of template single stranded DNA at the bottom of a zero-mode waveguide (ZMW). A ZMW can be a confinement structure which enables observation of incorporation of a single nucleotide by DNA polymerase against the background of fluorescent nucleotides that can rapidly diffuse in an out of the ZMW (in microseconds). It can take several milliseconds to incorporate a nucleotide into a growing strand. During this time, the fluorescent labelcan be excited and produce a fluorescent signal, and the fluorescent tag can be cleaved off. The ZMW can be illuminated from below. Attenuated light from an excitation beam can penetrate the lower 20- 30 nm of each ZMW. A microscope with a detection limit of 20 zepto liters (10" liters) can be created. The tiny detection volume can provide 1000-fold improvement in the reduction of background noise. Detection of the corresponding fluorescence of the dye can indicate which base was incorporated. The process can be repeated.

[0139] In some cases, the next generation sequencing is nanopore sequencing (See e.g., Soni G V and Meller A. (2007) Clin Chem., 53: 1996-2001). A nanopore can be a small hole, of the order of about one nanometer in diameter. Immersion of a nanopore in a conducting fluid and application of a potential across it can result in a slight electrical current due to conduction of ions through the nanopore. The amount of current which flows can be sensitive to the size of the nanopore. As a DNA molecule passes through a nanopore, each nucleotide on the DNA molecule can obstruct the nanopore to a different degree. Thus, the change in the current passing through the nanopore as the DNA molecule passes through the nanopore can represent a reading of the DNA sequence. The nanopore sequencing technology can be from Oxford Nanopore Technologies, e.g., a GridlON system. A single nanopore can be inserted in a polymer membrane across the top of a microwell. Each microwell can have an electrode for individual sensing. The microwells can be fabricated into an array chip, with 100,000 or more microwells (e.g., more than 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, or 1,000,000) per chip. An instrument (or node) can be used to analyze the chip. Data can be analyzed in real-time. One or more instruments can be operated at a time. The nanopore can be a protein nanopore, e.g., the protein alpha-hemolysin, a heptameric protein pore. The nanopore can be a solid-state nanopore made, e.g., a nanometer sized hole formed in a synthetic membrane (e.g., SiNx, or SiO2). The nanopore can be a hybrid pore (e.g., an integration of a protein pore into a solid-state membrane). The nanopore can be a nanopore with an integrated sensors (e.g., tunneling electrode detectors, capacitive detectors, or graphene based nano-gap or edge state detectors (see e.g., Garaj et al. (2010) Nature, vol. 67, doi: 10.1038 / nature09379)). A nanopore can be functionalized for analyzing a specific type of molecule (e.g., DNA, RNA, or protein). Nanopore sequencing can comprise “strand sequencing” in which intact DNA polymers can be passed through a protein nanopore with sequencing in real time as the DNA translocates the pore. An enzyme can separate strands of a double stranded DNA and feed a strand through a nanopore. The DNA can have a hairpin at one end, and the system can read both strands. In some cases, nanopore sequencing is “exonuclease sequencing” in which individual nucleotides can be cleaved from a DNA strand by a processive exonuclease, and the nucleotides can be passed through a protein nanopore. The nucleotides can transiently bind to a molecule in the pore (e.g., cyclodextran). A characteristic disruption in current can be used to identify bases.

[0140] Nanopore sequencing technology from GENIA can be used. An engineered protein pore can be embedded in a lipid bilayer membrane. “Active Control” technology can be used to enableefficient nanopore-membrane assembly and control of DNA movement through the channel. In some cases, the nanopore sequencing technology is from NABsys. Genomic DNA can be fragmented into strands of average length of about 100 kb. The 100 kb fragments can be made single stranded and subsequently hybridized with a 6-mer probe. The genomic fragments with probes can be driven through a nanopore, which can create a current-versus-time tracing. The current tracing can provide the positions of the probes on each genomic fragment. The genomic fragments can be lined up to create a probe map for the genome. The process can be done in parallel for a library of probes. A genome-length probe map for each probe can be generated. Errors can be fixed with a process termed “moving window Sequencing By Hybridization (mwSBH).” In some cases, the nanopore sequencing technology is from IBM / Roche. An electron beam can be used to make a nanopore sized opening in a microchip. An electrical field can be used to pull or thread DNA through the nanopore. A DNA transistor device in the nanopore can comprise alternating nanometer sized layers of metal and dielectric. Discrete charges in the DNA backbone can get trapped by electrical fields inside the DNA nanopore. Turning off and on gate voltages can allow the DNA sequence to be read.

[0141] The next generation sequencing can comprise DNA nanoball sequencing (as performed, e.g., by Complete Genomics; see e.g., Drmanac et al. (2010) Science, 327: 78-81). DNA can be isolated, fragmented, and size selected. For example, DNA can be fragmented (e.g., by sonication) to a mean length of about 500 bp. Adaptors (Adi) can be attached to the ends of the fragments. The adaptors can be used to hybridize to anchors for sequencing reactions. DNA with adaptors bound to each end can be PCR amplified. The adaptor sequences can be modified so that complementary single strand ends bind to each other forming circular DNA. The DNA can be methylated to protect it from cleavage by a type IIS restriction enzyme used in a subsequent step. An adaptor (e.g., the right adaptor) can have a restriction recognition site, and the restriction recognition site can remain non- methylated. The non-methylated restriction recognition site in the adaptor can be recognized by a restriction enzyme (e.g., Acul), and the DNA can be cleaved by Acul 13 bp to the right of the right adaptor to form linear double stranded DNA. A second round of right and left adaptors (Ad2) can be ligated onto either end of the linear DNA, and all DNA with both adapters bound can be PCR amplified (e.g., by PCR). Ad2 sequences can be modified to allow them to bind each other and form circular DNA. The DNA can be methylated, but a restriction enzyme recognition site can remain non- methylated on the left Adi adapter. A restriction enzyme (e.g., Acul) can be applied, and the DNA can be cleaved 13 bp to the left of the Adi to form a linear DNA fragment. A third round of right and left adaptor (Ad3) can be ligated to the right and left flank of the linear DNA, and the resulting fragment can be PCR amplified. The adaptors can be modified so that they can bind to each other and form circular DNA. A type III restriction enzyme (e.g., EcoP15) can be added; EcoP15 can cleave the DNA 26 bp to the left of Ad3 and 26 bp to the right of Ad2. This cleavage can remove a large segment of DNA and linearize the DNA once again. A fourth round of right and left adaptors (Ad4) can beligated to the DNA, the DNA can be amplified (e.g., by PCR), and modified so that they bind to each other and form the completed circular DNA template.

[0142] Rolling circle replication (e.g., using Phi 29 DNA polymerase) can be used to amplify small fragments of DNA. The four adaptor sequences can contain palindromic sequences that can hybridize, and a single strand can fold onto itself to form a DNA nanoball (DNB™) which can be approximately 200-300 nanometers in diameter on average. A DNA nanoball can be attached (e.g., by adsorption) to a microarray (sequencing flowcell). The flow cell can be a silicon wafer coated with silicon dioxide, titanium and hexamethyldisilazane (HMDS) and a photoresist material. Sequencing can be performed by unchained sequencing by ligating fluorescent probes to the DNA. The color of the fluorescence of an interrogated position can be visualized by a high-resolution camera. The identity of nucleotide sequences between adaptor sequences can be determined.

[0143] In some embodiments, the single flow cell comprises a Pl flow cell. (NextSeq Pl flow cell). The Pl flow cell is capable of performing about 4MM paired end reads per index. In other embodiments, the single flow cell comprises a MiSeq v2 flow cell. In still other embodiments, the single flow cell comprises a P2 flow cell (e.g., NextSeq P2 flow cell). The P2 flow cell is capable of performing about 8 MM paired end reads per index. f. Step (f) detecting microbes and / or characterizing metagenomes

[0144] Aspects of the disclosure involve detecting microbes in one or more samples and / or characterizing metagenomes of one or more samples. In one embodiment, step (f) characterizing the metagenomes of the at least one sample obtained from the at least one subject suspected of having the microbial infection based on the next generation sequencing results.

[0145] After sequencing of the microbial nucleic acids (e.g., archaeal, bacterial, fungal, protozoan, and / or viral nucleic acids), the sequences are compared with a database comprising reference microbial nucleic acids, e.g., viral nucleic acid sequences to determine the identity of the viral nucleic acid in the sample. Comparison of sequences generally involves aligning the experimentally determined sequence with a reference sequence. Methods of aligning sequences are known in the art. In a specific embodiment, the alignment algorithm utilized may be BWA-MEM. BWA-MEM is an alignment algorithm for aligning sequence reads or long query sequences against a large reference genome. It automatically chooses between local and end-to-end alignments, supports paired end reads and performs chimeric alignment. The algorithm is robust to sequencing errors and applicable to a wide range of sequence lengths from 70 bp to a few megabases. For mapping 100 bp sequences, BWA-MEM shows better performance than several state-of-art read aligners to date. The sequence alignments may then be evaluated to determine the identity of the viral nucleic acid in the sample. Methods of evaluating sequence alignments are known in the art. In a specific embodiment, SAMtools is utilized to evaluate the sequence alignments. SAMtools is a set of utilities for interactingwith and post-processing short DNA sequence read alignments in the SAM (=Sequence Alignment / Map), BAM (=Binary Alignment / Map) and CRAM formats. Both simple and advanced tools are provided, supporting complex tasks like variant calling and alignment viewing as well as sorting, indexing, data extraction and format conversion. g. Exemplary embodiments

[0146] In a first exemplary embodiment, a method of characterizing metagenomes of samples obtained from subjects having or suspected of having a viral infection is provided, the method comprising: (a) preparing at least 48 total unenriched libraries for metagenomic next generation sequencing from total nucleic acids extracted from at least 42 samples obtained from at least 42 subjects having or suspected of having a viral infection, wherein the at least 48 total unenriched libraries comprise: (i) at least two first control unenriched nucleic acid libraries comprising nucleic acids of known viral origin; (ii) at least two second control unenriched nucleic acid libraries comprising nucleic acids of human origin obtained from at least one human subject who is not suspected of having the viral infection; (iii) at least two third control unenriched libraries that are substantially free of nucleic acids; and (iv) at least 42 test unenriched nucleic acid libraries each comprising nucleic acids of human origin and nucleic acids of unknown viral origin prepared from the at least 42 samples, wherein each unenriched nucleic acid library comprises fragmented double- stranded DNA that is barcoded for indexing; (b) partitioning the 48 total unenriched libraries, on an equimolar basis, into a first set of two pools of unenriched libraries and a second set of two pools of unenriched libraries, wherein each pool in the first and second sets of pools of unenriched libraries comprises an aliquot taken from: (i) a different one of the at least two first control nucleic acid libraries; (ii) a different one of the at least two second control libraries; (iii) a different one of the at least two third control libraries; and (iv) at least 21 different test libraries from the at least 42 test libraries; (c) enriching nucleic acids of unknown viral origin in each of the unenriched libraries in the first set of pools; (d) combining from about 1 to about 2 microliters from each pool in the first set of pools with from about 8 to about 9 microliter from each corresponding pool in the second set of pools into a single flow cell for next generation sequencing; (e) sequencing the libraries in the combined volumes of each corresponding pool in the single flow cell using a next generation sequencing platform; and (f) characterizing the metagenomes of the at least 42 samples obtained from the at least 42 subjects suspected of having the viral infection based on the sequencing results.

[0017] In a second exemplary embodiment, a method of characterizing metagenomes of samples obtained from subjects having or suspected of having a viral infection is provided, comprising: (a) preparing at least 96 total unenriched libraries for metagenomic next generation sequencing from total nucleic acids extracted from at least 84 samples obtained from at least 84 subjects having or suspected of having a viral infection, wherein the at least 96 total unenriched libraries comprise: (i) at least fourfirst control unenriched nucleic acid libraries prepared from total nucleic acids extracted from nucleic acids of known viral origin; (ii) at least four second control unenriched nucleic acid libraries comprising nucleic acids of human origin obtained from at least one human subject who is not suspected of having the viral infection; (iii) at least four third control unenriched libraries that are substantially free of nucleic acids; and (iv) at least 84 test unenriched nucleic acid libraries each comprising nucleic acids of human origin and nucleic acids of unknown viral origin prepared from the at least 84 samples,

[0148] wherein each unenriched nucleic acid library comprises fragmented double-stranded DNA that is barcoded for indexing; (b) partitioning the 96 total unenriched libraries, on an equimolar basis, into a first set of four pools of unenriched libraries and a second set of four pools of unenriched libraries, wherein each pool in the first and second sets of four pools of unenriched libraries comprises an aliquot taken from: (i) a different one of the at least four first control nucleic acid libraries; (ii) a different one of the at least four second control libraries; (iii) a different one of the at least four third control libraries; and (iv) at least 21 different test libraries from the at least 84 test libraries; (c) enriching nucleic acids of unknown viral origin in each of the unenriched libraries in the first set of pools; (d) combining from about 1 to about 2 microliters from each pool in the first set of pools with from about 8 to about 9 microliter from each corresponding pool in the second set of pools into a single flow cell for next generation sequencing; (e) sequencing the libraries in the combined volumes of each corresponding pool in the single flow cell using a next generation sequencing platform; and (f) characterizing the metagenomes of the at least 84 samples obtained from the at least 84 subjects suspected of having the viral infection based on the sequencing results. In some embodiments, step (d) comprises combining about 1 microliter from each pool in the first set of pools with about 9 microliters from each corresponding pool in the second set in the single flow cell for next generation sequencing. In some embodiments, step (d) comprises combining about 2 microliters from each pool in the first set of pools with about 8 microliters from each corresponding pool in the second set in the single flow cell for next generation sequencing.3. Compositions

[0019] Aspects of the disclosure relate to compositions for use in the methods and kits of the disclosure. The compositions contemplated by the disclosure include samples, controls, primers, blocking oligonucleotides, universal adapters, barcoded primers, universal blockers, hybridization buffers, and probes used in library preparation step (a) and / or target enrichment step (b). a. Samples 0150] Aspects of the disclosure involve detection of one or more microbes of interest in one or more samples. The disclosure is not limited to the type of sample in which one or more microbes canbe detected. In an embodiment, one or more archaea of interest is detected in at least one sample. In an embodiment, one or more bacteria of interest is detected in at least one sample. In an embodiment, one or more fungi (e.g., yeast) of interest is detected in at least one sample. In an embodiment, one or more protozoa of interest is detected in at least one sample. In an embodiment, one or more viruses of interest are detected in at least one sample.

[0151] The sample may be a sample from a subject, the environment, a laboratory, or any sample in which nucleic acid is present. In some embodiments, the at least one sample is a biological sample. In some embodiments, the at least one sample is selected from the group consisting of bone marrow, bronchoalveolar lavage fluid, cerebral spinal fluid, plasma, serum, sputum, urine, whole blood, a nasal sample, a nasopharyngeal sample, stool, a swab, tissue, and other bodily fluids. In an embodiment, the at least one sample is a bone marrow sample. In an embodiment, the at least one sample is a bronchoalveolar lavage sample. In an embodiment, the at least one sample is a cerebral spinal fluid sample. In an embodiment, the at least one sample is a plasma sample. In an embodiment, the at least one sample is a serum sample. In an embodiment, the at least one sample is a sputum sample. In an embodiment, the at least one sample is a urine sample. In an embodiment, the at least one sample is a whole blood sample. In an embodiment, the at least one sample is a bronchoaleveolar lavage fluid sample. In an embodiment, the at least one sample is a stool sample. In an embodiment, the at least one sample is a swab sample (e.g., nasopharyngeal swab or a nasal sample). In an embodiment, the at least one sample is a tissue sample. The tissue sample may be a tissue biopsy. The biopsied tissue may be fixed, embedded in paraffin or plastic, and sectioned, or the biopsied tissue may be frozen and cryosectioned. Alternatively, the biopsied tissue may be processed into individual cells or an explant, or processed into a homogenate, a cell extract, a membranous fraction, or a protein extract.

[0152] The sample may be used “as is”, processed for cell lysis or disruption of microbial particles (e.g., viral particles), or the nucleic acid may be purified from the sample prior to sample preparation. Methods of isolating nucleic acid from a sample are known in the art. In certain embodiments, the isolated nucleic acid is reverse transcribed and amplified after isolation. This allows detection of both RNA and DNA viruses. Specifically, RNA in the total nucleic acid may be reverse transcribed with reverse transcriptase. Random primers may then be tagged with a conserved sequence to be used for subsequent amplification. Second strand synthesis may then be carried out using DNA polymerase to generate cDNA for the RNA viruses.

[0153] The DNA / cDNA mixture may then be amplified using DNA polymerase. In general, amplification is carried out using polymerase chain reaction (PCR). A PCR reaction may comprise nucleic acid, primers, polymerase, water, buffer, and deoxynucleotide triphosphates (dNTPs). PCR may be performed according to standard methods in the art. By way of non-limiting example, the PCR reaction may comprise denaturation, followed by about 5-10 cycles of denaturation, annealing and extension, followed by a final extension. In one embodiment, the PCR reaction comprises denaturation at about 98° C. for about 30 seconds, followed by about 5 to about 10 cycles of (about98° C. for about 10 seconds, about 62-72° C. for about 30 seconds, about 72° C. for about 30 seconds), followed by a final extension at about 72° C. for about 5 minutes. Optionally, the amplified nucleic acid is then purified, for example, via column purification. The nucleic acid in the sample may then be sheared via methods known in the art to generate fragments. The fragments may be about 100 to about 2000 bp. For example, the fragments may be about 200 to about 1500 bp, about 400 to about 1000 bp, about 400 to about 800 bp, or about 500 bp.

[0154] The at least one subject from which the sample is obtained has or is suspected of having a microbial infection and / or a symptom, disease, condition, or disorder associated with the microbial infection. In some embodiments, the at least one sample is obtained from a subject that has or is suspected of having a bacterial infection and / or a symptom, disease, condition, or disorder associated with the bacterial infection. In some embodiments, the at least one sample is obtained from a subject that has or is suspected of having a fungal (e.g., yeast) infection and / or a symptom, disease, condition, or disorder associated with the fungal infection. In some embodiments, the at least one sample is obtained from a subject that has or is suspected of having a protozoa infection and / or a symptom, disease, condition, or disorder associated with the protozoa infection. In some embodiments, the at least one sample is obtained from a subject that has or is suspected of having a viral infection and / or a symptom, disease, condition, or disorder associated with the viral infection.

[0155] The sample or samples can be obtained from subjects having any type of microbial infection (e.g., archaeal, bacterial, fungal, protozoan, and / or viral infection) and / or symptom, disease, condition, or disorder associated with the microbial infection. In some cases, the at least one subject has or is suspected of having acute febrile illness (AFI). In some cases, the at least one subject has or is suspected of having viral hemorrhagic fever). In some cases, the at least one subject has or is suspected of having a severe acute respiratory illness.

[0016] Samples used for controls can be obtained from microbes of known origin to extract total nucleic acids used in the methods. In some embodiments, samples of microbial strains of known origin are used to extract nucleic acids of known microbial origin for use in control unenriched libraries. In some embodiments, samples of bacterial strains of known origin are used to extract nucleic acids of known bacterial origin for use in control unenriched libraries. In some embodiments, samples of fungal strains of known origin are used to extract nucleic acids of known fungal origin for use in control unenriched libraries. In some embodiments, samples of protozoan strains of known origin are used to extract nucleic acids of known protozoan origin for use in control unenriched libraries. In some embodiments, samples of viral strains of known origin are used to extract nucleic acids of known viral origin for use in control unenriched libraries.

[0157] Samples used to prepare control unenriched libraries can also be obtained from subjects to extract total nucleic acids in the disclosed methods. In some embodiments, nucleic acids of human origin are extracted from at least one human subject who does not have or is not believed to have the microbial infection to prepare control unenriched libraries. In some embodiments, nucleic acids ofhuman origin are extracted from at least one human subject who does not have or is not believed to have the bacterial infection to prepare control unenriched libraries. In some embodiments, nucleic acids of human origin are extracted from at least one human subject who does not have or is not believed to have the fungal infection to prepare control unenriched libraries. In some embodiments, nucleic acids of human origin are extracted from at least one human subject who does not have or is not believed to have the protozoan infection to prepare control unenriched libraries. In some embodiments, nucleic acids of human origin are extracted from at least one human subject who does not have or is not believed to have the viral infection to prepare control unenriched libraries. The viral infection can be from an RNA virus or DNA virus.

[0158] Samples that are substantially free of nucleic acids can be used to prepare control libraries. These controls are referred to herein as “no template controls.” In some embodiments, the no-template control comprises phosphate-buffered saline.

[0019] In some embodiments, the sample comprises or is suspected to comprise one or more nucleic acids of unknown microbial origin that is capable of analysis by the methods. Preferably, the samples comprise nucleic acids (e.g., DNA, RNA, cDNAs, microRNA, mitochondrial DNA, etc.). Samples may be complex samples or mixed samples, which contain nucleic acids comprising multiple different nucleic acid sequences (e.g. host and pathogen nucleic acids; mutant and wild-type species; heterogeneous tumor). Samples may comprise nucleic acids from more than one source (e.g. different species, different subspecies, etc.), subject, and / or individual.

[0160] In some aspects, the amount of sample obtained from the subject is less than about 3.9 mL, about 3.8 mL, about 3.7 mL, about 3.6 mL, about 3.5 mL, about 3.4 mL, about 3.3 mL, about 3.2 mL, about 3.1 mL, about 3.0 mL, about 2.9 mL, about 2.8 mL, about 2.7 mL, about 2.6 mL, about 2.5 mL, about 2.4 mL, about 2.3 mL, about 2.2 mL, about 2.1 mL, about 2.0 ml, about 1.9 mL, about 1.8 mL, about 1.7 mL, about 1.6 mL, about 1.5 mL, about 1.4 mL, about 1.3 mL, about 1.2 mL, about 1.1 mL, about 1.0 mL, about 0.9 mL, about 0.8 mL, about 0.7 mL, about 0.6 mL, or about 0.5 mL. b. Controls

[0161] Aspects of the disclosure involve the preparation of control unenriched libraries for use in the methods of the disclosure. The disclosure is not limited to any type of control.

[0162] In an embodiment, the control comprises a first control unenriched nucleic acid library comprising nucleic acids of known microbial origin.

[0163] The first control can be obtained from microbes of known origin to extract total nucleic acids used in the methods. In some embodiments, samples of microbial strains of known origin are used to extract nucleic acids of known microbial origin to produce any desired number of first control unenriched libraries. In some embodiments, samples of bacterial strains of known origin are used to extract nucleic acids of known bacterial origin to produce any desired number of first control unenriched libraries. In some embodiments, samples of fungal strains of known origin are used toextract nucleic acids of known fungal origin to produce any desired number of first control unenriched libraries. In some embodiments, samples of protozoan strains of known origin are used to extract nucleic acids of known protozoan origin to produce any desired number of first control unenriched libraries. In some embodiments, samples of viral strains of known origin are used to extract nucleic acids of known viral origin to produce any desired number of first control unenriched libraries.

[0164] In an embodiment, the control comprises a second control unenriched nucleic acid library comprising nucleic acids of human origin obtained from at least one human subject who does not have or is not believed to have the microbial infection. The second control can be obtained from samples of the human subjects to extract total nucleic acids used in the methods. In some embodiments, nucleic acids of human origin are extracted from at least one human subject who does not have or is not believed to have the microbial infection to produce any desired number of second control unenriched libraries. In some embodiments, nucleic acids of human origin are extracted from at least one human subject who does not have or is not believed to have the bacterial infection to produce any desired number of second control unenriched libraries. In some embodiments, nucleic acids of human origin are extracted from at least one human subject who does not have or is not believed to have the fungal infection to produce any desired number of second control unenriched libraries. In some embodiments, nucleic acids of human origin are extracted from at least one human subject who does not have or is not believed to have the protozoan infection to produce any desired number of second control unenriched libraries. In some embodiments, nucleic acids of human origin are extracted from at least one human subject who does not have or is not believed to have the viral infection to produce any desired number of second control unenriched libraries. The viral infection can be from an RNA virus or DNA virus.

[0165] In an embodiment, the control comprises a third control unenriched library that is substantially free of nucleic acids. Samples that are substantially free of nucleic acids can be used to produce any desired number of third control unenriched libraries. In some embodiments, the third control unenriched library comprises phosphate buffered saline (PBS).

[0166] Each of the first control unenriched nucleic acid libraries can comprise a microbial composition comprising at least one microbe that is diluted. The first control of the disclosure is used as a positive control. The positive control controls for the enrichment step in the control unenriched libraries in each pool in the first set of pools (except for bacteria) and in the control unenriched libraries in each pool in the second set of pools. The positive control is used to help ensure that microbes (e.g., viruses) are detected a particular threshold or sensitivity (e.g., if controls are detected at log 3, then the viral nucleic acids in the sample should also be detected at log 3). In some embodiments, the at least one microbe is at least one archaea. In some embodiments, the at least one microbe is at least one bacteria. In some embodiments, the at least one microbe is at least one fungi. In some embodiments, the at least one microbe is at least one protozoa. In some embodiments, the at least one microbe is at least one virus.

[0167] The at least one microbe can be at least one, at least two, at least three, at least four, or at least five or more different microbes each having different characteristics. In some embodiments, the at least one microbe comprises at least one, at least two, at least three, at least four, or at least five or more archaeal strains each having different characteristics. In some embodiments, the at least one microbe comprises at least one, at least two, at least three, at least four, or at least five or more bacterial strains each having different characteristics. In some embodiments, the at least one microbe comprises at least one, at least two, at least three, at least four, or at least five or more fungal strains each having different characteristics. In some embodiments, the at least one microbe comprises at least one, at least two, at least three, at least four, or at least five or more protozoan strains each having different characteristics. In some embodiments, the at least one microbe comprises at least one, at least two, at least three, at least four, or at least five or more viral strains each having different characteristics.

[0168] As an illustrative example using viral strains, the at least one, at least two, at least three, at least four, or at least five viral strains selected for use as a first control to prepare any desired number of first control unenriched libraries are selected from viruses having different sizes, different types of nucleic acid, viruses with and / or without an envelope, and / or linear and / or segmented viruses.

[0169] In some embodiments, at least one viral strain present in the microbial composition comprises a virus having a size of 5 kb, 10 kb, 25kb, 50kb, 75kb, lOOkb, 125kb, 150kb, 175kb, or 200kb or more. The at least one viral strain in the microbial composition can include a first strain having a size of no more than 5kb,10kb, 25kb, 50kb, 75kb, or no more thanlOOkb and a second strain having a size of no less than 125kb, 150kb, 175kb, or no less than 200kb.

[0170] In some embodiments, the at least one viral strain present in the microbial composition is selected from the group consisting of a +ssRNA virus, a dsRNA virus, a ssDNA virus, a dsDNA virus, a retrovirus, etc, and combinations thereof.

[0171] In some embodiments, the at least one viral strain present in the microbial composition is selected from the group consisting of a virus without an envelope, a virus with an envelope and a combination thereof.

[0172] In some embodiments, the at least one viral strain present in the microbial composition is selected from the group consisting of a linear virus, a segmented virus, and combinations thereof.

[0173] In some embodiments, the at least one viral strain present in the microbial composition is a BSL-2. In still other embodiments, the at least one viral strain present in the microbial composition is not a viral hemorrhagic fever viral strain.

[0174] In another embodiment, the microbial composition comprises a mixed bacterial and viral composition comprising a small, non-enveloped, circular DNA virus, a small (e.g., 7kb) +ssRNA virus, a medium size (e.g., 35kb) linear dsDNA, non-enveloped virus, a bacteria, and an enveloped retrovirus. In some embodiments, the small, non-enveloped, circular DNA virus is selected from the group consisting of BK polyomavirus, Parvovirus B19, JC polyomavirus, human papillomavirus, andcombinations thereof. In some embodiments, the small (e.g., 7kb) +ssRNA virus is selected from the group consisting of EMCV, Rhinovirus, enterovirus (non-enveloped), flaviviruses like Zika, Dengue, West Nile (enveloped), Togavirus like WEE, Chikungunya, and combinations thereof. In some embodiments, the medium size (e.g., 35 kb) linear dsDNA, non-enveloped virus is selected from the group consisting of Adenovirus 7, herpes viruses like CMV, EBV, and combinations thereof. In some embodiments, the bacteria is selected from the group consisting of Chlamydia trachomatis, E coli, Staphylococcus, Haemophilus, and combinations thereof. In some embodiments, the enveloped retrovirus is selected from the group consisting of HIV-1, HIV-2, HTLV-1,2, and combinations thereof.

[0175] In yet another embodiment, the microbial composition comprises a BK polyomavirus (B), EMCV (E), Adenovirus 7 (A), Chlamydia trachomatis (C), and HIV-1 (H) (referred to herein as “BEACH”). 0176] In another embodiment, the microbial composition comprises a viral composition comprising a Parachovirus (P), an adeno-associated virus 2 (AAV-2; (A)), a Rotavirus (R), a varicella zoster virus (VZV; (V), and an Adenovirus-5 (A) (referred to herein as “PARVA”).

[0177] In some cases, the microbial composition comprises a small, ssDNA, linear, helper- dependent virus (HDV). In some cases, the microbial composition comprises a dsRNA, segmented 8- 11 virus. In some cases, the microbial composition comprises Reoviruses. In some cases, the microbial composition comprises picobimavirus. In some cases, the microbial composition comprises an influenza virus. In some cases, the microbial composition comprises a large dsDNA linear non- segmented genome. In some cases, the microbial composition comprises herpes viruses, poxviruses, baculovirus, and / or adenovirus.

[0178] The skilled artisan will appreciate that classifications for the above-mentioned viruses can be used to identify suitable alternative viruses for use in the microbial composition. Such classifications can be found online at the ViralZone.

[0179] In some embodiments, each of the first control unenriched nucleic acid libraries comprises a microbial composition comprising at least one microbe that is diluted to at least 1.0 log copy / ml, at least 2 log copies / ml, at least 3.0 log copies / ml, at least 4.0 log copies / ml, or at least 5.0 log copies / ml of each at least one microbe in the composition. In an embodiment, each of the first control unenriched nucleic acid libraries comprises a microbial composition comprising at least one archaea that is diluted to at least 1.0 log copy / ml, at least 2 log copies / ml, at least 3.0 log copies / ml, at least 4.0 log copies / ml, or at least 5.0 log copies / ml of each at least one microbe in the composition. In an embodiment, each of the first control unenriched nucleic acid libraries comprises a microbial composition comprising at least one bacteria that is diluted to at least 1.0 log copy / ml, at least 2 log copies / ml, at least 3.0 log copies / ml, at least 4.0 log copies / ml, or at least 5.0 log copies / ml of each at least one microbe in the composition. In an embodiment, each of the first control unenriched nucleic acid libraries comprises a microbial composition comprising at least one fungi that is diluted to atleast 1.0 log copy / ml, at least 2 log copies / ml, at least 3.0 log copies / ml, at least 4.0 log copies / ml, or at least 5.0 log copies / ml of each at least one microbe in the composition. In an embodiment, each of the first control unenriched nucleic acid libraries comprises a microbial composition comprising at least one protozoa that is diluted to at least 1.0 log copy / ml, at least 2 log copies / ml, at least 3.0 log copies / ml, at least 4.0 log copies / ml, or at least 5.0 log copies / ml of each at least one microbe in the composition. In an embodiment, each of the first control unenriched nucleic acid libraries comprises a microbial composition comprising at least one virus that is diluted to at least 1.0 log copy / ml, at least 2 log copies / ml, at least 3.0 log copies / ml, at least 4.0 log copies / ml, or at least 5.0 log copies / ml of each at least one microbe in the composition.

[0180] In one example, each of the first control unenriched nucleic acid libraries comprises a microbial composition comprising a BK polyomavirus, a EMCV, a Adenovirus 7, a Chlamydia trachomatis bacteria, and a HIV-1 virus that is diluted to at least 1.0 log copy / ml, at least 2 log copies / ml, at least 3.0 log copies / ml, at least 4.0 log copies / ml, or at least 5.0 log copies / ml of each at least one microbe in the composition. In another example, each of the first control unenriched nucleic acid libraries comprises a microbial composition comprising a BK polyomavirus, an EMCV, an Adenovirus 7, a Chlamydia trachomatis bacterium, and a HIV-1 virus that is diluted to about 3.0 log copies / ml, of each at least one microbe in the composition. In another example, each of the first control unenriched nucleic acid libraries comprises a microbial composition comprising a Parachovirus, a AAV-2, a Rotavirus, a VZV, and a Adenovirus-5 that is diluted to at least 1.0 log copy / ml, at least 2 log copies / ml, at least 3.0 log copies / ml, at least 4.0 log copies / ml, or at least 5.0 log copies / ml of each at least one microbe in the composition. In another example, each of the first control unenriched nucleic acid libraries comprises a microbial composition comprising a Parachovirus, an AAV-2, a Rotavirus, a VZV, and an Adenovirus-5 that is diluted to about 3.0 log copies / ml of each at least one microbe in the composition.

[0181] In an embodiment, the microbial composition comprises strains of EMCV, ZIKV, HIV-1, and SARS-CoV-2 diluted to at least 3.0 log copies / ml. c. Probes

[0182] Aspects of the disclosure involve target enrichment of microbial nucleic acids in unenriched libraries. The disclosure is not limited to the method of target enrichment or type of microbial nucleic acids that can be enriched in an unenriched library. For some embodiments, the microbial nucleic acids can be nucleic acids of archaeal origin, bacterial origin, fungal origin (e.g., yeast), protozoan origin, and / or viral origin.

[0183] In some embodiments, a plurality of probes can be used to enrich target microbial nucleic acid sequences in a larger population of sample microbial nucleic acids. In some embodiments, the plurality of probes each comprise a target binding sequence complementary to one or more target sequences, one or more non-target binding sequences, and one or more primer binding sites, such asuniversal primer binding sites. In some embodiments, target binding sequences that are complementary or at least partially complementary bind (hybridize) to target sequences. In some embodiments, primer binding sites, such as universal primer binding sites facilitate simultaneous amplification of all probes in a library. In some embodiments, the probes or adapters further comprise a barcode or index sequence. Barcodes are nucleic acid sequences that allow some features of a polynucleotide with which the barcode is associated to be identified. After sequencing, the barcode region provides an indicator for identifying a characteristic associated with the coding region or sample source. Barcodes can be designed at suitable lengths to allow sufficient degree of identification, e.g., at least about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51 , 52, 53, 54, 55, or more bases in length. Multiple barcodes, such as about 2, 3, 4, 5, 6, 7, 8, 9, 10, or more barcodes, may be used on the same molecule, optionally separated by non-barcode sequences.

[0184] In some embodiments, each barcode in a plurality of barcodes differs from every other barcode in the plurality at least three base positions, such as at least about 3, 4, 5, 6, 7, 8, 9, 10, or more positions. Use of barcodes allows for the pooling and simultaneous processing of multiple libraries for downstream applications, such as sequencing (multiplex). In some embodiments, at least 4, 8, 16, 32, 48, 64, 128, 512, 1024, 2000, 5000, or more than 5000 barcoded libraries are used. In some embodiments, the polynucleotides are ligated to one or more molecular (or affinity) tags such as a small molecule, peptide, antigen, metal, or protein to form a probe for subsequent capture of the target sequences of interest. In some embodiments, only a portion of the polynucleotides are ligated to a molecular tag. In some embodiments, two probes that possess complementary target binding sequences which are capable of hybridization form a double stranded probe pair. Polynucleotide probes or adapters may comprise unique molecular identifiers (UMI). UMIs allow for internal measurement of initial sample concentrations or stoichiometry prior to downstream sample processing (e.g., PCR or enrichment steps) which can introduce bias. In some embodiments, UMIs comprise one or more barcode sequences.(0185] Probes described here may be complementary to target sequences which are sequences in a genome. Probes described here may be complementary to target sequences which are exome sequences in a genome. Probes described here may be complementary to target sequences which are intron sequences in a genome. In some embodiments, probes comprise a target binding sequence complementary to a target sequence (of the sample nucleic acid), and at least one non-target binding sequence that is not complementary to the target. In some embodiments, the target binding sequence of the probe is about 120 nucleotides in length, or at least 10, 15, 20, 25, 50, 75, 100, 110, 120, 125, 140, 150, 160, 175, 200, 300, 400, 500, or more than 500 nucleotides in length. The target binding sequence is in some embodiments no more than 10, 15, 20, 25, 50, 75, 100, 125, 150, 175, 200, or no more than 500 nucleotides in length. The target binding sequence of the probe is in someembodiments about 120 nucleotides in length, or about 10, 15, 20, 25, 40, 50, 60, 70, 80, 85, 87, 90, 95, 97, 100, 105, 110, 115, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 135, 140, 145, 150, 155, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 175, 180, 190, 200, 210, 220, 230, 240, 250, 300, 400, or about 500 nucleotides in length. The target binding sequence is in some embodiments about 20 to about 400 nucleotides in length, or about 30 to about 175, about 40 to about 160, about 50 to about 150, about 75 to about 130, about 90 to about 120, or about 100 to about 140 nucleotides in length. The non-target binding sequence(s) of the probe is in some embodiments at least about 20 nucleotides in length, or at least about 1, 5, 10, 15, 17, 20, 23, 25, 50, 75, 100, 110, 120, 125, 140, 150, 160, 175, or more than about 175 nucleotides in length. The non-target binding sequence often is no more than about 5, 10, 15, 20, 25, 50, 75, 100, 125, 150, 175, or no more than about 200 nucleotides in length. The non-target binding sequence of the probe often is about 20 nucleotides in length, or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 25, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, or about 200 nucleotides in length. The non-target binding sequence in some embodiments is about 1 to about 250 nucleotides in length, or about 20 to about 200, about 10 to about 100, about 10 to about 50, about 30 to about 100, about 5 to about 40, or about 15 to about 35 nucleotides in length. The non-target binding sequence often comprises sequences that are not complementary to the target sequence, and / or comprise sequences that are not used to bind primers. In some embodiments, the non-target binding sequence comprises a repeat of a single nucleotide, for example polyadenine or polythymidine. A probe often comprises none (e.g., zero) or at least one non-target binding sequence. In some embodiments, a probe comprises one or two non-target binding sequences. The non-target binding sequence may be adjacent to one or more target binding sequences in a probe. For example, a non-target binding sequence is located on the 5' or 3' end of the probe. In some embodiments, the non-target binding sequence is attached to a molecular tag or spacer.

[0186] In some embodiments, the non-target binding sequence(s) may be a primer binding site. The primer binding sites often are each at least about 20 nucleotides in length, or at least about 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, or at least about 40 nucleotides in length. Each primer binding site in some embodiments is no more than about 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, or no more than about 40 nucleotides in length. Each primer binding site in some embodiments is about 10 to about 50 nucleotides in length, or about 15 to about 40, about 20 to about 30, about 10 to about 40, about 10 to about 30, about 30 to about 50, or about 20 to about 60 nucleotides in length. In some embodiments the polynucleotide probes comprise at least two primer binding sites. In some embodiments, primer binding sites may be universal primer binding sites, where all probes comprise identical primer binding sequences at these sites. In some embodiments, a pair of polynucleotide probes targeting a particular sequence and its reverse complement (e.g., a region of genomic DNA), comprising a first target binding sequence, a second target binding sequence, a first non-target binding sequence, and a second non-target binding sequence. Forexample, a pair of polynucleotide probes complementary to a particular sequence (e.g., a region of genomic DNA).

[0187] Methods of designing oligonucleotide probes that hybridize to and are useful for enriching microbial nucleic acids of interest are well-known to those skilled in the art. For embodiment, the skilled artisan can utilize publicly available databases (e.g., EMBL, Uniprot, GenBank, etc.) comprising nucleic acid sequences of microbial origin (e.g., archaeal, bacterial, fungal (e.g., yeast), protozoan, and / or viral) to identify microbial sequences (e.g., protein domain, motif, or portion thereof) that are specific for microbes of interest.

[0188] Sequences comprising a coding sequence can be narrowed to a group of sequences that comprise a protein domain or portion thereof or a protein motif. For example, narrowing the coding sequence set to those sequences comprising a protein domain (or portion thereof) or a protein motif can involve selecting only those coding sequences (or portions thereof) that are included in Pfam or other protein domain / motif database. Many viral proteins will have a described Pfam domain, some more than one. If a coding sequence is not included in Pfam or other protein domain / motif database, then a pair-wise comparison can be conducted to identify the sequence with respect to homologous clusters. The subset or group of viral sequences with Pfam domains (or portions thereof) or with sequences homologous to Pfam domains is then selected.

[0189] The sequences comprising a protein domain (or portion thereof) or a motif can be analyzed to identify statistically overrepresented DNA or protein sequence motifs. Statistically overrepresented DNA or protein sequence motifs can be with respect to whether a DNA or protein sequence motif is statistically overrepresented in any level of a taxonomy (i.e., family, subfamily, genus, subgenus, serogroup, serotype, species, subspecies, isolate, etc.). The methods can use, for example, the information theoretical concept of Shannon's Entropy to identify regions of conservation and variability in a collection of sequences. Motif finding can be used to identify sequence patterns within protein domains that are overrepresented in a level of taxonomy, thereby identifying motifs that are conserved across taxa or within a taxon.

[0199] A probability calculating algorithm, for example, MEME, Gibbs Sampler, or Splash, can identify statistically overrepresented DNA or protein sequences. The algorithm can be run exhaustively, which means that the algorithm can be run until a desired number of significant motifs are identified. Statistically overrepresented protein motifs can be cross-referenced to the nucleotide coding sequences from a database to design the oligonucleotide probes.

[0191] Statistically overrepresented protein motifs can also be cross-referenced to the nucleotide coding sequences in order to identify conserved peptides that can be used as immunogens for the generation of antibodies. It is possible that such antibodies can have a wide range of cross-reactivity across a taxon, where the cross-reactivity may potentially correspond to the level of taxonomic conservation displayed by the peptide sequence. Oligonucleotides can also be designed in a degenerate fashion with respect to the motifs.

[9192] DNA that corresponds to the motifs can be analyzed to determine which sequences are suitable for hybridization. Factors suitable for hybridization include, but are not limited to, identifying sequences: (1) having a high melting temperature, (2) little or no secondary structure, and (3) few homopolymeric stretches. DNA (or RNA) that corresponds to the motifs can be analyzed to determine which sequences are suitable for target enrichment via hybridization capture.

[0193] Oligonucleotide probes can be designed according to the algorithms and instructions described in Example 6 of U.S. Patent Publication No. 2009 / 0105092A1 (referred to herein as “the ‘092 publication”, which is hereby incorporated by reference herein in its entirety). For example, oligonucleotides that comprise the minimal number of oligonucleotides required to hybridize to any virus species in a specified taxon. The minimal number of oligonucleotides required to hybridize to any virus species in a specified taxon (i.e., a set of oligos providing comprehensive coverage of a taxon) can be determined by a Set Covering Algorithm, such as those described in Example 6 of the ‘092 publication. The Set Covering Algorithm can be adapted by the skilled artisan to design a plurality of oligonucleotide probes that comprise the minimal number of oligonucleotides required to hybridize to any archaeal, bacterial, fungal (e.g., yeast), and / or protozoan species in a specified taxon.

[0194] A set of oligonucleotides can be designed as described above such that any organism of a particular taxon can be detected.

[0195] In an embodiment, the disclosure provides a plurality of oligonucleotide probes that can detect any archaea. In another embodiment, the disclosure provides a plurality of oligonucleotide probes that can detect any bacteria. In yet another embodiment, the disclosure provides a plurality of oligonucleotide probes that can detect any fungus (e.g., yeast). In a further embodiment, the disclosure provides a plurality of oligonucleotide probes that can detect any protozoa. In one embodiment, the disclosure provides a plurality of oligonucleotide probes that can detect any vertebrate virus. In another embodiment, the disclosure provides a plurality of oligonucleotide probes that can detect a virus that belongs to a particular vertebrate virus family. In another embodiment, the disclosure provides a plurality of oligonucleotide probes that can detect a virus that belongs to a particular vertebrate virus genus. In another embodiment, the disclosure provides a plurality of oligonucleotide probes that can detect a virus that belongs to a particular vertebrate virus species. In another embodiment, the disclosure provides a plurality of oligonucleotide probes that can detect a particular viral strain.

[0196] To ensure coverage of a taxon, oligonucleotides can be designed with respect to database sequences that are not represented by a motif. Sequences that are not represented by a motif can be analyzed to identify regions that may be conserved, where the analysis can be conducted by using probability matrices that can describe mutation rates, such as PAM250 and BLOSUM. A mutation matrix can be used to identify protein stretches that have a low probability of mutation. These stretches can then be translated and used to design oligonucleotides to complement the motif-based oligonucleotides thereby ensuring coverage.

[0197] The disclosure contemplates identifying regions of conservation and variability, for example, by conducting sequence alignments of portions of genomes from databases that have not been reduced to smaller sets comprising coding sequences and PF AMs.

[0198] The concept of Shannon's Entropy can also be used to identify regions of conservation and variability in microbial genomes (e.g., viral genomes). For example, a curated, aligned database of viral genomes from public sequences is maintained to reflect diversity and phylogeny. Oligonucleotides for typing and subtyping of viruses can be selected by software for specificity and minimal cross-reactivity. This method has general applicability for a wide variety of platforms. This approach can in be extended to the identification of any nucleic acid.

[8199] Shannon's Entropy is the measure of variability in a system. In the case of DNA, there are 4 discrete states, corresponding to the nucleotides. The Shannon Entropy is the shortest binary encoding of the states of a variable. The formula is H(x)-E px log 2 where px is the probability of a given state (Shannon 1948, Cover and Thomas 1991). Alignments of viral genomes (or alignments of viral PF AMs) are made using sequences deposited in public databases. Where these databases are inadequate (e.g., there is only one representative of a given serotype or genomic sequence is incomplete) additional sequences can be obtained from either infected animals or cultured cells. Recent examples where this has been required include flaviviruses, bunyaviruses, and enteroviruses. Subregions of the alignments can be chosen for conservation or variability based on the Entropy metric by evaluating each position of the alignment. Subregions can also be chosen by the knowledge of viral biology that will allow speciation. To create alignments which represent a known sequence in a particular area, a representative seed alignment can be used to query the database by BLAST for homologous viral sequences. In cases where automated retrieval is unsatisfactory, homologous sequences can be manually retrieved from the databases. The sequences can be classified according to virus phylogeny along the lines of the ICTV scheme.

[0200] The subregions can be analyzed in parallel by software which implements the Entropy metric to quantify variability. A key objective is to identify targets for specific purposes, such as forward and reverse primers for PCR, reporter oligonucleotides for PCR, and oligonucleotides for target enrichment via hybridization capture. A sliding window of 10, 25, 50, 60, 70, 75, 100, 110, 120, 125, 150, 160, 170, or 180 nucleotides can be used to evaluate every possible target.

[0201] Oligonucleotides can be chosen based on their potential to (i) capture broad microbial (e.g., viral) taxa including unknown microbes (e.g., viruses, e.g., genus specific targets) or (ii) allow discrete speciation of microbial (e.g., viral) taxa (e.g., serotype or strain-specific targets). Whereas the former approach facilitates broad range surveillance and pathogen discovery, the latter facilitates molecular epidemiology and microbial forensics. In selecting broad targets, regions which are highly conserved (low entropy scores, connoting a similar makeup of nucleotides among the strains) are chosen. A degenerate oligonucleotide is determined, similar to consensus PCR, through fullyautomated programs. The degenerate target design algorithm maximizes the chance for hybridization with members of a genus while minimizing the number of degenerate positions.

[0202] A refining algorithm can be used to choose minimized degeneracy based on the propensity for nucleotide changes to occur together between strains. Serotype and strain-specific oligonucleotides can be determined by an algorithm which identifies speciating areas. By evaluating the contribution of a family to the overall Entropy (variability) of a microbial (e.g., virus) taxon, regions can be selected which are conserved in the family but variable in the rest of the genus. The algorithm maximizes intrafamily similarity to the target while minimizing extra-family similarity. Filtering of the potential targets examines critical performance characteristics including Tm, hairpin formation and self-annealing.

[0203] The oligonucleotide probes can be designed for: detection and differentiation of microorganism and host transcripts in clinical, environmental, and food samples; genetic compatibility studies; screening of blood and transplantation products; and forensics. The oligonucleotide probes can also be designed for the sensitive, multiplex detection and characterization of genetic targets where precise target sequence might not be known. The oligonucleotide probes can detect both completely sequenced microbial (e.g., viral) genomes and incompletely sequenced species. The oligonucleotide probes can also be designed to include more than one species representative because, for example, sequences can be considerably divergent within a species.

[0204] The oligonucleotide sequences can be designed to hybridize to related but not necessarily identical sequence targets. Primers can also be designed for consensus PCR and targets for hybridization capture. Because oligonucleotides are designed based on sequence conservation, the oligonucleotide probes can be used for sensitive, multiplex detection and characterization of genetic targets where precise target sequence might not be known.

[0205] The disclosure contemplates designing a plurality of oligonucleotide probes to enrich microbes of interest in unenriched libraries comprising microbial nucleic acids. The plurality of oligonucleotide probes can be designed to tile genomes of microbes of interest. The plurality of oligonucleotide probes can designed as single-stranded DNA capture probes having a length of between about 60 nt and about 180 nt, between about 70 nt and 170 nt, between about 80 nt and about 160 nt, between about 90 nt and about 150 nt, between about 100 nt and about 140 nt, between about 110 nt and about 130 nt, or about 115 nt, about 116 nt, about 117 nt, about 118 nt, about 119 nt, about 120 nt, about 121 nt, about 122 nt, about 123 nt, about 124 nt, or about 125 nt. In some embodiments, the single-stranded DNA capture probes comprise a length of about 70 nt, about 80 nt, about 90 nt, about 100 nt, about 110 nt, about 120 nt, about 130 nt, about 140 nt, about 150 nt, about 160 nt, or about 170 nt. In an embodiment, the plurality of probes comprises single-stranded DNA capture probes having a length of about 120 nt. In an embodiment, the plurality of probes comprises single- stranded DNA capture probes having a length of about 140 nt. In an embodiment, the plurality of probes comprises single-stranded DNA capture probes having a length of about 160 nt.

[0206] In some embodiments, the plurality of oligonucleotide probes are labeled with at least one molecular tag. In some embodiments, PCR is used to introduce molecular tags (via primers comprising the molecular tag) onto the probes during amplification. In some embodiments, the molecular tag comprises one or more of biotin, folate, a polyhistidine, a FLAG tag, glutathione, or other molecular tag. In some embodiments probes are labeled at the 5' terminus. In some embodiments, the probes are labeled at the 3' terminus. In some embodiments, both the 5' and 3' termini are labeled with a molecular tag. In some embodiments, the 5' terminus of a first probe in a pair is labeled with at least one molecular tag, and the 3' terminus of a second probe in the pair is labeled with at least one molecular tag. In some embodiments, a spacer is present between one or more molecular tags and the nucleic acids of the probe. In some embodiments, the spacer may comprise an alkyl, polyol, or polyamino chain, a peptide, or a polynucleotide.[(1207] The solid support used to capture probe-target nucleic acid complexes in some embodiments, can be a bead or a surface. The solid support in some embodiments comprises glass, plastic, or other material capable of comprising a capture moiety that will bind the molecular tag. In some embodiments, a bead is a magnetic bead. For example, probes labeled with biotin are captured with a magnetic bead comprising streptavidin. The probes are contacted with a library of nucleic acids to allow binding of the probes to target sequences. In some embodiments, blocking oligonucleotides are added to prevent binding of the probes to one or more adapter sequences attached to the target nucleic acids. In some embodiments, blocking oligonucleotides comprise one or more nucleic acid analogues. In some embodiments, blocking oligonucleotides have a uracil substituted for thymine at one or more positions.[0208 A plurality of probes described herein may comprise complementary target binding sequences which bind to one or more target nucleic acid sequences. In some embodiments, the target sequences are any DNA or RNA nucleic acid sequence. In some embodiments, target sequences may be longer than the probe insert. In some embodiments, target sequences may be shorter than the probe insert. In some embodiments, target sequences may be the same length as the probe insert. For example, the length of the target sequence may be at least or about at least 2, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 150, 200, 300, 400, 500, 1000, 2000, 5,000, 12,000, 20,000 nucleotides, or more. The length of the target sequence may be at most or about at most 20,000, 12,000, 5,000, 2,000, 1,000, 500, 400, 300, 200, 150, 100, 50, 45, 35, 30, 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 2 nucleotides, or less. The length of the target sequence may fall from 2-20,000, 3-12,000, 5-5, 5000, 10-2,000, 10-1,000, 10-500, 9-400, 11-300, 12-200, 13-150, 14-100, 15-50, 16-45, 17-40, 18-35, and 19-25. The probe sequences may target sequences associated with specific genes, diseases, regulatory pathways, or other biological functions consistent with the specification.

[0209] In some embodiments, a single probe insert is complementary to one or more target sequences in a larger polynucleic acid (e.g., sample nucleic acid). An example target sequence is an exon. In some embodiments, one or more probes target a single target sequence. In someembodiments, a single probe may target more than one target sequence. In some embodiments, at least at least 2, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 150, 200, 300, 400, 500, 1000, 2000, 5,000, 12,000, 20,000, or more than 20,000 probes target a single target sequence. In some embodiments no more than 4 probes directed to a single target sequence overlap, or no more than 3, 2, 1, or no probes targeting a single target sequence overlap. In some embodiments, one or more probes do not target all bases in a target sequence, leaving one or more gaps. In some embodiments, the gaps are near the middle of the target sequence. In some embodiments, the gaps are at the 5' or 3' ends of the target sequence. In some embodiments, the gaps are 6 nucleotides in length. In some embodiments, the gaps are no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, or no more than 50 nucleotides in length. In some embodiments, the gaps are at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, or at least 50 nucleotides in length. In some embodiments, the gap length falls within 1-50, 1-40, 1-30, 1-20, 1-10, 2-30, 2-20, 2-10, 3-50, 3-25, 3-10, or 3-8 nucleotides in length. In some embodiments, a set of probes targeting a sequence do not comprise overlapping regions amongst probes in the set when hybridized to complementary sequence. In some embodiments, a set of probes targeting a sequence do not have any gaps amongst probes in the set when hybridized to complementary sequence. Probes may be designed to maximize uniform binding to target sequences. In some embodiments, probes are designed to minimize target binding sequences of high or low GC content, secondary structure, repetitive / palindromic sequences, or other sequence feature that may interfere with probe binding to a target. In some embodiments, a single probe may target a plurality of target sequences.

[0210] The skilled artisan will appreciate that the number of oligonucleotide probes utilized in any particular target enrichment can vary, depending on the number of microbes of interest to be detected, the length of the microbial genome, and the length of the probes. For example, for a 9kb virus like HIV, tiling its genome with 120 bp probes would require about 75 probes. The number of the plurality of probes used can be about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 75, about 80, about 85, about 90, about 95, or about 100 probes per microbial (e.g., viral) strain of interest. In some embodiments,

[0211] A plurality of probes described herein can comprise at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 250,000, 300,000, 350,000, 400,000, 450,000, 500,000, 550,000, 600,000, 650,000, 700,000, 750,000, 800,000, 850,000, 900,000, 950,000, 1,000,000 or more than 1,000,000 probes. A plurality of probes can comprise no more than 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 500,000, or no more than 1,000,000 probes. A plurality of probes can comprise 10 to 500, 20 to 1000, 50 to 2000, 100 to 5000, 500 to 10,000, 1,000 to 5,000, 10,000 to 50,000, 100,000 to 500,000, or 50,000 to 1,000,000 probes. A plurality of probes can comprise about 370,000; 400,000; 500,000 or more different probes. A probe library described herein can comprise at least 2000, 5000, 10,000, 50,000, 100,000, 200,000, 500,000, 1,000,000, 2,000,000, 5,000,000, 10,000,000, 20,000,000, 50,000,000, 75,000,000,100,000,000 or more than 200,000,000 probes. A plurality of probes described herein can comprise about 2000, 5000, 10,000, 50,000, 100,000, 200,000, 500,000, 1,000,000, 2,000,000, 5,000,000, 10,000,000, 20,000,000, 50,000,000, 75,000,000, 100,000,000 or at least 200,000,000 probes. A plurality of probes described herein can comprise no more than 2000, 5000, 10,000, 50,000, 100,000, 200,000, 500,000, 1,000,000, 2,000,000, 5,000,000, 10,000,000, 20,000,000, 50,000,000, 75,000,000, 100,000,000 or no more than 200,000,000 probes. A plurality of probes can comprise 10,000 to 500,000 20,000 to 100,000, 50,000 to 200,000, 100,000 to 5,000,000, 500,000 to 10,000,000, 1,000,000 to 5,000,000, 10,000,000 to 50,000,000, 100,000 to 5,000,000, or 500,000 to 10,000,000 probes.

[0212] In some embodiments, a plurality of probes can comprise at least 1000, 5000, 10,000, 100,000 500,000, 1 million, 10 million, 100 million, 200 million, or at least 500 million bases. In yet other embodiments, probe libraries comprise about 1000, 5000, 10,000, 100,000, 500,000, 1 million, 10 million, 100 million, 200 million, or about 500 million bases. In some embodiments, a plurality of probes can comprise 1000 to 1 million, 5000 to 1 million, 10,000 to 5 million, 100,000 to 5 million, 500,000 to 100 million, 1 million to 200 million, 10 million to 500 million, 100 million to 250 million, or 200 million to 500 million bases.

[0213] The plurality of probes can be designed to be complementary to any number of microbes of interest. In some embodiments, the plurality of probes are complementary to at least one, at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 microbes of interest.

[0214] In one embodiment, the at least one microbe of interest is at least one archaea of interest and the plurality of probes are designed to be complementary to at least one, at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 archaea of interest. In such embodiments, the number of plurality of probes can be at least 10,000, at least 50,000, at least 100,000, at least 150,000 at least 200,000, at least 250,000, at least 300,000, at least 350,000, at least 400,000, at least 450,000, at least 500,000, at least 550,000, at least 600,000, at least 650,000, at least 700,000, at least 750,000, at least800,000, at least 850,000, at least 900,000, at least 950,000, or at least 1,000,000 probes tiling the genomes of at least one, at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 archaea of interest.

[0215] In another embodiment, the at least one microbe of interest is at least one bacteria of interest and the plurality of probes are designed to be complementary to at least one, at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 bacteria of interest. In such embodiments, the number of plurality of probes can be at least 10,000, at least 50,000, at least 100,000, at least 150,000 at least 200,000, at least 250,000, at least 300,000, at least 350,000, at least 400,000, at least 450,000, at least 500,000, at least 550,000, at least 600,000, at least 650,000, at least 700,000, at least 750,000, at least 800,000, at least 850,000, at least 900,000, at least 950,000, or at least 1,000,000 probes tiling the genomes of at least one, at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 bacteria of interest.

[0216] The plurality of probes designed to enrich bacterial nucleic acids of interest can be used to enrich any type of bacteria of interest, include without limitation, pathogenic bacteria, antibiotic- resistant bacteria, environmental bacteria, probiotics and gut microbiota, industrial bacteria, plant associated bacteria, oral and skin microbiota, and combinations thereof. In some embodiments, the plurality of probes can be designed to enrich and detect bacteria that cause diseases in humans, such as Escherichia coli (various pathogenic strains like EHEC, EPEC), Staphylococcus aureus (including MRSA), Streptococcus pneumoniae, Mycobacterium tuberculosis, Salmonella spp., Neisseriagonorrhoeae, Clostridioides difficile, and combinations thereof. In some embodiments, the plurality of probes can be designed to enrich and detect bacteria affecting livestock and pets, such as Brucella spp., Bordetella bronchiseptica, Pasteurella multocida, Mycoplasma spp., and combinations thereof. In some embodiments, the plurality of probes can be designed to enrich and detect antibiotic-resistant bacteria, including, without limitation: MRSA (Methicillin-resistant Staphylococcus aureus) — probes can target mecA and mecC genes associated with resistance; VRE (Vancomycin-resistant Enterococcus) — probes can detect genes like vanA and vanB; ESBL-producing Enterobacteriaceae — probes targeting beta-lactamase genes such as bla_CTX-M, bla_SHV, and bla_TEM; Multidrug- Resistant Acinetobacter baumannii-probes could target resistance genes and specific strains and combinations thereof.

[0217] In some embodiments, the plurality of probes can be designed to enrich and detect environmental bacteria, including without limitation, soil bacteria — probes can target Pseudomonas spp., Bacillus spp., Streptomyces spp., and Rhizobium spp., which are important for nutrient cycling and bioremediation; water bacteria — probes can detect bacteria involved in water quality, such as Vibrio spp. (including V. cholerae), Legionella pneumophila, and cyanobacteria like Microcystis spp, and combinations thereof. In some embodiments, the plurality of probes can be designed to enrich and detect probiotics and gut microbiota — probes can be designed for specific species such as Lactobacillus spp., Bifidobacterium spp., Saccharomyces boulardii (a yeast used as a probiotic), and Streptococcus thermophilus; commensal gut bacteria— probes can target major gut microbiota members such as Bacteroides spp., Faecalibacterium prausnitzii, Prevotella spp., Akkermansia muciniphila, and Clostridium spp., and combinations thereof. In some embodiments, the plurality of probes can be designed to enrich and detect industrial bacteria, including without limitation, fermentation and bioprocessing-probes can be designed for bacteria used in industrial fermentation, such as Lactococcus lactis, Corynebacterium glutamicum, and Escherichia coli (industrial strains), and combinations thereof; bioremediation-probes for bacteria like Pseudomonas putida and Deinococcus radiodurans, which are used in bioremediation to degrade pollutants, and combinations thereof. In some embodiments, the plurality of probes can be designed to enrich and detect plant- associated bacteria, including without limigation, plant pathogens — probes can be designed for bacteria that infect plants, like Agrobacterium tumefaciens, Xanthomonas spp., Erwinia amylovora, and Ralstonia solanacearum; nitrogen-fixing bacteria-probes targeting bacteria such as Rhizobium spp., Azospirillum spp., and Frankia spp. involved in nitrogen fixation in symbiosis with plants, and combinations thereof. In some embodiments, the plurality of probes can be designed to enrich and detect oral and skin microbiota, including without limitation, oral bacteria-probes can be designed for common oral microbiota like Streptococcus mutans, Porphyromonas gingivalis, Fusobacterium, and combinations thereof.

[0218] In some embodiments, the plurality of probes can comprise the 4.2 million bacterial probes described in WO2019 / 226992A1 (the contents of which are incorporated herein by reference in theirentirety), which can be used in solution-based capture of pathogenic bacterial nucleic acids present in complex samples containing variable proportions of different pathogenic bacterial and host nucleic acids, including bacterial nucleic acids from the 307 most important known pathogenic bacterial species listed in Table 1 therein.

[0219] In another embodiment, the at least one microbe of interest is at least one fungus (e.g., yeast) of interest and the plurality of probes are designed to be complementary to at least one, at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 fungi of interest. In such embodiments, the number of plurality of probes can be at least 10,000, at least 50,000, at least 100,000, at least 150,000 at least 200,000, at least 250,000, at least 300,000, at least 350,000, at least 400,000, at least 450,000, at least 500,000, at least 550,000, at least 600,000, at least 650,000, at least 700,000, at least 750,000, at least 800,000, at least 850,000, at least 900,000, at least 950,000, or at least 1 ,000,000 probes tiling the genomes of at least one, at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 fungi (e.g., yeast) of interest.

[0220] In another embodiment, the at least one microbe of interest is at least one protozoa of interest and the plurality of probes are designed to be complementary to at least one, at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 bacteria of interest. In such embodiments, the number of plurality of probes can be at least 10,000, at least 50,000, at least 100,000, at least 150,000 at least 200,000, at least 250,000, at least 300,000, at least 350,000, at least 400,000, at least 450,000, at least 500,000, at least 550,000, at least 600,000, at least 650,000, at least 700,000, at least 750,000, at least800,000, at least 850,000, at least 900,000, at least 950,000, or at least 1,000,000 probes tiling the genomes of at least one, at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 protozoa of interest.

[0221] In another embodiment, the at least one microbe of interest is at least one virus of interest and the plurality of probes are designed to be complementary to at least one, at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 viruses of interest. In such embodiments, the number of plurality of probes can be at least 10,000, at least 50,000, at least 100,000, at least 150,000 at least 200,000, at least 250,000, at least 300,000, at least 350,000, at least 400,000, at least 450,000, at least 500,000, at least 550,000, at least 600,000, at least 650,000, at least 700,000, at least 750,000, at least 800,000, at least 850,000, at least 900,000, at least 950,000, or at least 1,000,000 probes tiling the genomes of at least one, at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 protozoa of interest.

[0222] The plurality of probes designed to enrich viral nucleic acids of interest can be used to enrich any type of virus of interest, including without limitation, a single-stranded RNA virus, a double-stranded RNA virus, a single-stranded DNA virus, and a double-stranded DNA virus. In some cases, the at least one virus of interest comprises a vertebrate virus species. In other cases, the at least one virus of interest comprises a human virus species.

[0223] In certain embodiments, the plurality of probes are designed to hybridize to and enrich nucleic acids of at least one, at least two, at least three, at least four, at least five, at least 10, at least25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000, at least 3500, at least 4000, at least 5000, at least 6000, at least 7000, at least 7500, at least 8000, at least 8500, at least 9000, at least 10000, at least 11,000, at least 11,500, at least 12,000, at least 12,500, at least 13,000, at least 13,500, at least 14,000, at least 14,500, or at least 15,000 different viral strains.

[0224] The disclosure contemplates designing oligonucleotides allow for the creation of a plurality of oligonucleotide probes that can detect or identify viruses in a sample. The probes can be designed to detect any vertebrate (e.g., human) virus in a sample.

[0225] In an example embodiment, the disclosure provides a plurality of oligonucleotide probes that can provide comprehensive coverage of vertebrate viruses, where the plurality of oligonucleotide probes comprise the nucleic acid sequences listed in the CD-ROM Table Appendix of U.S. Patent Publication No. 2009 / 0105092A1 (the sequences of which are hereby incorporated by reference herein in their entirety).

[0226] In another embodiment, the disclosure provides a plurality of oligonucleotide probes that can provide comprehensive coverage of vertebrate viruses, where the plurality of oligonucleotide probes comprise nucleic acid sequences derived or reverse-translated from the amino acid sequences listed in the CD-ROM Table Appendix of U.S. Patent Publication No. 2009 / 0105092A1 (the sequences of which are hereby incorporated by reference herein in their entirety).

[0227] In another embodiment, the disclosure provides sequence motifs that are derived or obtained from the amino acid sequences listed in the CD-ROM Table Appendix of U.S. Patent Publication No. 2009 / 0105092A1 (the sequences of which are hereby incorporated by reference herein in their entirety). These sequence motifs can be compared to nucleic acid sequences in NCBI or ICTV databases or the NCBI / ICTV integrated database to identify viral nucleic acid sequences that may code for the sequence motifs. Oligonucleotides can be designed based on the sequence motifs, where the design can include degenerate sequences, conservative mutations, and sequence variation such that the oligonucleotides are at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the nucleotide sequences in NCBI or ICTV databases that code for the motifs.

[0228] In one embodiment, the disclosure provides a plurality of oligonucleotide probes for detecting vertebrate viruses, where the plurality of oligonucleotide probes are reverse translated from amino acid sequences, where each amino acid sequence comprises a motif conserved in either: (a) a virus family, where the virus family is selected from the group consisting of: Adenoviridae, Arenaviridae, Arteriviridae, Asfarviridae, Astroviridae, Bimaviridae, Bomaviridae, Bunyaviridae,Caliciviridae, Circoviridae, Coronaviridae, Deltavirus, Filoviridae, Flaviviridae, Hepadnaviridae, Hepatitis E-like viruses, Herpesviridae, Infectious laryngotrachetitis-like viruses, Iridoviridae, Nodaviridae, Orthomyxoviridae, Papillomaviridae, Paramyxoviridae, Parvoviridae, Picomaviridae, Polyomaviridae, Poxyiridae, Reoviridae, Retroviridae, Rhabdoviridae, and Togaviridae; (b) a virus genus, where the virus genus is selected from the group consisting of: Asfivirus, Orthopoxvirus, Parapoxvirus, Avipoxvirus, Capripoxvirus, Leporipoxvirus, Suipoxvirus, Molluscipoxvirus, Yatapoxvirus, Entomopoxvirus A, Entomopoxvirus B, Entomopoxvirus C, Iridovirus, Chloriridovirus, Ranavirus, Lymphocystivirus, Simplexvirus, Varicellovirus, Cytomegalovirus, Muromegalovirus, Roseolovirus, Lymphocryptovirus, Rhadinovirus, Ichnovirus, Bracovirus, Polyomavirus, Papillomavirus, Mastadenovirus, Aviadenovirus, Orthoreovirus, Orbivirus, Rotavirus, Coltivirus, Aquareovirus, Cypovirus, Fijivirus, Phytoreovirus, Oryzavirus, Aquabimavirus, Avibimavirus, Entomobimavirus, Influenzavirus A, Influenzavirus B, Influenzavirus C, Influenzavirus D, Paramyxovirus, Morbillivirus, Rubulavirus, Pneumovirus, Bomavirus, Marburgvirus, Ebolavirus, Arenavirus, Alpharetrovirus, Betaretrovirus, Gammaretrovirus, Type D Retrovirus group, Deltaretrovirus, Epsilonretrovirus, Lentivirus, Spumavirus, Bunyavirus, Hantavirus, Nairovirus, Phlebovirus, Tospovirus, Calicivirus, Enterovirus, Rhinovirus, Hepatovirus, Cardiovirus, Aphthovirus, Astrovirus, Flavivirus, Pestivirus, Hepacivirus, Alphanodavirus, Coronavirus, Torovirus, Alphavirus, Arterivirus, and Deltavirus; and / or (c) a virus species from the virus family in (a) or the virus genus in (b); where the set of oligonucleotides as a whole can detect any vertebrate virus. In one embodiment, the amino acid sequences are selected from the CD-ROM Table Appendix of U.S. Patent Publication No. 2009 / 0105092A1 (the sequences of which are hereby incorporated by reference herein in their entirety).

[0022] .In one embodiment, the disclosure provides a plurality of oligonucleotide probes for detecting vertebrate viruses from a particular family, where the set of oligonucleotides are reverse translated from amino acid sequences, where each amino acid sequence comprises a motif conserved in the virus family to be detected, where the virus family is selected from the group consisting of: Adenoviridae, Arenaviridae, Arteriviridae, Asfarviridae, Astroviridae, Bimaviridae, Bomaviridae, Bunyaviridae, Caliciviridae, Circoviridae, Coronaviridae, Deltavirus, Filoviridae, Flaviviridae, Hepadnaviridae, Hepatitis E-like viruses, Herpesviridae, Infectious laryngotrachetitis-like viruses, Iridoviridae, Nodaviridae, Orthomyxoviridae, Papillomaviridae, Paramyxoviridae, Parvoviridae, Picomaviridae, Polyomaviridae, Poxyiridae, Reoviridae, Retroviridae, Rhabdoviridae, and Togaviridae. The motif conserved in the vims family can include motifs that are conserved in genera or in species of the family. In one embodiment, the amino acid sequences are selected from the appropriate virus family table from the CD-ROM Table Appendix of U.S. Patent Publication No. 2009 / 0105092 Al (the sequences of which are hereby incorporated by reference herein in their entirety).

[0230] In one embodiment, the disclosure provides a plurality of oligonucleotide probes for detecting vertebrate viruses from a particular genus, where the plurality of oligonucleotide probes are reverse translated from amino acid sequences, where each amino acid sequence comprises a motif conserved in the virus genus to be detected, where the virus genus is selected from the group consisting of: Asfivirus, Orthopoxvirus, Parapoxvirus, Avipoxvirus, Capripoxvirus, Leporipoxvirus, Suipoxvirus, Molluscipoxvirus, Yatapoxvirus, Entomopoxvirus A, Entomopoxvirus B, Entomopoxvirus C, Iridovirus, Chloriridovirus, Ranavirus, Lymphocystivirus, Simplexvirus, Varicellovirus, Cytomegalovirus, Muromegalovirus, Roseolovirus, Lymphocryptovirus, Rhadinovirus, Ichnovirus, Bracovirus, Polyomavirus, Papillomavirus, Mastadenovirus, Aviadenovirus, Orthoreovirus, Orbivirus, Rotavirus, Coltivirus, Aquareovirus, Cypovirus, Fijivirus, Phytoreovirus, Oryzavirus, Aquabimavirus, Avibimavirus, Entomobimavirus, Influenzavirus A, Influenzavirus B, Influenzavirus C, Influenzavirus D, Paramyxovirus, Morbillivirus, Rubulavirus, Pneumovirus, Bomavirus, Marburgvirus, Ebolavirus, Arenavirus, Alpharetrovirus, Betaretrovirus, Gammaretrovirus, Type D Retrovirus group, Deltaretrovirus, Epsilonretrovirus, Lentivirus, Spumavirus, Bunyavirus, Hantavirus, Nairovirus, Phlebovirus, Tospovirus, Calicivirus, Enterovirus, Rhinovirus, Hepatovirus, Cardiovirus, Aphthovirus, Astrovirus, Flavivirus, Pestivirus, Hepacivirus, Alphanodavirus, Coronavirus, Torovirus, Alphavirus, Arterivirus, and Deltavirus. The motif conserved in the virus genus can include motifs that are conserved in genus or in species of the genus. In one embodiment, the amino acid sequences are selected from the appropriate virus genus from the appropriate family table from the CD-ROM Table Appendix of U.S. Patent Publication No.2009 / 0105092 Al (the sequences of which are hereby incorporated by reference herein in their entirety).

[0231] The plurality of probes designed to enrich viral nucleic acids of interest can be used to enrich any type of virus of interest, including without limitation, a single-stranded RNA virus, a double-stranded RNA virus, a single-stranded DNA virus, and a double-stranded DNA virus. In some cases, the at least one virus of interest comprises a vertebrate virus species. In other cases, the at least one virus of interest comprises a human virus species.

[0232] In an embodiment, the disclosure provides a plurality of oligonucleotide probes for detecting a virus from an order listed in Table 1. In some embodiments, the disclosure provides a plurality of probes for detecting a virus from the order Bunyavirales. In some embodiments, the disclosure provides a plurality of probes for detecting a virus from the order Herpesvirales. In some embodiments, the disclosure provides a plurality of probes for detecting a virus from the order Mononegavirales. In some embodiments, the disclosure provides a plurality of probes for detecting a virus from the order Nidovirales. In some embodiments, the disclosure provides a plurality of probes for detecting a virus from the order Ortervirales. In some embodiments, the disclosure provides a plurality of probes for detecting a virus from the order Picomavirales.

[0233] In some embodiments, the plurality of probes can comprise at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 250,000, 300,000, 350,000, 400,000, 450,000, 500,000, 550,000, 600,000, 650,000, 700,000, 750,000, 800,000, 850,000, 900,000, 950,000, 1,000,000 or more than 1,000,000 probes designed to hybridize with and enrich nucleic acids of at least one, at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 viruses of interest from an order listed in Table 1.

[0234] In such embodiments, the number of plurality of probes can be at least 10,000, at least 50,000, at least 100,000, at least 150,000 at least 200,000, at least 250,000, at least 300,000, at least 350,000, at least 400,000, at least 450,000, at least 500,000, at least 550,000, at least 600,000, at least 650,000, at least 700,000, at least 750,000, at least 800,000, at least 850,000, at least 900,000, at least 950,000, or at least 1,000,000 probes designed to hybridize to and enrich nucleic acids of at least one, at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000, at least 3500, at least 4000, at least 5000, at least 6000, at least 7000, at least 7500, at least 8000, at least 8500, at least 9000, at least 10000, at least 11,000, at least 11,500, at least 12,000, at least 12,500, at least 13,000, at least 13,500, at least 14,000, at least 14,500, or at least 15,000 different viral strains from at least one order, at least two orders, at least three orders, at least four orders, or at least five virus orders listed in Table 1.

[0235] In an embodiment, the disclosure provides a plurality of oligonucleotide probes for detecting a virus from a family listed in Table 2. In some embodiments, the plurality of probes can comprise at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 250,000, 300,000, 350,000, 400,000, 450,000, 500,000, 550,000, 600,000, 650,000, 700,000, 750,000, 800,000, 850,000, 900,000, 950,000, 1,000,000 or more than 1,000,000 probes designed to hybridize with and enrich nucleic acids of at least one, at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 viruses of interest from at least one family, at least two families, at least three families, at least four families, at least five families, at least six families, at least seven families, at least eight families, at least nine families, at least 10 families, at least 15 families, at least 20 families, at least 25 families, at least 30 families, or at least 35 families listed in Table 2.

[0236] In some embodiments, the plurality of probes can comprise at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 250,000, 300,000, 350,000, 400,000, 450,000, 500,000, 550,000, 600,000, 650,000, 700,000, 750,000, 800,000, 850,000, 900,000, 950,000, 1,000,000 or more than 1,000,000 probes designed to hybridize with and enrich nucleic acids of at least one, at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000, at least 3500, at least 4000, at least 5000, at least 6000, at least 7000, at least 7500, at least 8000, at least 8500, at least9000, at least 10000, at least 11,000, at least 11,500, at least 12,000, at least 12,500, at least 13,000, at least 13,500, at least 14,000, at least 14,500, or at least 15,000 different viral strains of interest from at least one family, at least two families, at least three families, at least four families, at least five families, at least six families, at least seven families, at least eight families, at least nine families, at least 10 families, at least 15 families, at least 20 families, at least 25 families, at least 30 families, or at least 35 families listed in Table 2.

[0237] In an embodiment, the disclosure provides a plurality of oligonucleotide probes for detecting a virus from a genus listed in Table 3. In some embodiments, the plurality of probes can comprise at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 250,000, 300,000, 350,000, 400,000, 450,000, 500,000, 550,000, 600,000, 650,000, 700,000, 750,000, 800,000, 850,000, 900,000, 950,000, 1,000,000 or more than 1,000,000 probes designed to hybridize with and enrich nucleic acids of at least one, at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or atleast 3000 viruses of interest from at least one genus, at least two genera, at least three genera, at least four genera, at least five genera, at least six genera, at least seven genera, at least eight genera, at least nine genera, at least 10 genera, at least 15 genera, at least 20 genera, at least 25 genera, at least 30 genera, at least 35 genera, at least 40 genera, at least 45 genera, at least 50 genera, at least 55 genera, at least 60 genera, at least 65 genera, at least 70 genera, at least 75 genera, at least 80 genera, at least 85 genera, at least 90 genera, or at least 95 genera listed in Table 3.

[0238] In some embodiments, the plurality of probes can comprise at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 250,000, 300,000, 350,000, 400,000, 450,000, 500,000, 550,000, 600,000, 650,000, 700,000, 750,000, 800,000, 850,000, 900,000, 950,000, 1,000,000 or more than 1,000,000 probes designed to hybridize with and enrich nucleic acids of at least one, at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000, at least 3500, at least 4000, at least 5000, at least 6000, at least 7000, at least 7500, at least 8000, at least 8500, at least 9000, at least 10000, at least 11,000, at least 11,500, at least 12,000, at least 12,500, at least 13,000, at least 13,500, at least 14,000, at least 14,500, or at least 15,000 different viral strains of interest from at least one genus, at least two genera, at least three genera, at least four genera, at least five genera, at least six genera, at least seven genera, at least eight genera, at least nine genera, at least 10 genera, at least 15 genera, at least 20 genera, at least 25 genera, at least 30 genera, at least 35 genera, at least 40 genera, at least 45 genera, at least 50 genera, at least 55 genera, at least 60 genera, at least 65 genera, at least 70 genera, at least 75 genera, at least 80 genera, at least 85 genera, at least 90 genera, or at least 95 genera listed in Table 3.

[0239] In an embodiment, the disclosure provides a plurality of oligonucleotide probes for detecting a virus species listed in Table 4. In some embodiments, the plurality of probes can comprise at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 250,000, 300,000, 350,000, 400,000, 450,000, 500,000, 550,000, 600,000, 650,000, 700,000, 750,000, 800,000, 850,000, 900,000, 950,000, 1,000,000 or more than 1,000,000 probes designed to hybridize with and enrich nucleic acids of at least one, at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 viruses of interest from at least one, at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 species in Table 4.

[0240] In some embodiments, the plurality of probes can comprise at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 250,000, 300,000, 350,000, 400,000, 450,000, 500,000, 550,000, 600,000, 650,000, 700,000, 750,000, 800,000, 850,000, 900,000, 950,000, 1,000,000 or more than 1,000,000 probes designed to hybridize with and enrich nucleic acids of at least one, at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1 100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000, at least 3500, at least 4000, at least 5000, at least 6000, at least 7000, at least 7500, at least 8000, at least 8500, at least 9000, at least 10000, at least 11,000, at least 11,500, at least 12,000, at least 12,500, at least 13,000, at least 13,500, at least 14,000, at least 14,500, or at least 15,000 different viral strains of interest from at least one, at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 species in Table 4.

[0241] In an embodiment, the disclosure provides a plurality of oligonucleotide probes for detecting a viral strain listed in Table 5. In some embodiments, the plurality of probes can comprise at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 250,000, 300,000, 350,000, 400,000, 450,000, 500,000, 550,000, 600,000, 650,000, 700,000, 750,000, 800,000, 850,000, 900,000, 950,000, 1,000,000 or more than 1,000,000 probes designed to hybridize with and enrich nucleic acids of at least one, at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000 viruses of interest from at least one, at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000, at least 3500, at least 4000, at least 5000, at least 6000, at least 7000, at least 7500, at least 8000, at least 8500, at least 9000, at least 10000, at least 11,000, at least 11,500, at least 12,000, at least 12,500, at least 13,000, at least 13,500, at least 14,000, at least 14,500, or at least 15,000 different viral strains listed in Table 5. In certain embodiments, the plurality of probes can comprise at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 250,000, 300,000, 350,000, 400,000, 450,000, 500,000, 550,000, 600,000, 650,000, 700,000, 750,000, 800,000, 850,000, 900,000, 950,000, 1,000,000 or more than 1,000,000 probes designed to hybridize with and enrich nucleic acids of at least one, at least two, at least three, at least four, at least five, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1200, at least 1250, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1750, at least 1800, at least 1900, at least 2000, at least 2100, at least 2200, at least 2300, at least 2400, at least 2500, at least 2600, at least 2700, at least 2750, at least 2800, at least 2900, at least, or at least 3000, at least 3500, at least 4000, at least 5000, at least 6000, at least 7000, at least 7500, at least 8000, at least 8500, at least 9000, at least 10000, at least 11,000, at least 11,500, at least 12,000, at least 12,500, at least 13,000, at least 13,500, at least 14,000, at least 14,500, or at least 15,000 different viral strains listed in Table 5.

[0242] The plurality of probes can be designed based on a cohort of viral nucleic acid sequences. Specifically, the cohort of viral nucleic acid sequences can comprise viral nucleic acid sequences from NCBI's RefSeq collection, complementary representation of unique regions from Genome Neighbor targets, selected representation of NCBI Influenza Virus Resource sequences, and the entirety of the probe space represented on the Virochip microarray (Chen, E.C., Miller, S.A., DeRisi, J.L., Chiu, C.Y. Using a Pan-Viral Microarray Assay (Virochip) to Screen Clinical Samples for Viral Pathogens. J. Vis. Exp. (50), e2536, doi:10.3791 / 2536 (2011), the entirety of which is hereby incorporated herein by reference), GEO accession number GPL 15905 (the entirety of which is hereby incorporated hereinby reference including all supplementary files and probe sequences). The cohort of viral nucleic acid sequences can comprise viral nucleic acid sequences from all DNA and RNA viruses with sequenced genomes from vertebrate hosts, excluding human endogenous retroviruses and bacteriophages. The cohort of viral nucleic acid sequences can comprise more than 150,000, more than 160,000, more than 170,000, or more than 180,000 viral nucleic acid sequences. Specifically, the cohort of viral nucleic acid sequences comprises 185,835 viral nucleic acid sequences. The cohort of viral nucleic acid sequences comprises greater than 150 Mb, greater than 160 Mb, greater than 170 Mb, greater than 180 Mb, or greater than 190 Mb of viral nucleic acid sequences. Specifically, the cohort of viral nucleic acid sequences comprises 198.9 Mb of viral nucleic acid sequences. Even more specifically, the cohort of viral nucleic acid sequences comprises about 27 Mb of viral nucleic acid sequences from RefSeq, about 153 Mb of viral nucleic acid sequences from Genome Neighbor targets, about 16 Mb from Influenza Virus Resource sequences, and about 3 Mb of viral nucleic acid sequences from Virochip microarray. The cohort of viral nucleic acid sequences is used to design a panel of probes that specifically hybridize to the viral nucleic acid sequences of the cohort of viral nucleic acid sequences.

[0243] The plurality of probes of the disclosure are capable of specifically hybridizing to greater than 10,000 viral nucleic acid sequences. For example, a plurality of probes of the disclosure comprises probes capable of specifically hybridizing to greater than 10,000, greater than 15,000, greater than 20,000, greater than 25,000, greater than 30,000, greater than 35,000, greater than 40,000, greater than 45,000, greater than 50,000, greater than 55,000, greater than 60,000, greater than 65,000, greater than 70,000, greater than 75,000, greater than 80,000, greater than 85,000, greater than 90,000, greater than 95,000, greater than 100,000, greater than 110,000, greater than 120,000, greater than 130,000, greater than 140,000, greater than 150,000, greater than 160,000, greater than 170,000, greater than 180,000, greater than 190,000, or greater than 200,000 viral nucleic acid sequences. In one embodiment, the plurality of probes comprises probes capable of specifically hybridizing to 185,835 viral nucleic acid sequences (See, Wylie et al., Enhanced virome sequencing using targeted sequence capture. Genome Res. 2015; 24(12): 1910-20, the disclosure of which is hereby incorporated by reference in its entirety, including all supplemental information and zip files associated with the publication).

[0244] In certain embodiments, a plurality of probes of the disclosure comprises probes capable of specifically hybridizing to viral nucleic acid sequences from NCBI's RefSeq collection, complementary representation of unique regions from NCBI Genome Neighbor targets, selected representation of NCBI Influenza Virus Resource sequences, and the entirety of the probe space represented on the Virochip microarray, GEO accession number GPL15905. In certain embodiments, a plurality of probes of the disclosure comprises probes capable of specifically hybridizing to about 27 Mb of viral nucleic acid sequences from RefSeq, about 153 Mb of viral nucleic acid sequences fromGenome Neighbor targets, about 16 Mb from Influenza Virus Resource sequences, and about 3 Mb of viral nucleic acid sequences from the Virochip microarray.

[0245] In other embodiments, a plurality of probes of the disclosure comprises probes capable of specifically hybridizing to viral nucleic acid sequences from 34 viral families comprising 190 annotated viral genera and 337 species. Non-limiting examples of viral families with which the probes are capable of specifically hybridizing to include Adenoviridae, Alloherpesviridae, Asfarviridae, Herpesviridae, Iridoviridae, Malacoherpesviridae, Papillomaviridae, Polyomaviridae, Poxviridae, Bimaviridae, Picobimaviridae, Reoviridae, Retroviridae, Hepadnaviridae, Parvoviridae, Anelloviridae, Circoviridae, Coronaviridae, Bunyaviridae, Flaviviridae, Orthomyxoviridae, Caliciviridae, Togaviridae, Arenaviridae, Arteriviridae, Astroviridae, Bomaviridae, Filoviridae, Hepeviridae, Paramyxoviridae, Picomaviridae, and Rhabdoviridae.

[0246] In one embodiment, a plurality of probes of the disclosure comprises probes capable of specifically hybridizing to viral nucleic acid sequences from the viruses listed in Table 10 of U.S. Patent No. 10,597,736 B2 (which is incorporated herein by reference in its entirety, including the viral nucleic acid sequences referenced therein).

[0247] In another embodiment, the plurality of probes comprises at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 300,000, 400,000, or 500,000 or more probes designed to hybridize with and enrich nucleic acids of ssRNA viruses from a family selected from the group consisting of Picomawiridae, Paramyxoviridae, Rhabdoviridae, Coronaviridae, Orthomyxoviridae, Caliciviridae, Flaviviridae, Bunyaviridae, Filoviridae, Astroviridae, Hepeviridae, Arenaviridae, Togaviridae, Nodaviridae, Nyamiviridae, Arterviridae, and / or Bomaviridae. In an embodiment, the plurality of probes comprises at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 300,000, 400,000, or 500,000 or more probes designed to hybridize with and enrich nucleic acids of ssRNA viruses from the family Rhabdoviridae that are from a genera selected from the group consisting of Ephemerovirus, Lyssavirus, Novirhabdovirus, Perhabdovirus, Spirivirus, Tibrovirus, Tupavirus, Vesiculovirus, and combinations thereof. In an embodiment, the plurality of probes comprises at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 300,000, 400,000, or 500,000 or more probes designed to hybridize with and enrich nucleic acids of ssRNA viruses from the family Coronaviridae that are from a genera selected from the group consisting of Alphacoronavirus, Bafomovoris, Betacoronavirus, Deltacoronavirus, Gamacoronavirus, Torovirus, and combinations thereof. In an embodiment, the plurality of probes comprises at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 300,000, 400,000, or 500,000 or more probes designed to hybridize with and enrich nucleic acids of ssRNA viruses from the family Orthomyxoviridae that are from a genera selected from the group consisting of Influenzavirus A, Influenzavirus B, Influenzavirus C, Influenzavirus D, Isavirus, Thogotovirus, and combinations thereof. In an embodiment, the plurality of probes comprises at least 10, 20, 50, 100,200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 300,000, 400,000, or 500,000 or more probes designed to hybridize with and enrich nucleic acids of ssRNA viruses from the family Caliciviridae that are from a genera selected from the group consisting of Lagovirus, Nebovirus, Norovirus, Sapovirus, Vesivirus, and combinations thereof. In an embodiment, the plurality of probes comprises at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 300,000, 400,000, or 500,000 or more probes designed to hybridize with and enrich nucleic acids of ssRNA viruses from the family Flaviviridae that are from a genera selected from the group consisting of Flavivirus, Hepacivirus, Pegivirus, Pestivirus, and combinations thereof. In an embodiment, the plurality of probes comprises at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 300,000, 400,000, or 500,000 or more probes designed to hybridize with and enrich nucleic acids of ssRNA viruses from the family Bunyaviridae that are from a genera selected from the group consisting of Hantavirus, Nairovirus, Orthobunyavirus, Phlebovirus, and combinations thereof. In another embodiment, the plurality of probes comprises at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 300,000, 400,000, or 500,000 or more probes designed to hybridize with and enrich nucleic acids of dsDNA viruses from a family selected from the group consisting of Polyomaviridae, Malacoherpesviridae, Asfarviridae, Iridoviridae, Alloherpesviridae, Adenoviridae, Poxviridae, Herpesviridae, Papillomaviridae, and combinations thereof. In another embodiment, the plurality of probes comprises at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 300,000, 400,000, or 500,000 or more probes designed to hybridize with and enrich nucleic acids of ssDNA viruses from a family selected from the group consisting of Circoviridae, Parvoviridae, Anelloviridae, and combinations thereof. In another embodiment, the plurality of probes comprises at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 300,000, 400,000, or 500,000 or more probes designed to hybridize with and enrich nucleic acids of dsRNA viruses from a family selected from the group consisting of Picobimaviridae, Bimaviridiae, Reoviridae, and combinations thereof. In another embodiment, the plurality of probes comprises at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 300,000, 400,000, or 500,000 or more probes designed to hybridize with and enrich nucleic acids of retrotranscribing viruses from a family selected from the group consisting of Hepadnaviridae, Retroviridae, and combinations thereof. d. Blocking oligonucleotides

[0248] Aspects of the disclosure involve blocking oligonucleotides. Described herein are blocking oligonucleotides (or hybridization reagents) comprising polynucleotides (polynucleotide library). In some embodiments, such blocking oligonucleotides are configured to reduce undesired hybridization to sequences in a complex sample mixture (e.g., genome or collection of genomes). In some embodiments, blocking oligonucleotides are configured to bind to modified genomes. In someembodiments, blocking oligonucleotides comprise at least one modification relative to a genomic DNA. In some embodiments, the at least one modification comprises a different abundance of one or more polynucleotides relative to an abundance in the genome. In some embodiments, modified genomes comprise post-transcriptional modifications identified through a conversion process. In some embodiments, the post-transcriptional modification comprises methylation (e.g., 5-methylcytosine, 5- hydroxymethylcytosine, or other modification). In some embodiments, blocking libraries are configured to bind to samples from specific organisms, such as humans or plants. In some embodiments, organisms comprise highly repetitive genetic elements, such as those found in polyploid species. Hybridization reagents used for blocking (including synthetic blocking oligonucleotides) may contain repetitive sequences. e. Universal Adapters

[0249] Aspects of the disclosure involve the use of universal adapters. In some embodiments, the universal adapters disclosed herein may comprise a universal polynucleotide adapter comprising a first strand and a second strand. In some embodiments, a first strand comprises a first primer binding region, a first non-complementary region, and a first yoke region. In some embodiments, a second strand comprises a second primer binding region, a second non-complementary region, and a second yoke region. In some embodiments, a primer binding region allows for PCR amplification of a polynucleotide adapter. In some embodiments, a primer binding region allows for PCR amplification of a polynucleotide adapter and concurrent addition of one or more barcodes to the polynucleotide adapter. In some embodiments, the first yoke region is complementary to the second yoke region. In some embodiments, the first non-complementary region is not complementary to the second non- complementary region. In some embodiments, the universal adapter is a Y-shaped or forked adapter. In some embodiments, one or more yoke regions comprise nucleobase analogues that raise the Tm between a first yoke region and a second yoke region. Primer binding regions as described herein may be in the form of a terminal adapter region of a polynucleotide. In some embodiments, a universal adapter comprises one index sequence. In some embodiments, a universal adapter comprises one unique molecular identifier. In some embodiments, universal adapters are configured for use with barcoded primers, where after ligation, barcoded primers are added via PCR.

[0250] A universal adapter may be shortened relative to a typical barcoded adapter (e.g., full- length “Y adapter”). For example, a universal adapter strand can be 20-45 bases in length. In some embodiments, a universal adapter strand can be 25-40 bases in length. In some embodiments, a universal adapter strand can be 30-35 bases in length. In some embodiments, a universal adapter strand is no more than 50 bases in length, no more than 45 bases in length, no more than 40 bases in length, no more than 35 bases in length, no more than 30 bases in length, or no more than 25 bases in length. In some embodiments, a universal adapter strand is about 25, 27, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, or about 60 bases in length. In some embodiments, a universal adapterstrand is about 60 base pairs in length. In some embodiments, a universal adapter strand is about 58 base pairs in length. In some embodiments, a universal adapter strand is about 52 base pairs in length. In some embodiments, a universal adapter strand is about 33 base pairs in length.

[0251] A universal adapter may be modified to facilitate ligation with a sample nucleic acid. For example, the 5' terminus is phosphorylated. In some embodiments, a universal adapter comprises one or more non-native nucleobase linkages such as a phosphorothioate linkage. For example, a universal adapter comprises a phosphorothioate between the 3' terminal base, and the base adjacent to the 3' terminal base. In some embodiments, an adapter-ligated sample polynucleotide comprises a sample polynucleotide (e.g., sample nucleic acid) with adapters universal adapters ligated to both the 5' and 3' end of the sample polynucleotide to form an adapter-ligated polynucleotide. A duplex sample polynucleotide comprises both a first strand (forward) and a second strand (reverse).[0252 Universal adapters may contain any number of different nucleobases (DNA, RNA, etc.), nucleobase analogues, or non-nucleobase linkers or spacers. For example, an adapter comprises one or more nucleobase analogues or other groups that enhance hybridization (Tm) between two strands of the adapter. In some embodiments, nucleobase analogues are present in the yoke region of an adapter. Nucleobase analogues and other groups include but are not limited to locked nucleic acids (LNAs), bicyclic nucleic acids (BNAs), C5-modified pyrimidine bases, 2'-O-methyl substituted RNA, peptide nucleic acids (PNAs), glycol nucleic acid (GNAs), threose nucleic acid (TNAs), xenonucleic acids (XNAs) morpholino backbone-modified bases, minor grove binders (MGBs), spermine, G- clamps, or a anthraquinone (Uaq) caps.

[0253] Universal adapters may comprise any number of nucleobase analogues (such as LNAs or BNAs), depending on the desired hybridization Tm. For example, an adapter comprises 1 to 20 nucleobase analogues. In some embodiments, an adapter comprises 1 to 8 nucleobase analogues. In some embodiments, an adapter comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or at least 12 nucleobase analogues. In some embodiments, an adapter comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or about 16 nucleobase analogues. In some embodiments, the number of nucleobase analogous is expressed as a percent of the total bases in the adapter. For example, an adapter comprises at least 1%, 2%, 5%, 10%, 12%, 18%, 24%, 30%, or more than 30% nucleobase analogues. In some embodiments, adapters (e.g., universal adapters) described herein comprise methylated nucleobases, such as methylated cytosine. f. Barcoded primers

[0254] Aspects of the disclosure involve the use of barcodes or indices. Nucleic acid primers may comprise defined sequences, such as barcodes (or indices). Barcodes can be attached to universal adapters, for example, using PCR and barcoded primers to generate barcoded adapter-ligated sample nucleic acids. Primer binding sites, such as universal primer binding sites, facilitate simultaneousamplification of all members of a barcode primer library, or a subpopulation of members. In some embodiments, a primer binding site comprises a region that binds to a flow cell or other solid support during next generation sequencing. In some embodiments, a barcoded primer comprises a P5 (5'- AATGATACGGCGACCACCGA-3' (SEQ ID NO:1) or P7 (5'- CAAGCAGAAGACGGCATACGAGAT-3' (SEQ ID NO:2) sequence. In some embodiments, primer binding sites are configured to bind to universal adapter sequences, and facilitate amplification and generation of barcoded adapters. In some embodiments, barcoded primers are no more than 60 bases in length. In some embodiments, barcoded primers are no more than 55 bases in length. In some embodiments, barcoded primers are 50-60 bases in length. In some embodiments, barcoded primers are about 60 bases in length. In some embodiments, barcodes described herein comprise methylated nucleobases, such as methylated cytosine.

[0255] The number of unique barcodes available for a barcode set (collection of unique barcodes or barcode combinations configured to be used together to unique define samples) may depend on the barcode length. In some embodiments, a Hamming distance is defined by the number of base differences between any two barcodes. In some embodiments, a Levenshtein distance is defined by the number changes needed to change one barcode into another (insertions, substitutions, or deletions). In some embodiments, barcode sets described herein comprise a Levenshtein distance of at least 2, 3, 4, 5, 6, 7, or at least 8. In some embodiments, barcode sets described herein comprise a Hamming distance of at least 2, 3, 4, 5, 6, 7, or at least 8.

[0256] Barcodes may be incorrectly associated with a different sample than they were assigned. In some embodiments, incorrect barcodes occur from PCR errors (e.g., substitution) during library amplification. In some embodiments, entire barcodes “hop” or are transferred from one sample polynucleotide to another. In some embodiments, such transfers result from cross-contamination of free adapters or primers during a library generation workflow. In some embodiments a group of barcodes (barcode set) is chosen to minimize “barcode hopping”. In some embodiments, barcode hopping (for a single barcode) for a barcode set described herein is no more than 7%, 5%, 4%, 3%, 2%, 1%, 0.5%, or no more than 0.1%. In some embodiments, barcode hopping (for a single barcode) for a barcode set described herein is 0.1-6%, 0.1-5%, 0.2-5%, 0.5-5%, 1-7%, 1-5%, or 0.5-7%. In some embodiments, barcode hopping (for two barcodes) for a barcode set described herein is no more than 0.7%, 0.5%, 0.4%, 0.3%, 0.2%, 0.1%, 0.05%, or no more than 0.1%. In some embodiments, barcode hopping (for two barcodes) for a barcode set described herein is 0.01-0.6%, 0.01-0.5%, 0.02- 0.5%, 0.05-0.5%, 0.1-0.7%, 0.1-0.5%, or 0.05-0.7%.

[0257] Barcoded primers comprise one or more barcodes. In some embodiments, the barcodes are added to universal adapters through PCR reaction. Barcodes are nucleic acid sequences that allow some features of a polynucleotide with which the barcode is associated to be identified. In some embodiments, a barcode comprises an index sequence. In some embodiments, index sequences allow for identification of a sample, or unique source of nucleic acids to be sequenced. In someembodiments, a barcode or combination of barcodes identifies a specific microbe of interest. In yet other embodiments, a barcode or combination of barcodes identifies a specific archaea of interest. In still yet other embodiments, a barcode or combination of barcodes identifies a specific bacteria of interest. In still yet other embodiments, a barcode or combination of barcodes identifies a specific fungi of interest. In still yet other embodiments, a barcode or combination of barcodes identifies a specific protozoa of interest. In still yet further embodiments, a barcode or combination of barcodes identifies a specific virus of interest (e.g., a virus order, a virus family, a virus genus, a virus species, or a virus strain. In some embodiments, a barcode or combination of barcodes identifies each virus order listed in Table 1. In some embodiments, a barcode or combination of barcodes identifies each virus family listed in Table 2. In some embodiments, a barcode or combination of barcodes identifies each virus genus listed in Table 3. In some embodiments, a barcode or combination of barcodes identifies each virus species listed in Table 4. In some embodiments, a barcode or combination of barcodes identifies each virus strain listed in Table 5.

[0258] In some embodiments, a barcode or combination of barcodes identifies a specific patient.In some embodiments, a barcode or combination of barcodes identifies a specific sample from a patient among other samples from the same patient. After sequencing, the barcode (or barcode region) provides an indicator for identifying a characteristic associated with the coding region or sample source. Barcodes can be designed at suitable lengths to allow sufficient degree of identification, e.g., at least about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51 , 52, 53, 54, 55, or more bases in length. Multiple barcodes, such as about 2, 3, 4, 5, 6, 7, 8, 9, 10, or more barcodes, may be used on the same molecule, optionally separated by non-barcode sequences. In some embodiments, a barcode is positioned on the 5' and the 3' sides of a sample polynucleotide. In some embodiments, each barcode in a plurality of barcodes differ from every other barcode in the plurality at least three base positions, such as at least about 3, 4, 5, 6, 7, 8, 9, 10, or more positions. Use of barcodes allows for the pooling and simultaneous processing of multiple libraries for downstream applications, such as multiplexed sequencing.

[0259] In some embodiments, at least 4, 8, 16, 32, 48, 64, 128, 256, or more 512 barcoded libraries are used. In some embodiments, at least 400, 500, 800, 1000, 2000, 5000, 10,000, 12,000, 15,000, 18,000, 20,000, or at 25,000 barcodes are used. Barcoded primers or adapters may comprise unique molecular identifiers (UMI). In some embodiments, such UMIs uniquely tag all nucleic acids in a sample. In some embodiments, at least 60%, 70%, 80%, 90%, 95%, or more than 95% of the nucleic acids in a sample are tagged with a UMI. In some embodiments, at least 85%, 90%, 95%, 97%, or at least 99% of the nucleic acids in a sample are tagged with a unique barcode, or UMI. In some embodiments, barcoded primers comprise an index sequence and one or more UMI. UMIs allow for internal measurement of initial sample concentrations or stoichiometry prior to downstream sample processing (e.g., PCR or enrichment steps) which can introduce bias. In some embodiments,UMIs comprise one or more barcode sequences. In some embodiments, each strand (forward vs. reverse) of an adapter-ligated sample polynucleotide possesses one or more unique barcodes. Such barcodes are optionally used to uniquely tag each strand of a sample polynucleotide. In some embodiments, a barcoded primer comprises an index barcode and a UMI barcode. In some embodiments, after amplification with at least two barcoded primers, the resulting amplicons comprise two index sequences and two UMIs. In some embodiments, after amplification with at least two barcoded primers, the resulting amplicons comprise two index barcodes and one UMI barcode. In some embodiments, each strand of a universal adapter-sample polynucleotide duplex is tagged with a unique barcode, such as a UMI or index barcode.

[0260] Barcoded primers in a library comprise a region that is complementary to a primer binding region on a universal adapter. For example, universal adapter binding region is complementary to primer region of the universal adapter, and universal adapter binding region is complementary to primer region of the universal adapter. Such arrangements facilitate extension of universal adapters during PCR, and attach barcoded primers. In some embodiments, the Tm between the primer and the primer binding region is 40-65 degrees C. In some embodiments, the Tm between the primer and the primer binding region is 42-63 degrees C. In some embodiments, the Tm between the primer and the primer binding region is 50-60 degrees C. In some embodiments, the Tm between the primer and the primer binding region is 53-62 degrees C. In some embodiments, the Tm between the primer and the primer binding region is 54-58 degrees C. In some embodiments, the Tm between the primer and the primer binding region is 40-57 degrees C. In some embodiments, the Tm between the primer and the primer binding region is 40-50 degrees C. In some embodiments, the Tm between the primer and the primer binding region is about 40, 45, 47, 50, 52, 53, 55, 57, 59, 61, or 62 degrees C. g. Universal blockers

[0261] Aspects of the disclosure involve blockers (e.g., universal blockers). Universal blockers are used to prevent off-target binding of capture probes to adapters ligated to genomic fragments, or adapter-adapter hybridization. Adapter blockers used for preventing off-target hybridization may target a portion or the entire adapter. In some embodiments, specific blockers are used that are complementary to a portion of the adapter that includes the unique index sequence. In cases where the adapter-tagged genomic library comprises a large number of different indices, it can be beneficial to design blockers which either do not target the index sequence, or do not hybridize strongly to it. For example, a “universal” blocker targets a portion of the adapter that does not comprise an index sequence (index independent), which allows a minimum number of blockers to be used regardless of the number of different index sequences employed. In some embodiments, no more than 8 universal blockers are used. In some embodiments, 4 universal blockers are used. In some embodiments, 3 universal blockers are used.

[0262] In some embodiments, 2 universal blockers are used. In some embodiments, 1 universal blocker is used. In one arrangement, 4 universal blockers are used with adapters comprising at least 4, 8, 16, 32, 64, 96, or at least 128 different index sequences. In some embodiments, the different index sequences comprise at least or about 4, 6, 8, 10, 12, 14, 16, 18, 20, or more than 20 base pairs (bp). In some embodiments, a universal blocker is not configured to bind to a barcode sequence. In some embodiments, a universal blocker partially binds to a barcode sequence. In some embodiments, a universal blocker which partially binds to a barcode sequence further comprises nucleotide analogs, such as those that increase the Tm of binding to the adapter (e.g., LNAs or BNAs).

[6263] Blockers may contain any number of different nucleobases (DNA, RNA, etc.), nucleobase analogues (non-canonical), or non-nucleobase linkers or spacers. In some embodiments, such blockers may are described as a “set”, where the set comprises two or more blockers configured to prevent unwanted interactions with the same adapter sequence. In some embodiments, universal blockers prevent adapter-adapter interactions independent of one or more barcodes present on at least one of the adapters. For example, a blocker comprises one or more nucleobase analogues or other groups that enhance hybridization (Tm) between the blocker and the adapter. In some embodiments, a blocker comprises one or more nucleobases which decrease hybridization (Tm) between the blocker and the adapter (e.g., “universal” bases). In some embodiments, a blocker described herein comprises both one or more nucleobases which increase hybridization (Tm) between the blocker and the adapter and one or more nucleobases which decrease hybridization (Tm) between the blocker and the adapter.

[0264] Described herein are hybridization blockers comprising one or more regions which enhance binding to targeted sequences (e.g., adapter), and one or more regions which decrease binding to target sequences (e.g., adapter). In some embodiments, each region is tuned for a given desired level of off-bait activity during target enrichment applications. In some embodiments, each region can be altered with either a single type of chemical modification / moiety or multiple types to increase or decrease overall affinity of a molecule for a targeted sequence. In some embodiments, the melting temperature of all individual members of a blocker set are held above a specified temperature (e.g., with the addition of moieties such as LNAs and / or BNAs). In some embodiments, a given set of blockers will improve off bait performance independent of index length, independent of index sequence, and independent of how many adapter indices are present in hybridization.

[0265] Blockers may comprise moieties which increase and / or decrease affinity for a target sequencing, such as an adapter. In some embodiments, such specific regions can be thermodynamically tuned to specific melting temperatures to either avoid or increase the affinity for a particular targeted sequence. In some embodiments, this combination of modifications is designed to help increase the affinity of the blocker molecule for specific and unique adapter sequence and decrease the affinity of the blocker molecule for repeated adapter sequence (e.g., Y-stem annealing portion of adapter). In some embodiments, blockers comprise moieties which decrease binding of a blocker to the Y-stem region of an adapter. In some embodiments, blockers comprise moieties whichdecrease binding of a blocker to the Y-stem region of an adapter, and moieties which increase binding of a blocker to non- Y-stem regions of an adapter.

[0266] Blockers (e.g., universal blockers) and adapters may form a number of different populations during hybridization. In some embodiments, a population ‘A’ comprises blockers correctly bound to non-index regions of the adapters. In a population ‘B’, a region of the blockers is bound to the “yoke” region of the adapter, but a remaining portion of the blocker does not bind to an adjacent region of the adapter. In a population ‘C’, two blockers unproductively dimerize. In a population ‘D’, blockers are unbound to any other nucleic acids. In some embodiments, when the number of DNA modifications that decrease affinity in the Y-stem annealing region of the blocker are increased, the populations ‘A’ & ‘D’ dominate and either have the desired or minimal effect. In some embodiments, as the number of DNA modifications that decrease affinity in the Y-stem annealing region of the blocker are decreased, the populations ‘B’ & ‘C’ dominate and have undesired effects where daisy-chaining or annealing to other adapters can occur (‘B’) or sequester blockers where they are unable to function properly (‘C’).

[0267] The index on both single or dual index adapter designs may be either partially or fully covered by universal blockers that have been extended with specifically designed DNA modifications to cover adapter index bases. In some embodiments, such modifications comprise moieties which decrease annealing to the index, such as universal bases. In some embodiments, the index of a dual index adapter is partially covered (or is overlapped) by one or more blockers. In some embodiments, the index of a dual index adapter is fully covered by one or more blockers. In some embodiments, the index of a single index adapter is partially covered by one or more blockers. In some embodiments, the index of a single index adapter is fully covered by one or more blockers. In some embodiments, a blocker overlaps an index sequence by at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20 or more than 20 bases. In some embodiments, a blocker overlaps an index sequence by no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, or no more than 25 bases. In some embodiments, a blocker overlaps an index sequence by about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20 or about 30 bases. In some embodiments, a blocker overlaps an index sequence by 1-5, 1-3, 2-5, 2-8, 2-10, 3-6, 3-10, 4-10, 4-15, 1-4 or 5-7 bases. In some embodiments, a region of a blocker which overlaps an index sequences comprises at least one 2-deoxyinosine or 5-nitroindole nucleobase.

[0268] One or two blockers may overlap with an index sequence present on an adapter. In some embodiments, one or two blockers combined overlap with at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20 or more than 20 bases of the index sequence. In some embodiments, one or two blockers combined overlap with no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20 or no more than 20 bases of the index sequence. In some embodiments, one or two blockers combined overlap with about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20 or about 20 bases of the index sequence. In some embodiments, one or two blockers combined overlap by 1-5, 1-3, 2-5, 2-8, 2-10, 3- 6, 3-10, 4-10, 4-15, 1-4 or 5-7 bases of the index sequence. In some embodiments, a region of ablocker which overlaps index sequences comprises at least one 2-deoxyinosine or 5-nitroindole nucleobase.

[0269] In a first arrangement, the length of the adapter index overhang may be varied. When designed from a single side, the adapter index overhang can be altered to cover from 0 to n of the adapter index bases from either side of the index. This allows for the ability to design such adapter blockers for both single and dual index adapter systems.

[0270] In a second arrangement, the adapter index bases are covered from both sides. When adapter index bases are covered from both sides, the length of the covering region of each blocker can be chosen such that a single pair of blockers is capable of interacting with a range of adapter index lengths while still covering a significant portion of the total number of index bases. As an example, take two blockers that have been designed with 3 bp overhangs that cover the adapter index. In the context of 6 bp, 8 bp, or 10 bp adapter index lengths, these blockers will leave 0 bp, 2 bp, or 4 bp exposed during hybridization, respectively.

[0271] In a third arrangement, modified nucleobases are selected to cover index adapter bases. Examples of these modifications that are currently commercially available include degenerate bases (i.e., mixed bases of A, T, C, G), 2'-deoxyInosine, & 5-nitroindole.

[0272] In a fourth arrangement, blockers with adapter index overhangs bind to either the sense (i.e., ‘top’) or anti-sense (i.e., ‘bottom’) strand of a next generation sequencing library.

[0273] In a fifth arrangement, blockers are further extended to cover other polynucleotide sequences (e.g., a poly-A tail added in a previous biochemical step in order to facilitate ligation or other method to introduce a defined adapter sequence, unique molecular identifier for bioinformatic assignment following sequencing, etc.) in addition to the standard adapter index bases of defined length and composition. These types of sequences can be placed in multiple locations of an adapter and in this case the most widely utilized case (i.e., unique molecular index next to the genomic insert) is presented. Other positions for the unique molecular identifier (e.g., next to adapter index bases) could also be addressed with similar approaches.

[0274] In a sixth arrangement, all of the previous arrangements are utilized in various combinations to meet a targeted performance metric for off-bait performance during target enrichment under specified conditions.

[0275] Blockers may comprise moieties, such as nucleobase analogues. Nucleobase analogues and other groups include but are not limited to locked nucleic acids (LNAs), bicyclic nucleic acids (BNAs), C5-modified pyrimidine bases, 2'-O-methyl substituted RNA, peptide nucleic acids (PNAs), glycol nucleic acid (GNAs), threose nucleic acid (TNAs), inosine, 2'-deoxyInosine, 3 -nitropyrrole, 5- nitroindole, xenonucleic acids (XNAs) morpholino backbone-modified bases, minor grove binders (MGBs), spermine, G-clamps, or a anthraquinone (Uaq) caps. In some embodiments, nucleobase analogues comprise universal bases, where the nucleobase has a lower Tm for binding to a cognate nucleobase. In some embodiments, universal bases comprise 5-nitroindole or 2'-deoxyInosine. Insome embodiments, blockers comprise spacer elements that connect two polynucleotide chains. In some embodiments, such nucleobase analogues are added to control the Tm of a blocker. Blockers may comprise any number of nucleobase analogues (such as LNAs or BNAs), depending on the desired hybridization Tm. For example, a blocker comprises 20 to 40 nucleobase analogues. In some embodiments, a blocker comprises 8 to 16 nucleobase analogues. In some embodiments, a blocker comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or at least 12 nucleobase analogues. In some embodiments, a blocker comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or about 16 nucleobase analogues. In some embodiments, the number of nucleobase analogous is expressed as a percent of the total bases in the blocker. For example, a blocker comprises at least 1%, 2%, 5%, 10%, 12%, 18%, 24%, 30%, or more than 30% nucleobase analogues. In some embodiments, the blocker comprising a nucleobase analogue raises the Tm in a range of about 2° C. to about 8° C. for each nucleobase analogue. In some embodiments, the Tm is raised by at least or about 1° C., 2° C., 3° C., 4° C., 5° C., 6° C., 7° C., 8° C., 9° C., 10° C., 12° C„ 14° C., or 16° C. for each nucleobase analogue. Such blockers in some embodiments are configured to bind to the top or “sense” strand of an adapter. In some embodiments, blockers are configured to bind to the bottom or “anti-sense” strand of an adapter. In some embodiments a set of blockers includes sequences which are configured to bind to both top and bottom strands of an adapter. In yet other embodiments, additional blockers are configured to the complement, reverse, forward, or reverse complement of an adapter sequence. In some embodiments, a set of blockers targeting a top (binding to the top) or bottom strand (or both) is designed and tested, followed by optimization, such as replacing a top blocker with a bottom blocker, or a bottom blocker with a top blocker. In some embodiments, a blocker is configured to overlap fully or partially with the bases of an index or barcode on an adapter. In other embodiments, a set of blockers comprise at least one blocker overlapping with an adapter index sequence. In some embodiments, a set of blockers comprise at least one blocker overlapping with an adapter index sequence, and at least one blocker which does not overlap with an adapter sequence. In some embodiments, a set of blockers comprise at least one blocker which does not overlap with a yoke region sequence. In yet other embodiments, a set of blockers comprise at least one blocker which does not overlap with a yoke region sequence and at least one blocker which overlaps with a yoke region sequence. In some embodiments, a set of blockers comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 blockers.

[0276] Blockers may be any length, depending on the size of the adapter or hybridization Tm. For example, blockers are 20 to 50 bases in length. In some embodiments, blockers are 25 to 45 bases, 30 to 40 bases, 20 to 40 bases, or 30 to 50 bases in length. In some embodiments, blockers are 25 to 35 bases in length. In some embodiments, blockers are at least 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or at least 35 bases in length. In some embodiments, blockers are no more than 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or no more than 35 bases in length. In some embodiments, blockers are about 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or about 35 bases in length. In some embodiments, blockers are about 50bases in length. In some embodiments, a set of blockers targeting an adapter-tagged genomic library fragment comprises blockers of more than one length. In some embodiments, two blockers are tethered together with a linker. Various linkers are well known in the art, and can comprise alkyl groups, polyether groups, amine groups, amide groups, or other chemical groups. In some embodiments, linkers comprise individual linker units, which are connected together (or attached to blocker polynucleotides) through a backbone such as phosphate, thiophosphate, amide, or other backbone. In one arrangement, a linker spans the index region between a first blocker that each targets the 5' end of the adapter sequence and a second blocker that targets the 3' end of the adapter sequence. In some embodiments, capping groups are added to the 5' or 3' end of the blocker to prevent downstream amplification. Capping groups variously comprise polyethers, polyalcohols, alkanes, or other non-hybridizable groups that prevent amplification. In some embodiments, such groups are connected through phosphate, thiophosphate, amide, or other backbone. In some embodiments, one or more blockers are used. In some embodiments, at least 4 non-identical blockers are used. In some embodiments, a first blocker spans a first 3' end of an adaptor sequence, a second blocker spans a first 5' end of an adaptor sequence, a third blocker spans a second 3' end of an adaptor sequence, and a fourth blockers spans a second 5' end of an adaptor sequence. In some embodiments, a first blocker is at least 20, 1 , 22, 23, 24, 25, 26, 27, 28, 29, 30, 31 , 32, 33, 34, or at least 35 bases in length. In some embodiments, a second blocker is at least 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or at least 35 bases in length. In some embodiments, a third blocker is at least 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or at least 35 bases in length. In some embodiments, a fourth blocker is at least 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or at least 35 bases in length. In some embodiments, a first blocker, second blocker, third blocker, or fourth blocker comprises a nucleobase analogue. In some embodiments, the nucleobase analogue is LNA.

[0277] The design of blockers may be influenced by the desired hybridization Tm to the adapter sequence. In some embodiments, non-canonical nucleic acids (for example locked nucleic acids, bridged nucleic acids, or other non-canonical nucleic acid or analog) are inserted into blockers to increase or decrease the blocker's Tm. In some embodiments, the Tm of a blocker is calculated using a tool specific to calculating Tm for polynucleotides comprising a non-canonical amino acid. In some embodiments, a Tm is calculated using the Exiqon online prediction tool. In some embodiments, blocker Tm described herein are calculated in-silico. In some embodiments, the blocker Tm is calculated in-silico, and is correlated to experimental in-vitro conditions. Without being bound by theory, an experimentally determined Tm may be further influenced by experimental parameters such as salt concentration, temperature, presence of additives, or other factors. In some embodiments, Tm described herein are in-silico determined Tm that are used to design or optimize blocker performance. In some embodiments, Tm values are predicted, estimated, or determined from melting curve analysis experiments. In some embodiments, blockers have a Tm of 70 degrees C. to 99 degrees C. In some embodiments, blockers have a Tm of 75 degrees C. to 90 degrees C. In someembodiments, blockers have a Tm of at least 85 degrees C. In some embodiments, blockers have a Tm of at least 70, 72, 75, 77, 80, 82, 85, 88, 90, or at least 92 degrees C. In some embodiments, blockers have a Tm of about 70, 72, 75, 77, 80, 82, 85, 88, 90, 92, or about 95 degrees C. In some embodiments, blockers have a Tm of 78 degrees C. to 90 degrees C. In some embodiments, blockers have a Tm of 79 degrees C. to 90 degrees C. In some embodiments, blockers have a Tm of 80 degrees C. to 90 degrees C. In some embodiments, blockers have a Tm of 81 degrees C. to 90 degrees C. In some embodiments, blockers have a Tm of 82 degrees C. to 90 degrees C. In some embodiments, blockers have a Tm of 83 degrees C. to 90 degrees C. In some embodiments, blockers have a Tm of 84 degrees C. to 90 degrees C. In some embodiments, a set of blockers has an average Tm of 78 degrees C. to 90 degrees C. In some embodiments, a set of blockers has an average Tm of 80 degrees C. to 90 degrees C. In some embodiments, a set of blockers has an average Tm of at least 80 degrees C. In some embodiments, a set of blockers have an average Tm of at least 81 degrees C. In some embodiments, a set of blockers have an average Tm of at least 82 degrees C. In some embodiments, a set of blockers have an average Tm of at least 83 degrees C. In some embodiments, a set of blockers has an average Tm of at least 84 degrees C. In some embodiments, a set of blockers have an average Tm of at least 86 degrees C. In some embodiments, blocker Tm are modified as a result of other components described herein, such as use of a fast hybridization buffer and / or hybridization enhancer.

[0278] The molar ratio of blockers to adapter targets may influence the off-bait (and subsequently off-target) rates during hybridization. The more efficient a blocker is at binding to the target adapter; the less blocker is required. In some embodiments, the blockers described herein achieve sequencing outcomes of no more than 20% off-target reads with a molar ratio of less than 20: 1 (blocker:target). In some embodiments, no more than 20% off-target reads are achieved with a molar ratio of less than 10: 1 (blocker:target). In some embodiments, no more than 20% off-target reads are achieved with a molar ratio of less than 5:1 (blocker:target). In some embodiments, no more than 20% off-target reads are achieved with a molar ratio of less than 2: 1 (blocker:target). In some embodiments, no more than 20% off-target reads are achieved with a molar ratio of less than 1.5:1 (blocker:target). In some embodiments, no more than 20% off-target reads are achieved with a molar ratio of less than 1.2: 1 (blocker:target). In some embodiments, no more than 20% off-target reads are achieved with a molar ratio of less than 1.05:1 (blocker:target).

[0279] Universal blockers may be used with unenriched libraries of varying sizes. In some embodiments, the unenriched libraries comprise at least or about 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, 1.0, 2.0, 4.0, 8.0, 10.0, 12.0, 14.0, 16.0, 18.0, 20.0, 22.0, 24.0, 26.0, 28.0, 30.0, 40.0, 50.0, 60.0, or more than 60.0 megabases (Mb).

[9280] Blockers as described herein may improve on-target performance. In some embodiments, on-target performance is improved by at least or about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more than 95%. In some embodiments, the on-target performance is improved by at least or about 5%, 10%, 15%, 20%, 25%,30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more than 95% for various index designs. In some embodiments, the on-target performance is improved by at least or about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more than 95% is improved for various panel sizes. h. Hybridization buffers

[0028] ] Any number of buffers may be used with the hybridization methods described herein. For example, a buffer comprises numerous chemical components, such as polymers, solvents, salts, surfactants, or other components. In some embodiments, hybridization buffers decrease the hybridization times (e.g., “fast” hybridization buffers) required to achieve a given sequencing result or level of quality. In some embodiments, such components lead to improved hybridization outcomes, such as increased on-target rate, improved sequencing outcomes (e.g., sequencing depth or other metric), or decreased off-target rates. Such components may be introduced at any concentration to achieve such outcomes. In some embodiments, buffer components are added in specific order. For example, water is added first. In some embodiments, salts are added after water. In some embodiments, salts are added after thickening agents and surfactants. In some embodiments, hybridization buffers such as “fast” hybridization buffers described herein are used in conjunction with universal blockers and liquid polymer additives. In some embodiments, use of fast hybridization buffers reduces hybridization times to no more than 4, 3, 2, 1, 0.5, 0.2, or 0.1 hours.

[0282] Hybridization buffers described herein may comprise solvents, or mixtures of two or more solvents. In some embodiments, a hybridization buffer comprises a mixture of two solvents, three solvents or more than three solvents. In some embodiments, a hybridization buffer comprises a mixture of alcohol and water. In some embodiments, a hybridization buffer comprises a mixture of a ketone containing solvent and water. In some embodiments, a hybridization buffer comprises a mixture of an ethereal solvent and water. In some embodiments, a hybridization buffer comprises a mixture of a sulfoxide-containing solvent and water. In some embodiments, a hybridization buffer comprises a mixture of an amide-containing solvent and water. In some embodiments, a hybridization buffer comprises a mixture of an ester-containing solvent and water. In some embodiments, hybridization buffers comprise solvents such as water, ethanol, methanol, propanol, butanol, other alcohol solvents, or a mixture thereof. In some embodiments, hybridization buffers comprise solvents such as acetone, methyl ethyl ketone, 2-butanone, ethyl acetate, methyl acetate, tetrahydrofuran, diethyl ether, or a mixture thereof. In some embodiments, hybridization buffers comprise solvents such as DMSO, DMF, DMA, HMPA, or a mixture thereof. In some embodiments, hybridization buffers comprise a mixture of water, HMPA, and alcohol. In some embodiments, two solvents are present at a 1:1, 1:2, 1:3, 1:4, 1:5, 1:8, 1:9, 1:10, 1:20, 1:50, 1:100, or 1:500 ratio.

[0283] Hybridization buffers described herein may comprise polymers. Polymers include but are not limited to thickening agents, polymeric solvents, dielectric materials, or other polymers. In some embodiments, polymers are hydrophobic or hydrophilic. In some embodiments, polymers are silicon polymers. In some embodiments, polymers comprise repeating polyethylene or polypropylene units, or a mixture thereof. In some embodiments, polymers comprise polyvinylpyrrolidone or polyvinylpyridine. In some embodiments, polymers comprise amino acids. For example, in some embodiments, polymers comprise proteins. In some embodiments, polymers comprise casein, milk proteins, bovine serum albumin, or other proteins. In some embodiments, polymers comprise nucleotides, for example, DNA or RNA. In some embodiments, polymers comprise polyA, polyT, Cot-1 DNA, or other nucleic acids. In some embodiments, polymers comprise sugars. For example, in some embodiments, a polymer comprises glucose, arabinose, galactose, mannose, or other sugar. In some embodiments, a polymer comprises cellulose or starch. In some embodiments, a polymer comprises agar, carboxyalkyl cellulose, xanthan, guar gum, locust bean gum, gum karaya, gum tragacanth, gum Arabic. In some embodiments, a polymer comprises a derivative of cellulose or starch, or nitrocellulose, dextran, hydroxyethyl starch, ficoll, or a combination thereof. In some embodiments, mixtures of polymers are used in hybridization buffers described herein. In some embodiments, hybridization buffers comprise Denhardt's solution. Polymers described herein may be present at any concentration suitable for reducing off-target binding. Such concentrations are often represented as a percent by weight, percent by volume, or percent weight per volume. For example, a polymer is present at about 0.0001%, 0.0002%, 0.0005%, 0.0008%, 0.001%, 0.002%, 0.005%, 0.008%, 0.01%, 0.02%, 0.05%, 0.08%, 0.1%, 0.2%, 0.5%, 0.8%, 1%, 1.2%, 1.5%, 1.8%, 2%, 5%, 10%, 20%, or about 30%. In some embodiments, a polymer is present at no more than 0.0001%, 0.0002%, 0.0005%, 0.0008%, 0.001%, 0.002%, 0.005%, 0.008%, 0.01%, 0.02%, 0.05%, 0.08%, 0.1%, 0.2%, 0.5%, 0.8%, 1%, 1.2%, 1.5%, 1.8%, 2%, 5%, 10%, 20%, or no more than 30%. In some embodiments, a polymer is present in at least 0.0001%, 0.0002%, 0.0005%, 0.0008%, 0.001%, 0.002%, 0.005%, 0.008%, 0.01%, 0.02%, 0.05%, 0.08%, 0.1%, 0.2%, 0.5%, 0.8%, 1%, 1.2%, 1.5%, 1.8%, 2%, 5%, 10%, 20%, or at least 30%. In some embodiments, a polymer is present at 0.0001%- 10%, 0.0002%-5%, 0.0005%-1.5%, 0.0008%-l%, 0.001%-0.2%, 0.002%-0.08%, 0.005%-0.02%, or 0.008%-0.05%. In some embodiments, a polymer is present at 0.005%-0.1%. In some embodiments, a polymer is present at 0.05%-0.1%. In some embodiments, a polymer is present at 0.005%-0.6%. In some embodiments, a polymer is present at l%-30%, 5%-25%, 10%-30%, 15%-30%, or 1%-15%.Liquid polymers may be present as a percentage of the total reaction volume. In some embodiments, a polymer is about 10%, 20%, 30%, 40%, 50%, 60%, 75%, or about 90% of the total volume. In some embodiments, a polymer is at least 10%, 20%, 30%, 40%, 50%, 60%, 75%, or at least 90% of the total volume. In some embodiments, a polymer is no more than 10%, 20%, 30%, 40%, 50%, 60%, 75%, or no more than 90% of the total volume. In some embodiments, a polymer is 5%-75%, 5%-65%, 5%- 55%, 10%-50%, 15%-40%, 20%-50%, 20%-30%, 25%-35%, 5%-35%, 10%-35%, or 20%-40% of thetotal volume. In some embodiments, a polymer is 25%-45% of the total volume. In some embodiments, hybridization buffers described herein are used in conjunction with universal blockers and liquid polymer additives.

[0284] Hybridization buffers described herein may comprise salts such as cations or anions. For example, hybridization buffer comprises a monovalent or divalent cation. In some embodiments, a hybridization buffer comprises a monovalent or divalent anion. In some embodiments, cations comprise sodium, potassium, magnesium, lithium, tris, or other salt. In some embodiments, anions comprise sulfate, bisulfate, hydrogensulfate, nitrate, chloride, bromide, citrate, ethylenediaminetetraacetate, dihydrogenphosphate, hydrogenphosphate, or phosphate. In some embodiments, hybridization buffers comprise salts comprising any combination of anions and cations (e.g. sodium chloride, sodium sulfate, potassium phosphate, or other salt). In some embodiments, a hybridization buffer comprises an ionic liquid. Salts described herein may be present at any concentration suitable for reducing off-target binding. Such concentrations are often represented as a percent by weight, percent by volume, or percent weight per volume. For example, salt is present at about 0.0001%, 0.0002%, 0.0005%, 0.0008%, 0.001%, 0.002%, 0.005%, 0.008%, 0.01%, 0.02%, 0.05%, 0.08%, 0.1%, 0.2%, 0.5%, 0.8%, 1%, 1.2%, 1.5%, 1.8%, 2%, 5%, 10%, 20%, or about 30%. In some embodiments, a salt is present at no more than 0.0001%, 0.0002%, 0.0005%, 0.0008%, 0.001%, 0.002%, 0.005%, 0.008%, 0.01%, 0.02%, 0.05%, 0.08%, 0.1%, 0.2%, 0.5%, 0.8%, 1%, 1.2%, 1.5%, 1.8%, 2%, 5%, 10%, 20%, or no more than 30%. In some embodiments, a salt is present in at least 0.0001%, 0.0002%, 0.0005%, 0.0008%, 0.001%, 0.002%, 0.005%, 0.008%, 0.01%, 0.02%, 0.05%, 0.08%, 0.1%, 0.2%, 0.5%, 0.8%, 1%, 1.2%, 1.5%, 1.8%, 2%, 5%, 10%, 20%, or at least 30%. In some embodiments, a salt is present at 0.0001%-10%, 0.0002%-5%, O.OOO5%-1.5%, 0.0008%-l%, 0.001%-0.2%, 0.002%-0.08%, 0.005%-0.02%, or 0.008%-0.05%. In some embodiments, salt is present at 0.005%-0.1%. In some embodiments, salt is present at 0.05%-0.1%. In some embodiments, salt is present at 0.005%-0.6%. In some embodiments, salt is present at l%-30%, 5%-25%, 10%-30%, 15%-3O%, or 1%- 15%. Liquid polymers may be present as a percentage of the total reaction volume. In some embodiments, salt is about 10%, 20%, 30%, 40%, 50%, 60%, 75%, or about 90% of the total volume. In some embodiments, a salt is at least 10%, 20%, 30%, 40%, 50%, 60%, 75%, or at least 90% of the total volume. In some embodiments, salt is no more than 10%, 20%, 30%, 40%, 50%, 60%, 75%, or no more than 90% of the total volume. In some embodiments, salt is 5%-75%, 5%- 65%, 5%-55%, 10%-50%, 15%-40%, 20%-50%, 20%-30%, 25%-35%, 5%-35%, 10%-35%, or 20%- 40% of the total volume. In some embodiments, salt is 25%-45% of the total volume.

[0285] Hybridization buffers described herein may comprise surfactants (or emulsifiers). For example, a hybridization buffer comprises SDS (sodium dodecyl sulfate), CTAB, cetylpyridinium, benzalkonium tergitol, fatty acid sulfonates (e.g., sodium lauryl sulfate), ethyloxylated propylene glycol, lignin sulfonates, benzene sulfonate, lecithin, phospholipids, dialkyl sulfosuccinates (e.g., dioctyl sodium sulfosuccinate), glycerol diester, polyethoxylated octyl phenol, abietic acid, sorbitanmonoester, perfluoro alkanols, sulfonated polystyrene, betaines, dimethyl polysiloxanes, or other surfactant. In some embodiments, a hybridization buffer comprises a sulfate, phosphate, or tetralkyl ammonium group. Surfactants described herein may be present at any concentration suitable for reducing off-target binding. Such concentrations are often represented as a percent by weight, percent by volume, or percent weight per volume. For example, a surfactant is present at about 0.0001%, 0.0002%, 0.0005%, 0.0008%, 0.001%, 0.002%, 0.005%, 0.008%, 0.01%, 0.02%, 0.05%, 0.08%, 0.1%, 0.2%, 0.5%, 0.8%, 1%, 1.2%, 1.5%, 1.8%, 2%, 5%, 10%, 20%, or about 30%. In some embodiments, a surfactant is present at no more than 0.0001%, 0.0002%, 0.0005%, 0.0008%, 0.001%, 0.002%, 0.005%, 0.008%, 0.01%, 0.02%, 0.05%, 0.08%, 0.1%, 0.2%, 0.5%, 0.8%, 1%, 1.2%, 1.5%, 1.8%, 2%, 5%, 10%, 20%, or no more than 30%. In some embodiments, a surfactant is present in at least 0.0001%, 0.0002%, 0.0005%, 0.0008%, 0.001%, 0.002%, 0.005%, 0.008%, 0.01%, 0.02%, 0.05%, 0.08%, 0.1%, 0.2%, 0.5%, 0.8%, 1%, 1.2%, 1.5%, 1.8%, 2%, 5%, 10%, 20%, or at least 30%. In some embodiments, a surfactant is present at0.0001%-10%, 0.0002%-5%, O.OOO5%-1.5%, 0.0008%- 1%, 0.001%-0.2%, 0.002%-0.08%, 0.005%-0.02%, or 0.008%-0.05%. In some embodiments, a surfactant is present at 0.005%-0.1%. In some embodiments, a surfactant is present at 0.05%-0.1%. In some embodiments, a surfactant is present at 0.005%-0.6%. In some embodiments, a surfactant is present at l%-30%, 5%-25%, 10%-30%, 15%-3O%, or 1%-15%. Liquid polymers may be present as a percentage of the total reaction volume. In some embodiments, a surfactant is about 10%, 20%, 30%, 40%, 50%, 60%, 75%, or about 90% of the total volume. In some embodiments, a surfactant is at least 10%, 20%, 30%, 40%, 50%, 60%, 75%, or at least 90% of the total volume. In some embodiments, a surfactant is no more than 10%, 20%, 30%, 40%, 50%, 60%, 75%, or no more than 90% of the total volume. In...

Claims

CLAIMS What is claimed is:

1. A method of characterizing metagenomes of at least one sample obtained from at least one subject suspected of having a microbial infection, comprising: (a) preparing at least four total unenriched libraries for metagenomic next generation sequencing from total nucleic acids extracted from at least one sample obtained from at least one subject suspected of having a microbial infection, wherein each unenriched library comprises: (i) a first control unenriched nucleic acid library comprising nucleic acids of known microbial origin; (i) a second control unenriched nucleic acid library comprising nucleic acids of human origin obtained from at least one human subject who is not suspected of having the microbial infection; (ii) a third control unenriched library that is substantialy free of nucleic acids; or (iv) at least one test unenriched nucleic acid library comprising nucleic acids of human origin and nucleic acids of unknown microbial origin prepared from the at least one sample; wherein each unenriched nucleic acid library comprises fragmented double-stranded DNA that is barcoded for indexing; (b) partitioning the unenriched libraries, on an equimolar basis, into a first set of pools of unenriched libraries and a second set of pools of unenriched libraries; (c) enriching each of the unenriched libraries in the first set of pools for nucleic acids of microbial origin to produce a first set of pools of enriched libraries that are enriched for nucleic acids of microbial origin; (d) combining a volume from each pool in the first set of pools of enriched libraries, in a predetermined volume ratio, with a volume from each corresponding pool in the second set of pools of unenriched libraries into a single flow cel for next generation sequencing; (e) sequencing the libraries in the combined volumes of each coresponding pool in the single flow cel using a next generation sequencing platform; and (f) characterizing the metagenomes of the at least one sample obtained from the at least one subject suspected of having the microbial infection based on the sequencing results.

2. The method of claim 1, wherein step (a) comprises preparing at least five, at least six, at least eight, at least 12, at least 24, at least 48, at least 96, at least 384, or at least 1536 unenriched libraries.

3. The method of claim 2 or 3, wherein the at least one sample comprises at least 2 samples, at least 3 samples, at least 5 samples, at least 6 samples, at least 9 samples, at least 12 samples, at least 18samples, at least 21 samples, at least 24 samples, at least 30 samples, at least 36, at least 42 samples, at least 45 samples, at least 72 samples, at least 78 samples, at least 84 samples, at least 90 samples, at least 312 samples, at least 336 samples, at least 348 samples, at least 354 samples, at least 366 samples, at least 1152 samples, at least 1248 samples, at least 1344 samples, at least 1392 samples, or at least 1440 samples obtained from subjects suspected of having the microbial infection.

4. The method of any one of claims 1-3, wherein the each of the first set of pools and the second set of pools comprises at least one pool, at least two pools, at least three pools, at least four pools, at least six pools, at least eight pools, at least 12 pools, at least 16 pools, at least 24 pools, at least 32 pools, at least 48 pools, at least 64 pools, at least 96 pools, or at least 128 pools of unenriched libraries.

5. The method of claim 4, wherein: (i) each set of the at least one pool comprises at least 4, at least 5, at least 6, at least 12, at least 24, or at least 48 unenriched libraries; (i) each set of the at least two pools comprises at least 12, at least 24, at least 48, or at least 96 unenriched libraries; (ii) each set of the at least three pools comprises at least 12, at least 24, at least 48, or at least 96 unenriched libraries; (iv) each set of the at least three pools comprises at least 12 nucleic acid libraries; (v) each set of the at least four pools comprises at least 24, at least 48, or at least 96 unenriched libraries; (vi) each set of the at least six pools comprises at least 24, at least 48, at least 96, or at least 384 unenriched libraries; (vi) each set of the at least eight pools comprises at least 24, at least 48, at least 96, or at least 384 unenriched libraries; (vii) each set of the at least 12 pools comprises at least 48, at least 96, or at least 384 unenriched libraries; (ix) each set of the at least 16 pools comprises at least 384 unenriched libraries; (x) each set of the at least 24 pools comprises at least 384 unenriched libraries; (xi) each set of the at least 32 pools comprises at least 1536 unenriched libraries; (xi) each set of the at least 48 pools comprises at least 1536 unenriched libraries; (xii) each set of the at least 64 pools comprises at least 1536 unenriched libraries; (xiv) each set of the at least 96 pools comprises at least 1536 unenriched libraries; or (xv) each set of the at least 128 pools comprises at least 1536 unenriched libraries.

6. The method of claim 5, wherein each set of pools comprises: (i) a first control unenriched nucleic acid library comprising nucleic acids of known microbial origin;(i) a second control unenriched nucleic acid library comprising nucleic acids of human origin obtained from at least one human subject who is not suspected of having the microbial infection; (ii) a third control unenriched library that is substantialy free of nucleic acids; and (iv) at least X test unenriched nucleic acid libraries prepared from at least a Y samples obtained from at least Y subjects suspected of having the microbial infection; wherein X is equal to a variable number Z that is three less than the total number of unenriched libraries prepared divided by the number of pools in each set, wherein Y is equal to the number of pools in each set times Z, wherein each test unenriched nucleic acid library comprises nucleic acids of human origin and nucleic acids of unknown microbial, and wherein each unenriched nucleic acid library comprises fragmented double-stranded DNA that is barcoded for indexing.

7. The method of any one of claims 1-6, further comprising, prior to step (a): (i) extracting total nucleic acids from at least one sample of at least one microbe to prepare the first control unenriched nucleic acid library; (i) extracting total nucleic acids from at least one sample obtained from at least one human subject who is not suspected of having the microbial infection to prepare the second control unenriched nucleic acid library; and / or (ii) extracting total nucleic acids from the at least one sample obtained from the at least one subject suspected of having the microbial infection to prepare the at least one test unenriched nucleic acid library.

8. The method of any one of claims 1-7, where step (a) comprises: preparing at least eight total unenriched libraries for metagenomic next generation sequencing from total nucleic acids extracted from at least five samples obtained from at least five subjects suspected of having a microbial infection, wherein the at least eight total unenriched libraries comprise: (i) a first control unenriched nucleic acid library comprising nucleic acids of known microbial origin; (i) a second control unenriched nucleic acid library comprising nucleic acids of human origin obtained from at least one human subject who is not suspected of having the microbial infection; (ii) a third control unenriched library that is substantialy free of nucleic acids; and (iv) at least five test unenriched nucleic acid libraries each comprising nucleic acids of human origin and nucleic acids of unknown microbial origin prepared from the at least five samples,wherein each unenriched nucleic acid library comprises fragmented double-stranded DNA that is barcoded for indexing.

9. The method of any one of claims 1-8, wherein step (b) comprises: partitioning the at least eight total unenriched libraries, on an equimolar basis, into a first set of one pool of unenriched libraries and a second set of one pool of unenriched libraries, wherein the one pool of unenriched libraries in each of the first and second sets comprises an aliquot taken from: (i) the first control unenriched nucleic acid library; (i) the second control unenriched nucleic acid library; (ii) the third control unenriched library; and (iv) the at least five test unenriched nucleic acid libraries.

10. The method of any one of claims 1-9, wherein step (a) comprises: preparing at least 48 total unenriched libraries for metagenomic next generation sequencing from total nucleic acids extracted from at least 42 samples obtained from at least 42 subjects suspected of having a microbial infection, wherein the at least 48 total unenriched libraries comprise: (i) at least two first control unenriched nucleic acid libraries comprising nucleic acids of known microbial origin; (i) at least two second control unenriched nucleic acid libraries comprising nucleic acids of human origin obtained from at least one human subject who is not suspected of having the microbial infection; (ii) at least two third control unenriched libraries that are substantialy free of nucleic acids; and (iv) at least 42 test unenriched nucleic acid libraries each comprising nucleic acids of human origin and nucleic acids of unknown microbial origin prepared from the at least 42 samples, wherein each unenriched nucleic acid library comprises fragmented double-stranded DNA that is barcoded for indexing.

11. The method of any one of claims 1-10, wherein step (b) comprises: partitioning the 48 total unenriched libraries, on an equimolar basis, into a first set of two pools of unenriched libraries and a second set of two pools of unenriched libraries, wherein each pool in the first and second sets of pools of unenriched libraries comprises an aliquot taken from: (i) a diferent one of the at least two first control nucleic acid libraries; (i) a different one of the at least two second control libraries; (ii) a diferent one of the at least two third control libraries; and (iv) at least 21 diferent test libraries from the at least 42 test libraries.

12. The method of any one of claims 1-11, wherein step (a) comprises: preparing at least 96 total unenriched libraries for metagenomic next generation sequencing from total nucleic acids extracted from at least 84 samples obtained from at least 84 subjects suspected of having a microbial infection, wherein the at least 96 total unenriched libraries comprise: (i) at least four first control unenriched nucleic acid libraries prepared from total nucleic acids extracted from nucleic acids of known microbial origin; (i) at least four second control unenriched nucleic acid libraries comprising nucleic acids of human origin obtained from at least one human subject who is not suspected of having the microbial infection; (ii) at least four third control unenriched libraries that are substantialy free of nucleic acids; and (iv) at least 84 test unenriched nucleic acid libraries each comprising nucleic acids of human origin and nucleic acids of unknown microbial origin prepared from the at least 84 samples, wherein each unenriched nucleic acid library comprises fragmented double-stranded DNA that is barcoded for indexing.

13. The method of any one of claims 1-12, wherein step (b) comprises partitioning the 96 total unenriched libraries, on an equimolar basis, into a first set of four pools of unenriched libraries and a second set of four pools of unenriched libraries, wherein each pool in the first and second sets of four pools of unenriched libraries comprises an aliquot taken from: (i) a diferent one of the at least four first control nucleic acid libraries; (i) a different one of the at least four second control libraries; (ii) a diferent one of the at least four third control libraries; and (iv) at least 21 diferent test libraries from the at least 84 test libraries.

14. The method of any one of claims 1-13, wherein enriching each of the unenriched libraries comprises enriching nucleic acids of unknown microbial origin from at least one microbe of interest in the libraries in each pool of the first set of pools.

15. The method of claim 14, further comprising diluting each pool of enriched nucleic acid libraries.

16. The method of claim 15, wherein each pool is diluted to between about .20 nM to about 20 nm, about 0.30 nM to about 19 nM, about 0.4 nM to about 18 nM, about 0.5 nm to about 17 nM, about 0.6 nm to about 16 nm, about 0.7 to about 15 nM, about 0.8 to about 14 nM, about 0.9 to about 13 nM,about 1 nM to about 12 nM, about 1.1 nM to about 11 nM, about 1.2 nM to about 10 nM, about 1.3 nM to about 9 nM, about 1.4 nM to about 8 nM, about 1.5 nM to about 7 nM, about 1.6 nM to about 6 nM, about 1.7 nM to about 5 nM, about 1.8 nM to about 4 nM, about 1.9 nM to about 3 nM, or about 2 nM.

17. The method of any one of claims 9-16, wherein each aliquot taken from each unenriched library in the second set of pools is diluted.

18. The method of claim 17, wherein each aliquot is diluted to between about .20 nM to about 20 nm, about 0.30 nM to about 19 nM, about 0.4 nM to about 18 nM, about 0.5 nm to about 17 nM, about 0.6 nm to about 16 nm, about 0.7 to about 15 nM, about 0.8 to about 14 nM, about 0.9 to about 13 nM, about 1 nM to about 12 nM, about 1.1 nM to about 11 nM, about 1.2 nM to about 10 nM, about 1.3 nM to about 9 nM, about 1.4 nM to about 8 nM, about 1.5 nM to about 7 nM, about 1.6 nM to about 6 nM, about 1.7 nM to about 5 nM, about 1.8 nM to about 4 nM, about 1.9 nM to about 3 nM, or about 2 nM.

19. The method of any one of claims 1-18, wherein the predetermined volume ratio in microliters is 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 2:8, 3:7, 4:6, 9:1, 8:1, 7:1, 6:1, 5:1, 4:1, 3:1, 2:1, 8:2, 7:3, or 6:

4.

20. The method of claim 19, wherein the volume ratio in microliters is 1:

9.

21. The method of claims 19, wherein the volume ratio in microliters is 2:

8.

22. The method of any one of claims 19-21, wherein between about 0.10 microliters and about 10 microliters from each pool in the first set of pools is combined with between about 0.10 microliters and about 10 microliters from each corresponding pool in the second set of pools.

23. The method of any one of claims 19-21, wherein about 1 microliter or about 2 microliters from each pool in the first set of pools is combined with about 8 or about 9 microliters from each corresponding pool in the second set of pools.

24. The method of any one of claims 19-21, wherein about 1 microliter from each pool in the first set of pools is combined with about 9 microliters from each corresponding pool in the second set of pools.

25. The method of any one of claims 19-21, wherein about 2 microliters from each pool in the first set of pools is combined with about 8 microliters from each corresponding pool in the second set of pools.

26. The method of any one of claims 1-25, wherein the at least one subject is suspected of having a viral infection.

27. The method of any one of claims 1-26, wherein the at least one subject has or is suspected of having acute febrile ilness (AFI).

28. The method of any one of claims 1-27, wherein the at least one subject has or is suspected of having a severe acute respiratory ilness.

29. The method of any one of claims 1-28, wherein the nucleic acids of known microbial origin are nucleic acids of known viral origin.

30. The method of any one of claims 1-29, wherein the nucleic acids of human origin are obtained from at least one human subject who does not have or is not believed to have the viral infection.

31. The method of any one of claims 1-30, wherein the nucleic acids of unknown microbial origin comprise nucleic acids of viral origin.

32. The method of any one of claims 29-31, wherein the viral infection or viral origin is a virus selected from the group consisting of a virus from an order listed in Table 1, a virus family listed in Table 2, a virus from genus listed in Table 3, a virus from a species listed in Table 4, and a virus strain listed in Table 5.

33. The method of any one of claims 1-32, wherein the at least one sample is a biological sample.

34. The method of any one of claims 1-33, wherein the at least one sample is selected from the group consisting of bronchoalveolar lavage, cerebral spinal fluid, plasma, serum, sputum, urine, whole blood, stool, and a swab.

35. The method of any one of claims 1-34, further comprising detecting a viral infection in the at least one subject based on the next generation sequencing results.

36. The method of any one of claims 1-35, further comprising treating the at least one subject for the viral infection or a symptom, condition, disease or disorder caused by the viral infection.

Citation Information

Patent Citations

  • Compositions and methods for detecting viruses in a sample

    US10597736B2

  • Viral database methods

    US20090105092A1

  • Process for amplifying, detecting, and / or-cloning nucleic acid sequences

    US4683195A

  • Process for amplifying nucleic acid sequences

    US4683202A

  • Bacterial capture sequencing platform and methods of designing, constructing and using

    WO2019226992A1