Methods and systems for quantifying targeted enrichment efficiency

The method addresses the inefficiency in quantifying nucleic acid enrichment by using biotinylated oligonucleotides and control polynucleotides to determine enrichment efficiency, providing a rapid and cost-effective solution for improved quality control in molecular biology.

WO2025244841A1PCT designated stage Publication Date: 2025-11-27ILLUMINA INC
View PDF 15 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/028156
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-23
Filing Date
2025-05-07
Publication Date
2025-11-27

AI Technical Summary

Technical Problem

Existing methods for quantifying the efficiency of nucleic acid enrichment are inefficient, particularly in the detection and quantification of specific nucleic acid sequences, especially in the field of molecular biology and biochemistry, where existing methods lack a standard metric or method for determining enrichment efficiency.

Method used

A method for quantifying the efficiency of nucleic acid enrichment using biotinylated oligonucleotides, which are designed to enrich specific nucleic acid sequences by hybridizing with homologous sequences in the sample, and using control polynucleotides to determine the efficiency of enrichment by comparing targeted and untargeted regions.

Benefits of technology

The method provides a rapid, cost-effective, and flexible way to quantify the efficiency of nucleic acid enrichment, allowing for improved quality control and reliable reporting of results from enrichment assays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025028156_27112025_PF_FP_ABST
    Figure US2025028156_27112025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments include methods and systems for quantifying the efficiency of enriching for a target nucleic acid in a nucleic acid sample. Some embodiments include methods and electronic systems for implementing such methods. Some embodiments include adding control polynucleotides to a nucleic acid sample; hybridizing a probe set to the control polynucleotides and nucleic acid sample; amplifying or separating any hybridized nucleic acids, thereby providing an enriched library; sequencing the enriched library; and quantifying the efficiency of the enrichment for the targeted regions based on the proportion of targeted control sequence reads to untargeted control sequence reads.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND SYSTEMS FOR QUANTIFYING TARGETED ENRICHMENTEFFICIENCYCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Prov. App. No. 63 / 651108 filed May 23, 2024 entitled “METHODS AND SYSTEMS FOR QUANTIFYING TARGETED ENRICHMENT EFFICIENCY” which is incorporated by reference in its entirety.REFERENCE TO SEQUENCE LISTING

[0002] The present application is filed with a Sequence Listing in Electronic format. The Sequence Listing is provided as a file entitled “SEQLISTFNG ILLINC836WO”, created April 30, 2025, which is 435,694 bytes in size. The information in the electronic format of the sequence listing is incorporated herein by reference in its entirety.BACKGROUNDField

[0003] The present disclosure relates to methods and systems for quantifying the efficiency of enriching for a target nucleic acid in a nucleic acid sample.Description

[0004] Target-capture enrichment sequencing is a method for detecting nucleic acid sequences present in low concentrations, generally from environmental and clinical samples. [Koehler et al. Development and evaluation of a panel of filovirus sequence capture probes for pathogen detection by next-generation sequencing. PLoS One. 2014 Sep 10;9(9): el07007]. Target-capture enrichment uses biotinylated oligonucleotides (“bait probes”) which can be bound to avidin and which are designed to enrich the signal of pre-selected low concentration nucleic acid sequences. By hybridizing with homologous sequences in the sample, the bait probes allow the targeted sequences to be separated from untargeted nucleic acids. [Lovett et al. Direct selection: a method for the isolation of cDNAs encoded by large genomic regions. Proc Natl Acad Sci USA. 1991 Nov l;88(21):9628-32]. An enrichment step may be added into a conventional sequencing library preparation process to increase thespecificity and sensitivity for detecting rare gene variants and pathogenic organisms in complex communities.SUMMARY

[0005] Disclosed herein are methods for quantifying the efficiency of enriching for a target nucleic acid in a nucleic acid sample. In some embodiments, the method includes (a) adding control polynucleotides to a nucleic acid sample comprising sample nucleic acids; (b) providing a probe set comprising probes complementary to targeted regions of the sample nucleic acids, and probes complementary to targeted regions of the control polynucleotides; (c) hybridizing the probes to the targeted regions, and amplifying or separating any hybridized nucleic acids from the nucleic acid sample, thereby providing an enriched library; (d) sequencing the enriched library, thereby providing sequence reads from targeted regions of sample nucleic acids, targeted control sequence reads from targeted regions of control polynucleotides, and untargeted control sequence reads from untargeted regions of control polynucleotides; and (e) quantifying the efficiency of the enrichment for the targeted regions based on the proportion of targeted control sequence reads to untargeted control sequence reads.

[0006] In some embodiments, the control polynucleotides comprise polynucleotides from a microorganism. In some embodiments, the control polynucleotides comprise polynucleotides from Enterobacteria phage T7, Escherichia virus T4, Allobacillus halotolerans, Escherichia virus MS2, Escherichia virus Qbeta, Imtechella halotolerans, Phocid alphaherpesvirus 1, Phocine morbillivirus, or Truepera radiovictrix. In some embodiments, the probes complementary to targeted regions of the control polynucleotides comprise any one of SEQ ID NOs: 1-462.

[0007] In some embodiments, quantifying targeted enrichment efficiency comprises normalizing a count of targeted control sequence reads by a sequence length of targeted regions of the control polynucleotides, or normalizing a count of untargeted control sequence reads by a sequence length of untargeted regions of the control polynucleotides.

[0008] In some embodiments, quantifying targeted enrichment efficiency comprises binning sequence reads from targeted regions of control polynucleotides, and binning sequence reads from untargeted regions of control polynucleotides. In someembodiments, the method includes binning sequence reads by taxonomic kingdom of origin. In some embodiments, targeted enrichment efficiency is quantified based on binned sequence reads. In some embodiments, quantifying targeted enrichment efficiency is based on a count of binned sequence reads from targeted regions of control polynucleotides, and a count of binned sequence reads from untargeted regions of control polynucleotides. In some embodiments, quantifying targeted enrichment efficiency is based on an average sequence read length.

[0009] In some embodiments, the method comprises aligning sequence reads to a reference sequence, and wherein quantifying targeted enrichment efficiency is based on the alignment of sequence reads to the reference sequence. In some embodiments, quantifying targeted enrichment efficiency is based on a depth of aligned sequence reads at each position in the reference sequence. In some embodiments, quantifying targeted enrichment efficiency is based on a median or mean depth of aligned sequence reads at each position in the reference sequence.

[0010] In some embodiments, the method comprises separating any hybridized nucleic acids from the nucleic acid sample.

[0011] In some embodiments, the method comprises: providing a probe set comprising at least two probes complementary to one or more target nucleic acids, wherein the probes are affixed to a support; capturing the one or more target nucleic acids on the support; using the one or more captured target nucleic acids as a template strand to produce one or more nucleic acid duplexes immobilized on the support, wherein the one or more target nucleic acids hybridize to one or more probes of the probe set on the support; tagmenting and extending to produce one or more tagged nucleic acid duplexes; amplifying the one or more tagged nucleic acid duplexes to produce a plurality of tagged nucleic acid strands; contacting the plurality of tagged nucleic acid strands with a probe set to create an enriched library; and amplifying the enriched library.

[0012] In some embodiments, the nucleic acid sample comprises nucleic acids derived from a human. In some embodiments, the nucleic acid sample comprises nucleic acids derived from a microorganism. In some embodiments, the microorganism is a virus or a bacterium. In some embodiments, the nucleic acid sample comprises DNA or RNA.

[0013] Some embodiments also include modifying a method for the enriching for a target nucleic acid in a nucleic acid sample, wherein the modifying is based on the efficiency of the enrichment for the targeted regions. In some embodiments, the modifying comprises selecting a capture probe for the target nucleic acid; adjusting an initial concentration of a nucleic acid sample; adjusting a concentration of a reagent in a step of the method for the enriching; and / or removing a step in the method for the enriching.

[0014] Some embodiments also include discarding the nucleic acid sample based on the efficiency of the enrichment for the targeted regions.

[0015] Some embodiments also include analyzing the nucleic acid sample based on the efficiency of the enrichment for the targeted regions. In some embodiments, the analyzing comprises sequencing the nucleic acid sample.

[0016] Some embodiments also include providing a quality control metric to an end user for the enriching for a target nucleic acid in a nucleic acid sample based on the efficiency of the enrichment for the targeted regions.

[0017] Some embodiments include an electronic system for quantifying the efficiency of targeted enrichment sequencing, comprising a processor configured to perform any one of the foregoing methods.

[0018] Further disclosed herein are electronic systems for quantifying the efficiency of targeted enrichment sequencing. In some embodiments, the electronic system comprises a processor configured to perform a method comprising: receiving sequence reads from targeted regions of sample nucleic acids, targeted control sequence reads from targeted regions of control polynucleotides, and untargeted control sequence reads from untargeted regions of control polynucleotides; and quantifying targeted enrichment efficiency based on the proportion of targeted control sequence reads to untargeted control sequence reads.

[0019] Further disclosed herein are non-transitory computer-readable media. In some embodiments, the non-transitory computer-readable medium comprises a plurality of instructions, which when executed by at least one processor, cause the at least one processor to: receive sequence reads from targeted regions of sample nucleic acids, targeted control sequence reads from targeted regions of control polynucleotides, and untargeted control sequence reads from untargeted regions of control polynucleotides; and quantify targetedenrichment efficiency based on the proportion of targeted control sequence reads to untargeted control sequence reads.BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Features of examples of the present disclosure will become apparent by reference to the following detailed description and drawings, in which like reference numerals correspond to similar, though perhaps not identical, components. For the sake of brevity, reference numerals or features having a previously described function may or may not be described in connection with other drawings in which they appear. In addition to the features described herein, additional features and variations will be readily apparent from the following descriptions of the drawings and exemplary embodiments. It is to be understood that these drawings depict typical embodiments and are not intended to be limiting in scope.

[0021] FIG. 1 is a flow diagram that schematically illustrates a computer- implemented method for quantifying efficiency of targeted enrichment sequencing.

[0022] FIG. 2 is a flow diagram that schematically illustrates a process of quantifying targeted enrichment efficiency that may take place within the method of FIG. 1.

[0023] FIG. 3A is a block diagram of an exemplary sequencing system that may be used to perform the disclosed methods.

[0024] FIG. 3B is a block diagram of an exemplary computing device that may be used in connection with the exemplary sequencing system of FIG. 3 A.

[0025] FIG. 4 is a set of two histograms of the log 10 transformed enrichment factor values, with a first histogram showing the Binned Read Ratio and a second histogram showing the Aligned Mean Depth.DETAILED DESCRIPTION

[0026] The foregoing and other aspects of the present disclosure will now be described in more detail with respect to the description and methodologies provided herein. This description is not intended to be a detailed catalogue of all the ways in which the embodiments of the present disclosure may be implemented, or of all the features that may be added to the present disclosure. For example, features illustrated with respect to one embodiment may be incorporated into other embodiments, and features illustrated with respectto a particular embodiment may be deleted from that embodiment. In addition, numerous variations and additions to the various embodiments suggested herein, which do not depart from the instant disclosure, will be apparent to those skilled in the art in light of the instant detailed description, figures and claims. Hence, the following specification is intended to illustrate some particular embodiments, and not to exhaustively specify all permutations, combinations and variations thereof.

[0027] All patents, patent applications, and other publications, including all sequences disclosed within these references, referred to herein are expressly incorporated herein by reference, to the same extent as if each individual publication, patent or patent application was specifically and individually indicated to be incorporated by reference. All documents cited are, in relevant part, incorporated herein by reference in their entireties for the purposes indicated by the context of their citation herein. However, the citation of any document is not to be construed as an admission that it is prior art with respect to the present disclosure.

[0028] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, similar symbols typically identify similar components, unless context dictates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are explicitly contemplated herein.

[0029] Embodiments relate to systems and methods for determining the efficiency of a nucleic acid enrichment process. This ability to determine the efficiency of the enrichment process allows a user to minimize the concentration at which an organism or allele of interest can be detected and sequenced. Enrichment sequencing data can be analyzed in similar approaches as conventional metagenomic shotgun libraries for taxonomic classification, variant detection, and quality control. However, there is no standard metric or method for determining enrichment efficiency.

[0030] Enrichment efficiency is a measure of the increase in signal from targeted regions which have been subject to enrichment above any background sequences. This is typically done by preparing a conventional shotgun library (that has not undergone targetcapture) for each enriched library of nucleic acids. By comparing the amount of data derived from the enrichment library preparations as compared to the data from the conventional library preparations, the efficiency of the enrichment steps can be quantified. The quantification can be based on some aligned read count / fraction, an estimate of depth of coverage, the breadth of coverage, and / or the number and length of gaps in coverage. [Wylie et al. Enhanced virome sequencing using targeted sequence capture. Genome Res. 2015 Dec;25(12):1910-20], However, as this approach doubles the cost per sample, it is often done during development and validation of an assay and analyzed in an ad-hoc manner. [Grillova L, Cokelaer T, Mariet JF, da Fonseca JP, Picardeau M. Core genome sequencing and genotyping of Leptospira interrogans in clinical samples by target capture sequencing. BMC Infect Dis. 2023 Mar 14;23(1):157],

[0031] To make the process of quantification of targeted enrichment efficiency able to be performed with a single library preparation, enrichment probes can be designed to target regions of control polynucleotides. These control polynucleotides may then be added at a low fixed concentration to all samples, which avoids the possibility of complete signal loss in unenriched libraries and can speed up the processing time. Aspects of the invention include both a metric and a method for quantifying the efficiency of a target capture enrichment reaction that is simple, rapid, cost-effective, and flexible to implement. It may include specific features in probe design and data analysis that allow an enrichment factor to be calculated alongside normal enrichment data for every sample processed. For example, by designing enrichment probes which target regions of control polynucleotides added to a sample, depth of coverage (or sequence read count) of targeted regions may be compared with depth coverage (or sequence read count) of untargeted regions of control polynucleotides within a single sample. Some embodiments of the method described may use one of two approaches to calculate an enrichment factor, calculation based on depth of coverage via alignment to a reference sequence (alignment- based) or based on counts of binned sequence reads (alignment- free), which provides a trade-off between accuracy (alignment-based) and speed (alignment-free). The usage of the methods disclosed herein may allow for improved quality control and ensure reliable reporting of results from enrichment assays.Definitions

[0032] Although the following terms are believed to be well understood by one of skill in the art, the following definitions are set forth to facilitate understanding of the presently disclosed subject matter.

[0033] All technical and scientific terms used herein, unless otherwise defined below, are intended to have the same meaning as commonly understood by one of ordinary skill in the art. References to techniques employed herein are intended to refer to the techniques as commonly understood in the art, including variations on those techniques or substitutions of equivalent techniques that would be apparent to one of skill in the art.

[0034] As used herein, the terms “a” or “an” or “the” may refer to one or more than one. For example, “a” marker can mean one marker or a plurality of markers.

[0035] As used herein, the term “about,” when used in reference to a measurable value such as an amount of mass, dose, time, temperature, and the like, is meant to encompass variations of 20%, 10%, 5%, 1%, 0.5%, or even 0.1% of the specified amount.

[0036] As used herein, the term “and / or” refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative (“or”).

[0037] Throughout this specification, unless the context requires otherwise, the words “comprise,” “comprises,” and “comprising” will be understood to imply the inclusion of a stated step or element or group of steps or elements but not the exclusion of any other step or element or group of steps or elements.

[0038] As used herein, the term “consists essentially of’ (and grammatical variants thereof), as applied to the compositions and methods of the present disclosure, means that the compositions / methods may contain additional components so long as the additional components do not materially alter the composition / method.

[0039] The term “nucleic acid” or “polynucleotide” refers to a deoxyribonucleotide or ribonucleotide polymer in either single- or double-stranded form, and unless otherwise limited, encompasses known analogs of natural nucleotides that hybridize to nucleic acids inmanner similar to naturally occurring nucleotides, such as peptide nucleic acids (PNAs) and phosphorothioate DNA. Unless otherwise indicated, a particular nucleic acid sequence includes the complementary sequence thereof. Nucleotides include, but are not limited to, ATP, dATP, CTP, dCTP, GTP, dGTP, UTP, TTP, dUTP, 5-methyl-CTP, 5-methyl-dCTP, ITP, diTP, 2-amino-adenosine-TP, 2-amino-deoxyadenosine-TP, 2-thiothymidine triphosphate, pyrrolo-pyrimidine triphosphate, and 2-thiocytidine, as well as the alphathiotriphosphates for all of the above, and 2'-O-methyl-ribonucleotide triphosphates for all the above bases. Modified bases include, but are not limited to, 5-Br-UTP, 5-Br-dUTP, 5-F-UTP, 5-F-dUTP, 5-propynyl dCTP, and 5-propynyl-dUTP.

[0040] As used herein, the term “reference genome” or “reference sequence” refers to any particular known genome sequence, whether partial or complete, of any organism or virus which may be used to reference identified sequences from a subject. For example, a reference genome used for human subjects as well as many other organisms is found at the National Center for Biotechnology Information at ncbi.nlm.nih.gov. In various embodiments, the reference sequence is significantly larger than the reads that are aligned to it. For example, it may be at least about 100 times larger, or at least about 1000 times larger, or at least about 10,000 times larger, or at least about 105times larger, or at least about 106times larger, or at least about 107times larger. In one example, the reference sequence is that of a full-length genome. Such sequences may be referred to as genomic reference sequences. Other examples of reference sequences include genomes of other species, such as control organisms as disclosed herein, as well as chromosomes, sub-chromosomal regions (such as strands), etc., of any species. In various embodiments, the reference sequence is a consensus sequence or other combination derived from multiple individuals. However, in certain applications, the reference sequence may be taken from a particular individual.

[0041] The term “nucleic acid sample” herein may refer to a sample, typically derived from one or more biological fluids, cells, tissues, organs, or organisms, comprising a nucleic acid or a mixture of nucleic acids comprising at least one nucleic acid sequence that is to be screened for copy number variation. In certain embodiments, the nucleic acid sample comprises at least one nucleic acid sequence whose copy number is suspected of having undergone variation. Such samples may include, but are not limited to sputum / oral fluid, amniotic fluid, blood, a blood fraction, or fine needle biopsy samples (such as surgical biopsy,fine needle biopsy, etc.), urine, peritoneal fluid, pleural fluid, and the like. Although the sample is often taken from a human subject (such as a patient), the sample may be from any mammal, including, but not limited to dogs, cats, horses, goats, sheep, cattle, pigs, etc. The sample may be used directly as obtained from the biological source or following a pretreatment to modify the character of the sample. For example, such pretreatment may include preparing plasma from blood, diluting viscous fluids and so forth. Methods of pretreatment may also involve, but are not limited to, filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation of interfering components, the addition of reagents, lysing, etc. If such methods of pretreatment are employed with respect to the sample, such pretreatment methods are typically such that the nucleic acid(s) of interest remain in the test sample, sometimes at a concentration proportional to that in an untreated test sample (such as namely, a sample that is not subjected to any such pretreatment method(s)). Such “treated” or “processed” samples are still considered to be biological “test” samples with respect to the methods described herein. A “nucleic acid sample” may also include nucleic acid sequence information stored in a memory, and which was originally obtained from a source such as one or more biological fluids, cells, tissues, organs, or organisms.

[0042] The term “read” or “sequence read” (or sequencing reads) refers to a sequence obtained from a portion of a nucleic acid sample. A read may be represented by a string of nucleotides sequenced from any part or all of a nucleic acid molecule. Typically, though not necessarily, a read represents a short sequence of contiguous base pairs in the sample. The read may be represented symbolically by the base pair sequence (in A, T, C, or G) of the sample portion. It may be stored in a memory device and processed as appropriate to determine whether it matches a reference sequence or meets other criteria. A read may be obtained directly from a sequencing apparatus or indirectly from stored sequence information concerning the sample. In some cases, a read is a DNA sequence of sufficient length (such as at least about 25 bp) that can be used to identify a larger sequence or region, for example, that can be aligned and specifically assigned to a chromosome or genomic region or gene. For example, a sequence read may be a short string of nucleotides (such as 20-150 bases) sequenced from a nucleic acid fragment, a short string of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of the entire nucleic acid fragment that exists in thebiological sample. Sequence reads may be obtained by any method known in the art. For example, a sequence read may be obtained in a variety of ways, such as using sequencing techniques or using probes, such as in hybridization arrays or capture probes, or amplification techniques, such as the polymerase chain reaction (PCR) or linear amplification using a single primer or isothermal amplification. Sequence reads can be generated by techniques such as sequencing by synthesis, sequencing by binding, or sequencing by ligation. Sequence reads can be generated using instruments such as MINISEQ, MISEQ, NEXTSEQ, HISEQ, and NOVASEQ sequencing instruments from Illumina, Inc. (San Diego, CA).

[0043] The term “sequencing depth,” as used herein, generally refers to the number of times a locus is covered by a sequence read aligned to the locus. The locus may be as small as a nucleotide, or as large as a chromosome arm, or as large as the entire genome. Sequencing depth can be expressed as 50 , 100 , etc., where “x” refers to the number of times a locus is covered with a sequence read. Sequencing depth can also be applied to multiple loci, or the whole genome, in which case x can refer to the mean number of times the loci or the haploid genome, or the whole genome, respectively, is sequenced. When a mean depth is quoted, the actual depth for different loci included in the dataset spans over a range of values. Ultra-deep sequencing can refer to at least 100 / in sequencing depth.

[0044] As used herein, the terms “aligned,” “alignment,” or “aligning” refer to the process of comparing a read or tag to a reference sequence and thereby determining the likelihood of the reference sequence contains the read sequence. If the reference sequence contains the read, the read may be mapped to the reference sequence or, in certain embodiments, to a particular location in the reference sequence. For example, the alignment of a read to the reference sequence for human chromosome 13 will tell the likelihood of the read is present in the reference sequence for chromosome 13. In some cases, an alignment additionally indicates a location where the read or tag maps to in the reference sequence. For example, if the reference sequence is the whole human genome sequence, an alignment may indicate that a read is present on chromosome 13, and may further indicate that the read is on a particular strand and / or site of chromosome 13. A “site” may be a unique position on a polynucleotide sequence or a reference genome (such as chromosome ID, chromosome position and orientation). In some embodiments, a site may provide a position for a residue, a sequence tag, or a segment on a sequence.

[0045] Aligned reads or tags are one or more sequences that are identified as a match in terms of the order of their nucleic acid molecules to a known sequence from a reference genome. Alignment can be done manually, although it is typically implemented by a computer algorithm, as it would be impossible to align reads in a reasonable time period for implementing the methods disclosed herein. The matching of a sequence read in aligning can be a 100% sequence match or less than 100% (non-perfect match).

[0046] Alignment may be performed by modifications and / or combinations of methods such as Burrows-Wheeler Aligner (BWA), iSAAC, BarraCUDA, BFAST, BLASTN, BLAT, Bowtie, CASHX, Cloudburst, CUDA-EC, CUSHAW, CUSHAW2, CUSHAW2-GPU, drFAST, ELAND, ERNE, GNUMAP, GEM, GensearchNGS, GMAP and GSNAP, Geneious Assembler, LAST, MAQ, mrFAST and mrsFAST, MOM, MOSAIK, MPscan, Novoaligh & NovoalignCS, NextGENe, Omixon, PALMapper, Partek, PASS, PerM, PRIMEX, QPalma, RazerS, REAL, cREAL, RMAP, rNA, RT Investigator, Segemehl, SeqMap, Shrec, SHRiMP, SLIDER, SOAP, SOAP2, SOAP3 and SOAP3-dp, SOCS, SSAHA and SSAHA2, Stampy, SToRM, Subread and Subjunc, Taipan, UGENE, VelociMapper, XpressAlign, and ZOOM.

[0047] The term “mapping” used herein refers to specifically assigning a sequence read to a larger sequence, such as a reference genome, by alignment.

[0048] As used herein, a “file” includes digital files. In some embodiments, a file is on a computer storage medium (such as a computer hard drive, for example a spinning magnetic disk drive or a solid state drive). In some embodiments, the digital file is stored in the format of a BAM, FASTQ, SAM, CRAM, JSON, CIGAR, or VCF file.Methods

[0049] Disclosed herein are methods for quantifying the efficiency of enriching for a target nucleic acid in a nucleic acid sample. The nucleic acid sample may be of any origin, for example from a biological fluid, from wastewater, or any other source. In some embodiments, the nucleic acid sample comprises DNA. In some embodiments, the nucleic acid sample comprises RNA.

[0050] In some embodiments, the nucleic acid sample comprises nucleic acids derived from a human. In some embodiments, the nucleic acid sample comprises nucleic acids derived from a microorganism. In some embodiments, the nucleic acid sample comprisesnucleic acids derived from a virus (including a phage) or a bacterium. In some embodiments, the nucleic acid sample comprises both nucleic acids derived from a human and comprises nucleic acids derived from one or more microorganisms as described herein.Adding Control Polynucleotides

[0051] In some embodiments, the method includes adding control polynucleotides to a nucleic acid sample comprising sample nucleic acids to be enriched. In some embodiments, the control polynucleotides comprise polynucleotides from a specified microorganism, such as a virus or a bacteria. Control polynucleotides may be DNA or RNA. Any type of control polynucleotides may be used so long as they are distinguishable from sequences in the sample to be analyzed.

[0052] In some embodiments, the control polynucleotides are derived from any of the organisms disclosed in TABLE 1 below:TABLE 1

[0053] Accordingly, in some embodiments, the control polynucleotides comprise polynucleotides from Enterobacteria phage T7, Escherichia virus T4, Allobacillus halotolerans, Escherichia virus MS2, Escherichia virus Qbeta, Imtechella halotolerans, Phocid alphaherpesvirus 1 , Phocine morbillivirus, or Truepera radiovictrix.

[0054] A person of skill in the art will readily be able to determine the amount of control polynucleotides that may be added. For example, control polynucleotides may be addedto a sample at a low, fixed amount. Control polynucleotides may be added in an amount relative to the concentration of sample nucleic acids. A person of skill in the art may determine an amount of control polynucleotides that provides ample signal for sequencing, but does not overwhelm sample nucleic acids. For example, control polynucleotides may be added to a concentration of about IxlO7copies per mL, about IxlO6to IxlO8copies per mL, or about IxlO5to IxlO9copies per mL control polynucleotides.Providing a Probe Set

[0055] In some embodiments, the method includes providing a probe set comprising probes complementary to targeted regions of the sample nucleic acids, and probes complementary to targeted regions of the control polynucleotides.

[0056] In some embodiments, the probes complementary to targeted regions of the control polynucleotides comprise any one of SEQ ID NOs: 1-462. Probes complementary to targeted regions of the control polynucleotides may target areas of the control polynucleotides that are spaced apart with gaps, or may be tiled, as is known to those of skill in the art. Probes complementary to targeted regions of the control polynucleotides may be configured to have about the same length as probes complementary to targeted regions of the sample nucleic acids. Probes complementary to targeted regions of the control polynucleotides may be configured to have approximately the same GC content as probes complementary to targeted regions of the sample nucleic acids.Hybridizing the Probes and Providing an Enriched Library

[0057] An enriched library may be created using any method known to those of skill in the art. In some embodiments, the enrichment method uses probes to capture and separate the target of interest. In other embodiments, the enrichment method uses primers to amplify the target of interest. In other embodiments, the enrichment method depletes unwanted nucleic acids.

[0058] In some embodiments, an enriched library is created by hybridizing capture probes to the target nucleic acid and separating out any hybridized nucleic acids from the nucleic acid sample. In this embodiment, hybridized capture probes may comprise biotinylated probes which are separate out hybridized nucleic acids using streptavidin. As another example, probe capture may be accomplished using charged probes and magnetic beads.

[0059] In some embodiments, an enriched library may be created using an amplification-based method. In this embodiment, target- specific amplification primers are used to amplify the target region of interest. Any amplification-based method known to those of skill in the art may be used.

[0060] In some embodiments, an enriched library may be created using tagmentation. Tagmentation has been described in U.S. 9,080,211, U.S. 10,035,992, U.S. 10,443,087, U.S. 10,920,219, U.S. 11,685,946, WO2010048605 Al, WO2015160895A2, WO2015189636A1 , WO2018156519A1 , and W02020144373 Al which are each incorporated by reference in its entirety. Briefly, tagmentation uses a transposase enzyme to fragment double stranded nucleic acid and add an adaptor to the 5' ends of the fragmented double stranded nucleic acid. In embodiments where a single-stranded adaptor is added to the 5' end, a polymerase may be used to extend the 3' ends of the fragmented double stranded nucleic acid. In embodiments where a double stranded or forked adaptor is added to the 5' end, a polymerase and / or ligase may be used to extend the 3' ends of the fragmented double stranded nucleic acid to the double stranded / forked adaptor.

[0061] Following tagmentation and extension, the double stranded nucleic acid is denatured and the various enrichments methods described in previous paragraphs may be used to enrich the target of interest. In some embodiments, the method comprises: providing a probe set comprising at least two probes complementary to one or more target nucleic acids, wherein the probes are affixed to a support; capturing the one or more target nucleic acids on the support; using the one or more captured target nucleic acids as a template strand to produce one or more nucleic acid duplexes immobilized on the support, wherein the one or more target nucleic acids hybridize to one or more probes of the probe set on the support; tagmenting and extending the nucleic acid duplexes to produce one or more tagged nucleic acid duplexes; amplifying the tagged nucleic acid duplexes to produce a plurality of tagged nucleic acid strands; contacting the plurality of tagged nucleic acid strands with a probe set to create an enriched library; and amplifying the enriched library.

[0062] In some embodiments, the target sequence of interest is enriched using target specific primers and amplifying the target sequences of interest. The amplified target sequences of interest are tagmented and extended to produce an enriched library. In some embodiments, the tagmentation step uses transposase attached to a solid surface such as a bead.

[0063] In some embodiments, the target sequence of interest is enriched by depleting off-target sequences. Enrichment by depletion has been described in U.S. 11,421,216, U.S. 20240 / 191288, W02020132304A1 and WO2022212589A1 which are each incorporated by reference in its entirety. In some embodiments, depletion methods are combined with the enrichment methods described above. In a non-limiting example, enrichment may include depleting ribosomal RNA and amplifying and / or capturing the target of interest. In these embodiments, the depleted / enriched targets may undergo tagmentation / extension to produce a sequencing library.Sequencing the Enriched Library

[0064] In some embodiments, the method includes sequencing the enriched library, thereby providing sequence reads from targeted regions of sample nucleic acids, targeted control sequence reads from targeted regions of control polynucleotides, and untargeted control sequence reads from untargeted regions of control polynucleotides.

[0065] Sequencing may be accomplished by any method known to those of skill in the art. For example, sequence reads can be generated by techniques such as sequencing by synthesis, sequencing by binding, or sequencing by ligation.Quantifying the Efficiency of the Enrichment

[0066] In some embodiments, the method includes quantifying the efficiency of the enrichment for the targeted regions based on the proportion of targeted control sequence reads to untargeted control sequence reads.

[0067] In some embodiments, quantifying the efficiency of the enrichment comprises calculating an enrichment factor based on the ratio of sequencing reads derived from regions of control polynucleotides targeted by enrichment probes to the proportion derived from untargeted regions.

[0068] In some embodiments, an enrichment factor is calculated during the normal course of a targeted enrichment analysis pipeline, which may include software applications paired with enrichment panels. An enrichment panel may target, for example, microorganisms and AMR (antimicrobial resistance) markers. A targeted enrichment analysis workflow mayinclude steps such as: read quality control, pre-sorting (binning) of reads by a k-mer classifier, and alignment to genomic regions targeted by enrichment probes, as further described below.

[0069] In some embodiments, an enrichment factor may be calculated based on counts of sequence reads deriving from both targeted and untargeted genomic regions of control polynucleotides. This data can be obtained from, for example, a binning step or a read alignment step. Each of these two exemplary methods are further described below.

[0070] Binning'. In some embodiments, quantifying targeted enrichment efficiency comprises binning sequence reads from targeted regions of control polynucleotides, and binning sequence reads from untargeted regions of control polynucleotides. In some embodiments, the method includes binning sequence reads by taxonomic kingdom of origin. For example, in some embodiments, a sample composition binner reports a breakdown of the sequence reads into various bins. In some embodiments, sequence reads are binned according to their taxonomic kingdom of origin and whether they were targeted by enrichment probes. Notably, in some embodiments, two of the bins that reads are assigned to include the targeted and untargeted control polynucleotide bins. In some embodiments, the classifier is trained to identify sequence reads from polynucleotides derived from the genomes of commercially available spike-in internal control (IC) organisms, such as those disclosed herein, and distinguish between targeted and untargeted regions.

[0071] In some embodiments, the enrichment factor based on to the ratio of sequencing reads derived from regions targeted by enrichment probes to the proportion derived from untargeted regions. Accordingly, in some embodiments, targeted enrichment efficiency is quantified based on binned sequence reads. In some embodiments, quantifying targeted enrichment efficiency is based on a count of binned sequence reads from targeted regions of control polynucleotides, and a count of binned sequence reads from untargeted regions of control polynucleotides.

[0072] In some embodiments, quantifying targeted enrichment efficiency comprises normalizing a count of targeted control sequence reads by a sequence length of targeted regions of the control polynucleotides and / or normalizing a count of untargeted control sequence reads by a sequence length of untargeted regions of the control polynucleotides. In some embodiments, the sequence read counts are normalized by sequencelengths of the targeted and untargeted regions of a given control genome, for example as shown in the below equation:

[0073] In this formula, I is the length of a read, L is the length of a region of the genome. The lengths of the untargeted and targeted reads and genomic regions are distinguished by u and t, respectively.

[0074] In some embodiments, quantifying targeted enrichment efficiency is based on an average sequence read length. In some embodiments, for methods involving the sample composition binner, equal read lengths are assumed, and so the counts of the reads can be used as opposed to their product with read length. The formula is modified so that both summation terms are replaced with the product of the average read length and the counts of reads binned to the untargeted (w) and targeted ( / ) regions. Because the average read length is the same for untargeted and targeted regions, this term cancels out and the equation becomes the following:

[0075] Aligning. In some embodiments, the method comprises aligning sequence reads to a reference sequence. In some embodiments, genomic regions from the same internal control organisms are included in the database used for sample sequence alignment. In some embodiments, quantifying targeted enrichment efficiency is based on the alignment of sequence reads to the reference sequence.

[0076] In some embodiments, for the alignment method, the provenance of individual reads is lost, and fractions of sequence reads may be aligned while others are ignored. Instead, a vector of the counts of aligned bases across the length of a sequence may be obtained. In that case, the two fractions in the formula above can be replaced with the mean or median of this depth vector. The formula can be represented slightly differently in this case:

[0077] In this formula, d represents the aligned depth at position n along a sequence. Accordingly, in some embodiments, quantifying targeted enrichment efficiency isbased on a depth of aligned sequence reads at each position in the reference sequence. In some embodiments, quantifying targeted enrichment efficiency is based on a median or mean depth of aligned sequence reads at each position in the reference sequence.

[0078] In some embodiments, multiple discrete regions within a genome are considered as the untargeted and targeted fraction. For example, enrichment probes may be designed for the targeted regions of control polynucleotides and samples of the remainder of the genome of the control polynucleotides may be used to measure the overall unenriched signal. Discontinuities may be ignored and treated as a single concatenated sequence when using the above formula.

[0079] Effectively prepared enrichment libraries can often produce thousands of times more reads in targeted regions versus untargeted regions. This may produce libraries that have zero reads that align to or are classified as untargeted internal control. In these cases, to prevent division by zero in the untargeted ratio within the formula above, the denominator may be set to 1. Furthermore, in some embodiments, the user is provided with the loglO value such that the range of possible values seen constrained between about zero to about 10.

[0080] Quantifying targeted enrichment efficiency based on sequence read counts at an alignment step versus a binning step may have certain advantages and disadvantages. For example, alignment may be an inherently more precise method because it provides location information for each read along the genome. By doing so, one can generate a per-base estimate of sequencing depth along a genomic region. Current methods of short read sequencing library preparation introduce various forms of bias in the number of reads generated from a given part of the genome, some of which are highly non-uniform. These can skew estimates of localized coverage. The advantage of having a per-base estimate of depth is that this bias can be removed by using a percentile-based estimator (median), however the arithmetic or geometric mean can also be used. In either case, the lower bound of the estimate of the depth in the untargeted region may be set algorithmically to 1. This allows for a more reasonable upper bound on the value of the enrichment factor, which can be useful for user interpretation.

[0081] An advantage of using sequence read counts from a binning step is that it can be performed much more quickly. However, it may suffer from two drawbacks. First, it may be a less specific process, and even reads with extremely low misclassification rates will still err. Second, without the ability to assign a read to a specific location on a referencegenome, one may assume a uniform distribution. In doing so, the estimated depth may simplify the arithmetic mean, which can be easily skewed by outlier regions of very high or low coverage.

[0082] A practical consideration that applies to alignment-based methods is that both the targeted and untargeted regions of all possible internal control organisms available to users may need to be included in any alignment database. For internal control organisms with longer genomes, the total applicable sequence length may be subsampled to reduce computational complexity at the expense of a proportional degree of precision.

[0083] There are also some considerations that apply to binning-based methods that can improve accuracy. These may be implemented if some conditions and assumptions described above do not apply.

[0084] The first relates to how a binner used for this calculation is constructed. In the exemplary implementation described above, the binner may be an ensemble of binary classifiers and each read can be assigned to multiple classes. The classes may be highly imbalanced, for example, one classifier may be trained on all high-quality bacterial genomes and another may be trained on only internal control genomes. The classifier itself may also generate a reduced representation of genomes during training and during read classification. These two facts taken together may contribute to misclassifications of reads with minor variations in sequence composition. Implementing this method with a dedicated classifier designed specifically to identify internal control sequences may produce more accurate counts. Because of the smaller size of such a classifier, the need for hashing sequences to expedite classification may also be reduced.

[0085] It may also be beneficial to include sequence regions neighboring those targeted by enrichment probes because of the likelihood that these regions are inadvertently enriched, especially when using read chemistry featuring longer insert sizes and read lengths.

[0086] Finally, in calculations of the enrichment factor from classified read counts, a uniform read length may be is assumed. However, this is only an estimate and may not be exactly accurate for real data after read trimming. In some embodiments, the exact read lengths of reads classified as targeted versus untargeted internal control can be used to shift the ratio to a more accurate value.Computer-Implemented Methods

[0087] In another aspect, disclosed herein are computer-implemented methods for quantifying efficiency of targeted enrichment sequencing. FIG. 1 is a flow diagram that schematically illustrates an exemplary method 100 for quantifying the efficiency of a targeted enrichment sequencing process. In some embodiments, the method 100 is implemented on a computer. The method 100 may be embodied in a set of executable program instructions stored on a computer-readable medium, such as one or more disk drives, of a computing system. When the method 100 is initiated, the executable program instructions can be loaded into a memory and executed by one or more processors of a server device.Receiving Sequence Reads

[0088] As shown in FIG. 1, the method 100 for quantifying the efficiency of targeted enrichment sequencing may start from start block 110. The method 100 may proceed to block 120, wherein sequence reads from targeted regions of sample nucleic acids, targeted control sequence reads from targeted regions of control polynucleotides, and untargeted control sequence reads from untargeted regions of control polynucleotides are received. The sequence reads may be generated as further described above. The sequence reads may be in any digital file format, such as FASTQ.Quantifying Targeted Enrichment Efficiency

[0089] The method 100 may proceed to process block 130, wherein targeted enrichment efficiency is quantified based on the proportion of targeted control sequence reads to untargeted control sequence reads. The process 100 then ends at an end block 140.

[0090] The process block 130 is explained in more detail with reference to FIG. 2, which is a block diagram that illustrates the details on the methods taking place within the process block 130.

[0091] As shown in FIG. 2, the method 130 may include a first block 210, wherein the sequence reads are controlled for read quality. For example, sequence reads which do not fit specified quality parameters such as a predefined length or quality score may be discarded.

[0092] The method 130 may then proceed to block 220, wherein sequence reads are binned. For example, block 220 may include binning sequence reads from targeted regionsof control polynucleotides, and binning sequence reads from untargeted regions of control polynucleotides. Block 220 may include binning sequence reads by taxonomic kingdom of origin. This may be accomplished as further described above with respect to “Quantifying the Efficiency of the Enrichment.”

[0093] In FIG. 2, two alternative methods for calculating an enrichment factor are shown in boxes with dashed lines. For example, one method includes blocks 221 and 222, while the other includes blocks 241 and 242. For example, the method 130 may proceed from block 220 to block 221, wherein sequence reads from control polynucleotides are counted by bin. For example, binned sequence reads from targeted regions of control polynucleotides may be counted, and binned sequence reads from untargeted regions of control polynucleotides may be counted. This may be accomplished as further described above with respect to “Quantifying the Efficiency of the Enrichment.”

[0094] The method 130 may proceed to block 222, wherein an enrichment factor is calculated based on a binned read ratio. For example, an enrichment factor may be calculated based on a ratio of the count of binned sequence reads from targeted regions of control polynucleotides, and the count of binned sequence reads from untargeted regions of control polynucleotides. The calculation may be based on assuming an equal average sequence read length for both targeted and untargeted regions. This may be accomplished as further described above with respect to “Quantifying the Efficiency of the Enrichment.”

[0095] The method 130 may proceed to block 260, wherein a detection report is generated. The detection report may include information related to a quantified target enrichment efficiency, such as based on a calculated enrichment factor. The detection report may also include information related to taxonomic groups detected in a nucleic acid sample through targeted enrichment of targeted regions of sample nucleic acids.

[0096] Returning to block 220, the method 130 may instead (or additionally) proceed from block 220 to block 230, wherein sequence reads are split by bin. From block 230, the method 130 may proceed to block 231, wherein undesirable sequence reads are discarded. For example, sequence reads with low complexity (such as sequence reads covering simple sequence repeats) may be discarded. As another example, sequence reads binned to a taxonomic kingdom origin that is not of interest may be discarded. For example, when targetedenrichment is being used to detect non-human sequences in a nucleic acid sample, sequence reads that are binned as “human” may be discarded at block 231.

[0097] Returning to block 230, if some sequences are not discarded the method 130 may proceed from block 230 to block 240, wherein sequence reads are aligned to one or more databases of reference genome sequences. This may be accomplished as further described above with respect to “Quantifying the Efficiency of the Enrichment.” Alignment may be accomplished by any method known to those of skill in the art.

[0098] The method 130 may proceed from block 240 to block 241, wherein depth vectors for targeted and untargeted regions of control polynucleotides are calculated. For example, a median or mean depth of aligned sequence reads at each position in a reference sequence of a control polynucleotide may be calculated. This may be accomplished as further described above with respect to “Quantifying the Efficiency of the Enrichment.”

[0099] The method 130 may proceed to block 242, wherein an enrichment factor is calculated based on the depth vectors, such as an alignment mean depth. For example, an enrichment factor may be calculated based on a median or mean depth of aligned sequence reads at each position in the reference sequence, comparing depth at targeted regions to depth at untargeted regions. This may be accomplished as further described above with respect to “Quantifying the Efficiency of the Enrichment.”

[0100] The method 130 may proceed from block 242 to block 260, wherein a detection report is generated. As described previously, the detection report may include information related to a quantified target enrichment efficiency and information related to taxonomic groups detected in a nucleic acid sample.

[0101] Returning to block 240, the method 130 may proceed from block 240 to block 250, wherein alignments are interpreted. For example, alignments of sequence reads from targeted regions of sample nucleic acids to the one or more reference genome databases may be evaluated to determine one or more taxonomic groups that are present.

[0102] The method 130 may proceed from block 250 to block 260, wherein a detection report is generated. As described previously, the detection report may include information related to a quantified target enrichment efficiency and information related to taxonomic groups detected in a nucleic acid sample.Systems and Computer-Readable Media

[0103] Further disclosed herein are electronic systems for quantifying the efficiency of targeted enrichment sequencing. In some embodiments, the system includes a processor configured to perform a method comprising: receiving sequence reads from targeted regions of sample nucleic acids, targeted control sequence reads from targeted regions of control polynucleotides, and untargeted control sequence reads from untargeted regions of control polynucleotides; and quantifying targeted enrichment efficiency based on the proportion of targeted control sequence reads to untargeted control sequence reads.

[0104] Further disclosed herein are non-transitory computer-readable media. In some embodiments, the non-transitory computer-readable medium includes a plurality of instructions, which when executed by at least one processor, cause the at least one processor to: receive sequence reads from targeted regions of sample nucleic acids, targeted control sequence reads from targeted regions of control polynucleotides, and untargeted control sequence reads from untargeted regions of control polynucleotides; and quantify targeted enrichment efficiency based on the proportion of targeted control sequence reads to untargeted control sequence reads.

[0105] FIG. 3 A illustrates a diagram of an environment in which an enrichment efficiency quantification system can operate in accordance with one or more implementations. The following paragraphs describe the enrichment efficiency quantification system with respect to illustrative figures that portray example implementations and embodiments. For example, FIG. 3 A illustrates a schematic diagram of a computing system 3000 in which an enrichment efficiency quantification application 3106 operates in accordance with one or more implementations. As illustrated, the computing system 3000 includes one or more server device(s) 3102 connected to a user client device 3108, a local device 3118, and a sequencing device 3114 via a network 3112. The network 3112 can comprise any suitable network over which computing devices can communicate.

[0106] As shown in FIG. 3 A, the computing system 3000 includes the server device(s) 3102. In various implementations, the server device(s) 3102 may generate, receive, analyze, store, and transmit digital data, such as data for nucleobase calls or sequenced nucleic- acid polymers. In some implementations, the server device(s) 3102 receive various data from the sequencing device 3114, such as data from a sample genome and / or sequence reads. Theserver device(s) 3102 may also communicate with the user client device 3108. In particular, the server device(s) 3102 can send data for sequence reads, direct nucleobase calls, nucleobase calls, and / or sequencing metrics to the user client device 3108.

[0107] As shown, the server device(s) 3102 includes a sequencing application 3110. In general, the sequencing application 3110 analyzes the data (such as call data) received from the sequencing device 3114 or elsewhere to determine nucleobase sequences for nucleic- acid polymers. For example, the sequencing application 3110 can receive raw data from the sequencing device 3114 and determine a nucleobase sequence for a sample genome or a nucleic-acid segment. In some implementations, the sequencing application 3110 determines the sequences of nucleobases in DNA and / or RNA segments or oligonucleotides.

[0108] As also shown, the sequencing application 3110 includes the enrichment efficiency quantification application 3106. As described below, in some embodiments, the enrichment efficiency quantification 3106 can quantify the efficiency of targeted enrichment sequencing. For example, in some embodiments, the enrichment efficiency quantification application 3106 receives sequence reads from targeted regions of sample nucleic acids, targeted control sequence reads from targeted regions of control polynucleotides, and untargeted control sequence reads from untargeted regions of control polynucleotides. In some embodiments, the enrichment efficiency quantification application 3106 quantifies targeted enrichment efficiency based on the proportion of targeted control sequence reads to untargeted control sequence reads.

[0109] While the sequencing application 3110 has been described as including the enrichment efficiency quantification application 3106, other systems or methods may be included within the sequencing application 3110, such as an application to determine taxonomic groups associated with sample nucleic acids (not illustrated).

[0110] Moreover, while the enrichment efficiency quantification application 3106 is described being implemented on the server device(s) 3102, as part of the sequencing application 3110, in some implementations, the enrichment efficiency quantification application 3106 is implemented by (such as located entirely or in part) on the user client device 3108, the sequencing device 3114, and / or the local device 3118. As mentioned, in some implementations, the enrichment efficiency quantification application 3106 is implemented by one or more other components of the computing system 3000, such as the sequencing device3114. In particular, the enrichment efficiency quantification application 3106 can be implemented in a variety of different ways across the server device(s) 3102, the network 3112, the user client device 3108, the local device 3118, and the sequencing device 3114.

[0111] As further shown in FIG. 3A, the computing system 3000 includes the user client device 3108. In various implementations, the user client device 3108 can generate, store, receive, and send digital data. In particular, the user client device 3108 can receive the data from the sequencing device 3114. As further illustrated, the user client device 3108 includes a sequencing application 3110. The sequencing application 3110 may be a web application or a native application stored and executed on the user client device 3108 (for example, a mobile application, desktop application, or web application). The sequencing application 3110 can receive data from the sequencing application 3110 and / or the enrichment efficiency quantification application 3106. For example, the user client device 3108 can receive variant call files and / or alignment files from the sequencing application 3110.

[0112] The sequencing application 3110 can also include instructions that (when executed) cause the user client device 3108 to receive data from the enrichment efficiency quantification application 3106 and present data from the sequencing device 3114 and / or the server device(s) 3102. Furthermore, the sequencing application 3110 can instruct the user client device 3108 to display data for variant calls, such as nucleobase calls or an indication of a copy number variant. Indeed, the user client device 3108 can display nucleobase call results for a genome sample and / or an indication of a predicted copy number variant.

[0113] As further shown in FIG. 3 A, the computing system 3000 includes the sequencing device 3114. In various implementations, the sequencing device 3114 can sequence a genomic sample or other nucleic-acid polymer. For example, the sequencing device 3114 analyzes nucleic-acid segments or oligonucleotides extracted from genomic samples to generate data either directly or indirectly on the sequencing device 3114. More particularly, the sequencing device 3114 receives and analyzes, within nucleotide-sample slides (such as flow cells), nucleic-acid sequences extracted from genomic samples. In one or more implementations, the sequencing device 3114 utilizes sequencing by synthesis (SBS) to sequence a genomic sample or other nucleic-acid polymers. In addition to, or in the alternative to communicating across the network 3112, in some implementations, the sequencing device 3114 bypasses the network 3112 and communicates directly with the user client device 3108.

[0114] As further depicted in FIG. 3 A, in some implementations, the server device(s) 3102 includes a distributed collection of servers, where the server device(s) 3102 include several server devices distributed across the network 3112 and located in the same or different physical locations. For instance, the server device(s) 3102 can be implemented, in whole or in part, on the local device 3118. To illustrate, the local device 3118 may implement the sequencing application 3110 and / or the enrichment efficiency quantification application 3106. Further, the server device(s) 3102 and / or the local device 3118 can include a content server, an application server, a communication server, a web-hosting server, or another type of server.

[0115] The user client device 3108 illustrated in FIG. 3 A can include various types of client devices. For example, in some implementations, the user client device 3108 includes non-mobile devices, such as desktop computers or servers, or other types of client devices. In various implementations, the user client device 3108 includes mobile devices, such as laptops, tablets, mobile telephones, or smartphones.

[0116] Though FIG. 3 A illustrates the components of the computing system 3000 communicating via the network 3112, in certain implementations, the components of computing system 3000 can also communicate directly with each other, bypassing the network 3112. For instance, in some implementations, the user client device 3108 communicates directly with the sequencing device 3114. Additionally, in some implementations, the user client device 3108 communicates directly with the enrichment efficiency quantification application 3106 and / or the server device(s) 3102. In some implementations, the user client device 3108 communicates directly with the local device 3118. Moreover, the enrichment efficiency quantification application 3106 can access one or more databases housed on or accessed by the server device(s) 3102 or elsewhere in the computing system 3000.

[0117] FIG. 3B is a block diagram of an exemplary server device 3102 that may be used in connection with the computing system 3000 of FIG. 3 A. The server device 3102 may be configured to quantify targeted enrichment efficiency in a nucleic acid sample. The general architecture of the server device 3102 depicted in FIG. 3B includes an arrangement of computer hardware and software components. The server device 3102 may include many more (or fewer) elements than those shown in FIG. 3B. It is not necessary, however, that all of these generally conventional elements be shown in order to provide an enabling disclosure. Asillustrated, the server device 3102 includes a processing unit 310, a network interface 320, a computer readable medium drive 330, an input / output device interface 340, a display 350, and an input device 360, all of which may communicate with one another by way of a communication bus. The network interface 320 may provide connectivity to one or more networks or computing systems. The processing unit 310 may thus receive information and instructions from other computing systems or services via a network. The processing unit 310 may also communicate to and from memory 370 and further provide output information for an optional display 350 via the input / output device interface 340. The input / output device interface 340 may also accept input from the optional input device 360, such as a keyboard, mouse, digital pen, microphone, touch screen, gesture recognition system, voice recognition system, gamepad, accelerometer, gyroscope, or other input device.

[0118] The memory 370 may contain computer program instructions (grouped as modules or components in some embodiments) that the processing unit 310 executes in order to implement one or more embodiments. The memory 370 generally includes RAM, ROM and / or other persistent, auxiliary or non-transitory computer readable media. The memory 370 may store an operating system 372 that provides computer program instructions for use by the processing unit 310 in the general administration and operation of the server device 3102. The memory 370 may store a reference genome 373, such as for use by the sequencing application 3110. The memory 370 may further include computer program instructions and other information for implementing aspects of the present disclosure.

[0119] For example, in one embodiment, the memory 370 includes a sequencing application 3110, which may include an enrichment efficiency quantification application 3106. The enrichment efficiency quantification application 3106 can perform the methods disclosed herein. In addition, memory 370 may include or communicate with the data store 390 and / or one or more other data stores that store one or more inputs, one or more outputs, and / or one or more results (including intermediate results) of quantifying the efficiency of targeted enrichment of a nucleic acid sample of the present disclosure, such the sequencing reads, and / or one or more reference genomes.

[0120] In some embodiments, the disclosed systems and methods may involve approaches for shifting or distributing certain sequence data analysis features and sequence data storage to a cloud computing environment or cloud-based network. User interaction withsequencing data, genome data, or other types of biological data may be mediated via a central hub that stores and controls access to various interactions with the data. In some embodiments, the cloud computing environment may also provide sharing of protocols, analysis methods, libraries, sequence data as well as distributed processing for sequencing, analysis, and reporting. In some embodiments, the cloud computing environment facilitates modification or annotation of sequence data by users. In some embodiments, the systems and methods may be implemented in a computer browser, on-demand or on-line.

[0121] In some embodiments, software written to perform the methods as described herein is stored in some form of computer readable medium, such as memory, CD- ROM, DVD-ROM, memory stick, flash drive, hard drive, SSD hard drive, server, mainframe storage system and the like.

[0122] In some embodiments, the methods may be written in any of various suitable programming languages, for example compiled languages such as C, C#, C++, Fortran, and Java. Other programming languages could be script languages, such as Perl, MatLab, SAS, SPSS, Python, Ruby, Pascal, Delphi, R and PHP. In some embodiments, the methods are written in C, C#, C++, Fortran, Java, Perl, R, Java or Python. In some embodiments, the method may be an independent application with data input and data display modules. Alternatively, the method may be a computer software product and may include classes wherein distributed objects comprise applications including computational methods as described herein.

[0123] In some embodiments, the methods may be incorporated into pre-existing data analysis software, such as that found on sequencing instruments. Software comprising computer implemented methods as described herein are installed either onto a computer system directly, or are indirectly held on a computer readable medium and loaded as needed onto a computer system. Further, the methods may be located on computers that are remote to where the data is being produced, such as software found on servers and the like that are maintained in another location relative to where the data is being produced, such as that provided by a third party service provider.

[0124] An assay instrument, desktop computer, laptop computer, or server which may contain a processor in operational communication with accessible memory comprising instructions for the implementation of systems and methods. In some embodiments, a desktopcomputer or a laptop computer is in operational communication with one or more computer readable storage media or devices and / or outputting devices. An assay instrument, desktop computer and a laptop computer may operate under a number of different computer based operational languages, such as those utilized by Apple based computer systems or PC based computer systems. An assay instrument, desktop and / or laptop computers and / or server system may further provide a computer interface for creating or modifying experimental definitions and / or conditions, viewing data results and monitoring experimental progress. In some embodiments, an outputting device may be a graphic user interface such as a computer monitor or a computer screen, a printer, a hand-held device such as a personal digital assistant (such as PDA, Blackberry, iPhone), a tablet computer (such as iPAD), a hard drive, a server, a memory stick, a flash drive and the like.

[0125] A computer readable storage device or medium may be any device such as a server, a mainframe, a supercomputer, a magnetic tape system and the like. In some embodiments, a storage device may be located onsite in a location proximate to the assay instrument, for example adjacent to or in close proximity to, an assay instrument. For example, a storage device may be located in the same room, in the same building, in an adjacent building, on the same floor in a building, on different floors in a building, etc. in relation to the assay instrument. In some embodiments, a storage device may be located off-site, or distal, to the assay instrument. For example, a storage device may be located in a different part of a city, in a different city, in a different state, in a different country, etc. relative to the assay instrument. In embodiments where a storage device is located distal to the assay instrument, communication between the assay instrument and one or more of a desktop, laptop, or server is typically via Internet connection, either wireless or by a network cable through an access point. In some embodiments, a storage device may be maintained and managed by the individual or entity directly associated with an assay instrument, whereas in other embodiments a storage device may be maintained and managed by a third party, typically at a distal location to the individual or entity associated with an assay instrument. In embodiments as described herein, an outputting device may be any device for visualizing data.

[0126] An assay instrument, desktop, laptop and / or server system may be used itself to store and / or retrieve computer implemented software programs incorporating computer code for performing and implementing computational methods as described herein,data for use in the implementation of the computational methods, and the like. One or more of an assay instrument, desktop, laptop and / or server may comprise one or more computer readable storage media for storing and / or retrieving software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like. Computer readable storage media may include, but is not limited to, one or more of a hard drive, a SSD hard drive, a CD-ROM drive, a DVD-ROM drive, a floppy disk, a tape, a flash memory stick or card, and the like. Further, a network including the Internet may be the computer readable storage media. In some embodiments, computer readable storage media refers to computational resource storage accessible by a computer network via the Internet or a company network offered by a service provider rather than, for example, from a local desktop or laptop computer at a distal location to the assay instrument.

[0127] In some embodiments, computer readable storage media for storing and / or retrieving computer implemented software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like, is operated and maintained by a service provider in operational communication with an assay instrument, desktop, laptop and / or server system via an Internet connection or network connection.

[0128] In some embodiments, a hardware platform for providing a computational environment comprises a processor (such as CPU) wherein processor time and memory layout such as random access memory (such as RAM) are systems considerations. For example, smaller computer systems offer inexpensive, fast processors and large memory and storage capabilities. In some embodiments, graphics processing units (GPUs) can be used. In some embodiments, hardware platforms for performing computational methods as described herein comprise one or more computer systems with one or more processors. In some embodiments, smaller computer are clustered together to yield a supercomputer network.

[0129] In some embodiments, computational methods as described herein are carried out on a collection of inter- or intra- connected computer systems (such as grid technology) which may run a variety of operating systems in a coordinated manner. For example, the CONDOR framework (University of Wisconsin-Madison) and systems available through United Devices are exemplary of the coordination of multiple stand-alone computersystems for the purpose dealing with large amounts of data. These systems may offer Perl interfaces to submit, monitor and manage large sequence analysis jobs on a cluster in serial or parallel configurations.EXAMPLES

[0130] Some aspects of the embodiments discussed above are disclosed in further detail in the following examples, which are not in any way intended to limit the scope of the present disclosure. Those in the art will appreciate that many other embodiments also fall within the scope of the disclosure, as it is described herein above and in the claims.Example 1

[0131] To evaluate the binning-based and alignment-based implementations described above, data was processed from a pilot study in which paired shotgun libraries and enrichment libraries prepared with the Illumina® Urinary Pathogen ID / AMR Panel (UPIP) for 363 urine samples.

[0132] Residual urine samples (n=362) were prepared for sequencing by enrichment with UPIP probes and by traditional shotgun metagenomic methods after DNA extraction (Quick-DNA Urine Kit, ZymoBiomics). Libraries were sequenced (1x147 bp) on the NextSeq™ 550 platform (Illumina) to ~16- I7total reads per sample. Internal control (IC) DNA (Bacteriophage T7, Microbiologies) was spiked into samples prior to extraction at 1.217copies / mL to enable absolute quantification of targeted genitourinary pathogens. Sequencing results resampled to IM reads / sample were analyzed using the Explify™ UPIP Data Analysis app v2.0.0 (Illumina BaseSpace™ Sequence Hub). Each sample was assigned a Study ID (for example, ID147). Each of the 614 libraries was assigned a unique Accession (for example, IDBD-R116093). Files containing sequence reads for each library were processed with the UPIP enrichment app (vl.1.0) on BaseSpace, which detected the most abundant microorganisms in each sample (“Detected” column in TABLE 2 below).

[0133] The ratio of the organism median depth in the enrichment library to the shotgun library was used to calculate the “Organism Enrichment Factor” in TABLE 2 below. The “Organism Enrichment Factor” was used for comparing the two internal control-basedcalculations of the Enrichment Factor (based on binning or on aligning) in the enriched libraries.

[0134] A development version of the UPIP enrichment app was used to analyze the same libraries to generate the data to calculate the IC-based enrichment factors according to the methods described herein, including the equations described above. The version of the Enrichment Factor calculation that uses sequence alignment is the “Enrichment Factor (mean aligned)” column in TABLE 2 below.

[0135] The version of the Enrichment Factor calculation that uses the composition binner is the “Enrichment Factor (Binned Read Ratio)” column in TABLE 2 below. To calculate the enrichment factor based on the composition binner, binned internal control untargeted read counts, binned internal control targeted read counts, the IC organism-specific correction factor, the ratio of untargeted reads to targeted reads, the corrected ratio, and the log 10 of the corrected ratio were calculated as described herein, including the equations described above.TABLE 2

[0136] FIG. 4 shows histograms of the log 10 transformed enrichment factor values calculated for the data set with different hatch markings denoting the library prep method (shotgun or enrichment). Despite the difference in the scale of the calculated value, both implementations show a clear separation of the shotgun and enriched libraries and thus can provide immediate value to users trying to develop an enrichment-based workflow in their labs.Other Considerations

[0137] Conditional language used herein, such as, among others, “can,” “might,” “may,” “e.g.,” and the like, unless specifically stated otherwise, or otherwise understood within the context as used, is generally intended to convey that certain embodiments include, while other embodiments do not include, certain features, elements and / or states. Thus, such conditional language is not generally intended to imply that features, elements and / or states are in any way required for one or more embodiments or that one or more embodiments necessarily include logic for deciding, with or without author input or prompting, whether these features, elements and / or states are included or are to be performed in any particular embodiment. The terms “comprising,” “including,” “having,” “involving,” and the like are synonymous and are used inclusively, in an open-ended fashion, and do not exclude additional elements, features, acts, operations, and so forth. Also, the term “or” is used in its inclusivesense (and not in its exclusive sense) so that when used, for example, to connect a list of elements, the term “or” means one, some, or all of the elements in the list.

[0138] Disjunctive language such as the phrase “at least one of X, Y or Z,” unless specifically stated otherwise, is otherwise understood with the context as used in general to present that an item, term, etc., may be either X, Y or Z, or any combination thereof (such as X, Y and / or Z). Thus, such disjunctive language is not generally intended to, and should not, imply that certain embodiments require at least one of X, at least one of Y or at least one of Z to each be present.

[0139] The terms “about” or “approximate” and the like are synonymous and are used to indicate that the value modified by the term has an understood range associated with it, where the range can be ±20%, ±15%, ±10%, ±5%, or ±1%. The term “substantially” is used to indicate that a result (such as a measurement value) is close to a targeted value, where close can mean, for example, the result is within 80% of the value, within 90% of the value, within 95% of the value, or within 99% of the value.

[0140] Unless otherwise explicitly stated, articles such as “a” or “an” should generally be interpreted to include one or more described items.

[0141] While the above detailed description has shown, described, and pointed out novel features as applied to illustrative embodiments, it will be understood that various omissions, substitutions, and changes in the form and details of the devices or algorithms illustrated can be made without departing from the spirit of the disclosure. As will be recognized, certain embodiments described herein can be embodied within a form that does not provide all of the features and benefits set forth herein, as some features can be used or practiced separately from others. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.

[0142] It should be appreciated that all combinations of the foregoing concepts (provided such concepts are not mutually inconsistent) are contemplated as being part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the inventive subject matter disclosed herein.

[0143] The scope of the present disclosure is not intended to be limited by the specific disclosures of examples in this section or elsewhere in this specification, and may bedefined by claims as presented in this section or elsewhere in this specification or as presented in the future. The language of the claims is to be interpreted broadly based on the language employed in the claims and not limited to the examples described in the present specification or during the prosecution of the application, which examples are to be construed as nonexclusive.

Claims

WHAT IS CLAIMED IS:

1. A method for quantifying the efficiency of enriching for a target nucleic acid in a nucleic acid sample, comprising:(a) adding control polynucleotides to a nucleic acid sample comprising sample nucleic acids;(b) providing a probe set comprising probes complementary to targeted regions of the sample nucleic acids, and probes complementary to targeted regions of the control polynucleotides;(c) hybridizing the probes to the targeted regions, and amplifying or separating any hybridized nucleic acids from the nucleic acid sample, thereby providing an enriched library;(d) sequencing the enriched library, thereby providing sequence reads from targeted regions of sample nucleic acids, targeted control sequence reads from targeted regions of control polynucleotides, and untargeted control sequence reads from untargeted regions of control polynucleotides; and(e) quantifying the efficiency of the enrichment for the targeted regions based on the proportion of targeted control sequence reads to untargeted control sequence reads.

2. The method of claim 1, wherein the control polynucleotides comprise polynucleotides from a microorganism.

3. The method of claim 1 or 2, wherein the control polynucleotides comprise polynucleotides from Enterobacteria phage T7, Escherichia virus T4, Allobacillus halotolerans, Escherichia virus MS2, Escherichia virus Qbeta, Imtechella halotolerans, Phocid alphaherpesvirus 1 , Phocine morbillivirus, or Truepera radiovictrix.

4. The method of any one of claims 1-3, wherein the probes complementary to targeted regions of the control polynucleotides comprise any one of SEQ ID NOs: 1-462.

5. The method of any one of claims 1-4, wherein quantifying targeted enrichment efficiency comprises normalizing a count of targeted control sequence reads by a sequence length of targeted regions of the control polynucleotides, or normalizing a count of untargeted control sequence reads by a sequence length of untargeted regions of the control polynucleotides.

6. The method of any one of claims 1-5, wherein quantifying targeted enrichment efficiency comprises binning sequence reads from targeted regions of control polynucleotides, and binning sequence reads from untargeted regions of control polynucleotides.

7. The method of claim 6, wherein the method includes binning sequence reads by taxonomic kingdom of origin.

8. The method of claim 6 or 7, wherein targeted enrichment efficiency is quantified based on binned sequence reads.

9. The method of any one of claims 6-8, wherein quantifying targeted enrichment efficiency is based on a count of binned sequence reads from targeted regions of control polynucleotides, and a count of binned sequence reads from untargeted regions of control polynucleotides.

10. The method of claim 8 or 9, wherein quantifying targeted enrichment efficiency is based on an average sequence read length.

11. The method of any one of claims 1-7, wherein the method comprises aligning sequence reads to a reference sequence, and wherein quantifying targeted enrichment efficiency is based on the alignment of sequence reads to the reference sequence.

12. The method claim 11 , wherein quantifying targeted enrichment efficiency is based on a depth of aligned sequence reads at each position in the reference sequence.

13. The method of claim 12, wherein quantifying targeted enrichment efficiency is based on a median or mean depth of aligned sequence reads at each position in the reference sequence.

14. The method of any one of claims 1-13, wherein the method comprises separating any hybridized nucleic acids from the nucleic acid sample.

15. The method of any one of claims 1-14, wherein the method comprises: providing a probe set comprising at least two probes complementary to one or more target nucleic acids, wherein the probes are affixed to a support; capturing the one or more target nucleic acids on the support; using the one or more captured target nucleic acids as a template strand to produce one or more nucleic acid duplexes immobilized on the support, wherein the one or more target nucleic acids hybridize to one or more probes of the probe set on the support;tagmenting and extending to produce one or more tagged nucleic acid duplexes; amplifying the one or more tagged nucleic acid duplexes to produce a plurality of tagged nucleic acid strands; contacting the plurality of tagged nucleic acid strands with a probe set to create an enriched library; and amplifying the enriched library.

16. The method of any one of claims 1-15, wherein the nucleic acid sample comprises nucleic acids derived from a human.

17. The method of any one of claims 1-16, wherein the nucleic acid sample comprises nucleic acids derived from a microorganism.

18. The method of claim 17, wherein the microorganism is a virus or a bacterium.

19. The method of any one of claims 1-18, wherein the nucleic acid sample comprises DNA or RNA.

20. The method of any one of claims 1-19, further comprising modifying a method for the enriching for a target nucleic acid in a nucleic acid sample, wherein the modifying is based on the efficiency of the enrichment for the targeted regions.

21. The method of claim 20, wherein the modifying comprises selecting a capture probe for the target nucleic acid; adjusting an initial concentration of a nucleic acid sample; adjusting a concentration of a reagent in a step of the method for the enriching; and / or removing a step in the method for the enriching.

22. The method of any one of claims 1-21, further comprising discarding the nucleic acid sample based on the efficiency of the enrichment for the targeted regions.

23. The method of any one of claims 1-21, further comprising analyzing the nucleic acid sample based on the efficiency of the enrichment for the targeted regions.

24. The method of claim 23, wherein the analyzing comprises sequencing the nucleic acid sample.

25. The method of any one of claims 1-24, further comprising providing a quality control metric to an end user for the enriching for a target nucleic acid in a nucleic acid sample based on the efficiency of the enrichment for the targeted regions.

26. An electronic system for quantifying the efficiency of targeted enrichment sequencing, comprising a processor configured to perform the method of any one of claims 1- 25.

27. An electronic system for quantifying the efficiency of targeted enrichment sequencing, comprising a processor configured to perform a method comprising: receiving sequence reads from targeted regions of sample nucleic acids, targeted control sequence reads from targeted regions of control polynucleotides, and untargeted control sequence reads from untargeted regions of control polynucleotides; and quantifying targeted enrichment efficiency based on the proportion of targeted control sequence reads to untargeted control sequence reads.

28. A non-transitory computer-readable medium comprising a plurality of instructions, which when executed by at least one processor, cause the at least one processor to: receive sequence reads from targeted regions of sample nucleic acids, targeted control sequence reads from targeted regions of control polynucleotides, and untargeted control sequence reads from untargeted regions of control polynucleotides; and quantify targeted enrichment efficiency based on the proportion of targeted control sequence reads to untargeted control sequence reads.

Citation Information

Patent Citations

  • Modified transposases for improved insertion sequence bias and increased DNA input tolerance

    US10035992B2

  • Methods and compositions for preparing sequencing libraries

    US10443087B2

  • Tagmentation using immobilized transposomes with linkers

    US10920219B2

  • Nuclease-based RNA depletion

    US11421216B2

  • Complex surface-bound transposome complexes

    US11685946B2