Modifying sequencing cycles or imaging during sequencing run to meet customized coverage estimates
By estimating and adjusting sequencing cycle and flow cell region imaging using a sequence-to-coverage system, the problem of inaccurate coverage in existing sequencing systems is solved, achieving efficient resource utilization and accurate target coverage.
Patent Information
- Application Number
- CN202480041739.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-08-02
- Filing Date
- 2024-06-26
- Publication Date
- 2026-01-20
AI Technical Summary
Existing sequencing systems have inaccurate estimates of nucleotide read coverage, resulting in wasted computation time, memory, and materials. Furthermore, existing sequence-to-answer workflows have not been successful at commercial scale, and sequencing runs cannot be effectively tuned to meet target coverage.
By estimating read coverage of genome samples within the pool using a sequence-to-coverage system, adjusting the number of sequencing cycles and imaging of flow pool regions, we can ensure that each genome sample achieves the target coverage and reduce unnecessary sequencing cycles and imaging.
It improves sequencing efficiency, reduces computation time, memory and material consumption, and achieves more accurate coverage estimation and resource conservation.
Smart Images

Figure CN121368638A_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority and benefit to U.S. Provisional Patent Application No. 63 / 511,564, filed June 30, 2023, entitled “MODIFYING SEQUENCING CYCLES ORIMAGING DURING A SEQUENCING RUN TO MEET CUSTOMIZED COVERAGE ESTIMATION,” and U.S. Provisional Patent Application No. 63 / 517,160, filed August 2, 2023, entitled “MODIFYING SEQUENCING CYCLES OR IMAGING DURING A SEQUENCING RUN TO MEET CUSTOMIZED COVERAGE ESTIMATION.” Each of these applications is incorporated herein by reference in its entirety. Background Technology
[0003] In recent years, biotechnology companies and research institutions have improved the hardware and software used for nucleotide sequencing of genomic samples and for determining nucleotide base detection. For example, some existing sequencers and sequencing data analysis software (collectively referred to as "existing sequencing systems") predict individual nucleotide bases within a sequence using conventional Sanger sequencing or sequencing-by-synthesis (SBS) methods. When using SBS, existing sequencing systems can monitor thousands to billions of oligonucleotides synthesized in parallel from a template to predict the nucleotide base detection of an ever-growing number of nucleotide reads. During sequencing runs in many existing sequencing systems, a camera captures images of irradiated fluorescent tags incorporated into the oligonucleotides. After capturing such images, some existing sequencing systems determine the nucleotide base detection of nucleotide reads corresponding to the respective oligonucleotide clusters on a flow cell or other nucleotide sample substrate for a given sequencing run. For example, some existing sequencing systems utilize sequencing data analysis software to analyze image data captured during sequencing cycles to determine the nucleotide base detection of a given oligonucleotide cluster and sequence such detections throughout the sequencing cycle to determine the nucleotide reads of the given cluster.
[0004] As part of this improved genome sequencing, biotechnology companies and research institutions have also improved methods for simultaneously pooling and sequencing large numbers of genomic samples. Existing sequencing systems can pool genetic samples from different individuals to increase the number of samples analyzed in a single sequencing run. For example, existing sequencing systems can utilize sample multiplexing (or multiplex sequencing) to add a separate“barcode” or index sequence to each deoxyribonucleic acid (DNA) fragment during library preparation. The index sequence corresponds to individual genomic samples within a sample pool. After the index sequences are identified, existing sequencing systems can perform demultiplexing to identify which index sequences and which oligonucleotide clusters on a flow cell correspond to which genomic samples.
[0005] Despite recent advances in multiplexing and per-cycle image analysis, existing sequencing systems cannot accurately determine nucleotide read coverage for a given genomic sample prior to ending a sequencing run and face other technical shortcomings that alter the level of nucleotide read coverage for samples provided by a given sequencing run. In multiplex sequencing, for example, the number of nucleotide fragments from each genomic sample in a cluster can not be evenly distributed, resulting in variations in nucleotide read depth or coverage. This uneven representation sometimes causes sequencing devices to perform an insufficient number of sequencing cycles or images for a sequencing run (or otherwise under-sequence) to generate the required number or length of nucleotide reads to meet a target coverage level for a given sample. While sequencing devices can under-sequence DNA fragments extracted from some samples, sequencing devices sometimes can perform an excessive number of sequencing cycles or images for a sequencing run (or otherwise over-sequence) to generate the required number or length of nucleotide reads to meet a target coverage level.
[0006] Due to the uncertainty and variation in read data coverage for a given sample produced by a given sequencing run, existing sequencing systems inefficiently expend excess computational time, memory, and consumable materials to compensate for inter-run variations. Some existing sequencing systems inefficiently expend excess computational time and memory to address under-sequenced samples. For example, existing sequencing systems often perform additional sequencing cycles during a sequencing run to avoid under-sequencing some samples. The additional sequencing cycles require excess computational time, memory, and reagents. Due to performing additional sequencing cycles in a sequencing run, existing sequencing systems often over-sequence samples within a sample pool. While increasing sequencing cycles can reduce under-sequenced samples, existing sequencing systems cannot completely eliminate under-sequenced samples. Thus, in addition to expending excess computational time and memory to over-sequence samples, existing systems must expend additional computational time to perform one or more additional sequencing runs to compensate for under-sequenced samples of a previous sequencing run.
[0007] Likewise, due to the uncertainty and variability of read data coverage—and in addition to wasting processing time and memory—existing sequencing systems inefficiently consume and waste excess reagents, processing materials, and sample materials during additional sequencing cycles or runs. By extending sequencing cycles to compensate for coverage uncertainty and sometimes performing additional sequencing runs to compensate for under-sequenced samples, existing sequencing systems consume an excessive amount of processing materials, including sequencing reagents, library preparation kits, cluster amplification materials, flow cells or other nucleotide sample substrates, scarce space on such flow cells, and other materials. In addition to consuming such materials, existing sequencing systems sometimes require re-extraction of genomic material from an individual and re-performance of library preparation required to seed oligonucleotide clusters on additional flow cells to perform additional sequencing runs to compensate for previous sequencing runs that failed to produce target nucleotide read coverage for variant call (or other secondary analysis) of the individual. For many existing systems, the relationship between cycle number and consumed processing materials is a linear function. Thus, many existing sequencing systems consume excess processing materials and sample materials to compensate for the above-described coverage uncertainty and variability.
[0008] Despite inefficiently utilizing computing resources and sequencing materials, some existing sequencing systems have theoreticalized models to compensate for the uncertainty and variability of read data coverage by modeling or attempting to implement a sequence-to-answer workflow. In theory, a sequence-to-answer workflow includes mapping and aligning read data of a genomic sample, including oligonucleotide clusters of the same genomic sample, during a sequencing run to determine nucleotide read coverage of the sample in real-time and stopping the sequencing run when the determined coverage meets a target. Such a sequence-to-answer workflow would require existing sequencing systems to convert raw sequencing data into meaningful nucleotide read coverage determinations through secondary analysis before the end of a sequencing run. In practice, however, such a sequence-to-answer workflow has not been successful at commercial scale nor has it resulted in substantial improvements in reducing sequencing cycles or imaging. For example, existing sequencing systems have not developed a computational model or hardware that enables such systems to generate data using a sequencing device, pre-process the data, transfer the data from the sequencing device to a server device to complete secondary analysis or accurately determine nucleotide read coverage at a sufficient speed to obtain a coverage answer before ending a sequencing run. In other words, existing sequencing systems have not accurately and simultaneously performed mapping and alignment of nucleotide reads (or other types of secondary analysis) during a corresponding sequencing run on a sequencing device to enable timely adjustments to a sequencing run before ending.
[0009] These problems and challenges, and others, exist in existing sequencing systems. SUMMARY
[0010] The present disclosure describes one or more embodiments of systems, methods, and non-transitory computer-readable storage media that solve one or more of the problems set forth above or provide other advantages over the prior art. For example, the disclosed systems estimate read coverage of genomic samples in a pool and adjust the number of sequencing cycles based on the estimated read coverage to meet a target coverage. Additionally or alternatively, the disclosed systems can determine a customized set of flow cell regions to image from a flow cell to meet a target coverage. As part of generating the estimated read coverage, the disclosed systems can estimate variation caused by sample merging and by filter variation.
[0011] To illustrate, in some embodiments, the disclosed systems perform index cycles to efficiently estimate respective numbers of clusters among samples within a pool. The disclosed systems can also estimate by-filter variation by generating a by-filter map that includes an indication of whether an oligonucleotide cluster of a sample passed a purity filter (or other filter) of an initial cycle of a sequencing run. Based on the respective numbers of clusters belonging to respective samples and the estimated numbers of clusters that passed the filter, the disclosed systems can estimate read coverage levels for individual genomic samples. The disclosed systems can also determine, for a sequencing run, a customized number of sequencing cycles sufficient to generate nucleotide reads that meet a target read coverage level for each genomic sample based on the estimated read coverage levels. In some implementations, the disclosed systems determine a customized set of flow cell regions to image from a flow cell sufficient to generate nucleotide reads that meet a target read coverage level. The disclosed systems also perform a sequencing run on a sequencing device by (i) completing the customized number of sequencing cycles and / or (ii) capturing images of the customized set of flow cell regions (e.g., flow cell tiles) during sequencing cycles of the sequencing run.
[0012] Additional features and advantages of one or more embodiments of the present disclosure will be set forth in the description that follows, and in part will be apparent from the description, or can be learned by practice of such exemplary embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0013] The detailed description refers to the following drawings which are
[0014] FIG. 1 A sequencing device and a corresponding sequence-to-coverage system according to one or more embodiments of the present disclosure are illustrated, which can operate in a computing system.
[0015] FIGS. 2A-2B Potential read coverage level failures or other technical sequencing limitations caused by various sources of variation during a sequencing run are illustrated.
[0016] FIG. 3An overview of a sequence-to-coverage system modifying a number of sequencing cycles in a sequencing run or a number of images of flow cell regions in a sequencing run to meet a target read coverage level according to one or more embodiments of the present disclosure is illustrated.
[0017] FIG. 4 An overview of a sequence-to-coverage system performing a subset of sequencing cycles with index cycles performed before genomic sequencing cycles according to one or more embodiments of the present disclosure is illustrated.
[0018] FIG. 5 A sequence-to-coverage system determining a respective number of oligonucleotide clusters belonging to a respective genomic sample according to one or more embodiments of the present disclosure is illustrated.
[0019] FIG. 6 A sequence-to-coverage system determining filter metrics according to one or more implementations of the present disclosure is illustrated.
[0020] FIGS. 7A-7B A sequence-to-coverage system generating a custom number of sequencing cycles to meet a target read coverage level and performing a sequencing run until the custom number of sequencing cycles is completed according to one or more embodiments of the present disclosure is illustrated.
[0021] FIG. 8 A sequence-to-coverage system determining a custom set of flow cell regions to image during a sequencing run and performing the sequencing run by capturing images of the custom set of flow cell regions during sequencing cycles of the sequencing run according to one or more embodiments of the present disclosure is illustrated.
[0022] FIGS. 9A-9B An improvement in sequencing efficiency resulting from performing a custom number of sequencing cycles according to one or more embodiments of the present disclosure is illustrated.
[0023] FIG. 10 An improvement in sequencing efficiency resulting from imaging a custom set of flow cell regions during sequencing cycles according to one or more embodiments of the present disclosure is illustrated.
[0024] FIG. 11 A schematic diagram of an example of a system that can be used to provide biological or chemical analysis according to one or more embodiments of the present disclosure is illustrated.
[0025] FIG. 12 A schematic diagram of an example of a set of components that can cooperate to provide a fluidic path in a system according to one or more embodiments of the present disclosure is illustrated. FIG. 11
[0026] FIG. 13A A flowchart of a series of actions for performing a sequencing run until a custom number of sequencing cycles is completed is illustrated in accordance with one or more embodiments of the present disclosure.
[0027] FIG. 13B A flowchart of a series of actions for performing a sequencing run by capturing images of a custom set of flow cell regions is illustrated in accordance with one or more embodiments of the present disclosure.
[0028] FIG. 14 A block diagram of an example computing device in accordance with one or more embodiments of the present disclosure is illustrated. DETAILED DESCRIPTION
[0029] The present disclosure describes one or more embodiments of a sequence-to-coverage system that can efficiently modify and perform sequencing runs to meet target read coverage levels for genomic samples within a genomic sample pool. For example, the sequence-to-coverage system can determine base calls for index sequences within oligonucleotide clusters from a subset of sequencing cycles of a sequencing run for the genomic samples. The sequence-to-coverage system can also determine respective numbers of oligonucleotide clusters belonging to respective genomic samples in the genomic samples based on the index sequences. Based on the respective numbers of oligonucleotide clusters and a current selected number of sequencing cycles (e.g., a preset number) of the sequencing run, the sequence-to-coverage system can estimate read coverage levels for the genomic samples. The sequence-to-coverage system can also generate a custom number of sequencing cycles sufficient to generate nucleotide reads that meet the target read coverage levels for each of the genomic samples in the sequencing run. Additionally or alternatively, the sequence-to-coverage system determines a custom set of flow cell regions (e.g., flow cell tiles) of a flow cell to image sufficient to generate nucleotide reads that meet the target read coverage levels for each of the genomic samples in the flow cell based on the estimated read coverage levels. The sequence-to-coverage system can perform the sequencing run (i) before the custom number of sequencing cycles is completed and / or (ii) by capturing images of the custom set of flow cell regions during sequencing cycles of the sequencing run on a sequencing device.
[0030] As just noted, the sequence-to-coverage system can determine base calls for index sequences within oligonucleotide clusters from a subset of sequencing cycles of a sequencing run for the genomic samples. In some cases, the sequence-to-coverage system speeds up determining numbers of oligonucleotide clusters belonging to respective genomic samples within a flow cell pool (or other nucleotide sample substrate pool) by base calling index sequences for two read pairs before base calling genomic sequences in library templates for each sample.
[0031] After determining the base calls of the index sequences, the sequence-to-coverage system can determine the corresponding number of oligonucleotide clusters belonging to the corresponding genomic samples. By de-multiplexing the indexed reads to determine which index sequences belong to which genomic samples, the sequence-to-coverage system can quickly and efficiently estimate the corresponding number of clusters corresponding to each genomic sample within the pool. In some embodiments, the sequence-to-coverage system determines the base calls of the index sequences (e.g., in both mates of a paired-end read) prior to determining the base calls of the genomic sequences of the nucleotide reads. However, in some embodiments, the sequence-to-coverage system determines a custom number of sequencing cycles or a custom set of flow cell regions to image without completing the base calls of the index sequences of each read prior to the genomic sequences of each read.
[0032] As mentioned, the sequence-to-coverage system can estimate the read coverage level based on (i) the corresponding number of oligonucleotide clusters belonging to the corresponding genomic sample of the genomic samples in the sequencing run and (ii) the currently selected number of sequencing cycles of the sequencing run. Generally, the sequence-to-coverage system can utilize the corresponding number of oligonucleotide clusters belonging to the corresponding genomic sample to estimate variations caused by unbalanced sample pooling. As explained further below, in some cases, the sequence-to-coverage system estimates an average number of nucleotide reads from the sequencing run sufficient to cover the genomic regions of each genomic sample.
[0033] In some implementations, the sequence-to-coverage system further estimates the read coverage level based on determined filter metrics. During the sequencing run, the sequence-to-coverage system can determine which clusters pass a purity filter or otherwise determine other filter metrics indicative of a subset of oligonucleotide clusters that satisfy a filtering threshold of the signal of the oligonucleotide clusters. Based on determining such filter metrics, the sequence-to-coverage system can account for variations between the genomic samples that originate from low-quality or poor signal data.
[0034] After estimating the read coverage level, the sequence-to-coverage system can determine a custom number of sequencing cycles of the sequencing run sufficient to generate nucleotide reads that satisfy the target read coverage level of each of the genomic samples. For example, the sequence-to-coverage system can adjust the number of sequencing cycles during the sequencing run by increasing or decreasing the preset number of sequencing cycles of the sequencing run prior to the end of the sequencing run. By generating the custom number of sequencing cycles, the sequence-to-coverage system can efficiently eliminate under-sequenced genomic samples, thereby avoiding performing additional and unnecessary sequencing runs.
[0035] In combination with or independent of determining the custom number of sequencing cycles, the sequence-to-coverage system can also determine a custom set of pool regions of the flow cell to image from the flow cell. More specifically, the sequence-to-coverage system can determine a custom set of flow cell regions to image sufficient to generate nucleotide reads that meet the target read coverage level for each of the genomic samples in the genomic samples. For example, by de-multiplexing the nucleotide reads according to the index sequences and determining the clusters passing through the filter within the flow cell, the sequence-to-coverage system can estimate how many flow cell regions need to be imaged to meet the target read coverage level.
[0036] Based on one or both of the custom number of sequencing cycles and the custom set of flow cell regions to image, the sequence-to-coverage system can execute a sequencing run on the sequencing device until completion. For example, the sequence-to-coverage system can execute a sequencing run on the sequencing device until the custom number of sequencing cycles is completed. Additionally or alternatively, the sequence-to-coverage system can capture images of the custom set of flow cell regions during the sequencing cycles of the sequencing run. By customizing the number of sequencing cycles and / or the set of flow cell regions to image, the sequence-to-coverage system can reduce consumable materials, sequencing run time, and computing resources needed to meet the target read coverage level for each of the genomic samples.
[0037] As indicated above, the sequence-to-coverage system provides several technical advantages over existing sequencing systems by, for example, improving resource, sequencing run time, and computational efficiency relative to existing sequencing systems. In some implementations, for example, the sequence-to-coverage system saves sequencing cycles, imaging, consumables, and other physical resources, and reduces overuse of fluidic devices and other hardware within the sequencing device relative to existing sequencing systems. To compensate for read data coverage uncertainty and variation, and to otherwise meet the target read coverage level for the multiplexed samples described above, existing sequencing systems typically repeat sequencing cycles and, at times, perform additional sequencing runs. Such excessive sequencing cycles or runs can require additional run time and consume sequencing reagents, processing materials, and sample materials.
[0038] In contrast to such existing sequencing systems, the sequence-to-coverage system can efficiently generate a customized number of sequencing cycles and / or determine a customized set of flowcell regions to image prior to the end of a sequencing run, thereby performing a sequencing run according to the customized sequencing cycles or flowcell regions. In some examples, the sequence-to-coverage system can reduce one or both of (i) the number of sequencing cycles and (ii) the number of flowcell regions imaged in a given sequencing run to meet a target read coverage level. By customizing parameters of a sequencing run based on a target read coverage level, the sequence-to-coverage system can reduce the run time and physical resources (e.g., reagents) consumed to achieve the target read coverage level. Moreover, by customizing the number of cycles and / or the number of flowcell regions imaged, the sequence-to-coverage system can avoid unnecessary wear and tear on physical components of the sequencing device.
[0039] In addition to reducing the run time and resources consumed to achieve a target read coverage level, relative to existing sequencing systems, the sequence-to-coverage system also reduces the amount of compute time and memory consumed on the sequencing device to reach a target read coverage level for a given sequencing run. By estimating the read coverage level prior to completion of a sequencing run, the sequence-to-coverage system can accurately perform the number of sequencing cycles needed to reach a target read coverage level for each genomic sample. Additionally or alternatively, the sequence-to-coverage system can accurately estimate a set of flowcell regions that, when imaged during a sequencing run, facilitate a sequencing run that produces enough nucleotide reads for each genomic sample to reach a target read coverage level. Relative to existing sequencing systems running on existing sequencing devices, the sequence-to-coverage system can perform a smaller number of sequencing cycles and / or image fewer flowcell regions, which consume less processing and memory due to the reduced sequencing run time, while still achieving an acceptable read coverage level for each genomic sample. As a result of intelligently reducing the sequencing run time, the sequence-to-coverage system can also reduce the amount of compute time needed to perform a sequencing run that meets a target read coverage level for a genomic sample.
[0040] In addition to saving run time, physical resources, and computing resources on the sequencing device for a given sequencing run, in some embodiments, by determining real-time coverage estimates based on only or primarily data generated by the sequencing device and not based on (or based on relatively little data) from secondary analysis performed by another computing device, the sequence-to-coverage system also improves computational efficiency and real-time flexibility relative to existing sequencing systems. As mentioned, some existing sequencing systems have attempted to implement a sequence-to-answer workflow that requires secondary analysis and sometimes separates the computing device from the sequencing device to determine nucleotide read coverage for individual samples before the sequencing run is complete. But such sequence-to-answer workflows have failed to achieve success at a commercial scale and have failed to achieve substantial improvements in efficient sequencing runs (e.g., intelligently adjusting / reducing sequencing cycles or regions of the flowcell to be imaged, saving reagents or computer processing or memory). In contrast to the unsuccessful sequence-to-answer workflows, the sequence-to-coverage system utilizes data obtained from primary analysis on the sequencing device for customized determinations. For example, the sequence-to-coverage system can estimate read coverage levels for individual genomic samples based on data available during primary analysis on the sequencing device. By determining base calls for index sequences and determining the number of clusters for individual genomic samples that pass a filter as a basis for read-coverage estimates, the sequence-to-coverage system efficiently and instantaneously customizes sequencing runs on the sequencing device to avoid unnecessary sequencing cycles and / or unnecessary flowcell region image capture. In relying on data available through primary analysis, the sequence-to-coverage system can avoid the need for further processing and exchange of data that has slowed and proven unsuccessful for existing sequencing systems that have attempted sequence-to-answer workflows.
[0041] As exemplified by the foregoing discussion, the present disclosure utilizes various terms to describe features and advantages of the sequence-to-coverage system. As used herein, for example, the term “sequencing run” refers to an iterative process of determining primary structure of nucleotide sequences from samples (e.g., genomic samples) on a sequencing device. Specifically, a sequencing run includes cycles of sequencing chemistry and imaging by a sequencing device that incorporates nucleobases into growing oligonucleotides to determine nucleotide reads from nucleotide sequences extracted from samples (or other sequences within library fragments) and seeded throughout a flowcell. In some cases, a sequencing run includes replicating oligonucleotides obtained or extracted from one or more genomic samples seeded in clusters throughout a flowcell. Upon completion of a sequencing run, the sequencing device can generate base call data in a file, such as a binary base call (BCL) sequence file or a fast full quality (FASTQ) file.
[0042] Relatedly, as used herein, for example, the term“sequencing cycle” refers to an iteration of adding or incorporating one or more nucleobases into one or more oligonucleotides representing or corresponding to a sample sequence (e.g., a genomic or transcriptomic sequence of a sample) or to a corresponding adapter sequence. In some cases, a sequencing cycle includes both an iteration of incorporating nucleobases into an oligonucleotide cluster using sequencing chemistry and capturing an image of such a cluster attached to a flow cell. A sequencing cycle can include one or both of an indexing cycle and a genomic sequencing cycle. For example, one oligonucleotide cluster or set of oligonucleotide clusters can be undergoing a genomic sequencing cycle in which nucleobases corresponding to a sample genomic sequence are incorporated, while another oligonucleotide cluster or another set of oligonucleotide clusters can be simultaneously undergoing an indexing cycle in which nucleobases corresponding to an index sequence of a nucleotide read are incorporated.
[0043] As further used herein, the term“genomic sequencing cycle” refers to an iteration of adding or incorporating one or more nucleobases into one or more oligonucleotides representing or corresponding to a sample genomic sequence (or a cDNA sequence). Specifically, a genomic sequencing cycle can include repeatedly capturing and analyzing one or more images having data indicative of individual nucleobases that are added or incorporated into an oligonucleotide representing or corresponding to one or more sample genomic sequences or that are (in parallel) added or incorporated into an oligonucleotide representing or corresponding to one or more sample genomic sequences. For example, in one or more embodiments, each genomic sequencing cycle involves capturing and analyzing an image to determine a single read of a DNA (or RNA) strand representing a portion of a genomic sample (or a transcribed sequence from a genomic sample). However, as noted above, in some cases, a genomic sequencing cycle is specific to an oligonucleotide cluster or set of oligonucleotide clusters.
[0044] In contrast, the term“indexing cycle” refers to an iteration of adding or incorporating one or more nucleobases into one or more oligonucleotides representing or corresponding to one or more index sequences. Specifically, an indexing cycle can include repeatedly capturing and analyzing one or more images of an oligonucleotide cluster indicative of one or more nucleobases that are added or incorporated into the oligonucleotide (or in parallel are added or incorporated into the oligonucleotide) representing or corresponding to one or more index sequences. An indexing cycle differs from a genomic sequencing cycle in that an indexing cycle includes sequencing at least one nucleobase (or a majority of the nucleobases) from one or more index sequences that identify or encode one or more sample library fragments. Because a genomic sequencing cycle can be specific to one or more oligonucleotide clusters, an indexing cycle for one oligonucleotide cluster can be performed simultaneously with a genomic sequencing cycle for another oligonucleotide cluster.
[0045] Relatedly, the term“currently selected number of sequencing cycles” refers to an adjustable value representing a number of sequencing cycles to be performed during a sequencing run. Specifically, the currently selected number of sequencing cycles can be determined automatically, based on a user selection, or preset according to a default number. For example, the sequence-to-coverage system can determine that the currently selected number of sequencing cycles is equal to 150 sequencing cycles. The sequence-to-coverage system can adjust the number of sequencing cycles by increasing the number of sequencing cycles or decreasing the number of sequencing cycles.
[0046] As used herein, the term“genomic sample” refers to a target genome or portion of a genome that is subjected to an assay or sequencing. For example, a genomic sample includes one or more nucleotide sequences (or copies of such isolated or extracted sequences) isolated or extracted from a sample organism. Specifically, a genomic sample includes a whole genome (in whole or in part) isolated or extracted from a sample organism and composed of nitrogenous heterocyclic bases. A genomic sample can include fragments or molecules of deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or other polymeric forms of nucleic acids or chimeric or hybrid forms of nucleic acids described below. In some cases, a genomic sample is present in a sample prepared or isolated by a kit and received by a sequencing device.
[0047] As used herein, the term“nucleobase call” (or simply“base call”) refers to a determination or prediction of a particular nucleobase (or nucleobase pair) of an oligonucleotide (e.g., a read) during a sequencing cycle. Specifically, a nucleobase call can indicate a determination or prediction of a type of nucleobase that has been incorporated within an oligonucleotide on a flow cell (e.g., based on a nucleobase call of a read). In some cases, for a nucleotide read, a nucleobase call includes determining or predicting a nucleobase based on intensity values generated by a fluorescently tagged nucleotide added to one or more oligonucleotides in a cluster of a flow cell. Alternatively, a nucleobase call includes determining or predicting a nucleobase from a chromatographic peak or current change caused by a nucleotide through a nanopore of a flow cell. As set forth above, a single nucleobase call can be an adenine (A) call, a cytosine (C) call, a guanine (G) call, or a thymine (T) call, or a uracil (U) call.
[0048] Further, as used herein, the term“library fragment” refers to a sample genomic sequence (or cDNA sequence) that is ligated to include one or more adapter sequences or primer sequences that facilitate detection or isolation of the sample genomic sequence or cDNA sequence. For example, a library fragment can include, but is not limited to, a sample genomic sequence (or cDNA sequence) that is extracted from a sample and ligated to directly or indirectly bind to one or more of a binding adapter sequence, an index sequence, or a read primer sequence.
[0049] As used herein, the term“sample genomic sequence” refers to a nucleotide sequence that is extracted, copied, or complementary to a chromosome of a sample. For example, a sample genomic sequence includes a nucleotide sequence that has been isolated or copied from chromosomal DNA of a sample, or a nucleotide sequence that has been sequenced as complementary to the extracted or copied nucleotide sequence. Thus, a sample genomic sequence includes genomic DNA (gDNA) of a particular unknown sample. Accordingly, as described herein, in some embodiments, the sequence-to-coverage system can use sample complementary sequences including cDNA instead of sample genomic sequences including gDNA in sample library fragments or anywhere suitable cDNA can replace gDNA as understood by one of skill in the art. Indeed, any embodiment or nucleotide read in the present disclosure that uses or includes a sample genomic sequence can also use or include a cDNA sequence corresponding to the genomic sample.
[0050] As used herein, the term“index sequence” refers to a unique artificial nucleotide sequence that identifies a nucleotide read of a sample and is linked to a nucleotide sequence of a sample (e.g., a gDNA fragment or a cDNA fragment) or to another sequence within a sample library fragment. As indicated above, an index sequence can be part of a sample library fragment. Similarly, an index sequence can be used to order or sort nucleotide reads by sample, especially as part of a demultiplexing process. In some cases, a sample library fragment includes an index primer sequence that satisfies the following requirements: different from a read primer sequence and indicates a starting point or starting nucleobase for determining nucleobases of the index sequence.
[0051] As used herein, the term“oligonucleotide cluster” refers to a local collection of DNA or RNA molecules immobilized on a solid surface. In particular, an oligonucleotide cluster can refer to a collection of fragment nucleotide sequences immobilized on a flow cell region of a flow cell. For example, an oligonucleotide cluster can refer to a collection of nucleotide fragments derived from a genomic sample. An oligonucleotide cluster can be imaged with one or more optical signals. For example, an oligonucleotide cluster image can be captured by a camera during a sequencing cycle of light emitted from an irradiated fluorescent tag incorporated into an oligonucleotide from one or more clusters on a flow cell.
[0052] As used herein, the term“nucleotide read” refers to a sequence of one or more nucleobases (or nucleobase pairs) inferred from all or a portion of a sample nucleotide sequence (e.g., a sample genomic sequence, a complementary DNA). In particular, a nucleotide read includes a determined or predicted sequence of nucleobase calls from a nucleotide sequence (or monoclonal nucleotide sequence set) of a sample library fragment corresponding to a genomic sample. For example, in some cases, a sequencing device determines a nucleotide read by generating nucleobase calls for nucleobases determined by a flow cell nanopore, via a fluorescent tag, or from a cluster in a flow cell.
[0053] As used herein, the term“read coverage level” refers to a measure or value indicative of the depth or redundancy of nucleotide sequence information for a particular genomic coordinate or genomic region of a sample. Specifically, the read coverage level refers to the number of times a particular genomic coordinate or genomic region of a sample is covered or spanned by a nucleotide read. Read coverage levels can be relevant when describing the depth of sequencing data obtained for a particular genomic region of interest or a particular genomic sample. For example, a read coverage level can include a numerical value indicative of the average number of unique nucleotide reads of a genomic sample that span or cover a genomic coordinate or region of a human genomic sample (e.g., 10x, 30x, 45x). In some cases, the read coverage level is limited to the average number of unique nucleotide reads that span a non-N portion of a human genome (e.g., a non-N portion of a PAR-masked human genome).
[0054] As used herein, the term“target read coverage level” refers to a desired or expected depth of sequencing coverage for a particular genomic coordinate or genomic region within a genomic sample. Specifically, the target read coverage level represents the minimum number of times a location within a genomic sample should be sequenced to achieve a desired level of confidence in the accuracy of the obtained sequence data. For example, a target read coverage level can include a numerical value indicative of a desired read coverage level for a given location within a genomic sample (e.g., 40).
[0055] As further used herein, the term“genomic coordinate” (or sometimes simply“coordinate”) refers to a particular location or positioning of a nucleobase within a genome (e.g., a genome of an organism or a reference genome). In some cases, a genomic coordinate includes an identifier of a particular chromosome of a genome and an identifier of a nucleobase positioning within the particular chromosome. For example, one or more genomic coordinates can include a number, name, or other identifier of a chromosome (e.g., chr1, chrX, chrM) and one or more specific positionings, such as a band number positioning following the identifier of the chromosome (e.g., chr1:1234570 or chr1:1234570-1234870). In some cases, a genomic coordinate refers to a genomic coordinate on a sex chromosome (e.g., chrX or chrY) or mitochondrial DNA (e.g., chrM). Further, in certain implementations, a genomic coordinate refers to a source of a reference genome (e.g., mt for a mitochondrial DNA reference genome or SARS-CoV-2 for a reference genome of SARS-CoV-2 virus) and a positioning of a nucleobase within the source of the reference genome (e.g., mt:16568 or SARS-CoV-2:29001). In contrast, in certain cases, a genomic coordinate refers to a positioning of a nucleobase within a reference genome without reference to a chromosome or source (e.g., 29727).
[0056] As used herein, a “genomic region” refers to a range of genomic coordinates. Like genomic coordinates, in certain implementations, a genomic region can be identified by an identifier of a chromosome and one or more specific positions, such as a band number position following the chromosome identifier, e.g., chr1:1234570-1234870. In various implementations, genomic coordinates include a location within a reference genome. In some cases, genomic coordinates are specific to a particular reference genome.
[0057] As used herein, the term “sequencing device” refers to an instrument or platform used to perform a sequencing process. In particular, a sequencing device refers to an instrument or platform used to perform a sequencing process based on a sequencing by synthesis (SBS) technique, a single molecule real-time sequencing (SMRT) technique using magnetic beads or nanopores, or other suitable medium. For example, a sequencing device can include components including, but not limited to, a flow cell receptacle, a fluidic system, a laser, an imaging system, and computing power for acquiring, processing, and analyzing image data during a sequencing run.
[0058] As used herein, the term “filter metric” refers to a measure that indicates the quality and reliability of sequencing data for an oligonucleotide cluster. In particular, a filter metric can include a value that indicates the quality and / or brightness of sequencing data that has passed a particular filter criterion. A filter metric can indicate a subset of imaged oligonucleotide clusters that meet a filter threshold for the signal of the oligonucleotide cluster. For example, a filter metric can include a percent passing filter (%PF), which represents the percentage of oligonucleotide clusters that pass a purity filter.
[0059] As used herein, the term “filter threshold” refers to a predetermined value or range of values used to determine whether a parameter meets a filter criterion. In particular, a filter threshold can include a numerical value above (or below) which a filter metric indicates an acceptable quality. Clusters with filter values that exceed the filter threshold can be considered to pass the filter. For example, a filter threshold can include a threshold purity value. A purity value can include a ratio of the brightest base intensity within an oligonucleotide cluster divided by the sum of the brightest base intensity and the second brightest base intensity. A sequence-to-coverage system can determine that oligonucleotide clusters with purity values below the filter threshold do not pass the filter and remove them from image analysis results. To illustrate, a cluster can pass a filter threshold if no more than 1 base calls have a purity value below 0.6.
[0060] As used herein, the term "nucleotide sample substrate" refers to a plate or substrate, e.g., a flowcell, that includes oligonucleotides for sequencing nucleotide sequences from genomic or other sample nucleic acid polymers. In particular, a flowcell can refer to a substrate containing fluidic channels through which reagents and buffers can travel as part of sequencing. For example, in one or more embodiments, a flowcell (e.g., a patterned flowcell or a non-patterned flowcell) can include small fluidic channels and oligonucleotide samples that can bind to adapter sequences on the substrate. In other particular implementations, a flowcell can be an open substrate with one or more regions for oligonucleotide samples to be analyzed, and can use a charged mat or other means to position the oligonucleotide samples. In yet another particular implementation, a nucleotide sample substrate can be a membrane with a nanopore through which one or more oligonucleotide samples can pass. As indicated above, a flowcell can include cells and wells (e.g., nanopores) containing clusters of oligonucleotides.
[0061] As set forth above, a flowcell or other nucleotide sample substrate can (i) include a device with a lid that extends over the reaction structures to form a flow channel therebetween in communication with the plurality of reaction sites of the reaction structures, and can (ii) include a detection device configured to detect a specified reaction occurring at or near the reaction sites. A flowcell or other nucleotide sample substrate can include a solid-state light detection or imaging device, such as a charge-coupled device (CCD) or complementary metal-oxide semiconductor (CMOS) (light) detection device. As one particular example, a flowcell can be configured to be fluidically and electrically coupled to a cartridge (with integrated pumps) that can be configured to be fluidically and / or electrically coupled to a bioassay system. The cartridge and / or bioassay system can deliver reaction solutions to the reaction sites of the flowcell according to a predetermined protocol (e.g., sequencing by synthesis) and perform a plurality of imaging events. For example, the cartridge and / or bioassay system can direct one or more reaction solutions through the flow channel of the flowcell, flowing along the reaction sites. At least one of the reaction solutions can include four types of nucleotides with the same or different fluorescent labels. The nucleotides can bind to the reaction sites of the flowcell, such as to corresponding oligonucleotides at the reaction sites. The cartridge and / or bioassay system then illuminates the reaction sites using an excitation light source (e.g., a solid-state light source, such as a light-emitting diode (LED)). The excitation light can provide an emission signal (e.g., light at one or more wavelengths that is different from the excitation light and can be different from one another) that can be detected by a light sensor of the flowcell.
[0062] As used herein, the term "flow cell region" refers to a region of a nucleotide sample substrate. Specifically, a flow cell region refers to a region or portion of a flow cell that contains one or more oligonucleotide clusters. For example, a flow cell region can refer to a cell of a flow cell. More specifically, flow cell regions can be organized in a grid-like pattern across a nucleotide sample substrate, and each flow cell region corresponds to a particular location on the surface of the nucleotide sample substrate. A flow cell region can also contain wells (e.g., nanowells) that contain individual compartments in which oligonucleotide clusters are amplified, denatured, and sequenced.
[0063] As used herein, the term "nucleotide read" (or simply "read") refers to a sequence of one or more nucleobases (or nucleobase pairs) inferred from all or a portion of a sample nucleotide sequence (e.g., a sample genomic sequence, cDNA). Specifically, a nucleotide read includes a determined or predicted sequence of nucleobase calls from a nucleotide sequence (or monoclonal nucleotide sequence set) of a sample library fragment corresponding to a genomic sample. For example, in some cases, a sequencing device determines a nucleotide read by generating nucleobase calls of nucleobases determined via fluorescent tags, nanowells through a nucleotide-sample matrix, or from a cluster in a flow cell.
[0064] As used herein, the term "nucleobase" refers to a nitrogenous base. Specifically, a nucleobase comprises a component of a nucleotide. For example, a nucleobase can be adenine (A), cytosine (C), guanine (G), thymine (T), or uracil (U).
[0065] The following paragraphs describe a sequence-to-coverage system with respect to the illustrative figures depicting example embodiments and specific implementations. For example, FIG. 1 A schematic diagram of a computing system 100 in which a sequence-to-coverage system 106 according to one or more embodiments operates is illustrated. As illustrated, the computing system 100 includes a local server device 102 connected to one or more server devices 110, a sequencing device 108, and a client device 114 via a network 112. Although FIG. 1 An embodiment of the sequence-to-coverage system 106 is shown, although alternative embodiments and configurations are described below.
[0066] As FIG. 1 illustrated, the local server device 102, the sequencing device 108, the server devices 110, and the client device 114 can communicate with each other via the network 112. The network 112 includes any suitable network through which computing devices can communicate. Example networks are discussed in more detail below. FIG. 14 Example networks are discussed in more detail below.
[0067] As FIG. 1As shown, the sequencing device 108 includes a device for sequencing a genomic sample or other nucleic acid polymer. In some embodiments, the sequencing device 108 directly or indirectly analyzes nucleic acid fragments or oligonucleotides extracted from a genomic sample on the sequencing device 108 with computer-implemented methods and systems (described herein) to generate nucleotide reads or other data. More specifically, the sequencing device 108 receives a nucleotide sample substrate (e.g., a flow cell) that includes nucleotide fragments extracted from a sample, which are then copied and the nucleotide base sequence of such extracted nucleotide fragments is determined. In one or more embodiments, the sequencing device 108 sequences nucleic acid polymers into nucleotide reads with SBS. Additionally, the sequencing device 108 can determine base calls of index sequences. In addition to or as an alternative to communicating over the network 112, in some embodiments, the sequencing device 108 bypasses the network 112 and communicates directly with the local server device 102 or the client device 114.
[0068] As shown, the sequencing device 108 includes a device for sequencing a genomic sample or other nucleic acid polymer. In some embodiments, the sequencing device 108 directly or indirectly analyzes nucleic acid fragments or oligonucleotides extracted from a genomic sample on the sequencing device 108 with computer-implemented methods and systems (described herein) to generate nucleotide reads or other data. More specifically, the sequencing device 108 receives a nucleotide sample substrate (e.g., a flow cell) that includes nucleotide fragments extracted from a sample, which are then copied and the nucleotide base sequence of such extracted nucleotide fragments is determined. In one or more embodiments, the sequencing device 108 sequences nucleic acid polymers into nucleotide reads with SBS. Additionally, the sequencing device 108 can determine base calls of index sequences. In addition to or as an alternative to communicating over the network 112, in some embodiments, the sequencing device 108 bypasses the network 112 and communicates directly with the local server device 102 or the client device 114. FIG. 1 Further shown, the local server device 102 is located at or near the same physical location as the sequencing device 108. Indeed, in some embodiments, the local server device 102 and the sequencing device 108 are integrated into the same computing device, as shown by the dashed line 122. The local server device 102 can run the sequencing system 104 to generate, receive, analyze, store, and transmit digital data, such as by receiving base call data or determining index sequence data or filter metric data based on analyzing such base call data. As shown, FIG. 1 As shown, the sequencing device 108 can send (and the local server device 102 can receive) base call data generated during a sequencing run of the sequencing device 108. By executing software in the form of the sequencing system 104, the local server device 102 can estimate read coverage levels of genomic samples in a pool of genomic samples. The local server device 102 can also communicate with the client device 114. In particular, the local server device 102 can send data to the client device 114, including read coverage information of genomic samples, filter metric data, estimated read coverage levels, variant call files (VCFs) indicating nucleobase calls, genotype calls, sequencing metrics, error data, or other metrics.
[0069] As shown, the sequencing device 108 includes a device for sequencing a genomic sample or other nucleic acid polymer. In some embodiments, the sequencing device 108 directly or indirectly analyzes nucleic acid fragments or oligonucleotides extracted from a genomic sample on the sequencing device 108 with computer-implemented methods and systems (described herein) to generate nucleotide reads or other data. More specifically, the sequencing device 108 receives a nucleotide sample substrate (e.g., a flow cell) that includes nucleotide fragments extracted from a sample, which are then copied and the nucleotide base sequence of such extracted nucleotide fragments is determined. In one or more embodiments, the sequencing device 108 sequences nucleic acid polymers into nucleotide reads with SBS. Additionally, the sequencing device 108 can determine base calls of index sequences. In addition to or as an alternative to communicating over the network 112, in some embodiments, the sequencing device 108 bypasses the network 112 and communicates directly with the local server device 102 or the client device 114. FIG. 1Further shown, the server device 110 is located remotely from the local server device 102 and the sequencing device 108. The sequencing device 108 can transmit (and the server device 110 can receive) base call data from the sequencing device 108. The server device 110 can also communicate with the client device 114. In particular, the server device 110 can transmit data to the client device 114, including estimated read coverage levels for a genomic sample, VCF, or other sequencing related information.
[0070] In some embodiments, the server device 110 includes a distributed set of servers, where the server device 110 includes many server devices distributed across the network 112 and located in the same or different physical locations. Further, the server device 110 can include a content server, an application server, a communication server, a network hosting server, or another type of server.
[0071] As FIG. 1 Further exemplified and pointed out in the background, the client device 114 can generate, store, receive, and transmit digital data. In particular, the client device 114 can receive read coverage data from the local server device 102, or receive sequencing metrics from the sequencing device 108. Further, the client device 114 can communicate with the local server device 102 or the server device 110 to receive a VCF including variant or genotype calls and / or other metrics, such as base call quality metrics or pass filter metrics. The client device 114 can accordingly present or display information related to the variant calls or other genotype calls to a user associated with the client device 114 within a graphical user interface. For example, the client device 114 can present a target read coverage interface including elements indicative of potential target read coverage levels for a genomic sample.
[0072] Although FIG. 1 Although the client device 114 is depicted as a desktop or laptop computer, the client device 114 can include various types of client devices. For example, in some embodiments, the client device 114 includes a non-mobile device, such as a desktop computer or a server, or other type of client device. In other embodiments, the client device 114 includes a mobile device, such as a laptop computer, a tablet computer, a mobile phone, or a smart phone. Additional details regarding the client device 114 are discussed below with respect to FIG. 2. FIG. 6
[0073] As FIG. 1 In further illustration, the client device 114 includes a sequencing application 116. The sequencing application 116 can be a web application or a native application (e.g., mobile application, desktop application) stored and executed on the client device 114. The sequencing application 116 can include instructions that, when executed, cause the client device 114 to receive data from the sequencing-to-coverage system 106 and present data regarding read coverage data for a sequencing run, data from a VCF, or other information for display at the client device 114. Further, the sequencing application 116 can instruct the client device 114 to display a graphical user interface for receiving input indicative of a target read coverage level.
[0074] As FIG. 1 In further illustration, versions of the sequencing-to-coverage system 106 can be located on the client device 114 as part of the sequencing application 116. Thus, in some embodiments, the sequence-to-coverage system 106 is implemented by being located (e.g., entirely or partially located) on the client device 114. In other embodiments, the sequence-to-coverage system 106 is implemented by one or more other components of the computing system 100, such as the server device 110. In particular, the sequence-to-coverage system 106 can be implemented in a variety of different ways across the local server device 102, the sequencing device 108, the client device 114, and the server device 110. For example, the sequence-to-coverage system 106 can be downloaded from the server device 110 to the local server device 102 and / or the client device 114, with all or part of the functionality of the sequence-to-coverage system 106 being executed at each respective device within the computing system 100.
[0075] As previously mentioned, sample multiplexing provides several advantages, but is also associated with some potential technical challenges. FIGS. 2A-2B Illustrated are read coverage level failures or other technical sequencing limitations caused by various sources of variation during a sequencing run. FIG. 2A Illustrated are charts depicting various sources of variation within a sequencing run. FIG. 2B Illustrated is how existing sequencing systems can over- or under-sequence a genomic sample due to sources of variation.
[0076] FIG. 2AFigure 200 illustrates various sources of variation within a sequencing run. Figure 200 includes sectors 202, 204, and 206. As shown in sector 202, the most significant source of variation within a sequencing run is sample pooling. Sample pooling refers to the practice of combining multiple individual genome samples into a single genome pool before performing a sequencing reaction. As previously described, sample pooling improves sequencing and computational efficiency by sequencing multiple genome samples during a single sequencing run. However, some variations that may be caused by sample pooling include unequal representation of genome samples. Furthermore, sample pooling can introduce additional contamination into the sequencing run. For example, genetic material from one genome sample may unintentionally cross-contaminate other samples in the pool. Therefore, sample pooling accounts for a large portion of the total variation within a sequencing run, with a coefficient of variation (CV) of approximately 10%–15%.
[0077] like FIG. 2A As illustrated in sector 206, filter failures account for a significant portion of the variation in sequencing runs. As shown, oligonucleotide clusters that failed to pass through the filter account for approximately 2%–5% of the total variation between sequencing runs. Sector 206 represents variations caused by sources related to the quality of sample preparation. Specifically, filter-through variations are caused by a variety of factors, including read quality (e.g., base quality fraction, read alignment score, etc.) caused by sequencing chemistry, cycle-specific bias, or differences in the quality of the input genome samples. Filter-through variations can also be caused by different sequencing platforms. For example, different sequencing platforms may use unique sequencing chemistry that leads to variations in data quality. Furthermore, filter-through variations can be caused by different experimental conditions, such as reagent batches, laboratory protocols, and differences in the quality performance of sequencing reagents, equipment, or other environmental factors. Additionally, filter-through variations can be affected by sample heterogeneity, where individual genome samples within a genome sample pool may have different quality or sequencing complexity, which affects the observed filter-through metric.
[0078] FIG. 2A Sector 204 within diagram 200 is also illustrated. For example... FIG. 2A As shown, sector 204 includes bioinformatics efficiency. Generally, bioinformatics efficiency refers to the ability to perform accurate secondary analysis of sequencing data. Specifically, bioinformatics efficiency involves using efficient algorithms, optimized computational resources, and simplified processes to interpret sequencing data in a cost-effective manner. For example, problems arising when aligning nucleotide reads with a reference genome can lead to variations in bioinformatics efficiency. In some cases, bioinformatics efficiency is measured by (i) dividing the unique, aligned nucleotide reads corresponding to one or more genome samples by (ii) the total number of nucleotide reads from clusters filtered through one or more genome samples. FIG. 2BAs shown, bioinformatics efficiency accounts for about 2-5% of the total variation between sequencing runs. In some examples, bioinformatics efficiency improves with (slightly lower) % filter values, and thus can compensate for lower filter metrics to some extent.
[0079] Existing sequencing systems typically attempt to compensate for sequencing data variation by performing additional sequencing cycles, reducing the number of genomic samples in a genomic sample pool, and / or (worse yet) performing additional sequencing runs. FIG. 2B A graph of sequencing data generated by an existing sequencing system is illustrated. FIG. 2B A graph 208 illustrating how an existing system often over-sequences or under-sequences genomic samples within a genomic sample pool is illustrated. The graph 208 includes a bar chart depicting a distribution of the number of sequencing runs corresponding to the level of read coverage of the worst-performing sample in each sequencing run. As shown, the x-axis includes unique aligned reads in gigabases (Gb) of the worst-performing sample in each sequencing run.
[0080] To avoid under-sequencing of certain genomic samples, existing sequencing systems tend to over-sequence most samples by performing additional sequencing cycles. As FIG. 2B As shown, the target read coverage level is equal to 40x, which corresponds to about 120 Gb. As FIG. 2B As further shown, sequencing of the worst-performing genomic samples by existing systems is typically about 15% higher than the target read coverage level, resulting in about 138 Gb. More specifically, about 95% of sequencing runs using existing systems result in more than the target 40x read coverage level.
[0081] As FIG. 3 As further shown, while most sequencing runs are over-sequenced, about 5% of sequencing runs still result in poorly-performing genomic samples. Existing systems often must perform additional sequencing runs to re-sequence poorly-performing genomic samples. Thus, not only do existing systems utilize excess resources to over-sequence most genomic samples, existing sequencing systems often need to perform additional sequencing runs to correct under-sequenced samples.
[0082] As previously mentioned, the sequence-to-coverage system 106 can tailor sequencing runs by increasing or decreasing sequencing cycles or by increasing or decreasing the region of the flowcell to be imaged, thereby efficiently meeting the target read coverage level before the end of the sequencing run. FIG. 3 An overview of the sequence-to-coverage system 106 modifying sequencing runs to meet a target read coverage level in accordance with one or more embodiments of the present disclosure is illustrated. FIG. 3An example of a series of actions 300 is given, including action 302 to determine the base detection of the index sequence, action 304 to determine the corresponding number of clusters belonging to the corresponding genome sample, action 306 to determine the filter metric, action 308 to estimate the read coverage level of the genome sample, action 310 to generate a customized number of sequencing cycles, and action 312 to determine a customized set of flow pool regions to be imaged.
[0083] like FIG. 3 As shown, the sequence-to-coverage system 106 performs action 302 to determine the base detection of the index sequence. The sequence-to-coverage system 106 determines the base detection of the index sequence within the oligonucleotide cluster. By determining the base detection of the index sequence, the sequence-to-coverage system 106 can accurately assign nucleotide reads to their respective genomic samples in multiplex sequencing. The sequence-to-coverage system 106 can determine the base detection of the index sequence at different times relative to determining the base detection of nucleotide reads in the genomic sample. FIG. 3 Examples of non-index-first workflow 314 and index-first workflow 316 include indexing cycles and genome sequencing cycles with different orders.
[0084] like FIG. 3 As further shown, the sequence-to-coverage system 106 can perform action 302 according to the order of index cycles between genome sequencing cycles. The sequence-to-coverage system 106 can perform sequencing cycles according to the order of the non-index-first workflow 314. In the non-index-first workflow 314, the sequence-to-coverage system 106 determines base detections for paired-end reads in the following order: (i) a first nucleotide read corresponding to a first portion of the sample genome sequence, (ii) a first index sequence appended to the sample genome sequence, (iii) a second index sequence appended to the sample genome sequence, and (iv) a second nucleotide read corresponding to a second portion of the sample genome sequence. During the sequencing process, the sequence-to-coverage system 106 performs paired-end switching between determining base detections of the first and second index sequences. In the non-index-first workflow 314, the sequence-to-coverage system 106 does not complete the detection of the first and second index sequences until base detections of at least a portion of the sample genome sequence have been determined. Therefore, the sequence-to-coverage system 106 does not obtain index sequence data until relatively deeper into running the index.
[0085] In contrast, in some implementations, the sequence-to-coverage system 106 utilizes an index-first workflow that enables the sequence-to-coverage system 106 to identify the genomic sample corresponding to the nucleotide reads before sequencing the reads. FIG. 4The index-first workflow 316 illustrated depicts the sequence-to-coverage system 106 performing an indexing cycle prior to the genome sequencing cycle. As shown, the sequence-to-coverage system 106 determines base detection for paired-end reads in the following order: (i) a first index sequence appended to the sample genome sequence, (ii) a second index sequence appended to the sample genome sequence, (iii) a first nucleotide read corresponding to a first portion of the sample genome sequence, and (iv) a second nucleotide read corresponding to a second portion of the sample genome sequence. In the index-first workflow 316, the sequence-to-coverage system 106 performs paired-end switching between determining the base detection of the first and second nucleotide reads. By utilizing the index-first workflow 316, the sequence-to-coverage system 106 can determine which nucleotide reads originate from which genome samples relatively early within the sequencing run. FIG. 3 The corresponding discussion further details the use of an index-first workflow 316 for the sequence to coverage system 106 according to one or more implementation schemes.
[0086] like FIG. 3 As shown, sequence-to-coverage system 106 performs action 304 to determine the corresponding number of clusters belonging to the corresponding genome sample. Typically, sequence-to-coverage system 106 determines the balance among genome samples within a genome sample pool. More specifically, based on the index sequence, sequence-to-coverage system 106 determines the corresponding number of oligonucleotide clusters belonging to the corresponding genome sample within the genome sample. In some embodiments, sequence-to-coverage system 106 compares the index sequences of nucleotide reads in oligonucleotide cluster 318 with a reference of a known index to determine the genome sample origin of each nucleotide read. Sequence-to-coverage system 106 can then sort oligonucleotide clusters 318 based on the original sample. FIG. 5 As further illustrated, the sequence-to-coverage system 106 determines the number of clusters belonging to each genome sample in the genome pool. FIG. 3 The corresponding discussion further details how the sequence-to-coverage system 106, based on one or more implementation schemes, determines the corresponding number of oligonucleotide clusters belonging to the corresponding genomic sample.
[0087] like FIG. 3Further to the series of actions 300, the series of actions 300 optionally includes an action 306 of determining a filter metric. By determining a filter metric, the sequence-to-coverage system 106 estimates variation in the sequencing data caused by passing through a filter problem. In some implementations, the sequence-to-coverage system 106 determines a filter metric that indicates a subset of the oligonucleotide clusters that satisfy a filtering threshold for signal of the oligonucleotide cluster. The sequence-to-coverage system 106 evaluates the oligonucleotide clusters to identify oligonucleotide clusters that pass a filter. For example, the sequence-to-coverage system 106 can evaluate dim, low quality, or multi-clonal wells or clusters as oligonucleotide clusters that do not pass a filter. As shown, the sequence-to-coverage system 106 determines that the cluster 320 does not satisfy the filtering threshold. The sequence-to-coverage system 106 can aggregate filter data for the oligonucleotide clusters to estimate a subset of the oligonucleotide clusters that satisfy the filtering threshold originating from each genomic sample. FIG. 6 Further to the series of actions 300, the series of actions 300 optionally includes an action 306 of determining a filter metric. By determining a filter metric, the sequence-to-coverage system 106 estimates variation in the sequencing data caused by passing through a filter problem. In some implementations, the sequence-to-coverage system 106 determines a filter metric that indicates a subset of the oligonucleotide clusters that satisfy a filtering threshold for signal of the oligonucleotide cluster. The sequence-to-coverage system 106 evaluates the oligonucleotide clusters to identify oligonucleotide clusters that pass a filter. For example, the sequence-to-coverage system 106 can evaluate dim, low quality, or multi-clonal wells or clusters as oligonucleotide clusters that do not pass a filter. As shown, the sequence-to-coverage system 106 determines that the cluster 320 does not satisfy the filtering threshold. The sequence-to-coverage system 106 can aggregate filter data for the oligonucleotide clusters to estimate a subset of the oligonucleotide clusters that satisfy the filtering threshold originating from each genomic sample. FIG. 3 The corresponding discussion further details how the sequence-to-coverage system 106 determines a filter metric that indicates a subset of the oligonucleotide clusters that satisfy a filtering threshold. In some embodiments, the action 306 is an optional action.
[0088] The series of actions 300 also includes an action 308 of estimating a read coverage level for the genomic samples. By determining the respective number of clusters belonging to the respective genomic samples, the sequence-to-coverage system 106 can estimate variation caused by sample pooling in part. The sequence-to-coverage system 106 can estimate a read coverage level for the genomic samples more accurately based on the respective number of oligonucleotide clusters belonging to the respective genomic samples and a currently selected number of sequencing cycles for the sequencing run. Additionally, in some embodiments, the sequence-to-coverage system 106 estimates the read coverage level based on the filter metric. As shown, the sequence-to-coverage system 106 can generate an estimated read coverage level for a genomic sample by multiplying the number of clusters belonging to the genomic sample by the currently selected number of sequencing cycles (Read Coverage = Number of Clusters x Number of Sequencing Cycles). As described, the currently selected number of sequencing cycles includes a number of sequencing cycles to be performed during the sequencing run.
[0089] As FIG. 3 Further to the sequence-to-coverage system 106, the sequence-to-coverage system 106 can determine an estimated read coverage level for the genomic samples based on the filter metric (Read Coverage = Number of Clusters x Filter Metric x Number of Sequencing Cycles). In some implementations, the sequence-to-coverage system 106 accesses the filter metric for the clusters corresponding to a particular genomic sample. In some examples, the sequence-to-coverage system 106 determines an estimated read coverage level for a genomic sample by multiplying the number of clusters belonging to the genomic sample by the filter metric for the genomic sample and a currently selected number of sequencing cycles.
[0090] Based on the estimated read coverage levels of the genomic samples, the sequence-to-coverage system 106 modifies the sequencing process to meet the target read coverage levels. As FIG. 3 As illustrated, the sequence-to-coverage system 106 performs the action 310 of generating a custom number of sequencing cycles. Specifically, the sequence-to-coverage system 106 generates a custom number of sequencing cycles sufficient to generate nucleotide reads that meet the target read coverage level for each of the genomic samples. Generally, the sequence-to-coverage system 106 can generate a custom number of sequencing cycles by increasing or decreasing the currently selected number of sequencing cycles. For example, the sequence-to-coverage system 106 can utilize the following equation to determine the custom number of sequencing cycles (Nc) ).
[0091]
[0092] wherein represents the custom number of sequencing cycles, represents the read coverage level of the genomic sample with the minimum coverage, and represents the target read coverage level. FIG. 7 and the corresponding discussion further details the sequence-to-coverage system 106 generating a custom number of sequencing cycles and performing a sequencing run in accordance with one or more embodiments.
[0093] FIG. 8 The series of actions 300 as illustrated in the middle further includes the action 312 of determining a custom set of flowcell regions to image. In addition to or in lieu of generating a custom number of sequencing cycles, the sequence-to-coverage system 106 can determine a custom set of flowcell regions from the flowcell to image sufficient to generate nucleotide reads that meet the target read coverage level for each of the genomic samples. In one example, the sequence-to-coverage system 106 utilizes the following equation to determine the custom set of flowcell regions to image:
[0094]
[0095] wherein represents the read coverage level of the genomic sample with the minimum coverage, represents the custom number of sequencing cycles, represents the custom set of flowcell regions to image, represents the total number of flowcell regions in the nucleotide sample flowcell, and represents the target read coverage level. FIG. 4 and the corresponding paragraph illustrates the sequence-to-coverage system 106 determining a custom set of flowcell regions to image in accordance with one or more embodiments of the present disclosure.
[0096] As mentioned, in some implementations, the sequence-to-coverage system 106 performs a subset of sequencing cycles according to the order of index cycles preceding the genomic sequencing cycles. FIG. 4 A sequence-to-coverage system 106 according to one or more embodiments of the present disclosure is illustrated performing a subset of sequencing cycles according to the order of index cycles preceding the genomic sequencing cycles. FIG. 4 A series of actions 400 is illustrated, including an action 402 of determining base calls for a first index sequence, an action 404 of determining base calls for a second index sequence, an action 406 of determining base calls for a first nucleotide read, and an action 408 of determining base calls for a second nucleotide read.
[0097] In some implementations, the sequence-to-coverage system 106 utilizes an index-first workflow to determine the balance of genomic samples within a relatively early genomic sample pool within a sequencing run. By performing index cycles prior to genomic sequencing cycles, the sequence-to-coverage system 106 determines which nucleotide reads belong to which genomic samples and the relative balance of genomic samples. In a non-index-first workflow, index sequence data from both index sequences attached to sample genomic sequences is only available after double-end turning is complete. In contrast, the sequence-to-coverage system 106 can improve efficiency by obtaining index data prior to performing genomic sequencing cycles. Thus, in some implementations, the sequence-to-coverage system 106 can adjust genomic sequencing cycles in a dynamic manner based on index sequence information.
[0098] FIG. 4 A series of actions 400 is illustrated, including an action 402 of determining base calls for a first index sequence. A first index primer 412 is annealed to a primer binding site attached to a sample genomic sequence 410. After the first index primer 412 is annealed, the sequence-to-coverage system 106 determines base calls for a first index sequence 416. As shown, the first index sequence 416 is attached to the sample genomic sequence 410 of a genomic sample. FIG. 4
[0099] As further illustrated in FIG. 4 the sequence-to-coverage system 106 performs an action 404 of determining base calls for a second index sequence. The sequence-to-coverage system 106 anneals a second index primer 418 to a primer binding site attached to the sample genomic sequence 410. The sequence-to-coverage system 106 determines base calls for a second index sequence 420. As further shown, the second index sequence 420 is attached to the 5' end of the sample genomic sequence 410, while the first index sequence 416 is attached to the 7' end of the sample genomic sequence 410. FIG. 5
[0100] After determining the base calls of the first index sequence 416 and the second index sequence 420, the sequence-to-coverage system 106 performs the act 406 of determining base calls of the first nucleotide reads. More specifically, the sequence-to-coverage system 106 determines the base calls of the first nucleotide reads corresponding to the first portion of the sample genomic sequence 410. More specifically, in a paired-end sequencing run, the sample genomic sequence 410 is sequenced from both ends, providing complementary information about the sample genomic sequence 410. As part of performing the act 406, the sequence-to-coverage system 106 anneals the first nucleotide read primer 422 to the read primer binding site, and the sequence-to-coverage system 106 sequences the first portion of the sample genomic sequence 410.
[0101] In some embodiments, after performing the act 406, the sequence-to-coverage system 106 performs a paired-end turn-around. Generally, during paired-end turn-around, the P7 region is cleaved, and all fragments are attached through the P5 region. Prior to paired-end turn-around, the P7 region is annealed to the surface of the flow cell. After paired-end turn-around, the P5 region is attached to the flow cell.
[0102] After paired-end turn-around, the sequence-to-coverage system 106 performs the act 408 of determining base calls of the second nucleotide reads. The sequence-to-coverage system 106 anneals the second nucleotide read primer 424 to the second read primer binding site, and the sequence-to-coverage system 106 sequences the second portion of the sample genomic sequence 410. In some embodiments, the sequence-to-coverage system 106 utilizes specialized reagents as part of the index-first workflow.
[0103] As mentioned, the sequence-to-coverage system 106 can determine, based on the index sequences, the respective number of oligonucleotide clusters belonging to the respective genomic samples. FIG. 5 An example is illustrated of the sequence-to-coverage system 106 determining, in accordance with one or more embodiments of the present disclosure, the respective number of oligonucleotide clusters belonging to the respective genomic samples in a pool of genomic samples.
[0104] After determining the base calls of the index sequences, the sequence-to-coverage system 106 determines which oligonucleotide clusters correspond to each of the genomic samples in the pool of genomic samples. The sequence-to-coverage system 106 can accomplish this through a process known as de-multiplexing. After determining the base calls of the index sequences, the sequence-to-coverage system 106 analyzes the raw sequencing data and assigns each read to its corresponding genomic sample using the index barcodes.
[0105] As FIG. 5As illustrated, the sequence-to-coverage system 106 accesses raw sequencing data that includes an index sequence 504 associated with a sample genomic sequence 518, an index sequence 506 associated with a sample genomic sequence 520, and an index sequence 508 associated with a sample genomic sequence 522. The index sequences 504-508 contain barcodes that serve as unique identifiers for each genomic sample, allowing reads to be distinguished and ordered during demultiplexing. For example, the index sequence 504 indicates that the sample genomic sequence 518 is from genomic sample 1. The index sequence 506 indicates that the sample genomic sequence 520 is derived from genomic sample 2.
[0106] In one or more implementations, the sequence-to-coverage system 106 demultiplexes nucleotide reads by utilizing a reference of known indexes. FIG. 5 A reference of enrolled indexes 514 is illustrated. The sequence-to-coverage system 106 compares the index sequences to known index sequences in the reference of enrolled indexes 514. The reference of enrolled indexes 514 associates each index barcode or sequence with its corresponding genomic sample. For example and as illustrated, the reference of enrolled indexes 514 stores index sequences and their corresponding genomic samples. As shown, a genomic sample can correspond to one or more unique barcodes.
[0107] In some implementations, the sequence-to-coverage system 106 can identify and distinguish between assigned index sequences and unassigned index sequences. Assigned index sequences match an index sequence that is enrolled for a particular run. Unassigned index sequences (e.g., index sequence 508) do not match an index sequence that is enrolled for a sequencing run. The sequence-to-coverage system 106 can identify unassigned index sequences based on determining that a given index sequence does not exist in the reference of enrolled indexes 514. To illustrate, the sequence-to-coverage system 106 can compare the index sequence 508 to enrolled indexes in the reference of enrolled indexes 514 and determine that the index sequence 508 is not in the reference of enrolled indexes 514. In one or more embodiments, the sequence-to-coverage system 106 identifies unassigned index sequences and removes a subset of oligonucleotide clusters corresponding to the unassigned index sequences from data of a sequencing run.
[0108] As mentioned, the sequence-to-coverage system 106 identifies a respective number of oligonucleotide clusters belonging to a respective one of the genomic samples. FIG. 5 A flow cell 502 including a flow cell is illustrated. The flow cell 502 includes a lane 510 that contains a flow cell region 512. The flow cell region 512 can represent a cell of the flow cell. As FIG. 5As shown, flow pool region 512 comprises several oligonucleotide clusters. Each cluster contains multiple copies of the same sample genome sequence. Sequence-to-coverage system 106 identifies clusters corresponding to genome sample 1 and genome sample 2. Sequence-to-coverage system 106 also identifies clusters with unassigned index sequences. As illustrated, sequence-to-coverage system 106 determines that cluster 524 corresponds to an unassigned index sequence that does not match the index sequence registered for the sequencing run.
[0109] When assigning clusters to genomic samples within a genomic sample pool, the sequence-to-coverage system 106 determines the number of oligonucleotide clusters belonging to each genomic sample. More specifically, the sequence-to-coverage system 106 counts the number of clusters belonging to each genomic sample corresponding to an assigned index sequence. FIG. 5 As illustrated in Table 516, the sequence-to-coverage system 106 determines that genome sample 1 corresponds to 275M clusters and genome sample 2 corresponds to 373M clusters.
[0110] Additionally, in some embodiments, the sequence-to-coverage system 106 generates and stores a genome sample map indicating the location of clusters corresponding to each genome sample. For example... FIG. 6 As illustrated, the sequence-to-coverage system 106 generates a genome sample map 526 indicating the location of clusters corresponding to each element in the genome sample. Furthermore, as shown, the sequence-to-coverage system 106 excludes data corresponding to unassigned index sequences from the genome sample map 526. For example, the sequence-to-coverage system 106 removes cluster 524 from the genome sample map 526.
[0111] As described, the sequence-to-coverage system 106 can estimate the read coverage level of a genome sample based on filter metrics. FIG. 6 An example is illustrated of a sequence-to-coverage system 106 according to one or more specific embodiments of this disclosure, which determines filter metrics. The sequence-to-coverage system 106 determines filter metrics for subsets of oligonucleotide clusters that indicate a filtering threshold for a signal satisfying the oligonucleotide cluster. Typically, filter metrics indicate the quality and reliability of sequencing reads generated during sequencing runs.
[0112] like FIG. 6 As shown, the sequence-to-coverage system 106 determines a base detection quality metric 602. More specifically, the sequence-to-coverage system 106 determines a base detection quality metric 602 for a subset of sequencing cycles. For illustration, during each sequencing cycle, the sequence-to-coverage system 106 images clusters within a flowpool region 612 (e.g., a small cell of the flowpool). The sequence-to-coverage system 106 evaluates the signals emitted from the oligonucleotide clusters to determine the base detection quality metric 602.
[0113] In some embodiments, the base call quality metrics 602 include a purity value. The term "purity value" refers to a quality metric used to assess the confidence or purity of a called nucleobase from a sequencing cycle. Specifically, the purity value is a measure of the confidence of a called base at each position within a sequencing read. For example, the purity value can be calculated based on the intensity of the fluorescent signal emitted from an oligonucleotide cluster. The sequence-to-coverage system 106 measures the intensity of each of the four nucleotide-specific fluorescent signals. The sequence-to-coverage system 106 can determine the purity value by determining the ratio of the brightest base intensity divided by the sum of the brightest base intensity and the second brightest base intensity. In some examples, the sequence-to-coverage system 106 can report the purity value as a percentage value ranging from 0-100%.
[0114] As FIG. 6 illustrated, the sequence-to-coverage system 106 utilizes the base call quality metrics 602 and a filter threshold to determine pass filter oligonucleotide clusters. Specifically, the sequence-to-coverage system 106 compares the quality metrics of the clusters to the filter threshold to determine whether the clusters are pass filter clusters. For example, the sequence-to-coverage system 106 compares the quality metrics of each cluster within the flowcell region 612 to the filter threshold. In some examples, the filter threshold includes a purity threshold (e.g., 80%). The sequence-to-coverage system 106 determines that clusters having a purity value that satisfies the purity threshold are eligible as pass filter clusters. As shown, the sequence-to-coverage system 106 determines that clusters 614a, 614b, and 614c all have quality metrics that do not satisfy the filter threshold. More specifically, the purity values of the clusters 614a-614c do not satisfy the purity threshold. Accordingly, the sequence-to-coverage system 106 determines that the clusters 614a-614c are not pass filter clusters. The sequence-to-coverage system 106 determines that clusters 616a-616c include pass filter clusters.
[0115] As mentioned, the sequence-to-coverage system 106 determines base call quality metrics 602 for a subset of sequencing cycles. To improve efficiency, the sequence-to-coverage system 106 utilizes images from early sequencing cycles to assess the reliability and accuracy of base calls within each oligonucleotide cluster. As FIG. 6 shown, the sequence-to-coverage system 106 determines base call quality metrics 602 for the flowcell region 612 within a subset of sequencing cycles. For example, the subset of sequencing cycles can include the first 25 sequencing cycles of a sequencing run. The sequence-to-coverage system 106 determines base call quality metrics 602 for each sequencing cycle within the subset of sequencing cycles. Further, the sequence-to-coverage system 106 determines pass filter clusters within each sequencing cycle. For example, when the sequence-to-coverage system 106 determines that cluster 616b is a pass filter cluster in a first sequencing cycle, the sequence-to-coverage system 106 can determine that cluster 616b is not a pass filter cluster in a second sequencing cycle.
[0116] As FIG. 5 illustrated, the sequence-to-coverage system 106 determines base call quality metrics 602 for clusters originating from each genomic sample by utilizing a genomic sample map 608. The genomic sample map 608 indicates locations of clusters corresponding to each of the genomic samples. For example, the genomic sample map 608 for flowcell region 612 indicates that clusters 614a-614b originate from genomic sample 1 and clusters 616a and 616c originate from genomic sample 2. As illustrated, the genomic sample map 608 also indicates that cluster 616b and cluster 614c result from an unregistered genomic sample. The sequence-to-coverage system 106 can utilize the processes described above with respect to FIG. 6 FIG. 6 to generate the genomic sample map 608.
[0117] By utilizing information from the genomic sample map 608, the sequence-to- coverage system 106 can identify a number of oligonucleotide clusters per genomic sample that satisfy the filter threshold. More specifically, the sequence-to-coverage system 106 utilizes the genomic sample map 608 to determine base call quality metrics for oligonucleotide clusters per genomic sample. By comparing the base call quality metrics to the filter threshold, the sequence-to-coverage system 106 can count a number of oligonucleotide clusters per genomic sample that qualify as oligonucleotide clusters that pass the filter.
[0118] As FIG. 6 illustrated, the sequence-to-coverage system 106 generates a pass-filter map 604. The sequence-to-coverage system 106 aggregates the base call quality metrics 602 across the subset of sequencing cycles to generate the pass-filter map 604. Generally, the pass-filter map 604 provides information about results of quality filtering applied to oligonucleotide clusters of the subset of sequencing cycles. The pass-filter map 604 indicates a percentage of clusters at locations across the subset of sequencing cycles that satisfy the filter threshold. For example, the sequence-to-coverage system 106 determines a percentage of clusters that pass the filter for each cluster in the flowcell region 612. For example, the sequence-to-coverage system 106 determines that 20% of cluster 614a includes clusters that pass the filter across the subset of sequencing cycles. The sequence-to-coverage system 106 performs this determination for the remaining clusters within the flowcell region 612. In some implementations, the sequence-to-coverage system 106 also indicates within the pass-filter map 604 a genomic sample corresponding to each cluster.
[0119] The sequence-to-coverage system 106 also aggregates information per genomic sample to generate a filter metric 610. The filter metric 610 indicates a subset of oligonucleotide clusters that satisfy the filter threshold for a signal of the oligonucleotide clusters. As FIG. 6 illustrated, the filter metric includes a percentage of clusters for a genomic sample that satisfy the filter threshold. For example and as FIG. 5As illustrated, the sequence-to-coverage system 106 determines that 83% of the clusters corresponding to genomic sample 1 satisfy the filtering threshold. In some implementations, the sequence-to-coverage system 106 determines a filter metric by the percentage of clusters corresponding to a genomic sample that pass the filter. For example, the sequence-to-coverage system 106 can average the percentage of clusters passing the filter corresponding to a genomic sample.
[0120] Further, in some embodiments, the sequence-to-coverage system 106 can utilize the filter metric 610 and the respective number of oligonucleotide clusters belonging to a respective genomic sample to determine the number of oligonucleotide clusters passing the filter for each genomic sample. For example, the sequence-to-coverage system 106 can utilize the filter metric 610 and the respective number of oligonucleotide clusters belonging to a respective genomic sample to determine the number of oligonucleotide clusters passing the filter for each genomic sample. FIG. 6 The described process to determine the number of clusters corresponding to a given genomic sample. The sequence-to-coverage system 106 multiplies the number of clusters for a given genomic sample by the percentage of clusters for the given genomic sample that satisfy the filtering threshold. As FIGS. 7A-7B As illustrated, the sequence-to-coverage system 106 determines that 275M clusters correspond to genomic sample 1. Based on determining that 83% of the clusters for genomic sample 1 pass the filter using the pass filter graph 606, the sequence-to-coverage system 106 determines that the number of clusters passing the filter for genomic sample 1 is equal to.83 x 275M or 228M.
[0121] As mentioned, the sequence-to-coverage system 106 can generate a custom number of sequencing cycles sufficient to generate nucleotide reads that satisfy a target read coverage level for each of the genomic samples. FIGS. 7A-7B A sequence-to-coverage system 106 according to one or more embodiments of the present disclosure is illustrated generating a custom number of sequencing cycles to satisfy a target read coverage level and performing a sequencing run. By estimating the read coverage level of the genomic samples, the sequence-to-coverage system 106 can adjust the number of sequencing cycles to ensure that all genomic samples receive at least the target read coverage level. FIG. 7A A series of actions 700 is illustrated, including an action 702 to start a sequencing run, an action 704 to determine base calls for index sequences, an action 706 to determine filter metrics, an action 710 to generate a custom number of sequencing cycles, and an action 712 to perform the sequencing run until the custom number of sequencing cycles is completed.
[0122] FIG. 7AThe series of actions 700 is further illustrated as including an action 702 of starting a sequencing run. In some implementations, as part of starting the sequencing run, the sequence-to-coverage system 106 determines a target read coverage level. For example, the sequence-to-coverage system 106 can provide a target read coverage level selection element for display via a client device (e.g., the client device 114). The sequence-to-coverage system 106 can receive user input indicating the target read coverage level. In some implementations, the sequence-to-coverage system 106 automatically determines the target read coverage level. In some implementations, the sequence-to-coverage system 106 determines a target read coverage level of 40x.
[0123] As FIG. 7A The sequence-to-coverage system 106 is further illustrated as performing an action 704 of determining base calls for index sequences. As previously described, by determining base calls for index sequences, the sequence-to-coverage system 106 determines variability between samples relatively early in a sequencing run. The sequence-to-coverage system 106 can utilize a non-index-first workflow and an index-first workflow at early stages of a sequencing cycle. More specifically, the sequence-to-coverage system 106 can improve efficiency of a sequencing run by utilizing an index-first workflow. As mentioned, the sequence-to-coverage system 106 can determine base calls for index sequences for a subset of sequencing cycles. For example, the sequence-to-coverage system 106 can determine base calls for index sequences for the first 5, 10, 25, etc. sequencing cycles of a sequencing run.
[0124] FIG. 6 The sequence-to-coverage system 106 is also illustrated as performing an action 706 of determining filter metrics. As previously described with respect to FIG. 7A the sequence-to-coverage system 106 determines filter metrics for a subset of oligonucleotide clusters that indicate a filter threshold is met for signal of the oligonucleotide cluster. Like determining base calls for index sequences, the sequence-to-coverage system 106 determines filter metrics for a subset of sequencing cycles. In some implementations, the sequence-to-coverage system 106 determines filter metrics for a second subset of sequencing cycles that is different than a first subset of sequencing cycles used to perform index cycles prior to a genomic sequencing cycle. For example, the sequence-to-coverage system 106 can determine base calls for index sequences for the first 10, 15, 20, 25, etc. sequencing cycles of a sequencing run. In some embodiments, the action 706 includes an optional action.
[0125] In some implementations, the sequence-to-coverage system 106 also determines PhiX loss early in a sequencing run. PhiX refers to a standard control library used in sequencing runs to monitor the sequencing process and assess sequencing platform performance. In some implementations, the PhiX control library is added to a sequencing run as a control sample. The amount of PhiX can be a small percentage (e.g., 1-2%) of the input sample. During PhiX alignment, the sequence-to-coverage system 106 maps nucleotide reads to the PhiX genome to determine the amount of PhiX loss. More specifically, PhiX loss occurs when the proportion of nucleotide reads originating from the PhiX control library is significantly reduced compared to the expected or anticipated amount. Larger PhiX loss can indicate problems with various parameters such as cluster density, signal strength, and base call accuracy. The sequence-to-coverage system 106 can determine PhiX loss early in a sequencing run by leveraging index and filter metric data. Examples of determining PhiX loss are also described in U.S. Patent No. 9,574,226 B2, the disclosure of which is incorporated by reference herein in its entirety.
[0126] As FIG. 7B illustrated, the sequence-to-coverage system 106 performs the act 708 of estimating a read coverage level based on determining base calls of the index sequence. For example, the sequence-to-coverage system 106 leverages the following equation to generate an estimated read coverage level for a given sample:
[0127]
[0128] wherein represents the estimated read coverage level for the given sample, “number of clusters” represents the number of clusters originating from the given sample, and “currently selected number of sequencing cycles” refers to the number of sequencing cycles expected within the sequencing run.
[0129] In some implementations, as part of performing the act 708, the sequence-to-coverage system 106 also leverages the filter metric determined as part of the act 706. The sequence-to-coverage system 106 can leverage the following equation to generate an estimated read coverage level for a given sample:
[0130]
[0131] wherein represents the estimated read coverage level for the given sample, “number of clusters” represents the number of clusters originating from the given sample, “filter metric” refers to the proportion or percentage of clusters produced by the given sample that satisfy a filtering threshold, and “currently selected number of sequencing cycles” refers to the number of sequencing cycles expected within the sequencing run. By leveraging index data and filter metric data, the sequence-to-coverage system 106 can capture about 80% of yield variation.
[0132] Based on the estimated read coverage levels, the sequence-to-coverage system 106 performs an action 710 of generating a customized number of sequencing cycles. The sequence-to-coverage system 106 can adjust the total number of sequencing cycles within a sequencing run. For example, if a genome sample is under-covered, the sequence-to-coverage system 106 can increase the number of sequencing cycles relative to the currently selected number of sequencing cycles. Alternatively, if a sequencing run can produce excess data, the sequence-to-coverage system 106 can decrease the total number of sequencing cycles relative to the currently selected number of sequencing cycles.
[0133] In one or more embodiments, the sequence-to-coverage system 106 utilizes the following equation to generate a customized number of sequencing cycles:
[0134]
[0135] wherein represents the customized number of sequencing cycles, represents the read coverage level of the genome sample with the lowest estimated read coverage level, and represents a target read coverage level. In some implementations, the sequence-to-coverage system 106 generates a customized number of sequencing cycles for a sequencing run by increasing or decreasing a preset number of sequencing cycles for the sequencing run.
[0136] As shown, the sequence-to-coverage system 106 utilizes data determined during the primary analysis to determine a customized number of sequencing cycles. More specifically, the sequence-to-coverage system 106 determines the customized number of sequencing cycles prior to completion of a sequencing run. In some implementations, the sequence-to-coverage system 106 can determine the customized number of sequencing cycles at a sequencing device (e.g., the sequencing device 108) or a local server device (e.g., the local server device 102). More specifically, the sequence-to-coverage system 106 can determine the customized number of sequencing cycles during the primary analysis, rather than the secondary analysis, which typically occurs at a server device (e.g., the server device 110). By utilizing data obtained during an early stage of a sequencing run, the sequence-to-coverage system 106 can efficiently determine a customized number of sequencing cycles.
[0137] FIG. 7B The sequence-to-coverage system 106 is illustrated as performing an action 712 of performing a sequencing run until completion of the customized number of sequencing cycles. More specifically, the sequence-to-coverage system 106 causes a sequencing device to perform the customized number of sequencing cycles. For example, the sequence-to-coverage system 106 can cause a fluidic device to perform additional sequencing cycles or fewer sequencing cycles based on the customized number of sequencing cycles.
[0138] FIG. 7BGraph 718 depicting over-sequencing results generated by existing systems using a custom number of sequencing cycles and graph 720 depicting results generated by the sequence-to-coverage system 106 using a custom number of sequencing cycles are illustrated. The x-axis of the graph 718 and the graph 720 represents the number of sequencing cycles. The y-axis of the graph 718 and the graph 720 represents the percentage of genomic samples that have reached the target read coverage level (40x).
[0139] As shown in the graph 718, about 95% of the genomic samples have been sequenced to the target read coverage level at 2x150 sequencing cycles. As shown, most of the genomic samples are over-sequenced at 2x150 sequencing cycles. Further, and as previously mentioned, about 5% of the genomic samples are still under-sequenced and have not reached the target read coverage level at 2x150 sequencing cycles.
[0140] As FIG. 7B illustrated, the sequence-to-coverage system 106 can adjust parameters of a sequencing run to improve efficiency. More specifically, the sequence-to-coverage system 106 not only considers the average read coverage level, the sequence-to-coverage system 106 also ensures that difficult-to-map genomic regions (e.g., repetitive regions) are not negatively impacted by reducing the number of sequencing cycles. Thus, in some implementations, the sequence-to-coverage system 106 evaluates a minimum number of sequencing cycles and a maximum number of sequencing cycles before the relevant metric for difficult-to-map regions starts to drop. Further, the sequence-to-coverage system 106 can design flow cell (FC) capacity and the number of genomic samples within a pool such that the maximum success rate is achieved with a minimum number of default sequencing cycles.
[0141] As FIG. 8As shown, the sequence-to-coverage system 106 reduces the size of the flow cell or increases the number of genomic samples in the genomic sample pool. The sequence-to-coverage system 106 can reduce the size of the flow cell by reducing the number of clusters per nucleotide sample substrate (e.g., flow cell). For example, the sequence-to-coverage system 106 can reduce the number of nanopores per flow cell. In some implementations, the sequence-to-coverage system 106 reduces the size of the flow cell by determining a reduced set of flow cell regions to image. As a result, approximately 50% of the genomic samples are sequenced to the target read coverage level at 2x 150 sequencing cycles. The sequence-to-coverage system 106 can determine a custom number of sequencing cycles that falls between a minimum number of sequencing cycles (2x 135c) and a maximum number of sequencing cycles (2x 185c). In some implementations, the sequence-to-coverage system 106 increases or decreases a preset number of sequencing cycles (e.g., a default number of sequencing cycles) within the minimum number of sequencing cycles and the maximum number of sequencing cycles. For example, the sequence-to-coverage system 106 can decrease or increase the preset number of sequencing cycles by 15, 35, etc.
[0142] As further illustrated in FIG. 8 As further illustrated in
[0143] In general, the sequence-to-coverage system 106 can determine a minimum number of sequencing cycles to ensure a baseline coverage for all genomic samples. The sequence-to- coverage system 106 can determine the minimum number of sequencing cycles based on a workflow or purpose of the sequencing run. The sequence-to-coverage system 106 can determine different minimum numbers of sequencing cycles for different assays. For example, some assays, such as enrichment assays, require a lower read coverage level. Other sequencing assays for sequencing hard-to-map genomic regions can require a higher read coverage level.
[0144] By generating a custom number of sequencing cycles, the sequence-to-coverage system 106 improves the efficiency of sequencing runs. Relative to existing systems, the sequence-to- coverage system 106 reduces the number of sequencing cycles required to meet a target read coverage level. For example, existing systems require an average of 316 sequencing cycles. In contrast, the sequence-to-coverage system 106 can achieve 120 Gb of coverage in 226 sequencing cycles, which is 28% fewer sequencing cycles than existing systems. The reduction in sequencing cycles also reduces the amount of sequencing reagents required for a sequencing run. More specifically, the sequence-to-coverage system 106 can perform sequencing runs that require 28% less reagents than existing systems. The sequence-to-coverage system 106 can also perform sequencing runs that require 11.5% less total material than existing systems. In addition to sequencing reagents, total material can also include library preparation kits, flow cells, cluster amplification material, and other processing materials. Additionally, the total run time of sequencing runs performed by the sequence-to-coverage system 106 is, on average, 90 minutes shorter than existing sequencing runs on existing sequencing devices. However, the run time savings depends on the per-cycle time metric of the sequencing device, and thus, the run time savings can be greater for sequencing devices with longer per-cycle time metrics and less for sequencing devices with shorter per-cycle time metrics. Moreover, the sequence-to-coverage system 106 improves genomic sample success rates from 96% to 99% by performing sequencing runs with a custom number of sequencing cycles.
[0145] In addition to or instead of generating a custom number of sequencing cycles, the sequence-to-coverage system 106 can determine a custom set of flow cell regions to image sufficient to generate nucleotide reads that meet a target read coverage level for each genomic sample. FIG. 7A An example of the sequence-to-coverage system 106 determining a custom set of flow cell regions to image is illustrated in accordance with one or more embodiments of the present disclosure. FIG. 8A series of actions 800 are illustrated, including an action 802 to start a sequencing run, an action 804 to determine base calls for an index sequence, an action 806 to determine filter metrics, an action 808 to estimate read coverage levels, an action 810 to determine a customized set of flow cell regions to image, and an action 812 to perform the sequencing run by capturing images of the customized set of flow cell regions.
[0146] The sequence-to-coverage system 106 can determine to utilize one or both of generating a customized number of sequencing cycles and determining a customized set of flow cell regions to image. In some applications, the sequence-to-coverage system 106 can determine to keep the number of sequencing cycles constant (e.g., 2x150c) and instead improve efficiency by adjusting the set of flow cell regions to image during a sequencing run. In other applications, the sequence-to-coverage system 106 performs a sequencing run with a customized number of sequencing cycles without adjusting the flow cell regions to image during those sequencing cycles. In some applications, the sequence-to-coverage system 106 determines to utilize both tools to improve efficiency of a sequencing run. More specifically, the sequence-to-coverage system 106 can adjust both the number of sequencing cycles and the number of flow cell regions to image within the same sequencing run.
[0147] The sequence-to-coverage system 106 can determine a customized set of flow cell regions to image based on the type of imaging process utilized by various sequencing devices. For example, the sequence-to-coverage system 106 can generate a customized number of sequencing cycles for a sequencing device with a fast imaging process. Some sequencing devices utilize a fast scanning process and are best improved for efficiency by reducing the number of sequencing cycles. Some sequencing devices utilize a slower imaging process. For example, sequencing devices that utilize a stop-and-rephotograph imaging system require more time in the imaging step than sequencing devices that rely on scanning. Thus, the sequence-to-coverage system 106 can improve turnaround time by adjusting the set of flow cell regions to image.
[0148] Actions 802-808 are similar to the actions 702-708 described above with reference to FIG. 7. FIG. 8 As with the actions 702-708 illustrated in FIG. 7, the sequence-to-coverage system 106 estimates read coverage levels for each genomic sample within a pool of genomic samples. The following paragraphs detail the changes between actions 802-808 and actions 702-708.
[0149] As FIG. 8As shown, the sequence-to-coverage system 106 performs an action 804 of determining base calls for the indexed sequences. In some implementations, the sequence-to- coverage system 106 determines a respective number of clusters belonging to the respective genomic sample for each flowcell region. More specifically, the sequence-to-coverage system 106 determines a respective oligonucleotide cluster belonging to the respective genomic sample for each flowcell region. The sequence-to-coverage system 106 stores the balance of genomic samples within each flowcell region. For example, the sequence-to-coverage system 106 can store the flowcell region data in a genomic sample map. The sequence-to-coverage system 106 can utilize the index data to identify flowcell regions that, when imaged, can compensate for imbalances in the genomic sample representation.
[0150] The sequence-to-coverage system 106 also performs an action 806 of determining filter metrics. Generally, the sequence-to-coverage system 106 stores filter metric data for each flowcell region. For example, in some implementations, the sequence-to-coverage system 106 stores a percent filter metric for each flowcell region. The sequence-to-coverage system 106 can utilize the stored filter metric data for each flowcell region to apply different weights to different flowcell regions. For example, the sequence-to-coverage system 106 can determine to image flowcell regions corresponding to higher %PFs compared to flowcell regions corresponding to lower %PFs.
[0151] FIGS. 9A-9B The sequence-to-coverage system 106 is illustrated as performing an action 810 of determining a customized set of flowcell regions to image. In some examples, the sequence-to-coverage system 106 utilizes the following equation to determine the customized set of flowcell regions to image during a sequencing cycle:
[0152]
[0153] wherein represents a read coverage level for the genomic sample with the lowest expected read coverage level, represents a number of customized sequencing cycles, represents a number of flowcell regions in the customized set of flowcell regions to image, represents a total number of flowcell regions in the flowcell, and represents a target read coverage level. In implementations in which the sequence-to- coverage system 106 performs a constant number of sequencing cycles, represents a constant number of sequencing cycles (e.g., 2 x 150c).
[0154] In addition to determining the raw number of flow cell regions to image, the sequence-to-coverage system 106 can also identify specific flow cell regions within the flow cell to image during the sequencing cycles. The sequence-to-coverage system 106 can utilize inter-region variation to improve the read coverage level of specific genomic samples and / or select flow cell regions that perform best. For example, the sequence-to-coverage system 106 can image flow cell regions with more clusters belonging to a given genomic sample to improve the read coverage level of the given genomic sample. Additionally or alternatively, the sequence-to-coverage system 106 can image flow cell regions with higher filter metrics and / or stop imaging flow cell regions with lower filter metrics.
[0155] The series of actions 800 includes an action 812 of performing a sequencing run by capturing images of a customized set of flow cell regions. FIG. 9A The flow cell 816 is illustrated as including a lane composed of flow cell regions 818. In some implementations, the flow cell regions include a well of the flow cell. As illustrated, the sequence-to-coverage system 106 captures images of a customized set of flow cell regions 814 during a sequencing cycle of a sequencing run. In other embodiments, the customized set of flow cell regions includes a plurality of flow cell regions to image within the lane 820. Some flow cells include addressable lanes, where a specific genomic sample is assigned to a specific lane of the flow cell. The sequence-to-coverage system 106 can generally determine to image a customized number of flow cell regions within the lane 820 to improve the read coverage level of the genomic sample corresponding to the lane 820.
[0156] Imaging a customized set of flow cell regions produces several improvements over existing systems. By imaging a customized set of flow cell regions, the sequence-to-coverage system 106 can reduce the number of sequenced flow cell regions from 324 to 233, a 28% reduction. The sequence-to-coverage system 106 also reduces the time required to complete a sequencing run relative to existing systems. For example, the sequence-to-coverage system 106 reduces the run time from 19 hours to 16 hours. The sequence-to-coverage system 106 improves efficiency with zero to very small computational cost.
[0157] The sequence-to-coverage system 106 improves the efficiency of a sequencing run by performing the sequencing run until a customized number of sequencing cycles is completed. FIG. 9B The improvement in sequencing efficiency resulting from performing a customized number of sequencing cycles is illustrated in accordance with one or more embodiments of the present disclosure. Specifically, FIGS. 9A-9B The improvement in efficiency is illustrated given a poor sample merge, and FIG. 9A The improvement in efficiency is illustrated given optimal sequence performance. FIG. 9A The graph in FIG. 1 illustrates simulated data.
[0158] FIG. 9B Figure 902 illustrates the read coverage levels of genome samples sequenced using existing systems, and Figure 904 illustrates the read coverage levels of genome samples sequenced using the sequence-to-coverage system 106 under poor sample pooling conditions. Figure 902 depicts the read coverage levels of genome samples 906a, 906b, and 906c. Figure 904 depicts the read coverage levels of genome samples 908a, 908b, and 908c.
[0159] like FIG. 9B As shown, genome sample 906a failed to achieve the target read coverage level of 40x after 2 × 150 sequencing cycles. Under-sequencing of genome sample 906a might require additional sequencing runs on existing systems to obtain sufficient data for genome sample 906a. In contrast, because the sequence-to-coverage system 106 determines and executes a customized number of sequencing cycles, it does not under-sequence any genome sample. For example, the sequence-to-coverage system 106 executes sequencing runs until a customized number of 2 × 160 sequencing cycles are completed. By increasing the number of sequencing cycles, the sequence-to-coverage system 106 ensures that samples 908a-908c are all sequenced to the target read coverage level.
[0160] FIG. 10 Chart 910 illustrates the read coverage levels of genome samples sequenced using existing systems, and Chart 912 illustrates the read coverage levels of genome samples sequenced using a coverage system 106 with sequences having optimal sequence performance. For example, variation is reduced because genome samples can be more balanced, and 80% or more clusters pass through the filter. Chart 910 depicts the read coverage levels of genome samples 914a, 914b, and 914c. Chart 912 depicts the read coverage levels of genome samples 916a, 916b, and 916c. As illustrated, genome samples 914a-914c and 916a-916c show the smallest variations in read coverage levels.
[0161] like FIG. 10As shown, the existing system over-sequences the genomic samples 914a-914c. For example, the existing system sequences the genomic samples 914a-914c to a read coverage level of about 72x when performing 2x150 sequencing cycles compared to a target read coverage level of 40x. In contrast and as shown in the chart 912, the sequence-to-coverage system 106 determines a custom number of 2x120 sequencing cycles, which is less than the default 2x150 sequencing cycles. By reducing the number of sequencing cycles, the sequence-to-coverage system 106 sequences the genomic samples 916a-916c to just meet and almost not exceed the target read coverage level of 40x.
[0162] The sequence-to-coverage system 106 also improves the efficiency of sequencing runs by imaging a custom set of flow cell regions during sequencing cycles. FIG. 10 An improvement in sequencing efficiency resulting from imaging a custom set of flow cell regions during sequencing cycles according to one or more embodiments of the present disclosure is illustrated. FIG. 10 The chart illustrated in the middle depicts simulated data.
[0163] FIG. 11 A chart 1002 depicting read coverage levels of genomic samples sequenced by an existing system and a chart 1004 depicting read coverage levels of genomic samples sequenced by the sequence-to-coverage system 106 are illustrated. The chart 1002 depicts read coverage levels of genomic samples 1006a, 1006b, and 1006c after imaging 100 flow cell regions for 2x150 sequencing cycles. The chart 1004 depicts read coverage levels of genomic samples 1008a, 1008b, and 1008c after imaging 70 flow cell regions for 2x150 cycles.
[0164] As FIG. 11 As shown, the existing system over-sequences the genomic samples 914a-914c. For example, the existing system sequences the genomic samples 1006a-1006c to a read coverage level of about 72x when imaging 100 flow cell regions during sequencing cycles compared to a target read coverage level of 40x. In contrast and as shown in the chart 1004, the sequence-to-coverage system 106 images a custom set of 70 flow cell regions during sequencing cycles. By reducing the number of flow cell regions imaged, the sequence-to-coverage system 106 sequences the genomic samples 916a-916c to just meet and almost not exceed the target read coverage level of 40x.
[0165] Aspects of the present disclosure generally relate to devices, systems, and methods that provide biological or chemical analysis. Various protocols in biological or chemical research involve conducting a large number of controlled reactions on a local support surface or within a predefined reaction chamber. The specified reactions can then be observed or detected, and subsequent analysis can help identify or reveal properties of chemicals involved in the reactions. For example, in some multiplex assays, unknown analytes with identifiable markers (e.g., fluorescent markers) can be exposed to thousands of known probes under controlled conditions. Each known probe can be placed into a corresponding well of a flowcell lane. Observing any chemical reactions that occur between the known probes and the unknown analytes within the well can help identify or reveal properties of the analytes. Other examples of such protocols include known DNA sequencing processes, such as sequencing by synthesis (SBS) or cyclic array sequencing.
[0166] While a variety of devices, systems, and methods have been made and used to perform biological or chemical analysis, it is believed that no one prior to the inventors has made or used the devices and techniques described herein.
[0167] FIG. 12 An illustration of an example of a system (1100) that can be used to analyze one or more samples of interest is shown. In some implementations, the samples can include one or more clusters of nucleotides (e.g., DNA) that have been linearized to form single-stranded DNA (sstDNA). In the illustrated implementation, the system (1100) is configured to receive a flowcell cartridge assembly (1102) that includes a flowcell assembly (1103) and a sample cartridge (1104). The system (1100) includes a flowcell receptacle (1122) that receives the flowcell cartridge assembly (1102), a vacuum chuck (1124) that supports the flowcell assembly (1103), and a flowcell interface (1126) for establishing fluidic coupling between the system (1100) and the flowcell assembly (1103). The flowcell interface (1126) can include one or more manifolds. The system (1100) also includes a pipettor manifold assembly (1106), a sample loading manifold assembly (1108), and a pump manifold assembly (1110). The system (1100) also includes a drive assembly (1112), a controller (1114), an imaging system (1116), and a waste reservoir (1118). The controller (1114) is electrically and / or communicatively coupled to the drive assembly (1112) and the imaging system (1116); and is configured to cause the drive assembly (1112) and / or the imaging system (1116) to perform various functions as disclosed herein.
[0168] In the present example, the flow cell assembly (1103) includes a flow cell (1128) having a channel (1130) and defining a plurality of first openings (1132) fluidically coupled to the channel (1130) and arranged on a first side (1134) of the channel (1130). The flow cell (1128) also includes a plurality of second openings (1136) fluidically coupled to the channel (1130) and arranged on a second side (1138) of the channel (1130). Thus, fluid can flow through the flow cell (1128) via the channel. While the flow cell (1128) is shown as including one channel (1130), the flow cell (1128) can include two or more channels (1130). The flow cell assembly (1103) also includes a flow cell manifold assembly (1140) coupled to the flow cell (1128) and having a first manifold fluidic line (1142) and a second manifold fluidic line (1144). The flow cell manifold assembly (1140) can be in the form of a laminate including a plurality of layers, as discussed in greater detail below.
[0169] In the illustrated implementation, the first manifold fluidic line (1142) has first fluidic line openings (1146) and is fluidically coupled to each of the plurality of first openings (1132) of the flow cell (1128); and the second manifold fluidic line (1144) has second fluidic line openings (1148) and is fluidically coupled to each of the second openings (1136). As shown, the flow cell assembly (1103) includes a gasket (1150) coupled to the flow cell manifold assembly (1140) and fluidically coupled to the fluidic line openings (1146, 1148). In some implementations in which the flow cell (1128) includes a plurality of channels (1130), the flow cell manifold assembly (1140) can include additional fluidic lines (1152) coupling the first fluidic line openings (1146) to a single manifold port (1154). In such implementations, a single gasket (1150) can be coupled to the flow cell manifold assembly (1140) that encloses the manifold port (1154) and is in fluid communication with the plurality of channels (1130). In operation, the flow cell interface (1126) engages with the corresponding gasket (1150) to establish fluidic coupling between the system (1100) and the flow cell (1128). Engagement between the flow cell interface (1126) and the gasket (1150) reduces or eliminates fluid leakage between the flow cell interface (1126) and the flow cell (1128).
[0170] In the illustrated implementation, the first manifold fluidic line (1142) has a portion (1156) that is substantially parallel to the longitudinal axis (1158) of the channel (1130); and the second manifold fluidic line (1144) has a portion (1160) that is substantially parallel to the longitudinal axis (1158) of the channel (1130). Additionally, the first manifold fluidic line (1142) is shown as being at least partially adjacent to a first end (1162) of the flow cell (1128) and spaced apart from a second end (1164) of the flow cell (1128); and the second manifold fluidic line (1144) is shown as being at least partially adjacent to the second end (1164) of the flow cell (1128) and spaced apart from the first end (1162). However, other arrangements of the manifold fluidic lines (1142, 1144) can prove suitable.
[0171] In the illustrated implementation, the system (1100) includes a sample cartridge receptacle (1166) that receives a sample cartridge 1104 that carries one or more samples of interest (e.g., analytes). The system (1100) also includes a sample cartridge interface (1168) that establishes a fluidic connection with the sample cartridge (1104). The sample loading manifold assembly (1108) includes one or more sample valves (1170). The pump manifold assembly (1110) includes one or more pumps (1172), one or more pump valves (1174), and a cache (1176). The valves (1170, 1174) and pumps (1172) can take any suitable form. The cache (1176) can include a serpentine cache and can temporarily store one or more reaction components, e.g., during bypass maneuvering of the system (1100). While the cache (1176) is shown as being included in the pump manifold assembly (1110), the cache (1176) can alternatively be located elsewhere (e.g., in the aspirator manifold assembly (1106) or in another manifold downstream of the bypass fluidic line (1178), etc.).
[0172] The sample loading manifold assembly (1108) and the pump manifold assembly (1110) cause the one or more samples of interest to flow from the sample cartridge (1104) through a fluidic line (1180) to the flow cell cartridge assembly (1102). In some implementations, the sample loading manifold assembly (1108) can individually load / address each channel (1130) of the flow cell (1128) with a respective sample of interest. The process of loading the channels 1130 with samples of interest can occur automatically using the system (1100). As FIG. 11As shown, a sample cartridge (1104) and a sample loading manifold assembly (1108) are positioned downstream of the flow cell cartridge assembly (1102). In the illustrated implementation, the sample loading manifold assembly (1108) is coupled between the flow cell cartridge assembly (1102) and a pump manifold assembly (1110). To draw a sample of interest from the sample cartridge (1104) and toward the pump manifold assembly (1110), the sample valve (1170), the pump valve (1174), and / or the pump (1172) can be selectively actuated to push the sample of interest toward the pump manifold assembly (1110). The sample cartridge (1104) can include a plurality of sample reservoirs that are fluidly accessible via corresponding sample valves (1170). To individually flow the sample of interest toward a channel (1130) of a flow cell (1128) and away from the pump manifold assembly (1110), the sample valve (1170), the pump valve (1174), and / or the pump (1172) can be selectively actuated to push the sample of interest toward the flow cell cartridge assembly (1102) and into a respective channel (1130) of the flow cell (1128).
[0173] The drive assembly (1112) interfaces with the aspirator manifold assembly (1106) and the pump manifold assembly (1110) to flow one or more reagents that interact with a sample within the flow cell (1128). In some cases, a reversible terminator is attached to a reagent to allow a single nucleotide to bind to a growing DNA strand. In some such implementations, one or more nucleotides have a unique fluorescent label that emits a color when excited. The color (or absence of color) is used to detect the corresponding nucleotide. In the illustrated implementation, an imaging system (1116) excites one or more of the identifiable labels (e.g., fluorescent labels) and then obtains image data of the identifiable labels. The labels can be excited by incident light and / or laser light, and the image data can include one or more colors emitted by the respective labels in response to the excitation. The image data (e.g., detection data) can be analyzed by the system (1100). Examples of features and functionality that can be incorporated into the imaging system (1116) are described in more detail below.
[0174] After obtaining the image data, the drive assembly (1112) interfaces with the aspirator manifold assembly (1106) and the pump manifold assembly (1110) to flow another reaction component (e.g., a reagent) through the flow cell (1128), which is then received by the waste reservoir (1118) via the main waste fluid line (1182) and / or otherwise depleted by the system (1100). Some reaction components can perform a wash operation that cleaves fluorescent labels and reversible terminators from the sstDNA. Subsequently, the sstDNA can be prepared for another cycle.
[0175] A main waste fluid line (1182) is coupled between the pump manifold assembly (1110) and a waste reservoir (1118). In some implementations, the pump (1172) and / or the pump valve (1174) of the pump manifold assembly (1110) selectively flow reaction components from the flow cell cartridge assembly (1102) through the fluid line (1180) and the sample loading manifold assembly (1108) to the main waste fluid line (1182). The flow cell cartridge assembly (1102) is coupled to a central valve (1184) via a flow cell interface (1126). The central valve (1184) is coupled to the flow cell interface (1126) via a fluid line (1185). An auxiliary waste fluid line (1186) is coupled to the central valve (1184) and the waste reservoir (1118). In some implementations, the auxiliary waste fluid line (1186) receives excess fluid from a sample of interest of the flow cell cartridge assembly (1102) via the central valve (1184) and flows the excess fluid of the sample of interest to the waste reservoir (1118) when the sample of interest is being reverse loaded into the flow cell (1128), as described herein.
[0176] The aspirator manifold assembly (1106) includes a shared line valve (1188) and a bypass valve (1190). The shared line valve (1188) can be referred to as a reagent selection valve. The central valve (1184) and the valves (1188, 1190) of the aspirator manifold assembly (1106) can be selectively actuated to control fluid flow through the fluid lines (1192, 1194, 1196). The aspirator manifold assembly (1106) can be coupled to a corresponding number of reagent reservoirs (1198) via reagent aspirators (1200). The reagent reservoirs (1198) can hold fluid (e.g., reagents and / or another reaction component). In some implementations, the aspirator manifold assembly (1106) includes a plurality of ports. Each port of the aspirator manifold assembly (1106) can receive one of the reagent aspirators (1200). The reagent aspirators (1200) can be referred to as fluid lines. Some forms of the reagent aspirators (1200) can include an array of aspirator tubes that extend downward along a z-dimension from a port in a body of the aspirator manifold assembly (1106). The reagent reservoirs (1198) can be disposed in a cartridge, and the tubes of the reagent aspirators (1200) can be configured to be inserted into the corresponding reagent reservoirs (1198) in the reagent cartridge such that liquid reagents can be aspirated from each reagent reservoir (1198) into the aspirator manifold assembly (1106).
[0177] The shared line valve (1188) of the aspirator manifold assembly (1106) is coupled to the central valve (1184) via a shared reagent fluid line (1196). Different reagents can flow through the shared reagent fluid line (1196) at different times. In some versions, the pump manifold assembly (1110) can draw a wash buffer through the shared reagent fluid line (1196), the central valve (1184), and the flow cell cartridge assembly (1102) when a flush operation is performed prior to changing between one reagent and another reagent.
[0178] The bypass valve (1190) of the aspirator manifold assembly (1106) is coupled to the central valve (1184) via a dedicated reagent fluid line (1194, 1196). Each of the dedicated reagent fluid lines (1194, 1196) can be associated with a single reagent. Fluids that can flow through the dedicated reagent fluid lines (1194, 1196) can be used during sequencing operations and can include cleavage reagents, integration reagents, scanning reagents, cleavage washes, and / or wash buffers.
[0179] The bypass valve (1190) is also coupled to the cache (1176) of the pump manifold assembly (1110) via a bypass fluid line (1178). One or more reagent priming operations, hydration operations, mixing operations, and / or transfer operations can be performed using the bypass fluid line (1178). The priming operations, hydration operations, mixing operations, and / or transfer operations can be performed independent of the flow cell cartridge assembly (1102). Thus, operations using the bypass fluid line (1178) can occur during, for example, incubation of one or more samples of interest within the flow cell cartridge assembly (1102). That is, the shared line valve (1188) can be utilized independent of the bypass valve (1190) such that the bypass valve (1190) can perform one or more operations using the bypass fluid line (1178) and / or the cache (1176) while the shared line valve (1188) and / or the central valve (1184) perform other operations simultaneously, substantially simultaneously, or offset synchronously.
[0180] The drive assembly (1112) includes a pump drive assembly (1202) and a valve drive assembly (1204). The pump drive assembly (1202) can be adapted to interface with one or more pumps (1172) to pump fluid through the flow cell (1128) and / or to load one or more samples of interest into the flow cell (1128). The valve drive assembly (1204) can be adapted to interface with one or more of the valves (1170, 1174, 1184, 1188, 1190) to control positioning of the corresponding valve (1170, 1174, 1184, 1188, 1190).
[0181] FIGS. 1-12An example of a fluidic arrangement (1220) that can be incorporated into a variation of the system (1100) is shown. The fluidic arrangement (1220) of this example includes a pump manifold assembly (1222) that can operate similarly to the pump manifold assembly (1110) described above, a sample loading manifold assembly (1228) that can operate similarly to the sample loading manifold assembly (1108) described above, a flow cell interface (1240) that can operate similarly to the flow cell interface (1126) described above, an aspirator manifold assembly (1250) that can operate similarly to the aspirator manifold assembly (1106) described above, and a waste reservoir (1270) that can operate similarly to the waste reservoir (1118) described above. The pump manifold assembly (1222) is coupled with a port assembly (1258) of the aspirator manifold assembly (1250) via a fluidic line (1224) that can be similar to the fluidic line (1178); and is coupled with the sample loading manifold assembly (1228) via a fluidic line (1226). The sample loading manifold assembly (1228) is coupled with the flow cell interface (1240) via a fluidic line (1230) that can be similar to the fluidic line (1180); and is coupled with the port assembly (1258) via fluidic lines (1232, 1234). The flow cell interface (1240) is coupled with the aspirator manifold assembly (1250) via a fluidic line (1242) that can be similar to the fluidic line (1185). The aspirator manifold assembly (1250) includes a manifold body (1252) and a common output port (1256) that provides fluid communication via the fluidic line (1185). A valve assembly (1254) controls fluid flow through the common output port (1256) and can operate similarly to the central valve (1184). The port assembly (1258) of the aspirator manifold assembly (1250) is coupled with the waste reservoir (1270) via a fluidic line (1272) that can be similar to the fluidic line (1186).
[0182] A plurality of reagent aspirators (1260) extend from the manifold body (1252) and are fluidically coupled with the valve assembly (1254) via respective fluid channels (1262) in the manifold body (1252). The reagent aspirators (1260) can operate similarly to the reagent aspirators (1200). The valve assembly (1254) is operable to selectively couple the fluid channels (1262) with the flow cell interface (1240) via the common output port (1256) and the fluid line (1230), thereby selectively providing various reagents to the flow cell interface (1240). In other words, a flow cell (e.g., similar to the flow cell (1128)) coupled with the flow cell interface (1240) can selectively receive different reagents based on control of the valve assembly (1254) when each reagent aspirator (1260) is disposed in a different respective reagent (e.g., in a respective reagent reservoir (1198)).
[0183] A plurality of reagent aspirators (1260) extend from the manifold body (1252) and are fluidically coupled with the valve assembly (1254) via respective fluid channels (1262) in the manifold body (1252). The reagent aspirators (1260) can operate similarly to the reagent aspirators (1200). The valve assembly (1254) is operable to selectively couple the fluid channels (1262) with the flow cell interface (1240) via the common output port (1256) and the fluid line (1230), thereby selectively providing various reagents to the flow cell interface (1240). In other words, a flow cell (e.g., similar to the flow cell (1128)) coupled with the flow cell interface (1240) can selectively receive different reagents based on control of the valve assembly (1254) when each reagent aspirator (1260) is disposed in a different respective reagent (e.g., in a respective reagent reservoir (1198)).
[0184] Referring back to FIGS. 13A-13B The controller (1114) of the present example includes a user interface (1206), a communication interface (1208), one or more processors (1210), and a memory (1212) storing instructions executable by the one or more processors (1210) to perform various functions including the disclosed implementations. The user interface (1206), the communication interface (1133), and the memory (1212) are electrically and / or communicatively coupled to the one or more processors (1210). The user interface (1206) can be adapted to receive input from a user and to provide information associated with the operation of the system (1100) and / or an ongoing analysis to the user. The user interface (1206) can include a touchscreen, a display, a keyboard, a speaker, a mouse, a trackball, and / or a voice recognition system.
[0185] The communication interface (1208) is adapted to enable communication between the system (1100) and a remote system (e.g., a computer) via a network (e.g., the Internet, an intranet, a local area network (LAN), a wide area network (WAN), a coaxial cable network, a wireless network, a wired network, a satellite network, a digital subscriber line (DSL) network, a cellular network, a Bluetooth connection, a near-field communication (NFC) connection, etc.). Some of the communications provided to the remote system can be associated with analysis results, imaging data, etc. generated or otherwise obtained by the system (1100). Some of the communications provided to the system (1100) can be associated with fluid analysis operations, patient records, and / or protocols to be performed by the system (1100).
[0186] The one or more processors (1210) and / or the system (1100) can include one or more of a processor-based system or a microprocessor-based system. In some implementations, the one or more processors (1210) and / or the system (1100) include one or more of a programmable processor, a programmable controller, a microprocessor, a microcontroller, a graphics processing unit (GPU), a digital signal processor (DSP), a reduced instruction set computer (RISC), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a field-programmable logic device (FPLD), a logic circuit, and / or another logic-based device that performs a variety of functions including the functions described herein.
[0187] The memory 1212 can include one or more of a semiconductor memory, a magnetic readable memory, an optical memory, a hard disk drive (HDD), an optical storage drive, a solid-state storage device, a solid state drive (SSD), a flash memory, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a random access memory (RAM), a non-volatile RAM (NVRAM) memory, a compact disk (CD), a compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a Blu-ray disk, a redundant array of independent disks (RAID) system, a cache, and / or any other storage device or storage disk in which information is stored for any duration (e.g., for a permanent period, for a temporary period, for a long period, for a buffering, for a caching).
[0188] FIG. 13A The corresponding text and examples provide a number of different methods, systems, devices, and non-transitory computer-readable media of sequences to coverage systems 106. In addition to the foregoing, one or more implementations can be described according to a flowchart FIG. 13B as shown in FIG. 10. FIGS. 13A-13BA flowchart of a series of acts 1300 for performing a sequencing run until a custom number of sequencing cycles is completed is illustrated in accordance with one or more embodiments of the present disclosure. FIGS. 13A-13B A flowchart of a series of acts 1362 for performing a sequencing run by capturing images of a custom set of flow cell regions is illustrated in accordance with one or more embodiments of the present disclosure. While FIGS. 13A-13B Acts in accordance with an embodiment are illustrated, but alternative embodiments can omit, add, reorder, and / or modify FIGS. 13A-13B any of the acts shown. FIG. 13A The acts depicted in FIG. 13 can be performed as part of a method. Alternatively, a non-transitory computer-readable storage medium can include instructions that, when executed by one or more processors, cause a computing device or system to perform the acts depicted in FIG. 13. In still further embodiments, a system includes an imaging system, a fluidic system, and a computer, the system comprising: at least one processor; and a non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the system to perform the acts depicted in FIG. 13. FIG. 13B The acts depicted in FIG. 13 can be performed as part of a method. Alternatively, a non-transitory computer-readable storage medium can include instructions that, when executed by one or more processors, cause a computing device or system to perform the acts depicted in FIG. 13. In still further embodiments, a system includes an imaging system, a fluidic system, and a computer, the system comprising: at least one processor; and a non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the system to perform the acts depicted in FIG. 13. FIG. 14 The acts depicted in FIG. 13 can be performed as part of a method. Alternatively, a non-transitory computer-readable storage medium can include instructions that, when executed by one or more processors, cause a computing device or system to perform the acts depicted in FIG. 13. In still further embodiments, a system includes an imaging system, a fluidic system, and a computer, the system comprising: at least one processor; and a non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the system to perform the acts depicted in FIG. 13.
[0189] As FIG. 14 illustrated, the series of acts 1300 includes an act 1310 of determining base calls for index sequences, an act 1320 of determining respective numbers of clusters belonging to genomic samples, an act 1330 of estimating read coverage levels, an act 1340 of generating a custom number of sequencing cycles, and an act 1350 of performing a sequencing run. For example, the series of acts 1300 can include acts for performing any of the operations described in the following clauses:
[0190] Clause 1. A method comprising:
[0191] determining, from a subset of sequencing cycles of a sequencing run of genomic samples, base calls for index sequences within oligonucleotide clusters;
[0192] based on the index sequences, determining respective numbers of oligonucleotide clusters belonging to respective ones of the genomic samples;
[0193] based on the respective numbers of oligonucleotide clusters belonging to respective genomic samples and a currently selected number of sequencing cycles of the sequencing run, estimating read coverage levels for the genomic samples;
[0194] for the sequencing run and based on the estimated read coverage levels, generating a custom number of sequencing cycles sufficient to generate nucleotide reads that satisfy a target read coverage level for each of the genomic samples; and
[0195] performing the sequencing run until the custom number of sequencing cycles is completed.
[0196] Clause 2. The method of clause 1, further comprising estimating the read coverage level of the genomic samples by:
[0197] determining a filter metric for a subset of oligonucleotide clusters indicative of a filter threshold that is satisfied by signals of the oligonucleotide clusters; and
[0198] estimating the read coverage level of the genomic samples based on the filter metric and the respective number of oligonucleotide clusters belonging to the respective genomic sample.
[0199] Clause 3. The method of clause 2, further comprising determining the filter metric by determining, in a filter plot, a percentage of clusters belonging to each genomic sample that satisfy a purity filter for signals emitted from the oligonucleotide clusters.
[0200] Clause 4. The method of clause 2, further comprising estimating the read coverage level of the genomic samples by:
[0201] determining, for each of the genomic samples, a number of oligonucleotide clusters that pass a filter that satisfies the filter threshold based on the filter metric and the respective number of oligonucleotide clusters belonging to the respective genomic sample; and
[0202] estimating a minimum number of nucleotide reads that cover a genomic region of each genomic sample based on the number of oligonucleotide clusters that pass the filter.
[0203] Clause 5. The method of clause 1, further comprising:
[0204] determining, based on the estimated read coverage level, a customized set of pool regions of the flow cell to image from the flow cell sufficient to generate the nucleotide reads that satisfy the target read coverage level for each of the genomic samples; and
[0205] performing the sequencing run by capturing images of the customized set of pool regions of the flow cell for the customized number of sequencing cycles using the imaging system.
[0206] Clause 6. The method of clause 1, further comprising performing the subset of sequencing cycles according to an order of index cycles prior to genomic sequencing cycles by:
[0207] determining a base call of a first index sequence appended to a sample genomic sequence of a genomic sample;
[0208] determining base calls of a second index sequence appended to the sample genomic sequence of the genomic sample; and
[0209] After determining the base calls of the first index sequence and the second index sequence, determining base calls of first nucleotide reads corresponding to a first portion of the sample genomic sequence, and determining base calls of second nucleotide reads corresponding to a second portion of the sample genomic sequence.
[0210] Clause 7. The method of clause 1, further comprising determining the respective number of oligonucleotide clusters belonging to the respective genomic sample by:
[0211] identifying, from the index sequences, assigned index sequences that match index sequences registered for the sequencing run and unassigned index sequences that do not match the index sequences registered for the sequencing run;
[0212] removing, from data of the sequencing run, a subset of oligonucleotide clusters corresponding to the unassigned index sequences;
[0213] determining a respective subset of assigned index sequences corresponding to the respective genomic sample; and
[0214] determining, from the respective subset of assigned index sequences, a number of oligonucleotide clusters belonging to each genomic sample.
[0215] Clause 8. The method of clause 1, further comprising generating the custom number of sequencing cycles for the sequencing run by increasing or decreasing a preset number of sequencing cycles of the sequencing run.
[0216] Clause 9. The method of clause 1, further comprising generating the custom number of sequencing cycles for the sequencing run by:
[0217] identifying a minimum number of sequencing cycles and a maximum number of sequencing cycles of the sequencing run; and
[0218] increasing or decreasing a preset number of sequencing cycles of the sequencing run to the custom number of sequencing cycles within the minimum number of sequencing cycles and the maximum number of sequencing cycles.
[0219] Clause 10. The method of clause 1, further comprising estimating the level of read coverage by:
[0220] determining, from the sequencing run, a number of unique nucleotide reads aligned to a reference genome;
[0221] The number of filtered nucleotide reads from the oligonucleotide clusters that have signals that meet the filtering threshold is determined from the sequencing run.
[0222] The bioinformatics efficiency metric is determined by dividing the number of unique nucleotide reads by the number of nucleotide reads that pass through the filter; and
[0223] The read coverage level of the genome sample is estimated based on the bioinformatics efficiency metric and the corresponding number of oligonucleotide clusters belonging to the corresponding genome sample.
[0224] Clause 11. The method according to Clause 1, the method further comprising detecting the reagent volume of the kit in fluid communication with the fluid system, and operating the fluid system to perform one or more additional sequencing cycles relative to the currently selected number of sequencing cycles until the customized number of sequencing cycles is completed by aspirating one or more reagents from the kit.
[0225] Clause 12. The method according to Clause 1, the method further comprising terminating the operation of the fluid system from one or more sequencing cycles of the currently selected number of sequencing cycles to complete the sequencing run after the customized number of sequencing cycles has been executed.
[0226] like FIG. 14 As shown, the series of actions 1362 includes actions 1360 to determine the base detection of the index sequence, actions 1370 to determine the corresponding number of clusters belonging to the genome sample, actions 1380 to estimate the read coverage level, actions 1390 to determine the customized set of flow pool regions to be imaged, and actions 1392 to perform the sequencing run. For example, the series of actions 1362 may include actions for performing any of the operations described in the following clauses:
[0227] Clause 13. A method comprising:
[0228] Base detection of index sequences within oligonucleotide clusters was determined from a subset of sequencing cycles in a sequencing run of a genome sample.
[0229] Based on the index sequence, determine the corresponding number of oligonucleotide clusters belonging to the corresponding genome sample in the genome sample;
[0230] Based on the corresponding number of oligonucleotide clusters belonging to the corresponding genome sample and the currently selected number of sequencing cycles in the sequencing run, the read coverage level of the genome sample is estimated;
[0231] determining, from a flow cell and based on estimated read coverage levels, a customized set of flow cell regions to image sufficient to generate a target number of nucleotide reads that meet a target read coverage level for each of the genomic samples in the sequencing run; and
[0232] performing the sequencing run by capturing images of the customized set of flow cell regions during sequencing cycles of the sequencing run.
[0233] Clause 14. The method of clause 13, further comprising determining the customized set of flow cell regions by determining a customized number of flow cell regions to image sufficient to generate the target number of nucleotide reads that meet the target read coverage level for each genomic sample.
[0234] Clause 15. The method of clause 13, further comprising determining the customized set of flow cell regions by determining, from a flow cell, a set of tiles to image sufficient to generate the target number of nucleotide reads that meet the target read coverage level for each genomic sample.
[0235] Clause 16. The method of clause 13, further comprising capturing the images of the customized set of flow cell regions without adjusting a currently selected number of sequencing cycles of the sequencing run.
[0236] Clause 17. The method of clause 13, further comprising determining the customized set of flow cell regions by increasing or decreasing a number of flow cell regions from an initial set of flow cell regions selected for the sequencing run.
[0237] Clause 18. The method of clause 13, further comprising estimating the read coverage levels by:
[0238] determining a filter metric for a subset of oligonucleotide clusters that indicate a filtering threshold for signals that meet the oligonucleotide clusters; and
[0239] estimating the read coverage levels for the genomic samples based on the filter metric and the respective number of oligonucleotide clusters belonging to a respective genomic sample.
[0240] Clause 19. The method of clause 18, further comprising determining the filter metric by determining, in a filter graph, a percentage of clusters belonging to each genomic sample that meet a purity filter for signals emitted from the oligonucleotide clusters.
[0241] Clause 20. The method of clause 18, further comprising estimating the read coverage levels for the genomic samples by:
[0242] determining, based on the filter metric and the respective number of oligonucleotide clusters belonging to a respective genomic sample, a number of oligonucleotide clusters that pass the filter for each of the genomic samples that satisfy the filtering threshold; and
[0243] estimating, based on the number of oligonucleotide clusters that pass the filter, a minimum number of nucleotide reads that cover a genomic region for each genomic sample.
[0244] The methods described herein can be used in conjunction with a variety of nucleic acid sequencing technologies. Particularly suitable technologies are those in which nucleic acids are attached to fixed positions in an array such that their relative positions do not change and in which the array is repeatedly imaged. Embodiments in which images are obtained in different color channels (e.g., in accord with different labels used to distinguish one type of nucleobase from another) are particularly suitable. In some embodiments, the process of determining the nucleotide sequence of a target nucleic acid (i.e., a nucleic acid polymer) can be an automated process. Preferred embodiments include sequencing by synthesis (SBS) technologies.
[0245] SBS technologies generally include the enzymatic extension of a nascent nucleic acid strand by repeated addition of nucleotides against a template strand. In traditional SBS methods, a single nucleotide monomer can be provided to a target nucleotide in the presence of a polymerase in each delivery. However, in the methods described herein, more than one type of nucleotide monomer can be provided to a target nucleic acid in the presence of a polymerase in a delivery.
[0246] SBS can utilize nucleotide monomers with a terminator moiety or nucleotide monomers that lack any terminator moiety. Methods that utilize nucleotide monomers that lack a terminator include, for example, pyrophosphate sequencing and sequencing using gamma-phosphate labeled nucleotides, as described in further detail below. In methods that use nucleotide monomers that lack a terminator, the number of nucleotides added in each cycle is generally variable and depends on the template sequence and the manner of nucleotide delivery. For SBS technologies that utilize nucleotide monomers with a terminator moiety, the terminator can be effectively irreversible under the sequencing conditions used, as is the case for traditional Sanger sequencing with dideoxynucleotides, or the terminator can be reversible, as is the case for sequencing methods developed by Solexa (now Illumina, Inc.).
[0247] SBS techniques can utilize nucleotide monomers with a label moiety or nucleotide monomers lacking a label moiety. Thus, incorporation events can be detected based on the properties of the label, such as fluorescence of the label; properties of the nucleotide monomer, such as molecular weight or charge; release of a byproduct of the incorporation of the nucleotide, such as pyrophosphate; and the like. In embodiments where two or more different nucleotides are present in the sequencing reagent, the different nucleotides can be distinguishable from one another, or alternatively, two or more different labels can be indistinguishable under the detection technique used. For example, different nucleotides present in the sequencing reagent can have different labels, and they can be distinguished using appropriate optics, as exemplified by the sequencing method developed by Solexa (now Illumina, Inc.).
[0248] A preferred embodiment includes pyrosequencing techniques. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) as specific nucleotides are incorporated into a nascent strand (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M., and Nyren, P. (1996) “Real-time DNA sequencing using detection of pyrophosphate release.” Analytical Biochemistry 242(1), 84-9; Ronaghi, M. (2001) “Pyrosequencing sheds light on DNA sequencing.” Genome Res. 11(1), 3-11; Ronaghi, M., Uhlen, M., and Nyren, P. (1998) “A sequencing method based on real-time pyrophosphate.” Science 281(5375), 363; U.S. Patent No. 6,210,891; U.S. Patent No. 6,258,568; and U.S. Patent No. 6,274,320, the disclosures of which are incorporated herein by reference in their entireties). In pyrosequencing, released PPi can be detected by being immediately converted to ATP by an adenosine triphosphate (ATP) sulfurylase enzyme, and the level of ATP produced is detected by the photons produced by a luciferase. The nucleic acid to be sequenced can be attached to features in an array, and the array can be imaged to capture the chemiluminescent signal produced as nucleotides are incorporated at the features of the array. An image can be obtained after the array is treated with a particular nucleotide type (e.g., A, T, C, or G). The images obtained after each nucleotide type is added will differ in terms of which features of the array are detected. These differences in the images reflect the different sequence content of the features on the array. However, the relative positions of each feature will remain the same in the images. The images can be stored, processed, and analyzed using the methods described herein. For example, the images obtained after the array is treated with each different nucleotide type can be processed in the same manner as exemplified herein for images obtained from different detection channels for a sequencing-by- synthesis method.
[0249] In another exemplary type of SBS, cyclic sequencing is accomplished by stepwise addition of reversible terminator nucleotides that comprise, for example, cleavable or photobleavable dye labels, as described, for example, in WO 04 / 018497 and U.S. Patent No. 7,057,026, the disclosures of which are incorporated herein by reference. This method is commercialized by Solexa (now Illumina Inc.) and is also described in WO 91 / 06678 and WO 07 / 123,744, the disclosure of each of which is incorporated herein by reference. The availability of fluorescently labeled terminators, where the termination can be reversible and the fluorescent label can be cleaved, facilitates efficient cyclic reversible termination (CRT) sequencing. Polymerases can also be co-engineered to efficiently incorporate and extend from these modified nucleotides.
[0250] Preferably, in reversible terminator-based sequencing embodiments, the labels do not substantially inhibit extension under SBS reaction conditions. However, the detection labels can be removable, for example, by cleavage or degradation. Images can be captured after the labels are incorporated into the arrayed nucleic acid features. In particular embodiments, each cycle involves simultaneous delivery of four different nucleotide types to the array, and each nucleotide type has a label that is spectrally distinct. Four images can then be obtained, each using a detection channel selective for one of the four different labels. Alternatively, different nucleotide types can be added sequentially, and an image of the array can be obtained between each addition step. In such embodiments, each image will show the nucleic acid features that have incorporated a particular type of nucleotide. Different features are present or absent in different images due to the different sequence content of each feature. However, the relative positions of the features will remain unchanged in the images. Images obtained by such reversible terminator-SBS methods can be stored, processed, and analyzed as described herein. After the image capture step, the labels can be removed and the reversible terminator moieties can be removed for subsequent cycles of nucleotide addition and detection. Removal of the labels after they have been detected in a particular cycle and before the subsequent cycle can provide the advantage of reducing background signal and cross-talk between cycles. Examples of labels and removal methods that can be used are set forth below.
[0251] In particular embodiments, some or all of the nucleotide monomers can include reversible terminators. In such embodiments, the reversible terminator / cleavable fluorophore can include a fluorophore attached to the ribose moiety via a 3' ester linkage (Metzker, Genome Res. 15: 1767-1776 (2005), which is incorporated herein by reference). Other methods have decoupled terminator chemistry from cleavage of fluorescent labels (Ruparel et al., Proc Natl Acad Sci USA 102: 5932-7 (2005), which is incorporated by reference in its entirety). Ruparel et al. describe the development of reversible terminators that use a small 3' allyl group to block extension, but which can be easily deblocked by a short treatment with a palladium catalyst. The fluorophore is attached to the base via a photocleavable linker that can be easily cleaved by exposure to long wavelength UV light for 30 seconds. Thus, disulfide reduction or photocleavage can be used as a cleavable linker. Another approach to reversible termination is to use natural termination that occurs after placing a bulky dye on the dNTP. The presence of a charged bulky dye on the dNTP can act as a highly efficient terminator by steric and / or electrostatic blockage. The presence of one incorporation event prevents further incorporation unless the dye is removed. Cleavage of the dye removes the fluorophore and effectively reverses the termination. Examples of modified nucleotides are also described in U.S. Patent No. 7,427,673 and U.S. Patent No. 7,057,026, the disclosures of which are incorporated by reference herein in their entireties.
[0252] Additional exemplary SBS systems and methods that can be utilized with the methods and systems described herein are described in U.S. Patent Application Publication No. 2007 / 0166705, U.S. Patent Application Publication No. 2006 / 0188901, U.S. Patent No. 7,057,026, U.S. Patent Application Publication No. 2006 / 0240439, U.S. Patent Application Publication No. 2006 / 0281109, PCT Publication No. WO 05 / 065814, U.S. Patent Application Publication No. 2005 / 0100900, PCT Publication No. WO 06 / 064199, PCT Publication No. WO 07 / 010,251, U.S. Patent Application Publication No. 2012 / 0270305, and U.S. Patent Application Publication No. 2013 / 0260372, the disclosures of which are incorporated by reference herein in their entireties.
[0253] Some embodiments can use fewer than four different labels to use detection of four different nucleotides. For example, SBS can be performed using the methods and systems described in the incorporated material of U.S. Patent Application Publication No. 2013 / 0079232. As a first example, a pair of nucleotide types can be detected at the same wavelength, but distinguished based on a difference in intensity of one member of the pair relative to the other, or based on a change (e.g., by chemical modification, photochemical modification, or physical modification) of one member of the pair that results in a clear signal appearance or disappearance compared to the signal of the other member of the pair that is detected. As a second example, three of the four different nucleotide types are capable of being detected under certain conditions, while the fourth nucleotide type lacks a label that can be detected under those conditions or that is minimally detected under those conditions (e.g., minimal detection due to background fluorescence, etc.). The first three nucleotide types can be determined to be incorporated into the nucleic acid based on the presence of their respective signals, and the fourth nucleotide type can be determined to be incorporated into the nucleic acid based on the absence of any signal or minimal detection of any signal. As a third example, one nucleotide type can include labels that are detected in two different channels, while the other nucleotide types are detected in no more than one channel. The three exemplary configurations described above are not considered to be mutually exclusive, and can be used in various combinations. An exemplary embodiment that combines all three examples is a fluorescence-based SBS method that uses a first nucleotide type that is detected in a first channel (e.g., dATP with a label that is detected in the first channel when excited by a first excitation wavelength), a second nucleotide type that is detected in a second channel (e.g., dCTP with a label that is detected in the second channel when excited by a second excitation wavelength), a third nucleotide type that is detected in both the first channel and the second channel (e.g., dTTP with at least one label that is detected in both channels when excited by the first excitation wavelength and / or the second excitation wavelength), and a fourth nucleotide type that lacks a label that is detected in either channel (e.g., dGTP without a label).
[0254] Additionally, as described in the incorporated material of U.S. Patent Application Publication No. 2013 / 0079232, sequencing data can be obtained using a single channel. In such a so-called single-dye sequencing method, a first nucleotide type is labeled, but the label is removed after the first image is generated, and only a second nucleotide type is labeled after the first image is generated. A third nucleotide type retains its label in both the first image and the second image, and a fourth nucleotide type remains unlabeled in both images.
[0255] Some embodiments can utilize edge-sequencing-by-ligation techniques. Such techniques utilize DNA ligase to incorporate oligonucleotides and recognize the incorporation of such oligonucleotides. The oligonucleotides typically have different labels related to the identity of specific nucleotides in the sequence to which the oligonucleotides hybridize. As with other SBS methods, images can be obtained after treating an array of nucleic acid features with labeled sequencing reagents. Each image will show nucleic acid features that have incorporated a particular type of label. Different features will be present or absent in different images due to the different sequence content of each feature, but the relative positioning of the features will remain unchanged in the images. Images obtained by connection-based sequencing methods can be stored, processed, and analyzed as described herein. Exemplary SBS systems and methods that can be used with the methods and systems described herein are described in U.S. Patent No. 6,969,488, U.S. Patent No. 6,172,218, and U.S. Patent No. 6,306,597, the disclosures of which are incorporated by reference herein in their entireties.
[0256] Some embodiments can utilize nanopore sequencing (Deamer, D. W. and Akeson, M., "Nanopores and nucleic acids: prospects for ultrarapid sequencing." Trends Biotechnol. 18, 147-151 (2000); Deamer, D. and D. Branton, "Characterization of nucleic acids by nanopore analysis." Acc. Chem. Res. 35:817-825 (2002); Li, J., M. Gershow, D. Stein, E. Brandin, and J. A. Golovchenko, "DNA molecules and configurations in a solid-state nanopore microscope", Nat. Mater., 2:611-615 (2003), the disclosures of which are incorporated herein by reference in their entireties). In such embodiments, the target nucleic acid is passed through a nanopore. The nanopore can be a synthetic pore or a biological membrane protein, such as a-hemolysin. As the target nucleic acid is passed through the nanopore, each base pair can be identified by measuring fluctuations in the electrical conductivity of the pore. (U.S. Patent No. 7,001,792; Soni, G. V. and Meller, "A. Progress toward ultrafast DNA sequencing using solid-state nanopores." Clin. Chem. 53, 1996-2001 (2007); Healy, K., "Nanopore-based single-molecule DNA analysis." Nanomed. 2, 459-481 (2007); Cockroft, S. L., Chu, J., Amorin, M., and Ghadiri, M. R., "A single-molecule nanopore device detects DNA polymerase activity with single-nucleotide resolution." J. Am. Chem. Soc. 130, 818-820 (2008), the disclosures of which are incorporated herein by reference in their entireties). Data obtained from nanopore sequencing can be stored, processed, and analyzed as described herein.In particular, data can be processed as images according to the exemplary processing of optical images and other images described herein.
[0257] Some embodiments can utilize methods involving real-time monitoring of DNA polymerase activity. Nucleotide incorporation can be detected by fluorescence resonance energy transfer (FRET) interactions between a fluorophore-bearing polymerase and a gamma-phosphate labeled nucleotide, as described in, for example, U.S. Patent No. 7,329,492 and U.S. Patent No. 7,211,414 (each of which is incorporated herein by reference), or nucleotide incorporation can be detected with zero-mode waveguides, as described in, for example, U.S. Patent No. 7,315,019 (which is incorporated herein by reference), and nucleotide incorporation can be detected using fluorescent nucleotide analogs and engineered polymerases, as described in, for example, U.S. Patent No. 7,405,281 and U.S. Patent Application Publication No. 2008 / 0108082 (each of which is incorporated herein by reference). Illumination can be limited to a volume on the order of femtoliters around surface-tethered polymerases, such that incorporation of fluorescently labeled nucleotides can be observed with low background (Levene, M. J. et al., "Zero-mode waveguides for single-molecule analysis at high concentrations." Science 299, 682-686 (2003); Lundquist, P. M. et al., "Parallel confocal detection of single molecules in real time." Opt. Lett. 33, 1026-1028 (2008); Korlach, J. et al., "Selective aluminum passivation for targeted immobilization of single DNA polymerase molecules in zero-mode waveguide nanostructures." Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008), the disclosures of which are incorporated herein in their entireties by reference). Images obtained by such methods can be stored, processed, and analyzed as described herein.
[0258] Some SBS embodiments include detecting protons released upon nucleotide incorporation extension products. For example, sequencing based on detection of released protons can use electrical detectors and related technology commercially available from Ion Torrent, Inc. (Guilford, CT, which is a subsidiary of Life Technologies) or sequencing methods and systems described in US 2009 / 0026082 Al, US 2009 / 0127589 Al, US 2010 / 0137143 Al, or US 2010 / 0282617 Al, each of which is incorporated herein by reference. The methods set forth herein using kinetic exclusion to amplify target nucleic acids can be readily applied to substrates for detecting protons. More specifically, the methods set forth herein can be used to generate clonal populations of amplicons for detecting protons.
[0259] The SBS methods described above can advantageously be performed in a variety of formats such that multiple different target nucleic acids are manipulated simultaneously. In particular embodiments, different target nucleic acids can be processed in a common reaction vessel or on the surface of a particular substrate. This allows for convenient delivery of sequencing reagents, removal of unreacted reagents, and detection of incorporation events in a variety of ways. In embodiments using surface-bound target nucleic acids, the target nucleic acids can be in an array format. In an array format, the target nucleic acids can typically be bound to a surface in a spatially distinguishable manner. The target nucleic acids can be bound by direct covalent attachment, attachment to a bead or other particle, or binding to a polymerase or other molecule attached to the surface. An array can include a single copy of a target nucleic acid at each site, also referred to as a feature, or multiple copies of the same sequence can be present at each site or feature. Multiple copies can be generated by amplification methods such as bridge amplification or emulsion PCR, as described in further detail below.
[0260] The methods described herein can use arrays having features at any of a variety of densities, including, for example, at least about 10 features / cm2, 100 features / cm2, 500 features / cm2, 1,000 features / cm2, 5,000 features / cm2, 10,000 features / cm2, 50,000 features / cm2, 100,000 features / cm2, 1,000,000 features / cm2, 5,000,000 features / cm2, or higher.
[0261] The advantages of the methods described herein are that they provide rapid and efficient detection of multiple target nucleic acids in parallel. Therefore, this disclosure provides an integrated system capable of preparing and detecting nucleic acids using techniques known in the art, such as those exemplified above. Thus, the integrated system of this disclosure may include fluid components capable of delivering amplification reagents and / or sequencing reagents to one or more immobilized DNA fragments, the system including components such as pumps, valves, reservoirs, fluid lines, etc. A flow cell in the integrated system may be configured for and / or for detecting target nucleic acids. Exemplary flow cells are described, for example, in US 2010 / 0111768A1 and US Serial No. 13 / 273,666, each of which is incorporated herein by reference. As illustrated with respect to a flow cell, one or more fluid components of the integrated system may be used for amplification and detection methods. Taking a nucleic acid sequencing implementation as an example, one or more fluid components of the integrated system may be used for the amplification methods described herein and for delivering sequencing reagents in sequencing methods, such as those exemplified above. Alternatively, the integrated system may include separate fluid systems to perform amplification and detection methods. Examples of integrated sequencing systems capable of generating amplified nucleic acids and determining their sequences include, but are not limited to, MiSeq. ™ The platform (Illumina, Inc., San Diego, CA) and the device described in U.S. Serial No. 13 / 273,666 are incorporated herein by reference.
[0262] The sequencing system described above sequences nucleic acid polymers present in samples received by the sequencing equipment. As defined herein, “sample” and its derivatives are used in their broadest sense and include any specimen, culture, etc., suspected of containing a target. In some embodiments, a sample includes nucleic acids in the form of DNA, RNA, PNA, LNA, chimeric, or hybrid forms. A sample may include any biological, clinical, surgical, agricultural, atmospheric, or aquatic plant or animal specimen containing one or more nucleic acids. The term also includes any isolated nucleic acid sample, such as genomic DNA, freshly frozen, or formalin-fixed paraffin-embedded nucleic acid specimens. It is also envisioned that a sample may originate from: a single individual, a collection of nucleic acid samples from genetically related members, nucleic acid samples from genetically unrelated members, a (matched) nucleic acid sample from a single individual (such as tumor samples and normal tissue samples), or a sample from a single source containing two different forms of genetic material (such as maternal DNA and fetal DNA obtained from a maternal subject), or a sample containing contaminating bacterial DNA in the presence of plant or animal DNA. In some embodiments, the source of nucleic acid material may include nucleic acids obtained from newborns, such as those commonly used for newborn screening.
[0263] A nucleic acid sample can include high molecular weight material, such as genomic DNA (gDNA). A sample can include low molecular weight material, such as nucleic acid molecules obtained from FFPE samples or archived DNA samples. In another embodiment, low molecular weight material includes enzymatically fragmented or mechanically fragmented DNA. A sample can include cell-free circulating DNA. In some embodiments, a sample can include nucleic acid molecules obtained from biopsy tissue, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture microdissection, surgical resections, and other clinically or laboratory obtained samples. In some embodiments, a sample can be an epidemiological sample, an agricultural sample, a forensic sample, or a pathogenic sample. In some embodiments, a sample can include nucleic acid molecules obtained from an animal, such as a human or a mammalian source. In another embodiment, a sample can include nucleic acid molecules obtained from a non-mammalian source, such as a plant, a bacterium, a virus, or a fungus. In some embodiments, the source of the nucleic acid molecules can be an archived or extinct sample or species.
[0264] Additionally, the methods and compositions disclosed herein can be used to amplify nucleic acid samples having low quality nucleic acid molecules, such as degraded and / or fragmented genomic DNA from a forensic sample. In one embodiment, the forensic sample can include nucleic acid obtained from a crime scene, nucleic acid obtained from a missing persons DNA database, nucleic acid obtained from a laboratory associated with a forensic investigation, or include a forensic sample obtained by law enforcement, one or more military services, or any such personnel. The nucleic acid sample can be a purified sample or a lysate containing crude DNA, for example, derived from a buccal swab, paper, fabric, or other substrate that can be impregnated with saliva, blood, or other bodily fluid. Thus, in some embodiments, the nucleic acid sample can include small amounts of DNA, such as genomic DNA, or fragmented portions of DNA. In some embodiments, the target sequence can be present in one or more bodily fluids, including but not limited to blood, sputum, plasma, semen, urine, and serum. In some embodiments, the target sequence can be obtained from a victim's hair, skin, tissue sample, autopsy, or remains. In some embodiments, the nucleic acid comprising one or more target sequences can be obtained from a deceased animal or human. In some embodiments, the target sequence can include nucleic acid obtained from non-human DNA, such as microbial, plant, or insect DNA. In some embodiments, the target sequence or amplified target sequence is directed toward the purpose of human identity identification. In some embodiments, the disclosure generally relates to methods for identifying characteristics of a forensic sample. In some embodiments, the disclosure generally relates to methods of human identity identification using one or more target-specific primers disclosed herein or one or more target-specific primers designed using the primer design criteria outlined herein. In one embodiment, a forensic sample or human identity identification sample containing at least one target sequence can be amplified using any one or more target-specific primers disclosed herein or using the primer criteria outlined herein.
[0265] The components of the sequence-to-coverage system 106 can include software, hardware, or both. For example, the components of the sequence-to-coverage system 106 can include one or more instructions stored on a computer-readable storage medium and executable by a processor of one or more computing devices (e.g., the local server device 102). When executed by the one or more processors, the computer-executable instructions of the sequence-to-coverage system 106 can cause the computing device to perform the bubble detection methods described herein. Alternatively, the components of the sequence-to-coverage system 106 can include hardware, such as a specialized processing device for performing a certain function or set of functions. Additionally or alternatively, the components of the sequence-to-coverage system 106 can include a combination of computer-executable instructions and hardware.
[0266] Additionally, the components of the sequence-to-coverage system 106 that perform the functions described herein with respect to the sequence-to-coverage system 106 can be implemented, for example, as part of a standalone application, a module of an application, a plug-in of an application, one or more library functions callable by other applications, and / or a cloud computing model. Thus, the components of the sequence-to-coverage system 106 can be implemented as part of a standalone application on a personal computing device or mobile device. Additionally or alternatively, the components of the sequence-to-coverage system 106 can be implemented in any application that provides sequencing services, including but not limited to Illumina, BaseSpace, Illumina MiSeq, Illumina NovaSeq, Illumina NextSeq, Illumina Trusq, or Illumina TruSight software. “Illumina,” “BaseSpace,” “MiSeq,” “NovaSeq,” “NextSeq,” “TruSeq,” and “TruSight” are registered trademarks or trademarks of Illumina, Inc. in the United States and / or other countries.
[0267] As discussed in greater detail below, embodiments of the present disclosure can include or utilize a special-purpose or general-purpose computer that includes computer hardware, such as, for example, one or more processors and system memory. Embodiments within the scope of the present disclosure also include physical and other computer-readable media for carrying or storing computer- executable instructions and / or data structures. In particular, one or more of the processes described herein can be implemented at least in part as instructions embodied in a non-transitory computer- readable medium and executable by one or more computing devices (e.g., any of the media content access devices described herein). Generally, a processor (e.g., a microprocessor) receives instructions, from a non-transitory computer-readable medium (e.g., a memory, etc.), and executes those instructions, thereby performing one or more processes, including one or more of the processes described herein.
[0268] A computer-readable medium can be any available medium or means that can be accessed by a general purpose or special purpose computing system. A computer-readable medium that stores computer-executable instructions is a non-transitory computer-readable storage medium (device). A computer-readable medium that carries computer-executable instructions is a transmission medium. Accordingly, embodiments of the present disclosure can include at least two distinct kinds of computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.
[0269] Non-transitory computer-readable storage media (devices) include RAM, ROM, EEPROM, CD-ROM, solid state drives (SSDs) (e.g., based on RAM), Flash memory, phase- change memory (PCM), other types of memory, other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other non-transitory storage medium which can be used to store desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer.
[0270] A "network" is defined as one or more data links that enable the transport of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided over a network or another communications connection (either hardwired, wireless, or a combination of hardwired or wireless) to a computer, the computer properly views the connection as a transmission medium. Transmission media can include a network and / or data links which can be used to carry desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer. Combinations of the above should also be included within the scope of computer-readable media.
[0271] Further, upon reaching various computer system components, program code means in the form of computer-executable instructions or data structures can be transferred automatically from transmission media to non-transitory computer-readable storage media (devices) (or vice versa). For example, computer-executable instructions or data structures received by way of network or data links can be buffered in RAM within a network interface module (e.g., a NIC), and then eventually transferred to computer system RAM and / or to less volatile computer storage media (devices) at a computer system. Accordingly, it should be understood that non-transitory computer-readable storage media (devices) can be included in computer system components that also (or even primarily) utilize transmission media.
[0272] Computer-executable instructions include, for example, instructions and data which, when executed at a processor, cause a general purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions. In some embodiments, computer-executable instructions are executed on a general purpose computer to transform the general purpose computer into a special purpose computer that implements elements of the present disclosure. Computer-executable instructions can be, for example, binaries, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.
[0273] Those skilled in the art will appreciate that the disclosure can be practiced in network computing environments with many types of computer system configurations, including, personal computers, desktop computers, laptop computers, message processors, hand-held devices, multi-processor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile telephones, PDAs, tablets, pagers, routers, switches, and the like. The disclosure can also be practiced in distributed system environments where local and remote computer systems, which are linked (either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) through a network, both perform tasks. In a distributed system environment, program modules can be located in both local and remote memory storage devices.
[0274] Embodiments of the disclosure can also be implemented in a cloud computing environment. In this description and the following claims, "cloud computing" is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, a consumer can request and be granted near-instantaneous access to or sharing of
[0275] A cloud computing model can be composed of various characteristics such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, and the like. A cloud computing model can also exhibit various service models such as, for example, Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (IaaS). A cloud computing model can also be deployed using different deployment models such as private cloud, community cloud, public cloud, hybrid cloud, and the like. In this description and the following claims, a "cloud computing environment" is an environment in which cloud computing is employed.
[0276] FIG. 14 A block diagram of a computing device 1400 that can be configured to perform one or more of the processes described above is illustrated. One will appreciate that one or more computing devices such as computing device 1400 can implement sequence-to-cover system 106. As As shown, computing device 1400 can include a processor 1402, memory 1404, storage 1406, an I / O interface 1408, and a communication interface 1410 that can be communicatively coupled via a communication infrastructure 1412. In certain embodiments, computing device 1400 can include fewer or more components than those shown. The components of computing device 1400 shown. The components of computing device 1400 shown.
[0277] In one or more embodiments, the processor 1402 includes hardware for executing instructions, such as those that make up a computer program, As an example and not by way of limitation, to execute instructions for dynamically modifying a workflow, the processor 1402 can retrieve (or fetch) the instructions from an internal register, an internal cache, memory 1404, or storage 1406 and decode and execute them. The memory 1404 can be a volatile or non-volatile memory for storing data, metadata, and program instructions to be executed by the processor. The storage 1406 includes storage such as a hard disk, flash drive, or other digital storage device for storing data or instructions for executing the methods described herein.
[0278] The I / O interface 1408 allows a user to provide input to the computing device 1400, receive output from the computing device, and otherwise transfer data to and from the computing device. The I / O interface 1408 can include a mouse, a keypad, or a keyboard, a touch screen, a camera, an optical scanner, a network interface, a modem, other well-known I / O devices, or combinations of such I / O devices. The I / O interface 1408 can include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (for example, a display screen), one or more output drivers (for example, display drivers), one or more audio speakers, and one or more audio drivers. In certain embodiments, the I / O interface 1408 is configured to provide graphical data to a display for presentation to a user. The graphical data can represent one or more graphical user interfaces and / or any other graphical content serving a particular implementation.
[0279] The communication interface 1410 can include hardware, software, or both. In any case, the communication interface 1410 can provide one or more interfaces for communication (such as, for example, packet-based communication) between the computing device 1400 and one or more other computing devices or networks. As an example and not by way of limitation, the communication interface 1410 can include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a WI-FI network.
[0280] Additionally, the communication interface 1410 can facilitate communication with various types of wired or wireless networks. The communication interface 1410 can also facilitate communication using various communication protocols. The communication infrastructure 1412 can also include hardware, software, or both providing communication between elements of computing device 1400. For example, the communication interface 1410 can enable multiple computing devices connected via a particular infrastructure to communicate with one another to perform one or more aspects of the processes described herein, using various communication protocols. To illustrate, a sequencing process can allow multiple devices (e.g., client devices, sequencing devices, and server devices) to exchange information such as sequencing data and error notifications.
[0281] In the foregoing specification, the disclosure has been described with reference to specific exemplary embodiments thereof. Various embodiments and aspects of the disclosure are described with reference to details discussed, and the accompanying drawings illustrate various embodiments. The description above and the drawings are illustrative of the disclosure and are not to be construed as limiting the disclosure. Numerous specific details are described to provide a thorough understanding of various embodiments of the disclosure.
[0282] The disclosure can be embodied in other specific forms without departing from the spirit or essential characteristics thereof. The described implementations are to be considered in all respects only as illustrative and not restrictive. For example, the methods described herein can be performed with fewer or additional steps / actions, or in different orders. Additionally, steps / actions described herein can be repeated or performed in parallel with one another or with different instances of the same or similar steps / actions. Accordingly, the scope of the application is indicated by the appended claims, rather than by the foregoing description. All changes that come within the meaning and range of equivalency of the claims are to be embraced within their scope.
Claims
1. A system comprising: an imaging system; a fluidic system; and a computing engine comprising: at least one processor; and a non-transitory computer-readable medium comprising instructions that, when executed by the at least one processor, cause the system to: determine, from a subset of sequencing cycles of a sequencing run of genomic samples, base calls of index sequences within oligonucleotide clusters; based on the index sequences, determine respective quantities of oligonucleotide clusters belonging to respective ones of the genomic samples; based on the respective quantities of oligonucleotide clusters belonging to respective ones of the genomic samples and a currently selected number of sequencing cycles of the sequencing run, estimate a level of read coverage for the genomic samples; for the sequencing run and based on the estimated level of read coverage, generate a customized number of sequencing cycles sufficient to generate nucleotide reads that satisfy a target level of read coverage for each of the genomic samples; and execute the sequencing run until the customized number of sequencing cycles is completed.
2. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to estimate the level of read coverage by: determining a filter metric for a subset of oligonucleotide clusters that indicates a filtering threshold for signals that satisfy the oligonucleotide clusters; and based on the filter metric and the respective quantities of oligonucleotide clusters belonging to respective ones of the genomic samples, estimate the level of read coverage for the genomic samples.
3. The system of claim 2, further comprising instructions that, when executed by the at least one processor, cause the system to determine the filter metric by determining, in a filter graph, a percentage of clusters belonging to each of the genomic samples that satisfy a purity filter for signals emitted from the oligonucleotide clusters.
4. The system of claim 2, further comprising instructions that, when executed by the at least one processor, cause the system to estimate the level of read coverage for the genomic samples by: based on the filter metric and the respective quantities of oligonucleotide clusters belonging to respective ones of the genomic samples, determining a number of oligonucleotide clusters that satisfy the filtering threshold that pass a filter for each of the genomic samples; and based on the number of oligonucleotide clusters that pass a filter, estimating a minimum number of nucleotide reads that cover a genomic region for each of the genomic samples.
5. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to: based on the estimated level of read coverage, determine a customized set of flowcell regions to image from a flowcell sufficient to generate the nucleotide reads that satisfy the target level of read coverage for each of the genomic samples; and execute the sequencing run by capturing images of the customized set of flowcell regions for the customized number of sequencing cycles using the imaging system. 6. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to perform the subset of sequencing cycles according to an order of index cycles prior to a genomic sequencing cycle by: determining base calls of a first index sequence appended to a sample genomic sequence of a genomic sample; determining base calls of a second index sequence appended to the sample genomic sequence of the genomic sample; and after determining the base calls of the first index sequence and the second index sequence, determining base calls of a first nucleotide read corresponding to a first portion of the sample genomic sequence and determining base calls of a second nucleotide read corresponding to a second portion of the sample genomic sequence.
7. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to determine the respective number of oligonucleotide clusters belonging to the respective genomic sample by: identifying from the index sequences, assigned index sequences that match an index sequence registered for the sequencing run and unassigned index sequences that do not match the index sequence registered for the sequencing run; removing from data of the sequencing run, a subset of oligonucleotide clusters corresponding to the unassigned index sequences; determining a respective subset of assigned index sequences corresponding to the respective genomic sample; and determining from the respective subset of assigned index sequences, a number of oligonucleotide clusters belonging to each genomic sample.
8. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to generate the customized number of sequencing cycles for the sequencing run by increasing or decreasing a preset number of sequencing cycles for the sequencing run.
9. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to generate the customized number of sequencing cycles for the sequencing run by: identifying a minimum number of sequencing cycles and a maximum number of sequencing cycles for the sequencing run; and increasing or decreasing a preset number of sequencing cycles for the sequencing run to the customized number of sequencing cycles within the minimum number of sequencing cycles and the maximum number of sequencing cycles.
10. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to estimate the read coverage level by: determining a number of unique nucleotide reads aligned to a reference genome from the sequencing run; determining a number of pass-filtered nucleotide reads from pass-filtered oligonucleotide clusters having a signal that satisfies a filtering threshold from the sequencing run; determining a bioinformatics efficiency metric by dividing the number of unique nucleotide reads by the number of pass-filtered nucleotide reads; and estimating the read coverage level of the genomic samples based on the bioinformatics efficiency metric and the respective number of oligonucleotide clusters belonging to respective genomic samples.
11. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to detect a reagent volume of a cartridge in fluid communication with the fluidic system and operate the fluidic system to perform one or more additional sequencing cycles relative to the currently selected number of sequencing cycles to complete the custom number of sequencing cycles by aspirating one or more reagents from the cartridge.
12. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to terminate operation of the fluidic system from performing one or more sequencing cycles of the currently selected number of sequencing cycles to complete the sequencing run after performing the custom number of sequencing cycles.
13. A system comprising: an imaging system; a fluidic system; and a computing engine comprising: at least one processor; and a non-transitory computer-readable medium comprising instructions that, when executed by the at least one processor, cause the system to: determine, from a subset of sequencing cycles of a sequencing run of genomic samples, base calls of index sequences within oligonucleotide clusters; determine, based on the index sequences, respective numbers of oligonucleotide clusters belonging to respective ones of the genomic samples; estimate read coverage levels of the genomic samples based on the respective numbers of oligonucleotide clusters belonging to respective genomic samples; determine, from a flowcell and based on the estimated read coverage levels, a custom set of flowcell regions to image sufficient to generate nucleotide reads that meet a target read coverage level for each of the genomic samples; and perform the sequencing run by capturing images of the custom set of flowcell regions during sequencing cycles of the sequencing run.
14. The system of claim 13, further comprising instructions that, when executed by the at least one processor, cause the system to determine the custom set of flowcell regions by determining a custom number of flowcell regions to image sufficient to generate the nucleotide reads that meet the target read coverage level for each genomic sample.
15. The system of claim 13, further comprising instructions that, when executed by the at least one processor, cause the system to determine the custom set of flowcell regions by determining, from a flowcell, a set of fields to image sufficient to generate the nucleotide reads that meet the target read coverage level for each genomic sample.
16. The system of claim 13, further comprising instructions that, when executed by the at least one processor, cause the system to capture the images of the custom set of flowcell regions without adjusting a currently selected number of sequencing cycles of the sequencing run. 17. The system of claim 13, further comprising instructions that, when executed by the at least one processor, cause the system to determine the customized set of flow cell regions by increasing or decreasing a number of flow cell regions from an initial set of flow cell regions selected for the sequencing run.
18. The system of claim 13, further comprising instructions that, when executed by the at least one processor, cause the system to estimate the read coverage levels of the genomic samples by: determining a filter metric for a subset of oligonucleotide clusters that indicate a filter threshold of signal that satisfies the oligonucleotide cluster; and estimating the read coverage level of the genomic samples based on the filter metric and the respective number of oligonucleotide clusters belonging to the respective genomic sample.
19. The system of claim 18, further comprising instructions that, when executed by the at least one processor, cause the system to determine the filter metric by determining, in a filter plot, a percentage of clusters belonging to each genomic sample that satisfy a purity filter of signal emitted from the oligonucleotide cluster.
20. The system of claim 18, further comprising instructions that, when executed by the at least one processor, cause the system to estimate the read coverage levels of the genomic samples by: determining, based on the filter metric and the respective number of oligonucleotide clusters belonging to the respective genomic sample, a number of oligonucleotide clusters per genomic sample of the genomic samples that satisfy the filter threshold; and estimating, based on the number of oligonucleotide clusters that pass the filter, a minimum number of nucleotide reads that cover a genomic region of each genomic sample.
21. The system of claim 20, further comprising instructions that, when executed by the at least one processor, cause the system to determine the number of oligonucleotide clusters that pass the filter by: determining, in a filter plot, a percentage of clusters belonging to each genomic sample that satisfy a purity filter of signal emitted from the oligonucleotide cluster; and determining, based on the percentage of clusters that satisfy the purity filter, a number of oligonucleotide clusters per genomic sample of the genomic samples that satisfy the filter threshold.
22. The system of claim 20, further comprising instructions that, when executed by the at least one processor, cause the system to estimate the minimum number of nucleotide reads that cover a genomic region of each genomic sample by: determining, based on the number of oligonucleotide clusters that pass the filter, a number of nucleotide reads that cover a genomic region of each genomic sample; and determining, based on the number of nucleotide reads that cover the genomic region of each genomic sample, a minimum number of nucleotide reads that cover the genomic region of each genomic sample.
Citation Information
Patent Citations
Method of nucleic acid amplification
US20050100900A1
Labelled nucleotides
US20060188901A1
Modified polymerases for improved incorporation of nucleotide analogues
US20060240439A1
Polymerases
US20060281109A1
Modified nucleotides
US20070166705A1